← All open positionsCAREERS / ENGINEERING

Deep Learning Engineer

Ramat Gan · Full-time · Hybrid

Apply now →

doubleAI is building next-generation reasoning and code-optimization agents that push the frontier of artificial expert intelligence. We are looking for a Deep Learning Engineer to help design, implement, optimize, and scale the training and inference systems behind our models.

This is not a pure research role. We are looking for an exceptional engineer who can work close to the model stack, understand deep learning systems end-to-end, and turn ambitious ideas into reliable, high-performance infrastructure. You will work across large-scale training, distributed systems, inference optimization, model evaluation, and deployment — helping us build AI systems that are scalable, efficient, and robust.

In this role, you will:

  • Collaborate with cross-functional teams to bring research breakthroughs into production environments.
  • Build tools, abstractions, and infrastructure that make model training, evaluation, and deployment faster and more reliable.
  • Analyze and overcome performance and scalability challenges in large-scale distributed training.
  • Debug complex issues across the full ML stack: data loading, model code, distributed execution, GPU utilization, memory usage, networking, checkpointing, and serving.
  • Work with modern training and inference frameworks such as Megatron, TorchTitan, PyTorch, and vLLM.

What we're looking for

  • MSc or PhD in AI, Machine Learning or Algorithms.
  • Strong software engineering skills, with experience designing and building complex systems.
  • Practical understanding of distributed training concepts: data parallelism, tensor parallelism, pipeline parallelism, sharding, checkpointing, synchronization, and fault tolerance.
  • Experience working with inference systems and serving frameworks, ideally including vLLM or similar high-throughput LLM serving infrastructure.
  • Solid understanding of machine learning fundamentals, including optimization, transformers, attention mechanisms, numerical stability, and model evaluation.
  • Low-level systems intuition: performance profiling, memory layout, GPU utilization, networking bottlenecks, and tradeoffs between latency, throughput, and reliability.
  • Ability to write clean, maintainable, well-tested code in fast-moving environments.
  • Strong ownership mindset, good judgment, and the ability to operate with high agency under uncertainty.

Bonus points for

  • Experience training or serving large language models at scale.
  • Experience with Mixture-of-Experts, model parallelism, or large distributed GPU clusters.
  • Familiarity with CUDA kernels, Triton, NCCL, RDMA, Kubernetes, Slurm, or cloud/HPC infrastructure.
  • Experience optimizing inference systems for throughput, latency, batching, KV-cache efficiency, or cost.
  • Experience building internal ML platforms, experiment infrastructure, observability tools, or model evaluation pipelines.
  • Background in reinforcement learning, post-training, synthetic data, or agentic systems.
  • Contributions to open-source ML systems or deep learning infrastructure.
  • Strong publication record or demonstrable research contributions.

Logistics

  • Location: Ramat Gan, Bursa
  • Workplace type: Hybrid
  • Employment type: Full-time
Apply now →