doubleAI is building next-generation reasoning and code-optimization agents that push the frontier of artificial expert intelligence. We are looking for a Deep Learning Engineer to help design, implement, optimize, and scale the training and inference systems behind our models.
This is not a pure research role. We are looking for an exceptional engineer who can work close to the model stack, understand deep learning systems end-to-end, and turn ambitious ideas into reliable, high-performance infrastructure. You will work across large-scale training, distributed systems, inference optimization, model evaluation, and deployment — helping us build AI systems that are scalable, efficient, and robust.
In this role, you will:
- Collaborate with cross-functional teams to bring research breakthroughs into production environments.
- Build tools, abstractions, and infrastructure that make model training, evaluation, and deployment faster and more reliable.
- Analyze and overcome performance and scalability challenges in large-scale distributed training.
- Debug complex issues across the full ML stack: data loading, model code, distributed execution, GPU utilization, memory usage, networking, checkpointing, and serving.
- Work with modern training and inference frameworks such as Megatron, TorchTitan, PyTorch, and vLLM.
What we're looking for
- MSc or PhD in AI, Machine Learning or Algorithms.
- Strong software engineering skills, with experience designing and building complex systems.
- Practical understanding of distributed training concepts: data parallelism, tensor parallelism, pipeline parallelism, sharding, checkpointing, synchronization, and fault tolerance.
- Experience working with inference systems and serving frameworks, ideally including vLLM or similar high-throughput LLM serving infrastructure.
- Solid understanding of machine learning fundamentals, including optimization, transformers, attention mechanisms, numerical stability, and model evaluation.
- Low-level systems intuition: performance profiling, memory layout, GPU utilization, networking bottlenecks, and tradeoffs between latency, throughput, and reliability.
- Ability to write clean, maintainable, well-tested code in fast-moving environments.
- Strong ownership mindset, good judgment, and the ability to operate with high agency under uncertainty.
Bonus points for
- Experience training or serving large language models at scale.
- Experience with Mixture-of-Experts, model parallelism, or large distributed GPU clusters.
- Familiarity with CUDA kernels, Triton, NCCL, RDMA, Kubernetes, Slurm, or cloud/HPC infrastructure.
- Experience optimizing inference systems for throughput, latency, batching, KV-cache efficiency, or cost.
- Experience building internal ML platforms, experiment infrastructure, observability tools, or model evaluation pipelines.
- Background in reinforcement learning, post-training, synthetic data, or agentic systems.
- Contributions to open-source ML systems or deep learning infrastructure.
- Strong publication record or demonstrable research contributions.
Logistics
- Location: Ramat Gan, Bursa
- Workplace type: Hybrid
- Employment type: Full-time