doubleAI is building next-generation reasoning and code-optimization agents that push the frontier of artificial expert intelligence. Our models are built on world-class infrastructure: large-scale GPU clusters, fast networking, reliable scheduling, deep observability, secure cloud environments, and CI/CD pipelines that allow our team to move quickly without compromising reliability.
We are looking for a DevOps Engineer to own and scale the infrastructure behind our AI systems. This role combines hands-on management of our large-scale GPU cluster with classical DevOps responsibilities across cloud infrastructure, CI/CD, Kubernetes, monitoring, security, and developer productivity.
We are looking for someone deeply technical, hands-on, and AI-native in mindset — someone who understands that compute, automation, reliability, and security are foundational to building an AI company.
In this role, you will:
- Own cluster reliability, availability, monitoring, alerting, capacity planning, and incident response.
- Configure, maintain, and troubleshoot GPU servers, networking equipment, switches, InfiniBand fabrics, storage, and related infrastructure.
- Deploy, maintain, and improve infrastructure around Slurm, Kubernetes, and GPU workload orchestration.
- Manage and improve our cloud environment, including compute, networking, storage, IAM, security groups, cost visibility, and production infrastructure.
- Build, maintain, and improve CI/CD pipelines.
- Work closely with engineers to understand infrastructure bottlenecks and improve development, training, inference, and deployment velocity.
- Contribute to security best practices across cloud, network, CI/CD, access control, secrets management, and production systems.
What we're looking for
- Strong hands-on experience managing Linux-based infrastructure in production environments.
- Experience with classical DevOps responsibilities: CI/CD, cloud infrastructure, monitoring, automation, deployment systems, and production operations.
- Familiarity with Kubernetes in real-world production environments.
- Experience operating GPU clusters, HPC environments, AI infrastructure, or other high-performance compute systems.
- Familiarity with Slurm or similar workload schedulers.
- Practical knowledge of networking, including switch configuration, routing, VLANs, firewalls, high-throughput networking, and debugging network issues.
- Experience with InfiniBand, RDMA, high-performance interconnects, or similar low-latency networking technologies.
- Ability to troubleshoot across the full stack: hardware, firmware, BIOS, Linux kernel, drivers, CUDA, containers, networking, storage, cloud infrastructure, CI/CD, and schedulers.
- Strong scripting and automation skills using Bash, Python, Ansible, Terraform, or similar tools.
- A strong ownership mindset: you notice problems, fix root causes, automate repetitive work, and leave systems more reliable than you found them.
Bonus points for
- Experience managing NVIDIA GPU clusters at scale.
- Familiarity with CUDA, NVIDIA drivers, NCCL, DCGM, MIG, GPUDirect RDMA, or GPU performance monitoring.
- Experience with Mellanox/NVIDIA networking equipment, InfiniBand fabric management, subnet managers, or RDMA debugging.
- Experience supporting large-scale ML training workloads using PyTorch, Megatron, TorchTitan, DeepSpeed, or similar frameworks.
- Experience in a startup or fast-moving research environment where priorities shift quickly and infrastructure must keep up.
Logistics
- Location: Ramat Gan, Bursa
- Workplace type: Hybrid
- Employment type: Full-time