← All open positionsCAREERS / INFRASTRUCTURE

DevOps Engineer

Ramat Gan · Full-time · Hybrid

Apply now →

doubleAI is building next-generation reasoning and code-optimization agents that push the frontier of artificial expert intelligence. Our models are built on world-class infrastructure: large-scale GPU clusters, fast networking, reliable scheduling, deep observability, secure cloud environments, and CI/CD pipelines that allow our team to move quickly without compromising reliability.

We are looking for a DevOps Engineer to own and scale the infrastructure behind our AI systems. This role combines hands-on management of our large-scale GPU cluster with classical DevOps responsibilities across cloud infrastructure, CI/CD, Kubernetes, monitoring, security, and developer productivity.

We are looking for someone deeply technical, hands-on, and AI-native in mindset — someone who understands that compute, automation, reliability, and security are foundational to building an AI company.

In this role, you will:

  • Own cluster reliability, availability, monitoring, alerting, capacity planning, and incident response.
  • Configure, maintain, and troubleshoot GPU servers, networking equipment, switches, InfiniBand fabrics, storage, and related infrastructure.
  • Deploy, maintain, and improve infrastructure around Slurm, Kubernetes, and GPU workload orchestration.
  • Manage and improve our cloud environment, including compute, networking, storage, IAM, security groups, cost visibility, and production infrastructure.
  • Build, maintain, and improve CI/CD pipelines.
  • Work closely with engineers to understand infrastructure bottlenecks and improve development, training, inference, and deployment velocity.
  • Contribute to security best practices across cloud, network, CI/CD, access control, secrets management, and production systems.

What we're looking for

  • Strong hands-on experience managing Linux-based infrastructure in production environments.
  • Experience with classical DevOps responsibilities: CI/CD, cloud infrastructure, monitoring, automation, deployment systems, and production operations.
  • Familiarity with Kubernetes in real-world production environments.
  • Experience operating GPU clusters, HPC environments, AI infrastructure, or other high-performance compute systems.
  • Familiarity with Slurm or similar workload schedulers.
  • Practical knowledge of networking, including switch configuration, routing, VLANs, firewalls, high-throughput networking, and debugging network issues.
  • Experience with InfiniBand, RDMA, high-performance interconnects, or similar low-latency networking technologies.
  • Ability to troubleshoot across the full stack: hardware, firmware, BIOS, Linux kernel, drivers, CUDA, containers, networking, storage, cloud infrastructure, CI/CD, and schedulers.
  • Strong scripting and automation skills using Bash, Python, Ansible, Terraform, or similar tools.
  • A strong ownership mindset: you notice problems, fix root causes, automate repetitive work, and leave systems more reliable than you found them.

Bonus points for

  • Experience managing NVIDIA GPU clusters at scale.
  • Familiarity with CUDA, NVIDIA drivers, NCCL, DCGM, MIG, GPUDirect RDMA, or GPU performance monitoring.
  • Experience with Mellanox/NVIDIA networking equipment, InfiniBand fabric management, subnet managers, or RDMA debugging.
  • Experience supporting large-scale ML training workloads using PyTorch, Megatron, TorchTitan, DeepSpeed, or similar frameworks.
  • Experience in a startup or fast-moving research environment where priorities shift quickly and infrastructure must keep up.

Logistics

  • Location: Ramat Gan, Bursa
  • Workplace type: Hybrid
  • Employment type: Full-time
Apply now →