Description:
As a Staff Software Engineer on the Inference team, you will operate as a technical leader across multiple teams and services, driving architecture, performance, and reliability for CoreWeave’s Kubernetes-native inference platform. You will define and lead complex, cross-cutting design initiatives spanning request routing, adaptive scheduling, GPU resource management, and cost-per-token optimisation under strict P99 SLAs. In this high-impact role, you will implement advanced inference optimisations—such as speculative decoding and KV-cache reuse—while establishing performance benchmarking frameworks, guiding cross-functional alignment across infrastructure boundaries, and raising the bar for engineering rigour and observability practices across the organisation.
Who You Are
- Bachelor’s degree in Computer Science, Engineering, or a related technical field (or equivalent practical experience).
- 8+ years of experience building large-scale distributed systems or cloud platforms, with a proven track record of leading cross-team or organisation-level technical initiatives at scale.
- Strong coding proficiency in Go, Python, or C++.
- Deep expertise in Kubernetes at production scale, including orchestration, scheduling, and service design.
- Strong understanding of networked systems, performance optimisation, and distributed system design.
- Hands-on engineering experience with inference systems, including batching/micro-batching strategies, caching, memory optimisation, mixed precision (BF16/FP8), and streaming token delivery.
- Demonstrated ability to systematically improve tail latency (P95/P99) and platform reliability through metrics-driven engineering.
- Proven experience owning system-wide SLIs/SLOs, capacity planning, autoscaling strategies, and mentoring senior and mid-level engineers.
Preferred
- Direct open-source or production contributions to modern inference frameworks (e.g., vLLM, Triton, TensorRT-LLM, Ray Serve, or TorchServe).
- Deep experience with GPU systems engineering and hardware performance optimisation (e.g., CUDA, NCCL, RDMA, NUMA, or GPU interconnects).
- Direct exposure to large-scale AI/ML infrastructure or hyperscale cloud environments.
Wondering if you're a good fit?
We believe in investing in our people, and value candidates who can bring their own diversified experiences to our teams—even if you aren't a 100% skill or experience match.
- You love to: Scale highly complex distributed architectures and mentor engineering cohorts to elevate technical standards across an organisation.
- You're curious about: Pioneering low-latency inference optimisations and finding innovative shortcuts to optimise cost-per-token performance.
- You're an expert in: Metrics-driven engineering, troubleshooting micro-bottlenecks, and delivering stable cloud platforms under strict multi-tenant constraints.