Description:
As an MLOps Engineer on the Physical AI team, you will serve as the hands-on owner for our machine learning operations surface across the end-to-end model lifecycle—from experimentation and training through to packaging, deployment, serving, and retirement. You will define and roll out MLOps practices, establish operational SLOs/SLAs, and build automated CI/CD and continuous training pipelines to accelerate the path from experiment to supported production deployment. In this role, you will implement comprehensive model observability, data versioning, and drift monitoring while ensuring robust security and governance controls. Additionally, you will partner closely with product, data science, and core infrastructure teams to optimize GPU compute utilization, resolve cross-boundary platform incidents, and mentor engineers on production-grade ML practices.
Who You Are
- 5–6+ years of professional experience in MLOps, ML platform engineering, ML infrastructure, or SRE/DevOps for production machine learning systems.
- Proven experience building, operating, and automating production ML pipelines covering experiment tracking, model registries, artifact versioning, dataset management, and deployment workflows.
- Deep hands-on experience implementing observability for ML systems, including monitoring inference availability, latency, throughput, GPU/resource utilization, and data or model drift.
- Strong background in reliability engineering, including defining operational SLOs, writing runbooks, building automated remediation, and managing incident response for ML workloads.
- Proficient in Python for platform tooling, infrastructure integration, and pipeline automation, alongside strong infrastructure-as-code and CI/CD practices.
- Comfortable operating containerised, cloud-native environments using Kubernetes and public cloud platforms.
- Excellent technical communication and cross-functional collaboration skills, with a track record of bridging data science and platform engineering teams.
Preferred
- Experience as an early or founding MLOps engineer establishing ML platform architecture, standards, and operating models from the ground up.
- Hands-on experience operating ML workloads on Kubernetes with GPU infrastructure, distributed training, or large-scale inference engines.
- Experience with ML platforms handling test, simulation, or time-series data (e.g., physical test benches, battery labs, automotive/aerospace R&D) within multi-tenant SaaS environments.