Senior ML Ops Engineer is a MLOps Engineer role (full-time). with Kayak. in BERLIN, DE. Imported listing (source: clever.cv). Apply on the employer's site (clever.cv).
Imported listing (source: clever.cv) · Apply on clever.cv
The sections below reproduce the third-party job description for reference. AIEngineer.careers does not write or control this text.
Imported job description
Sourced from clever.cv
KAYAK is seeking a Senior MLOps Engineer to design and implement machine learning infrastructure and production lifecycle. You will join the Machine Learning Platform team, building scalable infrastructure and automated pipelines for model training, deployment, and monitoring. This is a senior, hands-on role bridging data science and production engineering, requiring commuting to the Berlin office 3 times a week.
Responsibilities
Build and maintain ML infrastructure end-to-end, including CI/CD pipelines, model orchestration, and automated training pipelines.
Own model deployment and serving, defining standards and tooling for low latency and high availability.
Develop core MLOps capabilities such as feature stores, model registries, and automated monitoring for performance and data drift.
Operationalize infrastructure for the ML team, enabling Kubernetes autoscaling and GPU provisioning as self-service tools.
Improve platform reliability and performance by designing resilient monitoring, defining SLOs, and implementing automation.
Empower Data Scientists through standardized, optimized workflows and golden paths.
Requirements
Experience building and operating ML platforms in production environments.
KAYAK is a leading global travel search engine that helps users compare and book flights, hotels, car rentals, and vacation packages. It processes billions of searches annually by aggregating data from hundreds of travel providers.
Solid working knowledge of containerization and orchestration (Docker, Kubernetes), Linux internals, and model serving at scale.
Familiarity with ML lifecycle tooling, including orchestration frameworks, feature stores, model registries, and drift or performance monitoring.
Experience owning production systems: defining SLOs, building observability (e.g., Prometheus, Grafana, Datadog), participating in incident response, and diagnosing large-scale failures.
Comfort writing production-quality code in Python or a comparable language.
Experience modernizing production infrastructure with attention to reliability, risk, and cost.
Ability to take ownership of technical outcomes, advocate for decisions using data, and communicate clearly.