Scale AI is seeking a Senior Software Engineer for their Agent Oversight team to build platform infrastructure for observing, evaluating, and improving production agentic AI systems. The role involves designing APIs, data pipelines, and monitoring tools to ensure reliable AI deployments. Ideal candidates have backend/distributed systems experience, familiarity with ML/LLM production systems, and a track record of owning features end-to-end.
Responsibilities
Design and build core platform capabilities for deploying, monitoring, and evaluating agentic applications in production
Build reliable APIs and data pipelines that capture agent telemetry, evaluation signals, and performance metrics at scale
Work alongside ML engineers where platform work intersects with evaluation or improvement systems
Own the reliability, scalability, and observability of platform components serving multiple concurrent enterprise and government customers
Work cross-functionally with product, forward deployed engineering, and customers to translate real-world deployment requirements into platform features
Build features end-to-end: system design, implementation, debugging, and testing
Participate in high-velocity experimentation to validate platform capabilities against real customer usage
Requirements
4+ years of professional software engineering experience, with strong fundamentals in backend/distributed systems, APIs, and data pipeline design
Scale AI is an American artificial intelligence infrastructure and software company. It provides high-quality training data, annotations, and RLHF services to power AI models, and offers full-stack technologies for building and deploying AI applications.
Hands-on experience building production software for ML/LLM-powered products or platforms
Working knowledge of how LLM or ML systems behave in production: evaluation signals, failure modes, prompt/tool-calling workflows, experiment results, data quality issues, and tradeoffs
Experience partnering closely with ML engineers or applied researchers to turn prototypes into reliable platform capabilities
Experience building infrastructure or platforms that other engineering teams build on top of
Track record of taking ownership of features or components end-to-end
Comfortable operating in an ambiguous, fast-changing domain
Strong problem-solving skills and ability to work independently or as part of a cross-functional team
Excited to work directly with ML engineers and customer-facing teams
Gives direct, substantive feedback on designs and code, and mentors others
Nice to Have
Deep experience building or maintaining observability, monitoring, or evaluation systems for ML/LLM-powered products in production
Familiarity with agent architectures — tool use, planning, multi-agent orchestration
Exposure to MLOps, feature stores, model serving, or experiment infrastructure
Experience working in regulated or enterprise contexts
Experience reviewing others’ technical designs or mentoring engineers at a senior/staff level