Staff Software Engineer- AI Agent Evaluations is a LLM Evaluation Engineer role (full-time). with ID.me. in MOUNTAIN VIEW, US. Compensation shown: $218K–$271K. Imported listing (source: jobs-radar.com). Apply on the employer's site (jobs-radar.com).
Imported listing (source: jobs-radar.com) · Apply on jobs-radar.com
The sections below reproduce the third-party job description for reference. AIEngineer.careers does not write or control this text.
Imported job description
Sourced from jobs-radar.com
ID.me seeks a Staff Software Engineer to define and lead the discipline of testing AI agents, evaluating LLM behavior, and ensuring reliability of agentic systems. You will build eval infrastructure, production observability, and developer tooling for AI features, while mentoring engineers and establishing quality standards across the org.
Responsibilities
Define AI Quality Standards and own the evaluation framework for AI agents
Build and maintain evaluation pipelines for LLM outputs and agent behavior
Instrument agentic systems for production observability and behavioral drift detection
Lead the design of test suites for non-deterministic AI outputs
Champion developer experience by building internal tooling and feedback loops
Drive AI-first engineering culture and mentor engineers on AI testing best practices
Collaborate with Security, Platform, Product, and AI/ML teams to embed quality gates
Requirements
Bachelor's degree in Computer Science, Engineering, or equivalent experience
8+ years building and operating production software systems
Demonstrated experience evaluating or testing LLM-powered features or autonomous agents in production
ID.me is a digital identity network that provides secure identity verification services for government agencies, businesses, and healthcare organizations. It allows users to verify their identity once and use that credential to sign in securely across multiple platforms.