GRAIL seeks a Staff Site Reliability Engineer to lead reliability, scalability, and security of their cloud-native platform supporting large-scale data processing and cancer detection technologies. This role involves designing fault-tolerant cloud infrastructure, managing Kubernetes, improving observability, and mentoring engineers.
Responsibilities
Design and operate highly available cloud infrastructure across AWS, GCP, Azure
Architect and maintain CI/CD pipelines
Lead infrastructure-as-code adoption with Terraform, CloudFormation, Ansible
Own Kubernetes reliability across multi-cluster environments
Establish observability platforms and SLO/SLI frameworks
Lead incident response and root cause analysis
Optimize infrastructure for cost and performance
Define and enforce DevOps and security best practices
Mentor engineers and contribute to technical leadership
Requirements
BS in CS or equivalent, 8+ years SRE/DevOps experience
Strong hands-on with AWS, GCP, or Azure
Experience with Terraform, CloudFormation, or similar
GRAIL, Inc. is an innovative commercial-stage healthcare company focused on the early detection of cancer. The company utilizes next-generation sequencing, clinical studies, and machine learning to develop blood-based multi-cancer early detection (MCED) tests.