ML Systems Engineer is a ML Infrastructure Engineer role (full-time). with Periodic Labs. in MENLO PARK, US. Compensation shown: $250K–$350K. Imported listing (source: aihiringboard.com). Apply on the employer's site (aihiringboard.com).
Imported listing (source: aihiringboard.com) · Apply on aihiringboard.com
The sections below reproduce the third-party job description for reference. AIEngineer.careers does not write or control this text.
Imported job description
Sourced from aihiringboard.com
Periodic Labs is an AI and physical sciences company building state-of-the-art models to accelerate breakthroughs in materials, energy, and beyond. As an ML Systems Engineer, you will build and optimize the agentic infrastructure powering large-scale training, inference, and reinforcement learning. You will own critical pieces of the ML systems stack to maximize performance, scalability, reliability, and productivity for both engineers and AI agents.
Responsibilities
Build and optimize large-scale training and reinforcement learning infrastructure while ensuring its correctness
Develop high-performance inference and serving systems
Design distributed runtimes and scheduling systems for complex ML workloads
Build secure and large-scale sandboxing and execution environments
Optimize memory, GPU kernels and communication for maximum throughput and end-to-end efficiency
Improve scalability, reliability, and efficiency across the ML systems stack
Requirements
Strong systems programming and performance engineering skills
Experience building high-performance ML infrastructure at scale
Ability to own complex technical problems end-to-end
Strong coding ability and engineering judgment, including the ability to work effectively with AI agents to design, implement, test, and debug complex systems
Periodic Labs is a technology company founded in 2025 by former researchers from OpenAI and DeepMind that focuses on automating scientific discovery. The company is building AI scientists and autonomous laboratories that conduct physical experiments to generate high-quality experimental data for material discovery and other scientific research.
High ownership, fast execution, and a passion for pushing the frontier of AI systems and accelerating scientific discovery
Deep expertise in at least one of the following: Training (Megatron-LM), Distributed Runtime (Ray), Inference (SGLang), Sandboxing (secure execution environments), GPU Kernels (CUDA, Triton, CUTLASS, CuTe), GPU Communication (NCCL, NVLink, InfiniBand, RDMA, GPUDirect RDMA)
Nice to Have
Familiarity with TorchTitan, FSDP, veRL, Slime, or other distributed training systems
Familiarity with Monarch or other distributed execution frameworks
Familiarity with vLLM, TensorRT-LLM, or production LLM serving systems