Senior Software Engineer LLM Post-Training Platform is a Machine Learning Engineer role (full-time). with Streamlit. in BELLEVUE, US. Compensation shown: $160K–$230K. Imported listing (source: careerbuilder.com). Apply on the employer's site (careerbuilder.com).
Imported listing (source: careerbuilder.com) · Apply on careerbuilder.com
The sections below reproduce the third-party job description for reference. AIEngineer.careers does not write or control this text.
Imported job description
Sourced from careerbuilder.com
Snowflake's ML Platform team is building Cortex Training, an LLM post-training platform. They are seeking a Senior Software Engineer to design and build scalable distributed systems for GPU compute, productionize research building blocks, and drive performance at scale.
Responsibilities
Design and build across the full stack — from the public training APIs and SDK through the control plane to the GPU data plane.
Scale the distributed systems that make GPU compute serverless — multi-tenant scheduling, placement, and capacity-aware routing across regional GPU pools, with fault tolerance built in.
Drive end-to-end performance at scale — keep the training, inference, and RL loops fast and the data plane responsive under heavy concurrent load, with GPUs kept saturated.
Productionize research building blocks — partner with Snowflake Research to turn state-of-the-art training and inference techniques into reliable, composable components customers can run at enterprise scale.
Requirements
5+ years building and shipping production ML systems
Strong distributed systems and infrastructure foundation — designing scalable, fault-tolerant services and operating them on Kubernetes in production.
Streamlit is an open-source Python framework that allows data scientists and AI/ML engineers to transform data scripts into interactive web applications quickly. Acquired by Snowflake in 2022, it specializes in tools for machine learning model deployment and data visualization.
Familiarity with GPU and LLM infrastructure — e.g., PyTorch, DeepSpeed/FSDP, Ray, CUDA/NCCL, vLLM; able to debug across the data, infrastructure, and GPU layers.
Demonstrated ability to harden complex systems for reliability, throughput, and cost efficiency.
BS in Computer Science or a related field (MS/PhD a plus).
Nice to Have
Hands-on LLM post-training / modeling experience — the strongest candidates pair deep infra skills with real post-training intuition.
Tech Stack
PyTorchKubernetesCUDADeepSpeedRayvLLMFSDPNCCL
reputed company Machine Learning Engineer - Remote (US) or CA - Only W2