Model AI is building the Agent Cloud, a serving and training infrastructure platform for agentic AI workloads. This founding role focuses on optimizing high-performance LLM inference across accelerators, including batching, scheduling, KV cache management, and long-context inference. You'll work with frameworks such as PyTorch, vLLM, SGLang, and TensorRT-LLM to improve throughput, latency, and cost efficiency.
Responsibilities
Optimize large-scale LLM inference and serving systems.
Improve total tokens per second, decode tokens per second, latency, throughput, and cost efficiency.
Work on serving infrastructure for open-source models across different types of accelerators.
Model Ai provides professional service software that leverages artificial intelligence to assist industries like power plants and national parks in ensuring safety. The company focuses on professional service software solutions.
Experience optimizing inference or training workloads for large models.
Familiarity with TPUs, GPUs, or other accelerators.
Experience with one or more of CUDA, Triton, NCCL, JAX/XLA, PyTorch internals, vLLM, SGLang, TensorRT-LLM, distributed inference, or distributed training.
Strong systems debugging skills.
Comfort working across model code, runtime, infrastructure, and product requirements.
High ownership and the ability to operate effectively in an early-stage startup environment.