ElastixAI is an early-stage startup building next-generation AI inference infrastructure. This role focuses on designing and extending the low-level serving stack, including customizing open-source frameworks like vLLM and building model sharding and scheduling logic for proprietary AI accelerators. You will work at the intersection of ML systems, compiler/runtime engineering, and hardware-software co-design.
Responsibilities
Architect, extend, and optimize core components of the AI serving platform for throughput, latency, and scalability.
Customize open-source serving frameworks (e.g., vLLM, SGLang, TensorRT-LLM) for proprietary model ingestion and accelerator integration.
Develop efficient model partitioning, scheduling, and memory management strategies for multi-device inference.
Collaborate with ML engineers on model export and runtime optimization (quantization, graph transforms).
Work closely with hardware engineers to influence accelerator interface design and performance tuning.
Build APIs and runtime tools enabling flexible, PyTorch-native model deployment on our infrastructure.
Profile, debug, and optimize across the full stack — from Python orchestration to C++ kernels and PCIe drivers.
Requirements
BS/MS/PhD in Computer Science, Electrical/Computer Engineering, or related field.
3+ years of professional experience in systems programming, ML infrastructure, or distributed inference.
ElastixAI is a software development company that co-designs reconfigurable hardware, system software, and model-level optimization as a unified stack for AI inference. Their platform is designed to optimize large language models by matching compute demand to infrastructure at every layer.