Search
Go
Senior Machine Learning Engineer, LLM Inference Optimization is a ML Engineer role (full-time). with Nebius. in BERLIN, DE. Imported listing (source: databerlin.net). Apply on the employer's site (databerlin.net).
Imported listing (source: databerlin.net) · Apply on databerlin.net
The sections below reproduce the third-party job description for reference. AIEngineer.careers does not write or control this text.
Sourced from databerlin.net
Nebius is seeking a Senior Machine Learning Engineer to optimize LLM and VLM inference performance for their cloud platform. The role focuses on reducing latency, improving throughput, and lowering cost per token through advanced model compression and serving architecture optimizations. You will work closely with kernel and platform engineers to deploy and benchmark inference engines like vLLM and TensorRT-LLM.

Amsterdam · NL · 1590+ employees
Nebius is an AI cloud company providing full-stack infrastructure for AI developers and enterprises. Its platform supports the complete AI journey, including data storage, model training, tuning, and production runtime deployment.
MP Data