NVIDIA's Metropolis Vision AI team is seeking a Senior Software Engineer to develop and optimize high-performance, GPU-accelerated Vision AI pipelines that power intelligent spaces, smart cities, retail analytics, and digital twins. The role involves building large-scale distributed services for processing video, image, and 3D data, and bringing research into production. You will collaborate with experts in perception, simulation, and large multimodal models.
Responsibilities
Craft and implement high-performance Vision AI pipelines for real-time and streaming scenarios using new computer vision and deep learning models.
Develop and refine large-scale distributed services for processing video, image, and 3D data in edge and cloud settings.
Develop multi-modal perception capabilities combining 2D, 3D, and temporal information.
Use simulation and synthetic data tools to build, test, and validate perception algorithms at scale.
Profile and tune GPU-accelerated inference pipelines to meet latency, efficiency, and reliability targets.
Collaborate with product, research, and platform teams to translate requirements into technical builds.
Drive technical build reviews, promote code quality/testing standards, and mentor other engineers.
Requirements
BS, MS, or PhD in Computer Science, Electrical/Computer Engineering, or related field, or equivalent experience.
NVIDIA is a multinational technology company that designs GPUs, SoCs, and AI/data-science and high-performance computing platforms and APIs. The company reports 42,000+ employees across 38 countries and describes its corporate footprint globally.
12+ years of professional software development experience using modern C++ (14/17/20) and Python on Linux.
Strong computer science fundamentals: algorithms, data structures, concurrency, and distributed systems.
Demonstrated expertise in computer vision and deep learning with production systems deployed.
Experience building high-performance concurrent systems: multi-threading, async I/O, and efficient memory management.
Proficiency with Linux, containers, microservices, and integrating AI components into back-end services.
Ability to rapidly prototype vision models/pipelines and evolve them into production-quality services.
Practical experience with PyTorch for training, fine-tuning, and deploying vision models.
Strong analytical, problem-solving, and communication skills; success collaborating across time zones.
Nice to Have
Proven experience delivering end-to-end computer vision applications in production: video analytics, smart cities, autonomous systems, retail analytics, industrial inspection, or digital twins.
Practical experience with GPU acceleration such as CUDA, TensorRT, or comparable technologies, and low-level optimization for inference and pre/post-processing.
Experience in simulation and synthetic data creation with Omniverse, Unreal Engine, Unity, or similar digital-twin platforms.
Background in vision-language models or related multi-modal AI, including integrating these models into real products.
Background in multimedia: video-centric processing and delivery, codecs, video pipelines, or media frameworks.