Summary by AIEngineer.careers
Principal Solutions Architect is a AI Solutions Architect role (full-time). with ePlus Technology, inc. in IRVINE, US. Compensation shown: $170K–$190K. Imported listing (source: jobs-radar.com). Apply on the employer's site (jobs-radar.com).
Imported listing (source: jobs-radar.com) · Apply on jobs-radar.com
How our listings work
The sections below reproduce the third-party job description for reference. AIEngineer.careers does not write or control this text.
Imported job description
Sourced from jobs-radar.com
ePlus is seeking a Principal Solutions Architect to design and deploy NVIDIA AI Factory infrastructure for enterprise and hyperscale customers. This role involves architecting full-stack solutions including GPU compute, high-performance networking, parallel storage, and the NVIDIA AI software stack. You will serve as a trusted technical advisor, leading proof-of-concept deployments and delivering transformative AI infrastructure programs.
Responsibilities
- Lead discovery workshops to capture AI/ML workload requirements and architect full-stack AI Factory solutions aligned to NVIDIA reference architectures.
- Develop detailed Bills of Materials, rack elevation diagrams, network topology drawings, and power/cooling budgets.
- Design GPU cluster architectures using NVIDIA DGX, HGX, and MGX systems with Blackwell configurations.
- Specify colocation requirements including critical power load, cooling, and carrier-neutral telecom diversity.
- Design high-performance networking with InfiniBand and Spectrum-X, including RDMA and NCCL optimization.
- Architect parallel storage solutions using VAST Data, Hammerspace, and Pure Storage for AI training workloads.
- Deploy NVIDIA AI Enterprise stack, NIM microservices, and orchestration with Kubernetes and SLURM.
- Serve as primary technical point of contact throughout pre-sales and delivery lifecycle.
Requirements
- Bachelor's degree in Computer Science, Electrical Engineering, or related technical discipline.
- 8+ years of solutions architecture or systems engineering experience, with at least 4 years focused on GPU infrastructure or HPC.
- Proven track record designing and deploying NVIDIA DGX or HGX-based GPU clusters in production AI/ML environments.
- Deep understanding of distributed deep learning concepts (tensor parallelism, pipeline parallelism, etc.).
- Hands-on experience with InfiniBand or high-speed Ethernet fabric design and RDMA configuration.
- Experience sizing and deploying parallel storage systems for AI workloads.
- Strong working knowledge of Kubernetes, GPU Operator, and at least one GPU workload scheduler (Run:ai or SLURM).
- Excellent communication skills with ability to present to technical and executive audiences.
Nice to Have
- NVIDIA-certified professional credentials (DCA-Core, NCP-DS, or equivalent).
- Experience with NVIDIA Base Command Platform or Mission Control.
- Familiarity with sovereign AI, government cloud, or regulated industry AI infrastructure.
- Experience integrating AI Factory infrastructure with public cloud for hybrid architectures.
- Background in MLOps, LLMOps, or platform engineering for production AI model lifecycle management.
- Prior experience with colocation data center procurement and SLA negotiation.
- Contributions to open-source AI infrastructure projects or published technical content.
Tech Stack
PyTorchTensorFlowJAXKubernetesCUDADeepSpeedvLLMNCCLMegatron-LMInfiniBandNVIDIA GPU OperatorNVIDIA AI EnterpriseVAST DataPure Storage
ETHerndon · US · 2148+ employees
ePlus Technology, inc. is a wholly-owned subsidiary of ePlus inc., an American consultative technology solutions provider that helps organizations modernize, secure, and scale their IT infrastructure. The company delivers integrated IT solutions, including hardware, software, and specialized professional and managed services to a wide range of commercial and government clients.