G-Research is hiring an NLP Performance Engineer to optimize large-scale LLM inference workloads on advanced GPU infrastructure. This specialist Quantitative Developer will profile and benchmark model serving, implement inference optimizations, and build tooling to improve the performance and cost-efficiency of the firm's NLP platform.
Responsibilities
Profile, benchmark and optimize large-scale LLM inference workloads across compute infrastructure
Ensure efficient deployment of the latest models across GPU architectures and adapt the inference stack as hardware evolves
Design and implement inference optimizations while maintaining output quality
Develop reference implementations, libraries and tooling for efficient and reliable NLP workloads
Collaborate with researchers, senior stakeholders and engineers to design optimized solutions
Work with systems, architecture and platform teams to evolve the compute stack and influence long-term platform decisions
Requirements
Bachelor's, Master's or PhD in computer science or equivalent experience
Proven experience profiling, benchmarking and optimizing large-scale LLM inference workloads
Scientific, evidence-led approach to performance optimization using rigorous benchmarking and reproducible measurement
G-Research is a leading quantitative finance research and technology firm that applies scientific rigour, advanced machine learning, and cutting-edge technology to predict movements in global financial markets. It brings together researchers and engineers to build sophisticated trading strategies and robust infrastructure.