NLP Performance Engineer is a AI Infrastructure Engineer role (full-time). with G-Research. in GREATER LONDON, GB. Imported listing (source: gresearch.com). Apply on the employer's site (gresearch.com).
Imported listing (source: gresearch.com) · Apply on gresearch.com
The sections below reproduce the third-party job description for reference. AIEngineer.careers does not write or control this text.
Imported job description
Sourced from gresearch.com
G-Research is hiring an NLP Performance Engineer to optimize large-scale LLM inference workloads on advanced GPU infrastructure. This specialist Quantitative Developer will profile and benchmark model serving, implement inference optimizations, and build tooling to improve the performance and cost-efficiency of the firm's NLP platform.
Responsibilities
Profile, benchmark and optimize large-scale LLM inference workloads across compute infrastructure
Ensure efficient deployment of the latest models across GPU architectures and adapt the inference stack as hardware evolves
Design and implement inference optimizations while maintaining output quality
Develop reference implementations, libraries and tooling for efficient and reliable NLP workloads
Collaborate with researchers, senior stakeholders and engineers to design optimized solutions
Work with systems, architecture and platform teams to evolve the compute stack and influence long-term platform decisions
Requirements
Bachelor's, Master's or PhD in computer science or equivalent experience
Proven experience profiling, benchmarking and optimizing large-scale LLM inference workloads
Scientific, evidence-led approach to performance optimization using rigorous benchmarking and reproducible measurement
G-Research is a prominent quantitative finance research and technology firm that applies scientific rigour, machine learning, and advanced analytics to predict movements in global financial markets. Founded in 2001, the firm unites world-class researchers and engineers to build sophisticated trading strategies.