Preference Model is seeking a Research & Post-training Engineer to build automated ML research engineering systems. The role focuses on developing high-quality Reinforcement Learning (RL) training environments and optimizing post-training pipelines for large language models. You will blend research and engineering to push the boundaries of self-directed learning and model capability.
Search
Go
