Research Engineer, Data - Pika is a Data Engineer role (full-time). with Pika. in REMOTE, PALO ALTO, US. Imported listing (source: underprompt.com). Apply on the employer's site (underprompt.com).
Imported listing (source: underprompt.com) · Apply on underprompt.com
The sections below reproduce the third-party job description for reference. AIEngineer.careers does not write or control this text.
Imported job description
Sourced from underprompt.com
Pika is looking for a staff-level Research Engineer, Data to architect and scale data engineering systems supporting model training for multimodal foundation models. This role involves owning large-scale data pipelines, curating diverse datasets, and optimizing data processing for training. Ideal candidates have 5+ years experience in data pipelines for ML, expertise in distributed systems like Spark and Ray, and strong programming skills.
Responsibilities
Take ownership of large-scale data pipeline architecture and implementation to support model training and research workflows for text, image, audio, and video datasets.
Partner with research and engineering teams to curate, clean, and manage diverse, sensory-rich datasets for pre-training and mid-training of multimodal models.
Develop strategies and tools for scalable data ingestion, labeling, filtering, augmentation, and storage.
Optimize data processing, transformation, and delivery for large-scale distributed training pipelines.
Requirements
5+ years of experience building and scaling data pipelines for machine learning applications at staff or lead engineer level.
Strong background in data engineering and ML data curation for LLMs, VLMs, or other large-scale multimodal models.
Pika is an idea-to-video platform that helps bring creativity to motion. The company is headquartered in Palo Alto, California and works as a small, dynamic team focused on revolutionizing video creation.