Research Scientist, Data

Pika · Palo Alto HQ

Posted
59 days ago
Last confirmed live
1 day ago

What this role involves

Pika is seeking a staff or lead-level Research Engineer, Data to architect and scale data engineering systems supporting model training for multimodal foundation models. The role involves owning large-scale data pipelines, curating diverse datasets for text, image, audio, and video, and collaborating with research and engineering teams. The position requires 5+ years of experience in data pipelines for ML, strong skills in distributed data systems and cloud platforms, and a passion for enabling generative AI.

Skills this posting asks for

  • python
  • sql
  • pyspark
  • spark
  • hadoop
  • ray
  • aws
  • gcp
  • azure
  • data labeling
  • data filtering
  • deduplication
  • quality assurance
  • dataset management
  • etl
  • ml data curation
  • llm
  • vlm
  • multimodal models

Requirements

  • 5 years of experience
  • Level: lead
  • Remote policy: hybrid

From the employer’s posting

ABOUT THE ROLE At Pika, we are pioneering the next generation of creative infrastructure built around real-time, multimodal generation and intelligent agentic platforms. We are looking for a staff or lead-level Research Engineer, Data to architect and scale data engineering systems supporting mode…

Read the full description on Pika’s careers page

Apply without filling the form

Approve this role and the application is completed for you, including a résumé tailored to it. You get a confirmation when it lands, and a credit is only spent when a submission is confirmed.

Other roles at Pika

All 4 roles at Pika