← Back to list
Job

Research Engineer / Research Scientist, Data Understanding - Foundations

Other • Hybrid • Full-time • Switzerland Zurich, Switzerland

OpenAI's Data Understanding team treats data quality and curation for large-model pretraining as core research problems. The role develops new methods to select and transform data, then turns successful research into scalable data processing pipelines.

Stack

Responsibilities

  • ▹Develop new methods to select, combine and transform data
  • ▹Create datasets that improve model capabilities
  • ▹Design rigorous experiments to understand how data choices and interventions affect model learning and downstream behavior
  • ▹Work with frontier models and web-scale data to build evidence for which approaches work and why
  • ▹Translate successful research into scalable data processing pipelines

Requirements

  • ▹A strong track record of new or improved ML ideas, through publications, projects or applied research
  • ▹Own and drive a research agenda, from choosing the right problems to carrying long-running work through to impact
  • ▹Excitement about OpenAI's empirical, collaborative approach to research

Nice to have

  • ▹Thoughtfulness about AI's impact, including privacy, provenance and data quality
  • ▹Experience building high-performance deep learning or large-scale data processing systems

Soft skills

Owning a research agendaCollaboration

About the company

OpenAI is an AI research and deployment company dedicated to ensuring that general-purpose artificial intelligence benefits all of humanity. The Data Understanding team creates high-quality datasets and their quantized representation for model training: synthesizing data, building VQ representations, processing, filtering, deduplication, quality control and tokenization.

Similar jobs