← Back to list
Job

Research Crawling Engineer

Other • Remote • Full-time European Union EU/EMEA

Wynd Labs is hiring a Research Crawling Engineer to design and operate large-scale web data acquisition systems for research and model development.

Responsibilities

  • Build and maintain large-scale web crawlers across diverse domains
  • Design high-throughput, fault-tolerant systems for data collection (millions to billions of URLs/day)
  • Handle anti-bot systems, rate limits, and dynamic/JS-heavy sites
  • Develop pipelines for cleaning, deduplication, filtering, and normalization
  • Construct and maintain datasets for research and model training
  • Monitor crawl performance, coverage, and data quality; iterate quickly
  • Collaborate with research teams to align data collection with modeling needs
  • Optimize infrastructure for cost, latency, and reliability

Requirements

  • Strong programming experience in one or more of: Go, Rust, Python, Java, or C++
  • Experience building web crawlers or large-scale data pipelines
  • Solid understanding of HTTP, networking, and browser behavior
  • Familiarity with distributed systems and parallel processing
  • Experience working with large datasets (TB-PB scale preferred)
  • Ability to debug unstable or adversarial environments

Nice to have

  • Experience with NLP pipelines or dataset curation for ML
  • Familiarity with LLM pretraining data or retrieval systems
  • Experience with headless browsers (e.g., Chrome DevTools Protocol, Playwright, Puppeteer)
  • Knowledge of proxy systems, IP rotation, and large-scale request orchestration
  • Background in data quality evaluation or benchmarking
  • Experience running workloads on cloud or bare-metal infrastructure

Soft skills

Designing systems that scale without degrading qualityPractical problem-solving under real-world constraintsSpeed of iteration and ownership

What we offer

  • Compensation based on experience and demonstrated ability to operate at scale
  • Fully remote team

About the company

Wynd Labs builds infrastructure that delivers large volumes of web data to companies training the world's most powerful AI models. The team powers Grass, a bandwidth-sharing network enabling a large-scale distributed crawler, and builds pipelines for ingesting, segmenting, and annotating video, transcript, and audio data for frontier AI labs. Fully remote team.

Similar jobs