Wynd Labs is hiring a Research Crawling Engineer to design and operate large-scale web data acquisition systems for research and model development.
Responsibilities
- ▹Build and maintain large-scale web crawlers across diverse domains
- ▹Design high-throughput, fault-tolerant systems for data collection (millions to billions of URLs/day)
- ▹Handle anti-bot systems, rate limits, and dynamic/JS-heavy sites
- ▹Develop pipelines for cleaning, deduplication, filtering, and normalization
- ▹Construct and maintain datasets for research and model training
- ▹Monitor crawl performance, coverage, and data quality; iterate quickly
- ▹Collaborate with research teams to align data collection with modeling needs
- ▹Optimize infrastructure for cost, latency, and reliability
Requirements
- ▹Strong programming experience in one or more of: Go, Rust, Python, Java, or C++
- ▹Experience building web crawlers or large-scale data pipelines
- ▹Solid understanding of HTTP, networking, and browser behavior
- ▹Familiarity with distributed systems and parallel processing
- ▹Experience working with large datasets (TB-PB scale preferred)
- ▹Ability to debug unstable or adversarial environments
Nice to have
- ▹Experience with NLP pipelines or dataset curation for ML
- ▹Familiarity with LLM pretraining data or retrieval systems
- ▹Experience with headless browsers (e.g., Chrome DevTools Protocol, Playwright, Puppeteer)
- ▹Knowledge of proxy systems, IP rotation, and large-scale request orchestration
- ▹Background in data quality evaluation or benchmarking
- ▹Experience running workloads on cloud or bare-metal infrastructure
Soft skills
Designing systems that scale without degrading qualityPractical problem-solving under real-world constraintsSpeed of iteration and ownership
What we offer
- ▹Compensation based on experience and demonstrated ability to operate at scale
- ▹Fully remote team
About the company
Wynd Labs builds infrastructure that delivers large volumes of web data to companies training the world's most powerful AI models. The team powers Grass, a bandwidth-sharing network enabling a large-scale distributed crawler, and builds pipelines for ingesting, segmenting, and annotating video, transcript, and audio data for frontier AI labs. Fully remote team.
Similar jobs
Job
Founding Engineer (Distributed Systems) - Oodle - Bangalore, India
Pear Vc
BigqueryCppClickhouseDatabricks
+8
💰 Salary: not specified
🌍 Remote
🗣️ EN

Job
️ Location Services Engineer | Maps Platform (Remote in Europe)
MapTiler
Cpp
+3
💰 Salary: not specified
🌍 Remote
Anywhere in the World
🗣️ EN

Job
AI SWE / Agentic Handover Engineer (m/f/d)
T-Systems Iberia
ArgocdCpp
+10
💰 Salary: not specified
🏢 On-site
Granada
🗣️ EN

Job
Performance Engineer - Open Source
Canonical Ltd.
CppData Science
+8
💰 Salary: not specified
🌍 Remote
🗣️ EN

Job
Performance Engineer - Open Source
Canonical
CppData Science
+8
💰 Salary: not specified
🌍 Remote
🗣️ EN

Job
GPU acceleration engineer
GECI Int.
Cpp
💰 Salary: not specified
🏢 On-site
🗣️ EN
