Wynd Labs is looking for an experienced Machine Learning Engineer to develop data processing pipelines and ML solutions for large-scale NLP and LLM applications.
Responsibilities
- ▹Develop data processing pipelines and machine learning solutions for large-scale NLP and LLM applications, including improving the quality, filtering, and preparation of training datasets
- ▹Design and implement pipelines for processing and analyzing large datasets
- ▹Analyze and interpret complex time series data to provide actionable insights and solutions
- ▹Design, implement, and maintain data-driven models and algorithms
- ▹Develop techniques for dataset curation to improve the quality and efficiency of AI training data
- ▹Build scalable pipelines for filtering, deduplicating, and improving large-scale text datasets used for LLM training
- ▹Collaborate with cross-functional teams to understand data needs and deliver timely solutions
- ▹Ensure data quality and integrity throughout all processes
- ▹Utilize OCR technology to convert different types of documents into editable and searchable data
- ▹Continuously research and implement best practices in data science and machine learning
- ▹Contribute to the development and improvement of internal data processing tools and infrastructure
Requirements
- ▹Bachelor's, Master's, or Doctoral degree in Data Science, Computer Science, Statistics, or a related field
- ▹A minimum of 3 years of work or research experience dealing with large datasets
- ▹Strong coding skills in Python or other object-oriented programming languages
- ▹Graduate-level knowledge of statistics, including hypothesis testing, regression analysis, and probability
- ▹Excellent work ethic and ability to thrive in a fast-paced startup environment
- ▹Strong problem-solving skills and attention to detail
- ▹Good communication skills, able to articulate complex data concepts to non-technical stakeholders
- ▹Experience working in a high output team
Nice to have
- ▹Experience working with large-scale text datasets, NLP pipelines, or data preparation for LLM training
- ▹Experience with text deduplication, dataset filtering, corpus curation, or data distillation
Soft skills
Excellent work ethicProblem-solving mindsetAttention to detailCommunicating complex concepts clearlyTeamwork in a high-output environment
What we offer
- ▹Competitive salary, benefits, and equity package
- ▹Fully remote team
About the company
Wynd Labs builds infrastructure that delivers large volumes of web data to companies training the world's most powerful AI models. The team powers Grass, a bandwidth-sharing network enabling a large-scale distributed crawler, and builds pipelines for ingesting, segmenting, and annotating video, transcript, and audio data for frontier AI labs. Fully remote team.
Education: PhD, MSc vagy BSc Data Science, informatika, statisztika vagy kapcsolódó területen
Ähnliche Stellen
Stelle
Freelance AI Engineer
hived
AI/MLAzure DevopsCpp
+14
💰 Gehalt: keine Angabe
🏢 Vor Ort
Vlaanderen

Stelle
[R&D] AI Engineer (focus AI Agent)
Bosch
AI/MLAzure DevopsData ScienceLangchain
+2
💰 Gehalt: keine Angabe
🏢 Vor Ort
An Khanh Ward
🗣️ EN

Stelle
Machine Learning Engineer
Cuesta Partners
AI/MLData Science
+4
💰 Gehalt: keine Angabe
🌍 Remote
🗣️ EN

Stelle
Applied Scientist, Data Science
Prior Labs
AI/MLData ScienceLlm
+2
💰 Gehalt: keine Angabe
🏢 Vor Ort
New York
🗣️ EN

Stelle
Machine Learning Engineer - Fraud Risk
Rain
AI/MLData ScienceLlm
+2
💰 Gehalt: keine Angabe
🔀 Hybrid
New York
🗣️ EN
Stelle
AI / ML Engineer - Known
Pear Vc
AI/MLData ScienceHuggingfaceLlm
+2
💰 Gehalt: keine Angabe
🏢 Vor Ort
San Francisco
🗣️ EN
