← Zurück zur Liste
Stelle · Senior

AI Research Engineer (Multi-Modal Reinforcement Learning) - 100% Remote Worldwide

AI Research Scientist • Senior • Remote • Vollzeit • Europäische Union EU/EMEA

Tether's AI model team is looking for an engineer to advance multi-modal reinforcement learning (RL) in systems that integrate text, image and audio data. The role is research-driven and covers building new algorithms and training frameworks.

Stack

Responsibilities

  • ▹Conduct research on reinforcement learning algorithms for multimodal models, including diffusion-based approaches and unified frameworks that integrate multiple modalities
  • ▹Design and build RL infrastructure that supports scalable, distributed training across multimodal systems, efficiently and reliably
  • ▹Develop and refine reward modeling strategies that improve training stability, align model behavior with desired outcomes and mitigate reward hacking
  • ▹Create and curate multimodal simulation environments and datasets for training, evaluation and benchmarking of RL systems
  • ▹Design and conduct rigorous benchmarking and evaluation protocols to measure model performance
  • ▹Analyze and optimize policy performance across modalities by identifying bottlenecks in training, credit assignment and cross-modal alignment
  • ▹Investigate next-generation RL paradigms that learn more effectively from environment feedback, aiming for state-of-the-art (SOTA) performance
  • ▹Publish research findings in top-tier conferences (ICML, NeurIPS, ICLR, CVPR, ICCV, ECCV etc.)

Requirements

  • ▹Master's degree in Computer Science or a related field; a PhD in Machine Learning, NLP, Computer Vision or a closely related discipline is preferred, along with a strong track record of AI research and publications in top-tier conferences
  • ▹Proven experience running large-scale RL experiments in multimodal and vision-centric systems, including online RL settings, with measurable improvements in policy performance
  • ▹Deep understanding of RL algorithms and optimization methods applied to vision and multimodal problems, focusing on policy stability, exploration and sample efficiency
  • ▹Strong proficiency in PyTorch and deep learning frameworks for vision and multimodal AI, with hands-on experience building end-to-end RL pipelines (simulation, training, evaluation, deployment)
  • ▹Demonstrated ability to apply empirical research to core RL challenges (sample inefficiency, exploration-exploitation tradeoffs, training instability) and to design robust evaluation frameworks
  • ▹Proven track record of research publications in top-tier conferences such as ICML, NeurIPS, ICLR, CVPR, ICCV, ECCV

Soft skills

Research-driven, hands-on approachExcellent English communication

What we offer

  • ▹100% remote, worldwide

About the company

Tether works in digital finance: it offers USDT, one of the world's most trusted stablecoins, along with asset tokenization services, energy solutions, AI and peer-to-peer technology (Tether Data, the KEET app) and education. The team works remotely from around the world.

Languages: Angol: kiváló kommunikációs készség
Education: Mesterdiploma számítástechnikából vagy rokon területről (PhD előnyt jelent)

Ähnliche Stellen