← Back to list

Job
· Senior
AI Research Engineer (Multi-Modal Reinforcement Learning) - 100% Remote Worldwide
AI Research Scientist
• Senior
• Remote
• Full-time
•
EU/EMEA
Tether's AI model team is looking for an engineer to advance multi-modal reinforcement learning (RL) in systems that integrate text, image and audio data. The role is research-driven and covers building new algorithms and training frameworks.
Responsibilities
- ▹Conduct research on reinforcement learning algorithms for multimodal models, including diffusion-based approaches and unified frameworks that integrate multiple modalities
- ▹Design and build RL infrastructure that supports scalable, distributed training across multimodal systems, efficiently and reliably
- ▹Develop and refine reward modeling strategies that improve training stability, align model behavior with desired outcomes and mitigate reward hacking
- ▹Create and curate multimodal simulation environments and datasets for training, evaluation and benchmarking of RL systems
- ▹Design and conduct rigorous benchmarking and evaluation protocols to measure model performance
- ▹Analyze and optimize policy performance across modalities by identifying bottlenecks in training, credit assignment and cross-modal alignment
- ▹Investigate next-generation RL paradigms that learn more effectively from environment feedback, aiming for state-of-the-art (SOTA) performance
- ▹Publish research findings in top-tier conferences (ICML, NeurIPS, ICLR, CVPR, ICCV, ECCV etc.)
Requirements
- ▹Master's degree in Computer Science or a related field; a PhD in Machine Learning, NLP, Computer Vision or a closely related discipline is preferred, along with a strong track record of AI research and publications in top-tier conferences
- ▹Proven experience running large-scale RL experiments in multimodal and vision-centric systems, including online RL settings, with measurable improvements in policy performance
- ▹Deep understanding of RL algorithms and optimization methods applied to vision and multimodal problems, focusing on policy stability, exploration and sample efficiency
- ▹Strong proficiency in PyTorch and deep learning frameworks for vision and multimodal AI, with hands-on experience building end-to-end RL pipelines (simulation, training, evaluation, deployment)
- ▹Demonstrated ability to apply empirical research to core RL challenges (sample inefficiency, exploration-exploitation tradeoffs, training instability) and to design robust evaluation frameworks
- ▹Proven track record of research publications in top-tier conferences such as ICML, NeurIPS, ICLR, CVPR, ICCV, ECCV
Soft skills
Research-driven, hands-on approachExcellent English communication
What we offer
- ▹100% remote, worldwide
About the company
Tether works in digital finance: it offers USDT, one of the world's most trusted stablecoins, along with asset tokenization services, energy solutions, AI and peer-to-peer technology (Tether Data, the KEET app) and education. The team works remotely from around the world.
Languages: Angol: kiváló kommunikációs készség
Education: Mesterdiploma számítástechnikából vagy rokon területről (PhD előnyt jelent)
Similar jobs

Job
· Senior
AI Research Scientist
Your Personal AI
AI/MLCpp
+1
💰 Salary: not specified
🌍 Remote
🗣️ EN
Himalayas

Job
· Senior
Staff Research Engineer (Pre-training)
JetBrains
AI/MLHuggingfaceKubeflow
+4
💰 Salary: not specified
🌍 Remote
🗣️ EN
Himalayas

Job
· Senior
AI Research Engineer (Pre-training – LLM & Multi-Modal)
Tether Operations Limited
AI/MLHuggingfaceLlm
💰 Salary: not specified
🌍 Remote
🗣️ EN
Himalayas

Job
· Senior
Senior Applied Research Engineer - Video
Synthesia
AI/MLDatadog
+2
💰 Salary: not specified
🌍 Remote
🗣️ EN

Job
· Senior
AI Research Engineer (Model Compression & Quantization)
Tether Operations Limited
AI/MLCppLlm
💰 Salary: not specified
🌍 Remote
🗣️ EN
Himalayas

Job
· Senior
AI Research Engineer (Pre-training - LLM & Multi-Modal)
Tether Operations Limited
AI/MLHuggingfaceLlm
💰 Salary: not specified
🌍 Remote
🗣️ EN
Himalayas