← Back to list
Job · Senior

AI Research Engineer (Model Compression & Quantization) - 100% Remote Worldwide

AI Research Scientist • Senior • Remote • Full-time • European Union EU/EMEA

Tether's AI model team is looking for an engineer to drive model serving and inference architectures with high throughput, low latency and a small memory footprint, including on resource-constrained devices. The role is deeply technical and requires low-level kernel optimization on mobile devices.

Stack

Responsibilities

  • ▹Design and deploy state-of-the-art model serving architectures that deliver high throughput and low latency while optimizing memory usage
  • ▹Ensure pipelines run efficiently across diverse environments, including resource-constrained devices and edge platforms
  • ▹Establish clear performance targets such as reduced latency, improved token response and minimized memory footprint
  • ▹Build, run and monitor controlled inference tests in both simulated and live production environments
  • ▹Track key performance indicators such as response latency, throughput, memory consumption and error rates, with special attention to metrics specific to resource-constrained devices
  • ▹Document iterative results and compare outcomes against established benchmarks to validate performance across platforms
  • ▹Identify and prepare high-quality test datasets and simulation scenarios for real-world deployment challenges, especially on low-resource devices
  • ▹Analyze computational efficiency and diagnose bottlenecks in the serving pipeline (e.g. suboptimal batch processing, network delays, high memory usage)
  • ▹Work closely with cross-functional teams to integrate optimized serving and inference frameworks into production pipelines for edge and on-device applications
  • ▹Define clear success metrics and ensure continuous monitoring and iterative refinement

Requirements

  • ▹Degree in Computer Science or a related field; ideally a PhD in NLP, Machine Learning or a related field, with a solid track record in AI R&D (good publications in A* conferences)
  • ▹Knowledge of Metal Shading Language (MSL): comfortable writing custom compute shaders from scratch
  • ▹Proven experience in low-level kernel optimization and inference optimization on mobile devices, with measurable improvements in latency, throughput and memory footprint
  • ▹Deep understanding of modern model serving architectures and inference optimization techniques
  • ▹Strong expertise in writing GPU kernels for mobile devices (smartphones) and deep understanding of model serving frameworks and engines
  • ▹Practical experience developing and deploying end-to-end inference pipelines on resource-constrained devices
  • ▹Demonstrated ability to apply empirical research to serving challenges (latency optimization, computational bottlenecks, memory constraints) and to design robust evaluation frameworks
  • ▹Distributed inference systems: designing and optimizing high-performance inference engines using Tensor Parallelism, Pipeline Parallelism and Expert Parallelism on GPU clusters
  • ▹Deep understanding of the math and structure behind Diffusion Models and Vision Transformers
  • ▹Understanding of pruning, quantization, Flash attention, KV cache and speculative decoding (Eagle)

Soft skills

Research-driven, hands-on approachCross-functional collaborationExcellent English communication

What we offer

  • ▹100% remote, worldwide

About the company

Tether works in digital finance: it offers USDT, one of the world's most trusted stablecoins, along with asset tokenization services, energy solutions, AI and peer-to-peer technology (Tether Data, the KEET app) and education. The team works remotely from around the world.

Languages: Angol: kiváló kommunikációs készség
Education: Számítástechnikai vagy rokon területű diploma (PhD előnyt jelent)

Similar jobs