← Back to list
Job · Mid-level

AI Safety Specialist (Red Teaming)

Other • Mid-level • Remote • Full-time • Finland Finland

Mercor connects elite technical talent with leading AI research labs from its San Francisco base. This contract role involves red teaming conversational AI models to find jailbreaks, prompt injections and misuse cases, then documenting the findings.

Responsibilities

  • ▹Red team conversational AI models and agents for jailbreaks and prompt injections
  • ▹Annotate failures and classify vulnerabilities
  • ▹Apply structured testing using taxonomies and playbooks
  • ▹Document reproducible reports and attack cases
  • ▹Work independently and asynchronously to meet deadlines

Requirements

  • ▹Native fluency in English and Finnish
  • ▹Prior red teaming experience in AI adversarial work, cybersecurity or socio-technical testing
  • ▹Strong communication skills for technical and non-technical stakeholders

Nice to have

  • ▹Adversarial ML experience (jailbreak datasets, prompt injection, RLHF/DPO attacks)
  • ▹Cybersecurity background (penetration testing, exploit development, reverse engineering)
  • ▹Experience in socio-technical risk analysis

Soft skills

Creative, unconventional thinking for adversarial testingComfortable working independently across multiple projects and clients

What we offer

  • ▹Flexible, project-based contract work
  • ▹Hourly rate of $48-62/hr

About the company

Mercor connects elite creative and technical talent with leading AI research labs; its investors include Benchmark, General Catalyst and Peter Thiel.

Similar jobs