Handshake logo

Handshake

AI Red Teamer (LLM Generalist) - Remote

🇺🇸 Remote - US 🕑 Contract 💰 TBD 💻 Research 🗓️ September 19th, 2026
Python

Edtech.com's Summary

Handshake is hiring an AI Red Teamer to stress-test large language models by designing adversarial prompts that expose safety vulnerabilities. The red teamer probes models across risk categories such as content safety, CBRN, cybersecurity, persuasion, child safety, and self-harm, working across text, image, voice, and agentic capabilities depending on the project. This work directly supports frontier AI labs by identifying weaknesses in guardrails and model behavior before they reach real users.

Highlights
  • Craft creative, multi-turn adversarial prompts to stress-test AI guardrails across diverse risk categories.
  • Discover jailbreak, evasion, and prompt injection techniques that bypass safety filters and restrictions.
  • Evaluate and score model responses against structured harm taxonomies and severity rubrics.
  • Document experiments in detail, including methodology and findings, for the broader research team.
  • Review and refine adversarial prompts created by other red-teaming team members.
  • Contribute to harm taxonomy development and inter-rater reliability calibration exercises.
  • Collaborate with engineers, data scientists, and researchers to strengthen model defenses.
  • Work regularly with disturbing content, including violence, self-harm, and hate speech, as part of structured testing.
  • Full-time contract role at 40 hours per week, fully remote within the USA.
  • Strong hands-on experience with multiple LLMs (ChatGPT, Claude, Gemini, and open-source models) expected.

AI Red Teamer (LLM Generalist) - Remote Full Description

Location: Remote (USA)
Type: Contract, 40 hours per week

About the Role
As an AI Red Teamer, you will stress-test large language models by intentionally trying to break them. Rather than checking whether an answer is correct, you will design creative, adversarial prompts that expose vulnerabilities: unsafe content, bias, broken guardrails, hallucinations, prompt injection weaknesses, and unexpected behaviors. Your work directly supports AI safety and model robustness for leading research labs.
This is a generalist red teaming role. You will probe models across the full spectrum of risk categories, including content safety, CBRN, cybersecurity, persuasion and influence operations, child safety, self-harm, over-companionship, and regulatory compliance. Red teaming may span text, image, voice, and agentic model capabilities depending on project needs.

Day-to-Day Responsibilities
  • Craft creative prompts and multi-turn scenarios to stress-test AI guardrails across diverse risk categories.
  • Discover ways around safety filters, restrictions, and defenses using jailbreak, evasion, and prompt injection techniques.
  • Explore edge cases to provoke disallowed, harmful, or incorrect outputs.
  • Evaluate and score model responses against structured harm taxonomies and severity rubrics.
  • Document experiments clearly, including what you tried, why you tried it, and what it revealed.
  • Review and refine adversarial prompts generated by other team members.
  • Contribute to harm taxonomy development, calibration exercises, and inter-rater reliability work.
  • Collaborate with engineers, data scientists, and researchers to share findings and strengthen defenses.
  • Stay current on jailbreaks, attack methods, and evolving model behaviors.
Desired Capabilities
  • Strong hands-on experience using multiple LLMs (ChatGPT, Claude, Gemini, open-source models, etc.).
  • Intuition for crafting adversarial prompts; familiarity with jailbreak or evasion techniques is a strong plus.
  • Creative, adversarial problem-solving skills.
  • Clear and thoughtful written communication.
  • Strong ethical judgment and the ability to separate adversarial thinking from personal values.
  • Self-directed, collaborative, and comfortable in feedback-heavy environments.
Content Warning
This role involves regular and deliberate exposure to harmful content, including violence, self-harm, hate speech, sexually explicit material, and child safety scenarios, as part of structured adversarial testing. Candidates must be able to engage with this material professionally and sustainably. Support resources are available.