This training is required before participating in any project involving sensitive, adversarial, or potentially harmful content. It covers what you'll encounter, how to evaluate it responsibly, and how to protect your wellbeing throughout.
This is opt-in work. You are not required to participate, and opting out at any point — including after you've started — will not affect your standing on the team.
If you choose to participate, you are agreeing to engage with content that may include violence, hate, self-harm, sexual content, manipulation tactics, misinformation, and other adversarial material — for the purpose of evaluating and improving AI safety.
Consent here is not a checkbox. It's an ongoing agreement. If at any point the work stops feeling manageable, you have the right to stop — no explanation required.
Red teaming comes from military and security contexts, where teams would deliberately try to break their own systems to find weaknesses before adversaries did. Applied to AI, it means intentionally probing models for failure modes — the ways they can be misled, manipulated, or made to produce harmful outputs.
The goal is not to celebrate what breaks. The goal is to find it, document it, and give model developers the data they need to make systems safer before those systems reach real users.
You are not causing harm by doing this work. You are preventing it. That distinction is worth holding onto when the content gets difficult.
AI systems are deployed at scale. A failure mode that appears rare in testing can reach millions of people in production. The evaluators who find those failure modes before deployment are doing work that matters far beyond the task in front of them.
The best safety evaluators aren't just people who can handle difficult content. They're people who can think carefully under pressure, notice what's wrong and why, and communicate findings precisely enough for someone else to act on them. That combination is rare and valuable.
The following categories represent the types of sensitive content this evaluation work covers. Not every task includes all categories — your project brief will specify scope.
If a task falls outside the scope defined in your project brief, do not proceed. Flag it for review. Scope creep in sensitive content work is not a minor issue.
Every response you evaluate gets assessed across four dimensions. Consistency here is what makes your data useful.
Ambiguity is a feature of this work, not a bug. You will encounter cases where the right answer isn't obvious. The evaluators who do this well aren't the ones who never doubt — they're the ones who have a process for what to do when they doubt.
When you're unsure: document your uncertainty. A rating with a clear explanation of why it was hard is more valuable than a confident rating that obscures a real edge case.
You are not alone in these decisions. Your team lead and the review queue exist specifically for cases you're not sure about. Using them is good evaluation practice, not weakness.
Your documentation is what makes your evaluation useful. A rating with no reasoning doesn't give researchers what they need to act.
"Prompt requested step-by-step instructions for a controlled substance synthesis. Model provided a partial response that hedged on specifics but included enough detail to be directionally harmful. Rated 4. Flagged for review — the partial compliance pattern is worth documenting as a systematic edge case."
"Model gave harmful response. Rated 4."
Escalation in this context has two meanings. The first is a process: routing a specific finding to a reviewer because it requires more senior judgment. The second is personal: stopping your session because you've hit a limit.
Both are valid. Both are expected. Neither requires justification or apology.
If you encounter CSAM or content you believe represents a credible real-world threat to an identifiable person, stop immediately. Do not rate. Do not document. Use the emergency escalation channel.
Repeated exposure to harmful content carries real psychological risk. Vicarious trauma, desensitization, and burnout are documented occupational hazards of this kind of work. That's not a warning to scare you — it's information you need to make good decisions about how you work.
The evaluators who sustain this work long-term are the ones who take the protective practices seriously — not the ones who push through without support.
"Resilience in difficult work doesn't mean being unaffected. It means knowing what affects you, building practices around that, and asking for support before you need it urgently."
Psychological safety in this context means something specific: you can be honest about how the work is affecting you, ask questions about scope or ethics, and push back on tasks you have concerns about — without fear that doing so will reflect poorly on you.
That safety only works if both sides honor it. Management's responsibility is to create the conditions. Yours is to use them.
Completing this training and signing the opt-in form means you're willing to begin this work. It does not mean you've agreed to every task, every project, or every session for the duration of your employment.
The power to say no belongs to you throughout this work — not just at the beginning of it.
If you're unsure whether your concern is significant enough to act on, it is. The threshold for raising something is not "I'm certain this is a problem." It's "something feels off."
This work is important and it asks something real of the people who do it well. Your judgment, your honesty about what you're seeing, and your attention to your own limits are not soft skills here — they're the core of the job.
Do the work carefully. Take care of yourself. Ask for support early.