Summary
This is a non-engineering content-policy evaluation role. Applicants must demonstrate relevant depth in violent fiction or media, military or emergency response, crisis or threat assessment, trust and safety, content moderation, or closely related policy work. Software engineering or LLM product experience alone is not sufficient.About HandshakeHandshake was founded on a simple belief that everyone deserves a path to a great career, regardless of where they went to school or who they know. Today, we power 25 million job seekers, 1 million+ employers, and 1,600 educational institutions.In 2025, we started Handshake AI and built the fastest-growing AI data business in history. We work directly with frontier AI lab researchers to create evaluations, publish benchmarks, and push the boundary of data. We've grown from $0 to ~$1B run rate and pay ~$60M to over 30K individuals every month.Why join Handshake now:Shape how every career evolves in the AI economy, at global scale, with impact your friends, family and peers can see and feelPartner hand-in-hand with world-class AI labs, Fortune 500 partners and the world's top educational institutionsWork together with engineers, scientists, operators, and more from Palantir, Meta, Scale AI, and former YC foundersBuild a massive, fast-growing business with billions in revenueAbout Handshake AIHuman data is the core infrastructure to AI advancement. Frontier AI labs currently improve model capabilities with various data-intensive post-training techniques. We believe that data spend for AI training will increase by 3-5x in the next few years and continue for much longer as models take on new domains. Handshake AI supports all of the frontier AI labs, working on their most complex data at the largest scale.About the RoleAs an AI Policy Specialist on the Violence & Fiction team, you will help AI models learn where the line falls between depicting violence and enabling it.Violence is one of the hardest domains in AI safety because most violent content is legitimate. Novels, games, screenplays, history, journalism, self-defense, and ordinary human frustration all involve violence, and a model that refuses them is broken. A model that helps someone plan real harm is worse. Your job is to tell the difference, case by case, and to explain your reasoning clearly enough that it can train a model.You will read user requests, model responses, and conversation history, then decide which policy category applies and whether the model's response was appropriate. The interesting cases are the close ones: a torture scene that is either a chapter of a thriller or an interrogation manual with character names; a message that reads as venting about a boss or as a plan; a "realistic" combat question from a novelist that is also a real-world capability question. One word, one contextual detail, or one shift in intent changes the answer.We are looking for people who already have strong instincts about violence in at least one of these areas: how it works in fiction, how it works in the real world, or how it shows up in people who are struggling. You do not need all three. You need one deep and the judgment to learn the rest.This is not rote annotation. Policies cannot anticipate every edge case, and good evaluators do not apply them mechanically. You will balance policy text and intent with customer expectations, conversation context, precedent, and team calibration.What You Will DoEvaluate user requests and AI model responses involving violence, weapons, threats, and dark fiction within the full conversation contextDistinguish fictional, educational, historical, and defensive violence from requests that seek real-world uplift or express real intent to harmAssess whether a model's response gives meaningful real-world capability, regardless of how the request was framedDistinguish expressions of anger, frustration, or dark humor from credible threats or crisis indicatorsSelect the most defensible classification when a case is genuinely ambiguous, and write concise rationales that cite policy language and conversation detailsWrite and refine adversarial or borderline prompts that probe where a model draws the lineIdentify policy gaps, contradictions, and emerging edge cases, and raise them with project leads and policy teamsParticipate actively in calibration discussions; challenge interpretations respectfully and update your judgment when stronger reasoning emergesApply customer policy consistently without substituting personal beliefs for the policy standardMaintain accuracy and attention to detail across repeated, feedback-heavy evaluationsYou May Be a Fit IfYou have spent serious time in violent fiction as a writer, game master, game designer, screenwriter, or editor, and you know what good dark fiction looks like and what a story-shaped extraction attempt looks likeYou have real-world exposure to violence and its consequences through military, law enforcement, security, EMS, emergency medicine, or similar work, and you can tell movie logic from what actually worksYou have worked with people in distress through crisis lines, counseling, threat assessment, domestic violence advocacy, school safety, or trust and safety, and you know the difference between "I could kill him" and a planYou use AI tools heavily and have opinions about where they refuse too much, help too much, or miss the pointYou notice when one word, contextual detail, or change in intent materially affects the answerYou can hold a strong opinion without becoming attached to being rightYou explain judgment calls clearly enough that another person can audit your reasoningYou can separate your personal views from the standard a customer has asked you to applyYou remain careful and consistent during repetitive work with difficult materialYou communicate clearly and precisely in writingStrong candidates may come from fiction writing, game design or game mastering, film and TV, military or law enforcement, emergency medicine, crisis counseling, threat assessment, trust and safety, content moderation, journalism, or law. We care more about how you reason than where you learned to reason. A degree, a clearance, and a technical background are not required.Nice to HavePublished or produced work involving violence: novels, short fiction, screenplays, comics, tabletop or video game content, mods, or fan fiction with an audienceMilitary, law enforcement, corrections, private security, or armed professional experienceTraining or professional experience in crisis intervention, threat assessment, forensic or clinical psychology, or violence preventionExperience with firearms, martial arts, or other weapons disciplines as an instructor, competitor, or professionalPrior work in AI evaluation, red teaming, data annotation, RLHF, trust and safety, or content moderationExperience evaluating outputs from ChatGPT, Claude, Gemini, or other language models in a professional capacityFamiliarity with calibration sessions, inter-rater agreement, or adjudication workflowsPrior AI evaluation experience is helpful, but it is not required.Sensitive-Content NoticeThis role involves regular and deliberate engagement with graphic material. This is the core of the job, not an occasional part of it. Evaluations will routinely include depictions of violence, weapons, torture, abuse, threats, self-harm, and death, including violence against vulnerable people, along with the emotional distress that often surrounds it.The work is conducted within structured evaluation frameworks and professional guidelines, with exposure limits, content rotation, and access to mental health support. Candidates must be able to engage with this material carefully, responsibly, and sustainably while maintaining sound judgment and consistent work quality.Role DetailsLocation: Seattle, WA, onsite Monday-FridayCompensation: $55-$90/hrEmployment classification: W-2Schedule: 8:00 AM-5:00 PM PTWeekly commitment: Monday-FridayAssignment: OngoingPlanned start date: September 21, 2026Benefits eligibility: Benefits eligible#J-18808-Ljbffr
Job Description
This is a non-engineering content-policy evaluation role. Applicants must demonstrate relevant depth in violent fiction or media, military or emergency response, crisis or threat assessment, trust and safety, content moderation, or closely related policy work. Software engineering or LLM product experience alone is not sufficient.About HandshakeHandshake was founded on a simple belief that everyone deserves a path to a great career, regardless of where they went to school or who they know. Today, we power 25 million job seekers, 1 million+ employers, and 1,600 educational institutions.In 2025, we started Handshake AI and built the fastest-growing AI data business in history. We work directly with frontier AI lab researchers to create evaluations, publish benchmarks, and push the boundary of data. We've grown from $0 to ~$1B run rate and pay ~$60M to over 30K individuals every month.Why join Handshake now:Shape how every career evolves in the AI economy, at global scale, with impact your friends, family and peers can see and feelPartner hand-in-hand with world-class AI labs, Fortune 500 partners and the world's top educational institutionsWork together with engineers, scientists, operators, and more from Palantir, Meta, Scale AI, and former YC foundersBuild a massive, fast-growing business with billions in revenueAbout Handshake AIHuman data is the core infrastructure to AI advancement. Frontier AI labs currently improve model capabilities with various data-intensive post-training techniques. We believe that data spend for AI training will increase by 3-5x in the next few years and continue for much longer as models take on new domains. Handshake AI supports all of the frontier AI labs, working on their most complex data at the largest scale.About the RoleAs an AI Policy Specialist on the Violence & Fiction team, you will help AI models learn where the line falls between depicting violence and enabling it.Violence is one of the hardest domains in AI safety because most violent content is legitimate. Novels, games, screenplays, history, journalism, self-defense, and ordinary human frustration all involve violence, and a model that refuses them is broken. A model that helps someone plan real harm is worse. Your job is to tell the difference, case by case, and to explain your reasoning clearly enough that it can train a model.You will read user requests, model responses, and conversation history, then decide which policy category applies and whether the model's response was appropriate. The interesting cases are the close ones: a torture scene that is either a chapter of a thriller or an interrogation manual with character names; a message that reads as venting about a boss or as a plan; a "realistic" combat question from a novelist that is also a real-world capability question. One word, one contextual detail, or one shift in intent changes the answer.We are looking for people who already have strong instincts about violence in at least one of these areas: how it works in fiction, how it works in the real world, or how it shows up in people who are struggling. You do not need all three. You need one deep and the judgment to learn the rest.This is not rote annotation. Policies cannot anticipate every edge case, and good evaluators do not apply them mechanically. You will balance policy text and intent with customer expectations, conversation context, precedent, and team calibration.What You Will DoEvaluate user requests and AI model responses involving violence, weapons, threats, and dark fiction within the full conversation contextDistinguish fictional, educational, historical, and defensive violence from requests that seek real-world uplift or express real intent to harmAssess whether a model's response gives meaningful real-world capability, regardless of how the request was framedDistinguish expressions of anger, frustration, or dark humor from credible threats or crisis indicatorsSelect the most defensible classification when a case is genuinely ambiguous, and write concise rationales that cite policy language and conversation detailsWrite and refine adversarial or borderline prompts that probe where a model draws the lineIdentify policy gaps, contradictions, and emerging edge cases, and raise them with project leads and policy teamsParticipate actively in calibration discussions; challenge interpretations respectfully and update your judgment when stronger reasoning emergesApply customer policy consistently without substituting personal beliefs for the policy standardMaintain accuracy and attention to detail across repeated, feedback-heavy evaluationsYou May Be a Fit IfYou have spent serious time in violent fiction as a writer, game master, game designer, screenwriter, or editor, and you know what good dark fiction looks like and what a story-shaped extraction attempt looks likeYou have real-world exposure to violence and its consequences through military, law enforcement, security, EMS, emergency medicine, or similar work, and you can tell movie logic from what actually worksYou have worked with people in distress through crisis lines, counseling, threat assessment, domestic violence advocacy, school safety, or trust and safety, and you know the difference between "I could kill him" and a planYou use AI tools heavily and have opinions about where they refuse too much, help too much, or miss the pointYou notice when one word, contextual detail, or change in intent materially affects the answerYou can hold a strong opinion without becoming attached to being rightYou explain judgment calls clearly enough that another person can audit your reasoningYou can separate your personal views from the standard a customer has asked you to applyYou remain careful and consistent during repetitive work with difficult materialYou communicate clearly and precisely in writingStrong candidates may come from fiction writing, game design or game mastering, film and TV, military or law enforcement, emergency medicine, crisis counseling, threat assessment, trust and safety, content moderation, journalism, or law. We care more about how you reason than where you learned to reason. A degree, a clearance, and a technical background are not required.Nice to HavePublished or produced work involving violence: novels, short fiction, screenplays, comics, tabletop or video game content, mods, or fan fiction with an audienceMilitary, law enforcement, corrections, private security, or armed professional experienceTraining or professional experience in crisis intervention, threat assessment, forensic or clinical psychology, or violence preventionExperience with firearms, martial arts, or other weapons disciplines as an instructor, competitor, or professionalPrior work in AI evaluation, red teaming, data annotation, RLHF, trust and safety, or content moderationExperience evaluating outputs from ChatGPT, Claude, Gemini, or other language models in a professional capacityFamiliarity with calibration sessions, inter-rater agreement, or adjudication workflowsPrior AI evaluation experience is helpful, but it is not required.Sensitive-Content NoticeThis role involves regular and deliberate engagement with graphic material. This is the core of the job, not an occasional part of it. Evaluations will routinely include depictions of violence, weapons, torture, abuse, threats, self-harm, and death, including violence against vulnerable people, along with the emotional distress that often surrounds it.The work is conducted within structured evaluation frameworks and professional guidelines, with exposure limits, content rotation, and access to mental health support. Candidates must be able to engage with this material carefully, responsibly, and sustainably while maintaining sound judgment and consistent work quality.Role DetailsLocation: Seattle, WA, onsite Monday-FridayCompensation: $55-$90/hrEmployment classification: W-2Schedule: 8:00 AM-5:00 PM PTWeekly commitment: Monday-FridayAssignment: OngoingPlanned start date: September 21, 2026Benefits eligibility: Benefits eligible#J-18808-Ljbffr
Government Careers
Government jobs offer stability, competitive benefits, and the chance to make a meaningful impact on your community and country.
Whether you’re starting your career or seeking new opportunities, these roles provide pathways for growth, security, and service.
Explore positions across a wide range of fields and take the first step toward a rewarding future in public service.
MORE JOBS
-
Security Officer: Flexible Shifts & Customer Service
- Gainesville, Georgia
- Securitas
- Sep 07, 2026
-
Security Officer, Per Deim, Evening Shift (8-Hours)
- Santa Ana, California
- KPC Health
- Sep 07, 2026
-
Security Officer - Luxury Retail
- San Francisco, California
- GardaWorld
- Sep 07, 2026
-
Experienced Customs and Border Protection Officer
- Atascadero, California
- US Customs and Border Protection
- Sep 07, 2026
-
Dispatcher / Service Coordinator
- Tampa, Florida
- TUDI
- Sep 07, 2026
-
Combat Engineer
- Seattle, Washington
- U.S. Army
- Sep 07, 2026