AI Safety Expert (Seattle or Boston) at mpathic
mpathic is hiring a AI Safety Expert (Seattle or Boston) in Seattle, WA, US. On-site.
About mpathic
mpathic provides an AI-powered platform that analyzes conversational and contextual data to help AI builders evaluate, stress-test, and improve human-facing models. Its services include expert-led red teaming, ground truth benchmarking based on behavioral science, and AI-assisted annotation. The company also offers human services in red teaming, trust and safety, central rating and monitoring for clinical trials, and expert data annotation for LLM builders, with reviewers specialized in behavioral analysis, conversational design, mental health, psychiatry, social services, and clinical trial settings. It focuses on human-centered AI safety and clinical accuracy.
AI Safety Expert (Seattle or Boston) job description
About mpathic.ai
mpathic is keeping humans safe in the AI era through automated tools and expert datasets that are rooted in psychology and powered by clinicians.
We are a series A start-up backed by Tier 1 investors including Foundry.vc and Next Frontier Capital.
About the role
mpathic is seeking AI Safety Experts for a temporary project to support confidential projects evaluating and improving the safety, reliability, and real-world behavior of frontier AI systems.
This role requires on-site work in Seattle or Boston. Ideally, participants can work on-site between 20-40 hours per week.
This role is ideal for professionals with expertise in human behavior, communication, policy, education, healthcare, technology, trust & safety, or other domains where judgment, critical thinking, and nuanced decision-making matter. You'll help identify model strengths and weaknesses, uncover failure modes, and provide the high-quality human feedback that makes AI systems safer and more useful.
What You'll Be Doing
Responsibilities may include:
- Evaluating AI-generated conversations, responses, and reasoning for quality, safety, and usefulness
- Rating model outputs using structured evaluation rubrics and project guidelines
- Annotating conversational data to support AI training and benchmarking
- Identifying emerging risks, behavioral patterns, and opportunities for model improvement
- Providing written feedback that helps researchers and engineers improve model performance
- Maintaining strict confidentiality while working with proprietary AI systems and sensitive content
- Participating in calibration sessions and quality reviews to ensure consistent evaluations
What We're Looking For
Successful candidates are curious, analytical, thoughtful communicators who enjoy solving complex problems and exercising sound judgment. They are comfortable evaluating nuanced situations, following detailed guidelines, and contributing to the development of trustworthy AI.
Basic Qualifications
- Professional experience or subject matter expertise in a relevant field such as psychology, behavioral science, social work, trust & safety, research, or a related discipline
- Strong written communication skills with excellent attention to detail
- Comfortable learning structured evaluation frameworks and applying them consistently
- Strong critical thinking and problem-solving skills
- High ethical standards and sound judgment when working with sensitive or ambiguous content
- Comfortable using AI tools, Google Workspace, Slack, and other web-based collaboration platforms
- Willingness to sign NDAs and work on confidential projects

