Evaluation Science Lead

Austin Posted 21h Leaves the board in 7 days

Open in 2 locations. Pick one to apply to.

ATS Keyword Match See which key terms from this posting your resume already has. Free. Posting says: Works unsupervised · Always on

Language carries culture, nuance, and context — as Evaluation Science Lead, you will own the scientific rigor behind measuring how our AI-powered Globalization solutions perform across 50+ languages and dozens of markets. As a strategic individual contributor on the Globalization Quality and Operations team, you will design statistically grounded evaluation frameworks combining human judgment with scalable automation — partnering with Engineering, AI Strategy, and Production to turn evaluation into a strategic capability.

This role requires strong analytical and scientific thinking, hands-on execution, and sound judgment to communicate complex findings clearly. You thrive in ambiguity and know evaluation delivers value only when operationalized — if you believe rigorous, culturally-informed evaluation is one of the highest-leverage ways to improve AI language quality globally, this role was built for you.

Minimum Qualifications
5+ years of experience in evaluation science, data science, or ML systems development, with demonstrated experience owning evaluation systems at scale end-to-end
Experience applying statistical methodology — sampling, significance testing, and confidence intervals
Practical understanding of measurement validity principles
Hands-on experience measuring annotator agreement, diagnosing divergence, and improving annotation protocols
Proficiency with statistical tools and languages (R, Python, SQL) for analysis and reproducibility
Experience designing evaluation frameworks adopted and scaled by operational teams
Strong written and verbal communication skills for technical and non-technical audiences
Direct experience handling sensitive and confidential information with integrity and discretion
Ability to be onsite; this role is an in-person, onsite position
Availability to work occasional evenings and weekends, as business needs require
Up to 10% + travel; both domestic and international

Preferred Qualifications
Master’s, PhD, or comparable experience in Statistics, Computational Linguistics, Computer Science, Psychometrics, Data Science, or related quantitative field.
Experience evaluating generative AI systems—including hallucination detection, safety and cultural alignment, autograders / LLM-as-judge systems, benchmark design, and synthetic data evaluation.
Experience with A/B testing, causal inference, or experimental design
Experience in applied linguistics — translating cultural and linguistic nuances into quantitative evaluation framework

Share your thoughts