LLM Evaluation Engineer
Thirdlaw
- Location
- Remote, Global
- Type
- Full-time
- Posted
- Nov 05, 2025
Confirmed open on Thirdlaw's own job board on Oct 10, 2026. We re-check every day.
Job description
- Build an evaluation layer that determines whether LLM interactions comply with enterprise policies in real-time
- Design evaluation logic using semantic similarity, foundation model scoring, and rule-based systems to enforce AI safety
- Implement real-time guardrails and classifiers that connect with downstream enforcement actions like redaction or blocking
- Prototype and tune language models and prompt templates for classification and scoring purposes
- Build tools to observe, debug, and improve evaluator performance across diverse real-world data distributions
The difference you'll make
This role creates positive change by developing systems that ensure AI safety and compliance with enterprise policies, helping to prevent harmful or inappropriate AI interactions in real-world applications.
What makes you a great fit
- Experience with building evaluation systems for large language models
- Proficiency in semantic similarity techniques, foundation model scoring, and rule-based systems
- Ability to implement real-time guardrails and classifiers
- Experience with prototyping and tuning language models and prompt templates
- Skills in building tools for observing, debugging, and improving evaluator performance
Benefits
No benefits information provided in the job description.
About Thirdlaw
WebsiteNo organization information provided in the job description.
Share
Tweet Share WhatsApp Email FacebookCover letter and interview prep
Draft a cover letter for this role or practise the questions you are likely to be asked.
Open application toolsWant more roles like this?
Describe what you want and get your strongest matching jobs by email.
Find matching jobsRelated searches
LLM Evaluation Engineer
Thirdlaw