AI Chat and Search Evaluator
Job Description
Test chat and search tools consistently, using prompts and scoring rules that expose their strengths and failures.
What you will do
• Run repeatable evaluations of chat and search behavior.
• Write and stress-test prompts covering diverse user goals and edge cases.
• Rank and annotate responses against the rubric and correct inconsistent judgments.
• Record findings, resolve ambiguous cases and incorporate quality-review feedback.
Required background
• Accurate close reading and careful attention to generated content.
• Clear written and spoken communication.
• Independent organization, time management and analytical problem-solving.
Preferred background
• RLHF, model training, annotation or other evaluation experience.
• Practical AI-tool knowledge, prompt writing and understanding of model limitations.
• Rubric development or assessment of accuracy, helpfulness and safety.
• Experience preparing ML data or working with LLMs.
Location and eligibility
• Remote work from North America or Germany.
Availability
• Hours and duration are agreed before work; no fixed commitment is stated.
Pay and engagement
$15 • $38/hr
USD hourly compensation. The applicable rate and agreed work scope are confirmed before work begins.