LLM Application & Prompt Reliability Reviewer
Job Description
Work at IXO
Apply your experience in LLM application engineering to paid work at IXO. You will develop realistic evaluation examples and review AI responses for technical correctness, clear reasoning and practical usefulness. The work calls for explanations that identify the actual defect or trade-off and show how a better answer would address it.
Responsibilities
• Review prompt structure, examples and reasoning-oriented task design across major model providers.
• Evaluate application code, tool calls, structured outputs and JSON reliability in SDK and platform integrations.
• Assess context and token management, caching and evaluation harnesses, identifying grounding, injection-resilience and regression-testing gaps.
• Record the assumptions, supporting evidence and corrections needed for another specialist to follow your review.
Experience and expertise
• 3• years building production LLM applications with at least two major providers.
• Practical depth in prompt engineering patterns and failure modes.
• A thorough understanding of tool use, structured outputs, and agent loops.
• Experience with eval frameworks (lm-eval-harness, OpenAI Evals, custom harnesses).
• Ability to work confidently with Python or TypeScript SDKs for LLM development.
• Familiarity with fine-tuning and RAG architectures is an advantage.
Working arrangements and pay
Remote work in these eligible locations: USA, UK, Canada, Germany, Australia. Planning availability: Flexible, 10-25 hours/week. IXO will confirm the actual start and schedule before acceptance. Compensation is $95 • $160/hr USD. The agreed rate, delivery requirements and review criteria are confirmed before work begins.