Reliability, Observability & Incident Reviewer
Job Description
Work at IXO
Apply your experience in site reliability engineering to paid work at IXO. You will develop realistic evaluation examples and review AI responses for technical correctness, clear reasoning and practical usefulness. The work calls for explanations that identify the actual defect or trade-off and show how a better answer would address it.
Responsibilities
• Review SLO, error-budget and reliability calculations alongside monitoring with Prometheus, OpenTelemetry, Grafana or Datadog.
• Assess alert design, runbooks, incident procedures, on-call handoffs and post-incident analysis.
• Evaluate chaos experiments, load tests and capacity plans, explaining alert fatigue, excessive metric cardinality and silent failures.
• Record the assumptions, supporting evidence and corrections needed for another specialist to follow your review.
Experience and expertise
• 5• years in SRE, production engineering, or platform reliability.
• Practical depth in Prometheus/OpenTelemetry-based observability stacks.
• A thorough understanding of distributed-systems failure modes and incident management.
• Experience writing runbooks and leading post-incident reviews.
• Ability to work confidently with Go, Python, or another systems language.
• Familiarity with major cloud providers (AWS, GCP, Azure) at scale is required.
Working arrangements and pay
Remote work in these eligible locations: USA, UK, Canada, Germany, Australia. Planning availability: Flexible, 10-25 hours/week. IXO will confirm the actual start and schedule before acceptance. Compensation is $95 • $150/hr USD. The agreed rate, delivery requirements and review criteria are confirmed before work begins.