ML Infrastructure and Serving Evaluator
Job Description
IXO is engaging specialists to evaluate the reliability and performance of deployed machine-learning systems. Your contribution is technical work: make the reasoning inspectable, identify substantive errors and provide evidence that supports a reliable assessment.
Work you will do
• Evaluate model-serving and training infrastructure, looking at packaging, registries, rollout strategy and reproducible execution.
• Investigate accelerator utilization, distributed behavior, profiling results and serving bottlenecks with the tools relevant to the assignment.
• Review Python and framework-level implementation choices, and document correctness, deployment and observability findings with practical remedies.
• Evaluate registry and feature-store behavior, ML deployment infrastructure and reliable operational workflows.
Qualification routes
Qualify through the assessment route below or the additional specialist route. Each route has its own background and regional criteria; these routes are alternatives.
Assessment route requirements
• LLM infrastructure route: two or more professional years working directly on model infrastructure, serving systems or accelerator performance, with production experience in PyTorch or JAX.
• Demonstrate at least one specialty: kernel optimization using Pallas, Triton or CUDA; profiling with Nsight, Kineto, torch.profiler or XLA/JAX tooling; distributed-workload debugging; or serving with Ray Serve, vLLM, TensorRT-LLM or SGLang. Explain KV caches, continuous batching or paged attention where relevant.
• Reason about memory, latency and throughput on accelerators such as TPU, B200, H100 or A100; show increasing technical responsibility.
Preferred background
• Breadth across several systems specialties; custom operators, compiler/graph work or distributed training with Megatron, DeepSpeed, FSDP or DDP.
Additional specialist route
The senior MLOps route favors five or more years, Kubernetes serving, a major cloud-ML platform, Docker/Helm/IaC, model registries and Feast/Tecton-style feature stores. Use Python plus Go, Rust or Java; GPU scheduling and distributed training are helpful.
Deliverables
Submit the completed technical artifact or assessment with its supporting evidence, explicit assumptions, reproducible checks where applicable, and concise reasons for each material judgment. Address review findings within the agreed scope.
Location and schedule
The LLM-infrastructure regional track requires US residence; the broader MLOps track retains its listed eligible countries. Additional specialist route eligibility: USA, UK, Canada, Germany, Australia. Remote assignments are scheduled by agreement, with no guaranteed weekly volume. Availability planning can include 40 hours per week depending on the track. IXO confirms the applicable timing before work.
Pay and working terms
$95 • $150/hr USD. The agreed hourly rate, scope, schedule and acceptance criteria are confirmed before work starts. Applying does not guarantee an assignment. Use public, licensed or otherwise authorized material only; do not submit confidential employer information, personal data or restricted research.