Custom model evaluation

Measure capability with tasks, evidence, and clear criteria.

Start with the capability you need to measure

Oxyriz builds custom evaluation tasks across coding, STEM, professional reasoning, agent workflows, and robotics simulation. A useful evaluation begins with a concrete question: can a model complete the workflow, produce a correct solution, recover from an error, or satisfy a set of requirements? The task and grader should reflect that question.

Reference checks and expert rubrics

Executable checks can measure correctness when an outcome is objectively testable. Expert-authored rubrics can assess work that requires judgment across documents, spreadsheets, or competing evidence. Reference solutions help establish solvability. Neither a successful reference run nor one model failure establishes general model capability; the evaluation design and reporting context matter.

Robotics example: snow-course navigation

Our MuJoCo showcase includes a vehicle navigating a winding snow course with rocks and low-traction surfaces. The task requires steering and speed adjustments rather than a fixed sequence of actions. The recorded run’s final summary reports 24.74 metres of progress and zero collisions. Those figures describe that recorded run, not aggregate model performance.

Robotics example: underwater commissioning

A second simulation shows an underwater remotely operated vehicle performing a sequential commissioning task: alignment with ports, probe contact, a coded handshake, and safe retraction and holding. It illustrates how completion depends on both physical control and the correct sequence of actions. These are simulation examples, not demonstrations of deployment on physical robots. Watch both recorded tasks.

Keep evaluation separate from training

Held-out tasks and targeted failure cases help assess progress without training on the test cases. We discuss the task scope, success criteria, reference checks, and expected reporting before delivery. If the objective is to develop behavior instead, explore our training data and RL environments.

What should a project enquiry include?

Share the model or agent workflow being evaluated, the capability you care about, representative failure modes, constraints on tools or environments, and the evidence you need from a successful evaluation. We can then discuss an appropriate task format and review process. Discuss an evaluation project.

Discuss your project →