Oxariz · Training data

High-quality training data, built by domain experts.

We build bespoke evaluation benchmarks and training datasets for leading AI labs — RLHF, code, STEM, and safety data designed by working specialists.

We also source, clean, and prepare authorized business data through our partner network, turning raw information into reliable datasets ready for analysis and AI development.

Our motto

Connecting businesses, data, and talent through 200+ company partnerships.

Talk to us →

Selected work / MuJoCo robotics

From reasoning to real control.

Our robotics tasks turn complex physical challenges into repeatable simulation environments. Models must do more than describe a solution: they must produce control strategies that work when executed.

MuJoCo simulation · 23 seconds
A vehicle follows a winding course through snow, low-traction ice, and rock obstacles. The recorded run reports zero collisions.
In this recorded run 24.74 m progress 0 collisions Reported by the video’s final summary; not an aggregate evaluation.

01 / Adaptive navigation

Snow-Course Navigation

The vehicle must make progress toward a target while adjusting its steering and speed to a changing surface. Following the centerline alone is not enough: rocks demand detours, and icy patches change how the vehicle responds.

Why this challenges models

A controller that works on one stretch can drift off course on the next. The task tests whether a model can reason about feedback, anticipate obstacles, and recover from error rather than rely on a fixed sequence of actions.

What it helps develop

Testing these decisions in simulation supports research into more adaptable mobile robots, from field inspection to navigation in variable terrain. Progress, path error, and collisions provide concrete feedback for improving a controller.

MuJoCo simulation · 22 seconds · 3.3× review speed
An underwater vehicle visits five ports, makes probe contact, completes a coded handshake, then retracts the probe and holds position.
In this recorded run 5 ports final hold shown The rollout ends with “Complete: probe retracted + hold.”

02 / Contact-rich control

Acoustic ROV Physical Commissioning

This remotely operated vehicle (ROV) must navigate between physical ports and complete a sequence of interactions. The rollout combines acoustic cues, probe alignment, contact, and a final stable hold.

Why this challenges models

Getting close to a port is only the beginning. A controller must coordinate movement with the interaction sequence, maintain alignment during contact, and transition safely to the next port. Small errors can compound across the full mission.

What it helps develop

These tasks support research into robotic inspection and servicing: systems that must sense, move, interact, and verify completion. Simulation makes it possible to test those behaviors repeatedly before considering trials with physical hardware.

Learning through verifiable feedback

Attempt. Simulate. Measure. Improve.

A model proposes a controller, the simulator executes it, and a programmatic scorer measures the outcome. Those results can guide reinforcement learning or controller refinement. The objective is useful physical behavior, not an answer that merely sounds correct.

Human-designed reference

An expert-built controller using the same allowed observations provides a practical solvability check. A successful reference is distinct from a human teleoperating the robot.

Oracle solvability

An oracle controller may use privileged simulator state to establish what is physically achievable. Its success is an upper reference, not proof that an ordinary model can solve the task.

Challenging Fable 5.1

In our team’s evaluations, Fable 5.1 has struggled to solve these tasks. They expose the gap between describing a plausible strategy and executing reliable control. The videos show demonstration rollouts, not recordings of Fable 5.1’s attempts.

Discuss a robotics task
03/Why teams pick us

Built for quality, calibrated for scale.

01

Quality-first process

Every dataset is authored by working specialists in the domain, then graded by independent reviewers against an explicit rubric.

02

Rapid delivery pipeline

Tight feedback loops with the research team, daily batches in flight, and continuous calibration mean shorter iterations.

03

Domain expert network

STEM PhDs, software engineers, and safety researchers — vetted, calibrated, and matched to the work that fits their depth.

04

Cost-effective at scale

Tiered staffing and reusable rubric infrastructure keep per-example cost predictable as the program scales.

04/How it runs

A project, end to end.

01

Brief

We meet your research leads to scope the failure modes you want closed, the schema you need, and the success criteria you'll grade against.

02

Source & calibrate

Domain experts are staffed from our network, calibrated on a gold set, and onboarded into the project-specific rubric.

03

Deliver & iterate

Datasets are shipped in your preferred schema with rubric scores and reviewer trails. Continuous review keeps the dataset alive as your models evolve.

04 // Ready when you are

Join the data engine powering frontier AI Labs.

Tell us about the failure modes your model still has — we'll come back with a sourcing plan.