Real tasks. Clear evidence. More capable models.
We source and clean business data, and build tasks, reference solutions, and evaluation environments for model post-training. Our work spans code, STEM, professional reasoning, security, and robotics.
From business data to verifiable rewards and expert demonstrations, we support model development, training, and independent evaluation.
Business Data Sourcing & Cleaning
Through our business contacts and partner network, we source internal business data authorized for sharing and agreed use. We clean, deduplicate, standardize, and validate datasets so teams receive consistent, usable data for analysis and AI development.
RLVR
Coding, STEM, and structured reasoning tasks with executable checks and reference solutions. These environments can supply verifiable reward signals based on correctness, numerical accuracy, and compliance with task rules.
Rubric-based RL
Professional tasks that require judgment across documents, spreadsheets, and competing evidence. Expert-authored rubrics define how to score accuracy, completeness, and traceability, providing a foundation for reward-based training.
Agentic RL Environments
Multi-step software tasks in reproducible environments, where agents inspect code, use tools, implement changes, and test their work. Outcome-based graders support evaluation and can be integrated into an RL training loop.
Browser-use Environments
Web-app tasks that connect interface behavior with application state. Browser checks measure workflow completion, error handling, and UI correctness, supporting verifiable feedback for agents working on web applications.
Robotics RL
MuJoCo tasks for navigation, contact, and sequential control. Simulation outcomes provide feedback on progress, collisions, alignment, and completion, supporting controller refinement and RL experiments.
SFT
Expert-written reference solutions provide a starting point for supervised examples. After review and conversion into prompt–solution pairs, they can support SFT across code, scientific reasoning, and professional workflows.
Custom Evals
Held-out tasks, reference checks, and targeted failure cases for measuring model capability. Separate evaluation sets help teams assess progress across coding, security, STEM, and professional reasoning without training on the test cases.