Trusted by research and engineering teams across frontier AI
- Northstar
- ARC LABS
- Berkeley Research
- Meridian
- HAWTHORNE
- Frontier AI
Areas of focus
Robotics
Sensorimotor datasets and embodied evaluation environments for agents that act in the physical world, not just answer in text.
Domain-specific datasets & benchmarks
Built for the specific failure modes generic benchmarks miss — regulated, technical, and high-consequence domains.
Model behavior research
Systematic study of how frontier models reason, fail, and drift — the evidence behind every dataset we build.
AI infrastructure & tooling
The pipelines, labeling tools, and eval harnesses that turn research findings into repeatable measurement.
RL environments & agent evaluation
Interactive environments that measure how agents actually behave under real task pressure, not just static scoring.
