Everything on this page is read from a run artifact, the ledger, the frozen task battery, the benchmark outputs. Failed runs sit next to the passes.
A policy trained purely on pipeline-generated data, zero real demonstrations, zero fine-tuning, reaches 91.3% success in simulation (n=300, deploy-matched runtime). Deploy-matched means the exported model runs in the same control loop that would drive the physical robot, not the training harness.
You will also see a retired world-model track sitting at 0% on the same chart. We tried it, it measured zero, we killed it and pivoted to behavior cloning. The pivot is on the chart, not hidden in a footnote.
Replay re-integration error is 6.4×10⁻¹⁵ radians or less. An episode replays to the same physics, every time, forever.
The proof: v12 reproduced v11b seed for seed, across a storage-engine rewrite. We rewrote the engine underneath and the whole generate, train, eval chain came out bit-identical. Determinism here is not a slogan, it is a shipped feature.
5,000 episodes. 15 million physics steps. Zero generation errors. 93.2% expert success, and the 340 failures are kept as labeled negatives, not discarded. Storage runs 5.98× smaller, bit-exact.
We keep the failures. For recovery training and for eval, the failures are the product.
Five arms in one benchmark harness: Franka, UR5e, SO-ARM100, xArm7, and our lab arm. Prediction error falls on every one as pipeline data grows.
That is a benchmark registry, not five trained policies. Your robot drops into the same gates in about a day, assuming the model compiles and has a real gripper. A trained policy on your arm is what the pilot produces.
Do not look for the wins. Look for the refusal.
Six of seven prompts delivered. Eighteen of eighteen episodes gate-clean, up from zero of eleven at baseline. The seventh was refused: the validator caught books interpenetrating by 27.3 millimetres, the LLM failed three repair attempts, and it refused with the reason named.
The LLM is a sampler, not an oracle. The physics validator decides what ships.
Everything above is inside simulation. Here is our first measurement against physical reality, and it is not flattering, which is why it is here.
Our sim-to-real measurement is being re-measured right now. We had a figure, we stopped trusting it, and we caught that ourselves. Physical picks are zero today. That is the point, we are measuring, and the blocker is camera calibration, not the policy.
If our numbers do not mean exactly what they say, we do not have a product. That is the whole company.