Arithmancy generates physics-gated training data that closes the gap between what a robot perceives and what is physically real. URDF in. Validated state-action data out.
Every robot learns from data that never saw your floor. New lighting. New objects. New contact. The training set is the one thing that can get there first.
The question was never how much data you have. It is whether your data covers what is about to happen, and whether you can prove it does.
This is a failed pick, kept on purpose. 340 failed episodes from the last run sit next to the successes as labeled negatives. A robot that has only seen success has never learned to recover.
Most data companies hide their failure rate. We publish it. A data product that hides its failures is lying about its size.
Language models scraped an internet. Robots do not get one. The industry's answer is generated video that looks right and is physically false. Whoever builds the trusted, verifiable data layer wins the picks-and-shovels position for the entire robot economy.
That layer is what we build.
Drag to compare what the robot perceived against what was physically there. This is the gap we measure, and the gap we are closing in public.
Our sim-to-real number is being re-measured right now, because we stopped trusting the old one and caught it ourselves. We would rather say we are re-measuring it than quote a number we are about to retract.
The long tail of cases nobody thought to script is exactly where robots fail. We generate against the failure mode you name, not a demo reel.
A person in the workspace the policy never trained on
A grasp that looked right in render and was wrong in physics
The camera saw one thing, the world held another
The object is partly hidden and the policy has to infer
Mass, friction, or lighting shifted from the training set
This is the artifact. State-action trajectories with contact forces, friction, 6-DoF state, and a plain-language description of every state. 7,853 episodes on disk across 32 datasets. Deterministic replay at 6.4×10⁻¹⁵ radians. If a policy fails on our data, we reproduce the failure exactly and debug it.
Nobody debugs a diffusion video.
Every run appends to a ledger that is public and provenance-tracked. Failures are records too. Ask us to show you the failed ones.
We start where it is measurable: manipulation. Pick, place, sort, stack. That is the on-ramp. The destination is contact-rich dexterous manipulation, exactly where generative video fails hardest on physics.
5,000 episodes. 15 million physics steps. Zero generation errors. A policy trained purely on our data reaches 91.3% in simulation, n=300, deploy-matched, zero demos, zero fine-tuning.
Open the demo → sim.arithmancy.aiThis is the filmed physical pick, method disclosed. The camera policy has only ever seen rendered images, so the physical pick came through the coordinate-fed fallback path. The blocker is camera calibration, not the policy. We say it before you ask.
On the physical arm we are at repeatable contact. Every success rate above is in simulation. Full hardware record, failures included.
Send us a URDF and one failure mode that is costing you money. We generate a physics-gated dataset against it. You measure it on your baseline, your hardware, your threshold. If the number moves, we scope a pilot.