The world is not what the robot thinks it is.

Arithmancy generates physics-gated training data that closes the gap between what a robot perceives and what is physically real. URDF in. Validated state-action data out.

Reality
A real warehouse scene. Cardboard boxes, a shelf, an arm.
Robot Belief
The same scene as the robot perceives it. Segmented, rendered, labeled.
The Gap

Every robot learns from data that never saw your floor. New lighting. New objects. New contact. The training set is the one thing that can get there first.

The question was never how much data you have. It is whether your data covers what is about to happen, and whether you can prove it does.

The Failures

We kept the failures.

Perception drift Contact failure The pick drops

This is a failed pick, kept on purpose. 340 failed episodes from the last run sit next to the successes as labeled negatives. A robot that has only seen success has never learned to recover.

Most data companies hide their failure rate. We publish it. A data product that hides its failures is lying about its size.

The Thesis

The data wall for robotics is physical.

Language models scraped an internet. Robots do not get one. The industry's answer is generated video that looks right and is physically false. Whoever builds the trusted, verifiable data layer wins the picks-and-shovels position for the entire robot economy.

That layer is what we build.

The Pipeline

One pipeline. Five arms. One gate.

URDF + CONTROLLER PHYSICS-SUPERVISED SIMULATION QUALITY GATES VALIDATED STATE-ACTION DATASET
  1. 01
    You send a URDF and a controller. Two files you already have. No sim expertise required.
  2. 02
    Physics-supervised simulation. A validator audits every timestep: contact points, friction cones, joint torques, conservation laws.
  3. 03
    Quality gates. Scenes that fail are refused, with the reason written down. The validator decides what ships, not the language model.
  4. 04
    Validated state-action data out. Contact forces, friction, full 6-DoF state, plain-language annotations. Reproducible to the bit.
Compare

See the gap yourself.

Drag to compare what the robot perceived against what was physically there. This is the gap we measure, and the gap we are closing in public.

Reality
Robot belief

Our sim-to-real number is being re-measured right now, because we stopped trusting the old one and caught it ourselves. We would rather say we are re-measuring it than quote a number we are about to retract.

Hard Cases

Hard cases. Hard to get right.

The long tail of cases nobody thought to script is exactly where robots fail. We generate against the failure mode you name, not a demo reel.

Human interaction

A person in the workspace the policy never trained on

Contact failure

A grasp that looked right in render and was wrong in physics

Perception error

The camera saw one thing, the world held another

Occlusion

The object is partly hidden and the policy has to infer

Physics change

Mass, friction, or lighting shifted from the training set

The Data

URDF in. Ground truth out.


    

This is the artifact. State-action trajectories with contact forces, friction, 6-DoF state, and a plain-language description of every state. 7,853 episodes on disk across 32 datasets. Deterministic replay at 6.4×10⁻¹⁵ radians. If a policy fails on our data, we reproduce the failure exactly and debug it.

Nobody debugs a diffusion video.

The Ledger

The data compounds.

0 Episodes on disk across 32 datasets
0 Physics steps
0 Ledger runs, 46 passed, 23 failed, 7 informational
0 Generation errors

Every run appends to a ledger that is public and provenance-tracked. Failures are records too. Ask us to show you the failed ones.

Domains

Built for real-world robotics.

We start where it is measurable: manipulation. Pick, place, sort, stack. That is the on-ramp. The destination is contact-rich dexterous manipulation, exactly where generative video fails hardest on physics.

Medical
Logistics
Automation
Domestic
Industrial
Proof

This is not a plan. It runs.

5,000 episodes. 15 million physics steps. Zero generation errors. A policy trained purely on our data reaches 91.3% in simulation, n=300, deploy-matched, zero demos, zero fine-tuning.

Open the demo → sim.arithmancy.ai
The Result

Trained on the gap. Measured against the real world.

This is the filmed physical pick, method disclosed. The camera policy has only ever seen rendered images, so the physical pick came through the coordinate-fed fallback path. The blocker is camera calibration, not the policy. We say it before you ask.

On the physical arm we are at repeatable contact. Every success rate above is in simulation. Full hardware record, failures included.

Get Started

Train on the gap.

Send us a URDF and one failure mode that is costing you money. We generate a physics-gated dataset against it. You measure it on your baseline, your hardware, your threshold. If the number moves, we scope a pilot.