Everything a machine has never felt.

We are building the instrument that records the whole sensory field of physical work, and the first corpus that could give machines priors about the world.

Move through it

The premise

A person carries a lifetime of physical priors. A machine carries pixels and joint angles, and nothing else.

Every embodied model on earth is trained on a thin visual slice of the world. Not because vision is enough, but because vision is what anyone thought to record. Heat, slip, resistance, chemistry, the sound a material makes as it changes state: none of it exists in any dataset, anywhere.

This is the first time anyone has tried to record all of it at once, from a person actually doing the work, in one coherent stream. Not a lab bench with instruments pointed at an object. The complete sensory record of skilled physical action.

Where it goes

Anywhere skilled hands work.

The instrument is not built for one field. Wherever a person is doing something a machine cannot yet do, the same senses are carrying the same information, and none of it is being recorded. These are the places where the missing channels are not a nicety but the entire basis of the judgement.

Senses that carry the judgement here Not decisive in this sector Seven marks per card, in channel order

What we need

People who want the answer.

Places to record, teams who would train on channels that have never existed, and anyone working on sensing nobody has yet put on a working person.

The priors problem

A cook does not solve the kitchen. They recognise it.

A person walking into an unfamiliar kitchen already knows how things fall, how heat travels through a pan handle, how much a surface gives before it slips, what a material sounds like when it is about to break. One look is enough because almost all of the work was done before they arrived.

That accumulated stock of expectation is what makes a human competent somewhere new. It is called a prior, and every robot in the world is missing it.

Where a machine would get them

Only from data. And the data is wrong in three specific ways.

01

Wrong channels

Physical priors are not built by looking. They are built by contact: by burning yourself, by feeling something slip, by smelling a change before you could see it. Those are exactly the channels no dataset records. You cannot learn a prior about heat from a corpus that has never measured temperature.

02

Wrong axis of scale

Robot data has been scaled the wrong way, and this is now measured rather than argued. Generalisation scales as a power law in the number of environments and objects, not in the number of demonstrations. The field has been collecting more repetitions in the same few rooms.

03

Wrong source

Teleoperation data records what a robot's actuators did. It does not record what the person controlling it perceived. The expertise, and therefore the prior, lives in the human sensory stream, and that stream is thrown away at the moment of capture.

The field's three answers

Each one is real. None of them closes the gap.

These are the serious positions, stated at their strongest, with the numbers that actually decide them.

World models

Learn to imagine the world, then plan inside it.

The strongest evidence for this position is also the strongest evidence against relying on it. The flagship system is the only one in the field that genuinely demonstrates the target property: dropped into a lab it has never seen, with no data from that environment and no task-specific training, it works.

The catch16 seconds per action. A shipping humanoid's reflex tier runs at 1,000 per second. That is four orders of magnitude too slow to be the control loop, however good the representation is.

Scale

Stop engineering. Add compute and data.

The majority position, and it has seventy years of history behind it. General methods that leverage computation win in the long run, and hand-designed structure gets overtaken. This has happened repeatedly and it will happen again.

The catchRobotics has neither the data of vision nor the clean symmetry of chemistry. The largest robot pretraining corpora are around 10,000 hours, against more than a million hours of video for comparable vision systems.

Structural priors

Build the physics into the architecture.

Encode the symmetries of the physical world directly, so the model does not have to discover them. In the low-data regime this wins decisively and the effect is quantified, not hand-waved.

The catch200 demonstrations beat 1,000 without it, and success rose 21.9% at 100 demonstrations. But baked-in structure gets overtaken once the data arrives, and whole subfields have been caught out this way.

Where we are unique

The field is arguing about how to give machines priors. We think the priors were never recorded.

Every serious position in that debate is an argument about architecture: what shape the model should be, how much structure to build in, how much to let scale discover. All three take the input as given. But a human's physical priors are not architectural. They are experiential. They were assembled from a lifetime of contact, heat and chemistry, and if none of that was ever measured, no architecture recovers it and no amount of scale recovers it either. Scale multiplies whatever you sensed. It cannot multiply what you never sensed at all.

This is why we are not a model company arguing about architecture, and not a data company collecting more of the same. We are changing what gets measured.

Everyone gathering robot data records actuators and cameras. Everyone gathering human data records video and hand position. Nobody records what the person's body was actually sensing while they worked. Slip, heat and chemistry sit at zero across every public dataset in every language.

And the honest scope of the claim: this is an argument about capital efficiency, not about the eventual ceiling. We are competing in the regime of thousands to hundreds of thousands of trajectories, where the evidence for richer input is unambiguous and measured. If someone eventually reaches the asymptote with vision alone and enough compute, they may well get there. We think the cheaper road runs through sensing the world properly first, and we have designed the experiment that decides it.

Next

What the instrument actually records.

Seven channels, one moment, and the reason making them describe the same instant is the hard part.

Technology

Seven channels, one moment, one recording.

The instrument is worn by a person doing real work. It records what they see, hear, touch, feel as heat and detect as chemistry, at the moment they are doing it. Select a channel.

The hard part

These senses do not run at the same speed.

A microphone resolves microseconds. A chemical sensor takes seconds. Between them sit contact, motion, slip and heat, spread across roughly seven orders of magnitude. Recording them is not the difficulty. Making them describe the same instant, provably, is.

Every existing multi-sensor system resolves this the same way, and that way destroys the fast channels. Ours does not. How is not described on this page.

01

The instrument

Wearable capture built for this, because nothing you can buy records these channels together on a person who is working rather than on a bench with instruments aimed at an object.

02

The corpus

Real work in real places, not staged demonstrations. Deliberately small and dense. The aim is not the most hours, it is the first hours that contain what nobody has recorded.

03

The models

Trained on channels no embodied model has ever seen, and judged on one question: does it hold up somewhere it has never been. Deployed robots will not carry every sense the instrument does, and the models are designed for that from the start.

Where this stands

What is established, and what is not.

Instrument companies are judged on candour about their own error bars.

Established by others

Extra senses raise performance

Adding non-visual channels within touch improved policy success by 63%.

Sparsh-X · arXiv:2506.14754

Touch and hearing help together

Fusing vision, touch and audio improved success by 20%.

FuSe · arXiv:2501.04693

Training senses need not survive to deployment

Policies trained with richer sensing than the robot will ever carry still deploy on cameras alone.

Scaffolder · arXiv:2405.14853

Open, and ours to answer

Whether it helps where it counts

Published gains are averages. Our question is whether these senses help disproportionately more somewhere unfamiliar. Nobody has measured that. We are.

Whether chemistry carries usable signal

No dataset pairs it with physical work, so it cannot be settled with existing data. Our strongest claim and our thinnest evidence, simultaneously.

What it costs per usable hour

Unknown until the first sessions run. We will publish it.

About

Aisthe

From koinē aísthēsis, Aristotle's term for the faculty that binds the separate senses into a single perception. It is the oldest name for the thing we are building, and it is still the correct one.

Machines are being asked to act in the physical world with a fraction of the sensory access a person has. We think that is the constraint, not model scale, and that nobody has tested it because the recording has never existed. So we are making the recording.

Position

We would rather spend a small amount finding out we are wrong than a large amount finding out slowly.

The central question has a cheap answer. Before building anything, we are running it against public data, with the falsification criterion written down in advance. If the effect is not there, we will say so publicly and the instrument does not get built.

That discipline is the company. Everything else follows from it.

People

Who is behind it.

Mahmoud Omar

Chief Executive Officer

Origin of the sensory-priors thesis behind Aisthe. Owns the research programme, the falsification discipline that governs it, and the case for why this is worth building at all.

Noah Zerkin

Chief Technology Officer

Owns the instrument. The wearable capture hardware, the sensing stack that has to survive contact with real working environments, and the timing architecture that makes a multi-sensory recording mean one moment rather than seven.

Rulin Zhao

Product

Owns what the corpus becomes. Which sectors are recorded first, what a customer actually receives, and how the data turns into something a robotics team can build on rather than a pile of files.

Get in touch

If this is the problem you have been waiting for, say so.

Places to record, people to build with, and investors who fund questions rather than certainties. A technical briefing is available under NDA.