The World at Eye Level: Mapping at Human Scale to Solve AGI
This was originally published on Jeffrey Martin’s LinkedIn page and is part of a series.
This is a transcript from a talk I gave at the 2026 3DISE conference in Prague. You can watch the talk here as well:
Some of you know me as an experimental photographer; some of you know me as running Mosaic, where we make professional spatial camera systems for capturing scenes like these.
For the past few months, I’ve been thinking a lot about how what we’re doing intersects with what other people are doing in AI — specifically, the AI that is not the LLM stuff, but the other stuff. We call it world models. There are a couple of different things people mean when they talk about world models, so I want to clarify. I’m not talking about generative video like Sora, or even about things like World Labs [Marble and RTFM generative worlds], which is a kind of immersive feeling. I’m talking more about things related to embodied cognition — what allows our minds to develop and function so well using remarkably few resources and so little energy.
To put it in context: a rat brain has maybe 200 million neurons and maybe a trillion synapses. GPT-5 has 350 billion parameters or something like that. It’s not an apples-to-apples comparison, but even a rat brain is tiny and can do things we haven’t figured out how to solve. Nobody in the world has figured out how the brain of an animal, without using language, is able to perceive, think, and act in the world in remarkably sophisticated ways. That’s what I’ve been studying over the past few months.
People have been thinking about this for well over a century, trying to figure out what our mind does, what our eyes do, what exactly is going on. You have a baby who can do nothing, and then they eventually recognize faces, they have object permanence, they understand gravity, they understand physics to the extent that they can throw things and understand where they’re going to go. There’s something extremely sophisticated going on.
Hermann von Helmholtz, who invented the ophthalmoscope — a machine for looking into the eyeball — said: “The sensations of the senses are tokens for our consciousness. It is left to our intelligence to learn how to comprehend their meaning.” This was 1867. Already someone was proposing that we’re not actually seeing the world. We are using our eyes to see something, and then our brains are turning that into the thing we perceive. Probably not new to anyone, but really important.
Fast forward to 1963 and experiments with kittens. Two kittens in a carousel: one is walking and able to use its body, and the other is bundled up and not using its body. As the two kittens walk around, they both see the same thing. But one of them is not using its body to correlate what it’s seeing. What happens is that the visual system of that kitten doesn’t develop. It’s functionally blind. Its eyes work, but they’re not connected to the rest of its brain. So the eyes aren’t the whole story — they’re connected to the body, and there’s a process which requires the body for the eyes to actually work.
Generally, the idea is that the brain is not recording, the brain is not doing a 3D reconstruction of the world all the time, and then building up a model we use to interact with the world. That is almost certainly not what’s happening at all. There’s very much a top-down — you can call it a hallucination, you can call it an assumption — a top-down idea of what we think is happening all the time, and we’re only using our senses to do some error correction. We’re acting as part of this error correction in order to understand what’s really going on. Action and perception sit right next to each other.
We get a lot of glimpses into what’s happening from optical illusions, from people with brain damage, and from people with conditions like schizophrenia or autism. A classic optical illusion: square A and square B look different. They are not different. They are the same. It doesn’t matter how much you understand that they’re the same color — they look like different colors. Your brain is giving you this top-down, fairly reasonable assumption — almost always correct — that on a checkerboard, square A and square B are different colors. You can internalize the fact deeply, and it really doesn’t matter how many times you see it. There are other illusions, like the blue dress/gold dress, the hollow face illusion. They give you the general idea that your brain is making heavy predictions all the time, and even when your senses tell you otherwise, you still see it that way.
So the classical idea — there’s the world, you sense the world, your brain assembles it — is probably not right. Instead, you’ve got a cycle. Your brain generates predictions. That is the main job of your brain. This is much simpler and much less resource-intensive than bottom-up, generating an entire picture of the world. The world delivers signals that confirm or deny what’s going on, and these prediction errors go back to the brain to give you the hopefully true picture of what’s happening.
The whole thing is hierarchical. The system learns its own priors — basic ideas of what’s happening — and precision weighting decides which errors matter most. Sometimes you can just discard most of the information. In fact, most of the time, most people are doing that. If you’re driving home on a familiar route, you’re not really paying much attention. If you’re driving on a foggy night on an unfamiliar road, you’re exhausted afterward because you have to pay extreme attention and actually reconstruct what’s going on.
We’ve got this loop that uses the same neurons for perception, attention, and action. When you imagine something, you’re often using the same neurons as when you’re doing it for real. Action is not a response to the world. Action is how you select the next sample of what you’re doing, and therefore what you’re thinking about.
You can’t really remove the body, the self, from this mechanism. If you try to build a brain in a jar, it’s not going to work — just like with the kitten. If you take away the ability of the kitten to use its body, its eyes don’t work. The body is very central to the way we’re able to use our minds. The world itself is also part of that. You can’t take away someone’s body or the world outside and expect the mind to develop. They’re very much connected to each other, which is generally not how robotics and AI are happening right now. We’re trying to take the mind and put it into a world without having them develop together.
So when we’re talking about solving world models or solving embodied cognition, I’m talking about how we can build a machine that can act and perceive in the world the way we do. You probably need some training data. Take YouTube — a billion hours uploaded every hour or something like that. It’s a staggering amount of content, but it doesn’t have a lot of stuff that could be useful for training these things. It doesn’t have any geometric ground truth. It’s not multimodal. It’s not time-synchronized. Sometimes it’s captured through the scale, sometimes it’s not. Sometimes it’s not moving the way the body moves. It’s not really coupled to a person in any way.
So what we’re thinking about here is how to build capture devices that come closer to this kind of data, which is going to be very important in the next step of solving this aspect of AI.
So who am I, and why am I talking about this? I went through my photo collection and found all the pictures of me over the years. You can see by the lack of gray hair how long I’ve been doing this. I’ve been building various types of cameras attached to bodies, creating human-scale data, for a very long time. The first picture is from 2008 — I showed this at a Panotools conference. A couple of you in the audience were there for that. I’m not an electrical engineer or a software engineer. I’m just someone with persistence in building a human-body-scale camera system that can capture the world.
That turned into a project called Sphericam in 2015–2016. Some of you are familiar with it — a small VR camera. That morphed into Mosaic, who we are today. We’ve shipped cameras to 50 countries worldwide. Our cameras work really well, and we have a lot of very happy customers doing geospatial and 3D data collection for many different applications.
About the training data, and how this is going to come about: one of the co-founders of World Labs observed that robotics doesn’t have a large repository of training data. It’s not like LLMs. LLMs had all the words already there — trillions of words just there for the taking. That’s not the case for this.
So I’d like to show our next product. It’s shaped like a body, because the body is the best shape, and it’s great for capturing data at the human scale. This is not the same as the amazing Gaussian Splat datasets we’ve been looking at that were made with giant professional cameras. This is more — we can build a thousand of these and walk around and capture synchronized, multimodal, multisensor data that can produce the amazing Gaussian Splats I showed you earlier. But it’s with the idea that this could be part of building large datasets that can help solve the next chapter of AI.
Thank you very much.







