The ordinary
is extraordinary.
A simple handover.
A remarkable amount of intelligence.
Someone hands you a tool. You look at it, understand what they want, adjust your grip and take its weight. The whole exchange might take a second.
An exchange that lasts a moment. A world of decisions inside it.
You probably do not describe each finger movement to yourself. Yet you can explain the goal, anticipate what happens if you let go, and change your approach if the tool is heavier than expected. You draw on several abilities at once: perception, language, memory, prediction and practiced movement.
That effortless exchange is what makes intelligence feel alive. Understanding the request is only the beginning. The movement, the weight and the response all belong to the same moment. Nucleus AI begins with the ambition to bring those abilities together in the places where people work.
We want robots that can meet the world with that same flexibility. Robots that can take a little more weight off someone’s hands, carry a task through, and learn from the moments that ask something new of them.
Understanding.
Anticipation.
Action.
Three ideas worth understanding.
A robot needs more than a description of the scene. It needs a way to choose what to do. Researchers call that a policy: a learned rule that turns observations and a goal into actions. Different model designs give that policy different strengths.
Connect what you see with what you are asked to do.
A vision-language-action model, or VLA, connects camera observations and language instructions to robot actions. “Take the part on the left” becomes a problem of identifying the right object and acting on it, rather than matching a single fixed position.
The same instruction can mean something different in every scene. A part has moved. A hand is waiting. A space has opened up. Connecting language with what the robot actually sees gives an instruction its meaning in that moment.
Understanding what is being offered.
Finding the movement that follows.
Anticipate how the world could change.
A world model learns relationships between states of an environment. Instead of only asking “What should I do?”, it can help ask “What might happen next?” The prediction might be a future image or an internal representation that does not look like a picture.
That opens up a powerful possibility: exploring a movement before committing to it. A model can learn from how scenes unfold, anticipate possible outcomes and help evaluate a path forward. Experience becomes something a system can reason with, as well as something it has lived through.
The real world gets the final say. A predicted grasp may look convincing and still miss the object. Contact, geometry and movement have to hold up when the robot reaches out. That is why we connect anticipation to fresh observations and checks in the environment.
Every reach changes what the robot sees.
Turn an intention into a controlled movement.
A diffusion policy generates a sequence of actions by progressively refining a noisy initial proposal, conditioned on what the robot observes. Think of a rough movement becoming more specific, rather than a sentence being translated word by word into motor commands.
With receding-horizon control, the robot executes part of that sequence, observes again and plans the next part. The movement stays connected to what is happening. If an object shifts during a reach, the next observation gives the policy a chance to respond.
These capabilities can live inside the same system. A VLA can use diffusion to generate its actions, while a world model helps anticipate what comes next. Understanding, anticipation and movement become parts of one continuous exchange with the world.
Intelligence meets the world as it is.
Our approach.
Orchestrate the capability
the situation calls for.
Nucleus brings post-trained VLAs, world-model approaches and diffusion policies into an architecture organized around work.
Post-training means adapting an already trained model with additional experience and feedback. It gives us a way to connect broad capabilities to the particular objects, robot bodies and workflows of an industrial setting.
We call the organizing layer our Autonomy Compiler. The idea is to turn recorded work into callable skills: bounded capabilities with a purpose, a context for use and a way to check the result. Some steps can be expressed directly in code. Others need a learned action model to handle variation.
An agent and skill router connect the task to those capabilities. A skill registry keeps their versions organized. When the situation falls outside what the system can handle, the architecture includes a route to a human supervisor. That intervention also creates information for the next training cycle.
This is why orchestration sits at the heart of Nucleus AI. A task can call for careful reasoning in one moment and a practiced movement in the next. Bringing the right capability into that moment is how we aim to carry an intention all the way through to useful work.
Capture trajectories. Form callable skills. Route the task. Verify the result. Bring difficult moments back into learning.
Read the architecture in words +
Paid factory work, robot sensors and onboarding instructions feed trajectory capture. Synchronized trajectories are segmented and verified in the Autonomy Compiler, with human corrections feeding intervention learning. The resulting skills are selected by an agent and skill router, executed as code or action models, and checked in the real world. A registry tracks versions; the supervisor path handles situations that need human help.
Experience becomes
useful through context.
The most useful moment
may be the one that went wrong.
A successful demonstration shows a path to the goal. A correction can show where that path breaks. Perhaps the object slipped, the viewpoint changed, or the robot reached the right place with the wrong orientation.
For that event to become useful data, its pieces have to stay together: what the robot observed, the state of its body, the action it attempted, the correction and the outcome. A video alone cannot tell the whole story.
Our data pipeline is built around those connected episodes. The objective is to extract reusable experience from real work: which parts are repeatable, which vary, and which conditions require more support. Inside those hours are the moments that matter: a hesitation, an unexpected load, a better way to finish. Finding them is how experience starts to become a skill.
A movement becomes experience we can return to.
Experience can teach a system in different ways. Imitation learning learns from demonstrated actions. Self-supervised learning constructs learning targets from the data itself, such as predicting a missing observation. Reinforcement learning uses feedback about outcomes to improve the choice of actions. They address different questions and can be combined.
A demonstration can give the robot a starting point. Its own attempts reveal what holds up in practice. A person’s correction can show a better way through a difficult moment. Together, those experiences give post-training something richer to learn from.
For Nucleus, the loop begins and ends on the factory floor. We collect experience, find the lessons, train and evaluate, then return to work. We want the hard-won knowledge of one shift to help the next one go better.
A skill moves through the factory.
The next view brings the next decision.
Autonomy,
earned in practice.
A skill is only useful
if you know where it works.
The goal is for more work to be completed reliably, with less human intervention, across a broader range of conditions.
A factory needs to know that the work will get done. A familiar part under different lighting, a crowded aisle, a handover that comes a little late: ordinary changes become the real test. Reliability means carrying the task through those moments, and making it clear when a person’s help is needed.
Our intended path starts with guided work, turns recurring actions into verified skills, and expands their scope as evidence accumulates. Each stage asks a harder question: can it repeat the action, recover from a disturbance, and transfer to a new situation?
We look at whether a task reaches its intended outcome over repeated attempts, where people have to intervene, and how the robot recovers when something changes. Then we ask whether that skill carries across to another object, another layout or another day.
Built for the world around us.
Intelligence has to move with the body.
Each dependable skill gives people a little more room to focus on the work that needs their judgment. That is the progress we care about. More responsibility a robot can carry. More trust earned through experience.
A system
that can keep learning.
Many ways to learn.
One world to act in.
The physical world asks for different kinds of intelligence in the same moment. A robot may need language to interpret the goal, prediction to anticipate a consequence, a learned policy to handle contact, and a familiar procedure to finish the job.
We are building Nucleus to bring those abilities into the working world. Every shift should leave more behind than finished work. It should leave experience we can build on, skills we can improve, and a clearer path toward robots that people can depend on. That is the future we are working toward.
We don’t think in a single way.
Neither should AI.
