Knowledge · Fundamentals

Physical AI in eight chapters.

A compact introduction for management, operations, and IT: how cognitive robots learn, what they can do today, and what matters for safety and evaluation.

  • 8 chapters
  • approx. 7 min read

What is Physical AI?

Artificial intelligence that does not just process text or images, but acts in the physical world.

Physical AI -- often called embodied AI -- refers to AI systems with a body: robots that perceive their surroundings through cameras and sensors, derive decisions from that, and act through motors and grippers. We call them cognitive robots.

The key difference from classic robotics: behavior is no longer programmed line by line, but largely learned from data. That lets robots handle variation -- parts that sit slightly differently every time, changing tasks, and environments that were never built for robots.

Technically, the same loop runs continuously: perceive, decide, act -- and check the result to see whether the action worked. The faster and more robust that loop, the more useful the robot.

  • Perceive

    Cameras, depth sensors, LiDAR, and force sensors build a picture of the surroundings.

  • Decide

    An AI model derives the next movement from perception and the task at hand.

  • Act

    Motors, joints, and grippers carry out the decision -- often many times per second.

Classic automation or cognitive robot?

Not either-or: both approaches have their strengths -- the task decides.

Classic industrial robots are unbeatable when processes are high-volume, always identical, and well structured: welding, painting, palletizing in fixed patterns. They are precise, fast, and proven over decades -- but every change costs programming and retooling effort.

Cognitive robots play to their strengths where variation is part of daily work: changing parts, small batch sizes, unstructured environments, and tasks that only people could do so far. In return, they are usually slower and less predictable than classic systems today.

In practice, the two complement each other. The right question is not “Which robot is better?” but “How much variation is in this task -- and what does it cost us today?”

  • Classic: programmed

    Fixed movements, high precision and speed, high effort for every change.

  • Cognitive: learned

    Learned from demonstrations and data, tolerant of variation, still slower and less predictable today.

How robots learn

Show, practice, transfer: the main ways cognitive robots learn.

Much of the practical progress in recent years comes from imitation learning: a person guides the robot through a task via teleoperation while the robot records what it sees and how it moves. From many such demonstrations, a model -- the so-called policy -- learns to perform the task on its own.

In reinforcement learning, the robot tries things itself and is rewarded for success. Because that would be slow and risky in the real world, it mostly happens in simulation. The art then lies in sim-to-real transfer: what was learned must also work with real friction, real lighting, and real sensor noise. So far this approach has been especially successful for locomotion and balance.

How many demonstrations a task needs depends heavily on the task -- and on whether a pre-trained model serves as a starting point. Starting from a foundation model, a narrowly defined task often needs far fewer examples than training from scratch.

  • Imitation learning

    Learning from human demonstrations, usually recorded via teleoperation.

  • Reinforcement learning

    Learning by trial and reward, mostly in simulation.

  • Fine-tuning

    A pre-trained model is adapted to a specific task with a small amount of targeted data.

Foundation models for robots

What large language models are for text, vision-language-action models are becoming for robots.

Vision-language-action models (VLAs) combine three things: they see (camera images), understand language (the instruction, such as “Put the blue part in the box”), and output actions (joint movements or gripper commands). They are pre-trained on large datasets from many robots and tasks, then fine-tuned for the specific deployment.

Alongside closed models, there is a growing number of open models whose weights are freely available -- such as OpenVLA, π0, or SmolVLA. For businesses, that means models can be inspected, adapted, and run on your own hardware. That is exactly what we rely on.

A second line of research is world models: models that predict how the environment changes as a result of an action. They are meant to help robots plan ahead and cope with unfamiliar situations.

Data: the real bottleneck

Language models have the internet. Robots have no comparable source of data.

Robot data has to be generated: through teleoperation, in simulation, or during operation. Open collections such as Open X-Embodiment or the datasets of the LeRobot community help with pre-training -- but your own task almost always needs your own data from your own environment.

Quality matters more than quantity: clean, consistent demonstrations, well-covered variations, and documented failure cases. Operational data is just as important -- success rates, interventions, aborted runs -- because only that shows where a model needs retraining.

Because this data shows your processes, your products, and often your people, it is a competitive asset. Where it is stored and processed, and who may use it, should be settled before the first pilot.

Hardware and form factors

The best AI is of little use if the body does not fit the task.

Cognitive robots come in many form factors. Each has its own strengths and limits -- and not every task needs legs.

Beyond the form factor, what matters are the degrees of freedom, payload, sensing -- from depth cameras and LiDAR to force and tactile sensors -- and the compute on board. For mobile systems, battery life comes on top, and in continuous operation it is often the limiting factor.

Datasheets show best-case values under ideal conditions. Only a test shows what a system delivers in your environment.

  • Arms & cobots

    Stationary, precise, proven -- ideal for fixed workstations.

  • Mobile manipulators

    An arm on a wheeled base -- for tasks in changing locations.

  • Quadrupeds

    Robust and rough-terrain capable -- for inspection and transport on uneven ground.

  • Bipeds

    Human-like form for environments built for people. The most demanding technically.

Safety, security, and sovereignty

A robot that works with people and sits on your network needs two kinds of protection -- and clear control.

Functional safety protects people in the work area. It starts with a risk assessment; for industrial and collaborative robots, standards such as ISO 10218 and ISO/TS 15066 set the framework, for example for force and speed limits. Learning systems add a new question: what happens when the model is wrong? The answer is protective functions that act independently of the AI.

Cybersecurity protects the robot and your network. A cognitive robot is a networked computer with cameras and microphones -- like any other device, it belongs in an OT security concept: network segmentation, access rights, reviewed updates, audit logging.

Sovereignty, finally, means you stay in control of data, models, and updates. The regulatory framework is growing, too: the new EU Machinery Regulation applies from January 2027, alongside the EU AI Act and the Cyber Resilience Act.

  • Safety

    Protects people: risk assessment, force limits, emergency stop, fallback.

  • Security

    Protects systems: segmentation, access control, signed updates.

  • Sovereignty

    Keeps control: data, models, and operations in your hands.

Evaluate and adopt

A convincing demo is not a business case. Measurements from your own environment are.

Assess cognitive robots by the same standards as any other process: How often does the task succeed? How long does it take? How often does a person have to step in? And what does it all cost over its lifetime -- including integration, maintenance, and data preparation?

A step-by-step approach has proven itself: choose a task, measure the baseline, run a clearly scoped pilot with success criteria defined up front -- and only then decide on a rollout. That turns curiosity into a sound decision.

  • Success rate

    Share of runs that succeed without human help.

  • Cycle time

    Time per run -- compared with today’s process.

  • Interventions per hour

    How often a person has to help or correct.

  • Total cost

    Purchase, integration, energy, maintenance, and data over the lifetime.

Got the fundamentals? Then let’s get specific.

We apply this knowledge to your processes: in a workshop with your team or in a vendor-neutral assessment.