Skild AI uses NVIDIA tools to teach robots from video
Skild AI has launched S1, a robot foundation model built to learn previously unseen, long-horizon tasks from a video prompt rather than task-specific retraining. The model uses in-context learning to interpret demonstrated intent, objects and sequences, then translate them into robot actions without updating its weights.
The system is designed for changing manufacturing, warehouse and production environments where layouts, products and processes shift. S1 can perform unfamiliar tasks lasting up to 10 minutes, including plant potting, pancake making, pour-over coffee brewing and kit assembly. In one plant-potting test, Skild AI moved from demonstration recording to autonomous hardware execution in just 11 minutes.
Skild reported that S1 succeeded about 66% of the time at each step on new multistep tasks, compared with 9% for a similar AI system. The company estimates that showing one short video example can be as useful as providing roughly 380 hands-on training examples, which could take a person 50-100 hours to collect manually.
The work has moved onto factory floors through a Skild, NVIDIA and Foxconn deployment of Skild Brain on dual-arm manipulators for high-precision assembly of NVIDIA Blackwell systems. In a demonstrated workflow, a robot installs a busbar and limit block, fastens 16 screws and adapts to disturbances, while NVIDIA Cosmos, Omniverse, Isaac Sim, Isaac Lab, Nsight and TensorRT support the broader pipeline from data and simulation to inference.