Nvidia outlines synthetic data workflows for vision AI agents
Vision AI agents are moving closer to cameras, machines and sensors as enterprises look to turn video into operational insight across factories, cities, warehouses and transportation systems. Gartner projects that more than two-thirds of enterprise-managed data will be created and processed outside the data center or cloud by 2028, while over two-thirds of all enterprises globally will deploy edge AI by 2029, up from 10% in 2025. Yet as much as 90% of existing edge data goes unprocessed.
Nvidia is framing OpenUSD-based Omniverse, Metropolis agent skills and blueprints, TAO skills, and video search and summarization workflows as a lifecycle toolkit for building vision AI agents. The approach targets common bottlenecks including data gaps, limited fine-tuning expertise and the complexity of assembling video pipelines, alerts, reporting, search and system integrations.
Manufacturing examples include Roboflow’s integration of the Nvidia Defect Image Generation skill and Nvidia Cosmos world foundation models to create synthetic defect images when real examples are scarce. In a Corning benchmark, a model trained on just eight real defect images augmented with synthetic data reached an average precision of 95% and perfect recall on the most challenging defect class.
Other deployments highlight operational use cases. Linker Vision reduced development effort by 85% using the VSS blueprint in Kaohsiung and reduced incident response times by up to 80%. At Foxconn, DeepHow’s Live Standard Operating Procedure Verification agent has been used on NVIDIA GB300 server production lines to improve first-pass yield by 3% and achieve 99% task-level accuracy in micro-action understanding of critical SOP steps.