NVDA 228.45 ▲1.80%GOOGL 342.48 ▲1.59%MSFT 510.12 ▲2.68%AMD 456.16 ▼0.20%INTC 91.67 ▲1.80%TSMC 417.01 ▲0.36%AMZN 258.90 ▲1.54%META 610.68 ▲3.01%AAPL 328.21 ▲1.00%PLTR 182.53 ▲7.71%
Markets at last close

OpenAI · Research

OpenAI, Anthropic and Google converge on AI research automation

·1 min read

OpenAI, Anthropic and Google are pursuing systems that can help build more advanced AI, while also identifying that capability as a potential safety boundary in their own frameworks. The central concern is machine learning R&D: AI agents that can perform real research engineering tasks, potentially accelerating the development of future models.

Google DeepMind’s model card for Gemini 2.5 Pro includes a safety evaluation on RE-Bench, a METR benchmark that compares AI agents with human experts on practical machine-learning research engineering work. The evaluation found that the best agent solutions scored between 50% and 125% of the best expert-written solutions.

Despite those results, the table marked the capability threshold as “CCL not reached,” meaning Google DeepMind concluded under its published criteria that no additional mitigations were required. The gap between strong benchmark performance and an untriggered safety threshold highlights a key tension in the current AI race: labs are measuring potentially consequential capabilities while continuing to move them closer to deployment.

Originally reported by pub.towardsai.netRead the source →
Related coverage
All OpenAI news →