Monday, September 7, 2026

News

OpenAI Declares 'Era of AGI' With GPT-6 Astra, Researchers Push Back

ModelsPatryk Raba

OpenAI president Greg Brockman declared the arrival of the "era of AGI" at the launch of the GPT-6 Astra model, while Nvidia CEO Jensen Huang said general artificial intelligence has already arrived. The creators of the ARC-AGI benchmark and other researchers say it isn't there yet.

Contents
  1. What the Numbers Show
  2. Nvidia's Voice, and Its Cost
  3. Skeptics Push Back
  4. No Shared Definition

OpenAI unveiled its new GPT-6 Astra model on September 3, with company president Greg Brockman closing the press conference with the words "welcome to the era of AGI." A day later, Nvidia chief Jensen Huang went a step further, declaring on X that artificial general intelligence had already arrived. The creator of the ARC-AGI benchmark, which both companies cited, gave a short reply: we're not claiming that.

Brockman, OpenAI's president and co-founder, answered directly when reporters asked him at the briefing whether he believes the company has already achieved artificial general intelligence. He said that if someone looks back in a few years and asks when AGI was truly created, in his view this would be the moment, and probably this very model.

Personally, I think we're there - Greg Brockman, President of OpenAI

What the Numbers Show

OpenAI touts a string of results: 98 percent on the FrontierMath Tier 4 test, 99.9 percent on ARC-AGI-3 using its own customized test suite, and 100 percent on ExploitBench, a test of offensive cybersecurity capabilities. Those same numbers, set against the independent benchmark creator's standard score of 62-66 percent, immediately show how much the measurement method changes the result.

ARC-AGI's creator, François Chollet, admitted that progress is running roughly twice as fast as he expected, and moved his forecast for the arrival of true AGI to earlier than his previously assumed year of 2030. He highlighted what he called new model behavior: Astra builds symbolic models of unfamiliar environments on the fly and invents its own shorthand notation for itself, something that previously required extensive external tools.

Nvidia's Voice, and Its Cost

Jensen Huang, Nvidia's CEO, wrote on X on September 6 that artificial general intelligence had already arrived thanks to GPT-6 Astra, stressing that the journey from ChatGPT through the o1 model to the current model took just four years. The statement drew a cool reception from some researchers, who pointed out that Nvidia has a direct financial interest in the success of large training projects, since its chips power these models.

Skeptics Push Back

Technology critic Gary Marcus wrote that the ARC-AGI success is impressive but, despite the test's name, is not proof of AGI, and he expects the model to struggle with open-ended, real-world tasks. The team behind the ARC Prize itself stressed that it is not claiming Astra has achieved AGI - the 62.7 percent score is a serious leap, but not saturation of the scale, since capable humans score above 90 percent on the test.

GPT-6 Astra represents a step-change in model capability on interactive reasoning tasks. It reaches 66 percent on ARC-AGI-3 with our standard suite, and nearly 100 percent with the continuous-conversation suite and custom compaction, at a cost of about $360 per game - François Chollet, creator of ARC-AGI

No Shared Definition

Part of the confusion stems from the fact that the term AGI still lacks a single, widely accepted technical definition. OpenAI describes AGI as AI systems that are generally smarter than humans, and Brockman even suggested at the briefing that users should define the term for themselves. That flexibility lets one company announce a breakthrough while independent researchers use the very same numbers to question it.

For Polish companies and institutions following the AI race, what matters is the practical takeaway rather than the label itself. Regardless of whether Astra meets anyone's definition of AGI, the model genuinely pushes the boundaries of multi-step reasoning and planning tasks, capabilities that will sooner or later reach tools used by Polish companies through OpenAI's API or partner products.

The dispute also shows how far marketing now outpaces scientific verification in the AI industry. Executives and investors make their declarations faster than independent researchers can check results against new tasks the model hasn't seen before, and it is precisely those tests, not official company statements, that in practice settle disputes over systems' real capabilities.

Share: