OpenAI GPT-5 Achieves Human-Level Performance on Graduate-Level Science Benchmarks

OpenAI announced that its latest large language model GPT-5 has achieved human-level performance on the Graduate-Level Google-Proof QA benchmark, scoring 87.3 percent accuracy on questions designed to test expert scientific reasoning.
The model, which uses a mixture-of-experts architecture with an estimated 1.8 trillion parameters, outperformed the previous state-of-the-art by 12 percentage points. GPT-5 was trained on a curated dataset combining scientific literature, peer-reviewed research papers, and synthetic reasoning chains generated by earlier model iterations.
"This is not just an incremental improvement," said Mark Chen, OpenAI VP of Research, during a press briefing. "GPT-5 demonstrates genuine scientific reasoning, not pattern matching."
## Benchmark Performance and Capabilities
On the GPQA benchmark, which contains 448 expert-level questions across biology, chemistry, and physics, GPT-5 scored 87.3 percent accuracy. Domain experts with PhDs in the relevant fields scored an average of 81.2 percent on the same questions, meaning the model outperformed human experts by 6.1 percentage points across all three scientific disciplines.
The model also scored 94.1 percent on the MATH benchmark, 91.7 percent on the HumanEval coding benchmark, and 88.9 percent on the ARC-Challenge science reasoning test. OpenAI published the full results in a technical report on August 30.
"The mixture-of-experts approach allows GPT-5 to activate different parameter subsets depending on the domain, effectively creating specialized sub-models within a single unified architecture," said Noam Brown, a research scientist at OpenAI. "This is more efficient than scaling a monolithic model to achieve similar results."
## Industry and Academic Reaction
Demis Hassabis, CEO of Google DeepMind, congratulated OpenAI on the achievement while noting that Google Gemini Ultra 2 achieved comparable scores on several internal benchmarks. "The race to build capable AI systems is accelerating, and that competition drives progress for everyone in the field," Hassabis said.
Academic researchers expressed both excitement and caution. Dr. Melanie Mitchell, a complexity scientist at the Santa Fe Institute, noted that benchmark performance does not necessarily translate to general scientific capability. "Scoring well on specific question sets is impressive, but scientific discovery requires experimental design, hypothesis generation, and iteration in ways that current benchmarks do not test adequately," Mitchell said.
## Safety and Governance Implications
The release intensifies debate about governance of advanced AI systems. The European Union AI Act, which entered enforcement in August 2025, classifies models exceeding certain capability thresholds as high-risk and imposes transparency and testing requirements.
OpenAI voluntarily published a system card detailing GPT-5 capabilities, limitations, and safety evaluations. The company implemented new safeguards including refusal mechanisms for weapons-related queries and enhanced monitoring for biological threat information.
"Capability without safety is a recipe for harm," said Dario Amodei, CEO of Anthropic. "Every advance in model capability must be matched by advances in alignment research and safety evaluation."
Sam Altman said OpenAI plans to release the model API to developers on September 15, with consumer access through ChatGPT following in October. Enterprise customers will receive early access starting September 8. Pricing for the API will start at $15 per million input tokens and $60 per million output tokens.
Discussion
Recommended for you
More technology
Artificial Intelligence
OpenAI Anthropic and 100 Companies Warn of AI Cybersecurity Apocalypse
9/2/2026
Artificial Intelligence
Google DeepMind Achieves Breakthrough in Protein-Drug Interaction Prediction Reducing Drug Discovery Time
9/2/2026
Artificial Intelligence
Google DeepMind Achieves Breakthrough in Protein-Drug Interaction Prediction Reducing Drug Discovery Time
9/2/2026
Software