Anime, manga, and games, with a take · A Yukimedia publication

← all stories other 5 sources · 2d ago · · Updated

GPT-6 Astra Scores 67 on Coding Benchmark, Trails Fable 5.1

The results show GPT-6 Astra is a coding-focused release: it reaches top-tier agent scores at low per-task cost, while its general intelligence score stays flat against the previous generation.

Key Facts

  • GPT-6 Astra scored 67 on the Artificial Analysis Coding Agent Index, tying Fable 5 and trailing Fable 5.1's 70.
  • GPT-6 Astra's API pricing is $10 per million input tokens and $50 per million output tokens, 2.5 times GPT-5.6 Sol's rates.
  • GPT-6 Astra completed Portal without human help, using $571.18 in API tokens and under two hours of in-game time.
  • GPT-6 Astra was trained on more than 100,000 NVIDIA Grace Blackwell NVLink72 units, according to NVIDIA CEO Jensen Huang.
  • GPT-6 Astra's hallucination rate on the AA-Omniscience benchmark fell from 92 percent to 51 percent at maximum reasoning.

Reporting from 5 sources: Automaton, ASCII.jp, GIGAZINE, GameBusiness.jp, and 1 more.

GPT-6 Astra Scores 67 on Coding Benchmark, Trails Fable 5.1

OpenAI's GPT-6 Astra, released September 3, 2026, scored 67 on the Artificial Analysis Coding Agent Index, tying Fable 5 and trailing Fable 5.1's 70. The model's token efficiency improved sharply: at maximum reasoning on Codex, GPT-6 Astra consumed about a third of the tokens GPT-5.6 Sol used and a seventh of what Claude Opus 5 used. That efficiency carries into per-task cost, where GPT-6 Astra matches GPT-5.6 Sol's cost while scoring 2 points higher, and costs less than half of Fable 5 per task at the same 67 score. On the Artificial Analysis Intelligence Index, GPT-6 Astra scored 61, tied with GPT-5.6 Sol and Muse Spark 1.3 and 5 points below Fable 5.1's 66. Its API price rose 2.5 times to $10 per million input tokens and $50 per million output tokens, which offsets the token gains on general intelligence. The hallucination rate on the AA-Omniscience benchmark fell from 92 percent to 51 percent at maximum reasoning. Separately, GPT-6 Astra completed the puzzle FPS Portal without human help, at an API cost of $571.18.

The benchmark results paint GPT-6 Astra as a specialized tool rather than a general upgrade. Artificial Analysis found the model gained about 80 Elo points on AA-Briefcase, its long-horizon knowledge task, but lost about 80 Elo points on GDPval-AA v2, which measures economically valuable work across 44 occupations. It also slipped 2 to 3 points against GPT-5.6 Sol on banking support, SciCode, and long-context reasoning.

The model's 3D work drew attention separately: OpenAI staff posted examples of GPT-6 Astra building a walkable demo house in Blender and exporting it to Unreal Engine 5, and recreating San Francisco's Palace of Fine Arts from hundreds of reference photos. The Portal run, reported by GitHub user cozyblaze, used a modified SourcePauseTool to pause the game while the model reasoned, finishing in under two hours of game time at a cost of $571.18 in API tokens.

NVIDIA CEO Jensen Huang said on X that GPT-6 Astra was trained on more than 100,000 NVIDIA Grace Blackwell NVLink72 units and called it the arrival of AGI; OpenAI president Greg Brockman told Stratechery it was the first time the company trained on that scale. Sources differ on the Portal run's real-time length, with reports citing 23 hours 38 minutes and 23 hours 43 minutes.

Synthesized by Yomimono from the 5 cited sources below, including Japanese-language reporting where cited, then editorially reviewed before publishing.

Sources