Anime, manga, and games, with a take · A Yukimedia publication

← all stories other 1 sources · 2h ago ·

OpenAI Uses GPT-5.6 to Improve Its Own Inference Efficiency

OpenAI's use of GPT-5.6 to improve itself suggests that AI development may enter a self-accelerating loop, a possibility Anthropic has warned about.

Key Facts

  • OpenAI released the GPT-5.6 series on July 9, 2026, with three variants: GPT-5.6 Sol, GPT-5.6 Terra, and GPT-5.6 Luna.
  • GPT-5.6 Sol reduced end-to-end service costs by 20% by optimizing kernels for Triton and Gluon.
  • Autonomous improvements to speculative decoding increased token generation efficiency by 15%.
  • OpenAI reduced GPT-5.6 Luna's price by 80% and GPT-5.6 Terra's price by 20%.
  • Enabling two API settings (reasoning retention and compression) tripled GPT-5.6 Sol's ARC-AGI-3 score to 38.3%, compared to 13.3% with the official harness.

Reporting from 1 source: GIGAZINE.

OpenAI Uses GPT-5.6 to Improve Its Own Inference Efficiency

OpenAI has revealed that it used GPT-5.6 itself to reduce computational costs and improve GPU usage efficiency for the GPT-5.6 series, which was released on July 9, 2026. The series includes three variants: the high-performance GPT-5.6 Sol, the balanced GPT-5.6 Terra, and the low-cost GPT-5.6 Luna. According to OpenAI, GPT-5.6 Sol achieved high cost efficiency through self-improvement. The inference engine's kernel was improved with GPT-5.6 Sol, which learned writing methods that contribute to the efficiency of the programming languages Triton and Gluon, reducing end-to-end service costs by 20%. Speculative decoding, an inference acceleration technique using a small draft model, was also improved autonomously, boosting token generation efficiency by 15%. OpenAI also streamlined the agent system's harness, implementing changes such as calling MCP servers and skills only when needed, limiting tool output tokens to 10,000 by default, and fixing the order of tool calls, which improved cache utilization. OpenAI commented that since GPT-5.6 contributed significantly to improvements in inference and harness, the pace of optimization will likely accelerate in the future.

  • GPT-5.6 Sol is promoted as a model that outperforms Claude Fable 5 at half the cost.
  • OpenAI also autonomously improved KV cache handling and the system that allocates computational processing to GPUs.
  • OpenAI reduced prices for GPT-5.6 Terra by 20% and GPT-5.6 Luna by 80%, making Luna cheaper per task than Claude Sonnet 5, Gemini 3.6 Flash, DeepSeek V4 Pro, and GLM-5.2.
  • A new API "fast mode" doubles the price of GPT-5.6 Sol but increases processing speed by roughly 2.5 times.
  • On the ARC-AGI-3 benchmark, GPT-5.6 Sol scored 7.8% initially. Enabling two API settings used in ChatGPT and Codex, reasoning retention and compression, tripled the score on the public task set and cut output tokens to one-sixth.
  • With the official harness, GPT-5.6 Sol scored 13.3% on ARC-AGI-3; with the two improvements, the score rose to 38.3%. Human testers average 48%.
  • OpenAI said the experiments show that evaluations measure not just models in isolation but also "various less visible elements such as API settings, harness design, and prompt display."

Synthesized by Yomimono from the 1 cited source below, including Japanese-language reporting where cited, then editorially reviewed before publishing.

Sources