Agents' Last Exam
OpenAI launch table. This benchmark evaluates long-horizon browser and desktop work.
OpenAI's new generation reaches 72.6% on OSWorld 2.0, 92.7% on ScreenSpot-Pro, and 59.3% on Agents' Last Exam. In SPIRITT, run Astra inside the browser, terminal, files, integrations, and memory it needs to turn computer-use scores into completed work.
The official release and independent benchmark context, followed by eight hands-on builds spanning 3D landmarks and worlds, playable games, real estate, education, and native apps.
GPT-6 Astra is the most intelligent and aligned model in the world, and sets a new state of the art for computer use, browsing, software engineering, cybersecurity, science, and professional work. https://openai.com/index/gpt-6-astra/
— OpenAI (@OpenAI) September 3, 2026
GPT-6 Astra is rolling out today to a limited set of organizations and over the coming days will become available to all ChatGPT Plus, Pro, Business, and Enterprise users, as well as through the OpenAI API and AWS. Be ready to experience Astra at its best. Get the ChatGPT desktop app.
— OpenAI (@OpenAI) September 3, 2026
GPT-6 Astra is here. We hope it will begin to enable a new generation of entrepreneurship, scientific discovery, and building. We believe it is the best model in the world for computer use, professional work, science, coding, cybersecurity, and more. It took us some extra time to ensure that we could meet the safety and alignment standards required for this capability level, but we think you’ll find it worth the wait. It scores 98% on FrontierMath Tier 4, 99.9% on ARC-AGI 3, and 100% on ExploitBench.
— Sam Altman (@sama) September 3, 2026
GPT-6 Astra makes significant gains in the Artificial Analysis Coding Agent Index, scoring equal to Fable 5 at lower cost. In the Intelligence Index, it uses fewer tokens than GPT-5.6 Sol for similar performance, but this is outweighed by higher prices Pricing is 2.5x GPT-5.6 Sol’s current prices across the board, up from $4/$20 to $10/$50 per million input/output tokens, with the same 90% discount for cache reads and 25% premium for cache writes. We see distinct stories across our two flagship Indices. In the Artificial Analysis Coding Agent Index, GPT-6 Astra equals Fable 5 at less than half the cost, driven by significant token efficiency gains. In the Artificial Analysis Intelligence Index, GPT-6 Astra is more token efficient than its predecessor for similar performance, but this is offset by the price increase. Artificial Analysis Coding Agent Index - key takeaways: ➤ Rivals top models: In Codex, GPT-6 Astra scores 67 in the Index - approximately equal to Claude Opus 5 and Fable 5 in Claude Code, and Muse Spark 1.3 in Muse Code. Fable 5.1 in Claude Code leads the Index with a score of 70. ➤ 70% more token efficient than GPT-5.6 Sol: GPT-6 Astra sees a substantial improvement in token efficiency, using one third of the tokens compared to GPT-5.6 Sol (max) in the Codex harness, and one fifth of the tokens of Claude Opus 5 (xhigh). Various effort levels of the model occupy the Pareto frontier of token efficiency. ➤ Leads Coding Agent Index cost efficiency frontier: At max effort, GPT-6 Astra costs about the same as GPT-5.6 Sol (max) while scoring 2 points higher on the Index. Per task, the model is less than half the cost of Claude Fable 5, for the same score. Artificial Analysis Intelligence Index - key takeaways: ➤ Sits beside GPT-5.6 Sol in Intelligence: GPT-6 Astra scores equal to GPT-5.6 Sol in the Index at 61. This is 5 points lower than Claude Fable 5.1 (max with fallback). The model also trails Meta’s newly released Muse Spark 1.3 (max). ➤ ~10% fewer output tokens, offset by price increase: GPT-6 Astra defines a new Pareto frontier for Intelligence Index vs Output Tokens per Task - with a ~10% reduction in token use at max effort compared to GPT-5.6 Sol. However, due to the 2.5x increase in price, the model is 75% more expensive per task than its predecessor at max effort. ➤ Hallucinates half as much as GPT-5.6 Sol: GPT-6 Astra sees a large jump in AA-Omniscience, our knowledge and hallucination benchmark. This is driven by a significant decrease in hallucination rate from 92% to 51% at max effort. Unlike some models, this improvement does not come at the cost of accuracy - Astra increased accuracy by 4 points at the same time. ➤ ~80 point gain in AA-Briefcase Elo: GPT-6 Astra improves ~80 points in AA-Briefcase, our frontier long-horizon knowledge work evaluation. Models are tested on multi-week projects, with many linked tasks and thousands of source files. Astra sees a significant increase in both rubric scores and Analytical Quality Elo in AA-Briefcase compared to its predecessor. In the other direction, we observe a reduction in Presentation Quality Elo, where GPT-5.6 Sol (max) still leads all models. ➤ Mixed progress on other evaluations: The model sees a 6 point gain in Humanity’s Last Exam, a long-standing evaluation with emphasis on mathematics, science, and humanities. This is offset by a drop of ~80 Elo points in GDPval-AA v2 - a benchmark we adapted from OpenAI’s dataset measuring economically valuable tasks across 44 occupations. We also observe 2-3 point regressions on other evaluations across a mix of capabilities, including reductions in τ³-Banking (customer support), SciCode (Python problems in a scientific domain), and AA-LCR (long context reasoning over large documents). Congratulations @OpenAI and @sama on the launch!
— Artificial Analysis (@ArtificialAnlys) September 3, 2026
GPT-6 Astra recreated the Palace of Fine arts in Blender. This is favorite building in San Francisco because it was built for the World's Fair in 1914 where the steam locomotive and telephone were showed off. It was a time when technology gave us all a deep sense of optimism for the future. I feel like some of that sense has been lost since, but I hope models like astra can help restore it, and push us towards a more hopeful future.
— Sharif Shameem (@sharifshameem) September 3, 2026
GPT-6 Astra built this Manhattan world in Unreal Engine over the course of a week. It was literally able to go street by street to make each one perfect.
— Matt Shumer (@mattshumer_) September 3, 2026
GPT-6 Astra is world class at Blender and 3 dimensional reasoning. This was a 1-shot game it created, all running in browser.
— Theo - t3.gg (@theo) September 3, 2026
GPT-6 Astra is a beast. Give it a Zillow listing. It can 3D model the house based on the listing photos and create a cool promotion video. The video was just created in one shot and there are still some wrong details. But I am sure it will be much better if I ask it to polish further. OpenAI is so back.
— Yunfan Ye (@realYunfanYe) September 3, 2026
So how good really is GPT-6 Astra at 3D modeling? I took an old drawing of a steam train, gave it to Astra to reconstruct it in Blender. After few minutes it crafted 3,295 fully editable detailed objects with beautiful geometry. You can obviously tell it how detailed or simplified you want it. Go build your own transport tycoon 😎
— Tom Krcha (@tomkrcha) September 4, 2026
I’m in disbelief right now. 12 months ago I remember making flappy bird with AI in 2-3 prompts and being blown away. Now it can make a call of duty game. One that’s actually fun… I played for like 2 hours today and in between matches I would have GPT 6 make changes… I genuinely think gaming will be the mass onboarding to AI. It’s just too much fun at this fidelity.
— Riley Brown (@rileybrown) September 4, 2026
I am super excited to share this educational video that I had GPT-6 Astra create from a single prompt: “Create a 5 minute educational video about T cells,” cells that I had devoted most of my life studying! I also had Astra to use the Remotion plugin for the video generation. Astra also used Imagegen to create the visuals, it also suggested to use HeyGen for narration (have not used that before) and then assembled the entire educational video on its own one-shot, including what to say, how to explain it! The result is so incredibly well made that, narration, the animations are all so perfectly created that even with 35 years of experience studying T cells, I really don’t think I could have explained this any better myself! The quality is also better than anything I had seen before, Astra is far ahead of all models now! I’ve also been creating much longer and more advanced videos on T cells, immunology, and related topics, and I’m now planning to build an entire video lecture series and put it all on a dedicated website. I am just having the time of my life making these, it actually makes me emotional to be able to magically create these with few prompts!
— Derya Unutmaz, MD (@DeryaTR_) September 3, 2026
"Hey Astra, make me a Mac app. It should render a 3D iPod that you'll build in Blender, then using the original iPod UI and interactions I want to be able to visualize all my Codex threads on it." 15 minutes later: Boom. What's even in this model? What time we're living in!??
— Pietro Schirano (@skirano) September 3, 2026
Astra's clearest gains are in computer use, visual localization, long-running agents, and professional workflows. Independent testing keeps the overall ranking in perspective.
GPT-6 Astra is OpenAI's September 2026 frontier model for computer use, coding, professional knowledge work, and long-horizon agents. Official documentation lists a 1.05M-token context window, up to 128K output tokens, text and image input, configurable reasoning effort, and native support for web search, file search, code interpreter, image generation, skills, computer use, and remote MCP tools.
OpenAI's launch table reports 59.3% on Agents' Last Exam, 72.6% on OSWorld 2.0, 92.7% on ScreenSpot-Pro, 41.4% on AutomationBench, and 95.9% on BenchCAD. Those are vendor-reported results. Artificial Analysis independently scored Astra at 61 overall and 67 on its Coding Agent Index, below Fable 5.1 at 66 and 70 in the same launch snapshot.
Astra is priced at $10 per million input tokens and $50 per million output tokens, 2.5 times the standard GPT-5.6 Sol rate. API requests above 272K input tokens apply a 2x input and 1.5x output multiplier to the full session, so the largest context should be reserved for work that truly needs it.
GPT-6 Astra is live in the SPIRITT model picker and passed a real completion probe. Use it for computer-use and multimodal agent work where reliable execution matters more than the cheapest possible answer, then keep approvals and high-impact decisions with a person.
Artificial Analysis launch snapshot, September 4, 2026. Astra tied GPT-5.6 Sol within rounding and trailed the strongest Claude configurations on this broad composite.
The ranking above is independent. Seven capability cards below use OpenAI's published table; the coding-agent card uses Artificial Analysis.
OpenAI launch table. This benchmark evaluates long-horizon browser and desktop work.
OpenAI launch table using its reported computer-use harness.
OpenAI launch table. Visual grounding does not by itself prove full task completion.
OpenAI launch table. The absolute score still leaves substantial room for human review and recovery paths.
OpenAI launch table. Comparator rows marked by OpenAI as partly from other reporting are reproduced as published.
OpenAI launch table at the displayed high-capability settings.
OpenAI launch table. The top four models are separated by only 1.4 points.
Artificial Analysis launch snapshot. Astra was competitive and token-efficient, but did not lead the index.
Independent source: Artificial Analysis's September 4 GPT-6 Astra launch evaluation. Vendor source: OpenAI's GPT-6 Astra launch table and model documentation. OpenAI notes that some comparator results come from model providers or prior reporting and that its models used the displayed high-capability settings. Results vary with effort, harness, tools, context, and benchmark revision. Test your own computer-use and coding workflows before switching production work.
Choose GPT-6 Astra in the picker, then let it work through the browser, terminal, and files
Open a workspace and land in a fully equipped cloud computer: browser, files, terminal, integrations, and memory. No local setup. No thin chat box pretending to be an agent.

Open the SPIRITT model picker and select GPT-6 Astra. Use it for computer-use, multimodal, coding, and professional workflows where careful execution across many steps matters more than the lowest model cost.

Tell it what to ship or what to run. GPT-6 Astra can code, call tools, drive the browser, coordinate multi-step work, and keep going while you step away. The point is not another chat window. It is an agentic environment where GPT-6 Astra actually does the job.

Open a SPIRITT Workspace, choose GPT-6 Astra, and connect the model to the browser, terminal, files, tools, and acceptance checks required to finish real work.