AutomationBench 1.0.6
Business workflows across 47 tools. Efforts differ as labeled. Fable 5.1 includes Opus 5 fallback on about 40% of tasks. Source: OpenAI.
Tackle large code changes, longer research runs, and multi-step business work. GPT-6 Sol pairs capable coding and computer use with half the API token price of GPT-5.6 Sol, plus a 1.05M context window to keep more of your project in view.
See builders turn Sol into websites, explorable 3D scenes, games, videos, and real code changes, with firsthand comparisons along the way.
Please welcome GPT-6 Sol and GPT-6 Luna to the GPT-6 universe. GPT-6 Sol and Luna build on the advances behind GPT-6 Astra, bringing much of its strengths into faster and more affordable models to support work at scale. We’ve also made caching and inference more efficient, and we’re passing the savings directly to you: 50% lower API prices for Sol and Luna compared with GPT‑5.6 promotional pricing.
— OpenAI (@OpenAI) 2026-09-22T18:12:13.000Z
Built this site in just 15 mins using GPT-6 Sol with @EnterProAI. Meet an architecture-inspired space for every detail to breathe. Entirely prompted in Enter Pro. GPT-6 Sol hits the sweet spot between speed and cost. Luna offers an even lower-cost option for efficient smaller edits. The model got cheaper. The ideas don’t have to get smaller. #GPT6Sol #GPT6Luna @OpenAI
— Enter Pro by Converge.AI (@EnterProAI) 2026-09-23T05:39:54.000Z
I asked GPT 6 Sol to build a fully explorable 3D model of an apartment for sale in Paris (link in comments) I just gave it: 1. a short prompt 2. the URL of the listing with photos / floor plans 3. the Unreal Engine MCP The result is jaw dropping 👇 https://t.co/2fFXQrJqDK
— JC (@anakin) 2026-09-22T20:57:06.000Z
First impression of GPT-6-Sol Asked it to generate a single Three.js powered offline HTML game with an explorable pagoda valley, village, river, bridge. Generate it from web-chat ui, and took 8 min in total duration I guess it's fine, faster than Astra for sure, but you be the judge, what do you think?
— AJ (@ItsmeAjayKV) 2026-09-22T19:36:33.000Z
GPT-6 Sol rocks and cooks well video made with @HyperFrames_ with only a few prompts! pretty good results so far https://t.co/KAj8sRyke8
— Miguel Ángel (@Miguel07Code) 2026-09-22T19:27:34.000Z
GPT 6 Sol MAX can cook on design too. Solid frontend. https://t.co/myuvoCe3Ih
— Justin (@JustinGorya) 2026-09-22T19:02:54.000Z
I asked GPT-6 Sol to make a site for itself. https://t.co/9IdpoZ90PA
— zackbuilds (@zaxskylab) 2026-09-22T22:35:28.000Z
compared GPT-6 Sol and Astra on my "robot drawing on the board" bench Sol used ~25k output tokens vs ~22k for Astra. interestingly, Sol's controller drew twice as fast to be honest, i expected stronger spatial reasoning from new Sol here, but it was way cheaper to run (almost 5x)
— Dmytro Hrybov (@dimentary) 2026-09-22T23:24:01.000Z
GPT 6 Astra vs GPT 6 Sol vs GPT 6 Luna Spatial Intelligence Test: Dropped them into a city (Lublin) and asked to find a basketball court. Astra Completed in: 1m 9s Input tokens: 323,782 Output tokens: 1,340 Sol Completed in: 3m 7s Input tokens: 746,387 Output tokens: 6,031 Luna Failed to complete.
— Arun Kurian (@AKurian001) 2026-09-22T21:15:58.000Z
GPT-6 Sol and Luna are now available in Devin. On FrontierCode 1.1, GPT-6 Sol matches GPT-5.6 Sol’s score at 61% lower cost per task. GPT-6 Luna scores above GPT-5.6 Luna at about a quarter of the cost. At under $0.10 per task, it is the cheapest model on the leaderboard. https://t.co/Mct0ZGxMO2
— Cognition (@cognition) 2026-09-22T18:22:56.000Z
GPT-6 Sol debuts at 97% on Next.js evals, tying the highest success rate. Claude Opus 5.5 also reached 97%, joining Claude Fable 5.1. Opus ranks first among the three on average cost. View the leaderboard ↓ https://t.co/gjzCSLeMYE https://t.co/0Bn5U3b1Rg
— Next.js (@nextjs) 2026-09-22T21:52:02.000Z
GPT-6 Sol is #4 on BioMysteryBench, an open-source benchmark for agentic biology tasks. The top three models are tied at 79.3%, with GPT-6 Sol close behind at 74.8%. At $0.66 per test, it is also very cost-effective. https://t.co/iUpbRUEZll
— Vals AI (@ValsAI) 2026-09-23T05:23:18.000Z
GPT-6 Sol is now available to all Perplexity users. On our Wide-And-Deep-Research (WANDR) evals, it outperforms Opus 5 at one-fifth the price. Sol will become the “Light” Effort orchestrator for Computer users, while Astra remains the orchestrator for "High" effort. Congrats to @OpenAI for consistently launching pareto-optimal models!
— Aravind Srinivas (@AravSrinivas) 2026-09-22T19:06:52.000Z
GPT-6 Astra is OpenAI's flagship. GPT-6 Sol is the cheaper model positioned below it. On our writing benchmark, at each model's deepest effort setting: Sol scores 2182 Elo. Astra scores 2013. Sol costs $0.11 a script. Astra costs $0.85. Sol takes 171 seconds. Astra takes 526. The cheap model wins by 169 Elo while costing an eighth as much and running three times faster. Every effort setting of Sol we tested beats every effort setting of Astra. Astra's default setting sits at 1843, which is below where GPT-5.6 Sol was. this measures writing YouTube scripts in one editorial voice. Astra is built for agentic and engineering work, and nothing here says it is bad at those. It does say that paying the flagship price for prose is the wrong call. Check the cheap model against your own task before you assume the expensive one is better.
— Louis-François Bouchard 🎥🤖 (@Whats_AI) 2026-09-23T00:40:00.000Z
BREAKING: @OpenAI just dropped GPT-6 Sol. It’s my new daily driver in Codex: not quite Astra, but close enough for much of my everyday work, faster, and 50% cheaper than 5.6 Sol. We tested it across the work we actually do at @every and ran it head to head versus Opus 5.5. Here’s my vibe check: • Writing: On a paragraph-writing task drawn from my real work, Sol scored close to Astra, which is still my top model for writing. It writes clean, minimal prose and puts the important idea first. Opus 5.5 is pleasant to work with, but its drafts still tend to bury the point. • Computer use: If you love Astra’s computer use, you’ll like Sol. An earlier Sol preview matched Astra on 17 of 18 attempts across six of our simpler Hands tasks, at a much lower token price. • Coding: Sol improves on GPT-5.6—including better Ruby code in @kieranklaassen tests—but Opus 5.5 has the higher …
— Dan Shipper (@danshipper) 2026-09-22T18:14:11.000Z
Opus 5.5 absolutely beats GPT-6 Sol in 3D Despite being 13x more expensive than GPT-6 Sol, the high quality of Opus 5.5 is undeniable! cost: GPT-6 Sol: $0.34 Claude Opus 5.5: $4.37 Try both on https://t.co/wFX5XRmB47 https://t.co/wBHMizopsi
— AI/ML API (@aimlapi) 2026-09-22T21:49:33.000Z
Rocket launch, one prompt. Claude Opus 5.5 vs GPT-6 Sol, both day-one on OpenRouter. Opus: 18 min · 118K tokens · 79K thinking · $2.37 Sol: 1.7 min · 11.8K tokens · 2K thinking · $0.12 Opus: 20× the cost, 40× the thinking, the prettier launch pad. Sol: launched anyway. Zero fixes on either. Obviously Opus 5.5 is way way better here
— Wësche (@WescheNex1q) 2026-09-23T00:30:00.000Z
it cooked! 🥇 Opus 5.5 Fixed every source of the bug, made the smallest eval change, and backed it with before/after numbers. 🥈 GPT-6-Sol Best tests and docs, but left two skills contradicting the fix and never ran its own eval. 🥉 GPT-6-Astra Right idea at the wrong cost: guidance on every prompt and a separate 330-line eval harness. 4️⃣ GPT-6-Luna One line. It didn't touch the skills causing the bug, so the bug stays. 1 and 2 are close. 3 and 4 aren't.
— Hegar (@hegargarcia) 2026-09-23T00:44:57.000Z
Sol is the middle ground for ambitious projects: more capable than the smallest model, with a much lighter bill than the flagship.
OpenAI positions GPT-6 Sol for complex coding and agentic workflows. It supports text and image input, up to 922K input tokens within a 1.05M total context window, and up to 128K output tokens. Medium is the default reasoning effort, with options from none to max.
Standard API prices are $2/M input and $10/M output, both 50% below GPT-5.6 Sol. Cached input costs $0.20/M. Reusing project context can make a meaningful difference in long coding sessions and repeated agent work.
Independent testing gives the price cut practical context: Artificial Analysis reports $1.06 per Intelligence Index task versus $1.99 for GPT-5.6 Sol, both at max effort. Its Coding Agent Index rises from 55 to 57, while broad intelligence stays roughly level and some knowledge-work results regress.
Artificial Analysis, September 22. All rows use max effort; the two Claude models include fallback. A broad capability comparison, not a task-price ranking.
Compare code changes, business workflows, computer use, and the cost of keeping an agent working. Effort levels and evaluation setups are labeled so like-for-like comparisons stay clear.
Business workflows across 47 tools. Efforts differ as labeled. Fable 5.1 includes Opus 5 fallback on about 40% of tasks. Source: OpenAI.
Same settings as the score card. Fable’s shown cost excludes Opus 5 fallback; its actual total is higher. Source: OpenAI.
Long-horizon professional workflows. Opus 5 uses high, its best evaluated setting in this source; the OpenAI rows use max. Source: OpenAI.
Changes are judged for correctness and merge-readiness. Fable 5.1 medium is shown rather than a weaker higher-effort run. Source: OpenAI.
Complex repository work. Efforts are labeled; these official research/API runs differ from Artificial Analysis’s Codex evaluation below.
Partial reward on the v2026.08.08 offline set, not full-task completion. Both generations are shown at max; any other effort is labeled. Source: OpenAI.
Lower is better. Prompts were selected from user-flagged mistakes, so these are not ordinary-usage error rates. Max effort. Source: OpenAI.
Artificial Analysis, September 22. Codex for OpenAI; Claude Code for Claude, with fallback for Fable. All rows use max effort.
Artificial Analysis’s Coding Agent evaluation at max effort. Same Codex setup for both generations; separate from the official OpenAI comparison.
Repository questions in the same Codex evaluation at max effort. Published in Artificial Analysis’s September 22 launch report.
September 22 v4.3 launch snapshot, max effort. The models use more output tokens than their predecessors; the savings are driven by lower prices.
Lower is better. AA’s special hallucination metric is not an all-response error rate. Accuracy is shown separately; declining questions can change both.
Correct answers across all questions at max effort. Sol answers fewer questions, reducing hallucinations but also lowering accuracy. Source: Artificial Analysis.
Lower is better. OpenAI deliberately selected difficult, dishonesty-inducing coding tasks. Max effort; not a typical usage rate.
Lower is better. This challenge tests whether an agent admits that its search tool is broken. Max effort; not a typical usage rate.
OpenAI Standard rates for requests up to 272K input tokens. Both generations’ input price drops by 50%.
OpenAI Standard token rates. Sol’s output price drops by 50%; total task costs depend on token use and settings.
Cached reads cost 10% of uncached input. Repeated project context can use the discount when cache requirements are met.
Artificial Analysis snapshot, September 23, OpenAI API at max effort. Speed after the first chunk, not a guarantee of total task time.
Sources: OpenAI’s September 22 launch and GPT-6 system-card appendix, official API documentation, and Artificial Analysis. Each chart keeps its evaluator, version, effort, and fallback conditions. OpenAI’s comparator runs come from public reports; Fable 5 appears where 5.1 results were unavailable. The AA Intelligence and Coding Agent launch snapshots use v4.3 and v1.5 respectively. Prices are OpenAI API list rates, not SPIRITT plan prices; long-context, tool, and optional service charges can apply. Output speed is a separate September 23 observation.
Give the work a home, choose the right model, and keep the result moving.
Start with your files, repository, or business goal. A workspace brings your agent, browser, terminal, integrations, and memory together around the project.

Check the models and connection options enabled in your workspace. Choose the level of capability and cost that fits the work, then set a clear finish line.

Work toward a usable output: a tested change, a finished report, or a repeatable workflow. Keep the files and decisions together, and review important actions before they go live.

Bring your agent the tools, context, and workspace to turn a request into something useful.