AutomationBench 1.0.6
At high effort, Luna gains 5.4 points over its predecessor. This is OpenAI’s 47-tool workflow comparison, separate from AutomationBench-AA.
Extract data, triage queues, classify documents, and run focused agents without giving every request a flagship-sized bill. GPT-6 Luna brings a 1.05M context window at $0.10 input and $0.50 output per million tokens.
Explore low-cost websites, 3D scenes, CAD, and creative workflows, alongside firsthand notes on where Luna shines and where it needs more help.
Please welcome GPT-6 Sol and GPT-6 Luna to the GPT-6 universe. GPT-6 Sol and Luna build on the advances behind GPT-6 Astra, bringing much of its strengths into faster and more affordable models to support work at scale. We’ve also made caching and inference more efficient, and we’re passing the savings directly to you: 50% lower API prices for Sol and Luna compared with GPT‑5.6 promotional pricing.
— OpenAI (@OpenAI) 2026-09-22T18:12:13.000Z
GPT‑6 Luna and GPT‑6 Sol are now available in the NEXUS AI App Builder. Watch GPT‑6 Luna build a website from a prompt, check it in the preview, and refine it with a follow‑up instruction. From idea to working app to deploy with NEXUS AI. https://t.co/Z2JXu6UjnM
— NEXUS AI (@nexusai_apps) 2026-09-22T20:02:47.000Z
OpenAI’s new GPT-6 Luna is a sleeper model that truly impresses. Its ability to handle single-shot 3D game output is remarkable, especially given its compact size and accessible price point. Luna made the assets in Blender. My usage when using Luna just doesn't drop! https://t.co/96VnCnbvqf
— Matthew Lopez (@MattAperture) 2026-09-22T20:02:59.000Z
GPT-6 Luna might have just broken the cheap-model conversation. 27 coding-agent requests. 1.58M input tokens. 1.1M cached. Total bill: $0.08. I’ve already had it implementing substantial StackReplay work, with Sol auditing behind it. Frontier-level capability at DeepSeek-level economics? Maybe. Going to keep hammering it.
— Tyler (@btsouth) 2026-09-22T23:46:05.000Z
We had early access to GPT 6 Sol and Luna and ran them on Ramp Accounting Bench. Sol did better than GPT 5.6 Sol at 58% lower cost. Luna did better than GPT 5.6 Terra at 87% lower cost. https://t.co/sSsDKlOia2
— Ramp Labs (@RampLabs) 2026-09-22T20:06:20.000Z
I already like GPT-6 Luna a lot. So far most people seem to be disappointed, and some benchmarks put it below GPT-5.6 Luna. On my own eval it comes out ahead, at almost half the price (it's crazy cheap). I ran both models against a knowledge base of a fictional construction company, asking the kind of questions a site manager asks. GPT-6 Luna made four mistakes in 48 questions, GPT-5.6 Luna made seven. Execution time was basically the same, with GPT-6 slightly ahead.
— Timm Schäuble (@timmschaeuble) 2026-09-22T22:15:13.000Z
GPT-5.6 Luna vs GPT-6 Luna — same CAD prompt, different result. ⚙️ I asked both models to generate the same tilted phone holder in L2CAD. Watch them go from natural language → parametric 3D CAD. Which one did better? 👀 #L2CAD #AI #CAD #3DPrinting #GenerativeAI https://t.co/teAUAmUoOo
— L2CAD (@L2CAD_com) 2026-09-22T21:23:25.000Z
📝おお~!GPT-6 Luna、Hyper FramesとVOICEVOXを使った記事解説動画も安定して作れるようになってる! 今までOpusやAstraなど上位モデルで安定してでていたものが、GPT-6 Luna 極高で出せるようになるのは本当に嬉しい モデル性能が上がったら、まずは下位モデルでも出せるか試すのが良いです https://t.co/oztiQ9ct6V
— テツメモ|AI図解×検証|Newsletter (@tetumemo) 2026-09-23T03:43:05.000Z
Made with GPT-6 Luna Extra High. Time taken: 1 hour 38 minutes. #GPT6Luna #AI #MadeWithAI #newmodel #chatgpt @ChatGPT https://t.co/d38gUhN3ED
— Synapse keyboard (@Albatrozshadow) 2026-09-22T22:17:33.000Z
updated my unscientific design bench - Claude Opus 5.5 Low/High/Max - GPT-6 Sol High/Max - GPT-6 Luna High/Max quick thoughts: 1️⃣ OpenAI models are MUCH cheaper (and faster) 2️⃣ Claude looks very... Claude, but has a lot more detail/nuance across every page 3️⃣ Sol is unbeaten for the price/speed and I bet A LOT of people will wake up to the model in the next weeks 4️⃣ Luna is... unreasonably good for its speed and price... it has no business being this good at that cost 5️⃣ LAYOUT-wise I actually prefer the OpenAI models, but they lack much of the detail that Opus 5.5 adds across the board what do you think? https://t.co/QRNrh9ucmr
— Robin Ebers (@robinebers) 2026-09-23T03:28:57.000Z
I tested opus 5.5 vs gpt-6 sol / luna on my everyday coding tasks 10 tasks, domains: data engineering / backend / frontend / dev-ops / trajectory analysis results: - opus 5.5 high and astra low cost almost the same. - sol high matches their score at roughly half the cost. - luna is soo cheap: 29/30 passes, ~22× cheaper than sol honestly, i think most of the short-scoped coding tasks (<1h) are already solved by most models and now we just optimize speed and cost and scale up parallel execution / swarms
— Ibragim (@ibragim_bad) 2026-09-22T22:17:25.000Z
How is GPT-6 Luna even real… so insanely cheap to run, and look at the output. I was looking forward to Grok 4.7, but was disappointed. I figured Claude Opus 5.5 would be great. But GPT-6 Luna is blowing my mind. Far exceeds my expectations! https://t.co/l5XDA4JiVz
— Wayne Lowry (@wikiwayne) 2026-09-22T21:03:39.000Z
GPT 6 Luna gave me this ship render It is kinda okay for a model that is nearly free https://t.co/AO1bEZbuUz
— Saksham Saini (@im_sakshamsaini) 2026-09-22T19:49:51.000Z
We gave 4 models the same prompt: build a 3D rocket launch. GPT 6 Luna: <$0.01, 1 minute GPT 6 Sol: $0.11, 1 minute Grok 4.7: $0.29, 11 minutes Claude Opus 5.5: $1.52, 12 minutes GPT 6 Luna built a rocket launch scene in 65 seconds for less than a penny. The cost of frontier AI is collapsing in real time.
— Bridgebench (@bridgebench) 2026-09-22T19:15:13.000Z
Preliminary test results for GPT-6 models on music error detection: GPT-6 Astra (Extra High): 100% GPT-6 Sol (Medium): 86% GPT-6 Luna (High): 62% 6 of Luna’s 8 wrong answers were nearly correct -- Luna correctly identified the beats, voices, and type of error involved, but misidentified the measure. It only did this on chorales with pickup beats, which means it was almost certainly counting pickups as measure 1. I’m surprised no other LLM has made this mistake before, because it seems like an obvious failure point in retrospect.
— Auggie (@aug5thmusic) 2026-09-22T19:45:06.000Z
GPT-6 Luna on MAX first test was not impressive in terms of graphics. It made the models using few voxels ending up being too simple. The overall table could not be moved properly as well. On speed and price is where it shines: Cost $0.03 Time 15 Min https://t.co/XUXRwROdWN
— Daniel Zambrini (@DanielZambrini) 2026-09-22T19:45:41.000Z
GPT-6 Luna was not able to complete this task in our test. But it is much cheaper: 62 API calls cost $0.23, about $0.0037 per call. For comparison, GPT-6 Astra used 29 API calls and cost ~$3.20, about $0.11 per call. https://t.co/RVGEirXEJX
— Yu Xiang (@YuXiang_IRVL) 2026-09-22T21:24:09.000Z
I just tested GPT-6 Luna in FL Studio, it took 25 minutes to make a simple melody and didn't make a full beat. Only loaded Serum stock sound and never sound designed or found presets/drums from the folders and plugins I told it to use. Safe to say, this is nowhere near Astra for music production.
— Busy Works Beats (@BusyWorksBeats) 2026-09-22T19:28:26.000Z
Use Luna where the tasks are focused and the volume adds up. Save the heavier model budget for the decisions and projects that need it.
OpenAI positions GPT-6 Luna as its most efficient model for focused, high-volume work. It supports text and image input, up to 922K input tokens within a 1.05M context window, and up to 128K output tokens. Medium is the default effort; harder cases can use higher settings.
Standard API prices are $0.10/M input, $0.50/M output, and $0.01/M cached input. Compared with GPT-5.6 Luna, input is 50% cheaper and output is about 58% cheaper. Both standard rates are one twentieth of GPT-6 Sol’s.
Artificial Analysis reports roughly $0.07 per Intelligence Index task at max effort, versus $0.18 for GPT-5.6 Luna. Its Coding Agent Index falls from 43 to 41, so the savings are not a promise of better results on every task. Use Luna for focused work and check that the output is complete.
Artificial Analysis, September 22. OpenAI family comparison at max effort. Luna and its predecessor both round to 37, at different costs.
See the balance between coding, workflows, computer use, and cost. Lower prices are the headline; capability and completeness still depend on the task.
At high effort, Luna gains 5.4 points over its predecessor. This is OpenAI’s 47-tool workflow comparison, separate from AutomationBench-AA.
Same high-effort settings as above: 2.08 versus 5 cents per task, about 58% lower. Source: OpenAI.
Long-horizon professional workflows. Opus 5 uses high, its best evaluated setting in this source; the OpenAI rows use max. Source: OpenAI.
Changes are judged for correctness and merge-readiness. Fable 5.1 medium is shown rather than a weaker higher-effort run. Source: OpenAI.
Complex repository work. Efforts are labeled; these official research/API runs differ from Artificial Analysis’s Codex evaluation below.
Partial reward on the v2026.08.08 offline set, not full-task completion. Both generations are shown at max; any other effort is labeled. Source: OpenAI.
Lower is better. Prompts were selected from user-flagged mistakes, so these are not ordinary-usage error rates. Max effort. Source: OpenAI.
Artificial Analysis, September 22. Codex for OpenAI; Claude Code for Claude, with fallback for Fable. All rows use max effort.
Artificial Analysis’s Coding Agent evaluation at max effort. Same Codex setup for both generations; separate from the official OpenAI comparison.
Repository questions in the same Codex evaluation at max effort. Published in Artificial Analysis’s September 22 launch report.
September 22 v4.3 launch snapshot, max effort. The models use more output tokens than their predecessors; the savings are driven by lower prices.
Lower is better. AA’s special hallucination metric is not an all-response error rate. Accuracy is shown separately; declining questions can change both.
Correct answers across all questions at max effort. Luna’s accuracy is broadly level while its hallucination metric improves. Source: Artificial Analysis.
Lower is better. OpenAI deliberately selected difficult, dishonesty-inducing coding tasks. Max effort; not a typical usage rate.
Lower is better. This challenge tests whether an agent admits that its search tool is broken. Max effort; not a typical usage rate.
OpenAI Standard rates for requests up to 272K input tokens. Both generations’ input price drops by 50%.
OpenAI Standard token rates. Luna’s output falls from $1.20 to $0.50, about 58.3% lower rather than exactly half.
Cached reads cost 10% of uncached input. Repeated project context can use the discount when cache requirements are met.
Artificial Analysis snapshot, September 23, OpenAI API at max effort. Speed after the first chunk, not a guarantee of total task time.
Sources: OpenAI’s September 22 launch and GPT-6 system-card appendix, official API documentation, and Artificial Analysis. Each chart keeps its evaluator, version, effort, and fallback conditions. OpenAI’s comparator runs come from public reports; Fable 5 appears where 5.1 results were unavailable. The AA Intelligence and Coding Agent launch snapshots use v4.3 and v1.5 respectively. Prices are OpenAI API list rates, not SPIRITT plan prices; long-context, tool, and optional service charges can apply. Output speed is a separate September 23 observation.
Give the work a home, choose the right model, and keep the result moving.
Start with your files, repository, or business goal. A workspace brings your agent, browser, terminal, integrations, and memory together around the project.

Check the models and connection options enabled in your workspace. Choose the level of capability and cost that fits the work, then set a clear finish line.

Work toward a usable output: a tested change, a finished report, or a repeatable workflow. Keep the files and decisions together, and review important actions before they go live.

Bring your agent the tools, context, and workspace to turn a request into something useful.