Terminal Bench 2.1 · launch report
Fireworks’ launch evaluation, 89 tasks. Kimi K3 reasoning efforts are labeled. Reported pass rates; no confidence intervals are provided.
Meet Fireworks’ leaner take on Kimi K3. Ember-1 targets coding and multi-step work with shorter reasoning traces. Fireworks reports about 40% fewer tokens with comparable quality across its own evaluations.
Explore coding-tool support, developer tests, and independent comparisons. Shorter reasoning brings promising savings, alongside real trade-offs on quality, cost, and availability.
Ember-1 is a specialized model from Fireworks Research designed to make every token go further. Built on Kimi K3, it produces shorter reasoning traces, using roughly 40% fewer tokens while maintaining top-tier quality. https://t.co/PLD1l5aeh4 https://t.co/P8oNueFCQ4
— Fireworks (@FireworksAI_HQ) 2026-09-23T21:38:14.000Z
Most of the reasoning a model does isn't the answer. It's the model talking itself into the answer. With Ember-1 we went after that overhead: about 40% fewer tokens on Kimi K3, and the quality held up on live customer A/B tests, coding workloads and external benchmarks. I spent a lot of late nights watching models ramble. Glad to see it paid off. Try it on Fireworks.
— Pawel Garbacki (@pawelg) 2026-09-24T00:30:15.000Z
Excited to introduce Ember-1: Trained on @Kimi_Moonshot K3 to stop over reasoning. Same answer quality, 40% fewer tokens. - Problem: reasoning models spend lot of tokens on thinking/reasoning, which gets reread on every turn and stack up the costs. - Approach: 50+ training runs 200+ evals on @FireworksAI_HQ serverless training with new algorithms cut wasted reasoning. - Results: Ember-1 matches k3-max quality on 7 public benchmarks; sets the best cost-quality tradeoff against GPT-6 Astra Opus 5 on Bedside Bench - Business impact: one customer is moving to Ember-1 fully!
— Sophia Yang (@sophiamyang) 2026-09-24T14:46:19.000Z
Ember-1 by the Fireworks Research team is built on Kimi K3, and uses ~40% fewer tokens while achieving the same performance on benchmarks. This was accomplished by post-training K3 to think less repetitively. Reasoning models spend most of their output tokens (sometimes 90%+) on thinking before they answer, which gets expensive in agentic loops where the model tends to re-think the same thoughts on every step. Fireworks RL trained on real agentic coding task loops to teach the model which reasoning actually changes the answer vs which is just looping. In a live A/B test on coding traffic, Ember-1 used 71% fewer reasoning tokens and 39% fewer total tokens than K3 at the same success rate.
— Cline (@cline) 2026-09-27T21:58:58.000Z
Fireworks released Ember-1, their Kimi K3 based model, today. It makes a pretty good entrance into the BFI arena! Its performance is in line with Kimi K3 while being half the cost per task, due to using half the output tokens. https://t.co/sF1b3aFEYm
— aqui (@aquilesfd) 2026-09-24T22:57:34.000Z
Ember-1 scores are similar to Kimi K3's across the board, which makes sense because Kimi K3 is the base model used. It sacrifices only 8% of the time horizon in favor of 50% lower cost per task. That translates to a 40% higher productivity score https://t.co/t0dY8hv52Z
— aqui (@aquilesfd) 2026-09-24T23:01:45.000Z
If you combine the low amount of output tokens produced with the higher throughput speed of Ember-1, it ends up taking 60% less time to complete our benchmark. Even then, it's not in our time per task pareto frontier because of DeepSeek V4.1 Flash and Gemini 3.8 Flash https://t.co/o22eIfezCV
— aqui (@aquilesfd) 2026-09-24T23:03:09.000Z
Ember-1 is also not on our cost per task pareto frontier because it keeps the same per-token costs as Kimi K3, and is outparetoed by Step 5, MiMo V2.6 Pro, DeepSeek V4.1 Flash and GPT-6 Luna https://t.co/mNfzigdfVt
— aqui (@aquilesfd) 2026-09-24T23:04:04.000Z
One prompt, two OpenRouter slugs: Ember-1 vs MiMo-V2.6-Pro Same 13-word beauty brief. xiaomi/mimo-v2.6-pro returned Cosmica, a 45.4 KB landing page, in 101.29 s at $0.013697715 (runtime-gate PASS). fireworks/ember-1 returned no page: two HTTP 503s, Fireworks no healthy upstream. https://t.co/D8X8L35tds
— Mike Gannotti 🔳 (@MichaelGannotti) 2026-09-24T18:17:27.000Z
Fireworks' Ember-1 is Moonshot's Kimi K3, retrained to think less. K3's public tokenizer predicts all 50 of Ember-1's token counts exactly, and each request carries a hidden line telling it to think at maximum. Same price as K3. Read more here: https://t.co/CZAIl5ApvK
— YFarmX 🗞️ (@YFarmX) 2026-09-24T15:08:54.000Z
Ember 1 is on NanoGPT in research preview. Fireworks' Kimi K3-based model: images, tools and 1M context. Fireworks describes Ember as a research preview, initially offered for two weeks on its serverless platform. Continued availability depends on demand. https://t.co/O6Kg63g95G
— Nano-GPT (@NanoGPTcom) 2026-09-24T14:00:59.000Z
if you're using Kimi K3, give Ember-1 a spin to save on costs and e2e latency let me know if you want to keep using this beyond the research preview
— Richy Chen (@richychn) 2026-09-24T16:14:00.000Z
Ember-1 costs exactly what Kimi K3 costs per token on Fireworks. The whole saving is shorter reasoning, 49.3K to 29.9K output tokens on one coding workload. Good trade. It's also a two-week research preview with no published weights, so keep K3 wired in as the fallback. https://t.co/xkJ9FDNAFB
— Amit Shukla (@amitshuklabag) 2026-09-25T16:31:45.000Z
Long reasoning traces can make repeated tasks expensive. Ember-1 is a Kimi K3 specialization built to spend fewer tokens reaching useful answers.
Fireworks Research introduced Ember-1 on September 23, 2026. Its public catalog lists a 1,048,576-token context window, text and image input, text output, and function calling. It was trained for more concise reasoning, rather than simply running Kimi K3 at a lower reasoning-effort setting.
Fireworks lists $3 per million input tokens, $15 per million output tokens, and $0.30 per million cached input tokens. These match Kimi K3’s published standard rates. The potential saving comes from using fewer tokens, not paying less per token.
In the launch report, Ember-1 improves on Kimi K3 max in Terminal Bench and DeepSWE but scores slightly lower in SWE-bench Verified and SWE-Interact. The separate SII evaluation also shows quality trade-offs. Compare the exact task, not just the headline.
Ember-1 launched as a research preview. Fireworks describes two weeks of initial serverless access for its research releases, with continued availability based on community demand. Model access in SPIRITT depends on the models and connections enabled for your workspace.
Selected models in Fireworks’ publisher-run SII: 113 tasks, three runs, mini-swe-agent. This is a separate evaluation from the launch report below, not an independent intelligence ranking. Scores are rounded.
Compare coding, tool use, finance, and healthcare evaluations. Launch-report results and the separate Fireworks SII runs stay distinct, so cost and time savings do not hide quality trade-offs.
Fireworks’ launch evaluation, 89 tasks. Kimi K3 reasoning efforts are labeled. Reported pass rates; no confidence intervals are provided.
Fireworks’ launch evaluation, 500 tasks. Kimi K3 reasoning efforts are labeled. Reported pass rates; no confidence intervals are provided.
Fireworks’ launch evaluation, 75 tasks. Kimi K3 reasoning efforts are labeled. Reported pass rates; no confidence intervals are provided.
Fireworks’ launch evaluation, 113 tasks. Kimi K3 reasoning efforts are labeled. This reports different scores from the separate SII evaluation above.
Fireworks’ launch evaluation, 50 tasks. Kimi K3 reasoning efforts are labeled. Reported pass rates; no confidence intervals are provided.
Fireworks’ launch report: 5.9% to 51.9% lower evaluation cost across these five benchmarks. Higher means a larger saving. The report does not specify the aggregation behind its dollar deltas, so only its percentages are shown.
Fireworks’ coding A/B example: 49,300 versus 29,900 output tokens, about 39.4% fewer. Reported quality scores were 0.751 and 0.753. The score definition and sample size were not supplied.
Fireworks’ coding A/B example: 23.8 versus 21.4 steps. A step count is not a latency measurement. The customer and sample size were not disclosed.
Fireworks SII coding evaluation: 113 tasks, three runs, mini-swe-agent. Published per-task estimates are scaled to 100 tasks, not a bulk rate. Ember uses a Kimi K3 pricing proxy, matching its current published rates.
Fireworks SII mean end-to-end task duration, 113 tasks and three runs. Ember-1 takes less time here but has a lower quality score. This is not tokens per second or a speed guarantee.
Fireworks SII macro score: 500 cases including 250 training cases, three runs, a 16,384-token cap, and a GPT-5.6 Sol judge. This is not a pass rate or clinical validation.
Fireworks SII mean answer-model cost; judge cost is excluded. Published per-task estimates are scaled to 100 tasks, not a bulk rate. Ember uses a Kimi K3 pricing proxy, matching its current published rates.
Fireworks SII mean answer-only request time: 500 cases, three runs, a 16,384-token cap. Judge time is excluded. This is neither time to first token nor clinical readiness.
Fireworks SII: 139 private tasks, a Kimi K3 judge. Ember has one run with three trials per sample; Kimi has three independent runs. No confidence intervals are published.
Fireworks SII reported task-cost medians; Ember’s cost includes judging. Run and aggregation conditions differ, so this is not a controlled billing comparison. Published per-task estimates are scaled to 100 tasks, not a bulk rate. Ember uses a Kimi K3 pricing proxy, matching its current published rates.
Fireworks SII: the 97-task banking_knowledge subset v1.0.1, not full Tau3. Ember used default effort after max was dropped, with three runs versus Kimi’s two. Ember’s quality score is lower.
Fireworks SII banking subset; tested-model cost excludes judging and infrastructure. Ember used default effort after max was dropped. Published per-task estimates are scaled to 100 tasks, not a bulk rate. Ember uses a Kimi K3 pricing proxy, matching its current published rates.
Fireworks SII end-to-end trial duration on the 97-task banking subset. Ember is faster but scores lower. Effort and repetition differ: Ember default/three runs, Kimi two runs.
Sources: Fireworks’ September 23 Ember-1 announcement, public model catalog, and Specialized Intelligence Index snapshot from September 27. These charts are publisher-reported evaluations, not independent rankings. Launch and SII results are separate and are not combined. Scores are rounded; cost-per-100 charts scale the published per-task estimates. API rates are Fireworks list prices, not SPIRITT plan prices. Benchmark duration is not a universal speed promise.
Start with a real task, check your available model access, and compare the finished result rather than a token count alone.
Bring a code change, a document-heavy question, or a multi-step workflow into a SPIRITT workspace. Define what a good result needs to do before comparing models.

Ember-1 is a Fireworks research preview. Use it where it is available through your enabled models or connections. Access is not promised in every SPIRITT workspace.

Try representative tasks against your current model. Look at answer quality, actual cost, and completion time together. Shorter reasoning helps when the result still holds up.

Bring the work into SPIRITT, define the result you want, and use the model access available in your workspace. Compare Ember-1 on your own tasks wherever it is enabled.