SPIRITT logoSPIRITTFireworksEmber-1

Ember-1. Fewer tokens. More room to build.

Meet Fireworks’ leaner take on Kimi K3. Ember-1 targets coding and multi-step work with shorter reasoning traces. Fireworks reports about 40% fewer tokens with comparable quality across its own evaluations.

Ember-1 in the wild

Explore coding-tool support, developer tests, and independent comparisons. Shorter reasoning brings promising savings, alongside real trade-offs on quality, cost, and availability.

Keep useful reasoning. Trim the extra work.

Long reasoning traces can make repeated tasks expensive. Ember-1 is a Kimi K3 specialization built to spend fewer tokens reaching useful answers.

A smaller token footprint for agent work

Fireworks Research introduced Ember-1 on September 23, 2026. Its public catalog lists a 1,048,576-token context window, text and image input, text output, and function calling. It was trained for more concise reasoning, rather than simply running Kimi K3 at a lower reasoning-effort setting.

Fireworks lists $3 per million input tokens, $15 per million output tokens, and $0.30 per million cached input tokens. These match Kimi K3’s published standard rates. The potential saving comes from using fewer tokens, not paying less per token.

In the launch report, Ember-1 improves on Kimi K3 max in Terminal Bench and DeepSWE but scores slightly lower in SWE-bench Verified and SWE-Interact. The separate SII evaluation also shows quality trade-offs. Compare the exact task, not just the headline.

Ember-1 launched as a research preview. Fireworks describes two weeks of initial serverless access for its research releases, with continued availability based on community demand. Model access in SPIRITT depends on the models and connections enabled for your workspace.

Kimi K3 foundationLeaner reasoning1M contextText + image inputFunction callingResearch preview

DeepSWE v1.1: Fireworks SII

Claude Opus 5
72.27%
GPT-6 Astra
71.01%
Kimi K3
70.21%
GPT-6 Sol
67.26%
Ember-1
66.96%

Selected models in Fireworks’ publisher-run SII: 113 tasks, three runs, mini-swe-agent. This is a separate evaluation from the launch report below, not an independent intelligence ranking. Scores are rounded.

What the savings look like across real tasks

Compare coding, tool use, finance, and healthcare evaluations. Launch-report results and the separate Fireworks SII runs stay distinct, so cost and time savings do not hide quality trade-offs.

Fireworks launch evaluation
Winner

Terminal Bench 2.1 · launch report

Ember-1
82%
Kimi K3 max
80.9%
Kimi K3 high
77.6%
Kimi K3 low
76.4%

Fireworks’ launch evaluation, 89 tasks. Kimi K3 reasoning efforts are labeled. Reported pass rates; no confidence intervals are provided.

Fireworks launch evaluation

SWE-bench Verified · launch report

Kimi K3 max
93.2%
Ember-1
92.2%
Kimi K3 high
86%
Kimi K3 low
80.4%

Fireworks’ launch evaluation, 500 tasks. Kimi K3 reasoning efforts are labeled. Reported pass rates; no confidence intervals are provided.

Fireworks launch evaluation

SWE-Interact · launch report

Kimi K3 max
21.3%
Ember-1
20%
Kimi K3 high
13.3%
Kimi K3 low
6.7%

Fireworks’ launch evaluation, 75 tasks. Kimi K3 reasoning efforts are labeled. Reported pass rates; no confidence intervals are provided.

Fireworks launch evaluation
Winner

DeepSWE 1.1 · launch report

Ember-1
75.2%
Kimi K3 max
66.4%
Kimi K3 high
62.8%
Kimi K3 low
55.8%

Fireworks’ launch evaluation, 113 tasks. Kimi K3 reasoning efforts are labeled. This reports different scores from the separate SII evaluation above.

Fireworks launch evaluation
Winner

τ-2 Bench Airline · launch report

Ember-1
66%
Kimi K3 low
64%
Kimi K3 high
64%
Kimi K3 max
64%

Fireworks’ launch evaluation, 50 tasks. Kimi K3 reasoning efforts are labeled. Reported pass rates; no confidence intervals are provided.

Workload economics

Cost reduction vs. Kimi K3 max

Terminal Bench 2.1
51.9%
SWE-bench Verified
15.5%
SWE-Interact
32.5%
DeepSWE 1.1
23.7%
τ-2 Bench Airline
5.9%

Fireworks’ launch report: 5.9% to 51.9% lower evaluation cost across these five benchmarks. Higher means a larger saving. The report does not specify the aggregation behind its dollar deltas, so only its percentages are shown.

Publisher-reported customer A/B

Customer A/B output tokens

Ember-1
29900
Kimi K3
49300

Fireworks’ coding A/B example: 49,300 versus 29,900 output tokens, about 39.4% fewer. Reported quality scores were 0.751 and 0.753. The score definition and sample size were not supplied.

Publisher-reported customer A/B

Customer A/B steps per task

Ember-1
21.4
Kimi K3
23.8

Fireworks’ coding A/B example: 23.8 versus 21.4 steps. A step count is not a latency measurement. The customer and sample size were not disclosed.

Fireworks publisher evaluation

SII DeepSWE cost / 100 tasks

Ember-1
$418.92
Kimi K3
$539.71

Fireworks SII coding evaluation: 113 tasks, three runs, mini-swe-agent. Published per-task estimates are scaled to 100 tasks, not a bulk rate. Ember uses a Kimi K3 pricing proxy, matching its current published rates.

Fireworks publisher evaluation

SII DeepSWE time (seconds/task)

Ember-1
1712
Kimi K3
2340

Fireworks SII mean end-to-end task duration, 113 tasks and three runs. Ember-1 takes less time here but has a lower quality score. This is not tokens per second or a speed guarantee.

Fireworks publisher evaluation

Bedside Bench: SII quality

Claude Opus 5
94.6%
GPT-6 Astra
93.4%
Kimi K3
89.7%
Ember-1
89.1%
GPT-6 Sol
88.1%

Fireworks SII macro score: 500 cases including 250 training cases, three runs, a 16,384-token cap, and a GPT-5.6 Sol judge. This is not a pass rate or clinical validation.

Fireworks publisher evaluation

SII Bedside cost / 100 cases

Ember-1
$4.06
Kimi K3
$5.29

Fireworks SII mean answer-model cost; judge cost is excluded. Published per-task estimates are scaled to 100 tasks, not a bulk rate. Ember uses a Kimi K3 pricing proxy, matching its current published rates.

Fireworks publisher evaluation

SII Bedside time (seconds/case)

Ember-1
53.8
Kimi K3
82.9

Fireworks SII mean answer-only request time: 500 cases, three runs, a 16,384-token cap. Judge time is excluded. This is neither time to first token nor clinical readiness.

Fireworks publisher evaluation

Big Finance Benchmark: SII quality

Kimi K3
41.3%
Ember-1
39%

Fireworks SII: 139 private tasks, a Kimi K3 judge. Ember has one run with three trials per sample; Kimi has three independent runs. No confidence intervals are published.

Fireworks publisher evaluation

SII finance cost / 100 tasks

Ember-1
$19.50
Kimi K3
$22.00

Fireworks SII reported task-cost medians; Ember’s cost includes judging. Run and aggregation conditions differ, so this is not a controlled billing comparison. Published per-task estimates are scaled to 100 tasks, not a bulk rate. Ember uses a Kimi K3 pricing proxy, matching its current published rates.

Fireworks publisher evaluation

τ³-Banking: SII quality

Kimi K3
51.55%
Ember-1
38.14%

Fireworks SII: the 97-task banking_knowledge subset v1.0.1, not full Tau3. Ember used default effort after max was dropped, with three runs versus Kimi’s two. Ember’s quality score is lower.

Fireworks publisher evaluation

SII banking cost / 100 trials

Ember-1
$65.83
Kimi K3
$101.79

Fireworks SII banking subset; tested-model cost excludes judging and infrastructure. Ember used default effort after max was dropped. Published per-task estimates are scaled to 100 tasks, not a bulk rate. Ember uses a Kimi K3 pricing proxy, matching its current published rates.

Fireworks publisher evaluation

SII banking time (seconds/trial)

Ember-1
182.8
Kimi K3
422.6

Fireworks SII end-to-end trial duration on the 97-task banking subset. Ember is faster but scores lower. Effort and repetition differ: Ember default/three runs, Kimi two runs.

Sources: Fireworks’ September 23 Ember-1 announcement, public model catalog, and Specialized Intelligence Index snapshot from September 27. These charts are publisher-reported evaluations, not independent rankings. Launch and SII results are separate and are not combined. Scores are rounded; cost-per-100 charts scale the published per-task estimates. API rates are Fireworks list prices, not SPIRITT plan prices. Benchmark duration is not a universal speed promise.

Put leaner reasoning to work

Start with a real task, check your available model access, and compare the finished result rather than a token count alone.

01

Choose work worth repeating

Bring a code change, a document-heavy question, or a multi-step workflow into a SPIRITT workspace. Define what a good result needs to do before comparing models.

A soft 3D illustration of a workspace for practical tasks
02

Check your Ember-1 access

Ember-1 is a Fireworks research preview. Use it where it is available through your enabled models or connections. Access is not promised in every SPIRITT workspace.

A translucent peach and amber glass ember with a luminous core
03

Compare the finished task

Try representative tasks against your current model. Look at answer quality, actual cost, and completion time together. Shorter reasoning helps when the result still holds up.

A soft 3D illustration of an idea becoming a finished project

Make the next task count.

Bring the work into SPIRITT, define the result you want, and use the model access available in your workspace. Compare Ember-1 on your own tasks wherever it is enabled.

Questions

Know the trade-offs. Then build.

01What is Ember-1?+
Ember-1 is a Fireworks Research model built on Kimi K3, trained for shorter reasoning traces. Fireworks announced it on September 23, 2026 as a research preview aimed at more token-efficient coding and agentic work.
02Does Ember-1 really use 40% fewer tokens?+
Fireworks reports approximately 40% fewer tokens with comparable quality across its evaluations. Its customer A/B table shows output tokens falling from 49.3K to 29.9K, roughly 39.4%. The same table reports 71.3% fewer reasoning tokens and 39% fewer total tokens. These are different measures, not promises for every workload.
03What does Ember-1 cost?+
Fireworks lists $3 per million input tokens, $15 per million output tokens, and $0.30 per million cached input tokens. Kimi K3’s published standard rates match. Potential savings come from consuming fewer tokens, not a lower per-token rate. These are API list prices, not SPIRITT subscription prices; confirm any account, context, or service-mode conditions for your use.
04Is Ember-1 always faster or better than Kimi K3?+
No. The launch report shows gains and small regressions across coding and agent tests. Fireworks’ separate SII run reports lower Ember-1 quality scores on DeepSWE, finance, healthcare, and banking, alongside lower reported costs and shorter durations where timing is available. Test quality, cost, and time on the work you need.
05Why do the two DeepSWE comparisons differ?+
Fireworks’ launch report gives Ember-1 75.2%, while its separate SII DeepSWE evaluation reports 66.96%. They are different published evaluations, not one result. The available reports do not fully explain the discrepancy, so the page keeps their results separate.
06What inputs and tools does Ember-1 support?+
The public catalog lists text and image input, text output, and function calling, with a 1,048,576-token context window. An official hard maximum output limit was not established in the cited sources; a 128K request example is not proof of a universal output cap.
07Is Ember-1 permanently available?+
Not guaranteed. Fireworks describes two weeks of initial serverless access for its research releases and continued availability based on community demand. Confirm current availability before making a long-running workflow depend on this preview.
08Can I download or fine-tune Ember-1?+
Its Kimi K3 foundation does not establish a released Ember-1 checkpoint or license. No official weights download or Ember-specific license was found in the cited sources. The announcement discusses training support, while the current catalog marks fine-tuning as unsupported. Confirm access before planning a training workflow.
09Where do the comparisons come from?+
The charts use Fireworks’ launch report and its own Specialized Intelligence Index. They are not independent overall intelligence rankings. The X gallery also includes Aquiles’ separate BFI comparisons, including his finding that Ember-1 was not on that evaluation’s cost or latency frontier.
10Can I use Ember-1 in a SPIRITT workspace?+
Access depends on the models and connections enabled for your workspace. This page covers the Ember-1 research preview and its published results; it does not promise Ember-1 is enabled in every workspace. Start with an available model while checking the access options for Ember-1.
Buy from builders who use what they sellBuilt usingSPIRITT