top of page

Fireworks AI Ember-1 Cuts Kimi K3's Token Load, but the Evidence Is Still Vendor-Run

4 days ago
12 min read

Fireworks AI released Ember-1 with a direct promise: retain Kimi K3's task performance while generating about 40% fewer tokens. The Fireworks AI Ember-1 model targets a costly weakness in reasoning systems, especially coding agents that repeatedly carry earlier reasoning into later turns.

The important contest is not Ember-1 against an unrelated frontier model. It is trained efficiency against a simpler alternative, lowering Kimi K3's reasoning effort at inference time. Fireworks says the lower setting saves tokens but loses too much accuracy, while post-training teaches Ember-1 which reasoning to preserve.

That claim matters because token-heavy reasoning compounds inside long agent sessions. However, the evidence comes primarily from Fireworks itself. Ember-1 is also an API-only research preview, leaving outside researchers unable to inspect its weights or reproduce the training process independently.

Fireworks AI Ember-1 Changes the Cost Side of Kimi K3

Ember-1 turns reasoning length into a trainable behavior instead of treating it as a fixed cost of model quality.

Fireworks announced Ember-1 on September 23, 2026. According to the official Ember-1 release, the specialized model was post-trained from Moonshot AI's Kimi K3 open weights.

Post-training means additional training performed after a base model has learned its general capabilities. Here, the objective was narrower than building a new foundation model. Fireworks wanted K3 to use shorter reasoning traces without abandoning useful analysis.

The company says customers valued K3's coding ability but found its long reasoning expensive at scale. Its researchers concluded that lowering the available reasoning effort did not preserve enough quality. They instead trained Ember-1 to remove redundant loops while retaining productive reflection.

Fireworks reports that its team completed more than 50 training experiments and over 200 evaluations. The training set covered mathematics, coding, instruction following, conversation, search, tool use, and software engineering.

Those categories matter because token efficiency can overfit to one narrow workload. A model trained only on short coding problems might struggle when an agent must search, call tools, revise a plan, and recover from errors.

Fireworks says task feedback and environment feedback guided on-policy learning. On-policy learning uses behavior produced by the current model during training, allowing feedback to shape the reasoning patterns it actually generates.

The company also says it used its own data rather than customer data. It has not disclosed the complete dataset, training algorithms, training code, or Ember-1 weights.

Ember-1 is currently available through Fireworks Serverless as a research preview. Fireworks describes these previews as limited serverless releases that can become permanent when community demand justifies continued availability.

That distribution choice creates the first major limitation. Developers can test the model through an API, but they cannot self-host it or examine its parameters. Independent evaluators must therefore test the hosted endpoint under conditions controlled by Fireworks.

Kimi K3 provides the foundation for the comparison. Moonshot introduced K3 in July as a 2.8-trillion-parameter model with native vision and a one-million-token context window. Its official Kimi K3 specifications position it for long coding sessions, knowledge work, and reasoning.

Moonshot also exposes low, high, and maximum reasoning-effort settings. That makes K3 an unusually useful baseline for Fireworks' argument. The same underlying model family can be compared across inference settings and a separately post-trained variant.

The event is therefore more specific than another model launch. Fireworks is proposing that providers should train away waste instead of asking customers to tolerate it or manually reduce reasoning depth.

That idea puts pressure on inference platforms and model vendors alike. If token-efficient post-training works across production workloads, raw benchmark quality becomes only one part of the purchasing decision.

Why Long Reasoning Traces Become an Agent Cost Problem

A verbose reasoning trace is not merely a longer answer because an agent may carry that trace through every later step.

Reasoning models generate intermediate analysis before producing a final response. Fireworks says those internal reasoning tokens sometimes represent more than 90% of a model's generated output.

The percentage is a company-reported observation, not a universal property of every reasoning model. Still, it identifies a genuine architectural concern for multi-step agents.

A single long response incurs its cost once. An agent, however, often sends prior messages back to the model when it calls another tool, reviews a result, or attempts a correction.

Earlier reasoning can then be processed repeatedly. Fireworks describes this context growth as roughly quadratic with the number of turns, since each new turn can replay an expanding history.

The practical effect becomes visible in coding systems. An agent may inspect a repository, propose a patch, run tests, diagnose failures, and revise several files. Each additional turn can include earlier analysis that no longer helps the next decision.

Simply hiding reasoning from the user does not necessarily eliminate that burden. The provider still generates and processes reasoning tokens, depending on its API design and context-handling rules.

Reducing the reasoning-effort setting offers an obvious response. The model spends less time analyzing, produces fewer tokens, and returns faster.

Fireworks argues that this approach removes valuable reasoning alongside redundant reasoning. Its benchmark results show K3 at low effort trailing K3 at maximum effort across several evaluated coding tasks.

For example, Fireworks reports a 76.4% pass rate for K3 Low on Terminal-Bench 2.1. K3 Max reached 80.9%, while Ember-1 reached 82.0%.

The benchmark contains 89 terminal-based tasks covering software engineering, system administration, data processing, security, and related command-line work. Its maintainers revised 28 tasks when releasing Terminal-Bench 2.1.

The difference between lower effort and trained efficiency is the central mechanism. A lower setting provides the original model with less reasoning budget. Post-training attempts to change how the model allocates that budget.

Useful self-reflection can include checking an assumption, noticing a failed command, or revising a plan after environmental feedback. Redundant reasoning includes repeated summaries, abandoned loops, and lengthy deliberation that does not alter the final action.

Ember-1 is supposed to distinguish between those categories. Fireworks says the model retains reflection that improves task completion while curtailing unproductive loops, including during unsuccessful attempts.

That distinction is difficult to validate from final answers alone. Two models can produce the same correct patch while taking very different internal paths. They can also show comparable aggregate scores while failing on different tasks.

Production teams therefore need more than average token counts. They need distributions showing when compression works, which tasks lose accuracy, and whether rare failures become harder to detect.

Latency also deserves attention. Fewer output tokens usually shorten generation time, but tool execution and input processing can dominate some agent workflows. Fireworks has not published enough independent evidence to generalize the latency impact across deployments.

Even with those qualifications, the mechanism is strategically important. If post-training consistently removes waste, reasoning efficiency becomes a model property rather than an application-side compromise.

Trained Efficiency Beats the Lower-Effort Shortcut

Fireworks' strongest evidence is not that Ember-1 always wins, but that it approaches K3 Max quality with fewer generated tokens.

Fireworks evaluated Ember-1 against Kimi K3 at low, high, and maximum reasoning settings. The company calculated benchmark costs using the same public K3 rates, isolating savings created by shorter outputs.

The most favorable results appear on Terminal-Bench 2.1 and DeepSWE 1.1. Ember-1 scored 82.0% on Terminal-Bench, compared with 80.9% for K3 Max.

On DeepSWE 1.1, Fireworks reports 75.2% for Ember-1 and 66.4% for K3 Max. The evaluation covered 113 tasks.

DeepSWE tests long-horizon engineering work across active repositories and several programming languages. Its published DeepSWE methodology emphasizes original tasks intended to reduce exposure and contamination concerns.

The results are not uniformly favorable. Ember-1 scored 92.2% on SWE-bench Verified, while K3 Max scored 93.2%. It also reached 20.0% on SWE-Interact, compared with 21.3% for K3 Max.

Those losses are small in percentage points, but they matter. They show that the phrase "same quality" describes an aggregate judgment rather than identical capability on every test.

Ember-1 reached 66% on the airline portion of τ-2 Bench, against 64% for all three K3 effort settings. That result covered 50 samples, the minimum size Fireworks used for its published comparison.

Across seven benchmarks and two customer workloads, Fireworks says it shortened K3 reasoning by 35% to 50% without sacrificing overall accuracy. The company places Ember-1 on or near a quality-versus-cost Pareto frontier.

A Pareto frontier describes options where improving one dimension requires giving up another. In this case, a model belongs near the frontier when no alternative offers both better quality and lower task cost.

That framing is more useful than a single leaderboard rank. Enterprise users care about the cost of completing acceptable tasks, not just the maximum percentage attached to a model name.

Still, Pareto claims depend heavily on the selected tasks, agent scaffold, prompting, retry policy, and cost assumptions. Changing any of those inputs can move a model's position.

Fireworks also evaluated Ember-1 on Bedside Bench, a physician-validated set of 500 clinical cases across 10 categories. It says Ember-1 established a new cost-per-task frontier across the models included in its Specialized Intelligence Index.

That result broadens the model's story beyond coding. However, a medical benchmark does not establish that the model is suitable for clinical deployment, diagnosis, or unsupervised medical decisions.

The test is better understood as evidence about structured professional reasoning. It shows how Fireworks wants specialized models evaluated across real task categories, rather than only general academic tests.

Production buyers should also distinguish percentage-point differences from operational reliability. A small average improvement can conceal regressions on tasks that matter most to one company.

The proper comparison is therefore workload-specific. Teams should replay representative tasks with fixed agent scaffolds, identical tool permissions, and consistent success criteria.

They should measure total tokens, completed tasks, retries, time to completion, and failure severity together. Reducing tokens is valuable only when the system still reaches a usable outcome.

That evaluation discipline also supports a searchable knowledge base. Teams need retained prompts, evaluation notes, and failure examples when comparing rapidly changing model endpoints.

Ember-1 makes a credible argument against the lower-effort shortcut. It does not yet establish that one vendor's post-training recipe will generalize across every agent, repository, or professional domain.

Production Tests Make the Claim More Concrete

The customer A/B tests offer Ember-1's most practical evidence, although Fireworks has not identified the customers or released their evaluation data.

Fireworks says it tested Ember-1 with two customers using live production coding workloads. Both reportedly generated about 35% fewer tokens per task at comparable quality.

The company published more detailed figures for one comparison. The Kimi K3 arm scored 0.751, while the Ember-1 arm scored 0.753.

Average steps declined from 23.8 to 21.4. Output tokens fell from 49,300 to 29,900 per task.

Fireworks reports a 71.3% reduction in reasoning tokens and a 39% reduction in total tokens. It also says completion, success, and failure indicators generally held steady or improved.

One participating customer reportedly moved Ember-1 into live production and plans to scale it as a replacement for the base model. The customer's identity, sample size, workload composition, and evaluation rubric were not disclosed.

Those omissions prevent an independent assessment of statistical significance. A change from 0.751 to 0.753 might reflect equivalent performance, random variation, or a small improvement.

The token movement is harder to dismiss because its magnitude is much larger. Yet readers still need to know whether the averages were influenced by task length, failed runs, caching behavior, or changed stopping conditions.

Average steps provide one clue. Ember-1 used fewer steps, suggesting that some savings came from shorter trajectories rather than only shorter reasoning inside each step.

That can be beneficial when a model avoids unnecessary tool calls. It can also conceal early termination if a success metric fails to capture incomplete work.

Fireworks says unsuccessful Ember-1 attempts also use restrained token counts. This characteristic could limit spending on tasks that an agent cannot solve.

A cheap failure is not automatically a useful failure. Developers still need clear error states, trace visibility, and escalation rules so a shorter attempt does not silently pass incomplete work downstream.

Internal testing adds another signal. Fireworks routed part of its own coding and cowork traffic through Ember-1 before exposing the model to customers.

The company says its developers did not notice the switch while token consumption declined. That is a meaningful usability observation because an efficiency model should ideally feel uneventful.

It remains anecdotal. Fireworks did not publish the number of developers, duration of the internal test, task mix, or a controlled satisfaction measure.

Nevertheless, the production framing separates Ember-1 from models optimized only for a public leaderboard. Live agent workloads contain messy repositories, changing requirements, failed tools, and repeated interactions.

That is exactly where reasoning bloat becomes expensive. It is also where compressed reasoning can create hidden risks if the model skips checks that a clean benchmark does not require.

The most defensible conclusion is narrower than Fireworks' marketing line. Ember-1 produced substantial token reductions in the company's evaluations while broadly preserving measured task quality.

That result deserves attention from teams operating coding agents at volume. It does not remove the need for workload-level testing before switching a production system.

The Missing Weights Limit Independent Verification

Ember-1 inherits an open-weight foundation, but its API-only release prevents outsiders from reproducing the model or auditing the claimed training mechanism.

Fireworks calls Ember-1 its own model and the first in a planned series of specialized releases. However, the company has not published Ember-1's weights, training code, or exact algorithms.

That creates tension with the broader open-model story. Kimi K3 gives researchers and infrastructure teams more control, while Ember-1 converts the specialized derivative into a hosted service.

Developers can compare endpoint behavior, but they cannot inspect the checkpoint. They also cannot confirm whether the reported improvements survive different inference infrastructure or decoding implementations.

The preview format adds another uncertainty. Fireworks says research models receive limited serverless access and can become permanent based on community demand.

A temporary endpoint complicates adoption for teams that require stable model identifiers, repeatable evaluations, or long procurement cycles. A successful test does not guarantee continuing availability under the same conditions.

The benchmark evidence also needs careful interpretation. Fireworks ran the evaluations and selected the comparison settings.

Its Terminal-Bench and DeepSWE results appear strong, but independent leaderboard submissions would carry more weight. Repeated tests from outside groups could reveal variance, scaffold sensitivity, or workload-specific regressions.

SWE-bench Verified presents an additional problem. In February 2026, OpenAI said it had stopped reporting the benchmark because exposure was weakening its ability to measure frontier coding progress.

The benchmark contamination warning recommends newer evaluations for claims about current coding capability. Fireworks' 92.2% result remains descriptive, but it should not anchor the entire case.

DeepSWE and Terminal-Bench help diversify the evidence. Even so, no benchmark fully reproduces a production agent with private code, organization-specific tools, and business consequences.

The production A/B tests partially address that gap, yet their anonymity limits scrutiny. Fireworks has not provided task-level results, confidence intervals, or customer-authored accounts.

There is also a semantic question around "40% fewer tokens." Fireworks sometimes describes the reduction as reasoning tokens and elsewhere discusses output or total tokens.

Its detailed production example reports a 71.3% reasoning-token reduction, a 39% total-token reduction, and output falling from 49,300 to 29,900. Those metrics overlap, but they are not interchangeable.

Buyers should ask which measure applies to their workload. A model could sharply reduce hidden reasoning while leaving visible output unchanged, or reduce total output through shorter final responses.

Quality must also be defined before evaluation. Exact task completion, human preference, test passage, and business success can produce different conclusions from the same run.

Security-sensitive teams should examine whether shorter reasoning changes tool decisions. An agent that calls fewer tools might save tokens while performing fewer validations or bypassing defensive checks.

None of these concerns invalidates Fireworks' results. They define the work required to move the claim from promising vendor evidence to a repeatable industry finding.

Independent endpoint tests are possible now. Full scientific reproduction will remain impossible unless Fireworks releases the specialized weights, training method, or enough experimental detail for another group to recreate the process.

Three Signals Will Determine Whether Ember-1 Lasts

The next phase depends on independent testing, permanent availability, and evidence that token savings survive outside Fireworks' preferred workloads.

The first signal is independent replication on Terminal-Bench 2.1 and DeepSWE 1.1. Evaluators should use documented scaffolds, publish full configurations, and report variation across repeated runs.

Results close to Fireworks' scores would strengthen the efficiency claim. Large regressions or unstable token savings would suggest that Ember-1 depends heavily on the company's evaluation setup.

The second signal is Fireworks' decision about availability. A permanent Ember-1 endpoint would indicate enough demand and operational confidence to support real deployments.

A discontinued preview would not prove that the technical idea failed. It would, however, limit Ember-1's importance as a product and redirect attention toward Fireworks' training platform.

The third signal is broader production evidence. Named customers, larger task samples, or third-party case studies should report completion rates alongside total tokens and latency.

Evidence across different repositories, languages, and agent frameworks would support Fireworks' claim that trained efficiency generalizes. Savings concentrated in one coding workflow would narrow the model's value.

Competitor behavior will provide supporting context. Moonshot can improve K3's native efficiency, while other inference providers can post-train open models for shorter reasoning.

If that pattern spreads, Ember-1 will matter as an early example of a larger shift. Model buyers would compare useful work per token, not only intelligence scores or context-window size.

The approach also changes how teams should evaluate agents. A long trace should no longer be treated as evidence that the system performed deeper or better reasoning.

Verbose analysis can reflect productive verification, repeated uncertainty, or simple inefficiency. Only task outcomes and controlled testing can distinguish among them.

Fireworks AI Ember-1 presents a focused and plausible response to that problem. The company has produced meaningful benchmark and production evidence, while keeping enough details private to prevent full verification.

Developers should test the model against K3 Max and K3 Low on their own task distribution. They should preserve the same prompts, tools, stopping rules, and scoring system across all arms.

Track failures as carefully as token totals. Examine whether Ember-1 skips validations, exits difficult tasks earlier, or changes the severity of mistakes.

The final question is practical: does Fireworks AI Ember-1 finish your real work with fewer tokens, without moving risk into less visible places? Run that comparison before the preview ends, and retain enough evidence to repeat it later.

Give every agent the context to do better work

Connect your agents to the knowledge, decisions, and history already organized in remio.

remio currently supports Windows 10+ (x64) and Macs with Apple silicon.

Your AI Partner at Work
Get more done with remio

Plan. Create. Deliver.
All in one place.

bottom of page