AI News Feed
Market watch
Large Language Models

Fireworks AI Releases Ember-1, a Post-Trained Kimi K3 That Uses About 40% Fewer Tokens

Fireworks AI has released Ember-1, a model built by post-training Moonshot AI's open-weight Kimi K3 to reason in fewer tokens. Fireworks says it keeps K3's quality with about 40% fewer tokens, but offers it only through its serverless API as a Research Preview.

Fireworks has not released Ember-1's weights, training code or exact training algorithms, so self-hosting is not an option at present. The company says the approach differs from lowering a reasoning effort setting at inference time.

According to Fireworks, reasoning models such as Kimi K3 sometimes spend more than 90% of generated tokens on internal reasoning, a cost that compounds in multi-turn agentic workloads because each turn replays prior reasoning back to the model. Context grows roughly quadratically with the number of turns, and long traces from early turns are re-read and re-billed on every later call. Fireworks says customers wanted K3's coding capability at lower cost, and that turning down K3's reasoning effort did not solve the problem because lower effort settings gave up too much quality. The team instead trained the model to reason more efficiently.

Fireworks states that not all of K3's reasoning is waste, and that Ember-1 keeps useful self-reflection, such as revisiting an assumption or reacting to feedback, while cutting redundant reasoning and unproductive loops. The training collection spans mathematics, coding, instruction following, conversation, search, tool use and software engineering, and covers both standalone problems and extended multi-step interactions. Task and environment feedback guides on-policy planning and learning. Fireworks says it ran more than 50 training experiments and over 200 evaluations, and developed new training algorithms that it has not published. All training ran on Fireworks Serverless Training, and the company says it used its own data and no customer data.

Fireworks compared Ember-1 with Kimi K3 at three reasoning effort levels, computing cost with public Kimi K3 API pricing. In the company's published evaluations, Ember-1 scored 82.0% on Terminal Bench 2.1 against 80.9% for K3 Max, and 75.2% on DeepSWE 1.1 against 66.4% for K3 Max. It trailed slightly on SWE-bench Verified, with 92.2% against 93.2%, and on SWE-Interact, with 20.0% against 21.3%. On τ-2 Bench Airline it scored 66% against 64%. Fireworks listed cost reductions versus K3 Max of 51.9% on Terminal Bench 2.1, 23.7% on DeepSWE 1.1, 15.5% on SWE-bench Verified, 32.5% on SWE-Interact and 5.9% on τ-2 Bench Airline.

Across seven benchmarks and two customers' production traffic, Fireworks says K3's reasoning was shortened by 35 to 50% without sacrificing accuracy. On Doximity's Bedside Bench, described as a physician-validated set of 500 clinical cases, Fireworks says Ember-1 set a new cost-per-task Pareto frontier under its Specialized Intelligence Index.

Fireworks also ran live A/B tests with two customers on production coding workloads. Both saw roughly 35% fewer tokens per task at comparable quality, according to the company. In the published run, output tokens fell from 49.3K to 29.9K per task, reasoning tokens dropped 71.3% and total tokens dropped 39%. The task score was essentially unchanged at 0.753 for Ember-1 versus 0.751 for K3, and average steps fell from 23.8 to 21.4. One customer now runs Ember-1 in production.

Ember-1 costs the same per token as Kimi K3 on Fireworks: $3.00 for input, $0.30 for cached input and $15.00 for output per 1M tokens. Fireworks says the savings come entirely from generating fewer tokens. At that output rate, the A/B figures work out to about $0.74 versus $0.45 in output cost per task, a calculation that covers output only.

Editor's Summary

Fireworks AI has released Ember-1, a post-trained version of Moonshot AI's open-weight Kimi K3 that the company says matches K3's accuracy with about 40% fewer tokens, addressing the cost of lengthy reasoning traces in multi-turn agentic work. Fireworks reports benchmark and production A/B results showing token reductions of 35% to 50%, though Ember-1 trails K3 Max on some coding evaluations. The model is limited to Fireworks' serverless API as a Research Preview at K3's per-token pricing, with weights and training code withheld.