Fireworks: Ember-1
About this model
Ember-1 is a proprietary specialized reasoning model from Fireworks Research, post-trained on Moonshot AI's Kimi K3 mixture-of-experts base. It is tuned to produce shorter internal reasoning traces—roughly 35–50% fewer tokens in Fireworks' evaluations—while preserving answer quality on coding, tool use, mathematics, and agentic software-engineering tasks. The model targets production agent loops where long reasoning chains inflate both latency and per-task cost because prior turns are re-sent on every call.
Fireworks trained Ember-1 with more than 50 experiments and 200 evaluations on its serverless training stack, using a broad curriculum spanning coding, instruction following, search, and multi-turn interactions without customer data. Public reporting emphasizes cost-quality Pareto gains on agent benchmarks (Terminal-Bench 2.1, SWE-Bench Verified, DeepSWE) and live A/B tests on customer coding traffic, where token use dropped materially with comparable success rates.
Ember-1 ships as a research preview on Fireworks Serverless (model path accounts/fireworks/models/ember-1), with optional enterprise fine-tuning on the same stack. It supports very long context (about 1.04M tokens), function calling, prompt caching, and text plus image inputs on some gateways. Pricing aligns with Kimi K3 API tiers at roughly $3 per million input tokens, $0.30 per million cached input, and $15 per million output tokens.
Benchmark Scores
Technical Specs
- Parameters: 2780.0B
- Architecture: Mixture-of-Experts
- Context Window: 1,040,000 tokens
- Input Modalities: text, image
Hardware Requirements
- Compute: API only
Pricing
| Input | Output | Currency |
|---|---|---|
| 3.00 / 1M tokens | 15.00 / 1M tokens | USD |