Release2026-08-31

Fireworks made its Training API generally available. It also launched Fireworks Lab, a consulting service with embedded researchers and engineers. Customers write their own training loop in Python. Fireworks runs the compute, weight syncing and sample generation. Serverless training covers low-rank adapters on shared machines and bills per token. Dedicated training covers full-parameter runs and bills per GPU-hour. Dedicated access requires a Fireworks Lab consultation. The company says it keeps training and serving numerically aligned to avoid corrupted learning signals. Compressed weight transfers cut bandwidth between training steps. Fireworks cites customer results from Harvey, Vercel, Heidi Health and Factory.

What changed

Fireworks training was in preview, announced in April 2026.

What it unlocks

Running a custom Python training loop on managed GPU infrastructure, with checkpoints deployed straight to inference.

  • 10x less weight-transfer bandwidth
  • 2-4x more iterations per budget
  • RL run across 10,000+ GPUs
  • Harvey: 19.7% vs 10.8% all-pass

What you need to act on it

  • a Fireworks account
  • Fireworks Lab consultation for dedicated training

Send this to someone who needs it

Shares the story and its sources. Nothing about you.

What does this mean for your job?

This is the story as everyone gets it. Once a week we send you the version written for your role — what changed, why it matters for the work you actually do, and one thing to try. Free while we tune it.