Fireworks made its Training API generally available. It also launched Fireworks Lab, a consulting service with embedded researchers and engineers. Customers write their own training loop in Python. Fireworks runs the compute, weight syncing and sample generation. Serverless training covers low-rank adapters on shared machines and bills per token. Dedicated training covers full-parameter runs and bills per GPU-hour. Dedicated access requires a Fireworks Lab consultation. The company says it keeps training and serving numerically aligned to avoid corrupted learning signals. Compressed weight transfers cut bandwidth between training steps. Fireworks cites customer results from Harvey, Vercel, Heidi Health and Factory.
What changed
Fireworks training was in preview, announced in April 2026.
What it unlocks
Running a custom Python training loop on managed GPU infrastructure, with checkpoints deployed straight to inference.
- 10x less weight-transfer bandwidth
- 2-4x more iterations per budget
- RL run across 10,000+ GPUs
- Harvey: 19.7% vs 10.8% all-pass
What you need to act on it
- a Fireworks account
- Fireworks Lab consultation for dedicated training
Sources