Uber said its total spending on AI has stayed flat since April even though weekly requests from its automated coding assistants grew more than ninefold since February. The company credits routing each task to the cheapest model that can handle it, capping how much text a session may consume, showing engineers the running cost in their terminal and extending its reuse of repeated prompts from five minutes to an hour. Uber had overrun its 2026 AI budget in the first quarter.
What changed
Uber had exhausted its 2026 AI budget in the first quarter, with costs rising alongside usage.
What it unlocks
A worked example of holding AI spending flat while usage grows, using task-based model routing, session token caps, visible per-session costs and longer prompt caching.
- 9.4x more weekly agent requests since February
- 34% lower cost per 1,000 requests
- 52% lower cost per session
- 70%+ of code changes from agents
Sources