Microsoft opened public previews of two of its own models: a higher-quality image generator called MAI-Image-2.5-Pro and a faster, cheaper speech model called MAI-Voice-2-Flash, both with published per-use prices. The company also said its in-house models now run by default behind Bing Image Creator, PowerPoint image editing, OneDrive image editing and the Dynamics 365 Contact Center. The models are in preview rather than general availability, and the efficiency and accuracy figures come from Microsoft's own testing.
What changed
MAI-Image-2.5-Pro and MAI-Voice-2-Flash had only been shown as previews at Build, and Bing Image Creator and several Microsoft products used other models.
What it unlocks
Building image generation and voice applications on Microsoft's own image and speech models through Azure AI Foundry, choosing between a higher-quality image model and a faster, cheaper voice model.
- MAI-Image-2.5-Pro: $5 per 1M text input tokens, $8 per 1M image input tokens, $106 per 1M image output tokens
- MAI-Voice-2-Flash: $15 per 1M characters, 2x faster and 32% cheaper than MAI-Voice-2
- MAI-Image-2.5 in PowerPoint cuts GPU costs up to 84% versus GPT-Image-2
- OneDrive image editing: save rates up 26%, P95 latency down about 25%
- MAI-Transcribe-1.5 covers 58 languages with a 50% relative cut in transcription error in internal tests
What you need to act on it
- Azure AI Foundry access
- public preview status
Sources