ByteDance's Seed team introduced SeedRealtime on August 5, 2026, described as a native audio-visual full-duplex LLM that unifies audio, video and text in a single architecture for real-time interaction over continuous multimodal streams. Perception, understanding, decision-making and response generation happen inside one model rather than a cascaded pipeline, with chunked audio-visual input, streaming generation, quantization and inference optimization used to reduce latency. ByteDance reports that end-to-end human evaluations show roughly half as many conversational pacing issues versus cascaded models, plus fewer interruptions, false triggers and lower latency, and higher single-turn usability. The company says the model has been fully rolled out. Stated next steps include lower latency, more proactive perception, robustness in multi-party scenes, and tool use for booking and search.
- seed.bytedance.com2026-08-05