Skip to content
  1. Inworld AI Blog57

    Inworld AI releases Realtime TTS-2 realtime conversational voice model

    Inworld AI has released Realtime TTS-2, a new-generation realtime conversational voice model, now fully available in the Inworld API and the Inworld Realtime API.

    Why it matters: The original post covers multi-turn audio context, natural-language voice instructions and a first-audio latency of under 200ms, which readers can use to judge whether realtime voice conversation fits their own projects.

  1. Inworld AI Blog40

    Inworld cuts prices across its voice and LLM stack, by more than half for most developers

    Inworld announced price cuts across its entire voice stack, with most developers seeing reductions of more than half at each layer of text-to-speech, speech recognition, LLMs and compute. The new prices are already in effect for all developers. Realtime TTS-2 drops to about $10/1M characters, speech recognition to about $0.10/hour, and Router can connect to 100+ models with no markup on third-party models.

    Why it matters: Inworld has lowered prices at every layer of its voice stack and published tiered per-unit rates for TTS, speech recognition and LLMs, which can be used to estimate inference costs for consumer-grade applications.

  1. Inworld AI Blog49

    Inworld pairs Realtime TTS-2 with Stream Vision Agents for a realtime voice agent that can see and listen

    Inworld teamed up with Stream to build a reference implementation, Crashout Buddy, combining the latest speech model Realtime TTS-2 with Stream's open-source Vision Agents framework, so it can see facial expressions, hear speech and adjust its delivery in real time.

    Why it matters: The post shows how to turn visual signals and emotional context into speech delivery instructions, which readers can use to judge whether this realtime voice pipeline can be plugged into their own projects.

You've reached the end