Skip to content
Oct 8Thu
  1. Inworld AI27

    Inworld breaks down TTS-2 voice steering and how to measure it

    Inworld AI Head of Research Aleksey Tikhonov explains how voice steering in TTS-2 works: using instructions to tell a speech model how to deliver a line. The post focuses on how to measure whether those instructions actually take effect — automated audio judging fell short in its tests, so the team turned to measuring changes in the speech itself. He also discusses why instruction wording matters and why the same direction produces different results across voices; these metrics are used to track changes between model versions, and they have their limits.

Oct 7Wed
  1. Inworld AI30

    Muse Spark 1.3 goes live on Inworld Realtime Router

    Inworld AI announced that Meta's reasoning and coding model Muse Spark 1.3 is now live on Inworld Realtime Router, with closed weights, support for a 1M token context, and pricing of $1.25 per 1M input tokens and $4.25 per 1M output tokens. According to its disclosure, in Meta's internal coding tests, this version uses about 25% fewer tokens than Spark 1.2.

Oct 2Fri
Oct 1Thu
Sep 30Wed
  1. Inworld AI Blog51

    Inworld acquires realtime voice agent platform Ultravox

    Inworld has announced the acquisition of Ultravox, a realtime voice agent platform; the Ultravox team has joined Inworld and continues to develop the platform. The first update after the acquisition is that the built-in Inworld voice in Ultravox now uses Realtime TTS-2, existing voice IDs remain usable with no code changes, and the upgrade costs nothing extra.

    Why it matters: The acquisition brings voice understanding, reasoning and voice generation into one realtime inference infrastructure, letting readers judge how the voice agent stack is consolidating.

Aug 31Mon
  1. Inworld AI Blog57

    Inworld AI releases Realtime TTS-2 realtime conversational voice model

    Inworld AI has released Realtime TTS-2, a new-generation realtime conversational voice model, now fully available in the Inworld API and the Inworld Realtime API.

    Why it matters: The original post covers multi-turn audio context, natural-language voice instructions and a first-audio latency of under 200ms, which readers can use to judge whether realtime voice conversation fits their own projects.

Aug 28Fri
  1. Inworld AI Blog56

    Inworld AI open-sources the TTS Open Evaluation Toolkit

    Inworld AI has open-sourced the TTS Open Evaluation Toolkit, a customizable TTS evaluation framework for sampling TTS systems, evaluating audio offline and generating comparable reports.

    Why it matters: The post explains why the same TTS model can produce different WER in different teams' hands, and covers how to run this evaluation tool and its first set of metrics.

Jun 15Mon
  1. Inworld AI Blog47

    Building realtime conversational voice agents with Mastra and the Inworld Realtime API

    Inworld introduces @mastra/voice-inworld-realtime, which merges STT, LLM and realtime TTS into a single WebSocket session bound to a Mastra Agent.

    Why it matters: The post contrasts cascaded voice pipelines with a single WebSocket session and provides a reference CLI in under 100 lines of TypeScript, which readers can use to judge the integration cost.

Jun 10Wed
  1. Inworld AI Blog40

    Inworld cuts prices across its voice and LLM stack, by more than half for most developers

    Inworld announced price cuts across its entire voice stack, with most developers seeing reductions of more than half at each layer of text-to-speech, speech recognition, LLMs and compute. The new prices are already in effect for all developers. Realtime TTS-2 drops to about $10/1M characters, speech recognition to about $0.10/hour, and Router can connect to 100+ models with no markup on third-party models.

    Why it matters: Inworld has lowered prices at every layer of its voice stack and published tiered per-unit rates for TTS, speech recognition and LLMs, which can be used to estimate inference costs for consumer-grade applications.

May 14Thu
  1. Inworld AI Blog49

    Inworld pairs Realtime TTS-2 with Stream Vision Agents for a realtime voice agent that can see and listen

    Inworld teamed up with Stream to build a reference implementation, Crashout Buddy, combining the latest speech model Realtime TTS-2 with Stream's open-source Vision Agents framework, so it can see facial expressions, hear speech and adjust its delivery in real time.

    Why it matters: The post shows how to turn visual signals and emotional context into speech delivery instructions, which readers can use to judge whether this realtime voice pipeline can be plugged into their own projects.