Skip to content
Oct 8Thu
  1. Inworld AI27

    Inworld breaks down TTS-2 voice steering and how to measure it

    Inworld AI Head of Research Aleksey Tikhonov explains how voice steering in TTS-2 works: using instructions to tell a speech model how to deliver a line. The post focuses on how to measure whether those instructions actually take effect — automated audio judging fell short in its tests, so the team turned to measuring changes in the speech itself. He also discusses why instruction wording matters and why the same direction produces different results across voices; these metrics are used to track changes between model versions, and they have their limits.

Oct 1Thu
Sep 30Wed
  1. Youxituoluo43

    Vidu S2 Powers AI Streamer Ziying's Mid-Autumn Broadcast, Drawing Over 300,000 Viewers

    During the Mid-Autumn Festival, Zaomeng Ciyuan's AI streamer Ziying took on a time-limited broadcast challenge in her live room; the real-time interaction model behind her is Vidu S2, released not long before, and the room drew as many as 20,000 viewers at one point. Ziying has no fixed script: viewers can request songs, vote on her look and send gifts, and she has to change outfits and perform while interacting, then pick up where she left off after being interrupted. The author also tested Vidu S2's Avatar and Editing: uploading a real person's photo or an anime image and adding some settings creates a digital human, and switching on the camera allows real-time outfit changes, art style changes and background swaps while filming.

    Why it matters: Using a livestream hosted by an AI streamer with no fixed script as a case study, it tests Vidu S2's real-time interaction and outfit-switching abilities, and discusses how they could be applied to game NPCs.

  2. Inworld AI Blog51

    Inworld acquires realtime voice agent platform Ultravox

    Inworld has announced the acquisition of Ultravox, a realtime voice agent platform; the Ultravox team has joined Inworld and continues to develop the platform. The first update after the acquisition is that the built-in Inworld voice in Ultravox now uses Realtime TTS-2, existing voice IDs remain usable with no code changes, and the upgrade costs nothing extra.

    Why it matters: The acquisition brings voice understanding, reasoning and voice generation into one realtime inference infrastructure, letting readers judge how the voice agent stack is consolidating.

Sep 23Wed
  1. NVIDIA GameDev39

    NVIDIA announces this month's updates for RTX game developers, including DLSS 5, ACE upgrades and RTX Kit 2026.3

    NVIDIA announced this month's updates for RTX game developers, including DLSS 5 with 3D-Guided Neural Rendering controls, speech pipeline extensions and inference framework updates for NVIDIA ACE, and RTX Kit 2026.3, which includes RTX Mega Geometry 2.0. The official tweet links to the developer blog post for this update; capabilities and integration details are in the original on developer.nvidia.com.

    Why it matters: The official post lines up DLSS 5, NVIDIA ACE and RTX Kit updates side by side, so developers can compare integration priorities and upgrade order for their own projects.

Aug 31Mon
  1. Inworld AI Blog57

    Inworld AI releases Realtime TTS-2 realtime conversational voice model

    Inworld AI has released Realtime TTS-2, a new-generation realtime conversational voice model, now fully available in the Inworld API and the Inworld Realtime API.

    Why it matters: The original post covers multi-turn audio context, natural-language voice instructions and a first-audio latency of under 200ms, which readers can use to judge whether realtime voice conversation fits their own projects.

Aug 28Fri
  1. Inworld AI Blog56

    Inworld AI open-sources the TTS Open Evaluation Toolkit

    Inworld AI has open-sourced the TTS Open Evaluation Toolkit, a customizable TTS evaluation framework for sampling TTS systems, evaluating audio offline and generating comparable reports.

    Why it matters: The post explains why the same TTS model can produce different WER in different teams' hands, and covers how to run this evaluation tool and its first set of metrics.

Jun 15Mon
  1. Inworld AI Blog47

    Building realtime conversational voice agents with Mastra and the Inworld Realtime API

    Inworld introduces @mastra/voice-inworld-realtime, which merges STT, LLM and realtime TTS into a single WebSocket session bound to a Mastra Agent.

    Why it matters: The post contrasts cascaded voice pipelines with a single WebSocket session and provides a reference CLI in under 100 lines of TypeScript, which readers can use to judge the integration cost.

Jun 10Wed
  1. Inworld AI Blog40

    Inworld cuts prices across its voice and LLM stack, by more than half for most developers

    Inworld announced price cuts across its entire voice stack, with most developers seeing reductions of more than half at each layer of text-to-speech, speech recognition, LLMs and compute. The new prices are already in effect for all developers. Realtime TTS-2 drops to about $10/1M characters, speech recognition to about $0.10/hour, and Router can connect to 100+ models with no markup on third-party models.

    Why it matters: Inworld has lowered prices at every layer of its voice stack and published tiered per-unit rates for TTS, speech recognition and LLMs, which can be used to estimate inference costs for consumer-grade applications.

May 14Thu
  1. Inworld AI Blog49

    Inworld pairs Realtime TTS-2 with Stream Vision Agents for a realtime voice agent that can see and listen

    Inworld teamed up with Stream to build a reference implementation, Crashout Buddy, combining the latest speech model Realtime TTS-2 with Stream's open-source Vision Agents framework, so it can see facial expressions, hear speech and adjust its delivery in real time.

    Why it matters: The post shows how to turn visual signals and emotional context into speech delivery instructions, which readers can use to judge whether this realtime voice pipeline can be plugged into their own projects.