AI Models

GPT-Live vs Gemini 3.8 Live vs Muse: Voice AI War

Aditya Kumar JhaAditya Kumar JhaLinkedInAmazon·September 16, 2026·11 min read

GPT-Live, Gemini 3.8 Live, and Muse Voice Transcribe all launched in September 2026 - here's how their real prices and latency claims compare.

Real-time voice AI had its busiest stretch of 2026 between September 1 and September 16. Meta Superintelligence Labs, OpenAI, and Google each shipped a new voice model within about two weeks of each other, and for once the fight isn't only about who sounds the most human. It's about who can respond fast enough to feel like a real conversation, and who can afford to let you talk to it all day.

OpenAI opened its GPT-Live-1 model to developers on September 10, 2026, at $0.05 per minute of voice. Five days earlier, Meta AI Research had introduced Muse Voice Transcribe, a real-time transcription and diarization model built for a different job entirely. Google followed on September 15 with Gemini 3.8 Live and a heavier Gemini 3.8 Live Extended Thinking model that Google says beats rival voice models on quality, at what independent cost trackers describe as a fraction of GPT-Live-1's price. None of the three companies has published a directly comparable, independently verified latency number, and that gap between marketing copy and measured reality is most of this story.

Insight

Quick summary: Between September 1 and September 15, 2026, three AI labs shipped competing voice models. OpenAI's GPT-Live-1 reached its developer API on September 10 at $0.05/minute for voice alone, roughly $4.47-$5.83/hour all-in once a backend reasoning model is attached. Google's Gemini 3.8 Live followed on September 15 at $0.005/minute audio-in plus $0.018/minute audio-out; independent cost benchmarking puts it at roughly a sixth to half of GPT-Live-1's all-in hourly rate. Meta Superintelligence Labs' Muse Voice Transcribe, released September 1, isn't a talking assistant at all - it's a streaming transcription and speaker-diarization model that processes audio in 80-millisecond chunks and reports a 3.1% word error rate. None of the three vendors has published an independently audited, apples-to-apples end-to-end latency figure. The one independent test available, from Agora Media Lab, found GPT-Live's real-world response latency closer to 1.3 seconds median - not the sub-300ms figure that circulates in casual coverage.

What Each Company Actually Shipped

OpenAI: GPT-Live goes from ChatGPT feature to API product

OpenAI's GPT-Live story actually started on July 8, 2026, when the company replaced ChatGPT's Advanced Voice Mode with GPT-Live-1 and a smaller GPT-Live-1 mini, both full-duplex models that listen and speak at the same time so users can interrupt mid-sentence. OpenAI said at the time that more than 150 million people already talk to ChatGPT using Voice and Dictation. Sources: [Introducing GPT-Live](https://openai.com/index/introducing-gpt-live/), [TechCrunch](https://techcrunch.com/2026/07/08/openai-releases-new-voice-models-for-more-natural-live-conversations/). GPT-Live-1 mini became the default for free ChatGPT users, with the larger GPT-Live-1 reserved for paying tiers; by September it was reportedly the default voice model across ChatGPT Go, Plus, and Pro.

The bigger news came on September 10, when OpenAI opened GPT-Live-1 to developers through the API at $0.05 per minute of voice, billed per second, with backend reasoning and tool calls billed separately on top of that. OpenAI says the model posts a roughly 30-point jump on its internal Full Duplex Bench versus the prior GPT-Realtime-2.1 - but per OpenAI's own announcement, only that relative gap is published, not the two absolute scores, and OpenAI has not published any absolute turn-taking or end-to-end latency figure for GPT-Live-1 at all. It ships with 12 voices out of the box, though custom voices require contacting OpenAI's sales team, and one enterprise customer running a phone-based support/sales line (the kind of setup used for reservations or claims calls) told OpenAI that moving off a cascaded speech-to-text-to-speech pipeline onto GPT-Live-1 removed roughly 23,000 lines of code that had handled caller interruptions. Sources: [DataNorth](https://datanorth.ai/news/openai-launches-gpt-live-1-in-the-api), [AlphaSignal](https://alphasignal.ai/news/openai-s-gpt-live-1-tops-voice-benchmark-by-splitting-speech-from-reasoning).

OpenAI still has not published an official end-to-end latency figure for GPT-Live-1 itself. The only public number in the sub-second range attached to OpenAI's voice stack - a 300-to-600-millisecond median time-to-first-audio-chunk - describes the older GPT-4o Realtime v2 model, not GPT-Live. When Agora Media Lab independently measured GPT-Live inside the ChatGPT app, it found a median response latency of roughly 1.3 seconds, about 205 milliseconds faster than the old Advanced Voice Mode's 1.51-second median, but nowhere near sub-300ms, and interruption handling that was actually somewhat slower than the previous generation. A separate developer benchmark put GPT-Live-1's time-to-first-audio at 1.24 to 1.34 seconds. Sources: [Agora](https://www.agora.io/en/blog/openai-didnt-publish-gpt-lives-latency-so-we-measured-it/), [OrcaRouter](https://www.orcarouter.ai/blog/gemini-3-8-live-vs-gpt-live-1).

Google: Gemini 3.8 Live undercuts on price, tops a quality leaderboard

Google shipped Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking on September 15, rolling them out through the Gemini API and Google AI Studio, with Google's materials also referencing an enterprise/Vertex agent deployment path. The standard model is pitched at scale and cost efficiency; Extended Thinking adds multi-step reasoning for more complex voice-driven tasks. Extended Thinking topped Artificial Analysis' Speech-to-Speech Quality Index with a score of 82.6 and posted 97.7% on the Big Bench Audio reasoning benchmark. Source: [MarkTechPost](https://www.marktechpost.com/2026/09/15/google-releases-gemini-3-8-live-and-3-8-live-extended-thinking-for-production-grade-voice-agents/).

On price, Google's own documentation lists Gemini 3.8 Live at $0.005 per minute of audio input and $0.018 per minute of audio output. Using Artificial Analysis' Big Bench Audio cost methodology - a standardized hour of input audio - independent trackers vary: OfficeChai's write-up of the Big Bench Audio cost subset puts Gemini 3.8 Live at roughly $0.84 per hour and Extended Thinking at about $3.50 per hour, while OrcaRouter's own cost-per-hour-of-input-audio comparison puts standard Gemini 3.8 Live closer to $1.50 per hour - both well under the $5.83 per hour reported for a comparably configured GPT-Live-1 setup (labeled "GPT-Live-1 Astra" in that benchmark) and the $4.80 per hour reported for Grok's Voice Think Fast 2.0. That puts standard Gemini 3.8 Live at roughly a sixth to a quarter of GPT-Live-1's all-in cost depending on the source, and even Extended Thinking, the model actually competing with GPT-Live-1 on quality, still comes in noticeably cheaper. Sources: [OfficeChai](https://officechai.com/ai/google-releases-gemini-3-8-live-extended-conversational-model-claims-better-performance-than-gpt-live-1-astra-and-grok-voice-think-fast-2-0-at-lower-price/), [OrcaRouter](https://www.orcarouter.ai/blog/gemini-3-8-live-vs-gpt-live-1).

As with OpenAI, Google hasn't published an independently verifiable, millisecond-level end-to-end latency figure for Gemini 3.8 Live. Google's own materials describe the model as built for "ultra-low latency audio-to-audio interactions" without attaching a specific number, and no third-party lab has yet run the kind of head-to-head latency test against GPT-Live-1 that Agora ran for OpenAI's model. Gemini 3.8 Live also adds asynchronous function calling mid-conversation, near-real-time visual grounding, and automatic language detection across 97 languages.

Meta: Muse Voice Transcribe is a different product entirely

Meta Superintelligence Labs' Muse Voice Transcribe, which rolled out around September 1 to Meta AI for Mac, Muse Code, and the Meta Model API, isn't a conversational voice assistant at all - it doesn't generate speech. It's a real-time audio perception model: streaming speech-to-text with speaker diarization for 20-plus speakers, endpoint detection, and code-switching between languages mid-sentence. Meta says it processes audio in 80-millisecond chunks (12.5 Hz), converting each chunk into a "soft token," with an adaptive-delay mechanism that trades a little extra latency for accuracy on harder words. Sources: [Meta AI Research](https://research.meta.ai/blog/introducing-muse-voice-transcribe), [9to5Mac](https://9to5mac.com/2026/09/01/meta-launches-muse-voice-transcribe-for-real-time-voice-dictation-on-mac/).

On Meta's own published benchmarks, Muse Voice Transcribe posts a 3.1% streaming word error rate - the lowest of the systems Meta tested - and reaches roughly 3.0% error at about 0.16 seconds to final transcript. Diarization error rate averages about 17.5% across the AMI-IHM, AMI-SDM, and VoxConverse test sets. The model trained on more than 70 languages, with 25 "extensively verified" at launch. Meta says it ranks first on Artificial Analysis' streaming speech-to-text leaderboard and on public diarization benchmarks - claims that, unlike OpenAI's and Google's latency figures, come with concrete, checkable numbers, because transcription accuracy is far easier to benchmark objectively than "how natural does this feel."

That makes Muse Voice Transcribe less a direct rival to GPT-Live and Gemini 3.8 Live than a building block other companies could use to build their own voice agents. It's the ears, not the mouth. It's still part of the same story: Meta, OpenAI, and Google are all racing to own the audio layer sitting underneath every AI chat product, and whoever controls the fastest, cheapest, most accurate version of it has leverage over everyone building on top of it.

Vendor Claims vs. What's Actually Confirmed

Insight

The honest scoreboard: OpenAI has published relative benchmark gains, like Full Duplex Bench and turn-taking scores, but no absolute end-to-end latency number for GPT-Live-1 - and the one independent measurement available found real-world response times around 1.3 seconds, not sub-300ms. Google has published pricing and quality-benchmark scores but no absolute latency number either, and no independent lab has yet stress-tested Gemini 3.8 Live the way Agora tested GPT-Live. Meta's numbers are the most independently checkable of the three, because word-error-rate and diarization-error-rate are standardized, reproducible metrics - but Meta is also the only one of the three grading its own homework so far; the "ranks first" claims are Meta's, not a neutral third party's.

Pro Tip

When a company publishes a relative benchmark gain, such as "30 points better than our last model," without an absolute number attached, treat it as marketing copy until a neutral third party reproduces it. Relative gains can be entirely real even when the underlying absolute performance is unremarkable.

ModelMakerLaunchedPrice (vendor-reported)Latency (vendor-reported)Independently confirmed?
GPT-Live-1OpenAIJul 8 in ChatGPT; Sep 10 in API$0.05/min voice-only; ~$4.47-$5.83/hr all-inNo absolute figure; 0.798s turn-taking vs. prior modelNo - independent test found ~1.3s median response
Gemini 3.8 LiveGoogleSep 15, 2026$0.005/min in + $0.018/min out (~$0.84-$1.50/hr, source-dependent)"Ultra-low latency"; no ms figure publishedNo independent latency test found yet
Gemini 3.8 Live Extended ThinkingGoogleSep 15, 2026~$3.50/hrSame as above, plus reasoning overheadQuality confirmed via Artificial Analysis (82.6 score); latency not independently tested
Muse Voice TranscribeMeta Superintelligence Labs~Sep 1, 2026$3 per 1,000 audio-minutes (~$0.18/hr)80ms audio chunks; ~0.16s to final transcriptPartially - Meta's own published WER/DER benchmarks, not yet a neutral third party

What This Means If You're Choosing a Voice Assistant Right Now

  • If you're a consumer picking between ChatGPT and Gemini for everyday voice chat, the practical experience today is shaped more by which app you're already in and which subscription you're paying for than by launch-day benchmark numbers. Both GPT-Live-1 and Gemini 3.8 Live are still mid-rollout, and consumer-facing latency will keep shifting as both companies tune production infrastructure.
  • If you're a developer building a voice agent, price is the clearest, most verifiable difference right now. Gemini 3.8 Live's published per-minute rate is meaningfully lower than GPT-Live-1's, even before backend model costs are added on either side.
  • If your use case is transcription, captions, or multi-speaker meeting notes rather than a talking assistant, Muse Voice Transcribe is the one actually built for that job. GPT-Live-1 and Gemini 3.8 Live are optimized for dialogue, not batch or streaming transcription accuracy.
  • Treat every latency claim made this month, including the ones in this article, as vendor-reported until an independent lab has measured it. As of publication, only GPT-Live has been independently latency-tested by a third party, and it did not hit the sub-300ms mark sometimes attached to it in casual coverage.

This is a launch story, not a verdict; all three models are days old and still rolling out. For our broader, ongoing comparison of Claude's and ChatGPT's voice modes, including Anthropic's July 2026 update that let Claude's voice finally use Sonnet and Opus instead of only Haiku, see [Claude vs ChatGPT Voice: Which Wins in 2026?](https://lumichats.com/blog/claude-vs-chatgpt-voice-mode-2026-which-is-better). This piece is deliberately narrower: what OpenAI, Google, and Meta actually shipped this month, and what their own numbers do and don't prove.

Frequently Asked Questions
01Is GPT-Live's latency really under 300 milliseconds?

Not according to the only independent measurement available. OpenAI hasn't published an end-to-end latency figure for GPT-Live-1; the 300-600ms number sometimes cited for it actually describes the older GPT-4o Realtime v2 model. When Agora Media Lab tested GPT-Live directly, it measured a median response latency of about 1.3 seconds.

02Is Gemini 3.8 Live actually cheaper than GPT-Live-1?

By every pricing comparison we found, yes. Google lists $0.005/minute for audio input and $0.018/minute for audio output; using a standardized cost benchmark, independent trackers put Gemini 3.8 Live at roughly a sixth to a quarter of GPT-Live-1's all-in hourly cost depending on methodology, and even the pricier Extended Thinking variant still undercuts it.

03Does Muse Voice Transcribe compete with ChatGPT Voice or Gemini Live?

Not directly. It's a transcription and speaker-diarization model, not a conversational voice assistant - it converts speech to text with speaker labels in real time but doesn't generate speech back. It's more comparable to tools like Whisper than to GPT-Live or Gemini Live.

04Which one should I actually use right now?

For everyday voice chat, it mostly depends on which app and subscription you're already using - ChatGPT's GPT-Live-1 or Gemini's Live mode. For building a voice agent on a budget, Gemini 3.8 Live's published pricing is currently the cheapest of the three. For transcription or meeting notes, Muse Voice Transcribe is purpose-built for that job.

05Where can I read a full head-to-head on Claude vs ChatGPT voice?

See our evergreen comparison, Claude vs ChatGPT Voice: Which Wins in 2026?, which covers Anthropic's July 2026 voice upgrade alongside ChatGPT's desktop voice rollout in more depth than this news piece does.

If you're trying to decide which of these models is worth your time, or your subscription budget, LumiChats lets you compare and chat with ChatGPT, Gemini, Claude, Grok, and other leading models side by side in one place, so you can judge voice latency, pricing, and quality for yourself instead of taking any single vendor's launch-day claims at face value.

Read Next

Or try LumiChats to access 40+ AI models in one place — including Claude Sonnet 4.6 and GPT-5.4 — and get your questions answered today.

Was this article helpful?

Found this useful? Share it with someone who needs it.

Free to get started

Claude, GPT-5.4, Gemini —
all in one place.

Switch between 40+ AI models in a single conversation. No juggling tabs, no separate subscriptions. Pay only for what you use.

Start for free No credit card needed
Aditya Kumar Jha
Written by
Aditya Kumar JhaLinkedIn

Published author of six books and founder of LumiChats. Writes about AI tools, model comparisons, and how AI is reshaping work and education.

Keep reading

More guides for AI-powered students.