GPT-Live's Interaction Ingenuity and Flaws

2177402732-0817f6ab820987bd8459e3b476d1dc1cca387fbf19a13dc85f84889b1c90fed8-d.webp

I've been playing with ChatGPT-Live for the past few days.

At first, I told GPT a few books I really liked and asked it to recommend similar works based on my reading preferences. The results were surprisingly good.

I had previously thought about building a small tool to recommend books based on reading preferences. Now it seems I don't need to build that tool anymore—I even save on tokens. Just open Live, name a few books I enjoy, and it can follow up with recommendations.

And it doesn't just generate a book list; you can continue discussing the content and ideas from those books with it, basically having a back-and-forth conversation. As long as the topic isn't too complex, this real-time voice interaction feels much more natural than typing.

But after a few days of use, I've come to realize that the most interesting aspect of GPT-Live and its most obvious flaw might come from the same interaction design.


It actively interrupts you.

Whenever I keep talking continuously, GPT-Live inserts interjections like "mm-hmm," "right," or "got it" in the middle.

This isn't a random verbal tic.

OpenAI specifically mentions in the introduction to GPT-Live that it uses short responses like "mm-hmm," "yeah," and "got it" to let users know the model is listening, and to make the voice conversation feel more like a real human dialogue. GPT-Live adopts a full-duplex architecture, which allows it to generate sound while continuing to listen to the user, and continuously determine when it should speak, wait, pause, or interrupt.

From an interaction experience perspective, this design is indeed effective.

During real conversations, we use nods, eye contact, or interjections to maintain the flow. GPT-Live incorporates this natural human behavior into the product.

But I suspect that beyond anthropomorphic interaction, these interjections also serve the purpose of actively controlling the length of each user input.

Whenever GPT suddenly inserts an "mm-hmm" or "got it" while you're speaking continuously, you instinctively pause for a moment, judging whether it's about to start answering. What was originally a continuous expression ends up being broken into many shorter input segments.

Checking the conversation text after exiting Live, it's clear that these interjections are inserted in the middle of what was originally my continuous speech.

This might also be one of the reasons GPT-Live responds so quickly.

It doesn't need to wait for the user to finish a long stretch of content before starting to process and generate a response. It can actively confirm, segment, and take over the conversational turn during the exchange, splitting one long input into multiple short rounds.

For the model, processing shorter inputs is easier, and generating short replies is faster. For the user, the model always seems to be responding, creating a sense of almost no delay.

So, GPT-Live's fast response likely comes not only from the model's generation speed but also from its active control over the conversation's rhythm.

It uses interjections to shorten individual inputs, making the interaction more anthropomorphic while reducing the amount of information each round needs to process.

This is a clever interaction design.


The cost of real-time feel is context.

But the problem with this design is also obvious. GPT-Live can quickly understand the last few sentences but has difficulty stably comprehending a long, continuous expression.

When I try to fully explain a slightly complex idea, it constantly inserts interjections in the middle. Sometimes I'm forced to stop; sometimes I continue, but its subsequent response often only captures the last few keywords, while the preceding conditions and reasoning process have disappeared.

This drift in attention is even more noticeable in long text conversations.

Once the same conversation already has a large amount of text input and output, switching to Live makes it hard to pick up on that content. It may completely fail to follow up on issues discussed earlier, only answering based on the most recent short exchanges.

Live also cannot carry over context across sessions. Opening a new session to enter a Live conversation means the established discussion state is lost—GPT's celebrated cross-session memory does not apply.

In the ChatGPT App's Project conversations, this performance limitation is confirmed. Voice can be enabled in Projects, but only the older version of Voice is available; you cannot switch to Live.

The main value of Projects is to centrally store files, project instructions, and historical conversations, allowing ChatGPT to continuously use this context.

GPT-Live's current capability conflicts with this heavy-context scenario.

It excels at handling ongoing short conversations but does not do well with carrying over large amounts of existing text, nor can it stably track information from earlier rounds. Its understanding of "context" is closer to the last few sentences rather than the complete state accumulated over the entire conversation.

OpenAI's positioning of GPT-Live also reveals its current priority.

The official focus is on natural turn-taking, interruptions, pauses, quick responses, and full-duplex communication. When encountering web searches, deeper reasoning, or more complex tasks, GPT-Live hands them over to the backend frontier model and brings the results back into the voice conversation. OpenAI also states that it hopes to support more complex, longer-lasting, and more agent-like work with voice in the future.


"More complex, longer-lasting" remains a future goal.

For now, GPT-Live's primary concern is making the model sound like it's having a real-time conversation with you.

It can interject, interrupt, use interjections to show it's still listening, and improve response efficiency by shortening conversational rounds. These designs make voice interaction noticeably more natural, but they don't solve the issues of long texts, cross-session context, or long-term context.

So, in its current state, GPT-Live is better suited for fragmented free talk.

You can use it to chat about books, practice speaking, discuss a spontaneous idea, or rapidly go back and forth on a simple topic. As long as the conversation requires referencing a lot of historical information, understanding long texts, or maintaining a complex task state, it tends to only remember the latest few sentences.

The cleverness of GPT-Live lies in shortening conversations to make the AI seem more human. But its current flaw is also here: conversations become faster, but the content that can be fully understood becomes less.

From GPT-Live's implementation, traditional TTS models are already obsolete. In the future, it should fully adopt an agent-based approach.

Let's talk

Tell me what you think