Tencent’s Gander separates real-time conversation from background agent work, promising fewer interruptions while exposing tradeoffs in accuracy and multimodal understanding.

Tencent has introduced Gander, a research model designed to keep a real-time conversation going while it performs slower, more complex work in the background. The system can process speech, images, and text together, respond to interruptions, and provide progress updates while an agent searches files, writes code, or handles another task.
The approach targets a persistent weakness in voice assistants: systems often pause, wait for a user to finish, or stop responding while they plan an action. According to reporting by The Decoder based on Tencent’s technical report, Gander improves conversational timing compared with several competitors, but gives up some task accuracy and multimodal performance in the process.
Gander divides the work between two components that Tencent’s researchers describe using human anatomy. A “cerebellum” manages the immediate exchange, deciding when to listen, speak, or stop if a user interrupts. A swappable “brain” performs more demanding reasoning and agent tasks.
The split is intended to avoid forcing one model to optimize for two conflicting requirements. Conversation needs low latency and accurate turn-taking, while tasks such as coding or file retrieval need more time for planning. The background component can reportedly be replaced with agent systems such as Codex or Claude Code without retraining the conversational model.
Gander breaks the interaction into one-second segments and uses roughly the previous two minutes of conversation as context for its timing decisions. The model does not rely on a separate speech-start and speech-stop detector, according to the technical report described by The Decoder. Users can interrupt while Gander is speaking, change the requested task, or answer follow-up questions as the work proceeds.
That architecture resembles a broader industry move toward orchestrated AI agents rather than a single model handling every stage of an interaction. OpenAI has also explored separating live conversation from background reasoning, while other research systems delegate work across multiple models.
The strongest evidence presented for Gander concerns turn-taking rather than general intelligence. On Full-Duplex-Bench v3, a benchmark for voice assistants across task scenarios, the Tencent system reportedly began speaking at the appropriate moment in all 100 tested scenarios and interrupted users in 8 percent of cases.
The Decoder reported comparison figures of 13.5 percent for GPT-Realtime and almost 48 percent for the weakest competitor in the same evaluation. These figures come from the research team’s reported benchmark results, not an independent assessment, and the benchmark is not a dedicated standard for systems that combine conversation with background agents.
Gander was weaker on task accuracy. The researchers attributed part of that gap to the way the full system is evaluated: speech recognition and generated-output errors count alongside the reasoning model’s performance. When the underlying “brain” receives text directly, it performs substantially better, according to the report.
The model also underperformed its base model on at least one audio and video understanding test. The researchers linked that result to training that prioritizes smooth conversation over precise perception, including tasks such as counting objects or locating them in an image. That limitation matters because a system marketed for continuous multimodal interaction must not only manage the dialogue but also interpret what users show and say accurately.
Tencent’s team reportedly trained Gander on about 2.7 million examples. Some examples teach the model to remain silent when there is background noise or when no one in a group appears to be addressing it. Those behaviors are important for practical deployments, where false activations can be as disruptive as delayed responses.
The researchers characterize the work as early and acknowledge that scaling the architecture remains unresolved. There is also no widely established evaluation framework for judging the combination of low-latency conversation, interruption handling, multimodal perception, and long-running task execution.
Tencent plans to release the model weights and training data after completing what it calls the open-source release process. A GitHub repository for the code and project demonstrations already exists, according to The Decoder. Until the weights and data are available, however, outside developers cannot fully reproduce the reported results or assess the system’s training choices.
The company’s broader AI activity provides context. Gander follows Tencent’s reported release of Hy3, an open language model used in products including WorkBuddy, Yuanbao, and WeChat. The company is also reported to be negotiating for a major stake in agent startup Manus, a move that would fit with Tencent’s interest in embedding agent capabilities into existing consumer platforms.
For developers building voice interfaces, Gander highlights a useful systems design principle: conversational responsiveness and task execution may be better handled by separate components. A lightweight interaction model can preserve the user’s sense of control while a more capable model works asynchronously. This could reduce the need to make users wait silently for a search, code-generation step, or workflow action.
The tradeoff is operational complexity. A split system must coordinate state between the conversational layer and the background agent, decide which interruptions should cancel work, and prevent stale results from being delivered after a user changes direction. It also needs clear status signals so users understand whether the system is listening, reasoning, waiting for data, or still executing an action.
Accuracy remains a central concern. The benchmark results suggest that fewer interruptions do not automatically produce better outcomes. Product teams deploying AI agents will need to test not only latency and turn-taking, but also transcription quality, visual grounding, task completion, recovery after interruption, and the cost of running multiple models at once.
For enterprise buyers, the planned release could make Gander more relevant as a reference architecture than as an immediately deployable product. Open weights and training data would allow teams to examine how Tencent handles noisy environments, overlapping speech, and context retention. They would also make it easier to compare the approach with commercial systems such as GPT-Realtime, Gemini Live, and Grok under consistent conditions.
The first signal will be whether Tencent releases the promised weights and training data, and whether the public package includes enough code and evaluation material for independent reproduction. Developers should also watch for results on longer tasks, where background planning and interruptions create more opportunities for state-management failures.
A second signal is whether Gander improves its audio and video understanding without losing its timing advantage. Better perception would determine whether the architecture can support genuinely multimodal assistants rather than primarily voice systems with visual input.
Finally, the market will need clearer benchmarks for continuous interaction. Results from Full-Duplex-Bench v3 are useful for turn-taking, but builders need evaluations that combine interruption handling with task accuracy, safety, reliability, and operating cost.
Gander’s significance is less about a single benchmark score than about making the control boundary between conversation and work explicit. Users expect an assistant to remain available while an action runs, but that experience depends on careful orchestration, not simply on a larger model.
Tencent’s results also show why product teams should resist treating fluid conversation as a substitute for reliable execution. Gander appears promising on timing, while its reported weaknesses in task accuracy and multimodal understanding leave open the harder question: can an assistant stay natural without becoming less dependable when the work matters?