Thinking Machines Lab has published a research preview of its first AI model, called the Interaction Model. The model is designed to break out of the traditional question-and-answer pattern of voice AI by processing audio, video, and text in parallel 200-millisecond chunks. The startup claims it outperforms OpenAI's GPT-Realtime-2 and Google's Gemini Live on interaction quality. The core innovation is time-aligned micro-turns: the model continuously processes and generates 200ms tokens in an interleaved fashion, eliminating artificial turn boundaries. It uses a fast interaction model paired with a background reasoning model to handle deeper tasks without breaking conversation flow. On the Audio MultiChallenge benchmark, it scores 43.4% — above fast variants of competitors but below GPT-Realtime-2 in thinking mode. Founded by former OpenAI researchers including Mira Murati, Thinking Machines Lab previously raised $2B at a $12B valuation. This model is the first in-house AI product backing Murati's ambition to compete with OpenAI, Anthropic, and Google DeepMind. If successful, it could reshape voice AI by making interactions feel more natural and interruptible, but the company faces execution risks and a rapidly evolving competitive landscape.
Thinking Machines Lab released its first AI model, the Interaction Model, a voice AI that natively handles interactivity using micro-turns instead of the traditional request-response harness used by OpenAI and Google.
Unchanged: The fundamental capabilities of large language models remain unchanged; Thinking Machines still relies on transformer architectures. Competitors' existing products continue to operate without disruption.
The news is mostly positive, highlighting a significant technical breakthrough in voice AI interactivity. However, tempered by the company's execution risks and the model's status as a research preview.
Advances the state of the art in voice AI with a novel interactive architecture, potentially making AI conversations more natural and useful.
Demonstrates that a well-funded startup can challenge incumbents on core AI capabilities, inspiring other startups in the space.
Releases first in-house AI model, demonstrating technology to back its high valuation.
Founder; her vision for interactive AI takes concrete form with this release.
Faces a new competitor claiming superior voice AI interactivity, especially in GPT-Realtime-2.
Gemini Live may be compared unfavorably; pressure to improve interactivity.
Outperformed on interaction quality by Thinking Machines' model according to their benchmarks.
Similar voice AI product from Google; claims of being less interactive.
Previous tool released by Thinking Machines; not directly related to this model.
Voice AI has largely been constrained to question-answer patterns with artificial turn-taking. Thinking Machines Lab's native interaction approach could make AI conversations feel more natural, proactive, and interruptible, unlocking new use cases in live translation, tutoring, and customer service. However, the model is still a research preview with limited benchmarks, and the company faces significant execution and competition risks. This release signals that interactivity is becoming a key differentiator in the AI race, potentially forcing established players to rethink their architectures.
Gain access to a potentially more interactive voice AI model for building conversational applications, but must evaluate performance and stability in a research preview.
May benefit from a more capable voice AI foundation for product development, but face uncertainty about the company's long-term viability and model maturity.
OpenAI and Google face increased competitive pressure in voice AI, especially if the Interaction Model proves superior in real-world conversations.
Positive for Thinking Machines Lab's validation of its technology, but concerns about recent employee departures and fundraising challenges temper enthusiasm.
Headquartered in the US, the company strengthens US leadership in AI innovation and potentially creates jobs.
Voice AI improvements can be applied worldwide, but regulatory and infrastructure differences may affect adoption.
No specific security vulnerabilities disclosed.
Processing user conversations may raise data handling and retention questions.
If model fails to meet claims, it could damage credibility.
Company has seen employee departures and fundraising challenges; turning research preview into a reliable product is uncertain.
Relies on standard cloud infrastructure.
No immediate geopolitical implications; company is US-based.
Voice AI interacting with users may face data privacy regulations (e.g., GDPR, CCPA).
No hardware supply chain dependencies mentioned.
Augments human work rather than displacing.
No high-stakes decisions made by model yet.