Thinking Machines Lab, the AI startup founded by former OpenAI CTO Mira Murati, announced a research preview of its interaction models, a new class of multimodal systems designed for near-realtime voice and video conversation. The models move beyond traditional turn-based interaction by using a full-duplex architecture that processes input and output in parallel. This allows the AI to listen while speaking and react to visual cues in real time. The initial model, TML-Interaction-Small, is a 276-billion parameter mixture-of-experts model with 12 billion active parameters. It achieved a turn-taking latency of 0.40 seconds on internal benchmarks, outperforming competitors like Google's Gemini and OpenAI's Realtime API. The system also uses a background model for deep reasoning tasks, streaming results back to the interaction model. The company plans to open a limited research preview in the coming months, with a wider release later this year. This development signals a shift from static AI interactions to more fluid, human-like conversations, with implications for customer service, real-time monitoring, and collaborative workflows.
Thinking Machines introduced a new class of AI models that treat interactivity as a first-class citizen of model architecture, enabling full-duplex real-time voice and video conversation with sub-second latency.
Unchanged: Traditional turn-based AI chat models (text, voice) remain the standard and are still widely used. The interaction models are not yet publicly available.
The article conveys cautious optimism: the technical achievement is impressive but the models are not yet publicly available, and there are uncertainties about safety and adoption.
Advances state-of-the-art in real-time multimodal interaction, pushing AI capabilities forward.
Demonstrates innovation from a well-funded startup, encouraging investment in AI startups.
Creates new business opportunities for real-time AI applications and strengthens Thinking Machines' market position.
Announced breakthrough interaction models, strengthening its position as an AI innovator.
Former OpenAI CTO, co-founder of Thinking Machines, leading high-profile AI research.
Competitor whose GPT-realtime-2.0 lags in latency and interaction quality.
Gemini-3.1-flash-live has higher latency; competition may pressure Google.
Partner for deploying the models on Vera Rubin systems; benefits from increased AI compute demand.
Lead investor in Thinking Machines, likely to see returns if models succeed.
Attempted to acquire Thinking Machines and hired several employees, but also lost talent to them.
This development moves AI from turn-based to continuous interaction, making AI more natural and effective in real-time tasks. It could redefine how humans collaborate with AI in areas like customer service, manufacturing, and healthcare. The technical approach of native interactivity may set a new standard for multimodal AI. However, widespread adoption depends on reliability, safety, and cost.
New API for real-time AI interaction will enable novel applications like proactive AI assistants and real-time monitoring tools.
Improved customer service bots and real-time process monitoring could reduce latency and increase efficiency.
Demonstrates progress in AI research and potential for new market applications, supporting high valuations.
No immediate impact as models are not yet public; long-term benefit depends on availability and use cases.
Thinking Machines jumps ahead in real-time interaction quality, pressuring competitors like OpenAI and Google to accelerate their own developments.
Thinking Machines is US-based; the technology may drive innovation and investment in US AI sector.
Impact is global as real-time AI applications can be deployed worldwide, but regulatory differences may affect adoption.
New attack surface for adversarial inputs in real-time streams.
Real-time processing of voice/video raises significant data handling and privacy concerns.
If models behave unsafely, it could damage company reputation.
Scaling from research preview to reliable product is challenging.
Cloud infrastructure is sufficient; no new risks identified.
No immediate geopolitical tension; but AI leadership competition may increase.
Real-time voice/video AI may face strict regulations on data privacy and consent.
Dependence on Nvidia hardware could be a bottleneck.
May replace some customer service jobs but also create new AI-related roles.
Real-time AI actions in physical domains could cause harm; liability unclear.