Sesame is the voice-AI research lab behind the viral Maya and Miles conversational demo and the open-weight Conversational Speech Model (CSM) — full-duplex voice that sounds uncannily human.

Overview

Sesame is a voice-AI research company that went viral for Maya and Miles — two voice personas that hold a genuinely free-flowing spoken conversation. Unlike traditional text-to-speech, Sesame’s Conversational Speech Model (CSM) is built for full-duplex dialogue: it can be interrupted mid-sentence, picks up on your tone, and uses natural filler, breath, and micro-pauses that make the exchange feel present rather than robotic. The base CSM model is open-source on Hugging Face under Apache 2.0, so developers can self-host it. As of 2026 Sesame monetizes mainly through research partnerships and a planned voice-first wearable device, which means the web demo and open model are best understood as a preview of where conversational voice is heading.

Key Features

  • Full-duplex conversational voice — interrupts and responds to tone in real time, no turn-taking latency.
  • Maya & Miles personas — two distinct, persistent voice characters with natural pacing and emotional inflection.
  • Open-weight CSM — the 1B-parameter Conversational Speech Model is free to download and self-host under Apache 2.0.
  • Free browser demo — talk to Maya or Miles with no account required, the fastest way to feel the technology.
  • Emotionally expressive, low-latency speech — 200–300 ms responses that cross much of the ‘uncanny valley’ of voice AI.
  • Wearable roadmap — Sesame’s long-term bet is a voice-first companion device, with the demo serving as a preview of that engine.

Pricing

PlanPriceFor
Web Demo (Maya/Miles)$0Free, no account; full-duplex voice chat in the browser
Open-Source CSM$0Self-host the Apache 2.0 model on your own GPU
Future HardwareTBAVoice-first wearable device, no public pricing as of 2026

Comparison

vs. ElevenLabs: ElevenLabs is the production-grade choice for studio TTS, voice cloning, and dubbing at scale, with clear per-credit pricing. Sesame wins on live conversational realism but lags on enterprise features and a turnkey API.

vs. Hume AI: Hume’s EVI is the closest rival in emotional intelligence and empathy detection. Sesame is often praised for smoother prosody and lower latency, while Hume offers a more mature commercial API.

vs. ChatGPT: ChatGPT’s Advanced Voice Mode brings reasoning and multimodal context, but reviewers consistently find Sesame’s dedicated speech model more natural and emotionally present for pure conversation.

Compare alternatives

Side-by-side with the 3 closest alternatives.

ToolCategoryPricingVisit
Sesame (this) voiceFrom $0/mo Site ↗
ElevenLabsvoiceFree $0 · From $5/mo Site ↗
Hume AIvoice, audio, agentsFree $0 · From $3/mo Site ↗
ChatGPTchat, write, codeFree $0 · From $20/mo Site ↗
Sesame Current

Sesame is the voice-AI research lab behind the viral Maya and Miles conversational demo and the open-weight Conversational Speech Model (CSM) — full-duplex voice that sounds uncannily human.

voice
From $0/mo

Text-to-speech and voice cloning platform with some of the most natural AI voices available. Most natural TTS voices on the market Read our hands-on

voice
Free $0 · From $5/mo

Emotion-aware voice AI platform with the empathic EVI conversational interface and expressive Octave text-to-speech APIs.

voiceaudioagents
Free $0 · From $3/mo

OpenAI's flagship conversational AI assistant for chat, writing, coding, analysis, and web search. Versatile across chat, writing, coding, and research

chatwritecode
Free $0 · From $20/mo
Editor’s Review
4.2/5
Pros
  • +Most natural-sounding real-time voice most reviewers have tested — natural pauses, interruptions, and emotional tone
  • +Full-duplex conversation: you can interrupt mid-sentence and it responds to your tone, not turn-by-turn
  • +Open-weight CSM model on Hugging Face (Apache 2.0) — free to self-host for research and prototypes
Cons
  • No mature commercial API or public pricing yet — not a production-ready voice platform
  • Long-term product is hardware-focused (eyewear), not a SaaS play
  • Fewer integrations and developer tools than established voice platforms

The most human-sounding voice AI you can try today, but treat it as a preview: the demo is free and the model is open-weight, yet there's no turnkey API to build a paid product on yet. For production voice, ElevenLabs or Hume remain the practical choice.

See all reviews →