Give it an identity
Choose a face, define personality, and assign a voice that sounds unmistakably yours.
Video, voice, and real-time conversation for any being you can imagine, humans, mascots, game avatars, and brand personalities. Built in one studio. Shipped anywhere.
Built for production
Proof, not just a demo.
Voice, video, and conversation, generate, preview, and ship from a single screen. No tabs. No context switching. Just flow.
Welcome to Sunesis, I can see, hear, and talk back in real time.
Build a character in seconds, choose its face, its look, and its voice, then start a real video call right in your browser. It listens with live speech recognition, replies out loud with natural text-to-speech, and holds the whole conversation in a running chat thread. No clips, no pre-renders. Press the mic and say hello.

Choose a face, define personality, and assign a voice that sounds unmistakably yours.
Open a low-latency voice or video call that listens, responds, and takes turns naturally.
Connect knowledge, memory, and tools so every conversation can lead to an outcome.
Ship on the web, in your product, or over the phone from one real-time API.
Explore voice, live interaction, and developer tools when this section comes into view.
Choose the path that fits your team today. Add characters, agents, and developer tooling later without changing the foundation underneath.
Move from a brief to a character, voice, or guided conversation without assembling five separate tools.
Ground an agent in your knowledge, give it a memory, and let it hand off with the context your team needs.
Use the same credentials, events, usage controls, and SDK conventions across speech, characters, live rooms, and agents.
One workspace · shared credentials · usage you can see.
Fast to start. Open by design.
Build speech, cloning, avatars, and live transcription into one API workflow. TypeScript, Python, and Go SDKs are ready out of the box, with sub-200ms response times.
import { SunesisClient } from "@sunesis/sunesis-js"; const sunesis = new SunesisClient(); // Stream lifelike speech, token by token const audio = await sunesis.textToSpeech.convert({ text: "Welcome to Sunesis.", voiceId: "aria", stream: true, });
Four production paths. One API key, one event model, and one usage view underneath them all.
Lip-syncable, emotion-aware TTS for AI avatar products. Your digital character speaks with natural prosody and visible emotion, video + voice fused.
Sub-200ms turn-taking over WebSocket. Streaming TTS and ASR so your agent can listen, think, and respond without the awkward pause.
Notes-to-audio, narration, and AI companions that actually talk back. Per-character pricing that scales with usage, not seats.
Voice cloning from 15 seconds of audio. Persistent memory, function calling, and RAG, build agents that remember and act.
Create the character and voice together, then deliver a response quickly enough to keep the moment moving.
You define

Aria
Warm · Female · English
Everything above travels together. Change the voice and the character keeps its role; change the role and it keeps its voice.
Start with speech, a character, a live session, or an agent, and connect them when the experience calls for more. One workspace, one set of credentials, one usage view.
Generate, stream, or tailor a voice with delivery controls that fit the moment.
DocsTurn a script or an audio track into a speaking character on camera.
DocsRun live voice and video exchanges with turn-by-turn events you can act on.
DocsGrounded conversations with memory, tool access, and an optional on-screen presence.
DocsTypeScript, Python, Go, REST, and WebSocket examples that match production behavior.
Scoped API keys, audit-friendly workspace controls, and predictable error handling.
Talk to the engineers building the platform when you are designing at scale.
Talk to usStart with a single capability, then bring voice, characters, agents, and developer tooling together on the same Sunesis foundation.
A shared workspace for shaping voice, characters, media, and live interaction.
ExploreConversational roles with memory, grounded knowledge, and carefully scoped actions.
ExploreOne developer foundation for generation, real-time sessions, webhooks, and typed SDKs.
Explore


Real-time personalities with visual presence, a signature voice, and agent intelligence.
ExploreClone your voice from one short take, then narrate anything — and keep that same voice across every language you publish in.






Put a face on the conversation. A visible character that listens, answers, and holds the thread in real time.
One agent across phone, chat, and video — listening properly, and handing over with full context when a person should take it.
Click any face to hear them speak.
240+ voices, infinite possibilities, voices uploaded and built by the Sunesis community.
| Capability | The others | Sunesis Labs |
|---|---|---|
| Latency | 500ms+ latency | 180ms latency |
| Voice quality | Robotic voices | MOS 4.6 naturalness |
| Memory | No memory | Persistent memory |
| Pricing / lock-in | Vendor lock-in | Open API, no lock-in |
| Regions / uptime | Single-region | 14 regions · 99.99% uptime |
Trusted by teams shipping voice and video in production
From new products finding their voice to public companies redesigning a customer journey, teams use Sunesis to make the interaction count.
“We replaced our entire IVR with Sunesis agents. Call deflection jumped overnight and our support costs dropped by 64%.”

Mara Lindqvist
VP Engineering, Northwind
“Onboarding completion went from 41% to 89% once we added a face-to-face digital character to the flow.”

Dev Patel
Head of Product, Koru Health
“We ran 40k concurrent voice sessions during launch week without a single dropped call. The streaming latency is unreal.”

Sofia Marchetti
Staff Engineer, Helio
Creators and developers share the moments that make a voice experience feel different.

“Cloned my voice in 15 seconds and it nailed my cadence on the first try. Wild.”

“The streaming TTS feels instant, my viewers think I'm live when I'm not.”

“Built a customer support agent in an afternoon. It's better than our offshore team.”

“Finally a voice API that doesn't sound like a robot from 2015.”
Explore credits, production features, and team options on the pricing page, without interrupting the product story.
Every layer is measured for the moment between a person speaking and the system responding.
Live transcription as they talk
Grounded reasoning and tools
Streaming speech, first token
Character video, lip synced
Voice in to response out
180msmedian, measured in production
Streaming voice · character video · grounded agents
Start with the product, plan, or team information that matters to you.
Sunesis Voice v2 speaks 32 languages out of the box with any of our 240+ voices. The API supports 30+ languages including English, Spanish, Japanese, French, German, Mandarin, and more, all with native-quality prosody and the ability to clone a voice and use it across languages.