Give it an identity
Choose a face, define personality, and assign a voice that sounds unmistakably yours.
Video, voice, and real-time conversation for any being you can imagine, humans, mascots, game avatars, and brand personalities. Built in one studio. Shipped anywhere.
Built for production
Proof, not just a demo.
Voice, video, and conversation, generate, preview, and ship from a single screen. No tabs. No context switching. Just flow.
Welcome to Sunesis, I can see, hear, and talk back in real time.
Build a character in seconds, choose its face, its look, and its voice, then start a real video call right in your browser. It listens with live speech recognition, replies out loud with natural text-to-speech, and holds the whole conversation in a running chat thread. No clips, no pre-renders. Press the mic and say hello.

Choose a face, define personality, and assign a voice that sounds unmistakably yours.
Open a low-latency voice or video call that listens, responds, and takes turns naturally.
Connect knowledge, memory, and tools so every conversation can lead to an outcome.
Ship on the web, in your product, or over the phone from one real-time API.
Explore voice, live interaction, and developer tools when this section comes into view.
Choose the path that fits your team today. Add characters, agents, and developer tooling later without changing the foundation underneath.

Move from a brief to a character, voice, or guided conversation without assembling five separate tools.
One workspace · shared credentials · usage you can see.
Build speech, cloning, avatars, and live transcription into one API workflow. TypeScript, Python, and Go SDKs are ready out of the box, with sub-200ms response times.
No sales call required. Grab your key, paste the snippet, and ship.
import { SunesisClient } from "@sunesis/sunesis-js"; const sunesis = new SunesisClient(); // Stream lifelike speech, token by token const audio = await sunesis.textToSpeech.convert({ text: "Welcome to Sunesis.", voiceId: "aria", stream: true, });

Lip-syncable, emotion-aware TTS for AI avatar products. Your digital character speaks with natural prosody and visible emotion, video + voice fused.

Sub-200ms turn-taking over WebSocket. Streaming TTS and ASR so your agent can listen, think, and respond without the awkward pause.

Notes-to-audio, narration, and AI companions that actually talk back. Per-character pricing that scales with usage, not seats.

Voice cloning from 15 seconds of audio. Persistent memory, function calling, and RAG, build agents that remember and act.
Create the character and voice together, then deliver a response quickly enough to keep the moment moving.
Character foundation
Set the role, tone, voice direction, and context a character needs before the conversation starts.
First response
Audio begins almost immediately, so a reply does not feel like a loading state.
Languages
Let the same character carry a consistent voice from one language to the next.
Voice sample
Use a short, approved sample to create a voice a character can keep.
Start with speech, a character, a live session, or an agent, and connect them when the experience calls for more. The same workspace, credentials, webhooks, usage view, and SDK conventions carry across every layer.
Read the platform docsGenerate, stream, or tailor a voice with delivery controls that fit the moment.
Script or audio to a speaking character production.
Run low-friction voice and video exchanges with turn-by-turn events.
Grounded conversations with memory, tool access, and an optional on-screen presence.
TypeScript, Python, Go, REST, and WebSocket examples that match production behavior.
Scoped API keys, audit-friendly workspace controls, and predictable error handling.
Talk to the engineers building the platform when you are designing at scale.
Talk to usStart with a single capability, then bring voice, characters, agents, and developer tooling together on the same Sunesis foundation.

Clone your voice and narrate anything.

Face-to-face with a digital character.

Voice agents that truly listen.
Click any face to hear them speak.
240+ voices, infinite possibilities, voices uploaded and built by the Sunesis community.
From new products finding their voice to public companies redesigning a customer journey, teams use Sunesis to make the interaction count.
Voice agents“We replaced our entire IVR with Sunesis agents. Call deflection jumped overnight and our support costs dropped by 64%.”

Mara Lindqvist
VP Engineering, Northwind
Digital characters“Onboarding completion went from 41% to 89% once we added a face-to-face digital character to the flow.”

Dev Patel
Head of Product, Koru Health
Real-time API“We ran 40k concurrent voice sessions during launch week without a single dropped call. The streaming latency is unreal.”

Sofia Marchetti
Staff Engineer, Helio
Creators and developers share the moments that make a voice experience feel different.

“Cloned my voice in 15 seconds and it nailed my cadence on the first try. Wild.”

“The streaming TTS feels instant, my viewers think I'm live when I'm not.”

“Built a customer support agent in an afternoon. It's better than our offshore team.”

“Finally a voice API that doesn't sound like a robot from 2015.”
Explore credits, production features, and team options on the pricing page, without interrupting the product story.
Every layer is measured for the moment between a person speaking and the system responding.
Start with the product, plan, or team information that matters to you.
Sunesis Voice v2 speaks 32 languages out of the box with any of our 240+ voices. The API supports 30+ languages including English, Spanish, Japanese, French, German, Mandarin, and more, all with native-quality prosody and the ability to clone a voice and use it across languages.