Thevideo,voiceandconversationlayerforAI.

Video, voice, and real-time conversation for any being you can imagine, humans, mascots, game avatars, and brand personalities. Built in one studio. Shipped anywhere.

Built for production

Proof, not just a demo.

2,400+teams building
180msmedian voice latency
32languages live
SOC 2Type II
TypeScriptPython · Go
2,400+teams building
180msmedian voice latency
32languages live
SOC 2Type II
TypeScriptPython · Go
Sunesis Studio

One canvas. Every layer.

Voice, video, and conversation, generate, preview, and ship from a single screen. No tabs. No context switching. Just flow.

app.sunesis.ai/studio

Welcome to Sunesis, I can see, hear, and talk back in real time.

Narrate a storyRecord an adGuide a meditation
Sunesis Voice v2
Generate
Generate a character. Talk to it live.

A digital character that sees, hears, and talks back.

Build a character in seconds, choose its face, its look, and its voice, then start a real video call right in your browser. It listens with live speech recognition, replies out loud with natural text-to-speech, and holds the whole conversation in a running chat thread. No clips, no pre-renders. Press the mic and say hello.

  • Hears you in real time, replies out loud, no server round-trip
  • Full chat thread and live captions as they speak
  • Turn on your camera for a face-to-face call, right in the browser
1 · Face
2 · Look
3 · VoiceAria's voice · Warm · Female
Photorealistic · Voice · VideoReady
Aria
Aria · Photorealistic
Call Aria, live, right in your browser
How it connects

One character. A complete interaction layer.

Build your first character
01

Give it an identity

Choose a face, define personality, and assign a voice that sounds unmistakably yours.

02

Let people talk to it

Open a low-latency voice or video call that listens, responds, and takes turns naturally.

03

Give it work to do

Connect knowledge, memory, and tools so every conversation can lead to an outcome.

04

Put it where it matters

Ship on the web, in your product, or over the phone from one real-time API.

Voice, video, conversation, and action work from the same character definition.SUNESIS / REAL-TIME STACK
Sunesis Voice

Make every interaction feel considered.

Explore voice, live interaction, and developer tools when this section comes into view.

Flexible ways to build

Start with the moment you want to improve.

Choose the path that fits your team today. Add characters, agents, and developer tooling later without changing the foundation underneath.

Product designer waving during a live video call in a creative studio

Turn one idea into a live interaction.

Move from a brief to a character, voice, or guided conversation without assembling five separate tools.

  • Prototype a character in Studio
  • Preview voice and video together
  • Share a working flow before you ship

One workspace · shared credentials · usage you can see.

Developers

Voice, video, and conversation APIs built for production.
Fast to start. Open by design.

Build speech, cloning, avatars, and live transcription into one API workflow. TypeScript, Python, and Go SDKs are ready out of the box, with sub-200ms response times.

Quickstart

From signup to your first response in 5 minutes.

No sales call required. Grab your key, paste the snippet, and ship.

import { SunesisClient } from "@sunesis/sunesis-js";
const sunesis = new SunesisClient();

// Stream lifelike speech, token by token
const audio = await sunesis.textToSpeech.convert({
  text: "Welcome to Sunesis.",
  voiceId: "aria",
  stream: true,
});
Use cases

What teams ship on Sunesis.

A person in a live video conversation
Sunesis
Use case
USE CASE 01

Avatar video

Lip-syncable, emotion-aware TTS for AI avatar products. Your digital character speaks with natural prosody and visible emotion, video + voice fused.

Frame-accurate lip syncEmotion-aware prosody
Explore avatar video
A support specialist wearing a headset
Sunesis
Use case
USE CASE 02

Voice agent

Sub-200ms turn-taking over WebSocket. Streaming TTS and ASR so your agent can listen, think, and respond without the awkward pause.

<200ms turn-takingStreaming TTS + ASR
Tap to feature
A studio microphone ready for recording
Sunesis
Use case
USE CASE 03

Audio content & companions

Notes-to-audio, narration, and AI companions that actually talk back. Per-character pricing that scales with usage, not seats.

Notes-to-audioNarration & dubbing
Tap to feature
A developer working with a laptop
Sunesis
Use case
USE CASE 04

Conversational apps

Voice cloning from 15 seconds of audio. Persistent memory, function calling, and RAG, build agents that remember and act.

15s voice cloningMemory + function calling
Tap to feature
Character, voice & live conversation

Make every interaction feel like it has a point of view.

Create the character and voice together, then deliver a response quickly enough to keep the moment moving.

Character foundation

Start with someone worth talking to.

Set the role, tone, voice direction, and context a character needs before the conversation starts.

Role & point of viewVoice directionContext & guardrails

First response

<200ms

Keep the exchange moving.

Audio begins almost immediately, so a reply does not feel like a loading state.

Languages

32

Work across markets.

Let the same character carry a consistent voice from one language to the next.

Voice sample

15s

Make the voice yours.

Use a short, approved sample to create a voice a character can keep.

4.6 MOS voice qualityMeasured naturalness, not a promise on a slide.
One foundation. Many ways to build.

Compose the interaction your product deserves.

Start with speech, a character, a live session, or an agent, and connect them when the experience calls for more. The same workspace, credentials, webhooks, usage view, and SDK conventions carry across every layer.

Read the platform docs

Typed SDKs

TypeScript, Python, Go, REST, and WebSocket examples that match production behavior.

Built to operate

Scoped API keys, audit-friendly workspace controls, and predictable error handling.

Human developer support

Talk to the engineers building the platform when you are designing at scale.

Talk to us
Real people

Made for the humans on both sides.

Creators
Creators

Clone your voice and narrate anything.

Live conversations
Live conversations

Face-to-face with a digital character.

Support & sales
Support & sales

Voice agents that truly listen.

SUB-200MSLATENCY · VIDEO + VOICE
32 LANGUAGESCLONE IN 10 SECONDS
4.6 MOSHUMAN-LEVEL NATURALNESS
Digital Characters

Real faces. Real conversations.

Click any face to hear them speak.

240+ voices, infinite possibilities, voices uploaded and built by the Sunesis community.

The difference

Not another laggy AI platform.

The others

  • Latency500ms+ latency
  • Voice qualityRobotic voices
  • MemoryNo memory
  • Pricing / lock-inVendor lock-in
  • Regions / uptimeSingle-region

Sunesis Labs

The upgrade
  • Latency180ms latency
  • Voice qualityMOS 4.6 naturalness
  • MemoryPersistent memory
  • Pricing / lock-inOpen API, no lock-in
  • Regions / uptime14 regions · 99.99% uptime
Customer outcomes

Powering voice and video for 2,400+ teams.

From new products finding their voice to public companies redesigning a customer journey, teams use Sunesis to make the interaction count.

Customer support specialist
Voice agents
We replaced our entire IVR with Sunesis agents. Call deflection jumped overnight and our support costs dropped by 64%.

Mara Lindqvist

VP Engineering, Northwind

64% cost reduction
Live video conversation
Digital characters
Onboarding completion went from 41% to 89% once we added a face-to-face digital character to the flow.

Dev Patel

Head of Product, Koru Health

41% → 89% onboarding
Developer working on a laptop
Real-time API
We ran 40k concurrent voice sessions during launch week without a single dropped call. The streaming latency is unreal.

Sofia Marchetti

Staff Engineer, Helio

40k concurrent sessions
In the wild

The product is part of the conversation.

Creators and developers share the moments that make a voice experience feel different.

Cloned my voice in 15 seconds and it nailed my cadence on the first try. Wild.

@creatorkai@YouTube

The streaming TTS feels instant, my viewers think I'm live when I'm not.

@novastreams@TikTok

Built a customer support agent in an afternoon. It's better than our offshore team.

@indiehacker@Twitter

Finally a voice API that doesn't sound like a robot from 2015.

@devjules@YouTube
Plans that grow with you

Start building now. Choose a plan when it matters.

Explore credits, production features, and team options on the pricing page, without interrupting the product story.

Explore pricing
The real-time backbone

The details that make an interaction feel instant.

Every layer is measured for the moment between a person speaking and the system responding.

0.0k
teams building on Sunesis
0ms
median latency, voice generation
0
languages for live conversations
0.0
MOS, human-level naturalness
Streaming voice · character video · grounded agentsMeasured in production, not just in a demo
FAQ

Questions, answered.

Sunesis Voice v2 speaks 32 languages out of the box with any of our 240+ voices. The API supports 30+ languages including English, Spanish, Japanese, French, German, Mandarin, and more, all with native-quality prosody and the ability to clone a voice and use it across languages.

AI Voice, Video & Conversation Agents | Sunesis Labs