Talking avatar in JS: make one from Node, play it in React
A talking avatar in JavaScript: create it and send it a script from Node with the Sume SDK, wait for the job, then play the returned MP4 in React.
To make a talking avatar in JavaScript, call an avatar video API from your Node server, wait for the render, and give your front end the finished MP4 to play. With Sume that is two jobs from the @sume-com/sdk package: generateAvatarV1 creates the avatar once, and createAvatarV1TalkingVideo makes it speak a script. The result is a video file, not a live or 3D avatar that talks back in the browser.
The SDK facts come from TypeScript SDK and Waiting for runs and jobs, the avatar rules from Create new avatar and Generate avatar video, and response fields from the Sume API reference, read on 2026-09-29. The same flow in Python is Create a talking avatar using Python.
How do I make a talking avatar from Node?
Install @sume-com/sdk (npm install @sume-com/sdk; it runs on Node 18+, Bun, Deno and Cloudflare Workers, and the SDK quickstart covers setup and errors). Create the avatar from a prompt (a profile or a public HTTPS photo also works), wait for that job, then send the handle a script. Avatar routes create jobs, not runs, so the wait is waitForJob. The talking-video submit also returns an avatar_video_id; the avatar-video resource carries the file's video_url.
import { createAvatarV1TalkingVideo, createSumeClient,
generateAvatarV1, getAvatarVideo, waitForJob } from "@sume-com/sdk";
const client = createSumeClient({ apiKey: process.env.SUME_API_KEY! });
const made = await generateAvatarV1({
client,
headers: { "idempotency-key": "demo-host-v1" },
body: { avatar_handle: "demo_host",
input: { type: "prompt", prompt: "A friendly presenter in a bright studio" } },
});
if (made.error) throw new Error(JSON.stringify(made.error));
await waitForJob(made.data!.data.request_id, { client });
const talk = await createAvatarV1TalkingVideo({
client,
headers: { "idempotency-key": "demo-video-v1" },
body: { avatar_handle: "demo_host", quality: "standard",
script: "Hi! A Node script made this whole video." },
});
if (talk.error) throw new Error(JSON.stringify(talk.error));
const job = await waitForJob(talk.data!.data.request_id, { client });
if (job.status !== "completed") throw new Error(job.status);
const video = await getAvatarVideo({ client, path: { id: talk.data!.data.avatar_video_id } });
console.log(video.data?.data.avatar_video.video_url);How long does waitForJob wait?
Up to 20 minutes by default; the docs note that avatar-video jobs routinely run minutes. It resolves for any terminal status, so check status before reading the file, as the sample does. A timeout doesn't cancel the job: it keeps running and still bills, so store the job id and read it later with getApiJob. Waiting for jobs and runs covers the options and errors.
How do I show a talking avatar in React?
Keep the Sume calls on your server and let the React app talk to your own endpoint. Sume's authentication page says browser and mobile clients should call your backend, which attaches the API key, and the SDK page says there is no browser-safe key: never put it in client JavaScript, a mobile bundle or a NEXT_PUBLIC_* variable. Your endpoint returns the video_url, and the component renders it in an HTML <video> element with controls. Completed avatar results include public media.sume.com video files, so the browser can load the URL directly.
What are the limits and costs?
Avatar creation costs $0.95 per avatar, and talking video is billed per second: $0.184/s standard, $0.245/s plus, $0.55/s max (no product image). Each is plus a 5.5% agent fee by default. quality defaults to plus when omitted, so send standard if you want that rate.
| Limit | Value |
|---|---|
| Script length | An estimated 4–60 seconds; split longer scripts |
aspect_ratio | 1:1, 3:4, 9:16 (default), 4:3, 16:9 |
| Resolution | 720p is the documented resolution |
| Avatars per video | One |
| Spoken language | English only in current code |
captions on the create | Refused in current code; caption the finished file separately |
Sources
Related posts
More in Developers
- Avatar video quality settings: standard, plus or max?
Sume's talking avatar video takes quality standard, plus (default) or max. What each means, the other output fields, and how the preview relates to final tier.
- C# text to speech: call a TTS API with HttpClient
Text to speech in C#: POST the text and a voice with HttpClient, poll the job until it finishes, then stream the MP3 from its audio_url to a file.
- Text to speech streaming API: what Sume returns instead
Sume's text to speech API doesn't stream audio chunks. It returns a finished file per job; split long scripts into sentence jobs to start playback sooner.
- Duck background music under a voiceover with the Timeline API
Set soundtrack.duck_db (0 to 20) on POST /v1/timeline-1.0/render so the music dips under your voiceover spine. It needs a real spine; silence mode is refused.
Written by Sume