Lip sync from text-to-speech timings: from timestamps to mouth shapes
Lip sync from text-to-speech depends on which timings the voice returns. ElevenLabs gives character timestamps, Cartesia can stream phonemes, and Azure and Amazon Polly return visemes. How each becomes a mouth shape, and where forced alignment fits.
Updated
The short version
- A mouth follows sounds, not letters, so the timing that drives one has to say which sound is playing and when.
- ElevenLabs returns timing per character, which is what captions need; for a mouth, your code turns those letters into sounds first. Phoneme timestamps and viseme events skip that step.
- Forced alignment works for any voice, once the audio exists.
- Keep time with the audio clock, and hold a still face when a timing is missing.
A mouth follows sounds, so lip sync needs a timing for every sound
Someone on the team has a character, a voice from a text-to-speech API and a demo on Thursday. The audio sounds right. The mouth, so far, opens and closes on a timer, and everyone in the room can see it.
Two words carry the whole problem. A phoneme is a sound in speech. A viseme is the shape the face makes for it — Microsoft's speech documentation calls it the visual description of a phoneme — and several sounds share one shape, because s and z look the same on a face.
Letters are a weak stand-in for either. English spells one sound many ways and gives one spelling many sounds: the k in knight is silent, and the ough in though sounds nothing like the ough in cough. So the useful question to ask of any voice is what its timings are attached to — letters, words, sounds or mouth shapes.
There are four answers in the wild, and they make four different amounts of work for you. Character timings need your code to work out the sounds. Phoneme timings need a table from sounds to shapes. Viseme events hand you the shapes. And forced alignment works out the timings after the fact, from the audio itself. The sections below take them in that order.
Every vendor detail below comes from that vendor's own documentation, read on 6 October 2026.
Character timestamps, such as ElevenLabs', say when each letter was spoken
ElevenLabs' text-to-speech API has a variant its reference calls Create speech with timing. It returns the audio together with an alignment: the characters of your text, each with a start and an end time in seconds. It gives that twice, once for the text as you sent it and once for the text as the model normalised it.
A streaming version, Stream speech with timing, sends the same thing in pieces: a stream of JSON chunks, each carrying a slice of the audio and the alignment that goes with it.
That's exactly what captions, word-by-word highlighting and trimming audio at a word boundary need.
For a mouth there's one more step, and it's yours. Your code turns the letters into sounds — a grapheme-to-phoneme step — and shares each letter's time among the sounds it becomes. That step is a prediction your code makes on top of the timing; the timing itself doesn't carry it. If you take this route, read the normalised alignment: normalisation is the step that turns what you typed into what gets spoken, so the dates and numbers in it are written the way the voice said them.
Try it
Ask for timestamps on the word knight. Every character comes back with a start and an end, the silent k included. That's right for highlighting the word, and a reminder that a mouth keyed to letters would shape a sound nobody made.
Phoneme timestamps say which sound, so the mouth shape becomes a lookup
Some voices return the sounds themselves. Cartesia's streaming API, for one, takes add_phoneme_timestamps and then interleaves phoneme timestamp events — the phonemes, each with a start and an end — with the audio chunks.
From there the mouth is a table. Each phoneme maps to one of a small set of mouth shapes, and many map to the same one. Microsoft publishes a mapping from phonemes to its own viseme IDs, which makes a sensible starting point even if you never call its API.
Check which phonetic alphabet a vendor uses before you write that table. IPA and ARPAbet name the same sounds with different symbols, and a table written for one quietly mismatches the other.
Viseme events, from Azure and Amazon Polly, hand you the mouth shape directly
A few speech services go one step further and return the shape. Azure's Speech SDK raises a VisemeReceived event as it synthesises, carrying a viseme ID and the audio offset where that shape starts, and it can also return SVG animation for a 2D face or blend-shape frames for a 3D one.
Amazon Polly returns speech marks, and one type of speech mark is the viseme. Its documentation notes that when you request speech marks, Polly returns that metadata instead of synthesised speech, so the audio and the timings are two requests you line up yourself.
One thing to plan for: each service has its own set of shapes. Draw your mouths for the set you'll actually receive, or map every vendor's set onto one of your own, so the art outlives any one vendor.
Streaming changes when the timings arrive, so schedule them against the sound
In a live conversation the voice is generated while it plays, and so are its timings. ElevenLabs' streaming endpoint sends alignment with each chunk of audio, Cartesia interleaves phoneme events with its audio chunks, and Azure raises viseme events during synthesis. The timings are the same kind of thing they were in a single request; what changes is that they turn up in pieces, alongside the audio they describe.
So keep one queue. Each timing goes in against the playback position of the audio it belongs to, and the mouth reads from that queue as the audio plays rather than acting on events as they land. Before adding anything up, check in the vendor's reference whether its times count from the start of the whole utterance or from the start of each chunk, because a mouth built on the wrong one drifts further out with every chunk.
Two more cases need deciding before the demo, not after it. Start playback once a little audio is buffered, so a timing that arrives late still arrives before its sound. And when someone interrupts the character, stop the audio and drop the rest of its timings in the same step, so the mouth stops with the voice instead of finishing a sentence nobody can hear.
Forced alignment fits any voice, once the audio exists
The other route ignores what the voice returns. Give an aligner the audio and the text that was spoken, and it works out when each part was said.
ElevenLabs offers this as a separate Forced Alignment API, which returns characters and words with their timings. Rhubarb Lip Sync, an open-source tool, analyses voice recordings and produces mouth animation directly, and its documentation recommends giving it the dialogue text for more reliable results.
The trade is order. An aligner runs after the speech exists, so in a live conversation it has to keep pace with the voice, which is engineering of its own. What it buys is independence: any voice can drive the mouth, including one that returns no timings at all.
That's the route Runner takes for its own agents. Runner aligns its own speech, so a voice doesn't have to supply its own timings to move an agent's mouth. How the aligner works stays in house.
Runner's agents, and the stack they run onTurn the timing track into a mouth that reads as speech
Timings are half the job. The other half is drawing and scheduling, and a few habits matter more than where the timings came from.
- Keep the set of mouths small. A handful of well-drawn shapes reads better than many that flicker.
- Close the lips fully for p, b and m, the sounds made with the lips together.
- Ease from one shape to the next rather than snapping between them.
- Return to a resting mouth in pauses, so silence looks like silence.
- Keep time with the audio clock — in a browser, the Web Audio API's currentTime — rather than with timers, which drift away from what's actually playing.
- When a timing is missing or late, hold the still face. A calm, closed mouth reads as a character listening; a mouth moving on a guess reads as broken.
The short version
Timings say when; your mouth art says what. Draw both from the same set of sounds, and let the audio clock keep time.
Ask these before you pick a voice for a talking face
Each answer is in the vendor's API reference, and it's worth finding before the character is drawn.
- What are the timings attached to: characters, words, phonemes or visemes?
- Do they arrive with the audio as it streams, or in a separate request?
- Which phonetic alphabet or viseme set do they use, and is it the same in every language you need?
- Are they given for the text as you sent it, or as the voice normalised it?
- When streaming, do the times count from the start of the utterance or from each chunk?
- If the voice returns no phonemes, which aligner will you run, and where?
- What does your face do when a timing is missing or late, and when someone interrupts?
One system for the whole path
The mechanism above is one of the jobs Runner does on one record, under one login. Access is by application.
Trademarks of their owners. No affiliation or endorsement.