LipsieSync®: multilingual video dubbing with one voice actor per language

LipsieSync® is a multilingual video dubbing service that uses a human performance recorded in each target language, followed by speech-to-speech voice transformation. One voice actor performs all dialogue in that language, providing the acting, timing, emotion and intent. Lipsie then transforms the recorded timbre to create a distinct vocal identity for each speaker or character. Speech-to-speech changes the timbre, not the actor’s performance. Dialogue is adapted before recording, and lip sync is applied only to shots where the dubbed speech requires an adjustment to match visible mouth movements.

Request a LipsieSync® pilot dub — Send a representative excerpt, the target languages and the number of speakers. Lipsie uses the excerpt to assess dialogue adaptation, human performance, speaker differentiation and lip sync on the relevant shots, then confirms the project scope, timeline and deliverables.

LipsieSync® demonstration: the same speaker in four languages

This demonstration shows the same speaker dubbed in French, English, Italian and German. Each version is recorded as a human performance in the target language, then processed with speech-to-speech voice transformation. Dialogue is adapted for each language, and lip sync is applied only to shots that require adjustment.

  • a human performance recorded in each target language

  • a vocal timbre adjusted toward the source speaker’s voice

  • dialogue timing adapted to each language

  • lip sync applied only to the relevant shots

Source language: French · Dubbed versions: French, English, Italian and German

Production: dialogue adaptation, human performance, speech-to-speech voice transformation and selective lip sync.

Watch the full 26-minute versions

What is LipsieSync®?

LipsieSync® is a multilingual video dubbing service based on a human performance recorded in the target language, followed by speech-to-speech voice transformation. The voice actor provides the timing, emotion, intent, breathing and changes in intensity for each line. Speech-to-speech then transforms the recorded timbre to create a distinct vocal identity for each speaker or character. Speech-to-speech processing is performed locally by Lipsie, not through a cloud-based voice transformation service. Dialogue is adapted before recording, and lip sync is applied only to the shots where it is needed. The production is supervised by Lipsie.

LipsieSync® at a glance

One human performance in each target language

For each target language, one voice actor performs all dialogue. The acting, timing, emotion, breathing and intent come from the actor’s performance, not from a voice generated by AI from text.

Human performance and vocal timbre are separate stages

The voice actor is selected to perform the content. Speech-to-speech transforms the timbre after recording to bring the resulting voice closer to the source voice. The transformation changes the timbre, not the actor’s performance.

A distinct vocal identity for each speaker

One voice actor can perform multiple speakers or characters in the same target language. Lipsie transforms those recordings into distinct, identifiable voices, reducing the number of voice actors and separate castings required for videos with multiple speakers.

Dialogue is adapted before recording

Dialogue is adapted for the target language before the recording session. Spoken phrasing, duration, timing, intent and pronunciation are adjusted based on meaning, on-screen content and synchronization requirements.

Lip sync is applied only where needed

Mouth movements are adjusted on shots where the dubbed speech and visible mouth movements require closer synchronization. Lip sync is not applied systematically to the entire video.

Vocal identities can be reused across videos

Vocal identities defined for a project can be reused across multiple videos or episodes so that the same speakers retain the same voices. Dialogue adaptation, human performance, voice transformation and synchronization are supervised by Lipsie within the defined production scope.

What types of videos can be dubbed with LipsieSync®?

LipsieSync® is used for videos with one or more on-screen speakers who need a distinct vocal identity in each target language. One voice actor performs all dialogue in a target-language version, reducing the number of separate castings required for videos with multiple speakers. For series, training programs and recurring content, the same vocal identities can be reused across videos so that the same speakers retain the same voices.

Corporate videos

Executive communications, corporate interviews, brand films, customer testimonials and internal communications. For recurring content, the same vocal identity can be reused from one video to the next.

Interviews and documentaries

Multi-speaker interviews, documentary series, magazine-style programs and factual content. Each speaker has a distinct vocal identity without requiring a separate voice actor for every speaker.

Training and expert-led content

Masterclasses, executive training, presenter-led e-learning, demonstrations, and technical or medical content. Each version is performed in the target language, and the same vocal identities can be reused across modules.

YouTube videos and creator content

Video podcasts, conversations, recurring shows and multi-speaker YouTube videos. One voice actor per target language can perform multiple speakers, with distinct vocal identities reused across episodes.

YouTube and creator dubbing  ➤

How do Sub2Dub®, LipsieSync® and Studio Premium differ?

Sub2Dub® creates dubbed audio from text using synthetic voices. LipsieSync® records a human performance in the target language, then uses speech-to-speech voice transformation to create distinct vocal identities from that recording. Studio Premium assigns a separate voice actor to each role under artistic direction.

Comparison of Sub2Dub, LipsieSync and Studio Premium video dubbing services
Criterion Sub2Dub® LipsieSync® Studio Premium
Source of the performance Voice generated from text Human performance recorded separately for each role
Performance No recorded human performance Each role is performed by a separate voice actor under artistic direction
Vocal timbre Determined by the selected synthetic voice Determined by the voice actor cast for the role
Number of voice actors No voice actor recording required One voice actor per role and per language
Vocal identities A synthetic voice is assigned to each speaker Each role uses the voice of the actor performing it
Corrections and retakes The text or generation settings can be adjusted A correction may require a new recording from the voice actor assigned to the role
Lip sync Available depending on the project Synchronization is handled as part of the dubbing and direction process
Indicative turnaround 10-minute video / 1 language 1 to 2 business days 7 to 15 business days
Indicative pricing From €30 excluding VAT per minute and per language From €180 excluding VAT per minute and per language
Typical use cases Informational content, high volumes and short turnaround times Fiction, animation, broadcast content and productions with dedicated casting

Indicative prices are per minute of source video and per language, excluding VAT, for standard speech density. Final pricing depends on the language pair, number of speakers, dialogue adaptation, shots requiring lip sync, usage rights and requested deliverables. A minimum project fee may apply to short videos.

How is a LipsieSync® dubbing project produced?

A LipsieSync® dubbing project has six stages: video analysis, transcription and speaker identification, dialogue adaptation for the target language, voice actor recording, locally processed speech-to-speech voice transformation with selective lip sync, and quality control, targeted revisions and delivery.

The six stages of a LipsieSync multilingual video dubbing project The LipsieSync process includes video analysis, transcription and speaker identification, dialogue adaptation for the target language, voice actor recording, locally processed speech-to-speech voice transformation with selective lip sync, and quality control, targeted revisions and delivery. STEP 1 Video analysis source format, speakers and visible shots target languages, distribution and usage rights STEP 2 Transcription and speaker ID verified dialogue and identified speakers terminology, pronunciations and timecodes STEP 3 Dialogue adaptation spoken dialogue adapted for the target language meaning, duration, timing and intent adjusted STEP 4 Voice actor recording acting, timing, emotion and intent breathing, nuance and approved takes STEP 5 Speech-to-speech and lip sync speech-to-speech processed locally by Lipsie timbre adjusted toward source voices, lip sync STEP 6 QA, revisions and delivery voices, synchronization, mix and targeted revisions vocal identity continuity and final file checks STEP 1 Video analysis format, speakers and visible shots target languages, distribution and usage rights STEP 2 Transcription and speaker ID verified dialogue and identified speakers terminology, pronunciations and timecodes STEP 3 Dialogue adaptation spoken dialogue adapted for the target language meaning, duration, timing and intent adjusted STEP 4 Voice actor recording acting, timing, emotion and intent breathing, nuance and approved takes STEP 5 Speech-to-speech and lip sync speech-to-speech processed locally by Lipsie timbre adjusted toward source voices, lip sync STEP 6 Quality control and delivery voices, synchronization, mix and targeted revisions vocal identity continuity and final file checks

Frequently asked questions about LipsieSync®

LipsieSync® is a multilingual video dubbing service based on human performance and speech-to-speech voice transformation. These answers explain how the service works, how speech-to-speech is processed locally by Lipsie, how many voice actors are used, when lip sync is applied, and how voice continuity, corrections, rights, pricing and turnaround are handled.

LipsieSync® is a multilingual video dubbing service in which a voice actor performs the dialogue in the target language before speech-to-speech voice transformation. Speech-to-speech then transforms the recorded timbre to create a distinct vocal identity for each speaker or character. LipsieSync® can be produced in more than 50 languages, and lip sync is applied only to shots that require adjustment between the dubbed speech and visible mouth movements.

Speech-to-speech transforms the timbre after recording without changing the voice actor’s performance. The acting, timing, emotion, breathing and intent come from the human performance recorded in the target language. The timbre is then transformed to create a distinct voice for each speaker and can be adjusted toward that speaker’s voice in the source content.

No. LipsieSync® speech-to-speech processing is performed locally by Lipsie, not through a cloud-based voice transformation service. Timbre transformation takes place in Lipsie’s production environment after the voice actor has been recorded.

LipsieSync® uses one voice actor per target language, regardless of the number of speakers to be dubbed. The actor performs all dialogue in that language, and speech-to-speech then creates a distinct vocal identity for each speaker or character. Compared with dubbing that uses one actor per role, the number of required voice actors and castings does not increase with the number of speakers.

Yes. A vocal identity defined for a speaker can be reused across multiple videos, episodes or modules. The same person or character can therefore retain the same identifiable voice across a series, training program or other recurring content.

Yes. A correction can be limited to the affected part of the dubbed version. Depending on the requested change, Lipsie can revise dialogue adaptation, a specific recording, voice processing, synchronization or lip sync without systematically redoing the entire version.

No. Lip sync is applied only to shots where mouth movements need additional adjustment to match the dubbed speech. Lipsie considers face visibility, camera angle, line duration and the synchronization already achieved through dialogue adaptation.

Voice and visual transformations must be covered by the rights and consents required for the project. The scope can specify the people concerned, languages, territories, media and duration of use. The client must hold the necessary rights to the submitted content, and the terms applicable to the voice actors are defined for the production.

LipsieSync® pricing starts at €70 excluding VAT per minute of source video and per language. Final pricing depends on video duration, speech density, number of speakers, target language, dialogue adaptation, lip sync, usage rights and requested deliverables. A minimum project fee may apply to short videos.

A 10-minute video in one target language typically takes 4 to 7 business days. Turnaround depends on speech density, number of speakers, target language, voice actor availability, approval stages and the number of shots requiring lip sync.

Request a quote for a LipsieSync® multilingual dubbing project

Provide the video duration, target languages, number of speakers and planned release date. Lipsie will provide the production scope, price and estimated turnaround for your project.

One voice actor per target language · Speech-to-speech processed locally by Lipsie · Lip sync applied only where needed