LipsieSync® is a multilingual video dubbing service that uses a human performance recorded in each target language, followed by speech-to-speech voice transformation. One voice actor performs all dialogue in that language, providing the acting, timing, emotion and intent. Lipsie then transforms the recorded timbre to create a distinct vocal identity for each speaker or character. Speech-to-speech changes the timbre, not the actor’s performance. Dialogue is adapted before recording, and lip sync is applied only to shots where the dubbed speech requires an adjustment to match visible mouth movements.
This demonstration shows the same speaker dubbed in French, English, Italian and German. Each version is recorded as a human performance in the target language, then processed with speech-to-speech voice transformation. Dialogue is adapted for each language, and lip sync is applied only to shots that require adjustment.
LipsieSync® is a multilingual video dubbing service based on a human performance recorded in the target language, followed by speech-to-speech voice transformation. The voice actor provides the timing, emotion, intent, breathing and changes in intensity for each line. Speech-to-speech then transforms the recorded timbre to create a distinct vocal identity for each speaker or character. Speech-to-speech processing is performed locally by Lipsie, not through a cloud-based voice transformation service. Dialogue is adapted before recording, and lip sync is applied only to the shots where it is needed. The production is supervised by Lipsie.
For each target language, one voice actor performs all dialogue. The acting, timing, emotion, breathing and intent come from the actor’s performance, not from a voice generated by AI from text.
The voice actor is selected to perform the content. Speech-to-speech transforms the timbre after recording to bring the resulting voice closer to the source voice. The transformation changes the timbre, not the actor’s performance.
One voice actor can perform multiple speakers or characters in the same target language. Lipsie transforms those recordings into distinct, identifiable voices, reducing the number of voice actors and separate castings required for videos with multiple speakers.
Dialogue is adapted for the target language before the recording session. Spoken phrasing, duration, timing, intent and pronunciation are adjusted based on meaning, on-screen content and synchronization requirements.
Mouth movements are adjusted on shots where the dubbed speech and visible mouth movements require closer synchronization. Lip sync is not applied systematically to the entire video.
Vocal identities defined for a project can be reused across multiple videos or episodes so that the same speakers retain the same voices. Dialogue adaptation, human performance, voice transformation and synchronization are supervised by Lipsie within the defined production scope.
LipsieSync® is used for videos with one or more on-screen speakers who need a distinct vocal identity in each target language. One voice actor performs all dialogue in a target-language version, reducing the number of separate castings required for videos with multiple speakers. For series, training programs and recurring content, the same vocal identities can be reused across videos so that the same speakers retain the same voices.
Executive communications, corporate interviews, brand films, customer testimonials and internal communications. For recurring content, the same vocal identity can be reused from one video to the next.
Multi-speaker interviews, documentary series, magazine-style programs and factual content. Each speaker has a distinct vocal identity without requiring a separate voice actor for every speaker.
Masterclasses, executive training, presenter-led e-learning, demonstrations, and technical or medical content. Each version is performed in the target language, and the same vocal identities can be reused across modules.
Video podcasts, conversations, recurring shows and multi-speaker YouTube videos. One voice actor per target language can perform multiple speakers, with distinct vocal identities reused across episodes.
YouTube and creator dubbing ➤
Sub2Dub® creates dubbed audio from text using synthetic voices. LipsieSync® records a human performance in the target language, then uses speech-to-speech voice transformation to create distinct vocal identities from that recording. Studio Premium assigns a separate voice actor to each role under artistic direction.
| Criterion | Sub2Dub® | LipsieSync® | Studio Premium |
|---|---|---|---|
| Source of the performance | Voice generated from text | Human performance recorded in the target language | Human performance recorded separately for each role |
| Performance | No recorded human performance | The acting, timing, emotion and intent come from the voice actor | Each role is performed by a separate voice actor under artistic direction |
| Vocal timbre | Determined by the selected synthetic voice | Speech-to-speech transforms the timbre after recording without altering the voice actor’s performance | Determined by the voice actor cast for the role |
| Number of voice actors | No voice actor recording required | One voice actor per target language, including videos with multiple speakers | One voice actor per role and per language |
| Vocal identities | A synthetic voice is assigned to each speaker | A distinct vocal identity is created for each speaker from the voice actor’s performance, with the timbre adjusted toward the source voice | Each role uses the voice of the actor performing it |
| Corrections and retakes | The text or generation settings can be adjusted | Depending on the requested change, corrections can address dialogue adaptation, a specific recording or voice processing | A correction may require a new recording from the voice actor assigned to the role |
| Lip sync | Available depending on the project | Applied to shots where the dubbed speech requires an adjustment to visible mouth movements | Synchronization is handled as part of the dubbing and direction process |
| Indicative turnaround 10-minute video / 1 language | 1 to 2 business days | 4 to 7 business days | 7 to 15 business days |
| Indicative pricing | From €30 excluding VAT per minute and per language | From €70 excluding VAT per minute and per language | From €180 excluding VAT per minute and per language |
| Typical use cases | Informational content, high volumes and short turnaround times | On-camera videos, interviews, training content and videos with multiple speakers | Fiction, animation, broadcast content and productions with dedicated casting |
Indicative prices are per minute of source video and per language, excluding VAT, for standard speech density. Final pricing depends on the language pair, number of speakers, dialogue adaptation, shots requiring lip sync, usage rights and requested deliverables. A minimum project fee may apply to short videos.
A LipsieSync® dubbing project has six stages: video analysis, transcription and speaker identification, dialogue adaptation for the target language, voice actor recording, locally processed speech-to-speech voice transformation with selective lip sync, and quality control, targeted revisions and delivery.
LipsieSync® is a multilingual video dubbing service based on human performance and speech-to-speech voice transformation. These answers explain how the service works, how speech-to-speech is processed locally by Lipsie, how many voice actors are used, when lip sync is applied, and how voice continuity, corrections, rights, pricing and turnaround are handled.
Provide the video duration, target languages, number of speakers and planned release date. Lipsie will provide the production scope, price and estimated turnaround for your project.
One voice actor per target language · Speech-to-speech processed locally by Lipsie · Lip sync applied only where needed