AI simultaneous interpretation with Gemini AI Translate: function, cost, limits
AI simultaneous interpretation is an automatically generated audio track — what it delivers, where it stops, and when a human interpreter is the better choice

Multilingual events rarely fail on technology. They fail on the invoice. A simultaneous interpreter costs €900 to €1,400 net per day at market rates, and simultaneous interpreting requires interpreters to work in pairs — so roughly €1,800 to €2,800 per language for a single event day, plus briefing, preparation time and booth technology (source: AIIC Germany). At three target languages, the interpreting budget outgrows the entire production. Every communications team knows what follows: the all-hands, the partner webinar, the product launch all run in one language — not because nobody wanted a second one, but because it was never in the budget.
Since 9 June 2026 there has been an option that simply did not exist before. On that day Google introduced Gemini 3.5 Live Translate, a model that translates spoken language inside a running stream instead of waiting for sentence or speaker boundaries. In SlideSync we built a complete language track on top of it: a translated voice over the quietly continuing original, plus subtitles carrying the identical wording — selectable in the player exactly like a human interpreter channel.
If you are wondering how to use Gemini AI Translate for your event with the least effort: through SlideSync. Google provides the translation but not a finished language track — in between sits integration work your team or your production partner would otherwise have to do. In SlideSync it is already done: you book the languages, everything else stays the webcast you know. This page covers what it delivers today, what it costs, where it stops, and when you should still book an interpreter.
At a glance (as of August 2026):
- What it is: an automatically generated language track — a translated voice over the attenuated original, plus subtitles with the same wording. The translation comes from Google’s Gemini AI Translate; the voice-over on top of it comes from SlideSync.
- What you do: nothing you do not already do for a webcast — one encoder, one signal. No interpreter booth, no second desk, no software on the presenter’s machine, no additional hardware. Languages can be added flexibly right up to the start of the broadcast.
- Cost: a low three-figure amount per language and event — against €1,800 to €2,800 for a pair of interpreters per language and day.
- Delay: the translation trails the picture by a few seconds and is not lip-synced; over longer passages the gap widens.
- Limits: no terminology management for the spoken translation, no selectable voices, no distinction between multiple speakers. The audio track is currently processed outside the EU for the translation.
- Does it replace interpreters? Where cost meant no interpreting happened at all: yes. For annual general meetings, contract negotiations and crisis communications: no.
What is Gemini AI Translate?
Gemini is Google’s AI system; “Gemini AI Translate” is the part of it that translates. For an event, exactly one stage of it matters: Gemini 3.5 Live Translate, introduced on 9 June 2026. It translates spoken language while someone is speaking — it does not wait for a sentence or a contribution to finish. That is the difference between a translation you can listen to live and one you read afterwards.
Machine live translation existed before. It only became usable for an event now. Until June 2026, Google’s live translation in meetings handled five languages, and everything routed through English — French into Italian was simply not provided for. Today the system recognises more than 70 languages by itself and covers over 2,000 language combinations. On top of that come three things you hear immediately: the translation keeps running continuously instead of restarting after every contribution, it carries over the tone, speaking rhythm and pitch of the person speaking, and it copes better with noise in the room (source: Google, 9 June 2026).
Two things to know beforehand
First, Google still runs the whole thing as a preview. That is the ordinary route at Google before general release, but it also means details may still change over the coming months. Second, the audio track of your event is currently processed outside the EU for the translation. Both will foreseeably change. For today that means: quality is already sufficient for real events, and this is precisely the right moment to try it once rather than buying it as a commodity in two years. If data processing inside the EU is non-negotiable for you, book an interpreter for that one event.
Why this runs through SlideSync
Google offers live translation as part of its Translate product line. It can be integrated into your own products on a cloud-hosted basis — but the direct route there is fiddly: what you get is a translation service, not a finished language track for an event. Everything in between has to be built by somebody.
That is exactly what we did. What you hear in the player — the translated voice over the attenuated original, sitting alongside the other languages, switchable mid-stream, plus subtitles carrying exactly the same wording — is built into SlideSync and ready to use. The voice-over is not a Google feature; it comes from us. For you that means a checkbox next to the language instead of an integration project.
What you need to prepare — and what you do not
The effort on your side is the real difference to an interpreter booth, and it is quickly described: you do nothing you would not do for any other webcast. One encoder, one signal into SlideSync, the same setup as always. Side by side:
| What falls to you | With interpreters | With an AI language track |
|---|---|---|
| Encoder and signal | as usual | as usual |
| Interpreter booth, second control desk | required | not needed |
| Additional audio recording | required | not needed |
| Briefing material, scripts, pre-recordings | weeks ahead, including material review | not needed |
| Software on the presenter’s machine | depending on setup | not needed |
| Devices, cabling, on-site personnel | required | not needed |
| Fixing the languages | weeks ahead, subject to availability | flexibly up to the start of the event |
There is nothing to prepare on the content side either: the system detects the spoken language itself. If a speaker switches language mid-presentation, that is recognised and the translation follows.
Two organisational points remain. First, languages are booked before the event, and that works flexibly right up to the start of the broadcast — you do not have to commit weeks in advance. What does not work: adding a language spontaneously during a running broadcast. Second, you decide per language between AI and a human. Both can be combined within the same event, say French as an AI track and English with an interpreter — but the assignment is fixed from the moment the broadcast starts. An interpreter who is meant to join later cannot be brought into a running show.
What your viewers get
In the player, your viewers find the same language selector they know from human interpreting — the interaction is identical, there is nothing to explain. Anyone who switches hears, a few seconds behind, a spoken translation over the quietly continuing original. The original stays audible in the background, the way a television voice-over works: tone, applause and the reactions in the room are not lost.
Alongside it run subtitles that reproduce exactly what is being spoken — not a second, diverging translation. Anyone moving between listening and reading sees the sentence they just heard. It sounds like a detail; in practice it is the difference between “helpful” and “confusing”.
Every language track is recorded as a matter of course. After the event the translated version is available to you as an on-demand video just like the original — so you can still publish multilingually after the fact.
Time-shifted, not lip-synced
The translation trails the picture by several seconds, and over longer passages that gap widens. You watch a speaker and hear the translation shortly afterwards — comparable to simultaneous interpreting, where the interpreter also runs behind the speaker. This is not lip-synced and will not be in a live setting. For talks, presentations and Q&A the offset is unproblematic; for a format that depends on exact picture-sound synchronicity it is the wrong choice.
How the speaker sounds
The translation is spoken in a voice that automatically leans towards the original speaker’s, rather than a neutral announcer voice. The result sits noticeably closer to the original than most people expect from synthetic speech.
What does not work belongs in the same answer: the voice is not selectable. There is no voice library, no choice of gender or speaking style, and no way to deposit a brand voice. Similarity to the original is also not guaranteed across a full event, and multiple speakers are not distinguished from one another. The defensible claim is therefore “in a voice close to the original” — not “consistent speaker identity”.
What it costs — against a human interpreter
The comparison that matters here is not “subtitles”. It is “one interpreter per language”: technology, booth, personnel and lead time on one side — a checkbox next to the language on the other.
| Per target language and event day | Simultaneous interpreters | AI language track in SlideSync |
|---|---|---|
| Fee / price | €1,800–2,800 net (working in pairs, €900–1,400 each) | low three-figure amount |
| Lead time | weeks — availability, briefing, material review | flexibly up to the start of the event |
| On-site technology | booth, second desk, additional audio recording | none — one encoder, one signal |
| On-site personnel | two interpreters per language | none |
| Subject terminology | can be briefed in advance and is absorbed | cannot be deposited for the spoken translation |
| Time offset | a few seconds | a few seconds, widening over long passages |
The decisive number is not in the table: an extra language does not cost a little less than a pair of interpreters. It costs many times less. That shifts the question. It is no longer: can we afford interpreters for this event? It becomes: which languages do we want to reach? Communications teams answer that with more languages than ever made it into a budget.
What AI simultaneous interpretation cannot do today
This list deliberately sits in front of the offer rather than in the small print. Every point here is a case where somebody would be disappointed without warning.
No terminology management for the spoken translation
Product names, technical terms, proper nouns or a maintained corporate terminology cannot be handed to the spoken translation — the technology simply does not provide for it. Anyone expecting an exact result on the first product name will not get one. This is the clearest dividing line against our existing subtitle offering: terminology management exists for the subtitles, not for the voice. If your terminology has to land word for word, the combination of original audio and maintained subtitles is the better route — or a briefed interpreter.
No selectable voices, no speaker distinction
The voice leans towards the original but is not configurable, and multiple speakers are not kept apart from one another. On a panel discussion with rapid turn-taking, that is audible. Google names the same limits in its own documentation: the voice can shift after longer pauses, and language detection becomes less reliable with strong accents or fast language switching.
Processing outside the EU
The audio track of your event is currently processed outside the EU for the translation. We say so actively, because for German and wider European enterprise customers it is the first question asked, and nobody is served by it surfacing in a sales conversation instead. The rest of our platform is unaffected — SlideSync hosts in the EU, and every language without an AI track runs unchanged. If the data from your event has to stay inside the EU, the AI language track is the wrong choice today. We expect Google to offer European processing once the model leaves preview.
Reverse translation: one-directional today, in development
An edge case that comes up regularly at international events: if someone in an English-language event suddenly speaks German, German viewers hear it correctly — English viewers hear the passage untranslated. The reverse direction is technically possible, because the model continuously identifies the spoken language and does not translate when source and target language are identical. In SlideSync this reverse translation is currently being implemented — so it belongs on the roadmap, but not yet in your event planning.
Sign language and subtitle-only variants
These do not receive a synthetic voice. For accessible livestreams, sign-language interpreters and subtitles remain the right tools — the AI language track complements them, it does not replace them.
Does it replace the interpreter?
The honest answer is: on quality no, in practice often yes — and only both halves together give a usable picture.
On quality the research is unambiguous, and it favours the human. A comparative case study by Prof. Anja Rütten found 44 major to substantial content deviations in a half-hour speech segment under machine translation, against 5 under human simultaneous interpreting (2026). A University of Warsaw study (Korybski et al., 2025) compared professionals, interpreting students and two AI systems under realistic conference conditions: the professionals achieved the highest quality, while the AI systems produced considerably more meaning-altering errors along with speech-recognition and context deficits. Kayo Matsushita showed in 2026 that listeners retained more content with professional interpreting and found the AI version more cognitively demanding. The interpreters’ association AIIC names latency, faulty sentence segmentation, confusion at speaker changes and hallucinations as the typical weaknesses.
Anyone turning that into “AI replaces interpreters” is selling you something. Anyone turning it into “AI is useless” is ignoring the arithmetic above.
Because in most cases the honest benchmark is not the interpreter — it is no translation at all. The AI language track replaces the interpreter above all where the interpreter was never booked for cost reasons: the town hall meeting across twelve countries, the internal all-hands, the partner webinar, the product launch for the sales regions. For those formats the real alternative reads “everyone listens in English and half the room misses the nuance” — and against that, a translation with occasional inaccuracies wins comfortably.
The converse holds just as clearly. Where every word counts legally, financially or personally, the human remains the right tool. For the annual general meeting, the IR call with balance-sheet terminology, contract negotiations, crisis communications and anywhere a mistranslation becomes the story, book interpreters. We arrange them for you — and combine both within the same event when your languages carry different demands.
AI language track, subtitles or interpreters — which when?
| Criterion | AI language track | Live subtitles | Interpreters |
|---|---|---|---|
| Spoken translation | yes (voice-over) | no | yes |
| Cost per language/day | low three-figure amount | low | €1,800–2,800 |
| Lead time | flexibly up to the start of the event | short | weeks |
| Terminology can be deposited | no | yes | yes (briefing) |
| Multiple speakers distinguished | no | — | yes |
| Additional technology / personnel | none | none | booth + two people per language |
| Data processing | currently outside the EU | EU | EU |
| Suitable for AGM, IR call, negotiation | no | complementary | yes |
| Suitable for town hall, all-hands, webinar | yes | yes | yes, but rarely budgeted |
| Scaling across many languages | yes, one checkbox per language | yes | linearly more expensive |
Frequently asked questions
What does AI simultaneous interpretation cost compared to an interpreter?
A simultaneous interpreter costs €900 to €1,400 net per day at market rates, and simultaneous interpreting requires interpreters to work in pairs — roughly €1,800 to €2,800 per language for one event day. The AI language track in SlideSync sits in the low three-figure range per language and event. So an extra language does not cost a little less than a pair of interpreters. It costs many times less.
What technology do I need for it?
None beyond what you already have. One encoder and one signal into SlideSync — the same setup as any webcast. No additional device, cable or on-site personnel arises on the customer side, and no software has to be installed on the presenters’ machines.
How late can I still add a language?
Flexibly right up to the start of the broadcast — you do not have to commit weeks in advance. The booked languages are then ready; adding a language spontaneously during an already running broadcast is not possible.
Is the translation lip-synced?
No. The translated voice trails the picture by a few seconds, and the gap widens over longer passages. This matches how simultaneous interpreting behaves, where the interpreter also runs behind the speaker. For talks, presentations and Q&A the offset is unproblematic.
Can I deposit my own terminology or product names?
Not for the spoken translation — the technology does not provide for it. Terminology management is available for live subtitles, not for the voice. If terminology has to land word for word, the combination of original audio and maintained subtitles is the better route, or a briefed interpreter.
Can I combine human interpreters and AI languages in the same event?
Yes. You decide per language between the two — French as an AI track and English with an interpreter, for example. The assignment is fixed from the start of the broadcast, however: an interpreter who is only meant to join during a running show can no longer be brought in.
Is my audio processed in the EU?
For the AI language track, currently not — the audio is processed outside the EU because Google still runs the model as a preview. The SlideSync platform itself hosts in the EU, and all languages without an AI track are unaffected. If the data has to stay inside the EU, book an interpreter for that event or wait for European availability.
Which languages are available?
SlideSync currently carries 18 languages. Productively proven so far are French, Italian, Spanish, German and Chinese, each out of English. We test further language pairs once before your event so nothing surprises anyone on the day — just come to us with your language list.
What happens if a speaker switches language mid-presentation?
The switch is detected without anyone intervening. If someone in an English-language event suddenly speaks German, German viewers hear it correctly; English viewers currently hear the passage untranslated. The reverse direction is technically possible and is being implemented on our side.
What is the difference between AI subtitles and an AI language track?
AI subtitles show the text on screen. The AI language track speaks the translation aloud. Subtitles are the right choice when technical terms have to be exact — terminology can be deposited there. The language track is the right choice when your viewers should listen rather than read along, for instance during long presentations. Both can be combined in the same event, and in SlideSync the subtitles carry the same wording as the spoken translation.
How do I make a livestream multilingual?
You need one audio track per target language, which your viewers select in the player. That track is delivered either by a pair of interpreters or by the AI. The video stream stays the same, and so do your encoder and your signal. In SlideSync you book the languages before the event; the language selector then appears in the player automatically. More on our page about multilingual events.
Is this suitable for investor relations and quarterly communications?
For the call itself we advise against it. Balance-sheet terminology has to land word for word, and that is precisely where the AI language track cannot take on your terms. Book interpreters for IR calls and annual general meetings. The AI language track makes sense around them: internal explanations for sales regions, employee briefings on the quarterly result, partner updates.
Conclusion: the right moment to try it
AI simultaneous interpretation today is neither a finished product nor a gimmick. It is a technology in preview that already does one thing well enough to use productively: making an event audible in several languages without anyone booking a booth.
Where we recommend it:
- Internal large formats across several countries — town halls, all-hands, leadership updates. Wherever the second language previously failed on budget.
- Partner and customer formats — webinars, product launches, training for several sales regions.
- Formats with unclear language demand — if you do not know whether anyone will use the French track at all, finding out costs a low three-figure amount instead of a pair of interpreters.
Where we advise against it: annual general meetings, IR calls, contract negotiations, crisis communications — and any format where terminology has to land word for word or the data has to stay inside the EU. For those we arrange interpreters, and both can be combined within the same event.
The preview status and the processing outside the EU are the two points that will foreseeably change — Google services take this route regularly, and regional availability usually follows general release promptly. Until then this is above all one thing: an unusually good moment to run your own event series multilingually and see who actually switches. Nobody has that number yet, because the question could never be asked.
Wondering which languages make sense for your next event? Talk to us — we advise vendor-independently, and we will tell you when an interpreter is the better choice.
Try SlideSync!
Let’s talk about your event!