Live translation only works if something processes your speech, which means your meeting audio leaves the room. That is not a scandal — it is how the technology works — but it does mean the questions below belong in your evaluation, not in a security review six months after rollout. Your meetings carry unreleased roadmap, customer names, salary discussions and pricing. Ask.

1. Is raw audio stored, or only processed?

There is a large difference between "we process your audio" and "we keep your audio". Ask specifically whether the raw stream is retained anywhere after the sentence is transcribed, and for how long.

VoxTranslate: speech is streamed for transcription and translation while you talk and is not stored. What can persist is text, not audio.

2. Who else is in the chain?

Nobody in this category runs the entire stack alone: speech recognition, translation and speech synthesis are usually separate providers. Ask for the named sub-processor list — not the phrase "trusted partners" — and check it against your own DPA obligations.

VoxTranslate: which provider is involved depends on the engine tier chosen for the session — speech recognition, translation and speech synthesis are separate services, and on one tier the browser streams audio directly to the provider without passing through our server at all. The providers, and what each one receives, are published in the privacy policy.

3. What happens to transcripts?

Transcripts are the part that actually accumulates, and they are a written record of everything said. Ask whether they are stored by default, who can read them, whether they can be exported, and whether they are deleted with the account.

VoxTranslate: transcripts are stored for calls in which a signed-in user took part, so participants can review, export (PDF/JSON) and correct them afterwards. Calls with only guests are not stored. Stored transcripts are deleted when you delete your account.

4. Does the video pass through their servers?

This one is easy to skip and worth asking, because architecture beats policy: media that never reaches a server cannot be retained by one.

VoxTranslate: in meetings, video and audio travel peer-to-peer between browsers over WebRTC and are not routed through or recorded by the server. When a direct connection is impossible, media is relayed encrypted in a form the relay cannot read. The server handles sign-in, signaling, the speech-to-text stream, translation and chat relay.

5. Is recording opt-in — and can participants tell?

Ask whether recording and transcript retention are per-session switches or account-wide defaults, and what participants see when either is on. In several jurisdictions notice is not a courtesy, it is the law.

VoxTranslate: for webinars, video recording and transcript capture are per-session options rather than always-on defaults.

6. Can you actually get data deleted?

"You can request deletion" and "here is the button" are different products. Ask what a deletion covers — account, transcripts, past sessions — and how long it takes.

VoxTranslate: deleting your account deletes your stored transcripts and your utterances along with it.

For EU deployments, ask which legal basis is claimed for processing meeting audio, where processing happens, and how international transfers are handled. A vendor that cannot answer quickly has not thought about it.

VoxTranslate: processing audio for live captions and translation is covered by contract plus the consent given at sign-up; the details and the transfer position are in the privacy policy rather than paraphrased in marketing copy.

Two questions worth adding for your own side

  • Who in your company can read a transcript? The vendor's answer is only half of it. A stored transcript of a compensation discussion is a document your own access rules now have to cover.
  • What should not be translated live? Decide the exceptions in advance — legal proceedings, incident calls, anything under embargo — rather than during them.

A note on AI output

Transcription and translation are AI-generated and can contain mistakes, and spoken translation is synthesized rather than a recording of the speaker. That has a privacy implication people miss: an imperfect transcript is still a record, and it can misattribute or garble what someone said. Treat transcripts as useful, not as evidence.