How we measure translation latency
Every latency number on this site comes from the same harness, run against the same fixed corpus, and is republished when it moves. This page describes exactly how, so the figures can be checked instead of taken on trust.
What we measure
- Time to first audio
- From the moment the speaker's utterance ends to the first sample of translated audio leaving the server. This is the delay a listener actually perceives as a pause.
- Completion time
- From end of utterance to the last sample of translated audio.
- Glossary output
- The translation the engine actually produces for a fixed list of business terms โ recorded verbatim, not scored by a model.
p50 and p95, never averages
An average hides the tail, and in a conversation the tail is the only part anyone notices: one 4-second pause in twenty exchanges is what people remember, not the mean. Both percentiles are published for every pair.
Direction matters
Each direction is measured separately. English to Japanese and Japanese to English are different problems โ verb-final target languages force the pipeline to buffer more of the sentence before it can emit anything โ and they get different numbers and different pages.
Freshness
Measurements expire after 180 days. When a file passes that age the site fails to build rather than serving a figure nobody has re-checked.
No measurement run has been recorded yet. The corpus, hardware and sample sizes appear here as soon as the harness has produced them โ this page will not carry estimated figures.