Measured data · first-party · published · measured 2026-09-26

How accurate is Apple's on-device translation? We measured it on 11 languages to and from Chinese.

Apple's on-device translation is about as accurate as a cloud large language model when translating into Chinese, a little behind it when translating out of Chinese, and it did not drop a single number in 440 test sentences. That is the result of one measured run on 2026-09-26: the Translation framework on macOS 26.5.2 against GLM (glm-5.3-flash, a cloud model) on 100 FLEURS sentences per language, plus 20 number-heavy sentences per language in each direction. Into Chinese, the two models' chrF scores were within 0.016 of each other on 7 of 8 languages, and Apple was ahead on English by 0.053. Out of Chinese, the cloud model was ahead on all 8, by 0.018 to 0.061. Apple translated one sentence in 48 to 87 ms. Apple's on-device speech recognizer, SpeechTranscriber, transcribed the same read speech with a 4.1% to 9.2% error rate. We are Voice Translator, a Mac and iPhone app that captions meetings in Chinese, and we ran these numbers to decide which languages to ship. Every table below says how it was measured and what it does not show.

Apple on-device vs cloud LLM translation, 8 languages, FLEURS, 2026-09-26
LanguageRecognition error (SpeechTranscriber)chrF into Chinese · Apple / cloudchrF out of Chinese · Apple / cloudNumbers lost · Apple (in / out, of 20)Apple per sentence p50 (in / out)
English8.9% WER0.419 / 0.3660.580 / 0.6410 / 059 / 48 ms
French7.3% WER0.331 / 0.3310.573 / 0.6000 / 077 / 76 ms
German7.6% WER0.335 / 0.3370.520 / 0.5430 / 084 / 83 ms
Spanish5.2% WER0.294 / 0.3100.484 / 0.5020 / 086 / 84 ms
Portuguese (Brazil)9.2% WER0.350 / 0.3510.549 / 0.5710 / 080 / 76 ms
Italian4.1% WER0.321 / 0.3250.491 / 0.5250 / 085 / 74 ms
Japanese6.7% CER0.286 / 0.2800.316 / 0.3630 / 049 / 57 ms
Korean5.7% CER0.325 / 0.3180.293 / 0.3160 / 049 / 49 ms

One Apple M3 Max, macOS 26.5.2, Translation framework language packs installed. FLEURS test set, the first 100 sentences of each language that also exist in Chinese. WER for Latin-script languages, CER (per character) for Japanese and Korean, pooled over the corpus. chrF is the sentence average (character n-grams up to 6, β = 2); compare along a row, not across rows. Cloud = GLM glm-5.3-flash with our production prompts.

Is Apple's on-device translation as good as cloud translation?

Into Chinese, yes, on this test. Apple's on-device model and the cloud LLM scored within 0.016 chrF of each other on French, German, Spanish, Portuguese, Italian, Japanese and Korean, and Apple scored 0.053 higher on English.

Out of Chinese, the cloud model was ahead on every language, by 0.018 (Spanish) to 0.061 (English). If you are writing Chinese and need the other language to read naturally, the cloud model is the better tool on this evidence.

The test sentences matter. FLEURS is read Wikipedia-style prose: long, well-formed sentences. On short spoken meeting sentences, an earlier run of ours (2026-09) had the cloud model ahead into Chinese as well, at 0.45–0.50 chrF against Apple's 0.377. Neither result transfers automatically to the other kind of text.

We did not measure Google Translate, DeepL or Apple's cloud mode. This page compares one on-device model with one cloud model on one machine.

Does on-device translation keep numbers, dates and amounts?

Apple's Translation framework kept every number in our zero-tolerance set: 20 sentences per language in each direction, each carrying a date, an amount, a percentage or a phone number, across 11 languages — 440 sentences, 0 numbers lost. The same check caught the cloud model dropping a number in 2 of the 160 into-Chinese sentences on the 8 Mac languages: an Italian "2,81 seconds" rounded to 2.8, and a Korean date that lost its year.

The cloud model also did something no accuracy score would flag. Asked to translate the French "La remise est de 30 %" into Chinese, it wrote 打七折 ("at 70% of the price") — correct in meaning, but the 30 is gone, and a reader checking figures against a contract would not find it. We fixed that with one line in the prompt; the example is why our apps check every number after translation instead of trusting the score.

Formatting is the second trap. A Chinese reader sees the Italian "2,81" as two thousand eight hundred and ten and the German "1.500" as one and a half. Our products rewrite a restored number in the reader's convention (2.81, 1,500) and copy phone numbers verbatim. The same traps, with a converter, are on our big-number converter (Chinese).

How accurate is Apple's on-device speech recognition?

SpeechTranscriber, the on-device recognizer in Apple's SpeechAnalyzer API on macOS 26, transcribed FLEURS read speech at 4.1% (Italian) to 9.2% (Portuguese) word error, and 6.7% and 5.7% character error for Japanese and Korean. Most of the errors we inspected were formatting, not mishearing: "11:35 pm" written as "1135 p.m.", "it is" as "it's".

For comparison we ran NVIDIA's Parakeet v3 (the model our iPhone app uses for European languages, through FluidAudio) on the same sentences:

LanguageSpeechTranscriber WERParakeet v3 WER
English8.9%6.94%
French7.3%6.08%
German7.6%6.29%
Spanish5.2%5.96%
Portuguese9.2%6.46%
Italian4.1%2.86%
Dutch · Polish · Russian · Ukrainian—6.91% · 7.79% · 8.58% · 8.32%

Parakeet had the lower error on five of the six shared languages; Apple's recognizer was lower on Spanish. Parakeet does not cover Japanese or Korean, which is why our iPhone app uses SpeechTranscriber for those two.

Dutch, Polish and Russian: translation only

Three more languages went through the same Translation framework checks for our iPhone app. Apple's translation into and out of Chinese scored 0.313 / 0.511 chrF for Dutch, 0.290 / 0.438 for Polish and 0.308 / 0.472 for Russian, with 0 numbers lost in either direction and 74–87 ms per sentence on the Mac. We did not run the cloud model on these three.

"Too many allocated locales, 5 maximum": what we hit while measuring

An app can hold speech-recognition assets for at most five locales at once. On macOS 26.5, AssetInventory.maximumReservedLocales is 5; after our measuring tool had reserved en-US, fr-FR, de-DE, es-ES and es-MX, downloading pt-BR, it-IT, ja-JP or ko-KR failed with SFSpeechErrorDomain code 11, "Too many allocated locales, 5 maximum."

Three details that are not obvious from the documentation. The limit is per app: a second app on the same Mac still had its own five slots, and it could use assets the first app had already downloaded. To release a locale you have to pass the Locale object that reservedLocales gives back; a freshly constructed Locale(identifier:) for the same identifier returns false and releases nothing. And the language packs are large: installing ten Chinese translation pairs pulled about 4.9 GB, of which 953 MB were the translation models and 3,868 MB were speech-recognition assets the system added.

Apple's references: SpeechAnalyzer and the Translation framework.

How we measured, and what this does not show

  • Sentences. FLEURS (Google, CC-BY-4.0) test split: the same sentence id is the same FLORES sentence in every language, so one set gives the audio, the source text and the Chinese reference. We took the first 100 ids each language shares with Chinese, and the first reading of each (not the best one).
  • Recognition. Each WAV file transcribed on device; error rate = total edit distance ÷ total reference length over the corpus, after normalization, per word for Latin and Cyrillic scripts and per character for Japanese and Korean.
  • Translation. Both directions from the FLEURS reference text (not from the recognizer output), scored with sentence-level chrF against the FLEURS translation.
  • Numbers. Our own zero-tolerance set, 20 sentences per language per direction; every translation went through the same number check our apps run, and each flag was read by a person to separate real losses from false alarms.
  • Latency. Per-sentence p50 of the Translation framework on the Mac. It is not iPhone latency; phones are slower.
  • Not shown. Accents, crosstalk, noisy rooms, domain vocabulary, other Macs, iPhone timing, and any engine we did not name. One run, one machine — compare it with your own before you rely on it.

Our latency numbers for live meeting captions are on a separate page, measured caption latency (Chinese). How these numbers shape the product: Voice Translator turns a language on for a device only after it passes our lines: recognition error under 15%, zero numbers lost on screen, and end-of-speech latency within 0.5 s of English.

Questions people ask

How accurate is Apple translation?+

On our 2026-09-26 test, Apple's on-device Translation framework scored within 0.016 chrF of a cloud LLM when translating French, German, Spanish, Portuguese, Italian, Japanese and Korean into Chinese, 0.053 higher on English, and 0.018 to 0.061 lower when translating out of Chinese. It lost no numbers in 440 number-heavy sentences.

Does Apple's translation work offline on a Mac?+

Yes, once the language pair is downloaded: every number on this page comes from on-device translation with the packs installed. The download is the catch. Ten Chinese language pairs took about 4.9 GB on our Mac, most of it speech-recognition assets the system added alongside 953 MB of translation models.

How fast is on-device translation?+

On an Apple M3 Max with macOS 26.5.2, the Translation framework took a median of 48 to 87 ms per sentence, depending on the language and direction. That is fast enough that recognition, not translation, sets the pace of live captions. We have not yet measured iPhone timing.

Which is more accurate, SpeechTranscriber or Parakeet?+

On the same 100 FLEURS sentences, Parakeet v3 had the lower word error rate in English, French, German, Portuguese and Italian, and Apple's SpeechTranscriber was lower in Spanish (5.2% against 5.96%). Parakeet does not cover Japanese or Korean; SpeechTranscriber scored 6.7% and 5.7% character error on those.

Can I reproduce these numbers?+

The sentences are public (FLEURS test split, first 100 ids shared with Chinese) and the metrics are standard, so the recognition and chrF columns can be rerun on any Mac with macOS 26. The zero-tolerance number sentences are our own set; write to us if you want them.

Related: translated captions in Teams, Zoom and Meet, compared · speech time calculator · what Voice Translator is