Google Ships Sign Language AI to Phones: SL2T Translates ASL to English
Google's SL2T model, trained on 100,000+ hours of signing data, brings ASL-to-English dictation to Gboard and Live Transcribe on Pixel 11 — the first consumer deployment of sign language AI.

Updated
Why it matters
- SL2T is trained on over 100,000 hours of data across more than 50 sign languages, with roughly a quarter in ASL, and achieves a zero-shot 70 BLEURT on FLEURS-ASL (sd-test).
- The model powers ASL-to-English sign-to-text dictation in Gboard and Live Transcribe on Pixel 11 at no additional cost, with more devices and languages planned.
- Privacy-preserving design: on-device MediaPipe Holistic tracking sends only pose landmark coordinates to the server, discarding the original video; the project is overseen by the AI Sign Language Advisory Committee (AISLAC).
Google has moved sign language AI out of the lab and into consumer products for the first time. A new sign-language-to-text (SL2T) translation model, trained on more than 100,000 hours of data across over 50 sign languages, now powers sign-to-text dictation in Gboard and Live Transcribe on Pixel 11, starting with American Sign Language (ASL) to English. Additional languages and devices will follow. The feature costs nothing extra.
The release targets a population that automatic translation, dictation and conversational interfaces have largely bypassed: an estimated 70 million Deaf and hard of hearing people who use the world's more than 200 sign languages. While speech technology has advanced rapidly for decades, sign language AI has lagged, and Google frames the new model as a step toward parity between signed and spoken language access — a market and accessibility gap with significant social stakes.
A dictation feature for signers
The feature works like dictation for hearing users, but with signing instead of speech. Deaf users can sign to their phone anywhere they would normally type: to search the web, draft messages or documents, or ask Gemini to solve queries and execute tasks. In Live Transcribe, users can sign responses in conversations rather than typing back and forth.
According to Google's testers, signing in ASL is "faster, more natural, and more delightful" than typing in English.
The stakes go beyond convenience. Sign languages are the primary languages of Deaf communities worldwide and the cornerstone of Deaf cultural identity. Google notes wide diversity among Deaf people in proficiency across signing, speaking, reading and writing, so supporting access in all modalities matters. The technology also opens possibilities for bridging communication between Deaf and hearing communities.
Why sign language translation is hard
Google identifies two core challenges that distinguish the task from ordinary speech transcription. First, transcribing speech is a sequential mapping from sound to text in the same language. Sign languages, by contrast, are independent natural languages with their own grammars and lexicons, requiring true machine translation rather than sign-to-word transformation.
Second, the model must understand physical movement. Sign languages convey meaning through simultaneous movements of the hands, arms, torso, head and face. Tracking these at high frame rates is a computationally demanding computer vision problem.
That background explains why early efforts like sign language gloves fell short. "Sign languages aren't simply 'English on the hands,'" Google writes. They require fine-grained whole-body visual perception plus full-fledged language translation. SL2T is built to do both.
How SL2T works
Google combined a user-centric, culturally informed approach with large-scale data. Training spans more than 100,000 hours across more than 50 sign languages, with roughly a quarter of the data in ASL. Joint training on diverse languages, dialects and proficiency levels teaches the model shared underlying structures; in Google's experiments, this multilingual training outperforms single-language models.
Privacy shaped the architecture. SL2T never sees a raw camera feed. An on-device model, MediaPipe Holistic, tracks pose landmark locations on the signer, and only those geometric coordinates go to the server for translation. The original video can be discarded immediately.
SL2T translates the coordinate sequence directly into text, skipping the intermediate annotations called "glosses" that dominate prior sign language translation research. Glosses fail to capture rich, non-linear aspects of sign languages such as non-manual markers and spatial constructions. Direct landmark-to-text translation removes artificial vocabulary limits and lets quality scale with data.
Benchmark results — and practical fixes
On FLEURS-ASL (sd-test), a benchmark for ASL-to-English translation quality, SL2T achieves a zero-shot score of 70 BLEURT, which Google says is significantly higher than any previously reported score. Google calls it the most capable sign language translation model to date.
The company also worked on issues that academic benchmarks don't capture: minimizing streaming latency, preventing hallucination on non-signing inputs, ensuring fairness for the roughly 10% of signers who are left-handed, and improving performance for one-handed signing — the mode people use while holding a smartphone in the other hand.
Sample translations published by Google show fluent, natural output. Given "The Cook Islands do not have any cities but are composed of 15 different islands. The main ones are Rarotonga and Aitutaki," SL2T produced: "The Cook Islands have no cities and consist of 15 islands. The two main islands are Rarotonga and Aitutaki." The model does stumble: for a passage about a feathered, warm-blooded bird of prey, it rendered "eats prey" as "eats grey." Overall, the outputs preserve meaning while smoothing English phrasing.
Built with the Deaf community
Google positions the project as built with the Deaf community, not just for it. The concept originated with Sam Sepah, a Deaf Googler. Deaf partners contributed to data collection, Deaf user studies handled evaluation, and Deaf experts assessed impact.
To govern real-world deployment, Google established the AI Sign Language Advisory Committee (AISLAC), which brings together global Deaf organizations and subject-matter experts. Through that participatory model, the communities most affected by the technology directly influence development priorities. Google and its partners co-authored a joint impact report for the SL2T 1.0 release in Gboard and Live Transcribe, detailing capabilities and current limitations — an approach Google says it will repeat for all major sign language releases.
The work was done jointly by teams from Google DeepMind and Android. The core model team includes Garrett Tanzer, Benoit Brard, Elizabeth Clark, Tim Dozat, Sebastian Ebert, Dan Garrette, Manfred Georg, Vicky Holgate, Shankar Kumar, Mohammad Saboorian, Miloš Stanojević, Megh Umekar, John Wieting, Andy Zhang and Chris Dyer, with a separate Android team handling the Gboard and Live Transcribe integration.
What comes next
ASL input on phones is only the start. Google says its team is working to expand the technology into additional sign languages, sign language generation, and frontier AI capabilities, with the stated goal of making sign language access standard across digital products and achieving full parity with spoken and written languages. How quickly that expands beyond English — and beyond Pixel hardware — will determine whether SL2T becomes the baseline for sign language computing or a single-vendor novelty.
Original: store.google.com
More from Marcus Bennett
Show full bio
Senior reporter covering consumer brands and retail at AI In Context.
117 articles
Related articles
- Google Launches Gemini 3.5 Live Translate Across Products
- Google Ships Gemini 3.8 TTS Models With Voice Cloning and Direction
- Google Launches Gemini 3.1 Flash TTS With Audio Tags and 70+ Languages
- Google Rolls Out Gemini 3.8 Live with Live Avatar for Enterprise
- Google Ships Upgraded Gemini 2.5 Flash Native Audio and Live Translation