Google's DeepMind has released a new multilingual sign language to text model, SL2T, and integrated it into the Pixel 11 phone system, marking the first time sign language AI has been introduced into mainstream consumer electronics. The model is first integrated with Gboard keyboard and Live Transcribe, allowing users to edit messages, search the web, and converse with Gemini by performing sign language gestures through the front-facing camera, replacing traditional keyboard input.
SL2T was trained on over 50 sign languages, with a total of 100,000 hours of sign language data, of which about a quarter comes from the American Sign Language dataset. Cross-lingual joint training helps the model learn common action logic across different sign languages and has set a new record for sign language transcription models in the FLEURS-ASL evaluation. Unlike traditional solutions that rely on fixed vocabulary tags, SL2T can directly interpret human movements to generate text, while also processing non-hand semantic information such as facial expressions, body position in space, and lip movements.
Regarding privacy, the phone-side MediaPipe Holistic tracks 130 key points on hands, face, and torso in real time, transmitting only movement coordinates externally, with original video being immediately deleted and not uploaded to the cloud. Currently, this feature supports only American Sign Language to English translation, and Google plans to expand to more sign languages and Android devices in the future.
DeepMind previously open-sourced a sign language translation model, SignGemma, in May 2025, providing an algorithmic foundation for the implementation of SL2T. In addition, its WaveNet, Euphonia, and Live Transcribe accessibility technologies have accumulated experience in multimodal audio and visual interaction. As Google, iFLYTEK, Huawei, and Apple continue to invest, AI accessibility interaction is accelerating from research validation to consumer-level products.


