Google is making a distribution play for on-device AI, not proving that phones now run serious LLM workloads. Tensor ML SDK Beta ties PyTorch/TFLite conversion, compilation, deployment, inference, Play Feature Delivery, AI Packs, and CPU/GPU fallback into LiteRT. That plumbing matters because edge ML usually dies on packaging, runtime support, and device fragmentation, not on a single demo latency number.
The 100+ model garden sounds broad, but the hard examples are Gemma 3 1B, Function Gemma 270M, and EmbeddingGemma 300M. That is useful for local actions, semantic features, camera tricks, and speech flows. It is not a cloud-agent replacement. Apple keeps its on-device path tighter and more closed; Qualcomm’s NPU story still leaves developers stitching vendor pieces together. Google’s advantage is the LiteRT + Hugging Face + Play delivery loop. Performance, power draw, and Pixel 10 install base are not disclosed, so the victory lap is premature.