Skip to content
r/LocalLLaMA

AndroidLife: Can an AI agent survive a day in the life of a real user? Qwen3.8-27b run: 56.7% SR

A Reddit post introduces AndroidLife, a benchmark that tests whether an AI agent can survive a day of real user tasks. The Qwen3.8-27b model achieved a 56.7% success rate. The post body is blocked by Reddit, so no test details, task list, or model comparisons are disclosed.

Read the original ↗Export Markdown