With Dazzle, Marissa Mayer bets your camera roll has more info on your life than your inbox
前 Yahoo CEO Marissa Mayer 推出个人 AI 助手 Dazzle,去年 12 月完成 800 万美元种子轮融资,其上下文来源只有手机相机胶卷,而非邮件和日历。Dazzle 可通过 App 或短信交互,能扫描近期照片填充日历、根据照片库推荐旅行和礼物,并声称会丢弃被标记为敏感的个人信息。
前 Yahoo CEO Marissa Mayer 推出个人 AI 助手 Dazzle,去年 12 月完成 800 万美元种子轮融资,其上下文来源只有手机相机胶卷,而非邮件和日历。Dazzle 可通过 App 或短信交互,能扫描近期照片填充日历、根据照片库推荐旅行和礼物,并声称会丢弃被标记为敏感的个人信息。
Claude Sonnet 5.5 beats Sonnet 5 on every benchmark, runs 30%+ faster, and costs up to 30% less for most work. The big move: it's now the free-tier default on claude.ai, which Simon Willison tested and got a solid WebGL 3D pelican on a bicycle. The 'max' thinking effort still hits the same bug as Opus 5.5—128K tokens of thought with no output, costing $1.28. 'xhigh' delivered a decent SVG in 41 seconds for 5.74 cents. Anthropic says Haiku 5.5 is coming in weeks; Simon hopes it's price-competitive with GPT-6 Luna.
Why it matters: Putting the latest Sonnet on the free tier is a real product strategy shift, not a routine model update. Simon's hands-on test delivers concrete numbers ($1.28 burned, 5.74 cents for the working render, 41-second latency), and the max-mode bug matching Opus 5.5 is a useful sig...
Anthropic launched Claude Sonnet 5.5, claiming over 30% speed gains and clearer writing for fast-turnaround tasks like bug fixes, docs, and slide decks. Opus 5.5 targets complex judgment work, and Haiku 5.5 is coming in a few weeks. The post doesn't disclose pricing or latency numbers.
Why it matters: Anthropic model line refresh with a concrete 30% speed claim for Sonnet 5.5 and clear product-line differentiation. Held below 85 because the post doesn't disclose pricing, latency benchmarks, or the baseline for the 30% figure.
404 Media obtained internal documents showing Microsoft hires at least hundreds of contractors to review Copilot users' prompts and uploaded images. They are not filtering for safety—they judge output quality, such as whether AI-enlarged breasts are big enough. Reviewers are flooded with sexual requests: shortening skirts, foot fetish images of children's cartoon characters, pro-anorexia content. Uploaded faces are never blurred, and many prompts are dubiously consensual. One contractor said it's hard to take the work seriously when the focus is which model generated the right bust size.
Why it matters: 404 Media obtained internal documents with solid evidence. The story exposes how Copilot's content review actually works, hitting all three HKR axes. Score capped below 85 because it's a single investigative piece, not a product launch or model release — industry shake-up is l...
The author built a Python wrapper inspired by Jev that makes LLMs answer multiple-choice questions quickly by forcing single-token output and reading logprobs. It also handles images. On an RTX 3090 with Gemma 4 12B, it processes webcam frames at 1 FPS with three questions per frame (person visible, indoor/outdoor, brightness). OpenAI gpt-6-luna runs at 0.2 FPS due to connection overhead per request. The tradeoff: specialized CV models are faster, but LLMs let you change conditions by editing plain text. Code supports llama.cpp and OpenAI backends.
Ben Tossell spent a morning building a site with Codex and Factory that shows 87 iconic devices from 1976 to 2026. All images were generated by Astra, using 40 messages and 7 subagents, and it has received 1,455 votes so far. The project started from a 'forgotten devices' idea, took inspiration from Cole's timeline scrubber demo, and had agents find reference photos, generate new images, and fill gaps after 2009. The post doesn't detail each device's model or generation cost.
PrismML showed a 1-bit Bonsai LLM at Qualcomm's Snapdragon Summit that runs locally on smart glasses using the Snapdragon AR1 Gen 1 platform. The 2-billion-parameter model handles vision and language so wearers can ask about what they see in real time. PrismML shrinks larger models by 4x while keeping nearly all benchmark performance. The startup's bigger aim is open-weight on-device AI that doesn't depend on cloud labs' privacy promises. No smart glasses shipping with PrismML have been announced yet.
Google DeepMind released Gemini 3.8 Live with Live Avatar, adding near-real-time video generation to its native real-time conversation model. The result is a dynamic visual avatar with lip sync, natural expressions and smooth turn-taking.
Why it matters: The post details Live Avatar's real-time video conversation, async tool calls and 97-language support, a useful read on enterprise multimodal interaction.
Meta teased the Muse Charm at the end of Connect, a dedicated hardware device for its Muse AI agent. It resembles a chunky strapless smartwatch with a lanyard. A fingerprint sensor on the top right activates voice input; the front has at least three mic holes and a small camera. Zuckerberg noted you don't need to unlock a phone to use it. The post doesn't disclose pricing, battery life, or a release date.
Why it matters: Meta teased a standalone Muse AI gadget at the end of Connect — a thick watch-face on a lanyard with fingerprint wake, voice, and a camera. Only looks and interaction logic are disclosed; no price, battery, or launch date, so substance is thin and the score sits right at the f...
Qwen released Qwen-Image-2.1 on Sep 20, a 7B open-weight image model that unifies text-to-image, local editing, and native RGBA output in a single pipeline. The version number rolled back from 3.0 to 2.1 reflects a 2026 split: 3.0 is a closed-source commercial API, while 2.1 continues the open research branch. A built-in RGBA VAE outputs PNGs with transparency, skipping external matting. The interface supports up to 10 reference images and three mask types at native 2K. The license shifted from Apache 2.0 to a research-only agreement; commercial use requires a separate license. Community tests show ~25s per megapixel image on RTX 5070/5080 at 25 steps, ~15.6GB VRAM with Q8 quantization. Text rendering remains a strength, but multi-subject consistency shows facial generalization on well-known public figures—official demos don't guarantee universal performance.
Why it matters: Qwen open-sourced a model that combines image generation, editing, and native transparency output into one pipeline — a clear engineering increment, not a reskin. The backward version jump is inherently clickable, and it resonates with both designers and developers. Not scorin...
Two weeks after launching Muse, Meta says it's working on bringing the agent to its smart glasses, including the new ones shown at Connect. You'll activate it by saying its name and can ask it to guide workouts, log meals, or help shop for products you're looking at. The glasses are also getting an FDA-cleared hearing enhancement feature for adults with mild to moderate hearing loss. The post doesn't specify a launch date or which models will get Muse.
Why it matters: Putting Muse on glasses is a key step in Meta's push to move AI assistants from phones to wearables, with three concrete use cases. But the post doesn't give a launch date or supported models, so the score sits right at the featured threshold.
Meta announced two key updates for Muse AI today: each agent gets its own email address to send/receive and complete tasks, and the Mac app will gain computer-use abilities. Muse will also support video calls soon, moving beyond text chat. The post doesn't disclose launch dates or pricing.
Xiaomi open-sourced two multimodal models: MiMo-V2.6 Pro and Flash. Pro scored 46 on the Artificial Analysis Intelligence Index—the highest among open-source models—and matches Claude Opus 5 and GPT-5.6 Sol on most agent benchmarks. The post doesn't disclose parameter counts, training cost, inference latency, or the exact open-source license, so I'd hold off on production assumptions for now.
Why it matters: Xiaomi open-sourced MiMo-V2.6 Pro, matching Claude Opus 5 and GPT-5.6 Sol on agent benchmarks and hitting the highest open-source score on the Intelligence Index. Domestic flagship model release gets full weight per policy. Missing parameter count is a gap, but the signal is s...
RxFilm Studio is a native macOS app that packs the entire product-video pipeline into one window. An AI agent acts as the "director" — it generates music cues via Lyria, multi-speaker narration, auto-captions with translation (exportable as VTT/SRT), still images, and final 4K 60fps renders via Remotion. Every edit is proposed for review before it lands. Version 1.9.0 is free and Apple Silicon only. The post doesn't specify which model Lyria is or how many languages the narration supports.
Ant Group open-sourced the Ming-Image-0.1-Design series, which includes two 6B-parameter models for design generation and layer editing. The body is unavailable due to a page error, so details like model architecture, training data, or benchmarks are not disclosed. What's confirmed: the models are open-source and aim to cover the full pipeline from design generation to layer editing.
Alibaba's Qwen-Image-2.1 ranks first among open-source models on Arena's Image Edit and Text-to-Image leaderboards. It scored 1367 in Image Edit Arena, placing 16th overall, just 3 points behind GPT-Image-1.5-high-fidelity at #15. The post doesn't disclose parameter count, architecture, or release timeline.
Why it matters: Qwen-Image-2.1 hitting #1 open-source on Arena's image editing leaderboard, just 3 points behind GPT-Image-1.5, is a concrete cross-model signal. Score stays at 78 rather than higher because the post doesn't disclose parameter count, architecture, or release timeline — the inf...
OpenRouter shortlisted embedding models from its 37-entry catalog for English RAG, multilingual, code, and text-image retrieval. The default pick is OpenAI text-embedding-3-small for its low price and 8,192-token context. For longer inputs, Voyage 4 large offers a 32,000-token window and index compatibility across Voyage 4 tiers. Qwen3-Embedding-8B is recommended for multilingual retrieval with public weights and 100+ language support. Code search goes to Voyage Code 4, while Gemini Embedding 2 and Voyage Multimodal 3.5 handle text-and-image. The free route is Nvidia Nemotron-3-Embed-1B; the cheapest paid option is Perplexity pplx-embed-v1-0.6b at $0.004 per million tokens. OpenRouter notes these checks confirm API behavior, not retrieval quality, and advises testing on your own data before building an index.
Qwen-Image-2.1 is out with open weights. It scored 1367 on the Arena Image Edit Arena, ranking #1 among open-source models and #16 overall — just 3 points behind GPT-Image-1.5-high-fidelity at #15. It also landed #1 open-source on the Text-to-Image Arena. The post doesn't disclose parameter count, architecture details, or the exact open license.
Why it matters: Qwen-Image-2.1 open weights dropped, hitting #1 open-source on Arena's image editing leaderboard at #16 overall, just 3 points behind GPT-Image-1.5. Score held back because the post doesn't disclose parameter count, architecture, or license — we're grading on the leaderboard n...
Xiaomi released the MiMo-V2.6 series. The Pro version ranks #1 among open weights models on the Artificial Analysis Intelligence Index with a score of 46, at a training cost of $3M. A Flash variant targets efficiency, and an UltraSpeed variant offers 20x faster output. The technical report details RL scaling across three axes: larger batches and throughput, richer multi-task environments, and more grader compute. Code and training recipes are open-sourced, but the 7k+ task datasets are not yet released. Former DeepSeek engineer Fuli Luo, now at Xiaomi, previously live-streamed the training runs.
Why it matters: Xiaomi's MiMo-V2.6-Pro hit #1 on the Artificial Analysis open-weights leaderboard with a $3M training budget — price-performance right at the frontier. Flash and UltraSpeed variants cover efficiency and speed use cases, and the tech report details an async RL architecture. Not...
Xiaomi released two omni-modal models: MiMo-V2.6-Pro and Flash. The Pro version scored 1628 on Code Arena WebDev, up 153 points from MiMo-V2.5-Pro's 1475, landing around 10th overall and ~3rd among open-weight models under MIT license. The post doesn't disclose Flash's benchmark numbers or parameter counts.
Why it matters: Xiaomi's multimodal model hits ~10th on Code Arena WebDev and ~3rd among MIT open-weight models, with a 153-point gain for Pro. Flags a domestic flagship release with concrete benchmark data. Flash variant lacks params and scores, capping it below 80.
Tencent Hunyuan released Hy Image3.5 preview, an image generation model. The body only contains the title and navigation bar, with no details on capabilities, parameters, or release timeline. Wait for official disclosure.
Qwen-Image-2.1 is a 7B native image generation and editing model. It uses a single checkpoint for both tasks and supports up to 10 reference images. The model includes a built-in prompt-enhancement LLM, integrates with diffusers and ComfyUI, and offers a no-install browser demo on Hugging Face Spaces. The post doesn't disclose training data, inference latency, or benchmark comparisons.
Why it matters: Qwen drops an image model with a 7B single-checkpoint design for both generation and editing, plus a built-in prompt optimizer — a fresh combo. Score stays at 78 rather than 85+ because the post doesn't disclose training data, inference speed, or real image-quality comparisons...
Alibaba Qwen released open weights for Qwen-Image-2.1 with native ComfyUI support. A single 7B checkpoint handles both image generation and editing, outputs up to 2K natively, accepts up to 10 reference images per instruction, and supports RGBA with alpha channel. The post doesn't spell out license terms or hardware requirements.
Why it matters: Alibaba Qwen drops a 7B unified generation/editing model with native ComfyUI support, 2K output, and RGBA transparency — a direct win for the local image-gen community. Held below 84 because hardware requirements and license terms aren't disclosed, so real-world adoption is st...
StepFun introduces Step 5 Preview, a 600B-parameter MoE model with 27B active per token, a 1M-token context window, and vision support. It scores 67.7 on DeepSWE v1.1, ahead of Kimi K3 and GLM-5.3 but behind GPT-6 Astra and Claude Opus 5. On the in-house StepCodeBench it hits 49.0, again leading domestic models and trailing the two US labs. On FrontierFinance it reaches 66.4, second only to Claude Opus 5. Artificial Analysis gives it an intelligence index of 44; StepFun claims substantially lower cost per task at comparable intelligence. The post does not disclose API pricing, release timeline, or training details.
Why it matters: StepFun's Step 5 Preview is a 600B MoE model that edges out Kimi K3 and GLM-5.3 on coding benchmarks but still trails GPT-6 Astra and Claude Opus 5 by 6-7 points. Scored 78 because it's a substantive domestic model push in agentic coding with real numbers, but not industry-sha...
OpenRouter sent the same prompt to 20 image models and read the actual billed cost. GPT Image 2 was cheapest at $0.006 per 1024×1024 PNG; Gemini 3 Pro Image was priciest at $0.134—a 22x spread. Pricing units differ across providers (tokens, megapixels, per image), so side-by-side list prices mislead; generate once and check usage.cost. Five of six models rendered text correctly, including the cheapest. Recraft V4.1 Vector outputs editable SVG at $0.08. The post also details formats, resolution caps, and seed support per model.
Why it matters: OpenRouter ran one prompt through 20 image models and posted the actual bills — the kind of real cost data pricing pages never show. All three HKR axes hit: the headline pulls you in, the billing breakdown is genuinely new info, and it nails a daily pain point for builders. No...
PrismML open-sourced Ternary Bonsai 2 27B, a quantized version of Qwen3.8 27B that uses {-1, 0, +1} weights with FP16 group-wise scaling, hitting 1.76 bits per weight and a 5.9GB footprint — over 9x smaller than the original. It retains 98.2% of the full-precision model's aggregate benchmark score (83.9 vs 85.4), with particularly strong retention in coding, agentic tool use, and vision. Throughput reaches 143 tok/s on an RTX 5090 and 46.8 tok/s on M5 Max; on an RTX 4090 it draws 0.714 mWh/token, 40% more efficient than a full-precision 8B model. The model supports a 262K-token context window, multimodal input, and ships under Apache 2.0. The post does not disclose training data or the specific quantization distillation recipe.
Pinterest is testing Restyle, an AI feature that swaps furniture, decor, and lighting in user-uploaded room photos. It turns saved inspiration into shoppable items, bridging browsing and purchase. The post doesn't disclose launch date or supported markets.
DeepSeek released V4.1-Flash, a 552B MoE multimodal model. The key feature is KV cache compression, which cuts memory usage during long-context inference. The paper just hit arXiv and doesn't disclose compression ratios or benchmarks yet, but the title says 'Pushing the Limits' — this is about inference efficiency.
DeepSeek V4 is a model family, not a single model. OpenRouter's guide confirms only V4.1 Flash and V4 Flash Vision Exp accept image input; all others (V4 Pro 0813, V4 Flash 0731, etc.) are text-only. V4.1 Flash is the recommended choice with native vision support at $0.15/$0.60 per million tokens. V4 Flash Vision Exp is the pricier experimental option. The post also covers two integration methods: direct image input or using a separate vision model as a front-end.
Meta is putting AI image generation, editing, video generation, and Instagram's Restyle tool behind a new paid subscription called Meta One, covering Facebook, Instagram, and WhatsApp. It follows Meta's $14.3 billion investment in Scale AI in 2025 and aims to monetize its Muse models. In March, Meta already added low-cost subscription tiers for profile customizations and super reactions; Meta One now locks AI usage behind a paywall. The post does not disclose pricing.
Former TikTok execs launched Superpose, a camera app that analyzes your selfies or photos and generates four possible poses using AI. It solves the awkward 'where do I put my hands' problem. Google had a similar feature called Camera Coach on Pixel phones last year, but Superpose focuses specifically on human posing. The post doesn't disclose which model it uses, whether it's free, or the exact launch date.
An open-source project that uses a Raspberry Pi and microphone to detect bird calls in real time, running fully local AI without internet. When a bird is heard, it displays a 1800s-style bird illustration on an e-ink screen. It uses BirdNET for audio recognition and Stable Diffusion for vintage engraving style. Code and models are open-source. The post doesn't specify how many bird species are supported or the detection latency.
mukel90 released jinfer, a pure-Java inference engine covering chat, vision, audio transcription, embeddings, reranking, and TTS. It runs without Python, ONNX, or sidecar processes, and reads gguf/safetensors natively. CPU performance is claimed to be competitive with llama.cpp; GPU support via the jota backend is still in progress. It integrates with Spring AI and LangChain4j and supports GraalVM Native Image. The author previously built llama3.java and gemma4.java. This is an early release—I'd wait to see how GPU pans out before getting too excited.
All three are early-stage. Runway Solaris generates interactive UIs frame-by-frame with no frontend code—only curated demos and a waitlist so far, no public testing, pricing, or API. Google WikiSkill distills agent failure logs into reusable skill manuals, lifting Gemini-3.5-Flash accuracy from 49.5% to 68.1%, but skills from a small model can hurt a larger one; no official code repo. GitHub GHES 3.22 lets enterprises self-host Copilot CLI inside air-gapped networks with admin-managed model endpoints, though many features are disabled and it's labeled a technical preview.
Why it matters: Three items bundled, with Solaris as the main hook. Runway's frame-by-frame interface rendering is genuinely novel, but there's only a curated demo and waitlist — no public access, no third-party testing, and the cost comparison dodges standard web rendering. That keeps it bel...
Tom Gally built a site with Claude Fable 5.1 that extends Simon Willison's pelican-riding-a-bicycle test into 30 SVG drawing prompts. The 2026 run covers six models—GPT-6 Astra, Claude Fable 5.1, Gemini 3.8 Flash, DeepSeek V4 Pro, Qwen3.8 Max, and Fugu Ultra v2—while the 2025 run includes ten models like Claude Sonnet 4.5 and GPT-5.1. Each image shows generation time and cost: DeepSeek V4 Pro finished in 1 min 36 s at $0.10, Qwen3.8 Max took over 12 minutes, and Fugu Ultra v2 cost $1.02. The post presents raw SVG outputs without subjective ratings, so you compare the drawings directly.
Why it matters: Simon Willison's pelican test is a community staple, and this expands it to 30 prompts across 6 models with timing and cost — dense, useful signal. The deliberate lack of subjective scoring means readers have to flip through images themselves, which costs it a bit of immediate...
DeepSeek dropped V4.1-Flash on Sep 10. Despite the 4.1 label, Sebastian Raschka called it a V5-level rewrite. It's a 763B total-parameter MoE with a causal encoder-decoder split: 8B active for prefill, 16B for decode, yielding 1–2% sparsity and up to 8× smaller KV cache vs V4 Flash. Native vision is built in, and V4 Pro has been quietly retired. The post doesn't include benchmark tables but argues current evals miss the point—the real advance is context efficiency for long-running agents.
Why it matters: DeepSeek drops V4.1-Flash with a 763B causal encoder-decoder MoE, 8B/16B active params, 1%-2% sparsity, and vision. Sebastian Raschka says it should've been V5. This is a major domestic flagship architecture update with a cross-source cluster forming. HKR all hit. Not 90+ yet ...
DeepSeek released open weights for V4.1-Flash, a 552B MoE model with a Causal Encoder-Decoder architecture tuned for coding agents. It splits compute asymmetrically: 8B active params during prefill, 16B during decode, plus improved KV cache efficiency. On Terminal Bench 2.1 it hits 90.6; on Automation-Bench it scores 54.8—better than V4-Pro but still failing roughly half of complex workflows, so keep a human in the loop. It is also DeepSeek's first non-experimental model with native image input. Chartography reaches 78.9, but ZeroBench logical reasoning over images is only 49. DeepSeek has already retired V4-Flash traffic and will reroute V4-Pro traffic to V4.1-Flash starting September 14.
Why it matters: DeepSeek open-sourced V4.1-Flash, a 552B MoE that splits prefill and decode via CED architecture, directly targeting coding agent latency. Terminal Bench 2.1 scores are concrete, and Baseten's analysis adds deployment perspective. Not 85+ because this is a third-party writeup ...
The article body is blocked by WeChat; only the title remains. It claims DeepSeek V4.1 Flash has a big price drop, native vision, and was tested on game and city generation tasks. The post does not disclose the exact price cut, vision specs, or generation quality.
WorkBuddy now offers DeepSeek V4.1-Flash on its platform with a two-week free trial. The model is available via DeepSeek API and supports native multimodal input. The post doesn't spell out improvements over prior versions or pricing.
Google released Pics, an image tool built on Nano Banana, now live at pics.new. It supports local object editing, in-image text editing and translation, multi-player collaboration, and generating multiple options from one prompt. The post doesn't disclose pricing or model parameter details.