Skip to content

#多模态

1 today

Aug 6, 2025Wednesday

OpenAI News

Providing ChatGPT to the Entire U.S. Federal Workforce

OpenAI partnered with the U.S. General Services Administration to offer ChatGPT Enterprise to the full federal executive workforce for $1 per agency for 1 year. Participating agencies also get 60 days of unlimited advanced models and features, including Deep Research and Advanced Voice Mode; federal business data will not be used for training. The key signal is centralized procurement access, while the post does not disclose agency count, budget size, or exact model list.

Why it matters: OpenAI’s GSA deal turns ChatGPT Enterprise into a federal-wide procurement channel at $1 for one year, which is a real distribution signal, not a routine discount. HKR-H/K/R all pass, but the post omits agency count, budget size, and full model scope, so I keep it at 84.

Aug 1, 2025Friday

Jul 17, 2025Thursday

Mistral AI

Mistral adds Deep Research, voice mode and more to Le Chat

Mistral rolled out a batch of new Le Chat features: a preview Deep Research mode, a voice mode powered by the new Voxtral speech model, a multilingual thinking mode backed by the Magistral reasoning model, Projects for organizing conversations, and advanced image editing built with Black Forest Labs.

Why it matters: Mistral announced five Le Chat features at once, so readers can see how its research, voice and image-editing abilities fit together.

Jun 1, 2025Sunday

OpenAI News

OpenAI bans China-origin accounts using ChatGPT to generate US polarization content

OpenAI banned a set of China-origin ChatGPT accounts, dubbed 'Uncle Spam,' after a tip from Meta. The accounts used models to generate pro- and anti-tariff posts, create fake US veteran profile images, and write code to scrape user data from X and Bluesky. The content pushed both sides of divisive topics but got almost no real engagement—most posts had zero likes or reposts. OpenAI rates the impact as Category 2 on the Brookings Breakout Scale: multi-platform activity with no breakout.

Why it matters: Official OpenAI disclosure with a codename and behavioral specifics, not a generic threat report. Hits all three HKR axes, but it's a safety incident notice rather than a product/model update, so it lands in the 78-84 'worth recommending' band.

May 8, 2025Thursday

OpenAI News

Introducing data residency in Asia

OpenAI launched data residency on May 7, 2025 in Japan, India, Singapore, and South Korea for ChatGPT Enterprise, ChatGPT Edu, and the API Platform. Eligible API customers must create a new Project and pick a country; new Enterprise and Edu workspaces can store customer content at rest in-region, including chats, uploads, and text, vision, and image data. The key limit: the post only states at-rest storage and does not disclose whether inference stays fully local.

May 7, 2025Wednesday

Mistral AI

Mistral AI releases Mistral Medium 3, targeting low cost and enterprise deployment

Mistral AI released Mistral Medium 3, which it says reaches or exceeds 90% of Claude Sonnet 3.7 across benchmarks. Pricing is $0.4 per million input tokens and $2 per million output tokens.

Why it matters: Mistral Medium 3 benchmarks against Claude Sonnet 3.7 at lower cost, and the post gives a path to private enterprise deployment and customization.

Apr 23, 2025Wednesday

OpenAI News

Introducing our latest image generation model in the API

OpenAI added gpt-image-1 to the Images API on April 23, 2025, after ChatGPT image generation reached 130 million users and 700 million images in its first week. Pricing is token-based: $5 per 1M text input tokens, $10 per 1M image input tokens, and $40 per 1M image output tokens, or about $0.02, $0.07, and $0.19 per square image by quality. The part to watch is operational: it keeps 4o image safety guardrails, adds C2PA metadata, and does not train on customer API data by default.

Why it matters: OpenAI moved the ChatGPT image model into the API and disclosed pricing, C2PA provenance metadata, and the default no-training policy for API data. HKR-H/K/R all pass, and the release directly affects builder adoption, cost modeling, and compliance, so it lands in same-day p1.

Apr 16, 2025Wednesday

OpenAI News

Introducing OpenAI o3 and o4-mini

OpenAI released o3 and o4-mini on April 16, 2025, and said its reasoning models can now use ChatGPT tools together, including web search, Python, files, and images. The post says o3 makes 20% fewer major errors than o1 in expert evals, while o4-mini reaches 99.5% pass@1 and 100% consensus@8 on AIME 2025 with Python. The real shift is RL-trained tool use, not just two new model names.

Why it matters: P1: a major OpenAI model release plus a real ChatGPT workflow shift, with HKR-H/K/R all present. The story includes concrete claims (-20% major errors vs o1; 99.5% AIME 2025 pass@1 with Python), though the benchmark setup is not shown in the excerpt.

OpenAI News

Thinking with images

OpenAI said on April 16, 2025 that o3 and o4-mini can process user images inside their internal reasoning chain, with native crop, zoom, and rotation actions. The post shows o3 taking 20 seconds to read upside-down handwriting and 1m44s to solve a maze and draw a path; it claims strong multimodal benchmark results, but the provided body does not disclose the scores. The key point is that image manipulation is folded into the same reasoning stack, not handed off to a separate vision model.

Why it matters: OpenAI confirms a meaningful capability step: o3 and o4-mini manipulate images inside the same reasoning process, so HKR-H/K/R all pass. I kept it below p1 because the provided text gives demo timings, but not the benchmark scores or rollout scope.

Mar 25, 2025Tuesday

OpenAI News

Introducing 4o Image Generation

OpenAI integrated 4o image generation into GPT-4o on March 25, 2025, focusing on native multimodal generation, accurate text rendering, and multi-turn image editing in chat. The post points to joint training on image-text distributions and shows a “transformer → diffusion → pixels” pipeline; examples are labeled best of 1, best of ~8, or best of 8. The real signal is consistency and editability, while pricing, API details, and quotas are not disclosed.

Why it matters: This is a major ChatGPT capability update: native image generation lands inside GPT-4o with explicit claims on text rendering and multi-turn editing. HKR-H/K/R all pass; price, API details, and quotas are not disclosed, so it stays below the top of the band.

OpenAI News

Addendum to GPT-4o System Card: 4o image generation

OpenAI published a GPT-4o system card addendum on March 25, 2025, covering 4o image generation capabilities and marginal risks. The post confirms native GPT-4o integration, photorealistic output, image-to-image edits, and reliable text rendering; specific eval scores and mitigations are not disclosed in the post.

Why it matters: This official OpenAI addendum sits near the major-product-update band for native GPT-4o image generation. HKR-H/K/R all pass on the multimodal hook and concrete capability facts, but missing eval scores and mitigation detail keep it below P1.

Mar 20, 2025Thursday

OpenAI News

Introducing next-generation audio models in the API

OpenAI released three API audio models on March 20, 2025: gpt-4o-transcribe, gpt-4o-mini-transcribe, and gpt-4o-mini-tts. The post says the STT models beat Whisper v2 and v3 on FLEURS and other benchmarks across 100+ languages, while the TTS model adds style control but stays limited to monitored preset synthetic voices. The key shift is controllable TTS plus lower WER; the post does not disclose pricing or latency figures.

Why it matters: OpenAI shipped 3 API audio models with concrete benchmark and mechanism details, so HKR-H/K/R all pass and it clears featured. I kept it at 84, not 85+, because price, latency, and a fuller benchmark table are not disclosed.

Mar 17, 2025Monday

Mistral AI

Mistral AI releases Mistral Small 3.1 with multimodal support and 128k context

Mistral AI released Mistral Small 3.1, which improves text performance and multimodal understanding over Mistral Small 3 and extends the context window to 128k tokens. Inference runs at 150 tokens per second, and the model is open-sourced under Apache 2.0.

Why it matters: Mistral gives the multimodal, 128k-context and 150 tokens/s figures for Mistral Small 3.1, letting readers compare it with small models of the same class.

Mar 12, 2025Wednesday

Hugging Face Blog

Welcome Gemma 3: Google's all new multimodal, multilingual, long context open LLM

Google announced an open LLM called Gemma 3 and named three traits in the title: multimodal, multilingual, and long context. The RSS snippet has no body, so parameter size, context length, license, and benchmark results are not disclosed. Watch the full post or model card; “open” does not equal open source from the title alone.

Why it matters: A new Gemma release from Google is inherently newsy, and the multimodal/long-context/open framing hits HKR-H and HKR-R. HKR-K misses because the feed gives no specs, context window, license, or benchmark data, so this lands at the low end of featured.

Mar 7, 2025Friday

Mistral AI

Mistral AI releases document-understanding OCR API Mistral OCR

Mistral AI released Mistral OCR, an optical character recognition API that takes images and PDFs and outputs interleaved text and images in order. It handles complex layouts such as tables, formulas and LaTeX.

Why it matters: Mistral gives benchmark comparisons, multilingual performance and pricing for Mistral OCR, showing how usable document parsing is in a RAG pipeline.

Feb 6, 2025Thursday

OpenAI News

Introducing data residency in Europe

OpenAI launched European data residency for the API, ChatGPT Enterprise, and ChatGPT Edu on February 5, 2025. New API Projects can select Europe for in-region processing with zero data retention, while existing Projects cannot be changed; new Enterprise and Edu workspaces can store chats, files, and text, vision, and image content at rest in Europe, but the post does not disclose the eligible endpoint list.

Why it matters: A solid enterprise/compliance update. HKR-K lands on concrete conditions—Europe region, zero data retention, new projects only, no migration for existing ones—and HKR-R lands on EU legal and procurement pressure. HKR-H is weak, so this sits at the low end of featured.

Jan 28, 2025Tuesday

OpenAI News

Introducing ChatGPT Gov

OpenAI launched ChatGPT Gov on January 28, 2025 for U.S. agencies to deploy in Microsoft Azure commercial or Azure Government cloud with access to models including GPT-4o. The post lists file upload, shared chats, custom GPTs, and an admin console, and ties the setup to IL5, CJIS, ITAR, and FedRAMP High requirements. The signal is adoption: since 2024, 90,000+ users across 3,500+ U.S. agencies have sent 18 million+ messages.

Why it matters: This clears HKR-H/K/R: the Gov-specific SKU is a real hook, and the post includes hard numbers plus compliance targets. Strong featured rather than p1 because this is a packaging/deployment launch with adoption proof, not a major frontier-model capability jump.

Jan 23, 2025Thursday

OpenAI News

Operator System Card

OpenAI published the Operator System Card on Jan 23, 2025 and said its Computer-Using Agent can be deployed only if its post-mitigation score is Medium or lower. The card rates CBRN, cybersecurity, and model autonomy as Low, and persuasion as Medium; it highlights harmful tasks, model mistakes, and prompt injection. The key mechanism is human confirmation plus task refusal: critical steps like financial transactions, emails, and calendar deletion need approval, while stock trading is fully restricted.

OpenAI News

Computer-Using Agent

OpenAI released a research preview of Computer-Using Agent on Jan 23, 2025, and is exposing it first through Operator to U.S. ChatGPT Pro users. The model combines GPT-4o vision with RL-based reasoning and acts through screenshots, a mouse, and a keyboard; it scored 38.1% on OSWorld, 58.1% on WebArena, and 87.0% on WebVoyager. The key point is API-free GUI control, while sensitive actions still require user confirmation.

Why it matters: This is a same-day OpenAI agent release: CUA powers Operator and ships first to US ChatGPT Pro users. HKR-H/K/R all pass because the GUI-control hook is novel, the post gives mechanism plus 38.1/58.1/87.0 benchmarks, and it raises concrete autonomy and safety questions.

OpenAI News

Introducing Operator

OpenAI released Operator on Jan 23, 2025 as a research preview for U.S. Pro users; it uses its own browser to click, type, and scroll through web tasks. It runs on Computer-Using Agent, combining GPT-4o vision with RL-based reasoning; the post says it sets SOTA on WebArena and WebVoyager but does not disclose scores. The key boundary is control: login, payment, and CAPTCHA flows hand control back to users, and a July 17 update says it was folded into ChatGPT agent.

Why it matters: OpenAI's Operator is a same-day, must-write product release: a browser-using agent moves ChatGPT from answering to acting. HKR-H/K/R all pass; the post gives the own-browser setup, GPT-4o+RL, and user handoff for login/payments, but US Pro limits and missing benchmark scores keep

Dec 9, 2024Monday

OpenAI News

Sora is here

OpenAI moved Sora out of research preview on December 9, 2024 and rolled it out to ChatGPT Plus and Pro users. Sora Turbo supports up to 1080p and 20-second videos; Plus includes up to 50 monthly 480p videos or fewer 720p generations. The key detail for practitioners is deployment scope: the UK, Switzerland, and the EEA are excluded, person uploads are limited, and OpenAI says physics and long complex actions remain weak.

Why it matters: OpenAI moved Sora from preview to paid availability, so HKR-H/K/R all pass: high-curiosity launch, concrete specs and limits, and clear impact on creator workflows. I stop below 95 because the post itself notes region blocks, restrictions on uploads with people, and instabilityon

Oct 23, 2024Wednesday

OpenAI News

Simplifying, stabilizing, and scaling continuous-time consistency models

OpenAI introduced sCM and scaled continuous-time consistency models to 1.5B parameters on ImageNet at 512×512. The post says sCM reaches sample quality comparable to leading diffusion models in 2 sampling steps, with about 50x wall-clock speedup. Its largest model generates one sample in 0.11s on a single A100 at batch size 1 without inference optimization.

Why it matters: This clears HKR-H/K/R: the hook is 2-step sampling with diffusion-like quality, and the paper gives concrete numbers—1.5B params, ImageNet 512x512, ~50x wall-clock speed, and 0.11s per sample on one A100. Strong research release, but not a shipped product, so featured fits better

Oct 1, 2024Tuesday

OpenAI News

Introducing the Realtime API

OpenAI launched a public beta of the Realtime API on Oct. 1, 2024 for all paid developers, using a persistent WebSocket to stream low-latency speech-to-speech interactions with GPT-4o. It supports function calling and interruption handling, priced at $5/1M text input tokens and $100/1M audio input tokens; the post also says audio I/O for Chat Completions would arrive in the following weeks.

Why it matters: OpenAI moved voice apps from stitched ASR+TTS calls to a persistent GPT-4o session, with function calling, interruption handling, and published audio/token pricing. HKR-H/K/R all pass, so this is a same-day must-write developer platform update and clears p1.

OpenAI News

Introducing vision to the fine-tuning API

OpenAI launched GPT-4o vision fine-tuning on Oct 1, 2024, letting paid-tier developers train with images plus text, starting from as few as 100 images. The post cites Grab improving lane-count accuracy by 20% and speed-limit sign localization by 13%, while Automat raised RPA success from 16.60% to 61.67%. The notable shift is multimodal customization in the main API; the pricing section is truncated, so full price details are not disclosed.

Why it matters: OpenAI shipped a substantive API update: GPT-4o vision fine-tuning with a 100-image floor and named gains from Grab and Automat, so HKR-H/K/R all pass. Scope is strong for builders, but the blast radius is narrower than a flagship model launch, and pricing is incomplete in the ex

Sep 26, 2024Thursday

OpenAI News

Upgrading the Moderation API with OpenAI's new multimodal moderation model

OpenAI released omni-moderation-latest on September 26, 2024, a GPT-4o-based Moderation API model for text and image inputs that is free for all developers. It adds illicit and illicit/violent text categories, supports image moderation in 6 subcategories, and improves 42% on an internal 40-language eval, with gains in 98% of languages tested.

Why it matters: Official OpenAI developer product update with strong HKR-K: new moderation classes, image coverage, and a concrete +42% result across 40 languages. HKR-R also lands because moderation and compliance affect shipping teams directly; HKR-H is weak, so this sits at the low end of the

Aug 8, 2024Thursday

OpenAI News

GPT-4o System Card

OpenAI published the GPT-4o System Card on August 8, 2024, reporting 3 of 4 Preparedness categories as low risk and persuasion as borderline medium. The post says GPT-4o accepts text, audio, image, and video inputs, responds to audio in as little as 232 ms with a 320 ms average, and is 50% cheaper than GPT-4 Turbo in the API. The key issue for practitioners is voice safety: the card names unauthorized voice generation, speaker identification, and sensitive trait attribution, and says only models with post-mitigation scores at medium or below can be deployed.

Why it matters: This is not a routine post: it adds concrete preparedness ratings, 232ms voice latency, and a clear deployment threshold. HKR-H/K/R all pass, but it is a safety disclosure rather than a new model or major launch, so it lands as featured, not p1.

Jul 23, 2024Tuesday

Hugging Face Blog

Llama 3.1: 405B, 70B & 8B with multilinguality and long context

Meta released Llama 3.1 with 405B, 70B, and 8B sizes, and the title says it adds multilingual support and long context. Only the title is available; the post does not disclose context length, languages, license terms, or benchmark results. Watch the 405B release terms and real inference cost.

Why it matters: Meta's Llama 3.1 is a major flagship open-model release, and the title already gives concrete sizes plus multilingual and long-context positioning. HKR-H/K/R all pass; missing license, exact context window, and benchmark detail keep it at the low end of the 85-94 band.

Jul 18, 2024Thursday

OpenAI News

GPT-4o mini: advancing cost-efficient intelligence

OpenAI released GPT-4o mini on July 18, 2024 at $0.15 per 1M input tokens and $0.60 per 1M output tokens, replacing GPT-3.5 in ChatGPT. It supports text and vision, offers a 128K context window and 16K max output, scores 82.0% on MMLU and 87.2% on HumanEval. The key detail for builders is that its API version is the first to use instruction hierarchy against jailbreaks and prompt injection.

Why it matters: This is a substantive OpenAI model launch, not a minor refresh: GPT-4o mini adds $0.15/$0.60 pricing, 128K context, 16K max output, benchmark details, and instruction hierarchy, then replaces GPT-3.5 in ChatGPT. HKR-H/K/R all pass, so it lands in P1.

Sep 25, 2023Monday

OpenAI News

ChatGPT can now see, hear, and speak

OpenAI says ChatGPT now supports seeing, hearing, and speaking. The post body is empty, so it does not disclose model versions, rollout timing, regional limits, pricing, or API scope. The real watchpoints are voice latency, vision limits, and access paths.

Why it matters: This is a substantive OpenAI product update: the title confirms vision input, voice input, and speech output for ChatGPT, so HKR-H/K/R all pass. The copy provided here omits tiers, rollout scope, latency, and pricing, which keeps it at 88 rather than the top of the band.