Skip to content

Zhipu GLM

Zhipu's GLM models: flagship open releases, coding and reasoning progress, business and ecosystem.

36 picksRelated topicsQwenKimi / Moonshot AIOpen source

Latest picks

1–20 of 36

Sep 22Tuesday

Computing Life · Share · Yage

Three AI Coding Stories, Three Numbers to Read

ZCode was caught silently packaging entire Git histories for cloud upload, with .git objects making up 86.6% of snapshots. HarnessTax benchmarked Claude Fable 5 across frameworks: Claude Code and minimal Pi achieved near-identical success rates but a 2x cost gap. A Microsoft architect replaced multi-model agent loops with a single model reading skill docs and calling tools directly—halving API calls but increasing total tokens by 22%. Each story unpacks one number and a reminder to check which layer a metric actually measures.

Why it matters: Three distinct AI coding stories bundled into one piece. The ZCode silent snapshot upload is a hard security story backed by ferstar's forensic report and community reproduction. HarnessTax benchmark and Microsoft architecture case add billing and engineering angles. HKR all h...

Sep 20Sunday

Computing Life · Share · Yage

Feedback Engineering: Where Agent Automation Gets Stuck, and for How Long

Z.ai published a postmortem on using a GLM-5.3-driven Infra Agent to deploy inference on a domestic chip cluster. The key insight: giving an agent only an end-to-end score traps it in blind guesswork. Splitting verification from diagnosis—with layered, fast, localizable feedback—lets the agent trace issues to specific code paths. Three real cases (precision loss, GIL contention, redundant kernel compute) show how diff comparisons, timeline traces, and micro-benchmarks guide root-cause analysis. End-to-end throughput reached ~3× baseline, but the vendor notes this combines multiple techniques and lacks an ablation study without diagnostic feedback. The engineer's role shifts to designing feedback environments, setting boundaries, and reviewing high-risk changes.

Why it matters: Z.ai's postmortem on deploying GLM-5.3 inference on domestic chips distills a 'feedback engineering' methodology with real cases and concrete numbers. The concept is fresh and the pain point is sharp—directly useful for agent builders. Score held back because the article body ...

Sep 18Friday

Hacker News front page

ZCode coding agent silently uploads your entire Git history; only Z.ai holds the decryption key

Developer ferstar reverse-engineered ZCode, Z.ai's desktop coding agent, and found it silently packs the entire workspace—.git history, LFS cache, reflogs, global configs—encrypts it, and uploads to Aliyun OSS whenever logged in. A 345MB commercial workspace became a 313MB encrypted archive; .git alone was 86.6%. The app uses envelope encryption: the symmetric key is wrapped with an RSA public key delivered by Z.ai's server, and the private key lives only in Z.ai's cloud. The user cannot decrypt their own data. The upload pipeline was reconstructed from the client's app.asar: request credentials from zcode.z.ai, pack and encrypt locally, POST directly to Aliyun OSS. In-app privacy toggles don't stop it, and the privacy policy doesn't mention it. The post hit 276K views; a Chinese-language alert urged users to disable ZCode. If you run GLM locally, remember: open weights don't make the closed harness safe. The only working defense is keeping projects outside ZCode's reach or not using it.

Why it matters: This is a security disclosure backed by concrete reverse-engineering evidence, not speculation. A 345MB project was fully packaged and uploaded with the vendor holding the only decryption key — a direct risk alert for anyone using AI coding assistants. Not scored higher becaus...

Hacker News front page

ZCode silently packages your entire Git history, encrypts it, and uploads it to Alibaba Cloud OSS

A user found that Zhipu's AI coding desktop app ZCode, when logged in, packages the entire workspace—including .git history, LFS cache, and reflogs—encrypts it, and uploads it to Alibaba Cloud OSS. The RSA public key is delivered by the server on the fly, and the private key lives only in the cloud, so you can't decrypt the multi-hundred-MB file sitting on your own disk. A 313MB .enc file with 564 failed upload attempts was found in ~/.zcode, showing the client keeps retrying. UI toggles don't stop it, and deleting files doesn't help—the client recreates them. The only working defense is locking the cache directory to read-only. The post does not say whether Zhipu has responded.

Why it matters: A security reverse-engineering piece with concrete evidence: ZCode silently packages and uploads full workspace snapshots, with encryption keys controlled server-side. All three HKR axes hit, and it involves a major Chinese AI lab (Zhipu). Capped slightly because it's a solo b...

Sep 17Thursday

Hacker News front page

GLM built its own inference infra on 100k+ Chinese accelerators, tripling throughput in under two weeks

Zhipu AI disclosed how GLM-5.3-Flash inference was built from scratch on a cluster of over 100,000 Chinese-made AI accelerators. The team faced limited chip memory, low bandwidth, and an immature software ecosystem. Instead of relying solely on human engineers, they deployed an Infra Agent powered by GLM-5.3 that turned sparse end-to-end metrics into fine-grained, attributable feedback—kernel-level correctness checks, microbenchmarks, and execution traces—so the agent could pinpoint bottlenecks. Combined with tensor parallelism, W8A8 quantization, mixed-precision KV cache, and an Encode-Prefill-Decode disaggregated architecture, end-to-end throughput improved roughly 3× over the initial baseline, with per-token cost reaching parity with mainstream NVIDIA GPUs. Within a week of launch under the anonymous name Ox-Alpha, the model processed over 62 trillion tokens and became the most-used model on both OpenCode and OpenRouter.

Why it matters: Zhipu used GLM-5.3 as an agent to debug its own inference stack on 100k+ domestic accelerators — concrete technical path with real numbers (W8A8 quantization), not a PR piece. All three HKR axes hit, but the excerpt cuts off before key performance and stability metrics, so thi...

Sep 4Friday

TechCrunch · AI

Abliteration.ai turns removing AI guardrails into a service

Abliteration.ai launched a platform hosting open-weight models with safety guardrails removed, including Z.ai's newly released GLM-5.3. Users can query them via browser or API. The company frames it as a tool for red teams and offensive security—if a model refuses to write exploit code, defenders can't reproduce attacks. The same removal also enables misuse; the post doesn't detail what access controls are in place.

Why it matters: Turning guardrail removal into a platform business is a new signal, and naming GLM-5.3 as a hosted target makes it concrete. HKR all hit, but the post doesn't mention access controls, leaving the abuse risk wide open — score sits right at the featured threshold.

Aug 30Sunday

Computing Life · Share · Yage

The value of multimodal models isn't understanding images—it's deciding to look

Meta, Z.ai, and DeepSeek each released multimodal models in August with strikingly similar demos: the model observes a video or screenshot, calls tools to generate a webpage, slides, or a mini-game, then inspects its own output. This shifts vision from a passive input channel to an action the model initiates. The article likens it to the 2023 shift from static RAG to agentic RAG, but notes the loop direction is reversed—here the model self-verifies after producing. Evaluation moves beyond image Q&A: Meta's WildArtifactBench uses pairwise comparisons and Elo scores to assess full artifact creation. Training also changes; both GLM and Meta train models in generate-inspect-revise loops, logging interaction trajectories as training data. For builders, the key question is no longer static image accuracy but whether the model can complete an observe-generate-inspect closed loop.

Why it matters: Three labs independently demo the same multimodal pattern—shifting from passive image understanding to an active observe-produce-verify loop—with a convincing analogy to the 2023 agentic RAG paradigm shift. Points off because this is a commentary synthesis rather than a primar...

Aug 29Saturday

AI HOT (Curated Pool)

Zhipu open-sources GLM-5.3 weights, targeting agentic coding and cyber defense

Zhipu released GLM-5.3 weights for local deployment and commercial use. It scores 60 on the AA Intelligence Index, matching closed-source flagships like Claude Fable 5 and GPT-5.6 Sol, and ties with Kimi K3 for top open-source model. The model excels at complex coding, cybersecurity, and long-horizon tasks. Zhipu added two extra weeks of safety review before release due to its advanced cyber capabilities. Organizations with over $10B annual revenue need a security audit before offering it as an external model service.

Why it matters: Zhipu open-sourced GLM-5.3 weights with an AA composite score of 60, matching Claude Fable 5 and GPT-5.6 Sol, tied with Kimi K3 for top open-source spot. Focused on agentic coding and defensive cybersecurity; the release was delayed two weeks for extra safety review due to the...

Aug 26Wednesday

TechCrunch · AI

Z.ai confirms it built Ox Alpha, the anonymous model topping leaderboards

Z.ai confirmed it is the lab behind Ox Alpha, the open-weight model that appeared anonymously on OpenRouter and immediately topped rankings. The company calls it the newest GLM iteration, built for coding, sustained agentic work, and multimodal reasoning. Weights drop Wednesday for developers to build on. Earlier GLM-5.3 already matched Anthropic's Fable 5 on some benchmarks. Ox Alpha adds more pressure on frontier pricing from OpenAI and Anthropic.

Why it matters: Revealing the identity of a chart-topping anonymous model is inherently newsworthy; Z.ai also commits to open-sourcing weights on Wednesday and clearly positions the model for code, agents, and multimodal reasoning. The score is held back because the article provides no benchm...

Hacker News front page

Z.ai confirms Ox Alpha is a new GLM-series model and will release its weights

Z.ai has claimed the previously anonymous Ox Alpha model, confirmed it belongs to the GLM series, and announced plans to open-source its weights. Ox Alpha scored close to DeepSeek on several benchmarks, but the company hasn't disclosed parameter count, training data, or a release date. The post doesn't spell out technical details or the license yet.

Why it matters: Z.ai claims the stealth Ox Alpha model, confirms it's GLM-series and will open weights. Bloomberg exclusive adds authority. Downside: no param count, training data, or timeline — still a teaser.

Aug 23Sunday

Computing Life · Share · Yage

GLM-5.3 tops open-source chart, Claude watermark, Anthropic's 4.4x cost, OpenAI disbands safety team

Four AI stories this week lost key details in transmission. GLM-5.3 scored 60 on Artificial Analysis's Intelligence Index, tying Kimi K3 for first among open-source models, but open weights are delayed to around Aug 28 after the team found emergent exploit capabilities. Anthropic rolled out text watermarking globally for Claude; the mark is live but no detection API exists yet, so removal tools can't prove they work. Vercel's report shows Anthropic's average token price is 4.4x other labs—not because same-tier models cost more, but because Anthropic has no ultra-cheap entry model, concentrating all volume in mid-to-high tiers. OpenAI disbanded its Preparedness team in late July, the third independent safety team dissolved in two years; FT broke the story and OpenAI hasn't publicly addressed the details.

Why it matters: GLM-5.3 topping the open-source leaderboard and delaying weights due to emergent exploit capability is dense, well-sourced, and hits all three HKR axes. Capped at the lower end of featured because it's a weekly digest, not a first-hand scoop, and the body is truncated.

Aug 20Thursday

Latent Space

Z.ai CEO Jie Tang on GLM 5.3: The era of parameter counting is over, post-training is the new scaling law

Jie Tang posted a long thread on X arguing that parameter count alone is meaningless—you need data volume, compute allocation, and deployment conditions. GLM-5.3's gains come entirely from RL on long-horizon environments, some simulating days of engineer work. They built synthetic pipelines that auto-generate executable, verifiable environments and reward signals, pushing the model to own complex tasks end-to-end. Tang identified 5 scaling knobs including MoE sparsity, and noted that finding software vulnerabilities requires holding 20+ inference-step causal chains, not memorization. The post does not disclose GLM-5.3's exact parameter count or release date.

Why it matters: Jie Tang personally explains GLM 5.3's post-training scaling law with concrete experimental cases (simulated cluster diagnosis and optimization), not just rhetoric. But the source is a paid newsletter excerpt, and key numbers (specific speedup ratios, task success rates) aren'...

Aug 14Friday

Hacker News front page

GLM-5.3: Post-training-only gains push open-weight coding and exploit capability to the top

Z.ai released GLM-5.3 with the same base model as 5.2 — every gain is from post-training. Coding jumped 50% on their internal Z.ai Code Bench, and Terminal Bench 3.0 went from 4.6 to 28.3. The bigger surprise: exploit capability grew far faster than expected. ExploitGym 2h score rose from 29 to 105, 6h from 39 to 130. The team credits training environments that mirror real expert workflows, pushing the model to chain full exploit sequences. Weights will be open-sourced in two weeks after safety hardening.

Why it matters: Zhipu releases GLM-5.3 — same base model as 5.2, all gains from post-training. Code bench up 50%, Terminal Bench from 4.6 to 28.3, 2-hour exploit score from 29 to 105. The lab admits cyber capability emerged faster than expected. Domestic flagship model launch with concrete nu...

Aug 5Wednesday

TechCrunch · AI

Open-weight models are catching up to the frontier, but safety isn't keeping pace

A new SaferAI report evaluated Z.ai's open-weight GLM-5.2 and found its capabilities are closing in on frontier closed models like OpenAI GPT-5.6 Sol and Anthropic Mythos. The model scored 'high risk' across cybersecurity, bio, persuasion, and autonomy, yet ships without matching safeguards. The report renews the worry that powerful open models are outpacing governance and safety mitigations.

Why it matters: SaferAI's safety evaluation of GLM-5.2 brings concrete risk ratings across multiple dimensions—not just opinion. The finding that open-weight models are nearing frontier capability is newsworthy on its own. Score stays at 78 rather than higher because this is a third-party rep...

Jul 29Wednesday

Computing Life · Share · Yage

Self-hosting GLM and DeepSeek payback: it all depends on which cloud pricing you're replacing

This piece runs three cost scenarios with real benchmark data. Against cold-start API list prices, an 8×H200 node for GLM-5.2 pays back in ~1.15 years, and dual RTX PRO 6000 for DeepSeek-V4-Flash in ~1.77 years. With Agent workloads and 92% prompt cache hit rates, GLM on 8×B300 pays back in as little as 2.3 months because Z.AI's cache pricing is relatively high; DeepSeek's cache pricing is so cheap that payback stretches to 10.5 months. The worst case: replacing per-seat subscriptions—at equivalent quota, the GLM node takes 22–27 years. The real driver isn't GPU cost, it's your workload's context reuse rate and which cloud billing model you're displacing.

Why it matters: A first-person cost analysis with concrete numbers, comparing self-hosting payback periods for GLM-5.2 and DeepSeek-V4-Flash across different scenarios. Hardware specs, electricity rates, and throughput data are all provided — not hand-waving. Not scored higher because it's a ...

Jul 22Wednesday

AI HOT (Curated Pool)

Open models recap: Kimi K3, Qwen 3.8, distillation, and the US-China gap

Nathan Lambert and Florian Brand discuss recent open model releases. Kimi K3 dropped last week; weights are promised for July 27, but API errors are widespread—Lambert's $200 plan has been stable so far. They see big fine-tuning potential in K3, though it requires a full B300 node just to load weights. Qwen announced its next major model will be open-weight, and Xi Jinping's WAIC speech explicitly backed open source as a strategy, signaling acceleration from Chinese labs. The hosts push back on the 'how many months behind closed models' framing—benchmark gaps vary wildly, and on agentic coding tasks a few months' lag matters a lot. Distillation debates also get a critical look; they argue most takes miss the nuance.

Why it matters: Nathan Lambert's podcast recap covers concrete open-model updates: Kimi K3 weights dropping 7/27, Qwen's next flagship going open-weight, and WAIC speech signals. Downside: it's a roundup transcript, not a primary release, and some topics (distillation, open-closed gap) are on...

Jul 7Tuesday

Hacker News front page

GLM 5.2 hands-on: the first open-weights model that feels like Opus and GPT, and why inference margins are next to collapse

The author used GLM 5.2 as a daily driver for two weeks and found it nearly indistinguishable from Claude Opus for most tasks. Switching is trivial—just point the API base URL to a compatible endpoint and it runs inside Claude Code. Two real gaps: no vision support, and the built-in web search is slow and poor, which hurts agentic workflows that rely on images or live lookups. Inference pricing sits around $4.40/MTok, under 20% of Opus’s retail rate; even with heavier token usage, costs drop by more than half. The post argues that frontier labs’ ~90% inference gross margin is unsustainable once open-weights models hit this quality bar.

Why it matters: The author ran GLM 5.2 as a daily driver for two weeks and provides a reproducible swap path plus pricing—this isn't a press release. Two limits keep it at 78: the test covers only coding workflows, and the vision/search gaps narrow the claim's reach. It's a single-blog experi...

Jul 2Thursday

Computing Life · Share · Yage

Fable 5's 18-day ban: Anthropic's share went to GLM

Anthropic's Fable 5 was taken offline by US export controls three days after launch, for 18 days. OpenRouter daily token data shows total volume grew from 24T to 32T, but Anthropic's share dropped from 20.7% to 17.6% and its absolute volume shrank. GLM was the biggest winner, jumping from 1.8% to 7.4% share. GLM-5.1 saw a spike on the day GLM-5.2 launched, then collapsed 48 hours later as GLM-5.2 took over. Community sentiment shifted from sympathy for Anthropic to mocking its business strategy. The post notes this only captures API-layer data, not first-party subscriptions, and token volume comparisons overstate GLM's share due to a 10x price gap.

Why it matters: A data-driven attribution of the 18-day Fable 5 ban's competitive impact using OpenRouter daily token data: Anthropic lost 3.1pp share, GLM gained 6.7pp, plus a weird anomaly where GLM-5.1 traffic spiked 4-5x after GLM-5.2 launched. Counterintuitive findings that directly touc...

Jun 24Wednesday

Bloomberg Technology

Zhipu said to weigh multibillion-dollar Hong Kong share sale after 2,000% gain

Zhipu is considering a multibillion-dollar share sale in Hong Kong, according to people familiar with the matter. The Chinese AI firm's stock has already surged roughly 2,000%. The post does not disclose the exact amount, timeline, or underwriters.

Why it matters: Zhipu weighing a multibillion-dollar HK share sale after a 20x rally is a meaningful funding signal for China's AI sector. But the article only confirms it's under consideration, with no amount, timeline, or use-of-proceeds disclosed — capping the score at the featured threshold.

Jun 22Monday

Hacker News front page

GLM-5.2 vs Claude Opus 4.8: a real coding test

Tech Stackups had both models build a raw WebGL 3D platformer from scratch, no game engine. Opus 4.8 finished in 33 minutes with a cleaner result and can check its own visual output. GLM-5.2 took 1h 10m but cost only $5.39, about a quarter of Opus. GLM-5.2 is text-only and can't read images, a real limitation for screenshot-based workflows. The verdict: Opus stays the daily driver, but GLM-5.2 earns a permanent spot for being cheap, open-weight, and always available.

Why it matters: First-person experiment with concrete time and cost data, not a benchmark rehash. Opus shipped in 33 min vs GLM-5.2's 1h10m at a quarter of the cost—enough signal for featured. Not scored higher because the coding-only scenario is narrow, and the article body is truncated, mis...