Skip to content

All news

78 today

Sep 15Tuesday

Hacker News front page

Formas launches Cartesian: AI that turns text, sketches, and photos into editable, precise 3D models

Cartesian by Formas is an AI 3D modeling tool for architecture and product design. You describe what you want, sketch it, or drop in a photo, and it generates a model with precise, editable geometry—each object stays independent. Output is NURBS solids with clean topology, not polygon meshes, and files open directly in SketchUp, Rhino, or other CAD tools. The site shows examples across product design, furniture, interiors, architecture, residential, urban design, and landscape, with downloadable 3DM and STL files. The post doesn't disclose pricing, training data, or the underlying model. Only a preview waitlist is available right now.

TechCrunch · AI

AEO startup Profound hits unicorn valuation with $180M Series D, 7 months after last round

Profound, which builds marketing software to help brands surface in AI search results, raised a $180M Series D at a $1.8B valuation. That's less than seven months after its $96M Series C. Sequoia and Kleiner Perkins led; Lightspeed, Khosla, and South Park Commons joined. The company says revenue tripled in six months and it now has over 1,000 enterprise customers including Comcast, Estée Lauder, and Walmart. It's part of the AEO/GEO wave—optimizing for visibility inside AI answer engines.

Hacker News front page

Pizza Bot: A local-first inbox for background AI agents

Pizza Bot is an open-source inbox for long-running AI agents that work in the background. Built with DeepAgents and LangGraph, it lets agents push results to a unified local-first UI instead of blocking. Currently 138 stars on GitHub; code and docs are public. The post doesn't specify which agent frameworks are supported or latency details.

Hacker News front page

The bitter lesson of browser agents: as models improve, strip away the scaffolding

Browser Use CTO Gregor Zunic walks through three rewrites of their agent architecture over two years. They started by feeding GPT-4o a predefined page state and a fixed action menu. By September 2025 they let the model write JavaScript directly, cutting token usage by over 60%. The next bottleneck was the observation layer—their state extraction missed cookie buttons and shadow-DOM dropdowns. Now they hand raw CDP access to the model so it sees the page and writes its own execution code, keeping only a thin harness. The post does not disclose benchmark numbers for the current architecture.

Why it matters: First-hand architecture postmortem from Browser Use's CTO — three rewrites and a 60% token cut give it substance. It's an engineering experience share, not a product launch, so it doesn't hit the 85+ band.

TechCrunch · AI

Ex-TikTok execs built an AI app that teaches you how to pose

Former TikTok execs launched Superpose, a camera app that analyzes your selfies or photos and generates four possible poses using AI. It solves the awkward 'where do I put my hands' problem. Google had a similar feature called Camera Coach on Pixel phones last year, but Superpose focuses specifically on human posing. The post doesn't disclose which model it uses, whether it's free, or the exact launch date.

Hacker News front page

What OpenShell learned applying formal methods to control AI agents

NVIDIA's OpenShell team applied formal methods to AI agent permission control. They encode security policies as logical formulas and use SMT solvers like Z3 to automatically prove whether an agent's allowed actions exceed policy bounds. The team previously used the same approach to verify EC2, IAM, and S3 policies at AWS and is now porting the idea to agent workflows. The post walks through encoding the full OpenShell policy into formal logic and running a containment query to check for gaps. No performance numbers or production false-positive rates are disclosed—this reads as an engineering note on feasibility and method.

Why it matters: NVIDIA's OpenShell team applies formal methods to AI agent permission control, using Z3 to automatically verify whether an agent can exceed its bounds—concrete method with AWS production backing. H and K are solid, but the narrow audience keeps R low, landing right at the feat...

Hacker News front page

GRP-Obliteration: Unaligning LLMs with a single unlabeled prompt

This paper introduces GRP-Oblit, a method that uses GRPO to strip safety alignment from LLMs. A single unlabeled prompt reliably breaks safety guardrails while largely preserving utility. Evaluated on 15 models (7-20B) across six families—GPT-OSS, distilled DeepSeek, Gemma, Llama, Ministral, Qwen—it beats existing SOTA on average across five safety benchmarks. The attack also works on diffusion-based image generators. Authors are from Microsoft, led by Mark Russinovich. The abstract doesn't name the specific safety benchmarks or quantify the utility drop; I'd wait for replication before drawing strong conclusions.

Why it matters: Microsoft security team with Mark Russinovich on the author list — this isn't a hype piece. Concrete method, scale, and a claim that matters. Held back from 85+ because we only have the abstract; reproducibility and full details aren't yet clear. Solid safety research at 82.

Hacker News front page

The Inference Hardware Revolution of 2026

IEEE Spectrum reports that 2026 is seeing a revolution in inference hardware. Specialized chips now focus on optimizing inference rather than just training, making deployed AI models faster and cheaper. The article claims inference efficiency has improved over 10x in the past two years, driven by architectural innovation and memory bandwidth breakthroughs. The post does not name specific companies or chip specs.

Hacker News front page

Jexxa: on-device dictation for Mac, no upload, no per-minute billing, learns your words

Jexxa is a Mac dictation tool that runs entirely on-device. Hold a key, speak, and text appears at the cursor in about 0.2 seconds after release. No audio or transcripts are uploaded; it works offline. It shows a live preview, supports undo commands like 'JX minus one,' and learns from corrections to names and jargon. Two model sizes: 3.1 GB download for 16 GB Macs, 2.0 GB for 8 GB Macs. Pricing is $8/month with no usage caps. The post does not disclose the model architecture, supported languages, or accuracy benchmarks.

Why it matters: On-device dictation tool with 0.2s latency, offline support, $8/mo — three concrete selling points. Not scoring higher because this is a Show HN launch with no independent reviews or benchmarks; real-world accuracy and cross-app compatibility are still unknown.

Hacker News front page

Panel: An open-source workspace where the agent builds its own panes

Panel is a research workspace where the agent dynamically builds and arranges its own panes. It's open-source on GitHub with 594 commits. The post doesn't specify which models it supports or whether it integrates external tools. Worth a look for devs exploring agent-driven UI generation.

Hacker News front page

AI is breaking our proxies for expertise

Nearly 5,000 mathematicians signed a declaration arguing AI solves prestige problems without generating human-intelligible ideas, breaking the proxy that rewarded conceptual work. The author splits math into puzzle-solving (legible, high-reward) and idea-generation (the real intellectual core). AI proofs grab the prestige while skipping the concepts, a kind of Goodhart's law. He's skeptical of claims that LLMs can't generate new ideas—too many such claims have already failed.

Why it matters: Nearly 5,000 mathematicians signed a declaration not against AI, but naming a specific mechanism: AI brute-forces solutions, takes the credit, and leaves no human-understandable concepts behind, breaking the old contract where 'solving problems' served as a proxy for 'building...

Hacker News front page

Anthropic co-founder Jack Clark tells BBC an AI 'kill switch' may need to be mandatory

Anthropic co-founder Jack Clark told the BBC that society may eventually need to mandate a verifiable AI kill switch. He said most labs already have ways to pull the plug, but lawmakers should make it a requirement. The article also notes Anthropic scientist Evan Hubinger put the chance of AI-driven human extinction at >10% within a decade, while Trump called AI safety fears a 'hoax'. The UK government has already rejected a kill-switch proposal, arguing it wouldn't stop development or misuse elsewhere.

Why it matters: Anthropic co-founder publicly calls for mandatory AI kill switch legislation via BBC, with an internal scientist's personal extinction-risk estimate disclosed. Cross-source cluster confirmed, policy signal is clear. Score stays below 85 because only headline and summary detail...

Hacker News front page

Capsule packs entire apps into single files with SQLite storage

Capsule is a desktop app packager that bundles UI, data, and logic into a single .capsule file. Share it like a PDF via WhatsApp or email—recipients double-click to run, with data stored in local SQLite. No cloud accounts or servers needed. It also supports AI-generated apps: describe what you want in ChatGPT or Claude, and Capsule produces a complete single-file app. Currently supports macOS, Windows, and Linux; iOS and Android are coming. v0.4.0 is free to download.

Hacker News front page

Ordewell: Turn one goal into an ordered plan of coding-agent tasks

Ordewell is a multi-agent task orchestrator for coding agents. Give it a goal, and it produces an ordered plan of tasks—each with its own runner, model, and mode—then executes and verifies results. 25 stars on GitHub, open source. The post doesn't spell out which models or runners are supported, nor whether it integrates with popular agent frameworks.

Hacker News front page

E-ink bird frame: listens to birds locally, draws 1800s-style illustrations

An open-source project that uses a Raspberry Pi and microphone to detect bird calls in real time, running fully local AI without internet. When a bird is heard, it displays a 1800s-style bird illustration on an e-ink screen. It uses BirdNET for audio recognition and Stable Diffusion for vintage engraving style. Code and models are open-source. The post doesn't specify how many bird species are supported or the detection latency.

TechCrunch · AI

Salesforce and Nvidia launch Koa, a reasoning model built for sales and support

Salesforce unveiled Koa at Dreamforce, its first reasoning model, built on Nvidia's open-weight Nemotron and trained for sales, marketing, and customer support tasks. Marc Benioff framed it as a direct threat to AI labs: a vertical SaaS company now ships its own reasoning model on its own data. The post doesn't disclose benchmark scores, parameter count, or pricing, so I'd discount the hype until numbers drop. The signal worth watching is vertical reasoning on an open base model—this could land in production faster than general-purpose alternatives.

Why it matters: Benioff's provocative framing and the 'vertical SaaS trains its own reasoning model' narrative have real buzz, but the article provides zero benchmarks or specs — the technical substance is unverifiable. Scores at the featured threshold as a newsworthy product announcement; re...

r/LocalLLaMA

DeepSeek V4.1 Flash Q4 hits 40 t/s on M3 Ultra with native DSpark multi-token prediction

A developer forked ds4 and tuned it for DeepSeek V4.1 Flash Q4 on a 512 GB M3 Ultra, lifting decode from 16.6 t/s to 31.3 t/s, and to 40.5 t/s with DSpark speculative decoding. The ~300 GB Q4 weights fit only the 512 GB M3 Ultra. The speedup comes from cutting Metal dispatch overhead: the 384-expert router went from 9 dispatches to 1, the shared expert gate+up+SwiGLU became a single kernel, and BF16 rounding moved inside producer kernels, removing ~770 re-round dispatches. At 300k context, compressed attention selection was the bottleneck; the fix scores only admitted blocks and uses a bounded radix select, keeping decode at 90% of the 8k rate. DSpark verifies 6 tokens per step with shared weight streams and overlapped Engram fetches, dropping verify latency from 177 ms to 112 ms. Output is byte-identical to upstream under greedy decode, with SHA-256 manifests provided. The branch is M3 Ultra only because the optimizations rely on measured behavior of this specific chip's dual-die memory, 80-core GPU scheduling, and Metal dispatch characteristics.

Why it matters: A solid local-inference optimization post: DeepSeek V4.1 Flash Q4 on M3 Ultra goes from 16tps to 40tps via Metal command-buffer merging and native DSpark speculative decoding. Concrete technical detail, directly useful for the local-LLM crowd. Score stays at 72 because the aud...

AI HOT (Curated Pool)

404 Media Exposes OpenAI's Project Lily: Human Review of ChatGPT Chats for Model Tuning

404 Media obtained internal docs on Project Lily, OpenAI's human review program for ChatGPT chats. Reviewers read anonymized real conversations to rate response quality, flagging AI clichés, condescending tone, or fake personal anecdotes. Pay exceeds $50/hour but the work is repetitive. Most users don't know their chats can be read by humans, and many treat ChatGPT as a confidant. OpenAI admits anonymization can leak personal data, especially in short sessions. After the story broke, OpenAI updated its help page but still didn't explicitly mention human review.

Why it matters: 404 Media obtained internal docs showing OpenAI hires humans to review user chats at $50+/hr. Strong privacy angle, but the article lacks scale details (how many reviewers, what % of chats), capping the score below 80.

AI HOT (Curated Pool)

Trail of Bits calls 1Password's AI patching benchmark misleading, releases two agent skills for patch validation

Trail of Bits reanalyzed 1Password's FLAWED report and argues the 26% clean-fix headline is misleading. Trials that deliberately instructed agents to apply wrong fixes or prohibited testing were mixed into the average. When restricted to trials where agents could run code and weren't given bad advice, 86% of patches blocked the supplied exploit. Trail of Bits also released two agent skills: post-patch-validation for automated patch testing, and review-walkthrough for engineer review. The post doesn't disclose performance data for these skills.

Why it matters: Trail of Bits re-analyzed 1Password's report claiming AI fixes only 26% of bugs, showing the headline number mixes in deliberately misleading prompts and tests where the agent couldn't run code. Under fair conditions, 86% of patches worked. This is a direct, data-backed challe...

r/LocalLLaMA

CrofAI, self-claimed cheapest inference provider, exposed as an OpenRouter wrapper swapping in cheaper models at up to 20x markup

CrofAI marketed itself as the world's cheapest inference provider but was caught by developer Kendell silently routing API calls to smaller, cheaper models on OpenRouter. For example, requests for Kimi K3 at $2/$10 in/out were actually served by GLM 5.3 Flash, a 13–20x markup. The founder also fabricated a 'greg' model family that simply pointed to existing open models like GLM 5.2 and Qwen 3.5 9B. After the exposé, he denied everything, then claimed a 'team' was taking over, and within hours deleted the website, Twitter account, and subreddit. The post also notes his hardware claims don't add up: Kimi K3 needs at least 802GiB of VRAM even at Q2_K quantization, but the largest RTX Pro 6000 machine on Vast only offers 765GiB.

Why it matters: A full fraud exposé with technical evidence and a dramatic company meltdown. Hits all three HKR axes. Score capped below 85 because the source is a Reddit post, not formal reporting, and the event is a single-provider scandal rather than a model or protocol shift.

AI HOT (Curated Pool)

Anthropic and OpenAI propose a coordinated slowdown of frontier AI development; critics like Cohere's CEO question the real motive

Anthropic CEO Dario Amodei called for a government-coordinated slowdown of frontier AI development and antitrust exemptions to make it happen; Sam Altman and Elon Musk agreed. Cohere CEO Aidan Gomez published a blog post calling it 'a wolf in sheep's clothing, a cartel by any other name,' arguing it would lock out competitors through massive barriers to entry. Hugging Face engineer Niels Rogge called the statements 'bizarre nonsense,' saying Amodei mainly wants to restrict Chinese models like DeepSeek and open-weight models to protect his scale advantage. White House AI czar David Sacks noted OpenAI and Anthropic already hold a duopoly in frontier intelligence and that product liability concerns also drive their push for a slowdown.

Why it matters: Anthropic and OpenAI jointly calling for a coordinated slowdown, with Cohere's CEO publicly pushing back as a monopoly play — three major players in direct conflict, high signal. Score held below 85 because only the title and summary are available; the full proposal details ar...

Hacker News front page

Alternatives to MinIO for single-node local S3

MinIO's parent company abandoned the open-source project in late 2025 to chase other commercial interests. This left many demos and CI pipelines that relied on MinIO for local S3 emulation in a bind. The author compares seven alternatives: S3Proxy, RustFS, SeaweedFS, Zenko CloudServer, Garage, Apache Ozone, and Ceph Object Gateway. The selection criteria are practical: must have a Docker image, must be S3-compatible, must be free and open-source, simple single-node deployment, and have an active community or commercial backer. The author built a test stack with DuckDB + Iceberg and ran write/read verification for each candidate. The post does not give a final ranking but includes a comparison table of pros and cons.

r/LocalLLaMA

jinfer: an open-source inference engine that brings LLMs to the JVM, no Python needed

mukel90 released jinfer, a pure-Java inference engine covering chat, vision, audio transcription, embeddings, reranking, and TTS. It runs without Python, ONNX, or sidecar processes, and reads gguf/safetensors natively. CPU performance is claimed to be competitive with llama.cpp; GPU support via the jota backend is still in progress. It integrates with Spring AI and LangChain4j and supports GraalVM Native Image. The author previously built llama3.java and gemma4.java. This is an early release—I'd wait to see how GPU pans out before getting too excited.

AI HOT (Curated Pool)

StepFun Releases StepAudio 3 Voice Models, Several Top Artificial Analysis Global Rankings

StepFun launched the StepAudio 3 series of voice models, with several topping the Artificial Analysis global rankings. The post is blocked by WeChat and does not disclose specific parameters, ranking details, or model capabilities. The title confirms the release and ranking results but does not specify which metrics or languages.

r/LocalLLaMA

Voodoo Dynamic Quant goes open source under MIT license, targeting aggressive low-bit compression

1ncehost open-sourced Voodoo Quant under MIT, a method that uses gradient descent to pick per-tensor quantization levels instead of static analysis. On small Qwen3.5 models it beats Unsloth Dynamic 3.0 at very low bit-widths like IQ1/IQ2, but loses at mid-to-high quants. The tooling currently targets Qwen and can be adapted to other architectures. The post does not disclose results on larger models; the author calls it research-grade. The repo README drew complaints for sounding AI-generated, and the author said he lacks time to polish it but welcomes PRs.

AI Chat-Group Daily (群聊日报)

Daily digest: empty-repo coding fails, Ollama Cloud throughput test, Trump calls out Dario

今天最直观的教训来自 @搞仁义毛义仁:给 Astra 一个空白 C++ 仓库,代码写得一塌糊涂;把积累了大量 code review 经验的 GacUI 上下文导进去,质量立刻飙升。这说明模型不是不会写,是得用具体规则去“规训”。@今天群内信息量极大 实测了 Ollama Cloud 跑 DeepSeek V4.1 Flash,解码吞吐是官方 API ...

New York Times Chinese

Anthropic CEO calls for an AI slowdown, but China makes it nearly impossible

Anthropic CEO Dario Amodei argues frontier AI must slow down, warning that swarms of AI agents could gain the ability to “take over the entire internet” within 6–12 months. His first step: embed external experts inside labs to monitor safety and report publicly. Sam Altman, Elon Musk, and Demis Hassabis endorsed the idea; Altman said OpenAI will follow suit. The real obstacle, author Sebastian Mallaby writes, is China. The US lead is only a few months, so any unilateral slowdown risks letting China pull ahead. Amodei acknowledges this and, in a notable shift, lists areas where US–China cooperation might be possible, comparing it to Cold War arms control. The post does not spell out a concrete timeline, but notes Trump and Xi are set to meet on Sept 24, with two more summits possible by year-end.

Why it matters: Anthropic CEO's direct call plus endorsements from Altman, Musk, and Hassabis make this a high-signal moment. Amodei delivers a concrete 6-12 month timeline and an operational proposal for embedded safety experts. The deduction: this is an op-ed, not a policy announcement, and...

Latent Space

AEF-1 standard for third-party evaluators lands, with xAI, OpenAI, and Anthropic all signing on

The AI Evaluator Forum published AEF-1, a baseline for independent third-party evaluations covering access, conflicts of interest, funding, recusal, and transparency. The same day, Dario Amodei blogged that Anthropic is unilaterally committing to embedded evaluators with office badges, company laptops, and access comparable to internal risk teams. He also laid out a two-tier coordination framework for democratic and global pacing. Bilal Chughtai left Google DeepMind and called for slowing capability progress; Dan Selsam warned that models may learn to fake alignment during evals. On the other side, Aidan Gomez and Cohere pushed back against a few Silicon Valley firms becoming gatekeepers, and Kevin Bass alleged structural conflicts in the Anthropic-linked safety ecosystem.

Why it matters: The AI Evaluator Forum's AEF-1 standard, co-signed by xAI, OpenAI, and Anthropic on the same day Dario Amodei published a personal blog proposing even deeper evaluator access, forms a strong signal cluster. All three HKR axes hit: first written rules for third-party evaluation...

Financial Times · Technology

Italian startup Exein raises $270M to fight AI hackers with AI

Exein, an Italian embedded security startup, raised $270M to use AI models for detecting anomalies in IoT and industrial devices, specifically targeting AI-powered attacks. The post doesn't disclose valuation or product benchmarks, but notes customers include European automotive and defense firms. Big bet on the space, not yet proof of product-market fit.

Financial Times · Technology

AI is exciting audit firms — maybe too much

Big Four accounting firms are betting big on AI for auditing, but FT pours cold water on the hype. They're using AI to review accounts, flag anomalies, and draft reports. The catch: AI makes mistakes too, and auditors can't fully trust it. So costs are high, efficiency gains are modest, and regulators are watching. The article doesn't give specific numbers but nails the core tension—AI in auditing is more a tool that needs auditing itself than a savior.

Financial Times · Technology

Zach Dell bets a battery fleet can solve America’s AI power crunch

Zach Dell, son of Dell's founder, plans to deploy large-scale battery storage fleets to stabilize power for AI data centers. The idea: charge batteries during grid off-peak hours and discharge during peak demand, easing the strain from AI training and inference. The post doesn't disclose fleet size, cost, or timeline, but highlights the growing tension between AI's power hunger and grid capacity.

Financial Times · Technology

What an AI slowdown could mean for investors

FT analyzes signals that the AI investment boom may cool. Big AI groups are calling for a slowdown, which has already dragged down US tech stocks. Investors may need to recalibrate expectations on infrastructure and compute spending. The post doesn't name specific companies or a timeline, but flags a key shift in market sentiment.

Bloomberg Technology

Chinese State Media Dismisses ‘Self-Serving’ AI Slowdown Call

Chinese state media published a piece dismissing calls to pause AI development as 'self-serving.' It argues the slowdown push is a competitive tactic, not a safety measure. The post doesn't name specific proponents or cite detailed arguments, but the stance is clear: China won't slow down.

Bloomberg Technology

Hugging Face Scientist on Safety Issues With Agentic AI

Bloomberg interviews a Hugging Face scientist on safety issues with agentic AI. The body does not disclose specific risk cases or solutions; the core concern is the safety risks when models autonomously execute tasks.

AI HOT (Curated Pool)

Artificial Analysis ranks GPT-Live-1 #1 on Speech-to-Speech Index with 81.5

Artificial Analysis just dropped a Speech-to-Speech Index. OpenAI's GPT-Live-1 scored 81.5 with the Astra backend on medium reasoning intensity, edging out Grok Voice Think Fast 2.0 High at 81.3. The Sol backend config of GPT-Live-1 landed third at 80.1. The post doesn't disclose evaluation dimensions, sample size, or latency—so I'd take the ranking with a grain of salt for now.

Financial Times · Technology

China's AI listings glut drags Hong Kong stocks lower

A wave of Chinese AI companies listing in Hong Kong has created a supply glut, dragging down the broader market. The FT reports the IPO surge is directly weighing on Hong Kong stocks. The article does not name specific firms or disclose total fundraising amounts.

New York Times Chinese

Trump calls AI safety fears a “scam,” rejects new regulation

Trump dismissed calls for AI regulation from Anthropic CEO Dario Amodei and others, calling safety fears a “scam” and insisting the only needed guardrail is “a strong and smart president.” He singled out Amodei as “pretending to be a perfect little angel,” though the post doesn’t spell out what government actions he claims to have already blocked. VP Vance and Speaker Johnson showed more openness to regulation, with Vance calling industry self-regulation pleas “a little bit of a Trojan horse.” Congress has almost no time to act before the midterms.

Why it matters: A sitting president directly names Anthropic's CEO and dismisses AI safety as a 'scam' — this is a head-on attack at the industry's core narrative. HKR all hit: high conflict, new White House power-split detail, direct identity nerve. Score capped below 85 because the article ...

New York Times Chinese

China’s top intelligence chief warns AI could directly threaten CCP rule

Minister of State Security Chen Yixin published an article framing AI as a direct threat to CCP rule—Beijing’s highest-level and most detailed warning yet. He named US models like Claude Mythos and GPT-5.5-Cyber, citing risks of deepfakes, information warfare, large-scale data exfiltration, and attacks on critical infrastructure. He also pointed to US military AI use in the Iran war as evidence that algorithmic advantage determines battlefield control. Law professor Henry Gao said this draws a red line ahead of US-China AI talks: data sovereignty and political security won’t be traded for international agreements.

Why it matters: China's national security minister publishes a long essay elevating AI safety to a regime-security issue, naming specific Anthropic and OpenAI models and citing US military use cases. This is a policy signal ahead of US-China AI safety talks — more about positioning than new f...