devtake.dev

#llm

RSS

Large language model releases, benchmarks, capability jumps, and the infrastructure that runs them.

OpenAI's GPT-6 Astra launch card, showing a white spiral galaxy of stars on black with the words GPT-6 and Astra set on either side.
AI·

OpenAI shipped GPT-6 Astra with an AGI claim and a build that refuses exploit work

Brockman called it the AGI era. Astra shipped to Daybreak enterprises first at $10/$50 per million tokens, and the cyber build stays gated.

Anthropic's launch card for Claude Fable 5.1 and Claude Mythos 5.1, white serif text over a pale blue sky with clouds and a daytime moon.
AI·

Six Claude 5 models in twelve weeks. Fable 5.1 is the first to watermark its output.

Anthropic's Fable 5.1 doubles a science benchmark and cuts cache reads 75%. Independent tests measured tasks costing up to 3x more than Fable 5.

Aerial photograph of the Pentagon's five concentric rings and central courtyard, ringed by parking lots and highway interchanges on a bare winter day.
Policy·

The Pentagon added ChatGPT and Grok to its own AI portal for 3 million staff

GenAI.mil now serves OpenAI's ChatGPT Mil and Starshield AI's Grok for Government at IL5, nine months after Google's Gemini opened the portal.

The Uber wordmark and square U logo in white and silver on a black background
AI·

Uber spent its entire 2026 AI budget in four months, and the rest of the industry is now metering tokens

Uber's year of AI budget lasted until April. A 15-company survey shows per-developer spend running from $200 to $3,000 a month, and finance has noticed.

The Rust gear-and-R logo in black beside the words The Rust Programming Language on a white background.
Open Source·

Rust's new rules require disclosing LLM use. A Linux maintainer gives AI patches three seconds.

Five Rust teams ratified an LLM policy that bans AI-written docs and mandates disclosure. Linux leaves it to each maintainer. Both are rationing review time.

An Alibaba Group office building with the company's orange logo sign above the entrance and cars parked outside.
AI·

Alibaba's Qwen3.8-Max beats Fable 5 on Terminal-Bench, and the weights go public next week

Qwen3.8-Max is a 2.4-trillion-parameter MoE that tops Claude Fable 5 on Terminal-Bench 2.1 and trails it badly on SWE-bench Pro. It's the first open Max-tier Qwen.

The GitHub social preview card for the openai/ten-proofs repository, described as Lean certificates accompanying proofs in mathematics and theoretical computer science, showing 386 stars and 35 forks.
AI·

Ten decade-old math problems fell to an unreleased OpenAI model, for about $2,000 of tokens each

OpenAI published ten results in math and theoretical CS from an internal build of Astra, with Lean 4 certificates for every proof. What that verification does and doesn't settle.

The Hugging Face model page card for deepseek-ai/DeepSeek-V4-Flash-0731, showing the DeepSeek whale logo and the repository name
AI·

DeepSeek's new 304B agentic model now runs on a single 128GB workstation

Salvatore Sanfilippo repacked DeepSeek V4 Flash into a lossless MXFP4 GGUF that streams from SSD at over 20 tokens a second. The hardware bill, and where hosted still wins.

GitHub repository card for songquanpeng/one-api, the open-source LLM API management and distribution gateway that most relay services run on
AI·

Matt Lenhard found 49 relays reselling OpenAI and Anthropic tokens. The cheapest runs 97.8% below list.

Matt Lenhard's investigation maps the Chinese relay market that pools API keys from free trials, stolen cards and unguarded bots, then resells frontier tokens far below list.

The Debian swirl logo in dark red above the lowercase debian wordmark on a near-black background.
Open Source·

Codeberg banned LLM-generated projects, and Debian is voting on the same question

Codeberg's terms now bar projects that mostly consist of AI-written code. Debian's open resolution puts three answers on one ballot. Provenance is the crux.

Anthropic's Claude Opus 5 announcement artwork: a large numeral 5 formed from an arrangement of vintage speckled bird-egg illustrations on a cream background.
AI·

Claude Opus 5 nears Fable 5's frontier intelligence at half the price

Anthropic shipped Claude Opus 5 at the same $5/$25 per million tokens as Opus 4.8. It nears Fable 5's intelligence at half the cost, with new effort and fallback controls.

Google Gemini key art showing the 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber model names
AI·

Google shipped three Gemini Flash models but held back its flagship Pro

Google released Gemini 3.6 Flash, 3.5 Flash-Lite, and a cyber specialist, and teased Gemini 4. The 3.5 Pro tier it promised in May still isn't out.

Two pelicans face each other with crossed beaks on dark water, mirrored in the surface, a nod to the informal 'pelican on a bicycle' LLM benchmark
AI·

Kimi K3 trades blows with Anthropic's Fable, and Moonshot is opening the weights

Kimi K3, GLM 5.2 and DeepSeek V4 put open-weight AI next to the frontier this month. What each model is good at, and why the benchmarks mislead.

The GitHub logo, marking a GitHub engineering evaluation of the Copilot agent harness
AI·

GitHub ran four frontier models through Copilot's harness. None won every task.

GitHub benchmarked Copilot's agent harness against Claude Code and Codex CLI on five tests. The token savings are real, and the best model depends on the task.

Students seated in a university lecture hall
AI·

A Dartmouth AI textbook is tied to final-exam gains of up to 1.30 standard deviations

Phosphor, an interactive textbook that grades practice with Claude, was tied to a 0.71 to 1.30 SD final-exam gain in a Dartmouth statistics course.

Anthropic Claude Sonnet 5 announcement graphic
AI·

Claude Sonnet 5: cheaper agents on paper, until you count the new tokenizer's tokens

Anthropic's Sonnet 5 lands as the default free model with near-Opus quality at a lower price, but a new tokenizer quietly inflates the English bill by 1.4x.

The north facade of the White House in Washington under a clear sky.
Policy·

The White House told OpenAI to gate GPT-5.6. Frontier models now need government sign-off.

The Trump administration asked OpenAI to limit GPT-5.6 to trusted partners, with the government vetting access customer by customer. Here's what that gatekeeping means.

University students filing into an examination hall to sit a written exam
AI·

A Brown professor caught 40 of 86 students cheating with AI. Now he wants take-home exams gone.

A Brown economist found AI fraud across a midterm. The scandal exposes how routine AI cheating has become, and why detectors can't reliably catch it.