
OpenAI shipped GPT-6 Astra with an AGI claim and a build that refuses exploit work
Brockman called it the AGI era. Astra shipped to Daybreak enterprises first at $10/$50 per million tokens, and the cyber build stays gated.
Large language model releases, benchmarks, capability jumps, and the infrastructure that runs them.

Brockman called it the AGI era. Astra shipped to Daybreak enterprises first at $10/$50 per million tokens, and the cyber build stays gated.

Anthropic's Fable 5.1 doubles a science benchmark and cuts cache reads 75%. Independent tests measured tasks costing up to 3x more than Fable 5.

GenAI.mil now serves OpenAI's ChatGPT Mil and Starshield AI's Grok for Government at IL5, nine months after Google's Gemini opened the portal.

Uber's year of AI budget lasted until April. A 15-company survey shows per-developer spend running from $200 to $3,000 a month, and finance has noticed.

Five Rust teams ratified an LLM policy that bans AI-written docs and mandates disclosure. Linux leaves it to each maintainer. Both are rationing review time.

Qwen3.8-Max is a 2.4-trillion-parameter MoE that tops Claude Fable 5 on Terminal-Bench 2.1 and trails it badly on SWE-bench Pro. It's the first open Max-tier Qwen.

OpenAI published ten results in math and theoretical CS from an internal build of Astra, with Lean 4 certificates for every proof. What that verification does and doesn't settle.

Salvatore Sanfilippo repacked DeepSeek V4 Flash into a lossless MXFP4 GGUF that streams from SSD at over 20 tokens a second. The hardware bill, and where hosted still wins.

Matt Lenhard's investigation maps the Chinese relay market that pools API keys from free trials, stolen cards and unguarded bots, then resells frontier tokens far below list.

Codeberg's terms now bar projects that mostly consist of AI-written code. Debian's open resolution puts three answers on one ballot. Provenance is the crux.

Anthropic shipped Claude Opus 5 at the same $5/$25 per million tokens as Opus 4.8. It nears Fable 5's intelligence at half the cost, with new effort and fallback controls.

Google released Gemini 3.6 Flash, 3.5 Flash-Lite, and a cyber specialist, and teased Gemini 4. The 3.5 Pro tier it promised in May still isn't out.

Kimi K3, GLM 5.2 and DeepSeek V4 put open-weight AI next to the frontier this month. What each model is good at, and why the benchmarks mislead.

GitHub benchmarked Copilot's agent harness against Claude Code and Codex CLI on five tests. The token savings are real, and the best model depends on the task.

Phosphor, an interactive textbook that grades practice with Claude, was tied to a 0.71 to 1.30 SD final-exam gain in a Dartmouth statistics course.

Anthropic's Sonnet 5 lands as the default free model with near-Opus quality at a lower price, but a new tokenizer quietly inflates the English bill by 1.4x.

The Trump administration asked OpenAI to limit GPT-5.6 to trusted partners, with the government vetting access customer by customer. Here's what that gatekeeping means.

A Brown economist found AI fraud across a midterm. The scandal exposes how routine AI cheating has become, and why detectors can't reliably catch it.