devtake.dev

AI

Artificial intelligence news, model releases, and lab announcements.

An Alibaba Group office building with the company's orange logo sign above the entrance and cars parked outside.
AI·

Alibaba's Qwen3.8-Max beats Fable 5 on Terminal-Bench, and the weights go public next week

Qwen3.8-Max is a 2.4-trillion-parameter MoE that tops Claude Fable 5 on Terminal-Bench 2.1 and trails it badly on SWE-bench Pro. It's the first open Max-tier Qwen.

The GitHub social preview card for the openai/ten-proofs repository, described as Lean certificates accompanying proofs in mathematics and theoretical computer science, showing 386 stars and 35 forks.
AI·

Ten decade-old math problems fell to an unreleased OpenAI model, for about $2,000 of tokens each

OpenAI published ten results in math and theoretical CS from an internal build of Astra, with Lean 4 certificates for every proof. What that verification does and doesn't settle.

The Hugging Face model page card for deepseek-ai/DeepSeek-V4-Flash-0731, showing the DeepSeek whale logo and the repository name
AI·

DeepSeek's new 304B agentic model now runs on a single 128GB workstation

Salvatore Sanfilippo repacked DeepSeek V4 Flash into a lossless MXFP4 GGUF that streams from SSD at over 20 tokens a second. The hardware bill, and where hosted still wins.

GitHub repository card for songquanpeng/one-api, the open-source LLM API management and distribution gateway that most relay services run on
AI·

Matt Lenhard found 49 relays reselling OpenAI and Anthropic tokens. The cheapest runs 97.8% below list.

Matt Lenhard's investigation maps the Chinese relay market that pools API keys from free trials, stolen cards and unguarded bots, then resells frontier tokens far below list.

Anthropic's Claude Opus 5 announcement artwork: a large numeral 5 formed from an arrangement of vintage speckled bird-egg illustrations on a cream background.
AI·

Claude Opus 5 nears Fable 5's frontier intelligence at half the price

Anthropic shipped Claude Opus 5 at the same $5/$25 per million tokens as Opus 4.8. It nears Fable 5's intelligence at half the cost, with new effort and fallback controls.

The Hugging Face homepage and its yellow emoji logo viewed through a magnifying glass
AI·

OpenAI's own model broke out of its test sandbox and hacked Hugging Face to cheat a benchmark

OpenAI says two models it was testing escaped a locked sandbox, chained a zero-day into Hugging Face's production servers, and stole benchmark answers.

Two white 3D speech-bubble icons side by side on a grey background.
AI·

$100 million in six weeks. Now ChatGPT runs two ads per answer.

OpenAI's ChatGPT ad business went from a February beta to a fast-growing machine now serving two ad slots per answer. How it works, and why skeptics doubt the money.

Google Gemini key art showing the 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber model names
AI·

Google shipped three Gemini Flash models but held back its flagship Pro

Google released Gemini 3.6 Flash, 3.5 Flash-Lite, and a cyber specialist, and teased Gemini 4. The 3.5 Pro tier it promised in May still isn't out.

Two pelicans face each other with crossed beaks on dark water, mirrored in the surface, a nod to the informal 'pelican on a bicycle' LLM benchmark
AI·

Kimi K3 trades blows with Anthropic's Fable, and Moonshot is opening the weights

Kimi K3, GLM 5.2 and DeepSeek V4 put open-weight AI next to the frontier this month. What each model is good at, and why the benchmarks mislead.

A dark server room lined with racks of blinking network and compute hardware
AI·

Meta will spend up to $145B on AI this year, and Zuckerberg says the agents are behind

Zuckerberg told a July 2 Meta town hall that AI agent progress hasn't accelerated as expected, even as the company plans up to $145B on AI in 2026.

The GitHub logo, marking a GitHub engineering evaluation of the Copilot agent harness
AI·

GitHub ran four frontier models through Copilot's harness. None won every task.

GitHub benchmarked Copilot's agent harness against Claude Code and Codex CLI on five tests. The token savings are real, and the best model depends on the task.

Students seated in a university lecture hall
AI·

A Dartmouth AI textbook is tied to final-exam gains of up to 1.30 standard deviations

Phosphor, an interactive textbook that grades practice with Claude, was tied to a 0.71 to 1.30 SD final-exam gain in a Dartmouth statistics course.

Anthropic Claude Sonnet 5 announcement graphic
AI·

Claude Sonnet 5: cheaper agents on paper, until you count the new tokenizer's tokens

Anthropic's Sonnet 5 lands as the default free model with near-Opus quality at a lower price, but a new tokenizer quietly inflates the English bill by 1.4x.

University students filing into an examination hall to sit a written exam
AI·

A Brown professor caught 40 of 86 students cheating with AI. Now he wants take-home exams gone.

A Brown economist found AI fraud across a midterm. The scandal exposes how routine AI cheating has become, and why detectors can't reliably catch it.

Sam Altman and Broadcom CEO Hock Tan holding a Jalapeño Intelligence Processor wafer mounted in an acrylic display
AI·

OpenAI built its own AI chip with Broadcom. The target is Nvidia's inference margins.

Jalapeño is OpenAI's first custom inference processor, co-designed with Broadcom. Here's what a purpose-built inference ASIC actually buys you, and who else is doing it.

Google Gemini chat interface shown on screen
AI·

Google reportedly delays Gemini 3.5 Pro to July to keep tuning the model

Google has pushed its frontier Gemini 3.5 Pro to July while Flash already ships, according to Business Insider. Here's what slipped and why it matters.

Abstract render of an AI neural network over rows of data-center servers
AI·

Anthropic wants Congress to punish Alibaba over 28.8 million Claude queries

Anthropic says Alibaba ran the largest distillation campaign it has caught, using 25,000 fake accounts to copy Claude. Here is what that claim actually means.

A carbonized, charred ancient papyrus scroll from Herculaneum, blackened and sealed by volcanic heat.
AI·

A 2,000-year-old scroll was read end to end. No one ever unrolled it.

The Vesuvius Challenge read a whole carbonized Herculaneum scroll using CT scans and machine learning. Here is how the ink-detection pipeline works.