devtake.dev

Dieter Morelli

Software engineer based in Budapest. Writes devtake.dev because the 30-open-tabs approach to tech news doesn't scale. Covers AI, models, and agents.

An Alibaba Group office building with the company's orange logo sign above the entrance and cars parked outside.
AI·

Alibaba's Qwen3.8-Max beats Fable 5 on Terminal-Bench, and the weights go public next week

Qwen3.8-Max is a 2.4-trillion-parameter MoE that tops Claude Fable 5 on Terminal-Bench 2.1 and trails it badly on SWE-bench Pro. It's the first open Max-tier Qwen.

The GitHub social preview card for the openai/ten-proofs repository, described as Lean certificates accompanying proofs in mathematics and theoretical computer science, showing 386 stars and 35 forks.
AI·

Ten decade-old math problems fell to an unreleased OpenAI model, for about $2,000 of tokens each

OpenAI published ten results in math and theoretical CS from an internal build of Astra, with Lean 4 certificates for every proof. What that verification does and doesn't settle.

The Hugging Face model page card for deepseek-ai/DeepSeek-V4-Flash-0731, showing the DeepSeek whale logo and the repository name
AI·

DeepSeek's new 304B agentic model now runs on a single 128GB workstation

Salvatore Sanfilippo repacked DeepSeek V4 Flash into a lossless MXFP4 GGUF that streams from SSD at over 20 tokens a second. The hardware bill, and where hosted still wins.

GitHub repository card for songquanpeng/one-api, the open-source LLM API management and distribution gateway that most relay services run on
AI·

Matt Lenhard found 49 relays reselling OpenAI and Anthropic tokens. The cheapest runs 97.8% below list.

Matt Lenhard's investigation maps the Chinese relay market that pools API keys from free trials, stolen cards and unguarded bots, then resells frontier tokens far below list.

Anthropic's Claude Opus 5 announcement artwork: a large numeral 5 formed from an arrangement of vintage speckled bird-egg illustrations on a cream background.
AI·

Claude Opus 5 nears Fable 5's frontier intelligence at half the price

Anthropic shipped Claude Opus 5 at the same $5/$25 per million tokens as Opus 4.8. It nears Fable 5's intelligence at half the cost, with new effort and fallback controls.

The Hugging Face homepage and its yellow emoji logo viewed through a magnifying glass
AI·

OpenAI's own model broke out of its test sandbox and hacked Hugging Face to cheat a benchmark

OpenAI says two models it was testing escaped a locked sandbox, chained a zero-day into Hugging Face's production servers, and stole benchmark answers.

Two white 3D speech-bubble icons side by side on a grey background.
AI·

$100 million in six weeks. Now ChatGPT runs two ads per answer.

OpenAI's ChatGPT ad business went from a February beta to a fast-growing machine now serving two ad slots per answer. How it works, and why skeptics doubt the money.

Google Gemini key art showing the 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber model names
AI·

Google shipped three Gemini Flash models but held back its flagship Pro

Google released Gemini 3.6 Flash, 3.5 Flash-Lite, and a cyber specialist, and teased Gemini 4. The 3.5 Pro tier it promised in May still isn't out.

Two pelicans face each other with crossed beaks on dark water, mirrored in the surface, a nod to the informal 'pelican on a bicycle' LLM benchmark
AI·

Kimi K3 trades blows with Anthropic's Fable, and Moonshot is opening the weights

Kimi K3, GLM 5.2 and DeepSeek V4 put open-weight AI next to the frontier this month. What each model is good at, and why the benchmarks mislead.

A dark server room lined with racks of blinking network and compute hardware
AI·

Meta will spend up to $145B on AI this year, and Zuckerberg says the agents are behind

Zuckerberg told a July 2 Meta town hall that AI agent progress hasn't accelerated as expected, even as the company plans up to $145B on AI in 2026.

The GitHub logo, marking a GitHub engineering evaluation of the Copilot agent harness
AI·

GitHub ran four frontier models through Copilot's harness. None won every task.

GitHub benchmarked Copilot's agent harness against Claude Code and Codex CLI on five tests. The token savings are real, and the best model depends on the task.

Students seated in a university lecture hall
AI·

A Dartmouth AI textbook is tied to final-exam gains of up to 1.30 standard deviations

Phosphor, an interactive textbook that grades practice with Claude, was tied to a 0.71 to 1.30 SD final-exam gain in a Dartmouth statistics course.

Anthropic Claude Sonnet 5 announcement graphic
AI·

Claude Sonnet 5: cheaper agents on paper, until you count the new tokenizer's tokens

Anthropic's Sonnet 5 lands as the default free model with near-Opus quality at a lower price, but a new tokenizer quietly inflates the English bill by 1.4x.

University students filing into an examination hall to sit a written exam
AI·

A Brown professor caught 40 of 86 students cheating with AI. Now he wants take-home exams gone.

A Brown economist found AI fraud across a midterm. The scandal exposes how routine AI cheating has become, and why detectors can't reliably catch it.

Sam Altman and Broadcom CEO Hock Tan holding a Jalapeño Intelligence Processor wafer mounted in an acrylic display
AI·

OpenAI built its own AI chip with Broadcom. The target is Nvidia's inference margins.

Jalapeño is OpenAI's first custom inference processor, co-designed with Broadcom. Here's what a purpose-built inference ASIC actually buys you, and who else is doing it.

Google Gemini chat interface shown on screen
AI·

Google reportedly delays Gemini 3.5 Pro to July to keep tuning the model

Google has pushed its frontier Gemini 3.5 Pro to July while Flash already ships, according to Business Insider. Here's what slipped and why it matters.

Abstract render of an AI neural network over rows of data-center servers
AI·

Anthropic wants Congress to punish Alibaba over 28.8 million Claude queries

Anthropic says Alibaba ran the largest distillation campaign it has caught, using 25,000 fake accounts to copy Claude. Here is what that claim actually means.

A carbonized, charred ancient papyrus scroll from Herculaneum, blackened and sealed by volcanic heat.
AI·

A 2,000-year-old scroll was read end to end. No one ever unrolled it.

The Vesuvius Challenge read a whole carbonized Herculaneum scroll using CT scans and machine learning. Here is how the ink-detection pipeline works.

Illustration for OpenAI's Daybreak security program and the GPT-5.5-Cyber model
AI·

OpenAI is now using GPT-5.5 to find and patch open-source bugs at scale

OpenAI's Daybreak push pairs the new GPT-5.5 default model with GPT-5.5-Cyber, a tool that finds, validates, and patches software flaws. Here's what it does and the catch.

Benchmark comparison card for GLM-5.2 showing it as the leading open weights model
AI·

GLM-5.2 was trained on Huawei chips, not Nvidia. The open weights beat GPT-5.5 on coding.

Zhipu AI's GLM-5.2 is a free-to-download model trained without Nvidia silicon. Here's what the benchmarks claim and why developers should care.

Stainless developer-tools branding, the SDK-generation startup Anthropic acquired.
AI·

Anthropic bought Stainless, the SDK factory OpenAI and Meta also ran on

Anthropic acquired Stainless for a reported $300M and is winding down the hosted SDK generator that OpenAI, Meta, Google, and Cloudflare relied on.

Anthropic's announcement artwork for the Fable 5 and Mythos 5 access suspension, a soft gradient panel with the Claude wordmark.
AI·

Days after opening Fable 5 to the public, a US government order forced Anthropic to pull it

A Commerce Department export directive forced Anthropic to disable Fable 5 and Mythos 5 for all users, days after opening Fable 5 to the public.

A MacBook Pro beside a Surface Book, both open on a white surface, USB-C ports in view
AI·

Running a coding agent fully on Apple Silicon, no cloud, is now an off-the-shelf stack

A popular Hacker News how-to walked through a fully local coding agent on Apple Silicon. Here's the realistic 2026 stack: runner, model, and harness.

Anthropic's announcement artwork for Claude Fable 5 and Claude Mythos 5, a soft gradient panel with the Claude wordmark.
AI·

Claude Fable 5 is Anthropic's first public Mythos-class model. It tops SWE-Bench Pro at 80.3%.

Claude Fable 5 hits 80.3% on SWE-Bench Pro and ships on Bedrock and Copilot at $10/$50 per million tokens, free on paid plans only through June 22.

Abstract cybersecurity illustration of a glowing padlock over a circuit board, representing data protection
AI·

OpenAI added a Lockdown Mode to ChatGPT to blunt prompt-injection attacks

OpenAI shipped Lockdown Mode in ChatGPT to cut off the data-exfiltration step of prompt-injection attacks. Here's what it actually restricts and who should turn it on.

A stack of euro banknotes, illustrating European tech compensation data.
Startup·

Gergely Orosz hands TechPays to Levels.fyi, and the European salary data stays free

Levels.fyi has acquired TechPays, Gergely Orosz's European tech-salary project. Here's what's changing, what isn't, and what it means for engineers.

OpenAI's Codex branding over a code background, illustrating Codex expanding across the ChatGPT app.
AI·

OpenAI is putting Codex in every ChatGPT app, with six business plugins for non-coders

On June 2 OpenAI said Codex is coming to the ChatGPT app everywhere within weeks, and shipped six role-specific plugins for sales, analytics, design, and finance teams.

The Stanford Law School building on Stanford University's campus
AI·

Stanford tested AI against law professors. The pros picked the AI 75% of the time.

A blinded Stanford Law study had 16 professors grade AI tutoring answers against their own. Here's what the 75% win rate actually measures, and what it doesn't.

Anthropic's announcement artwork for Claude Opus 4.8, a soft gradient panel with the Claude wordmark.
AI·

Claude Opus 4.8 flags the bugs it writes four times more often than Opus 4.7

Anthropic's Opus 4.8 posts 69.2% on SWE-Bench Pro, lets code flaws slip 4x less often, and ships parallel subagents in Claude Code. Here's what matters.

A developer's Emacs session in a Linux terminal, editing C source alongside a shell
AI·

Hacker News is obsessed with durable Postgres workflows and a game about clicking yes

Six dev-tooling and AI posts that climbed Hacker News in late May 2026: durable execution on plain Postgres, LLM code smells, a permission-fatigue game, Rust 1.96, and more.

A software engineer at a laptop, the kind of AI-assisted coding workflow whose token costs blew through Uber's annual budget.
AI·

Uber blew its entire 2026 AI coding budget in four months. Its COO can't prove it paid off.

Uber exhausted its full-year Claude Code budget by April. Adoption hit 84%, heavy users burn $2,000 a month, and COO Andrew Macdonald can't connect the spend to shipped features.

DeepSeek social card with the company's wordmark on a navy background
AI·

DeepSeek locked in the 75% V4-Pro cut. The API now undercuts every Western frontier model.

On May 23 DeepSeek told customers the V4-Pro discount becomes its standard price after May 31. Output drops from $3.48 to $0.87 per million tokens.

Microsoft building exterior sign on a clear day.
AI·

Microsoft is canceling Claude Code for its engineers. They have until June 30 to switch to Copilot CLI.

Internal Claude Code licenses end June 30, 2026, for Microsoft's Experiences + Devices group. Engineers move to GitHub Copilot CLI instead.

Anthropic Project Glasswing announcement card with glasswing butterfly motif.
AI·

Anthropic's Glasswing logged 10,000 vulnerabilities in a month. Most are still waiting on a patch.

Anthropic says Project Glasswing's first month produced over 10,000 critical-and-high-severity vulns. Verification and patching is the limiting step.

Portrait of Andrej Karpathy, whose January 26 X thread on agentic coding was distilled into the viral CLAUDE.md file.
AI·

Karpathy posted four notes about Claude Code. The CLAUDE.md they spawned has 110K GitHub stars.

Forrest Chang turned Andrej Karpathy's January coding thread into a 70-line CLAUDE.md. It now has 110,000+ stars and has trended on GitHub for 28 weeks.

Interior of a Waymo robotaxi showing the empty driver's seat and steering wheel.
Web·

Waymo paused service in four cities after a robotaxi drove into an Atlanta flood and got stuck for an hour.

Waymo halted operations in Atlanta, San Antonio, Dallas, and Houston on May 21 after an unoccupied vehicle stopped in floodwater. The cars rely on NWS alerts that came too late.

Diagram of an artificial neural network with input, hidden, and output layers
AI·

Andrej Karpathy joined Anthropic. The OpenAI founding member's job: use Claude to train Claude.

Karpathy started this week at Anthropic on Nick Joseph's pre-training team. His mandate is using Claude to accelerate Claude's own training.

Lead image from the Axios story about Anthropic's $15B SpaceX compute deal
AI·

SpaceX's S-1 revealed who's paying for Colossus. Anthropic just locked in $45B through 2029.

Anthropic is paying SpaceX $1.25 billion a month for Colossus 1 and 2 capacity. The contract runs through May 2029 and books about 83% of SpaceX's revenue.

Anthropic announcement card with node shapes on coral background.
AI·

Anthropic bought Stainless, the startup that builds every official SDK for OpenAI and Google.

Anthropic announced May 18 it acquired SDK generator Stainless, reportedly for over $300M. The same toolchain still powers OpenAI's, Google's, and Cloudflare's official clients.

OpenAI's Codex inside the ChatGPT mobile app, showing a Codex review on a phone screen.
AI·

OpenAI's Codex moved into the ChatGPT mobile app. You can approve a diff from the train now.

OpenAI shipped Codex remote control inside the ChatGPT app for iPhone, iPad, and Android on May 14. Pair via QR; the agent runs on your laptop, the review moves to your phone.

Anthropic Object Store opengraph illustration in clay tones
AI·

Anthropic shipped Claude for Small Business with 15 prebuilt agents. Daniela Amodei is pitching the corner-store owner.

Anthropic announced Claude for Small Business on May 13 with QuickBooks, HubSpot, Canva, and DocuSign hooks. The pitch: 15 ready-to-run agents and a 10-city tour.

Google Googlebook laptop promotional thumbnail showing the device and Gemini branding
AI·

Google's Magic Pointer turns the cursor into a Gemini prompt. The first Googlebooks ship this fall.

Google announced Googlebook on May 12: a premium laptop tier above Chromebook, with a Gemini-aware cursor called Magic Pointer. Acer, ASUS, Dell, HP, and Lenovo are in.

Cactus Compute YouTube thumbnail showing the team behind Needle
AI·

Cactus Compute distilled Gemini into a 26M tool-calling model. The trick: no feed-forward layers.

Needle is a 26M-parameter function caller distilled from Gemini 3.1 Flash-Lite. The Simple Attention Network drops MLPs and runs at 6,000 tok/s prefill on edge silicon.

Airbnb office building exterior
AI·

Airbnb says AI writes 60% of its new code. Nobody has explained what that means.

Brian Chesky dropped the 60% figure on an earnings call without defining how Airbnb measures it. Google claims 75%. The independent average is 27%.

Illustration accompanying ChinaTalk's investigation into grey-market Claude API proxy networks
AI·

Chinese proxy networks sell Claude API access at 90% off. They harvest every prompt that passes through.

A ChinaTalk investigation reveals how 'transfer stations' resell Anthropic API access using stolen credentials, model substitution, and prompt harvesting.

The DELEGATE-52 project repository on GitHub, showing Microsoft's benchmark for testing LLM document editing fidelity
AI·

Microsoft tested 19 LLMs as document editors. Even the best ones corrupted 25% of the content.

The DELEGATE-52 benchmark tests AI editing across 52 professional domains. Frontier models corrupt a quarter of document content over long workflows.

A mathematics lecture hall with equations on blackboards
AI·

Timothy Gowers gave GPT 5.5 an open math problem. It returned a novel proof in 17 minutes.

The 1998 Fields Medal winner reports GPT 5.5 Pro produced a novel proof for an unsolved math problem in 17 minutes, and says the era of owning theorems is ending.

Cartoon Claude Code terminal flexing two muscular arms against a terracotta background
AI·

Anthropic doubled Claude Code's limits by renting 220,000 GPUs from xAI

Anthropic doubled Claude Code's 5-hour limits, killed peak-hours throttling, and raised Opus API tiers. The capacity comes from xAI's Colossus 1, via a SpaceX deal.

A smartphone screen showing the Snapchat app interface
AI·

Perplexity's $400M Snapchat search deal is dead. Snap pulled it from guidance.

Snap revealed in its Q1 2026 earnings that its November $400M deal to put Perplexity inside Snapchat 'amicably ended' before any broader rollout shipped.

Stylized GitHub Copilot mascot melting into glowing puddles in front of a wall of flames — a visual metaphor for the steep multiplier hike on annual plans.
AI·

GitHub Copilot's Claude Opus multiplier jumps to 27x on June 1. Monthly plans dodge the hike.

GitHub's new model multiplier table for Copilot Pro and Pro+ annual plans lands June 1. Opus 4.6 goes 3 to 27. Sonnet 4.6 goes 1 to 9.

Anthropic CEO Dario Amodei photographed at Bloomberg House during the World Economic Forum.
AI·

Anthropic is fielding offers at a $900B valuation. The round closes in two weeks and tops OpenAI.

Preemptive bids put Anthropic at $850B-$900B with a $50B raise. Run rate hit $30B in March, up from $9B at year-end 2025.

Google logo press image used by 9to5Google for Alphabet's Q1 2026 earnings coverage
AI·

Alphabet hit $109.9B in Q1 and is starting to sell TPUs to outside data centers

Alphabet posted $109.9B Q1 2026 revenue with Cloud up 63% and a $460B backlog. Sundar Pichai said Google will sell TPUs to select customers running them in their own data centers.

Aerial view of a Meta data center site used by Fortune for AI infrastructure spending coverage
AI·

Hyperscalers are on track to spend $700B on AI infrastructure in 2026

Big-tech AI capex is projected at $700B in 2026, up from $410B in 2025. Microsoft alone guided $190B. Wall Street is split: Meta got punished for the spend, Alphabet rallied.

Title card for Boris Cherny's 'Mastering Claude Code in 30 Minutes' Anthropic workshop talk.
AI·

Anthropic just dropped its Claude Code workshop tapes. The playbook is better than the marketing.

Boris Cherny on Claude Code, Applied AI on prompting, Erik Schluntz on vibe coding in prod. Three Code with Claude tapes hit YouTube ahead of the 2026 conference.

AWS marketing illustration of an interconnected machine-learning workflow.
AI·

OpenAI's models are on AWS Bedrock the day after Microsoft lost exclusivity

Amazon shipped Bedrock Managed Agents powered by OpenAI on April 28, plus Codex on Bedrock. Altman tells Stratechery the runtime matters as much as the model.

Anthropic Claude generic brand graphic shown in promotional material for enterprise customers.
AI·

Disney built an AI leaderboard. One employee called Claude 460,000 times in nine days.

Leaked internal Disney screenshots show 4,800 product and tech staff burning 3.1 billion Claude tokens and 13.3 billion Cursor tokens across nine April workdays.

GitHub Octocat mark on a dark gradient, the cover graphic on the GitHub Blog post announcing the Copilot billing change.
AI·

GitHub Copilot kills premium requests on June 1. Token billing arrives, fallback models do not.

On June 1 every Copilot plan switches to GitHub AI Credits priced per token. Code completions stay free. Fallback models and credit rollover do not.

Microsoft and OpenAI logos paired on a navy gradient backdrop.
AI·

Microsoft and OpenAI just rewrote their deal. Exclusivity is dead, and so is the AGI clause.

Microsoft loses exclusive rights to OpenAI's models. The revenue share now caps at 2030 and stops depending on AGI. Here's what actually changed and who it benefits.

OpenAI just retired SWE-bench Verified. The headline coding benchmark of 2025 is officially saturated.
AI·

OpenAI just retired SWE-bench Verified. The headline coding benchmark of 2025 is officially saturated.

OpenAI says SWE-bench Verified is saturated and contaminated, and 60% of remaining problems are unsolvable. Here's what comes next, and why every coding leaderboard is suspect.

A padlock chained to a smartphone displaying a lock icon, illustrating data privacy.
AI·

OpenAI's Privacy Filter is a 1.5B PII redactor that ships under Apache 2.0. Here's what it actually does.

OpenAI released Privacy Filter on April 22 as an open-weight on-device model for masking eight types of PII. F1 of 96%. Runs in a browser. Here's the catch.

Illustration of an AI-driven chip design process from IEEE Spectrum's coverage.
AI·

An AI agent built a working RISC-V CPU from a 219-word prompt in 12 hours. Here's what it actually did.

Verkor's Design Conductor agent went from a 219-word spec to a tape-out-ready RISC-V core called VerCore in 12 hours. The catch: it's still a Celeron.

Anthropic Project Glasswing branding from Anthropic's news page.
AI·

A Discord group guessed Anthropic's URL pattern and walked into Claude Mythos

Bloomberg reports a small group accessed Anthropic's locked-down Mythos model the same day it launched, using credentials from a third-party contractor and educated URL guessing.

DeepSeek social card from the V4 API documentation release post.
AI·

DeepSeek V4 lands: 1.6T-param open MoE, 1M-token context, and SWE-bench within 0.2 of Opus 4.6

DeepSeek shipped V4-Pro and V4-Flash under MIT on April 24. V4-Pro hits 80.6% on SWE-bench Verified. V4-Flash is $0.14 in / $0.28 out.

Anthropic brand illustration used on the Anthropic newsroom.
AI·

Google is putting up to $40B into Anthropic. That's five days after Amazon's $5B.

Google committed $10B upfront and up to $40B total at a $350B valuation, plus five gigawatts of Google Cloud capacity. It's Anthropic's second nine-figure deal in a week.

Anthropic Engineering postmortem cover image.
AI·

Anthropic admits three Claude Code bugs quietly tanked quality for six weeks

Anthropic's April 23 postmortem names three bugs that degraded Claude Code between March 4 and April 20. Usage limits are being reset for every subscriber.

OpenAI's GPT-5.5 model launch with ChatGPT and Codex interfaces
AI·

OpenAI shipped GPT-5.5 seven weeks after 5.4. API tokens now cost twice as much.

OpenAI released GPT-5.5 (codename Spud) on April 23. The API runs at $5/$30 per million tokens, double GPT-5.4, with Pro at $30/$180.

OpenAI workspace agents launch graphic
AI·

OpenAI's Workspace Agents kill Custom GPTs and take the fight straight to Claude Code

Workspace Agents for ChatGPT Business, Enterprise, Edu, plus Teachers launched April 22. Team-shared, cloud-run, Codex-powered. Free until May 6, then credit-based.

World ID blog branding from Tools for Humanity, the Sam Altman-backed company behind the Orb
AI·

Worldcoin promoted a Bruno Mars partnership that doesn't exist. His team had to say so publicly.

Sam Altman's World pitched its 'Concert Kit' ticketing feature with a Bruno Mars partnership at its April 17 event. Live Nation told reporters no such deal exists.

GitHub Copilot announcement cover graphic
AI·

GitHub Copilot paused new signups and kicked Opus out of Pro. Here's what actually changed.

GitHub froze Copilot Pro/Pro+/Student signups on April 20 and moved Claude Opus 4.7 behind the $39 Pro+ tier. Agent workflows broke the old math.

Anthropic illustration from the Amazon compute deal announcement.
AI·

Amazon puts another $5B into Anthropic. Anthropic promises $100B back to AWS.

Amazon added $5B (up to $20B) to its Anthropic stake. Anthropic committed $100B+ to AWS over 10 years and 5 GW of Trainium capacity.

Illustration for Anthropic's Project Glasswing, a cybersecurity program powered by Claude Mythos Preview
AI·

NSA is running Anthropic's Mythos. The Pentagon says Anthropic is a supply-chain risk.

Axios reports the NSA is using Anthropic's unreleased Mythos model even though the Defense Department has blacklisted Anthropic. One government, two positions.

Google DeepMind Gemini Robotics-ER 1.6 architecture and demo illustration
AI·

Gemini Robotics-ER 1.6 is the brain, not the hands. Here's how the new two-model robotics stack works.

Google DeepMind's Gemini Robotics-ER 1.6 jumps instrument reading from 23% to 93% and slots in above a separate VLA model. Boston Dynamics' Spot is already running it.

Anthropic's Claude Design announcement illustration, a quill on a cactus-green background
AI·

Anthropic shipped Claude Design. Figma stock dropped 7% the same day.

Anthropic launched Claude Design on April 17, a prompt-to-prototype tool that exports to Canva, not Figma. Figma's stock closed down 7% on the same day.

The World Orb, a chrome sphere used to scan human irises for World ID verification
AI·

Sam Altman's Orb is now your Tinder badge. World goes mass-market.

World ID is rolling out to Tinder US, Zoom, Shopify, DocuSign, and Okta. Tools for Humanity is betting iris scans solve the bot problem.

Anysphere founder Michael Truell, the CEO behind the Cursor AI code editor
AI·

Cursor wants $50B for an AI editor that's burning cash on individuals

Cursor is in talks to raise $2B at a $50B valuation, nearly double its September mark. Revenue is up, but it's still losing money per indie seat.

Screenshot of the updated OpenAI Codex Mac app with background computer-use panel
AI·

OpenAI's Codex now drives your Mac, not just your code

OpenAI shipped a Codex update that can pilot desktop apps with a cursor, generate images in-line, and run parallel agents. It's the opening move in a real Claude Code fight.

Header card from Simon Willison's 'Qwen3.6 beats Opus' post comparing pelican SVGs
AI·

Qwen 3.6-35B-A3B: the open MoE beating Opus 4.7 on Simon Willison's laptop

Alibaba's Qwen 3.6-35B-A3B is a 35B-param mixture-of-experts with only 3B active. Apache 2.0, runs on consumer GPUs, and it's already winning real tasks.

Claude Opus 4.7 launch artwork from the Anthropic news post
AI·

Claude Opus 4.7 is here, and the long-context benchmarks got worse

Anthropic's Opus 4.7 is state-of-the-art on SWE-bench and CursorBench, but independent tests show regressions on long-context retrieval and thematic reasoning.

Google Gemini app running on a Mac desktop showing the mini chat interface
AI·

Google Gemini finally has a Mac app, and it's gunning for ChatGPT's desktop lead

Google shipped a native Swift Gemini app for macOS with screen sharing, voice, and Deep Research. Here's what it does, what it doesn't, and how it stacks up.

Abstract visualization of cybersecurity and AI defense systems
AI·

OpenAI launches GPT-5.4-Cyber for defensive security, opens access to thousands

OpenAI's new cybersecurity-tuned model can reverse-engineer binaries and analyze malware. It's restricted to verified defenders through the Trusted Access program.

Adobe Firefly expansion announcement social image from Adobe's blog
AI·

Adobe's Firefly AI Assistant can now drive Photoshop, Premiere, and Lightroom for you

Adobe renamed Project Moonlight to Firefly AI Assistant and opened a public beta. It runs multi-step workflows across Photoshop, Premiere, Lightroom, and more.

Claude wordmark on Anthropic's introducing-Routines announcement
AI·

Claude Code Routines: what they actually do, and when to use them over GitHub Actions

Anthropic just shipped Routines: Claude Code sessions as cron jobs, webhooks, and GitHub-event reactors. Here's what they replace, what they don't, and one rule to follow.