devtake.dev

#local-inference

RSS
The Hugging Face model page card for deepseek-ai/DeepSeek-V4-Flash-0731, showing the DeepSeek whale logo and the repository name
AI·

DeepSeek's new 304B agentic model now runs on a single 128GB workstation

Salvatore Sanfilippo repacked DeepSeek V4 Flash into a lossless MXFP4 GGUF that streams from SSD at over 20 tokens a second. The hardware bill, and where hosted still wins.

A MacBook Pro beside a Surface Book, both open on a white surface, USB-C ports in view
AI·

Running a coding agent fully on Apple Silicon, no cloud, is now an off-the-shelf stack

A popular Hacker News how-to walked through a fully local coding agent on Apple Silicon. Here's the realistic 2026 stack: runner, model, and harness.

Cyera Research disclosure illustration for the Bleeding Llama vulnerability in Ollama's model execution pipeline
Security·

A crafted Ollama model file leaks the whole server's memory. 300,000 instances are exposed.

Cyera disclosed CVE-2026-7482 on May 1, a CVSS 9.1 unauthenticated heap read in Ollama. Three API calls dump prompts, env vars, and API keys from any open instance.

Render of an AMD Ryzen AI Max+ 495 'Gorgon Halo' APU package, surrounded by labels for the Radeon 8065S integrated GPU and the chip's memory configuration.
Hardware·

AMD's 'Gorgon Halo' refresh leaks with 192GB memory. Strix Halo tops out at 128GB.

A leaked Geekbench listing puts AMD's Ryzen AI Max+ 495 on a 192GB platform with a Radeon 8065S iGPU. The Strix Halo chip it replaces capped at 128GB.

Apple Mac mini desktop computer on a clean background, showing the M4-era chassis.
Apple·

Apple killed the $599 Mac mini. The cheapest one is now $799 with 512GB.

Apple quietly pulled the 256GB Mac mini from its store on May 1. Tim Cook had warned the day before that demand was outpacing supply for months.

Header card from Simon Willison's 'Qwen3.6 beats Opus' post comparing pelican SVGs
AI·

Qwen 3.6-35B-A3B: the open MoE beating Opus 4.7 on Simon Willison's laptop

Alibaba's Qwen 3.6-35B-A3B is a 35B-param mixture-of-experts with only 3B active. Apache 2.0, runs on consumer GPUs, and it's already winning real tasks.