Hardware·
One model per chip: AMD bought Taalas, whose test part hit 17,000 tokens a second
Taalas casts a model's weights into transistors. Its 6nm test chip served Llama 3.1 8B at roughly 17,000 tokens per second, and AMD just bought the company.
Taalas casts a model's weights into transistors. Its 6nm test chip served Llama 3.1 8B at roughly 17,000 tokens per second, and AMD just bought the company.

Arcee released Trinity-Large-Thinking on April 1: a 399B-param sparse MoE with 13B active, Apache 2.0 weights, $0.88 per million output tokens, and PinchBench just behind Opus 4.6.