Watershed moment: I picked up a Strix Halo machine in mid-2025 with 128GB of RAM, shortly before the shortage hit. I've been running toy models on it to start, like Mistral 8B, then moved on to GPT OSS 20B/120B, and finally Qwen 3.6 35B-A3B/27B. But today, for the first, time, with the release of DeepSeek v4 Flash 0731, I'm able to run a 2-bit quant (Unsloth Dynamic) on the rig with a llama.cpp backend, running in the fantastic omp harness.
This is by far the best local LLM experience I've ever had...it's a touch slow (15t/s), but the tool-calling is top notch, and the thinking/planning seems roughly on par with the hosted version (though I know I'm taking the 2-bit hit).
Extremely excited to see if the results hold up over time. I can imagine selling an appliance running this with Hermes.