Qwen’s 125-billion-parameter preview activates only six billion at a time
This is a mixture-of-experts model: it has 125 billion parameters in total, while each token uses a much smaller active subset. That distinction helps explain the efficiency claim, but it does not turn the model into a six-billion-parameter download.
Simon tried quantized versions on a DGX Spark. The two files he tested were still substantial: 72.5GB and 78.9GB. His initial experiments used the familiar pelican-on-a-bicycle SVG prompt, with results shared alongside the note.
That makes this an early, inspectable local-model experiment. The drawings show what those particular settings produced; a wider judgment about coding, reasoning, or reliability would require more varied tests.