Product signal

Qwen opens weights for its 2.4T-parameter Max model

Qwen has released weights for Qwen3.8-2.4T-A95B, moving Qwen3.8-Max from limited hosted-preview channels to downloadable, self-deployable distribution.

Qwen has moved its largest Qwen3.8-Max model from restricted hosted preview to open-weight distribution. The official repository now provides Qwen3.8-2.4T-A95B for download, while Alibaba had previously limited the 2.4-trillion-parameter Qwen3.8-Max-Preview to Token Plan, Qoder, and QoderWork. The change expands model competition from API access to deployment control and infrastructure ownership.

Qwen's largest model enters open distribution

Qwen's official repository provides public weights and a model card for Qwen3.8-2.4T-A95B, confirming that the specific model can be downloaded and documenting its architecture and use limitations. NVIDIA identifies it as Qwen3.8-Max, a 2.4-trillion-parameter model, and describes configurable-reasoning deployment on GB300 NVL72. Previously, users could access the model through hosted preview channels; organizations with sufficient hardware and engineering capacity can now run and evaluate it themselves.

Open weights change who controls deployment

Alibaba said in July that Qwen3.8-Max-Preview was limited to Token Plan, Qoder, and QoderWork, while signaling that Max weights would be opened. Hosted access lets a model provider control pricing, queues, regions, and interfaces. Open weights let enterprises and cloud providers choose their inference stack, quantization strategy, data boundaries, and resource scheduling. For a model at this scale, that does not inherently mean cheaper deployment; it changes the cost mix from usage fees toward hardware, operations, and throughput optimization.

High-end infrastructure remains the practical constraint

NVIDIA's GB300 NVL72 deployment path indicates that cluster capacity and systems optimization remain the key constraints on running a model of this size. Open distribution broadens availability without automatically broadening affordability. Its more immediate industrial significance is that cloud providers, sovereign-compute projects, and enterprises with dense GPU clusters can build differentiated services around the same model rather than relying entirely on a vendor-hosted endpoint.

What to watch next

Watch for Qwen's formal benchmarks, recommended hardware configurations, and fuller production-deployment documentation; for managed offerings from major cloud platforms; and for independent reports on throughput, cost, and reliability across quantization methods, cluster sizes, and workloads. Those observations will test whether open weights translate into broad adoption.

Sources