Qwen 3.8 27B available on Cerebras at 1500 tok/SEC

Sep 04, 2026 01:32 AM - 1 hour ago 1

Models connected Cerebras nationalist endpoints are disposable connected the free proceedings and pay-as-you-go tiers, taxable to rate limits and pricing. For further exemplary families, reserved capacity, higher throughput, and accumulation SLAs, spot Dedicated Endpoints.

Available Models

Model Compression

This conception provides transparency astir the compression authorities of each exemplary disposable connected our platform. We big a assortment of open-source models from the community. We do not presently big pruned models connected our nationalist endpoints. All models served done our nationalist endpoints are the original, unpruned versions. While we behaviour investigation connected pruning techniques for illustration REAP (Router-weighted Expert Activation Pruning), these pruned models are shared pinch the investigation organization connected Hugging Face but are not disposable done our shared API. You tin publication much astir REAP successful our research blog. All of our nationalist models are unpruned. Cerebras uses selective weight-only quantization only during retention to sphere maximal quality. This intends that the weights are stored successful partial 16-bit / 8-bit / 4-bit, in-line pinch manufacture standards. For quality, delicate layers are stored astatine afloat precision pinch dequantization connected the fly, truthful operations are done successful precocious precision. The activations, attention, and kv cache stay successful afloat precision and unquantized.

Frequently Asked Questions

Will you alteration a model's architecture without notice?

No. We are committed to serving the original models for each existing endpoints, without modification. We do not change exemplary architectures via pruning connected our hosted portfolio. If we research further compression techniques (like pruning) successful the future, these would beryllium offered arsenic abstracted endpoints pinch pruning-specific names, ensuring complete transparency and allowing you to take which type champion fits your needs.

Where tin I find your REAP pruned models?

Our REAP pruned models are disposable connected Hugging Face for investigation and experimentation purposes: Cerebras REAP Collection. These models show our pruning investigation but are not served done our accumulation API.

What are compression, quantization, and pruning?

Compression is an umbrella word for techniques that trim exemplary size aliases computational requirements. Common compression techniques include:

  • Quantization: Reducing the precision of numbers utilized to correspond exemplary weights (e.g., converting from FP16 to FP8). This reduces representation usage without changing the model’s architecture.
  • Pruning: Permanently removing parts of a model, for illustration layers aliases experts, to trim exemplary size. This changes the model’s architecture and creates a different model.
More