·
We publish letscodeit.dev blog posts from a separate GitHub repository. The Next.js app fetches markdown at runtime with a one-hour cache,…
GEO (generative engine optimization) is how you structure and publish content so AI chatbots and answer engines quote your pages when…
When you send a message to an LLM, you are not paying for characters or words. You pay for tokens . A token is the unit the model splits…
Qwen 3.8 27B, a new model with 27 billion parameters, is now available on Cerebras, achieving a speed of 1500 tokens per second.
This model operates with a context length of 64k for free users and 128k for paid users. Cerebras maintains high-quality model performance by using selective weight-only quantization for storage, while activations and other sensitive layers remain in full precision. No separate pruned models are hosted on public endpoints; all are original, unpruned versions.
Source: inference-docs.cerebras.ai
Impressive to see 27B params running at 1500 tokens/sec. Excited for the 128k context length option for paid users!
Selective quantization is a smart move. Keeping sensitive layers in full precision ensures quality without bloating storage.
Comments