Quantization
My first book is out! 🤓
September 09. 2026 • Category: Announcement

I have been super-silent for a long time on these pages, but for a good reason! My first book is out!
During last year I was super focused on writing a book (my first) about the topic of quantization of language models.
Everything started just by chance, I was doing some research on ranking weight-only-quantization methods to write down a new series of articles for this blog, while the amount of notes, documentation and scientific papers become so huge that I had enough material that I said “Oh Wait. This could actually be a book!”. I had wanted to write one for as long as I can remember: I wrote it basically first of all for myself, a huge achievement both for me and my professional growth.
So, this is the genesis of "Weight-Only Quantization of Language Models: A Deep-Dive Analysis of Methods for Converting and Compressing Small and Large Language Models" is available in three editions: Kindle ebook, paperback and hardcover, as Amazon Exclusive, at Amazon EU and Amazon USA, or on other Amazon websites, for a worldwide distribution.
Continue Reading >>>
Igniting 2025 with tons of INT4 Quantizations!
January 01. 2025 • Category: Announcement

As we just ignited 2025, and 2024 came to an end, I am proud to share that I have successfully uploaded over 230 quantized SLM/LLM models to my HuggingFace account. These models were entirely quantized using the computational resources of my homelab, achieving approximately 72 TFLOPS of performance-powered solely by "domestic" hardware.
Continue Reading >>>
Advanced Weight-only Quantization Technique on CPU
May 05. 2024 • Category: Framework

When LLMs started spreading at the end of 2022, it sounded something really impossible: training or even just fine-tune a model on your modest customer-grade hardware was fantasy.
Now, in the middle of 2024, thanks to an intensive work of scientific research, considerable investment, open governance, open collaboration, and a good dose of human ingenuity, we are now able to fine-tune models directly on our devices. Incredibile!
Continue Reading >>>

