bare-metal.ai Articles

Generative AI on prem: secure, ethical, and accessible.

Advanced Weight-only Quantization Technique on CPU

image400

When LLMs started spreading at the end of 2022, it sounded something really impossible: training or even just fine-tune a model on your modest customer-grade hardware was fantasy.

Now, in the middle of 2024, thanks to an intensive work of scientific research, considerable investment, open governance, open collaboration, and a good dose of human ingenuity, we are now able to fine-tune models directly on our devices. Incredibile!
Continue Reading >>>

NeuralChat: deploy a local chatbot within minutes

013ca0a7-dd34-4e1a-ba8b-adf750ef3389

After showcasing Neural Speed in my past articles, my desire is to share a direct application of the theory: a tool developed using Neural Speed as very first brick, NeuralChat.

NeuralChat is highlighted as “A customizable framework to create your own LLM-driven AI apps within minutes”: it is available as part of the Intel® Extension for Transformers, a Transformer-based toolkit that makes possible to accelerate Generative AI/LLM inference both on CPU and GPU.
Continue Reading >>>

Neural Speed, Advanced Usage

74a82db1-be7b-4479-95f9-f0eddaa29683_1920x1080

After an initial presentation and a second follow up, here the third episode of my excursus about weight-only quantization, SignRound technique and their code implementation: the tensor parallelism library and inference engine, Neural Speed.

It is an amazing tool, and makes sense to explore a little more all the opportunities it offers through its multiple options.
Continue Reading >>>

Notice
This website does not track you. There are no analytics, no advertising, and no profiling cookies. The only things stored on your device are a session cookie used by the contact form, and a note that you dismissed this message. Web fonts are loaded from Google, which receives your IP address. Full detail is in the cookie policy and the privacy policy. There is nothing here to consent to. Dismiss simply closes this notice.

baremetalai on Substack

Read on Substack