SOTA low-bit LLM quantization (INT8/FP8/MXFP8/INT4/MXFP4/NVFP4) & sparsity; leading model compression techniques on PyTorch, TensorFlow, and ONNX Runtime
-
Updated
Sep 5, 2026 - Python
SOTA low-bit LLM quantization (INT8/FP8/MXFP8/INT4/MXFP4/NVFP4) & sparsity; leading model compression techniques on PyTorch, TensorFlow, and ONNX Runtime
LLM model quantization (compression) toolkit with HW acceleration support for Nvidia, AMD, Intel GPU and Intel/AMD/Apple CPU via HF, vLLM, and SGLang.
🦖 X—LLM: Cutting Edge & Easy LLM Finetuning
Run any Large Language Model behind a unified API
🪶 Lightweight OpenAI drop-in replacement for Kubernetes
A guide about how to use GPTQ models with langchain
Run gguf LLM models in Latest Version TextGen-webui and koboldcpp
Private self-improvement coaching with open-source LLMs
ChatSakura:Open-source multilingual conversational model.(开源多语言对话大模型)
This repository is for profiling, extracting, visualizing and reusing generative AI weights to hopefully build more accurate AI models and audit/scan weights at rest to identify knowledge domains for risk(s).
A.L.I.C.E (Artificial Labile Intelligence Cybernated Existence). A REST API of A.I companion for creating more complex system
Compress Any LLM Up to 6x in One Command. Unified CLI for GGUF, GPTQ, and AWQ quantization.
🎯 Fine-tune large language models and use them for text-related tasks. This repository provides a straightforward approach to fine-tuning models like Gemma, Llama 🦙, and Mistral 🌪️ for various NLP tasks. 🔧 It includes training 📚, fine-tuning 🛠️, and inference pipelines ⚙️. 🚀
本来叫 nano 的,后来发现装不下 Qwen3.5,就改名叫 big 了
Conversation AI model for open domain dialogs
Edge-device lion dance robot with MSPM0G3507 firmware, open hardware, and QWEN2_0_5B 4-bit GPTQ edge AI bridge.
To associate your repository with the gptq topic, visit your repo's landing page and select "manage topics."