Skip to content
#

model-quantization

Here are 84 public repositories matching this topic...

A curated collection of papers, benchmarks, surveys, and tools for model quantization, covering low-bit networks, LLMs, multimodal and generative models, vector and lattice quantization, and efficient deployment.

  • Updated Sep 28, 2026

离线可用的本地类型化决策:4 核 CPU 单题 15.6ms。Local & offline Jev / System One inference on CPU — ONNX + INT8, no torch at runtime. 支持 laya / kev / PlayJev

  • Updated Sep 21, 2026
  • Python

AI Engineering: Annotated NBs to dive into Self-Attention, In-Context Learning, RAG, Knowledge-Graphs, Fine-Tuning, Model Optimization, and many more.

  • Updated Apr 2, 2025
  • Jupyter Notebook

H.E.R.A. (Healing Evaluation and Recognition Architecture) — an on-device edge-AI system for offline wound assessment using YOLOv8 segmentation, HSV tissue analysis, and PUSH-based decision logic.

  • Updated Aug 30, 2026
  • Dart

🧠 A comprehensive toolkit for benchmarking, optimizing, and deploying local Large Language Models. Includes performance testing tools, optimized configurations for CPU/GPU/hybrid setups, and detailed guides to maximize LLM performance on your hardware.

  • Updated Mar 27, 2025
  • Shell

Add this topic to your repo

To associate your repository with the model-quantization topic, visit your repo's landing page and select "manage topics."

Learn more