AIbase
Product LibraryTool NavigationMCP

KVSplit

Public

Run larger LLMs with longer contexts on Apple Silicon by using differentiated precision for KV cache quantization. KVSplit enables 8-bit keys & 4-bit values, reducing memory by 59% with <1% quality loss. Includes benchmarking, visualization, and one-command setup. Optimized for M1/M2/M3 Macs with Metal support.

Creat2025-05-17T02:45:59
Update2025-06-07T09:13:53
353
Stars
1
Stars Increase

Related projects