UL-SMF-Cache-Compression
The Unified Latent-State Memory Fabric (UL-SMF) is a hardware-software co-designed memory compression fabric that solves the memory bottleneck in long-context Transformer inference. By combining FSQ with dynamic 16-dimensional latent mapping, UL-SMF compresses Key-Value (KV) cache tensors by up to 384x while maintaining >94% semantic retention.
File Explorer
- codeql.yml
- dependency-review.yml
- python-package.yml
- python-publish.yml
- .gitignore
- INTEGRATION.md
- LICENSE
- README.md
- setup.py
- ul_smf.py
- UL_SMF_Interceptor.ipynb
# Use via CDN
jsDelivrjsDelivr serves any public GitHub repository as a CDN with zero setup. Pick a version and a file to get a ready-to-paste link and snippet.
Link
Example
// repository documentation
Was this content helpful?
(0 ratings)
