lfm25-inference

(★ 9)

A Rust + CUDA inference runtime for LFM2.5-1.2B-Instruct, focused on low-latency GPU inference, explicit memory management, paged KV caching, and measured kernel/runtime optimization.

File Explorer

  • .gitignore
  • build.rs
  • Cargo.lock
  • Cargo.toml
  • README.md

# Use via CDN

jsDelivr

jsDelivr serves any public GitHub repository as a CDN with zero setup. Pick a version and a file to get a ready-to-paste link and snippet.

// repository documentation