llama_cu_awq

(β˜… 54)

llama INT4 cuda inference with AWQ

  • .gitattributes
  • CMakeLists.txt
  • common.h
  • convert_awq_to_bin.py
  • gpu_kernels.h
  • LICENSE
  • llama2_q4.cu
  • llama2_q4.vcxproj
  • perplexity.h
  • README.md
  • sampler.h
  • tokenizer.bin
  • tokenizer.h
  • weight_packer.cpp
// repository documentation