llama_cu_awq
llama INT4 cuda inference with AWQ
File Explorer
Download Latest Version (.zip)- common.h
- gpu_kernels.h
- LICENSE
- llama2_q4.sycl.cpp
- Makefile
- perplexity.h
- README.md
- sampler.h
- tokenizer.h
- .gitattributes
- CMakeLists.txt
- common.h
- convert_awq_to_bin.py
- gpu_kernels.h
- LICENSE
- llama2_q4.cu
- llama2_q4.vcxproj
- perplexity.h
- README.md
- sampler.h
- tokenizer.bin
- tokenizer.h
- weight_packer.cpp
// repository documentation
Was this content helpful?
(0 ratings)
