sorry-bench
Benchmark evaluation code for "SORRY-Bench: Systematically Evaluating Large Language Model Safety Refusal" (ICLR 2025)
File Explorer
Download Latest Version (.zip)- .dummy
- .dummy
- .dummy
- decode.py
- encode_experts.py
- mutate.py
- mutation_utils.py
- prompts_and_demonstrations.py
- README.md
- judge_prompts.jsonl
- benchmark-results.png
- meta-eval-demo-hf-original.jpg
- sorry-bench-logo-circle.png
- sorry-bench-mutation-demo.png
- sorry-bench-taxonomy-202503.png
- common.py
- gen_api_answer.py
- gen_judgment_safety.py
- gen_judgment_safety_vllm.py
- gen_model_answer.py
- gen_model_answer_vllm.py
- LICENSE
- README.md
- visualize_result.ipynb
// repository documentation
Was this content helpful?
(0 ratings)
