fastrl
[ASPLOS'26] Taming the Long-Tail: Efficient Reasoning RL Training with Adaptive Drafter
File Explorer
Download Latest Version (.zip)Showing a partial file list โ download the zip above to see everything.
- demo-first-frame.png
- Eurus_sample.json
- datagen_config.yaml
- deepspeed_config.json
- Eagle-DS-Qwen-7B.json
- Eagle-Llama-3-8B-Instruct.json
- Eagle-Llama-3.1-8B-Instruct.json
- Eagle-Llama-3.3-70B-Instruct.json
- Eagle-Qwen2.5-32B.json
- create_mixed_dataset.py
- generate_freq.py
- llama_eagle.py
- llama_eagle3.py
- qwen2_eagle.py
- qwen2_eagle3.py
- qwen3_eagle3.py
- datagen_eagle2.sh
- train_eagle2.sh
- train_eagle3.sh
- .gitignore
- __init__.py
- eagle3_trainer.py
- eagle_datagen.py
- eagle_trainer.py
- README.md
- requirements.txt
- utils.py
- bench_sd.sh
- grpo_32B_multi_nodes.sh
- grpo_7B.sh
- bench_speculative_decoding.py
- devcontainer.json
- Dockerfile
- 1-bug-report.yml
- 2-feature-request.yml
- bot-bump-kernel-version.yml
- bot-bump-sglang-version.yml
- cancel-all-pending-pr-test-runs.yml
- cancel-pr-workflow-on-merge.yml
- ci-monitor.yml
- close-inactive-issues.yml
- execute-notebook.yml
- label-pr.yml
- lint.yml
- nightly-release-router.yml
- nightly-test-amd.yml
- nightly-test.yml
- open-pr-copy-from-oss.yml
- open-pr-copy-to-oss.yml
- pr-benchmark-rust.yml
- pr-test-amd.yml
- pr-test-npu.yml
- pr-test-pd-router.yml
- pr-test-rust.yml
- pr-test-xeon.yml
- pr-test-xpu.yml
- pr-test.yml
- release-docker-amd-nightly.yml
- release-docker-amd.yml
- release-docker-dev.yml
- release-docker-npu-nightly.yml
- release-docker-npu.yml
- release-docker-router.yml
- release-docker-xeon.yml
- release-docker.yml
- release-docs.yml
- release-fake-tag.yml
- release-pypi-router.yml
- release-pypi.yml
- release-whl-kernel.yml
- vllm-dependency-test.yml
- CODEOWNERS
- pull_request_template.md
- REVIEWERS.md
- client.sh
- install_rpd.sh
- loadTracer.sh
- PROFILING.md
- rpd.patch
- rpd_profile_server_enable.patch
- rpd_profile_server_enable_wCPU_activities.patch
- server.sh
- torch_profiler.patch
- benchmark_moe_rocm.py
- TUNING.md
- logo.png
- logo.svg
- logo_square.svg
- bench_attention_sink_triton.py
- bench_in_batch_prefix.py
- benchmark_batch.py
- benchmark_tokenizer.py
- README.md
- 405b_sglang.sh
- 405b_trt.sh
- 405b_vllm.sh
- config.md
- README.md
- bench_sglang.py
- convert_parquet_to_json.py
- parquet_to_json.sh
- README.md
- bench_sglang.py
- README.md
- README.md
- bench_dspy_intro.py
- README.md
- agent_functions.py
- bench_other.py
- bench_sglang.py
- README.md
- README.md
- bench_other.py
- bench_sglang.py
- README.md
- bench_other.py
- bench_sglang.py
- README.md
- bench.sh
- bench_client.py
- bench_storage.py
- bench_zerocopy.py
- bench_long_context.py
- bench_mix.py
- bench_mix.sh
- bench_multiturn.py
- bench_serving.py
- data_processing.py
- download.sh
- nextqa.py
- README.md
- bench_other.py
- bench_sglang.py
- build_dataset.py
- README.md
- bench_other.py
- bench_sglang.py
- build_dataset.py
- dataset.txt
- README.md
- bench_sglang.py
- README.md
- benchmark_mscclpp.py
- benchmark_symm_mem.py
- triton_flashinfer_cudnn.py
- deepep_utils.py
- tuning_deepep.py
- benchmark_deepgemm_fp8_gemm.py
- benchmark_deepgemm_fp8_group_gemm.py
- README.md
- benchmark_concat_mla.py
- benchmark_fused_collective.py
- README.md
- benchmark_sglang_fused_moe_triton.py
- benchmark_torch_compile_fused_moe.py
- benchmark_vllm_vs_sglang_fused_moe_triton.py
- README.md
- tuning_fused_moe_triton.py
- benchmark_lightning_attention_decode.py
- benchmark_lightning_attention_prefill.py
- bench_fp4_quant.py
- bench_int8_quant.py
- tuning_block_wise_kernel.py
- benchmark_get_last_loc_triton.py
- benchmark_write_req_to_token_pool_triton.py
- bench_triton_swa_kernel.py
- bench_sglang.py
- gen_data.py
- README.md
- bench_hf_llava_bench.sh
- bench_hf_mme.sh
- bench_sglang.py
- bench_sglang_mme.sh
- download_images.py
- README.md
- bench_other.py
- bench_sglang.py
- README.md
- bench_other.py
- bench_sglang.py
- build_dataset.py
- README.md
- launch_server.py
- lora_bench.py
- bench_other.py
- bench_sglang.py
- download_data.sh
- README.md
- bench_hf.py
- bench_sglang.py
- data_utils.py
- eval_utils.py
- internvl_utils.py
- prompt_format.yaml
- README.md
- bench_other.py
- bench_sglang.py
- bench_sglang_eagle.py
- README.md
- bench_other.py
- bench_sglang.py
- README.md
- bench_other.py
- bench_sglang.py
- build_dataset.py
- README.md
- bench_other.py
- bench_sglang.py
- data_gen.py
- long_prompt_multi_turn.py
- README.md
- bench_embeddings.py
- bench_score.py
- util.py
- bench_other.py
- bench_sglang.py
- README.md
- answer_extraction.py
- bench_sglang.py
- eval_utils.py
- README.md
- .gitignore
- bench_other.py
- bench_sglang.py
- lmql_funcs.py
- README.md
- topic.jsonl
- bench_other.py
- bench_sglang.py
- lmql_funcs.py
- README.md
- bench_other.py
- bench_sglang.py
- README.md
- compose.yaml
- Dockerfile
- Dockerfile.b300
- Dockerfile.npu
- Dockerfile.rocm
- Dockerfile.router
- Dockerfile.sagemaker
- Dockerfile.xeon
- Dockerfile.xpu
- k8s-sglang-distributed-sts.yaml
- k8s-sglang-service.yaml
- serve
- custom_log.css
- readthedocs.css
- logo.ico
- logo.png
- attention_backend.md
- hicache.rst
- hicache_best_practices.md
- hicache_design.md
- hyperparameter_tuning.md
- lora.ipynb
- observability.md
- pd_disaggregation.md
- pd_multiplexing.md
- quantization.md
- router.md
- separate_reasoning.ipynb
- server_arguments.md
- speculative_decoding.ipynb
- structured_outputs.ipynb
- structured_outputs_for_reasoning_models.ipynb
- tool_parser.ipynb
- vlm_query.ipynb
- deepseek.md
- gpt_oss.md
- llama4.md
- native_api.ipynb
- offline_engine_api.ipynb
- openai_api.rst
- openai_api_completions.ipynb
- openai_api_embeddings.ipynb
- openai_api_vision.ipynb
- qwen3.md
- sampling_params.md
- send_request.ipynb
- bench_serving.md
- benchmark_and_profiling.md
- contribution_guide.md
- development_guide_using_docker.md
- release_process.md
- setup_github_runner.md
- install.md
- amd_gpu.md
- ascend_npu.md
- cpu_server.md
- nvidia_jetson.md
- tpu.md
- choices_methods.md
- frontend_index.rst
- frontend_tutorial.ipynb
- d-svc.yaml
- d.yaml
- lb.yaml
- p-svc.yaml
- p.yaml
- lws_pd_deploy.md
- deploy_on_k8s.md
- multi_node.md
- multi_node_index.rst
- custom_chat_template.md
- environment_variables.md
- faq.md
- learn_more.md
- production_metrics.md
- production_request_trace.md
- torch_compile_cache.md
- embedding_models.md
- generative_models.md
- modelscope.md
- multimodal_language_models.md
- rerank_models.md
- reward_models.md
- support_new_models.md
- transformers_fallback.md
- conf.py
- deploy.py
- index.rst
- Makefile
- README.md
- requirements.txt
- serve.sh
- wrap_run_llm.py
- tool_chat_template_deepseekr1.jinja
- tool_chat_template_deepseekv3.jinja
- tool_chat_template_deepseekv31.jinja
- tool_chat_template_deepseekv32.jinja
- tool_chat_template_llama4_pythonic.jinja
- vision_template_sarashina_vl.jinja
- cat.jpeg
- dog.jpeg
- anthropic_example_chat.py
- anthropic_example_complete.py
- azure_openai_example_chat.py
- gemini_example_chat.py
- gemini_example_complete.py
- gemini_example_multimodal_chat.py
- local_example_chat.py
- local_example_complete.py
- local_example_llava_next.py
- openai_example_chat.py
- openai_example_complete.py
- openai_example_n.py
- openai_example_o1.py
- openrouter_example_chat.py
- together_example_chat.py
- together_example_complete.py
- srt_example_llava_v.py
- srt_example_llava_v.sh
- trace_and_evaluate_rag_using_parea.ipynb
- config.pbtxt
- Dockerfile
- README.md
- chinese_regex.py
- choices_logprob.py
- cot_decoding.py
- json_decode.py
- json_logprobs.py
- openai_chat_speculative.py
- openai_speculative.py
- parallel_sample.py
- readme_examples.py
- sgl_gen_min_tokens.py
- streaming.py
- dashboard.yaml
- sglang-dashboard.json
- datasource.yaml
- docker-compose.yaml
- opentelemetry.yaml
- prometheus.yaml
- README.md
- tracing_compose.yaml
- gputrc2graph.py
- README.md
- sglang_engine_model.json
- custom_server.py
- embedding.py
- fastapi_engine_inference.py
- launch_engine.py
- offline_batch_inference.py
- offline_batch_inference_async.py
- offline_batch_inference_eagle.py
- offline_batch_inference_qwen_1m.py
- offline_batch_inference_vlm.py
- readme.md
- save_remote_state.py
- save_sharded_state.py
- hidden_states_engine.py
- hidden_states_server.py
- llama3_llava_server.py
- llava_onevision_server.py
- pixtral_server.py
- qwen_llava_server.py
- token_in_token_out_llm_engine.py
- token_in_token_out_llm_server.py
- token_in_token_out_vlm_engine.py
- token_in_token_out_vlm_server.py
- lora.py
- multimodal_embedding.py
- openai_chat_with_response_prefill.py
- README.md
- reward_model.py
- vertex_predict.py
- llama3_eval.py
- loogle_eval.py
- anthropic.py
- base_backend.py
- litellm.py
- openai.py
- runtime_endpoint.py
- vertexai.py
- api.py
- chat_template.py
- choices.py
- compiler.py
- interpreter.py
- ir.py
- tracer.py
- __init__.py
- batch_invariant_ops.py
- backend.py
- compilation_config.py
- compilation_counter.py
- compile.py
- compiler_interface.py
- cuda_piecewise_backend.py
- fix_functionalization.py
- fx_utils.py
- inductor_pass.py
- pass_manager.py
- piecewise_context_manager.py
- weak_ref_tensor.cpp
- weak_ref_tensor_jit.py
- __init__.py
- chatglm.py
- dbrx.py
- deepseekvl2.py
- device_config.py
- dots_ocr.py
- dots_vlm.py
- exaone.py
- falcon_h1.py
- internvl.py
- janus_pro.py
- kimi_vl.py
- kimi_vl_moonvit.py
- load_config.py
- longcat_flash.py
- mamba_utils.py
- model_config.py
- nemotron_h.py
- qwen3_next.py
- qwen3_vl.py
- step3_vl.py
- update_config.py
- utils.py
- __init__.py
- safe_serde.py
- serde.py
- __init__.py
- base_connector.py
- redis.py
- remote_instance.py
- s3.py
- utils.py
- bitmask_ops.py
- base_grammar_backend.py
- llguidance_backend.py
- outlines_backend.py
- outlines_jump_forward.py
- reasoner_grammar_backend.py
- xgrammar_backend.py
- __init__.py
- dump_comparator.py
- dump_loader.py
- dumper.py
- text_comparator.py
- __init__.py
- conn.py
- transfer_engine.py
- __init__.py
- conn.py
- __init__.py
- conn.py
- utils.py
- __init__.py
- conn.py
- __init__.py
- conn.py
- transfer_engine.py
- __init__.py
- conn.py
- decode.py
- decode_kvcache_offload_manager.py
- decode_schedule_batch_mixin.py
- kv_events.py
- mini_lb.py
- prefill.py
- utils.py
- all_reduce_utils.py
- cuda_wrapper.py
- custom_all_reduce.py
- custom_all_reduce_utils.py
- hpu_communicator.py
- npu_communicator.py
- pymscclpp.py
- pynccl.py
- pynccl_allocator.py
- pynccl_wrapper.py
- quick_all_reduce.py
- shm_broadcast.py
- symm_mem.py
- xpu_communicator.py
- __init__.py
- communication_op.py
- naive_distributed.py
- parallel_state.py
- utils.py
- __init__.py
- protocol.py
- serving_base.py
- serving_chat.py
- serving_completions.py
- serving_embedding.py
- serving_rerank.py
- serving_responses.py
- serving_score.py
- serving_tokenize.py
- tool_server.py
- usage_processor.py
- utils.py
- context.py
- engine.py
- EngineBase.py
- grpc_server.py
- harmony_utils.py
- http_server.py
- http_server_engine.py
- tool.py
- __init__.py
- deepseek.py
- deepseek_vec.py
- __init__.py
- reader.py
- __init__.py
- eplb_manager.py
- expert_distribution.py
- expert_location.py
- expert_location_dispatch.py
- expert_location_updater.py
- base_format_detector.py
- core_types.py
- deepseekv31_detector.py
- deepseekv3_detector.py
- ebnf_composer.py
- function_call_parser.py
- glm4_moe_detector.py
- gpt_oss_detector.py
- json_array_parser.py
- kimik2_detector.py
- llama32_detector.py
- mistral_detector.py
- pythonic_detector.py
- qwen25_detector.py
- qwen3_coder_detector.py
- step3_detector.py
- utils.py
- __init__.py
- compile_proto.py
- grpc_request_manager.py
- sglang_scheduler.proto
- sglang_scheduler_pb2.py
- sglang_scheduler_pb2.pyi
- sglang_scheduler_pb2_grpc.py
- chunk.py
- chunk_delta_h.py
- chunk_o.py
- chunk_scaled_dot_kkt.py
- cumsum.py
- fused_recurrent.py
- fused_sigmoid_gating_recurrent.py
- index.py
- l2norm.py
- layernorm_gated.py
- op.py
- solve_tril.py
- utils.py
- wy_fast.py
- __init__.py
- layernorm_gated.py
- mamba_ssm.py
- ssd_bmm.py
- ssd_chunk_scan.py
- ssd_chunk_state.py
- ssd_combined.py
- ssd_state_passing.py
- causal_conv1d.py
- causal_conv1d_triton.py
- mamba.py
- mamba2_metadata.py
- mixer2_rms_norm_gated.py
- mla_preprocess.py
- dequant_k_cache.py
- index_buf_accessor.py
- nsa_indexer.py
- quant_k_cache.py
- tilelang_kernel.py
- transform_index.py
- triton_kernel.py
- utils.py
- decode_attention.py
- double_sparsity_attention.py
- extend_attention.py
- merge_state.py
- prefill_attention.py
- rocm_mla_decode_rope.py
- decode_attention.py
- extend_attention.py
- prefill_attention.py
- aiter_backend.py
- ascend_backend.py
- attention_registry.py
- base_attn_backend.py
- cutlass_mla_backend.py
- double_sparsity_backend.py
- dual_chunk_flashattention_backend.py
- flashattention_backend.py
- flashinfer_backend.py
- flashinfer_mla_backend.py
- flashmla_backend.py
- hybrid_attn_backend.py
- hybrid_linear_attn_backend.py
- intel_amx_backend.py
- merge_state.py
- nsa_backend.py
- tbo_backend.py
- torch_flex_backend.py
- torch_native_backend.py
- triton_backend.py
- trtllm_mha_backend.py
- trtllm_mla_backend.py
- utils.py
- vision.py
- vision_utils.py
- wave_backend.py
- __init__.py
- kernels.py
- layer.py
- E=1,N=14336,device_name=NVIDIA_A100-SXM4-80GB,dtype=int8_w8a16.json
- E=1,N=14336,device_name=NVIDIA_A100-SXM4-80GB.json
- E=1,N=1792,device_name=NVIDIA_A100-SXM4-80GB,dtype=int8_w8a16.json
- E=1,N=1792,device_name=NVIDIA_A100-SXM4-80GB.json
- E=1,N=3072,device_name=NVIDIA_A100-SXM4-80GB,dtype=int8_w8a16.json
- E=1,N=3072,device_name=NVIDIA_H100_80GB_HBM3,dtype=int8_w8a16.json
- E=1,N=3072,device_name=NVIDIA_H100_80GB_HBM3.json
- E=1,N=3584,device_name=NVIDIA_A100-SXM4-80GB,dtype=int8_w8a16.json
- E=1,N=3584,device_name=NVIDIA_A100-SXM4-80GB.json
- E=1,N=7168,device_name=NVIDIA_A100-SXM4-80GB,dtype=int8_w8a16.json
- E=1,N=7168,device_name=NVIDIA_A100-SXM4-80GB.json
- E=144,N=512,device_name=NVIDIA_H100_80GB_HBM3.json
- E=16,N=1024,device_name=NVIDIA_H100_80GB_HBM3.json
- E=16,N=1024,device_name=NVIDIA_H200.json
- E=16,N=1344,device_name=NVIDIA_A100-SXM4-40GB.json
- E=16,N=1344,device_name=NVIDIA_A100-SXM4-80GB.json
- E=16,N=1344,device_name=NVIDIA_H100_80GB_HBM3.json
- E=16,N=14336,device_name=NVIDIA_A100-SXM4-80GB,dtype=int8_w8a16.json
- E=16,N=14336,device_name=NVIDIA_A100-SXM4-80GB.json
- E=16,N=1792,device_name=NVIDIA_A100-SXM4-80GB,dtype=int8_w8a16.json
- E=16,N=1792,device_name=NVIDIA_A100-SXM4-80GB.json
- E=16,N=2048,device_name=NVIDIA_H100_80GB_HBM3.json
- E=16,N=2688,device_name=NVIDIA_A100-SXM4-80GB.json
- E=16,N=2688,device_name=NVIDIA_H100_80GB_HBM3.json
- E=16,N=3072,device_name=NVIDIA_A100-SXM4-80GB,dtype=int8_w8a16.json
- E=16,N=3072,device_name=NVIDIA_H100_80GB_HBM3,dtype=int8_w8a16.json
- E=16,N=3200,device_name=NVIDIA_H100_80GB_HBM3,dtype=fp8_w8a8.json
- E=16,N=3584,device_name=NVIDIA_A100-SXM4-80GB,dtype=int8_w8a16.json
- E=16,N=3584,device_name=NVIDIA_A100-SXM4-80GB.json
- E=16,N=6400,device_name=NVIDIA_H100_80GB_HBM3,dtype=fp8_w8a8.json
- E=16,N=7168,device_name=NVIDIA_A100-SXM4-80GB,dtype=int8_w8a16.json
- E=16,N=7168,device_name=NVIDIA_A100-SXM4-80GB.json
- E=16,N=7168,device_name=NVIDIA_H100_80GB_HBM3,dtype=int8_w8a16.json
- E=16,N=800,device_name=NVIDIA_H100_80GB_HBM3,dtype=fp8_w8a8.json
- E=160,N=192,device_name=NVIDIA_A800-SXM4-80GB.json
- E=20,N=2048,device_name=NVIDIA_H100_80GB_HBM3.json
- E=24,N=1024,device_name=NVIDIA_H100_80GB_HBM3.json
- E=256,N=128,device_name=NVIDIA_A100-SXM4-80GB,dtype=int8_w8a8,block_shape=[128, 128].json
- E=256,N=128,device_name=NVIDIA_A100-SXM4-80GB,dtype=int8_w8a8.json
- E=256,N=128,device_name=NVIDIA_A800-SXM4-80GB,dtype=int8_w8a8,block_shape=[128, 128].json
- E=256,N=128,device_name=NVIDIA_A800-SXM4-80GB,dtype=int8_w8a8.json
- E=256,N=128,device_name=NVIDIA_H100_80GB_HBM3,dtype=fp8_w8a8,block_shape=[128, 128].json
- E=256,N=128,device_name=NVIDIA_H20,block_shape=[128, 128].json
- E=256,N=128,device_name=NVIDIA_H20,dtype=fp8_w8a8,block_shape=[128, 128].json
- E=256,N=128,device_name=NVIDIA_L20Y,dtype=fp8_w8a8,block_shape=[128, 128].json
- E=256,N=256,device_name=AMD_Instinct_MI300X,dtype=fp8_w8a8,block_shape=[128, 128].json
- E=256,N=256,device_name=AMD_Instinct_MI325X,dtype=fp8_w8a8,block_shape=[128, 128].json
- E=256,N=256,device_name=AMD_Radeon_Graphics,dtype=fp8_w8a8,block_shape=[128, 128].json
- E=256,N=256,device_name=NVIDIA_B200,dtype=fp8_w8a8,block_shape=[128, 128].json
- E=256,N=256,device_name=NVIDIA_H20,dtype=int8_w8a8,block_shape=[128, 128].json
- E=256,N=256,device_name=NVIDIA_H200,dtype=fp8_w8a8,block_shape=[128, 128].json
- E=256,N=64,device_name=NVIDIA_A800-SXM4-80GB.json
- E=256,N=64,device_name=NVIDIA_L20,dtype=int8_w8a8.json
- E=256,N=64,device_name=NVIDIA_L40S,dtype=int8_w8a8.json
- E=64,N=1024,device_name=NVIDIA_H100_80GB_HBM3,dtype=fp8_w8a8,block_shape=[128, 128].json
- E=64,N=1280,device_name=NVIDIA_A100-SXM4-80GB.json
- E=64,N=1280,device_name=NVIDIA_A800-SXM4-80GB.json
- E=64,N=1280,device_name=NVIDIA_H100_80GB_HBM3,dtype=fp8_w8a8.json
- E=64,N=1280,device_name=NVIDIA_H100_80GB_HBM3.json
- E=64,N=1280,device_name=NVIDIA_H200,dtype=fp8_w8a8.json
- E=64,N=1280,device_name=NVIDIA_H200.json
- E=64,N=2560,device_name=NVIDIA_H100_80GB_HBM3,dtype=fp8_w8a8.json
- E=64,N=2560,device_name=NVIDIA_H200,dtype=fp8_w8a8.json
- E=64,N=2560,device_name=NVIDIA_H200.json
- E=64,N=320,device_name=NVIDIA_H100_80GB_HBM3,dtype=fp8_w8a8.json
- E=64,N=320,device_name=NVIDIA_H100_80GB_HBM3.json
- E=64,N=320,device_name=NVIDIA_H200,dtype=fp8_w8a8.json
- E=64,N=320,device_name=NVIDIA_H200.json
- E=64,N=512,device_name=NVIDIA_H200,dtype=fp8_w8a8,block_shape=[128, 128].json
- E=64,N=640,device_name=NVIDIA_A100-SXM4-80GB.json
- E=64,N=640,device_name=NVIDIA_A800-SXM4-80GB.json
- E=64,N=640,device_name=NVIDIA_GeForce_RTX_4090,dtype=fp8_w8a8.json
- E=64,N=640,device_name=NVIDIA_H100_80GB_HBM3,dtype=fp8_w8a8.json
- E=64,N=640,device_name=NVIDIA_H100_80GB_HBM3.json
- E=64,N=640,device_name=NVIDIA_H200,dtype=fp8_w8a8.json
- E=64,N=640,device_name=NVIDIA_H200.json
- E=8,N=14336,device_name=AMD_Instinct_MI300X.json
- E=8,N=14336,device_name=AMD_Instinct_MI325X.json
- E=8,N=14336,device_name=AMD_Radeon_Graphics.json
- E=8,N=14336,device_name=NVIDIA_H100_80GB_HBM3,dtype=fp8_w8a8.json
- E=8,N=14336,device_name=NVIDIA_H200,dtype=fp8_w8a8.json
- E=8,N=14336,device_name=NVIDIA_H200.json
- E=8,N=1792,device_name=AMD_Instinct_MI300X.json
- E=8,N=1792,device_name=AMD_Instinct_MI325X.json
- E=8,N=1792,device_name=AMD_Radeon_Graphics.json
- E=8,N=1792,device_name=NVIDIA_A100-SXM4-40GB.json
- E=8,N=1792,device_name=NVIDIA_A100-SXM4-80GB.json
- E=8,N=1792,device_name=NVIDIA_H100_80GB_HBM3.json
- E=8,N=1792,device_name=NVIDIA_H200,dtype=fp8_w8a8.json
- E=8,N=1792,device_name=NVIDIA_H200.json
- E=8,N=2048,device_name=NVIDIA_A100-SXM4-80GB.json
- E=8,N=2048,device_name=NVIDIA_H100_80GB_HBM3,dtype=fp8_w8a8.json
- E=8,N=2048,device_name=NVIDIA_H100_80GB_HBM3.json
- E=8,N=2048,device_name=NVIDIA_H200,dtype=fp8_w8a8.json
- E=8,N=2048,device_name=NVIDIA_H200.json
- E=8,N=3584,device_name=AMD_Instinct_MI300X.json
- E=8,N=3584,device_name=AMD_Instinct_MI325X.json
- E=8,N=3584,device_name=AMD_Radeon_Graphics.json
- E=8,N=3584,device_name=NVIDIA_A100-SXM4-40GB.json
- E=8,N=3584,device_name=NVIDIA_A100-SXM4-80GB.json
- E=8,N=3584,device_name=NVIDIA_GeForce_RTX_4090,dtype=fp8_w8a8.json
- E=8,N=3584,device_name=NVIDIA_H100_80GB_HBM3,dtype=fp8_w8a8.json
- E=8,N=3584,device_name=NVIDIA_H100_80GB_HBM3.json
- E=8,N=3584,device_name=NVIDIA_H200,dtype=fp8_w8a8.json
- E=8,N=3584,device_name=NVIDIA_H200.json
- E=8,N=3584,device_name=NVIDIA_L40S.json
- E=8,N=4096,device_name=AMD_Instinct_MI300X,dtype=fp8_w8a8.json
- E=8,N=4096,device_name=AMD_Instinct_MI325X,dtype=fp8_w8a8.json
- E=8,N=4096,device_name=AMD_Radeon_Graphics,dtype=fp8_w8a8.json
- E=8,N=4096,device_name=NVIDIA_A100-SXM4-80GB.json
- E=8,N=4096,device_name=NVIDIA_H100_80GB_HBM3,dtype=fp8_w8a8.json
- E=8,N=4096,device_name=NVIDIA_H100_80GB_HBM3.json
- E=8,N=4096,device_name=NVIDIA_H200,dtype=fp8_w8a8.json
- E=8,N=4096,device_name=NVIDIA_H200.json
- E=8,N=7168,device_name=AMD_Instinct_MI300X.json
- E=8,N=7168,device_name=AMD_Instinct_MI325X.json
- E=8,N=7168,device_name=AMD_Radeon_Graphics.json
- E=8,N=7168,device_name=NVIDIA_A100-SXM4-80GB.json
- E=8,N=7168,device_name=NVIDIA_H100_80GB_HBM3,dtype=fp8_w8a8.json
- E=8,N=7168,device_name=NVIDIA_H100_80GB_HBM3.json
- E=8,N=7168,device_name=NVIDIA_H200,dtype=fp8_w8a8.json
- E=8,N=7168,device_name=NVIDIA_H200.json
- E=8,N=8192,device_name=AMD_Instinct_MI300X,dtype=fp8_w8a8.json
- E=8,N=8192,device_name=AMD_Instinct_MI325X,dtype=fp8_w8a8.json
- E=8,N=8192,device_name=AMD_Radeon_Graphics,dtype=fp8_w8a8.json
- E=8,N=8192,device_name=NVIDIA_H100_80GB_HBM3,dtype=fp8_w8a8.json
- E=8,N=8192,device_name=NVIDIA_H200,dtype=fp8_w8a8.json
- E=128,N=192,device_name=NVIDIA_A800-SXM4-80GB.json
- E=128,N=192,device_name=NVIDIA_H100_80GB_HBM3.json
- E=128,N=192,device_name=NVIDIA_H20.json
- E=128,N=192,device_name=NVIDIA_H200.json
- E=128,N=384,device_name=NVIDIA_H100_80GB_HBM3.json
- E=128,N=384,device_name=NVIDIA_H20,dtype=fp8_w8a8,block_shape=[128, 128].json
- E=128,N=384,device_name=NVIDIA_H20.json
- E=128,N=384,device_name=NVIDIA_H200,dtype=fp8_w8a8,block_shape=[128, 128].json
- E=128,N=384,device_name=NVIDIA_H200.json
- E=128,N=512,device_name=NVIDIA_H100_80GB_HBM3.json
- E=128,N=768,device_name=NVIDIA_A800-SXM4-80GB.json
- E=128,N=768,device_name=NVIDIA_H100_80GB_HBM3.json
- E=128,N=768,device_name=NVIDIA_H20,dtype=fp8_w8a8,block_shape=[128, 128].json
- E=128,N=768,device_name=NVIDIA_H20.json
- E=128,N=768,device_name=NVIDIA_H200,dtype=fp8_w8a8,block_shape=[128, 128].json
- E=128,N=768,device_name=NVIDIA_H200.json
- E=128,N=96,device_name=NVIDIA_H20.json
- E=129,N=352,device_name=NVIDIA_A800-SXM4-80GB,dtype=int8_w8a8.json
- E=160,N=320,device_name=NVIDIA_A800-SXM4-80GB,dtype=int8_w8a8.json
- E=161,N=192,device_name=NVIDIA_A800-SXM4-80GB,dtype=int8_w8a8.json
- E=257,N=128,device_name=NVIDIA_H100_80GB_HBM3,dtype=fp8_w8a8,block_shape=[128, 128].json
- E=257,N=128,device_name=NVIDIA_H20,dtype=fp8_w8a8,block_shape=[128, 128].json
- E=257,N=256,device_name=NVIDIA_B200,dtype=fp8_w8a8,block_shape=[128, 128].json
- E=257,N=256,device_name=NVIDIA_H20,dtype=fp8_w8a8,block_shape=[128, 128].json
- E=257,N=256,device_name=NVIDIA_H200,dtype=fp8_w8a8,block_shape=[128, 128].json
- E=264,N=128,device_name=NVIDIA_A800-SXM4-80GB,dtype=int8_w8a8.json
- E=264,N=256,device_name=NVIDIA_B200,dtype=fp8_w8a8,block_shape=[128, 128].json
- E=264,N=256,device_name=NVIDIA_H20,dtype=fp8_w8a8,block_shape=[128, 128].json
- E=264,N=256,device_name=NVIDIA_H200,dtype=fp8_w8a8,block_shape=[128, 128].json
- E=272,N=128,device_name=NVIDIA_A800-SXM4-80GB,dtype=int8_w8a8.json
- E=272,N=128,device_name=NVIDIA_H100_80GB_HBM3,dtype=fp8_w8a8,block_shape=[128, 128].json
- E=272,N=128,device_name=NVIDIA_H20,dtype=fp8_w8a8,block_shape=[128, 128].json
- E=272,N=64,device_name=NVIDIA_A800-SXM4-80GB.json
- E=288,N=64,device_name=NVIDIA_A800-SXM4-80GB.json
- E=8,N=7168,device_name=NVIDIA_H100_80GB_HBM3.json
- E=16,N=1024,device_name=NVIDIA_B200.json
- E=128,N=352,device_name=NVIDIA_RTX_6000_Ada_Generation,dtype=fp8_w8a8.json
- E=128,N=384,device_name=NVIDIA_H100_80GB_HBM3,dtype=fp8_w8a8,block_shape=[128, 128].json
- E=128,N=384,device_name=NVIDIA_H20,dtype=fp8_w8a8,block_shape=[128, 128].json
- E=128,N=768,device_name=NVIDIA_H20,dtype=fp8_w8a8,block_shape=[128, 128].json
- E=128,N=768,device_name=NVIDIA_H20.json
- E=160,N=192,device_name=NVIDIA_H200,dtype=fp8_w8a8.json
- E=160,N=320,device_name=NVIDIA_H20-3e.json
- E=160,N=384,device_name=NVIDIA_H200,dtype=fp8_w8a8.json
- E=160,N=640,device_name=NVIDIA_H200,dtype=fp8_w8a8,block_shape=[128, 128].json
- E=257,N=128,device_name=NVIDIA_H20,dtype=fp8_w8a8,block_shape=[128, 128].json
- E=257,N=128,device_name=NVIDIA_H200,dtype=fp8_w8a8,block_shape=[128, 128].json
- E=257,N=256,device_name=NVIDIA_B200,dtype=fp8_w8a8,block_shape=[128, 128].json
- E=257,N=256,device_name=NVIDIA_H20,dtype=fp8_w8a8,block_shape=[128, 128].json
- E=257,N=256,device_name=NVIDIA_H20-3e,dtype=fp8_w8a8,block_shape=[128, 128].json
- E=257,N=256,device_name=NVIDIA_H200,dtype=fp8_w8a8,block_shape=[128, 128].json
- E=384,N=128,device_name=NVIDIA_H20,dtype=fp8_w8a8,block_shape=[128, 128].json
- E=384,N=128,device_name=NVIDIA_H200,dtype=fp8_w8a8,block_shape=[128, 128].json
- E=384,N=256,device_name=NVIDIA_H20-3e,dtype=fp8_w8a8,block_shape=[128, 128].json
- E=385,N=128,device_name=NVIDIA_B200,dtype=fp8_w8a8,block_shape=[128, 128].json
- E=385,N=128,device_name=NVIDIA_H200,dtype=fp8_w8a8,block_shape=[128, 128].json
- E=8,N=7168,device_name=NVIDIA_H100_80GB_HBM3.json
- E=128,N=352,device_name=NVIDIA_RTX_5880_Ada_Generation,dtype=fp8_w8a8.json
- E=128,N=384,device_name=NVIDIA_B200,dtype=fp8_w8a8,block_shape=[128, 128].json
- E=129,N=352,device_name=NVIDIA_B200,dtype=fp8_w8a8.json
- E=129,N=352,device_name=NVIDIA_RTX_PRO_6000_Blackwell_Max-Q_Workstation_Edition,dtype=fp8_w8a8.json
- E=129,N=704,device_name=NVIDIA_B200,dtype=fp8_w8a8.json
- E=161,N=384,device_name=NVIDIA_RTX_PRO_6000_Blackwell_Max-Q_Workstation_Edition,dtype=fp8_w8a8.json
- E=256,N=256,device_name=NVIDIA_B200,dtype=fp8_w8a8,block_shape=[128, 128].json
- E=256,N=256,device_name=NVIDIA_B200.json
- E=256,N=256,device_name=NVIDIA_H800,dtype=fp8_w8a8,block_shape=[128, 128].json
- E=256,N=512,device_name=NVIDIA_H20.json
- E=257,N=128,device_name=NVIDIA_H100_80GB_HBM3,dtype=fp8_w8a8,block_shape=[128, 128].json
- E=257,N=64,device_name=NVIDIA_A100-SXM4-80GB.json
- E=384,N=256,device_name=NVIDIA_B200,dtype=fp8_w8a8,block_shape=[128, 128].json
- E=512,N=128,device_name=NVIDIA_H100_80GB_HBM3.json
- E=512,N=128,device_name=NVIDIA_H20-3e.json
- E=512,N=128,device_name=NVIDIA_H200.json
- E=512,N=128,device_name=NVIDIA_H800,dtype=fp8_w8a8,block_shape=[128, 128].json
- E=512,N=256,device_name=NVIDIA_B200.json
- E=512,N=256,device_name=NVIDIA_H20-3e.json
- E=512,N=256,device_name=NVIDIA_H200.json
- E=512,N=64,device_name=NVIDIA_H100_80GB_HBM3.json
- E=512,N=64,device_name=NVIDIA_H200.json
- README.md
- __init__.py
- fused_moe.py
- fused_moe_triton_config.py
- fused_moe_triton_kernels.py
- layer.py
- moe_align_block_size.py
- triton_kernels_moe.py
- __init__.py
- base.py
- deep_gemm.py
- runner.py
- triton.py
- __init__.py
- base.py
- deepep.py
- mooncake.py
- standard.py
- __init__.py
- cutlass_moe.py
- cutlass_moe_params.py
- cutlass_w4a8_moe.py
- flashinfer_cutedsl_moe.py
- fused_moe_native.py
- rocm_moe_utils.py
- router.py
- topk.py
- utils.py
- __init__.py
- compressed_tensors_scheme.py
- compressed_tensors_w8a16_fp8.py
- compressed_tensors_w8a8_fp8.py
- compressed_tensors_w8a8_int8.py
- __init__.py
- compressed_tensors.py
- compressed_tensors_moe.py
- README.md
- utils.py
- N=1536,K=1536,device_name=NVIDIA_A100-SXM4-80GB,dtype=int8_w8a8,block_shape=[128, 128].json
- N=1536,K=1536,device_name=NVIDIA_A800-SXM4-80GB,dtype=int8_w8a8,block_shape=[128, 128].json
- N=1536,K=1536,device_name=NVIDIA_H100_80GB_HBM3,dtype=fp8_w8a8,block_shape=[128, 128].json
- N=1536,K=1536,device_name=NVIDIA_H20,dtype=fp8_w8a8,block_shape=[128, 128].json
- N=1536,K=1536,device_name=NVIDIA_L20Y,dtype=fp8_w8a8,block_shape=[128, 128].json
- N=1536,K=7168,device_name=AMD_Instinct_MI300X,dtype=fp8_w8a8,block_shape=[128, 128].json
- N=1536,K=7168,device_name=AMD_Instinct_MI325X,dtype=fp8_w8a8,block_shape=[128, 128].json
- N=1536,K=7168,device_name=AMD_Radeon_Graphics,dtype=fp8_w8a8,block_shape=[128, 128].json
- N=1536,K=7168,device_name=NVIDIA_A100-SXM4-80GB,dtype=int8_w8a8,block_shape=[128, 128].json
- N=1536,K=7168,device_name=NVIDIA_A800-SXM4-80GB,dtype=int8_w8a8,block_shape=[128, 128].json
- N=1536,K=7168,device_name=NVIDIA_H100_80GB_HBM3,dtype=fp8_w8a8,block_shape=[128, 128].json
- N=1536,K=7168,device_name=NVIDIA_H20,dtype=fp8_w8a8,block_shape=[128, 128].json
- N=1536,K=7168,device_name=NVIDIA_H200,dtype=fp8_w8a8,block_shape=[128, 128].json
- N=1536,K=7168,device_name=NVIDIA_L20Y,dtype=fp8_w8a8,block_shape=[128, 128].json
- N=2048,K=512,device_name=NVIDIA_A100-SXM4-80GB,dtype=int8_w8a8,block_shape=[128, 128].json
- N=2048,K=512,device_name=NVIDIA_A800-SXM4-80GB,dtype=int8_w8a8,block_shape=[128, 128].json
- N=2048,K=512,device_name=NVIDIA_H100_80GB_HBM3,dtype=fp8_w8a8,block_shape=[128, 128].json
- N=2048,K=512,device_name=NVIDIA_H20,dtype=fp8_w8a8,block_shape=[128, 128].json
- N=2048,K=512,device_name=NVIDIA_H200,dtype=fp8_w8a8,block_shape=[128, 128].json
- N=2048,K=512,device_name=NVIDIA_L20Y,dtype=fp8_w8a8,block_shape=[128, 128].json
- N=2304,K=7168,device_name=NVIDIA_A100-SXM4-80GB,dtype=int8_w8a8,block_shape=[128, 128].json
- N=2304,K=7168,device_name=NVIDIA_A800-SXM4-80GB,dtype=int8_w8a8,block_shape=[128, 128].json
- N=2304,K=7168,device_name=NVIDIA_H100_80GB_HBM3,dtype=fp8_w8a8,block_shape=[128, 128].json
- N=2304,K=7168,device_name=NVIDIA_H20,dtype=fp8_w8a8,block_shape=[128, 128].json
- N=2304,K=7168,device_name=NVIDIA_H200,dtype=fp8_w8a8,block_shape=[128, 128].json
- N=2304,K=7168,device_name=NVIDIA_L20Y,dtype=fp8_w8a8,block_shape=[128, 128].json
- N=24576,K=1536,device_name=AMD_Instinct_MI300X,dtype=fp8_w8a8,block_shape=[128, 128].json
- N=24576,K=1536,device_name=AMD_Instinct_MI325X,dtype=fp8_w8a8,block_shape=[128, 128].json
- N=24576,K=1536,device_name=AMD_Radeon_Graphics,dtype=fp8_w8a8,block_shape=[128, 128].json
- N=24576,K=7168,device_name=NVIDIA_A100-SXM4-80GB,dtype=int8_w8a8,block_shape=[128, 128].json
- N=24576,K=7168,device_name=NVIDIA_A800-SXM4-80GB,dtype=int8_w8a8,block_shape=[128, 128].json
- N=24576,K=7168,device_name=NVIDIA_B200,dtype=fp8_w8a8,block_shape=[128, 128].json
- N=24576,K=7168,device_name=NVIDIA_H100_80GB_HBM3,dtype=fp8_w8a8,block_shape=[128, 128].json
- N=24576,K=7168,device_name=NVIDIA_H20,dtype=fp8_w8a8,block_shape=[128, 128].json
- N=24576,K=7168,device_name=NVIDIA_H20,dtype=int8_w8a8,block_shape=[128, 128].json
- N=24576,K=7168,device_name=NVIDIA_H200,dtype=fp8_w8a8,block_shape=[128, 128].json
- N=24576,K=7168,device_name=NVIDIA_L20Y,dtype=fp8_w8a8,block_shape=[128, 128].json
- N=256,K=7168,device_name=NVIDIA_A100-SXM4-80GB,dtype=int8_w8a8,block_shape=[128, 128].json
- N=256,K=7168,device_name=NVIDIA_A800-SXM4-80GB,dtype=int8_w8a8,block_shape=[128, 128].json
- N=256,K=7168,device_name=NVIDIA_H100_80GB_HBM3,dtype=fp8_w8a8,block_shape=[128, 128].json
- N=256,K=7168,device_name=NVIDIA_H20,dtype=fp8_w8a8,block_shape=[128, 128].json
- N=256,K=7168,device_name=NVIDIA_L20Y,dtype=fp8_w8a8,block_shape=[128, 128].json
- N=3072,K=1536,device_name=AMD_Instinct_MI300X,dtype=fp8_w8a8,block_shape=[128, 128].json
- N=3072,K=1536,device_name=AMD_Instinct_MI325X,dtype=fp8_w8a8,block_shape=[128, 128].json
- N=3072,K=1536,device_name=AMD_Radeon_Graphics,dtype=fp8_w8a8,block_shape=[128, 128].json
- N=3072,K=1536,device_name=NVIDIA_B200,dtype=fp8_w8a8,block_shape=[128, 128].json
- N=3072,K=1536,device_name=NVIDIA_H20,dtype=int8_w8a8,block_shape=[128, 128].json
- N=3072,K=1536,device_name=NVIDIA_H200,dtype=fp8_w8a8,block_shape=[128, 128].json
- N=3072,K=7168,device_name=NVIDIA_B200,dtype=fp8_w8a8,block_shape=[128, 128].json
- N=3072,K=7168,device_name=NVIDIA_H100_80GB_HBM3,dtype=fp8_w8a8,block_shape=[128, 128].json
- N=3072,K=7168,device_name=NVIDIA_H20,dtype=int8_w8a8,block_shape=[128, 128].json
- N=3072,K=7168,device_name=NVIDIA_H200,dtype=fp8_w8a8,block_shape=[128, 128].json
- N=32768,K=512,device_name=AMD_Instinct_MI300X,dtype=fp8_w8a8,block_shape=[128, 128].json
- N=32768,K=512,device_name=AMD_Instinct_MI325X,dtype=fp8_w8a8,block_shape=[128, 128].json
- N=32768,K=512,device_name=AMD_Radeon_Graphics,dtype=fp8_w8a8,block_shape=[128, 128].json
- N=32768,K=512,device_name=NVIDIA_A100-SXM4-80GB,dtype=int8_w8a8,block_shape=[128, 128].json
- N=32768,K=512,device_name=NVIDIA_A800-SXM4-80GB,dtype=int8_w8a8,block_shape=[128, 128].json
- N=32768,K=512,device_name=NVIDIA_B200,dtype=fp8_w8a8,block_shape=[128, 128].json
- N=32768,K=512,device_name=NVIDIA_H100_80GB_HBM3,dtype=fp8_w8a8,block_shape=[128, 128].json
- N=32768,K=512,device_name=NVIDIA_H20,dtype=fp8_w8a8,block_shape=[128, 128].json
- N=32768,K=512,device_name=NVIDIA_H20,dtype=int8_w8a8,block_shape=[128, 128].json
- N=32768,K=512,device_name=NVIDIA_H200,dtype=fp8_w8a8,block_shape=[128, 128].json
- N=32768,K=512,device_name=NVIDIA_L20Y,dtype=fp8_w8a8,block_shape=[128, 128].json
- N=36864,K=7168,device_name=NVIDIA_H100_80GB_HBM3,dtype=fp8_w8a8,block_shape=[128, 128].json
- N=36864,K=7168,device_name=NVIDIA_H200,dtype=fp8_w8a8,block_shape=[128, 128].json
- N=4096,K=512,device_name=AMD_Instinct_MI300X,dtype=fp8_w8a8,block_shape=[128, 128].json
- N=4096,K=512,device_name=AMD_Instinct_MI325X,dtype=fp8_w8a8,block_shape=[128, 128].json
- N=4096,K=512,device_name=AMD_Radeon_Graphics,dtype=fp8_w8a8,block_shape=[128, 128].json
- N=4096,K=512,device_name=NVIDIA_B200,dtype=fp8_w8a8,block_shape=[128, 128].json
- N=4096,K=512,device_name=NVIDIA_H100_80GB_HBM3,dtype=fp8_w8a8,block_shape=[128, 128].json
- N=4096,K=512,device_name=NVIDIA_H20,dtype=int8_w8a8,block_shape=[128, 128].json
- N=4096,K=512,device_name=NVIDIA_H200,dtype=fp8_w8a8,block_shape=[128, 128].json
- N=4608,K=7168,device_name=AMD_Instinct_MI300X,dtype=fp8_w8a8,block_shape=[128, 128].json
- N=4608,K=7168,device_name=AMD_Instinct_MI325X,dtype=fp8_w8a8,block_shape=[128, 128].json
- N=4608,K=7168,device_name=AMD_Radeon_Graphics,dtype=fp8_w8a8,block_shape=[128, 128].json
- N=4608,K=7168,device_name=NVIDIA_B200,dtype=fp8_w8a8,block_shape=[128, 128].json
- N=4608,K=7168,device_name=NVIDIA_H100_80GB_HBM3,dtype=fp8_w8a8,block_shape=[128, 128].json
- N=4608,K=7168,device_name=NVIDIA_H20,dtype=int8_w8a8,block_shape=[128, 128].json
- N=4608,K=7168,device_name=NVIDIA_H200,dtype=fp8_w8a8,block_shape=[128, 128].json
- N=512,K=7168,device_name=AMD_Instinct_MI300X,dtype=fp8_w8a8,block_shape=[128, 128].json
- N=512,K=7168,device_name=AMD_Instinct_MI325X,dtype=fp8_w8a8,block_shape=[128, 128].json
- N=512,K=7168,device_name=AMD_Radeon_Graphics,dtype=fp8_w8a8,block_shape=[128, 128].json
- N=512,K=7168,device_name=NVIDIA_B200,dtype=fp8_w8a8,block_shape=[128, 128].json
- N=512,K=7168,device_name=NVIDIA_H20,dtype=int8_w8a8,block_shape=[128, 128].json
- N=512,K=7168,device_name=NVIDIA_H200,dtype=fp8_w8a8,block_shape=[128, 128].json
- N=576,K=7168,device_name=AMD_Instinct_MI300X,dtype=fp8_w8a8,block_shape=[128, 128].json
- N=576,K=7168,device_name=AMD_Instinct_MI325X,dtype=fp8_w8a8,block_shape=[128, 128].json
- N=576,K=7168,device_name=AMD_Radeon_Graphics,dtype=fp8_w8a8,block_shape=[128, 128].json
- N=576,K=7168,device_name=NVIDIA_A100-SXM4-80GB,dtype=int8_w8a8,block_shape=[128, 128].json
- N=576,K=7168,device_name=NVIDIA_A800-SXM4-80GB,dtype=int8_w8a8,block_shape=[128, 128].json
- N=576,K=7168,device_name=NVIDIA_B200,dtype=fp8_w8a8,block_shape=[128, 128].json
- N=576,K=7168,device_name=NVIDIA_H100_80GB_HBM3,dtype=fp8_w8a8,block_shape=[128, 128].json
- N=576,K=7168,device_name=NVIDIA_H20,dtype=fp8_w8a8,block_shape=[128, 128].json
- N=576,K=7168,device_name=NVIDIA_H20,dtype=int8_w8a8,block_shape=[128, 128].json
- N=576,K=7168,device_name=NVIDIA_H200,dtype=fp8_w8a8,block_shape=[128, 128].json
- N=576,K=7168,device_name=NVIDIA_L20Y,dtype=fp8_w8a8,block_shape=[128, 128].json
- N=7168,K=1024,device_name=NVIDIA_A100-SXM4-80GB,dtype=int8_w8a8,block_shape=[128, 128].json
- N=7168,K=1024,device_name=NVIDIA_A800-SXM4-80GB,dtype=int8_w8a8,block_shape=[128, 128].json
- N=7168,K=1024,device_name=NVIDIA_H100_80GB_HBM3,dtype=fp8_w8a8,block_shape=[128, 128].json
- N=7168,K=1024,device_name=NVIDIA_H20,dtype=fp8_w8a8,block_shape=[128, 128].json
- N=7168,K=1024,device_name=NVIDIA_H200,dtype=fp8_w8a8,block_shape=[128, 128].json
- N=7168,K=1024,device_name=NVIDIA_L20Y,dtype=fp8_w8a8,block_shape=[128, 128].json
- N=7168,K=1152,device_name=NVIDIA_A100-SXM4-80GB,dtype=int8_w8a8,block_shape=[128, 128].json
- N=7168,K=1152,device_name=NVIDIA_A800-SXM4-80GB,dtype=int8_w8a8,block_shape=[128, 128].json
- N=7168,K=1152,device_name=NVIDIA_H100_80GB_HBM3,dtype=fp8_w8a8,block_shape=[128, 128].json
- N=7168,K=1152,device_name=NVIDIA_H20,dtype=fp8_w8a8,block_shape=[128, 128].json
- N=7168,K=1152,device_name=NVIDIA_H200,dtype=fp8_w8a8,block_shape=[128, 128].json
- N=7168,K=1152,device_name=NVIDIA_L20Y,dtype=fp8_w8a8,block_shape=[128, 128].json
- N=7168,K=128,device_name=NVIDIA_A100-SXM4-80GB,dtype=int8_w8a8,block_shape=[128, 128].json
- N=7168,K=128,device_name=NVIDIA_A800-SXM4-80GB,dtype=int8_w8a8,block_shape=[128, 128].json
- N=7168,K=128,device_name=NVIDIA_H100_80GB_HBM3,dtype=fp8_w8a8,block_shape=[128, 128].json
- N=7168,K=128,device_name=NVIDIA_H20,dtype=fp8_w8a8,block_shape=[128, 128].json
- N=7168,K=128,device_name=NVIDIA_L20Y,dtype=fp8_w8a8,block_shape=[128, 128].json
- N=7168,K=16384,device_name=AMD_Instinct_MI300X,dtype=fp8_w8a8,block_shape=[128, 128].json
- N=7168,K=16384,device_name=AMD_Instinct_MI325X,dtype=fp8_w8a8,block_shape=[128, 128].json
- N=7168,K=16384,device_name=AMD_Radeon_Graphics,dtype=fp8_w8a8,block_shape=[128, 128].json
- N=7168,K=16384,device_name=NVIDIA_A100-SXM4-80GB,dtype=int8_w8a8,block_shape=[128, 128].json
- N=7168,K=16384,device_name=NVIDIA_A800-SXM4-80GB,dtype=int8_w8a8,block_shape=[128, 128].json
- N=7168,K=16384,device_name=NVIDIA_B200,dtype=fp8_w8a8,block_shape=[128, 128].json
- N=7168,K=16384,device_name=NVIDIA_H100_80GB_HBM3,dtype=fp8_w8a8,block_shape=[128, 128].json
- N=7168,K=16384,device_name=NVIDIA_H20,dtype=fp8_w8a8,block_shape=[128, 128].json
- N=7168,K=16384,device_name=NVIDIA_H20,dtype=int8_w8a8,block_shape=[128, 128].json
- N=7168,K=16384,device_name=NVIDIA_H200,dtype=fp8_w8a8,block_shape=[128, 128].json
- N=7168,K=16384,device_name=NVIDIA_L20Y,dtype=fp8_w8a8,block_shape=[128, 128].json
- N=7168,K=18432,device_name=NVIDIA_A100-SXM4-80GB,dtype=int8_w8a8,block_shape=[128, 128].json
- N=7168,K=18432,device_name=NVIDIA_A800-SXM4-80GB,dtype=int8_w8a8,block_shape=[128, 128].json
- N=7168,K=18432,device_name=NVIDIA_B200,dtype=fp8_w8a8,block_shape=[128, 128].json
- N=7168,K=18432,device_name=NVIDIA_H100_80GB_HBM3,dtype=fp8_w8a8,block_shape=[128, 128].json
- N=7168,K=18432,device_name=NVIDIA_H20,dtype=fp8_w8a8,block_shape=[128, 128].json
- N=7168,K=18432,device_name=NVIDIA_H20,dtype=int8_w8a8,block_shape=[128, 128].json
- N=7168,K=18432,device_name=NVIDIA_H200,dtype=fp8_w8a8,block_shape=[128, 128].json
- N=7168,K=18432,device_name=NVIDIA_L20Y,dtype=fp8_w8a8,block_shape=[128, 128].json
- N=7168,K=2048,device_name=AMD_Instinct_MI300X,dtype=fp8_w8a8,block_shape=[128, 128].json
- N=7168,K=2048,device_name=AMD_Instinct_MI325X,dtype=fp8_w8a8,block_shape=[128, 128].json
- N=7168,K=2048,device_name=AMD_Radeon_Graphics,dtype=fp8_w8a8,block_shape=[128, 128].json
- N=7168,K=2048,device_name=NVIDIA_B200,dtype=fp8_w8a8,block_shape=[128, 128].json
- N=7168,K=2048,device_name=NVIDIA_H100_80GB_HBM3,dtype=fp8_w8a8,block_shape=[128, 128].json
- N=7168,K=2048,device_name=NVIDIA_H20,dtype=int8_w8a8,block_shape=[128, 128].json
- N=7168,K=2048,device_name=NVIDIA_H200,dtype=fp8_w8a8,block_shape=[128, 128].json
- N=7168,K=2304,device_name=AMD_Instinct_MI300X,dtype=fp8_w8a8,block_shape=[128, 128].json
- N=7168,K=2304,device_name=AMD_Instinct_MI325X,dtype=fp8_w8a8,block_shape=[128, 128].json
- N=7168,K=2304,device_name=AMD_Radeon_Graphics,dtype=fp8_w8a8,block_shape=[128, 128].json
- N=7168,K=2304,device_name=NVIDIA_B200,dtype=fp8_w8a8,block_shape=[128, 128].json
- N=7168,K=2304,device_name=NVIDIA_H100_80GB_HBM3,dtype=fp8_w8a8,block_shape=[128, 128].json
- N=7168,K=2304,device_name=NVIDIA_H20,dtype=int8_w8a8,block_shape=[128, 128].json
- N=7168,K=2304,device_name=NVIDIA_H200,dtype=fp8_w8a8,block_shape=[128, 128].json
- N=7168,K=256,device_name=AMD_Instinct_MI300X,dtype=fp8_w8a8,block_shape=[128, 128].json
- N=7168,K=256,device_name=AMD_Instinct_MI325X,dtype=fp8_w8a8,block_shape=[128, 128].json
- N=7168,K=256,device_name=AMD_Radeon_Graphics,dtype=fp8_w8a8,block_shape=[128, 128].json
- N=7168,K=256,device_name=NVIDIA_B200,dtype=fp8_w8a8,block_shape=[128, 128].json
- N=7168,K=256,device_name=NVIDIA_H20,dtype=int8_w8a8,block_shape=[128, 128].json
- N=7168,K=256,device_name=NVIDIA_H200,dtype=fp8_w8a8,block_shape=[128, 128].json
- __init__.py
- compile_utils.py
- configurer.py
- entrypoint.py
- __init__.py
- quark_scheme.py
- quark_w4a4_mxfp4.py
- __init__.py
- quark.py
- quark_moe.py
- utils.py
- __init__.py
- awq.py
- awq_triton.py
- base_config.py
- blockwise_int8.py
- fp8.py
- fp8_kernel.py
- fp8_utils.py
- fpgemm_fp8.py
- gptq.py
- int8_kernel.py
- int8_utils.py
- kv_cache.py
- marlin_utils.py
- marlin_utils_fp8.py
- modelopt_quant.py
- moe_wna16.py
- mxfp4.py
- mxfp4_tensor.py
- petit.py
- petit_utils.py
- qoq.py
- rocm_mxfp4_utils.py
- unquant.py
- utils.py
- w4afp8.py
- w8a8_fp8.py
- w8a8_int8.py
- activation.py
- amx_utils.py
- communicator.py
- dp_attention.py
- elementwise.py
- flashinfer_comm_fusion.py
- layernorm.py
- linear.py
- logits_processor.py
- model_parallel.py
- modelopt_utils.py
- multimodal.py
- parameter.py
- pooler.py
- radix_attention.py
- rocm_linear_utils.py
- rotary_embedding.py
- sampler.py
- torchao_utils.py
- utils.py
- vocab_parallel_embedding.py
- base_backend.py
- chunked_backend.py
- triton_backend.py
- __init__.py
- chunked_sgmv_expand.py
- chunked_sgmv_shrink.py
- gate_up_lora_b.py
- qkv_lora_b.py
- sgemm_lora_a.py
- sgemm_lora_b.py
- eviction_policy.py
- layers.py
- lora.py
- lora_config.py
- lora_manager.py
- lora_registry.py
- mem_pool.py
- utils.py
- async_dynamic_batch_tokenizer.py
- cache_controller.py
- configure_logging.py
- data_parallel_controller.py
- detokenizer_manager.py
- disagg_service.py
- io_struct.py
- mm_utils.py
- multi_tokenizer_mixin.py
- multimodal_processor.py
- overlap_utils.py
- schedule_batch.py
- schedule_policy.py
- scheduler.py
- scheduler_input_blocker.py
- scheduler_metrics_mixin.py
- scheduler_output_processor_mixin.py
- scheduler_profiler_mixin.py
- scheduler_recv_skipper.py
- scheduler_update_weights_mixin.py
- session_controller.py
- template_manager.py
- tokenizer_communicator_mixin.py
- tokenizer_manager.py
- tp_worker.py
- utils.py
- .clang-format
- common.h
- radix_tree.py
- tree_v2.cpp
- tree_v2.h
- tree_v2_binding.cpp
- tree_v2_debug.cpp
- tree_v2_impl.h
- tree_v2_node.h
- aibrix_kvcache_storage.py
- README.md
- unit_test.py
- eic_storage.py
- README.md
- test_unit.py
- deploy_sglang_3fs_multinode.md
- README.md
- setup_usrbio_client.md
- hf3fs_client.py
- hf3fs_usrbio_client.py
- hf3fs_utils.cpp
- mini_3fs_metadata_server.py
- storage_hf3fs.py
- test_hf3fs_utils.py
- example_config.yaml
- lmc_radix_cache.py
- README.md
- unit_test.py
- mooncake_store.py
- README.md
- test_mooncake_store.py
- hicache_nixl.py
- nixl_utils.py
- README.md
- test_hicache_nixl_storage.py
- __init__.py
- backend_factory.py
- allocator.py
- allocator_ascend.py
- base_prefix_cache.py
- chunk_cache.py
- common.py
- evict_policy.py
- flush_cache.py
- hicache_storage.py
- hiradix_cache.py
- mamba_radix_cache.py
- memory_pool.py
- memory_pool_host.py
- multimodal_cache.py
- radix_cache.py
- radix_cache_cpp.py
- swa_radix_cache.py
- collector.py
- func_timer.py
- startup_func_log_and_timer.py
- utils.py
- cpu_graph_runner.py
- cuda_graph_runner.py
- forward_batch_info.py
- model_runner.py
- npu_graph_runner.py
- piecewise_cuda_graph_runner.py
- __init__.py
- loader.py
- remote_instance_weight_loader_utils.py
- utils.py
- weight_utils.py
- apertus.py
- arcee.py
- baichuan.py
- bailing_moe.py
- bailing_moe_nextn.py
- bert.py
- chatglm.py
- clip.py
- commandr.py
- dbrx.py
- deepseek.py
- deepseek_janus_pro.py
- deepseek_nextn.py
- deepseek_v2.py
- deepseek_vl2.py
- dots_ocr.py
- dots_vlm.py
- dots_vlm_vit.py
- ernie4.py
- ernie4_eagle.py
- exaone.py
- falcon_h1.py
- gemma.py
- gemma2.py
- gemma2_reward.py
- gemma3_causal.py
- gemma3_mm.py
- gemma3n_audio.py
- gemma3n_causal.py
- gemma3n_mm.py
- glm4.py
- glm4_moe.py
- glm4_moe_nextn.py
- glm4v.py
- glm4v_moe.py
- gpt2.py
- gpt_bigcode.py
- gpt_oss.py
- granite.py
- granitemoe.py
- grok.py
- hunyuan.py
- idefics2.py
- internlm2.py
- internlm2_reward.py
- interns1.py
- internvl.py
- kimi_vl.py
- kimi_vl_moonvit.py
- llama.py
- llama4.py
- llama_classification.py
- llama_eagle.py
- llama_eagle3.py
- llama_embedding.py
- llama_reward.py
- llava.py
- llavavid.py
- longcat_flash.py
- longcat_flash_nextn.py
- mimo.py
- mimo_mtp.py
- minicpm.py
- minicpm3.py
- minicpmo.py
- minicpmv.py
- mistral.py
- mixtral.py
- mixtral_quant.py
- mllama.py
- mllama4.py
- nemotron_h.py
- nemotron_nas.py
- olmo.py
- olmo2.py
- olmoe.py
- opt.py
- persimmon.py
- phi.py
- phi3_small.py
- phi4mm.py
- phi4mm_audio.py
- phi4mm_utils.py
- phimoe.py
- pixtral.py
- qwen.py
- qwen2.py
- qwen2_5_vl.py
- qwen2_audio.py
- qwen2_classification.py
- qwen2_eagle.py
- qwen2_moe.py
- qwen2_rm.py
- qwen2_vl.py
- qwen3.py
- qwen3_classification.py
- qwen3_moe.py
- qwen3_next.py
- qwen3_next_mtp.py
- qwen3_vl.py
- qwen3_vl_moe.py
- registry.py
- roberta.py
- sarashina2_vision.py
- siglip.py
- solar.py
- stablelm.py
- starcoder2.py
- step3_vl.py
- torch_native_llama.py
- transformers.py
- utils.py
- vila.py
- xverse.py
- xverse_moe.py
- yivl.py
- base_processor.py
- clip.py
- deepseek_vl_v2.py
- dots_vlm.py
- gemma3.py
- gemma3n.py
- glm4v.py
- internvl.py
- janus_pro.py
- kimi_vl.py
- llava.py
- minicpm.py
- mlama.py
- mllama4.py
- phi4mm.py
- pixtral.py
- qwen_audio.py
- qwen_vl.py
- sarashina2_vision.py
- step3_vl.py
- vila.py
- mm_utils.py
- code_completion_parser.py
- conversation.py
- harmony_parser.py
- jinja_template_utils.py
- reasoning_parser.py
- __init__.py
- frequency_penalty.py
- min_new_tokens.py
- orchestrator.py
- presence_penalty.py
- custom_logit_processor.py
- sampling_batch_info.py
- sampling_params.py
- .clang-format
- ngram.cpp
- ngram.h
- ngram_cache.py
- ngram_cache_binding.cpp
- param.h
- queue.h
- draft_utils.py
- eagle_draft_cuda_graph_runner.py
- eagle_draft_extend_cuda_graph_runner.py
- eagle_info.py
- eagle_info_v2.py
- eagle_mab.py
- eagle_utils.py
- eagle_worker.py
- eagle_worker_v2.py
- ngram_info.py
- ngram_worker.py
- spec_info.py
- spec_utils.py
- standalone_worker.py
- tiktoken_tokenizer.py
- trace.py
- __init__.py
- aio_rwlock.py
- bench_utils.py
- common.py
- hf_transformers_utils.py
- host_shared_memory.py
- offloader.py
- patch_torch.py
- poll_based_barrier.py
- profile_merger.py
- rpd_utils.py
- slow_rank_detector.py
- torch_memory_saver_adapter.py
- tensor_bucket.py
- utils.py
- _custom_ops.py
- constants.py
- custom_op.py
- environ.py
- operations.py
- operations_strategy.py
- server_args.py
- server_args_config_parser.py
- single_batch_overlap.py
- two_batch_overlap.py
- warmup.py
- __init__.py
- test_flashattn_backend.py
- test_flashattn_mla_backend.py
- test_prefix_chunk_info.py
- test_trtllm_mla_backend.py
- __init__.py
- longbench_v2_evaluation.md
- test_longbench_v2_eval.py
- validate_longbench_v2.py
- validate_longbench_v2_standalone.py
- __init__.py
- doc_patch.py
- few_shot_gsm8k.py
- few_shot_gsm8k_engine.py
- get_logits_ut.py
- long_prompt.txt
- run_eval.py
- runners.py
- send_one.py
- simple_eval_common.py
- simple_eval_gpqa.py
- simple_eval_humaneval.py
- simple_eval_longbench_v2.py
- simple_eval_math.py
- simple_eval_mgsm.py
- simple_eval_mmlu.py
- simple_eval_mmmu_vlm.py
- test_activation.py
- test_block_fp8.py
- test_block_fp8_deep_gemm_blackwell.py
- test_custom_ops.py
- test_cutlass_moe.py
- test_cutlass_w4a8_moe.py
- test_deepep_utils.py
- test_deterministic.py
- test_deterministic_utils.py
- test_disaggregation_utils.py
- test_dynamic_grad_mode.py
- test_fp4_moe.py
- test_layernorm.py
- test_marlin_moe.py
- test_marlin_utils.py
- test_programs.py
- test_utils.py
- __init__.py
- bench_offline_throughput.py
- bench_one_batch.py
- bench_one_batch_server.py
- bench_serving.py
- check_env.py
- compile_deep_gemm.py
- global_config.py
- launch_server.py
- profiler.py
- README.md
- utils.py
- version.py
- pyproject.toml
- pyproject_cpu.toml
- pyproject_other.toml
- pyproject_xpu.toml
- amd_ci_exec.sh
- amd_ci_install_dependency.sh
- amd_ci_start_container.sh
- ci_install_deepep.sh
- ci_install_dependency.sh
- ci_install_rust.sh
- ci_start_disaggregation_servers.sh
- npu_ci_install_dependency.sh
- publish_traces.py
- ci_analyzer.py
- ci_analyzer_perf.py
- example.sh
- README.md
- copy_from_oss.py
- copy_to_oss.py
- guideline.md
- install_github_cli.sh
- convert_yi_vl.py
- convert_yi_vl.sh
- test_curl.sh
- test_flashinfer.py
- test_httpserver_concurrent.py
- test_httpserver_decode.py
- test_httpserver_decode_stream.py
- test_httpserver_llava.py
- test_httpserver_reuse.py
- test_jump_forward.py
- test_robust.py
- bump_kernel_version.py
- bump_sglang_version.py
- commit_and_pr.sh
- README.md
- test_utils.py
- utils.py
- check_vram_clear.sh
- ensure_vram_clear.sh
- export_deepseek_nextn.py
- killall_sglang.sh
- sort_testcases_alphabetically.py
- update_kernel_whl_index.py
- version_branch_to_tag.sh
- bench_activation.py
- bench_awq_dequant.py
- bench_cutlass_mla.py
- bench_dsv3_fused_a_gemm.py
- bench_dsv3_router_gemm.py
- bench_es_fp8_blockwise_grouped_gemm.py
- bench_fp4_gemm.py
- bench_fp8_blockwise_gemm.py
- bench_fp8_blockwise_group_gemm.py
- bench_fp8_gemm.py
- bench_int8_gemm.py
- bench_lightning_attention_decode.py
- bench_moe_align_block_size.py
- bench_moe_ep_post_reorder.py
- bench_moe_fused_gate.py
- bench_moe_topk_softmax.py
- bench_nvfp4_scaled_gemm.py
- bench_per_tensor_quant_fp8.py
- bench_per_token_group_quant_8bit.py
- bench_per_token_quant_fp8.py
- bench_qserve_w4a8_gemm.py
- bench_rmsnorm.py
- bench_rotary_embedding.py
- bench_sum_scale.py
- bench_top_k_top_p_sampling.py
- utils.cmake
- custom_all_reduce.cu
- custom_all_reduce.cuh
- custom_all_reduce.hip
- custom_all_reduce_hip.cuh
- mscclpp_allreduce.cu
- mscclpp_allreduce.cuh
- quick_all_reduce.cu
- quick_all_reduce.cuh
- quick_all_reduce.h
- quick_all_reduce_base.h
- test_mscclpp_allreduce.cu
- sm100_mla.hpp
- sm100_fmha_mla_reduction.hpp
- sm100_fmha_mla_tma_warpspecialized.hpp
- sm100_mla_tile_scheduler.hpp
- cascade.cu
- cutlass_mla_kernel.cu
- lightning_attention_decode_kernel.cu
- merge_attn_states.cu
- vertical_slash_index.cu
- activation.cpp
- bmm.cpp
- CMakeLists.txt
- common.h
- decode.cpp
- extend.cpp
- gemm.cpp
- gemm.h
- gemm_fp8.cpp
- gemm_int8.cpp
- interface.cpp
- moe.cpp
- moe_fp8.cpp
- moe_int8.cpp
- norm.cpp
- numa_utils.cpp
- qkv_proj.cpp
- rope.cpp
- shm.cpp
- shm.h
- topk.cpp
- torch_extension_cpu.cpp
- vec.h
- mixed_input_utils.hpp
- epilogue_per_row_per_col_scale.h
- sm90_gmma_builder_mixed_input.inl
- collective_builder_mixed_input.hpp
- collective_mma_array_mixed_input.hpp
- sm90_mma_array_tma_gmma_rs_warpspecialized_mixed_input_.hpp
- cutlass_gemm_caller.cuh
- dispatch_policy.hpp
- fp8_blockwise_gemm_sm90_dispatch.cuh
- gemm_universal_base_compat.h
- gemm_with_epilogue_visitor.h
- common.hpp
- activation.cu
- cast.cu
- concat_mla.cu
- copy.cu
- fused_add_rms_norm_kernel.cu
- pos_enc.cuh
- rope.cu
- topk.cu
- utils.cuh
- es_fp8_blockwise.cu
- es_fp8_blockwise_functor.cuh
- es_fp8_blockwise_launcher.cuh
- es_fp8_blockwise_traits.cuh
- compat.cuh
- gptq_kernel.cu
- matrix_view.cuh
- qdq_2.cuh
- qdq_3.cuh
- qdq_4.cuh
- qdq_8.cuh
- qdq_util.cuh
- awq_marlin_repack.cu
- dequant.h
- gptq_marlin.cu
- gptq_marlin_repack.cu
- kernel.h
- marlin.cuh
- marlin_dtypes.cuh
- marlin_template.h
- awq_kernel.cu
- bmm_fp8.cu
- dsv3_fused_a_gemm.cu
- dsv3_router_gemm_bf16_out.cu
- dsv3_router_gemm_entry.cu
- dsv3_router_gemm_float_out.cu
- fp8_blockwise_gemm_kernel.cu
- fp8_gemm_kernel.cu
- int8_gemm_kernel.cu
- math.hpp
- nvfp4_expert_quant.cu
- nvfp4_quant.cuh
- nvfp4_quant_entry.cu
- nvfp4_quant_kernels.cu
- nvfp4_scaled_mm_entry.cu
- nvfp4_scaled_mm_kernels.cu
- per_tensor_quant_fp8.cu
- per_token_group_quant_8bit.cu
- per_token_group_quant_8bit_v2.cu
- per_token_quant_fp8.cu
- qserve_w4a8_per_chn_gemm.cu
- qserve_w4a8_per_group_gemm.cu
- apply_token_bitmask_inplace_cuda.cu
- transfer.cu
- causal_conv1d.cu
- causal_conv1d.h
- store.cu
- scaled_mm_entry.cu
- w4a8_get_group_starts.cuh
- w4a8_grouped_mm_c3x.cu
- w4a8_grouped_mm_c3x.cuh
- w4a8_moe_data.cu
- generate_kernels.py
- kernel.h
- kernel_bf16_ku4.cuh
- kernel_bf16_ku4b8.cuh
- kernel_bf16_ku8b128.cuh
- kernel_fp16_ku4.cuh
- kernel_fp16_ku4b8.cuh
- kernel_fp16_ku8b128.cuh
- kernel_marlin.cuh
- marlin_template.h
- ops.cu
- cutlass_moe_helper.cu
- fp8_blockwise_moe_kernel.cu
- moe_align_kernel.cu
- moe_fused_gate.cu
- moe_sum.cu
- moe_sum_reduce.cu
- moe_topk_softmax_kernels.cu
- nvfp4_blockwise_moe.cu
- prepare_moe_input.cu
- dequantize.cuh
- ggml-common.h
- gguf_kernel.cu
- mmq.cuh
- mmvq.cuh
- moe.cuh
- moe_vec.cuh
- vecdotq.cuh
- cuda_utils.h
- greenctx_stream.cu
- greenctx_stream.h
- eagle_utils.cu
- ngram_utils.cu
- packbit.cu
- speculative_sampling.cu
- speculative_sampling.cuh
- common_extension.cc
- common_extension_rocm.cc
- flash_extension.cc
- spatial_extension.cc
- hip_vec_bf16_impl.h
- hip_vec_fp32_impl.h
- hip_vec_half_impl.h
- hip_act_and_mul.cuh
- hip_math_def.h
- hip_vec_dtypes.h
- pytorch_extension_utils_rocm.h
- scalar_type.hpp
- sgl_flash_kernel_ops.h
- sgl_kernel_ops.h
- sgl_kernel_torch_shim.h
- utils.h
- __init__.py
- gguf.py
- __init__.py
- rotary_embedding.py
- __init__.py
- _fa4_interface.py
- allreduce.py
- attention.py
- cutlass_moe.py
- elementwise.py
- expert_specialization.py
- flash_attn.py
- fused_moe.py
- gemm.py
- grammar.py
- hadamard.py
- kvcacheio.py
- mamba.py
- marlin.py
- memory.py
- moe.py
- sampling.py
- scalar_type.py
- sparse_flash_attn.py
- spatial.py
- speculative.py
- test_utils.py
- top_k.py
- utils.py
- version.py
- test_greenctx_stream.py
- test_eagle_utils.py
- test_ngram_utils.py
- test_speculative_sampling.py
- test_activation.py
- test_apply_token_bitmask_inplace.py
- test_awq_dequant.py
- test_bmm_fp8.py
- test_causal_conv1d.py
- test_custom_allreduce.py
- test_cutlass_mla.py
- test_cutlass_w4a8_moe_mm.py
- test_dsv3_fused_a_gemm.py
- test_dsv3_router_gemm.py
- test_es_fp8_blockwise_moe.py
- test_flash_attention.py
- test_flash_attention_4.py
- test_fp4_gemm.py
- test_fp4_quantize.py
- test_fp8_blockwise_gemm.py
- test_fp8_blockwise_moe.py
- test_fp8_gemm.py
- test_gguf.py
- test_gptq_kernel.py
- test_hadamard.py
- test_int8_gemm.py
- test_kvcacheio.py
- test_lightning_attention_decode.py
- test_marlin_gemm.py
- test_marlin_repack.py
- test_merge_state.py
- test_merge_state_v2.py
- test_moe_align.py
- test_moe_fused_gate.py
- test_moe_topk_softmax.py
- test_mscclpp.py
- test_norm.py
- test_per_tensor_quant_fp8.py
- test_per_token_group_quant_8bit.py
- test_per_token_quant_fp8.py
- test_qserve_w4a8_per_chn_gemm.py
- test_qserve_w4a8_per_group_gemm.py
- test_rotary_embedding.py
- test_sampling.py
- test_sparse_flash_attn.py
- test_topk.py
- utils.py
- .clang-format
- build.sh
- CMakeLists.txt
- LICENSE
- Makefile
- pyproject.toml
- pyproject_cpu.toml
- pyproject_rocm.toml
- README.md
- rename_wheels.sh
- setup_rocm.py
- THIRDPARTYNOTICES.txt
- config.toml
- request_processing.rs
- tokenizer_benchmark.rs
- tool_parser_benchmark.rs
- __init__.py
- launch_router.py
- launch_server.py
- mini_lb.py
- router.py
- router_args.py
- version.py
- conftest.py
- test_e2e_embeddings.py
- test_pd_router.py
- test_regular_router.py
- __init__.py
- mock_worker.py
- ports.py
- router_manager.py
- __init__.py
- test_cache_aware.py
- test_power_of_two.py
- test_random.py
- test_round_robin.py
- __init__.py
- conftest.py
- test_api_auth.py
- test_circuit_breaker.py
- test_fault_tolerance.py
- test_payload_size.py
- test_pd_routing.py
- test_rate_limiting.py
- test_retries.py
- test_service_discovery_shim.py
- test_worker_management.py
- __init__.py
- test_arg_parser.py
- test_router_config.py
- test_startup_sequence.py
- test_validation.py
- __init__.py
- conftest.py
- run_benchmarks.py
- setup-sccache.sh
- mod.rs
- types.rs
- validation.rs
- circuit_breaker.rs
- error.rs
- job_queue.rs
- mod.rs
- retry.rs
- token_bucket.rs
- worker.rs
- worker_builder.rs
- worker_manager.rs
- worker_registry.rs
- conversation_item_memory_store.rs
- conversation_item_oracle_store.rs
- conversation_items.rs
- conversation_memory_store.rs
- conversation_noop_store.rs
- conversation_oracle_store.rs
- conversations.rs
- mod.rs
- response_memory_store.rs
- response_noop_store.rs
- response_oracle_store.rs
- responses.rs
- mod.rs
- sglang_scheduler.rs
- client_manager.rs
- config.rs
- error.rs
- mod.rs
- oauth.rs
- cache_aware.rs
- factory.rs
- mod.rs
- power_of_two.rs
- random.rs
- registry.rs
- round_robin.rs
- sglang_scheduler.proto
- mod.rs
- spec.rs
- validated.rs
- worker_spec.rs
- base.rs
- deepseek_r1.rs
- glm45.rs
- kimi.rs
- mod.rs
- qwen3.rs
- step3.rs
- factory.rs
- mod.rs
- README.md
- traits.rs
- context.rs
- mod.rs
- pd_router.rs
- pipeline.rs
- processing.rs
- router.rs
- streaming.rs
- utils.rs
- mod.rs
- pd_router.rs
- pd_types.rs
- router.rs
- conversations.rs
- mcp.rs
- mod.rs
- responses.rs
- router.rs
- streaming.rs
- utils.rs
- factory.rs
- header_utils.rs
- mod.rs
- router_manager.rs
- chat_template.rs
- factory.rs
- hub.rs
- huggingface.rs
- mock.rs
- mod.rs
- README.md
- sequence.rs
- stop.rs
- stream.rs
- tests.rs
- tiktoken.rs
- traits.rs
- deepseek_parser.rs
- glm4_moe_parser.rs
- gpt_oss_harmony_parser.rs
- gpt_oss_parser.rs
- helpers.rs
- json_parser.rs
- kimik2_parser.rs
- llama_parser.rs
- mistral_parser.rs
- mod.rs
- passthrough_parser.rs
- pythonic_parser.rs
- qwen_parser.rs
- step3_parser.rs
- errors.rs
- factory.rs
- mod.rs
- partial_json.rs
- state.rs
- tests.rs
- traits.rs
- types.rs
- lib.rs
- logging.rs
- main.rs
- metrics.rs
- middleware.rs
- server.rs
- service_discovery.rs
- tree.rs
- mock_mcp_server.rs
- mock_openai_server.rs
- mock_worker.rs
- mod.rs
- streaming_helpers.rs
- test_app.rs
- chat_completion.rs
- chat_message.rs
- embedding.rs
- mod.rs
- rerank.rs
- api_endpoints_test.rs
- cache_aware_backward_compat_test.rs
- chat_template_format_detection.rs
- chat_template_integration.rs
- chat_template_loading.rs
- mcp_test.rs
- policy_registry_integration.rs
- request_formats_test.rs
- responses_api_test.rs
- streaming_tests.rs
- test_openai_routing.rs
- test_pd_routing.rs
- tokenizer_integration.rs
- tool_parser_deepseek.rs
- tool_parser_edge_cases.rs
- tool_parser_fallback.rs
- tool_parser_glm4_moe.rs
- tool_parser_gpt_oss.rs
- tool_parser_json.rs
- tool_parser_kimik2.rs
- tool_parser_llama.rs
- tool_parser_mistral.rs
- tool_parser_mixed_edge_cases.rs
- tool_parser_partial_json.rs
- tool_parser_pythonic.rs
- tool_parser_qwen.rs
- tool_parser_step3.rs
- tool_parser_streaming.rs
- .coveragerc
- build.rs
- Cargo.toml
- Makefile
- MANIFEST.in
- pyproject.toml
- README.md
- setup.py
- run_suite.py
- test_anthropic_backend.py
- test_bind_cache.py
- test_choices.py
- test_litellm_backend.py
- test_openai_backend.py
- test_separate_reasoning.py
- test_separate_reasoning_execution.py
- test_srt_backend.py
- test_tracing.py
- test_vertexai_backend.py
- test_ascend_deepep.py
- test_ascend_graph_tp1_bf16.py
- test_ascend_graph_tp2_bf16.py
- test_ascend_mla_fia_w8a8int8.py
- test_ascend_mla_w8a8int8.py
- test_ascend_tp1_bf16.py
- test_ascend_tp2_bf16.py
- test_ascend_tp2_fia_bf16.py
- test_ascend_tp4_bf16.py
- test_ascend_w8a8_quantization.py
- test_batch_invariant_ops.py
- deepseek_v3.yaml
- deepseek_v3_long_context.yaml
- llama_405b.yaml
- random_config.yaml
- random_flashinfer_vs_triton_config.yaml
- sharegpt_config.yaml
- test_activation.py
- test_binding.py
- test_cpu_graph.py
- test_decode.py
- test_extend.py
- test_gemm.py
- test_intel_amx_attention_backend_a.py
- test_intel_amx_attention_backend_b.py
- test_intel_amx_attention_backend_c.py
- test_mla.py
- test_moe.py
- test_norm.py
- test_qkv_proj_with_rope.py
- test_rope.py
- test_shared_expert.py
- test_topk.py
- utils.py
- test_abort_request.py
- test_deepep_internode.py
- test_deepep_intranode.py
- test_deepep_large.py
- test_deepep_low_latency.py
- test_deepep_small.py
- test_eplb.py
- test_hybrid_dp_ep_tp_mtp.py
- test_moe_deepep.py
- test_moe_deepep_eval_accuracy_large.py
- test_moe_ep.py
- test_mooncake_ep_small.py
- test_json_schema_constraint.py
- test_disaggregation_hicache.py
- test_hicache.py
- test_hicache_eagle.py
- test_hicache_mla.py
- test_hicache_page.py
- test_hicache_storage.py
- test_hicache_storage_3fs_backend.py
- test_hicache_storage_file_backend.py
- test_hicache_storage_mooncake_backend.py
- test_causal_conv1d.py
- test_mamba2_mixer.py
- test_mamba_ssm.py
- test_mamba_ssm_ssd.py
- test_act_quant_triton.py
- test_chunked_sgmv_backend.py
- test_lora.py
- test_lora_backend.py
- test_lora_cuda_graph.py
- test_lora_eviction.py
- test_lora_eviction_policy.py
- test_lora_llama4.py
- test_lora_qwen3.py
- test_lora_radix_cache.py
- test_lora_tp.py
- test_lora_update.py
- test_multi_lora_backend.py
- utils.py
- compare.py
- test_clip_models.py
- test_compressed_tensors_models.py
- test_cross_encoder_models.py
- test_dummy_grok_models.py
- test_embedding_models.py
- test_encoder_embedding_models.py
- test_falcon_h1_models.py
- test_generation_models.py
- test_glm4_moe_models.py
- test_gme_qwen_models.py
- test_grok_models.py
- test_llama4_models.py
- test_mtp_models.py
- test_nvidia_nemotron_nano_v2.py
- test_qwen3_next_models.py
- test_qwen_models.py
- test_reward_models.py
- test_transformers_models.py
- test_unsloth_models.py
- test_vlm_models.py
- __init__.py
- test_openai_embedding.py
- test_openai_server.py
- test_protocol.py
- test_serving_chat.py
- test_serving_completions.py
- test_serving_embedding.py
- __init__.py
- test_cache_report.py
- test_enable_thinking.py
- test_json_constrained.py
- test_json_mode.py
- test_openai_server_ebnf.py
- test_openai_server_hidden_states.py
- test_reasoning_content.py
- __init__.py
- test_openai_function_calling.py
- test_tool_choice.py
- __init__.py
- test_large_max_new_tokens.py
- test_matched_stop.py
- test_openai_server_ignore_eos.py
- test_request_length_validation.py
- __init__.py
- test_awq.py
- test_awq_dequant.py
- test_block_int8.py
- test_fp8_kernel.py
- test_fp8_kvcache.py
- test_int8_kernel.py
- test_triton_scaled_mm.py
- test_w4a8_deepseek_v3.py
- test_w8a8_quantization.py
- test_fp32_lm_head.py
- test_update_weights_from_disk.py
- test_update_weights_from_distributed.py
- test_update_weights_from_tensor.py
- test_verl_engine_2_gpu.py
- test_verl_engine_4_gpu.py
- test_intel_xpu_backend.py
- double-sparsity-config-Llama-3.1-8B-Instruct.json
- experiment_runner.py
- kv_cache_scales_llama3_1_8b.json
- kv_cache_scales_llama3_8b.json
- kv_cache_scales_qwen2_1_5b.json
- parse_results.py
- run_suite.py
- test_abort.py
- test_async_dynamic_batch_tokenizer.py
- test_bench_one_batch.py
- test_bench_serving.py
- test_bnb.py
- test_build_eagle_tree.py
- test_chunked_prefill.py
- test_config_integration.py
- test_cpp_radix_cache.py
- test_create_kvindices.py
- test_custom_allreduce.py
- test_cutedsl_flashinfer_8gpu.py
- test_data_parallelism.py
- test_deepseek_v32_basic.py
- test_deepseek_v3_basic.py
- test_deepseek_v3_fp4_4gpu.py
- test_deepseek_v3_mtp.py
- test_deterministic.py
- test_disaggregation_basic.py
- test_disaggregation_different_tp.py
- test_disaggregation_dp_attention.py
- test_disaggregation_hybrid_attention.py
- test_disaggregation_pp.py
- test_double_sparsity.py
- test_dp_attention.py
- test_eagle_infer_a.py
- test_eagle_infer_b.py
- test_eagle_infer_beta.py
- test_eagle_mab.py
- test_ebnf_constrained.py
- test_eval_accuracy_large.py
- test_eval_fp8_accuracy.py
- test_expert_distribution.py
- test_expert_location_updater.py
- test_fa3.py
- test_fim_completion.py
- test_flashmla.py
- test_forward_split_prefill.py
- test_function_call_parser.py
- test_fused_moe.py
- test_get_weights_by_name.py
- test_gguf.py
- test_gpt_oss_1gpu.py
- test_gpt_oss_4gpu.py
- test_gpt_oss_common.py
- test_gptqmodel_dynamic.py
- test_harmony_parser.py
- test_health_check.py
- test_hidden_states.py
- test_hybrid_attn_backend.py
- test_input_embeddings.py
- test_io_struct.py
- test_jinja_template_utils.py
- test_kv_events.py
- test_load_weights_from_remote_instance.py
- test_local_attn.py
- test_logprobs.py
- test_mamba_unittest.py
- test_metrics.py
- test_metrics_utils.py
- test_mla.py
- test_mla_deepseek_v3.py
- test_mla_flashinfer.py
- test_mla_fp8.py
- test_mla_int8_deepseek_v3.py
- test_mla_tp.py
- test_modelopt.py
- test_modelopt_fp8kvcache.py
- test_modelopt_loader.py
- test_models_from_modelscope.py
- test_moe_eval_accuracy_large.py
- test_mscclpp.py
- test_multi_instance_release_memory_occupation.py
- test_multi_tokenizer.py
- test_ngram_speculative_decoding.py
- test_nightly_gsm8k_eval_amd.py
- test_nightly_text_models_gsm8k_eval.py
- test_nightly_text_models_perf.py
- test_nightly_vlms_mmmu_eval.py
- test_nightly_vlms_perf.py
- test_no_chunked_prefill.py
- test_no_overlap_scheduler.py
- test_original_logprobs.py
- test_page_size.py
- test_patch_torch.py
- test_penalty.py
- test_piecewise_cuda_graph.py
- test_pp_single_node.py
- test_priority_scheduling.py
- test_profile_merger.py
- test_profile_merger_http_api.py
- test_pytorch_sampling_backend.py
- test_quick_allreduce.py
- test_radix_attention.py
- test_radix_cache_unit.py
- test_reasoning_parser.py
- test_regex_constrained.py
- test_release_memory_occupation.py
- test_request_queue_validation.py
- test_retract_decode.py
- test_rope_rocm.py
- test_sagemaker_server.py
- test_schedule_policy.py
- test_score_api.py
- test_server_args.py
- test_session_control.py
- test_skip_tokenizer_init.py
- test_srt_endpoint.py
- test_srt_engine.py
- test_srt_engine_with_quant_args.py
- test_standalone_speculative_decoding.py
- test_start_profile.py
- test_swa_unittest.py
- test_tokenizer_batch_encode.py
- test_tokenizer_manager.py
- test_torch_compile.py
- test_torch_compile_moe.py
- test_torch_flex_attention_backend.py
- test_torch_native_attention_backend.py
- test_torch_tp.py
- test_torchao.py
- test_tracing.py
- test_triton_attention_backend.py
- test_triton_attention_kernels.py
- test_triton_attention_rocm_mla.py
- test_triton_fused_moe.py
- test_triton_moe_channel_fp8_kernel.py
- test_triton_moe_wna16.py
- test_triton_sliding_window.py
- test_two_batch_overlap.py
- test_utils_update_weights.py
- test_vertex_endpoint.py
- test_vision_chunked_prefill.py
- test_vision_openai_server_a.py
- test_vision_openai_server_b.py
- test_vision_openai_server_common.py
- test_vllm_dependency.py
- test_vlm_accuracy.py
- test_vlm_input_format.py
- test_wave_attention_backend.py
- test_wave_attention_kernels.py
- test_weight_version.py
- README.md
- .clang-format-ignore
- .editorconfig
- .gitignore
- .isort.cfg
- .pre-commit-config.yaml
- CODE_OF_CONDUCT.md
- LICENSE
- Makefile
- package-lock.json
- README.md
- __init__.py
- agent_loop.py
- single_turn_agent_loop.py
- tool_agent_loop.py
- tool_parser.py
- __init__.py
- sampler.py
- __init__.py
- dynamicgen_dataset.py
- __init__.py
- __init__.py
- interaction_registry.py
- __init__.py
- base.py
- gsm8k_interaction.py
- __init__.py
- __main__.py
- base_model_merger.py
- fsdp_model_merger.py
- megatron_model_merger.py
- __init__.py
- llama_loader.py
- llama_loader_depracated.py
- llama_saver.py
- __init__.py
- parallel_attention.py
- parallel_decoder.py
- parallel_linear.py
- parallel_mlp.py
- parallel_rmsnorm.py
- __init__.py
- modeling_llama_megatron.py
- __init__.py
- __init__.py
- attention.py
- model.py
- rope_utils.py
- vision_config.py
- vision_model.py
- vision_transformer_block.py
- __init__.py
- config_converter.py
- loader.py
- mbridge.py
- model_forward.py
- model_forward_fused.py
- model_initializer.py
- patch_v012.py
- readme.md
- registry.py
- saver.py
- util.py
- weight_converter.py
- __init__.py
- qwen2_loader.py
- qwen2_loader_depracated.py
- qwen2_saver.py
- __init__.py
- parallel_attention.py
- parallel_decoder.py
- parallel_linear.py
- parallel_mlp.py
- parallel_rmsnorm.py
- __init__.py
- modeling_qwen2_megatron.py
- __init__.py
- __init__.py
- dense_common.py
- kimi_vl.py
- llama.py
- monkey_patch.py
- npu_patch.py
- qwen2.py
- qwen2_5_vl.py
- qwen2_vl.py
- __init__.py
- README.md
- registry.py
- weight_loader_registry.py
- __init__.py
- ray.py
- __init__.py
- decorator.py
- worker.py
- worker_group.py
- __init__.py
- base.py
- __init__.py
- __init__.py
- parallel_state.py
- __init__.py
- state_dict.py
- __init__.py
- _state_dict_utils.py
- __init__.py
- __init__.py
- __init__.py
- McpClientManager.py
- utils.py
- __init__.py
- search_r1_like_utils.py
- tool_registry.py
- __init__.py
- base_tool.py
- geo3k_tool.py
- gsm8k_tool.py
- mcp_base_tool.py
- mcp_search_tool.py
- sandbox_fusion_tools.py
- schemas.py
- search_tool.py
- actor.yaml
- dp_actor.yaml
- megatron_actor.yaml
- critic.yaml
- dp_critic.yaml
- megatron_critic.yaml
- legacy_data.yaml
- npu_profile.yaml
- dp_ref.yaml
- megatron_ref.yaml
- ref.yaml
- dp_reward_model.yaml
- megatron_reward_model.yaml
- reward_model.yaml
- rollout.yaml
- __init__.py
- _generated_ppo_megatron_trainer.yaml
- _generated_ppo_trainer.yaml
- algorithm.py
- config.py
- evaluation.yaml
- fastrl_trainer.yaml
- generation.yaml
- ppo_megatron_trainer.yaml
- ppo_trainer.yaml
- sft_trainer.yaml
- __init__.py
- core_algos.py
- metric_utils.py
- ray_trainer.py
- reward.py
- __init__.py
- constants_ppo.py
- fsdp_sft_trainer.py
- main_eval.py
- main_fastrl.py
- main_generation.py
- main_ppo.py
- runtime_env.yaml
- __init__.py
- checkpoint_manager.py
- fsdp_checkpoint_manager.py
- megatron_checkpoint_manager.py
- __init__.py
- multiturn_sft_dataset.py
- README.md
- rl_dataset.py
- rm_dataset.py
- sft_dataset.py
- vision_utils.py
- __init__.py
- metrics.py
- performance.py
- trajectory_tracker.py
- __init__.py
- torch_functional.py
- __init__.py
- kernels.py
- linear_cross_entropy.py
- __init__.py
- aggregate_logger.py
- __init__.py
- dist_checkpointing.py
- memory.py
- optimizer.py
- pipeline_parallel.py
- sequence_parallel.py
- tensor_parallel.py
- __init__.py
- utils.py
- __init__.py
- config.py
- empty_annotations.py
- mstx_profile.py
- nvtx_profile.py
- performance.py
- profile.py
- __init__.py
- ray_backend.py
- __init__.py
- README.md
- testing_util.py
- utils.py
- __init__.py
- grader.py
- math_normalize.py
- __init__.py
- utils.py
- __init__.py
- geo3k.py
- gsm8k.py
- math.py
- math_batch.py
- math_dapo.py
- math_verify.py
- search_r1_like_qa_em.py
- __init__.py
- patch.py
- utils.py
- __init__.py
- activation_offload.py
- config.py
- data_buffer.py
- device.py
- distributed.py
- flops_counter.py
- fs.py
- fsdp_utils.py
- hdfs_io.py
- import_utils.py
- logging_utils.py
- megatron_utils.py
- memory_buffer.py
- memory_utils.py
- model.py
- net_utils.py
- py_functional.py
- pyext.py
- ray_utils.py
- rollout_skip.py
- rollout_trace.py
- seqlen_balancing.py
- tokenizer.py
- torch_dtypes.py
- torch_functional.py
- tracking.py
- ulysses.py
- version
- __init__.py
- base.py
- dp_actor.py
- megatron_actor.py
- __init__.py
- actor.py
- critic.py
- engine.py
- optimizer.py
- __init__.py
- base.py
- dp_critic.py
- megatron_critic.py
- __init__.py
- llama_eagle.py
- llama_eagle3.py
- qwen2_eagle.py
- qwen2_eagle3.py
- qwen3_eagle3.py
- __init__.py
- eagle_background_trainer.py
- __init__.py
- engine_impl.py
- utils.py
- __init__.py
- engine_impl.py
- __init__.py
- base.py
- __init__.py
- abstract.py
- batch.py
- dapo.py
- naive.py
- prime.py
- registry.py
- __init__.py
- reward_model.py
- __init__.py
- base.py
- __init__.py
- actor.py
- critic.py
- __init__.py
- async_server.py
- base.py
- __init__.py
- fsdp_workers.py
- megatron_workers.py
- __init__.py
- base_config.py
- protocol.py
- .gitignore
- .pre-commit-config.yaml
- LICENSE
- README.md
- requirements.txt
- setup.py
# Installation Guide
git clone https://github.com/mit-han-lab/fastrl
Downloads the entire project code from GitHub to your computer.
cd fastrl
Moves into the project folder you just downloaded.
2. Official Install Script
Easy Recommended- Python 3 Python is required to use pip.
pip install https://github.com/Dao-AILab/flash-attention/releases/download/v2.8.3/flash_attn-2.8.3+cu12torch2.8cxx11abiTRUE-cp312-cp312-linux_x86_64.whl
Installs the package published on PyPI directly โ no need to clone the source.
Pulled directly from this repo's README.
3. Docker
Easy- Git Needed to download the project code from GitHub.
- Docker Desktop Needed to build and run containers. Install it and keep it running in the background.
docker compose -f third-party/sglang/docker/compose.yaml up -d --build
Runs the command against the services defined in the compose file.
4. CMake
Mediumcd third-party/sglang/sgl-kernel
This project's files live in a subfolder, so move into it first.
mkdir build && cd build
Creates a folder to hold the build output and moves into it.
cmake ..
Analyzes the source code and generates build configuration files (must be run inside the build folder).
make
Compiles the code based on the generated build configuration to produce an executable.
5. Python
Easypip install -e "python[all]"
Installs the Python libraries listed in requirements.txt (or similar).
pip install https://github.com/Dao-AILab/flash-attention/releases/download/v2.8.3/flash_attn-2.8.3+cu12torch2.8cxx11abiTRUE-cp312-cp312-linux_x86_64.whl
Installs the package published on PyPI directly โ no need to clone the source.
pip install -e .
Installs the Python libraries listed in requirements.txt (or similar).
Pulled directly from this repo's README.
6. Rust
Medium- Git Needed to download the project code from GitHub.
- Rust (rustup) Installing via rustup also installs cargo.
cd third-party/sglang/sgl-router
This project's files live in a subfolder, so move into it first.
cargo build --release
Compiles the Rust project.
cargo run
Builds and then immediately runs the program.
7. Make
Medium- Git Needed to download the project code from GitHub.
- Make Usually pre-installed on Linux/macOS. On Windows, install separately (e.g. via MSYS2 or WSL).
cd third-party/sglang
This project's files live in a subfolder, so move into it first.
make
Compiles the code based on the generated build configuration to produce an executable.
