Sim2Reason
Sim2Reason: Solving Physics Olympiad via Reinforcement Learning on Physics Simulators. We present a method for turning physics simulators into scalable generators of question–answer pairs for improving LLM physical reasoning.
파일 탐색기
최종 버전 다운로드 (.zip)- collision.yaml
- em.yaml
- orbital.yaml
- pulley.yaml
- rotation.yaml
- spring.yaml
- all.yaml
- celestial.yaml
- collision.yaml
- pulley.yaml
- rotation.yaml
- config.yaml
- auto_scoring_judge.py
- scoring_examples.json
- check_prompt.py
- custom_llm_evaluator.py
- deepseek_evaluator.py
- evaluator.py
- gemini_pro_vision.py
- gpt_4v.py
- qwen_vl.py
- text_only_gpt_4.py
- evaluate_all.py
- math_judger.py
- api_requirements.txt
- calculate_accuracy.py
- judge.py
- readme.md
- LICENSE
- README.md
- general_info.txt
- verify.txt
- basic_utils.py
- cost.json
- math_equivalence.py
- math_utils.py
- unicode_to_latex.py
- __init__.py
- preprocess_json_to_parquet.py
- __init__.py
- contact_utils.py
- recorder.py
- utils.py
- __init__.py
- base_bodies.py
- composed_bodies.py
- friction_bodies.py
- geom_bodies.py
- mass.py
- plane.py
- pulley_bodies.py
- rotation_bodies.py
- spring_bodies.py
- __init__.py
- base_entities.py
- collision_entities.py
- magnetic_electric_entities.py
- mass_entities.py
- orbital_motion_entities.py
- plane_entities.py
- pulley_entities.py
- rotation_entities.py
- throwing_motion_entities.py
- rocket.obj1.mtl
- rocket.obj1.obj
- rocket.obj10.mtl
- rocket.obj10.obj
- rocket.obj11.mtl
- rocket.obj11.obj
- rocket.obj12.mtl
- rocket.obj12.obj
- rocket.obj13.mtl
- rocket.obj13.obj
- rocket.obj2.mtl
- rocket.obj2.obj
- rocket.obj3.mtl
- rocket.obj3.obj
- rocket.obj4.mtl
- rocket.obj4.obj
- rocket.obj5.mtl
- rocket.obj5.obj
- rocket.obj6.mtl
- rocket.obj6.obj
- rocket.obj7.mtl
- rocket.obj7.obj
- rocket.obj8.mtl
- rocket.obj8.obj
- rocket.obj9.mtl
- rocket.obj9.obj
- rocket.white-texture-background.jpg
- rocket.white-texture-background.png
- polygonal_cylinder.stl
- rocket.mtl
- rocket.obj
- round_arch.stl
- round_arch_old.stl
- round_bowl.stl
- check_normal.py
- main.py
- main.txt
- main.xml
- main.yaml
- main.py
- main.txt
- main.xml
- main.yaml
- main.py
- main.txt
- main.xml
- main.yaml
- main.py
- main.txt
- main.xml
- main.yaml
- main.py
- main.txt
- main.xml
- main.yaml
- main.py
- main.txt
- main.xml
- main.yaml
- main.py
- main.txt
- main.xml
- main.yaml
- main.py
- main.txt
- main.xml
- main.yaml
- main.py
- main.txt
- main.xml
- main.yaml
- main.py
- main.txt
- main.xml
- main.yaml
- main.py
- main.txt
- main.xml
- main.yaml
- main.py
- main.txt
- main.xml
- main.yaml
- main.py
- main.txt
- main.xml
- main.yaml
- main.py
- main.txt
- main.xml
- main.yaml
- main.py
- main.txt
- main.xml
- main.yaml
- main.py
- main.txt
- main.xml
- main.yaml
- main.py
- main.txt
- main.xml
- main.yaml
- main.py
- main.txt
- main.xml
- main.yaml
- main.py
- main.txt
- main.xml
- main.yaml
- main.py
- main.txt
- main.xml
- main.yaml
- times.ttf
- expected_numerical_question.txt
- expected_symbolic_question.txt
- scene.yaml
- expected_numerical_question.txt
- expected_symbolic_question.txt
- scene.yaml
- expected_numerical_question.txt
- expected_symbolic_question.txt
- scene.yaml
- expected_numerical_question.txt
- expected_symbolic_question.txt
- scene.yaml
- expected_numerical_question.txt
- expected_symbolic_question.txt
- scene.yaml
- expected_numerical_question.txt
- expected_symbolic_question.txt
- gt.py
- scene.yaml
- expected_numerical_question.txt
- expected_symbolic_question.txt
- scene.yaml
- expected_numerical_question.txt
- expected_symbolic_question.txt
- gt.py
- scene.yaml
- gt.py
- scene.yaml
- test.py
- __init__.py
- constants.py
- create_child_scenes.py
- geometry_utils.py
- logger_manager.py
- mesh_utils.py
- objects.py
- qa_gen_rule.py
- scene.py
- scene_generator.py
- utils.py
- write_json.py
- xml_body_unpacker.py
- config.yaml
- e2e_ppo_trainer.yml
- e2e_ppo_trainer_megatron_sglang.yml
- e2e_prime.yml
- e2e_prime.yml
- check-pr-title.yml
- checkpoint_converter.yml
- cpu_unit_tests.yml
- doc.yml
- e2e_ascend.yml
- e2e_dapo.yml
- e2e_eval_aime24.yml
- e2e_genrm_remote.yml
- e2e_one_step_off_policy.yml
- e2e_ppo_trainer.yml
- e2e_ppo_trainer_megatron.yml
- e2e_ppo_trainer_megatron_sglang.yml
- e2e_ppo_trainer_megatron_vllm.yml
- e2e_sft.yml
- e2e_spin.yml
- e2e_sppo.yml
- gpu_unit_tests.yml
- model.yml
- pre-commit-full.yml
- pre-commit.yml
- README.md
- sanity.yml
- scorecard.yml
- secrets_scan.yml
- sgl.yml
- type-coverage-check.yml
- vllm.yml
- CODEOWNERS
- dependabot.yml
- PULL_REQUEST_TEMPLATE.md
- Dockerfile.app.sglang.vllm.mcore0.12
- Dockerfile.app.sglang.vllm.mcore0.12.deepep
- Dockerfile.app.sglang.vllm.mcore0.13.preview
- Dockerfile.app.vllm.mcore0.12
- Dockerfile.app.vllm.mcore0.12.deepep
- Dockerfile.app.vllm.mcore0.13.preview
- Dockerfile.base
- README.md
- Dockerfile.app.sglang.mcore0.12
- Dockerfile.app.sglang0.4.10.post2.mcore0.13
- Dockerfile.app.sglang0.4.9.post6.mcore0.13
- Dockerfile.app.vllm.mcore0.12
- Dockerfile.app.vllm.mcore0.13
- Dockerfile.base.torch2.7.0
- Dockerfile.base.torch2.7.1
- README.md
- Dockerfile.app.sglang.mcore0.12
- Dockerfile.app.sglang.mcore0.13.preview
- Dockerfile.base
- README.md
- Dockerfile.app.sglang.megatron
- Dockerfile.base
- README.md
- Apptainerfile.rocm
- Dockerfile.awsefa
- Dockerfile.extention.awsefa
- Dockerfile.ngc.vllm
- Dockerfile.ngc.vllm0.8
- Dockerfile.ngc.vllm0.8.sagemaker
- Dockerfile.rocm
- Dockerfile.rocm_verl-0.3.0.post1
- Dockerfile.rocm_verl-0.4.1
- Dockerfile.sglang
- Dockerfile.vemlp.vllm.te
- Dockerfile.vllm.sglang.megatron
- Dockerfile.vllm.sglang.megatron.deepseek
- README.md
- resizable-sidebar.js
- runllm-widget.js
- custom.css
- agent_loop.rst
- checkpoint.rst
- dpo_extension.rst
- fsdp_extension.rst
- megatron_extension.rst
- one_step_off.md
- placement.rst
- ppo_lora.rst
- rollout_skip.rst
- rollout_trace.rst
- rope.rst
- baseline.md
- dapo.md
- entropy.md
- gpg.md
- grpo.md
- opo.md
- ppo.md
- spin.md
- sppo.md
- amd_build_dockerfile_page.rst
- amd_vllm_page.rst
- data.rst
- single_controller.rst
- trainer.rst
- utils.rst
- ascend_profiling.rst
- ascend_profiling_en.rst
- ascend_profiling_zh.rst
- ascend_quick_start.rst
- config.rst
- gsm8k_example.rst
- multi_modal_example.rst
- ppo_code_architecture.rst
- sandbox_fusion_example.rst
- faq.rst
- device_tuning.rst
- dpsk.md
- nsight_profiling.md
- perf_tuning.rst
- verl_profiler_system.md
- prepare_data.rst
- reward_function.rst
- interaction_system.rst
- multiturn.rst
- sandbox_fusion.rst
- search_tool_example.rst
- agentic_rl.rst
- install.rst
- more_resources.rst
- multinode.rst
- quickstart.rst
- ray_debug_tutorial.rst
- fsdp_workers.rst
- megatron_workers.rst
- ray_trainer.rst
- sglang_worker.rst
- conf.py
- hybrid_flow.rst
- index.rst
- Makefile
- README.md
- README_vllm0.7.md
- README_vllm0.8.md
- requirements-docs.txt
- single_controller.rst
- aime2024_multiturn_w_tool.py
- aime24.py
- dapo_multiturn_w_tool.py
- full_hh_rlhf.py
- geo3k.py
- geo3k_multiturn_w_tool.py
- gsm8k.py
- gsm8k_multiturn_w_interaction.py
- gsm8k_multiturn_w_tool.py
- gsm8k_tool_agent_loop.py
- hellaswag.py
- math_dataset.py
- multiturn.py
- preprocess_search_r1_dataset.py
- run_deepseek7b_mutli_node.sh
- run_deepseek_v2_lite_math.sh
- README.md
- run_qwen2_5-7b_math.sh
- test_dapo_7b_math.sh
- test_dapo_qwen3_30b_math.sh
- gpg.md
- run_qwen2-7b_math.sh
- run_qwen2-7b_math_megatron.sh
- README.md
- run_deepseek671b_math_megatron.sh
- run_deepseek671b_math_megatron_80gb.sh
- run_deepseek671b_math_megatron_96gb.sh
- run_deepseek7b_llm.sh
- run_deepseek7b_llm_math.sh
- run_deepseek7b_llm_math_megatron.sh
- run_deepseek7b_llm_seq_balance.sh
- run_minicpmo2_6.sh
- run_moonlight16b_math_megatron.sh
- run_qwen2-7b.sh
- run_qwen2-7b_math.sh
- run_qwen2-7b_math_megatron.sh
- run_qwen2-7b_seq_balance.sh
- run_qwen2-7b_seq_balance_math_megatron.sh
- run_qwen2-7b_sgl_megatron.sh
- run_qwen2_5-3b_gsm8k_grpo_lora.sh
- run_qwen2_5-7b_math_megatron_diff_tp.sh
- run_qwen2_5_32b_grpo_npu.sh
- run_qwen2_5_7b_grpo_discrete_prof_npu.sh
- run_qwen2_5_7b_grpo_e2e_prof_npu.sh
- run_qwen2_5_7b_grpo_npu.sh
- run_qwen2_5_vl-7b-megatron.sh
- run_qwen2_5_vl-7b-sglang.sh
- run_qwen2_5_vl-7b.sh
- run_qwen2_5_vl-7b_lora.sh
- run_qwen2_5_vl-7b_seq_balance.sh
- run_qwen2_5_vl_32b_npu.sh
- run_qwen2_5_vl_3b_npu.sh
- run_qwen2_5_vl_7b_npu.sh
- run_qwen3-235b_megatron_96gb.sh
- run_qwen3-236b_megatron.sh
- run_qwen3-8b.sh
- run_qwen3moe-30b_megatron.sh
- run_qwen3moe-30b_megatron_96gb.sh
- README.md
- run_deepseek7b_llm.sh
- run_deepseek7b_llm_modelscope.sh
- run_deepseek7b_llm_pfppo.sh
- run_deepseek7b_llm_sandbox_fusion.sh
- run_deepseek7b_llm_sp2.sh
- run_deepseek_full_hh_rlhf.sh
- run_deepseek_math_gsm8k_megatron.sh
- run_deepseek_math_gsm8k_megatron_nsys.sh
- run_gemma.sh
- run_moonlight16b_a3b_gsm8k_megatron.sh
- run_qwen1.5_moe_a2.7b-gsm8k_megatron.sh
- run_qwen2-7b_math_gsm8k_megatron.sh
- run_qwen2-7b_rm.sh
- run_qwen2-7b_rm_seq_balance.sh
- run_qwen2-7b_rm_seq_balance_fused_kernels.sh
- run_qwen2-7b_rm_seq_balance_nsys.sh
- run_qwen2-7b_seq_balance.sh
- run_qwen2-7b_sglang_seq_balance.sh
- run_qwen2.5-32b.sh
- tutorial.ipynb
- run_qwen2-7b_math_rf.sh
- run_qwen2-7b_math_rf_baseline.sh
- run_qwen2.5-3b_seq_balance.sh
- run_qwen2.5-7b_seq_balance.sh
- run_qwen2-7b.sh
- run_deepseek_6b7.sh
- run_gemma_2b.sh
- run_gemma_7b.sh
- run_qwen2_5_05b_sft_peft_sp2_npu.sh
- run_qwen3_8b_sft_peft_sp2_npu.sh
- run_qwen_05_peft.sh
- run_qwen_05_sp2.sh
- run_qwen_05_sp2_liger.sh
- run_qwen_05_sp2.sh
- gsm8k_interaction_config.yaml
- geo3k_tool_config.yaml
- gsm8k_tool_config.yaml
- mcp_server.json
- mcp_tool_config.yaml
- sandbox_fusion_tool_config.yaml
- search_tool_config.yaml
- geo3k_multiturn_grpo.yaml
- geo3k_multiturn_megatron_grpo.yaml
- gsm8k_multiturn_grpo.yaml
- gsm8k_multiturn_grpo_server.yaml
- gsm8k_multiturn_grpo_w_interaction.yaml
- gsm8k_multiturn_megatron_grpo.yaml
- retool_multiturn_grpo.yaml
- search_multiturn_grpo.yaml
- run_qwen2.5-3b_geo3k_multiturn.sh
- run_qwen2.5-3b_geo3k_multiturn_4xgpu.sh
- run_qwen2.5-3b_megatron_geo3k_multiturn.sh
- download.py
- retrieval_server.py
- run_qwen2.5-3b_instruct_search_multiturn.sh
- README.md
- run_qwen0.5b_gsm8k_multiturn_curriculum.sh
- run_qwen2.5-0.5b_gsm8k_multiturn_w_interaction.sh
- run_qwen2.5-3b_gsm8k_multiturn.sh
- run_qwen2.5-3b_gsm8k_multiturn_4xgpu.sh
- run_qwen2.5-3b_gsm8k_multiturn_4xgpu_server.sh
- run_qwen2.5-3b_gsm8k_multiturn_server.sh
- run_qwen2.5-3b_gsm8k_tool_agent_mlflow.sh
- run_qwen2.5-3b_megatron_gsm8k_multiturn.sh
- run_qwen3-4b_gsm8k_multiturn.sh
- run_qwen3_4b_dapo_multiturn.sh
- ray_on_slurm.slurm
- ppo_trainer_split.yaml
- main_ppo_split.py
- README.md
- run_deepseek7b_llm.sh
- split_monkey_patch.py
- qwen2-0.5b_grpo-lora_1_h100_fsdp_vllm.sh
- qwen2-1.5b_grpo-lora_1_h100_fsdp_vllm.sh
- qwen2-14b_grpo-lora_2_h100_fsdp_vllm.sh
- qwen2_14b_grpo_4_h800_fsdp_vllm.sh
- qwen2-32b_grpo-lora_4_h100_fsdp_vllm.sh
- qwen2_32B_grpo_8_h20_megatron_vllm.sh
- qwen2-3b_grpo-lora_1_h100_fsdp_vllm.sh
- qwen2-70b_grpo_32_h20_fsdp_vllm.sh
- qwen2-70b_grpo_32_h800_fsdp_vllm.sh
- qwen2-72b_grpo-lora_8_h100_fsdp_vllm.sh
- qwen2-7b_grpo-lora_1_h100_fsdp_vllm.sh
- qwen2-7b_grpo_2_h800_fsdp_vllm.sh
- create_dataset.py
- README.md
- reward_function.py
- train_grpo.sh
- train_sft.sh
- dapo_trainer.yaml
- dapo_ray_trainer.py
- JEEBench.py
- main_dapo.py
- olympiad_bench.py
- PHYSICS.py
- README.md
- run_dapo_3b_m.sh
- run_dapo_early_qwen2.5_32b.sh
- run_dapo_qwen2.5_32b.sh
- run_dapo_qwen2.5_32b_npu.sh
- run_dapo_qwen2.5_32b_tis.sh
- run_dapo_qwen3_14b_base_npu.sh
- run_dapo_qwen3_8b_base_npu.sh
- run_dapo_qwen3_moe_30b_base_npu_fsdp.sh
- run_dapo_qwen3_moe_30b_megatron_npu.sh
- run_dapo_wo_ds_qwen2.5_32b.sh
- runtime_env.yaml
- simple_eval.py
- test_dapo_7b.sh
- test_dapo_7b_debug.sh
- test_dapo_7b_math.sh
- test_dapo_7b_math_lora.sh
- test_dapo_7b_math_megatron.sh
- test_dapo_dspk_671b_megatron.sh
- test_dapo_dspk_671b_megatron_96gb.sh
- test_dapo_qwen3_30b_math.sh
- test_dapo_qwen3_30b_math_single_node.sh
- deepeyes_multiturn_grpo.yaml
- image_zoom_in_tool_config.yaml
- deepeyes.py
- README.md
- run_deepeyes_grpo.sh
- entropy_trainer.yaml
- __init__.py
- grader.py
- math_normalize.py
- __init__.py
- 32b_clip_cov.sh
- 32b_kl_cov.sh
- 32b_kl_cov_mininbsz.sh
- 7b_clip_cov.sh
- 7b_kl_cov.sh
- entropy_ray_trainer.py
- main_entropy.py
- README.md
- reward.py
- README.md
- reward_function.py
- run_genrm_remote.sh
- test_gspo_3b_math.sh
- test_gspo_3b_math_slurm.sh
- README.md
- reward_fn.py
- run_3b.sh
- run_7b.sh
- agent.yaml
- create_dataset.py
- math_expression.py
- README.md
- run_qwen2.5_3b.sh
- __init__.py
- chat_model.py
- react_agent_loop.py
- test_react_agent_loop.py
- rl_dataset.py
- one_step_off_ppo_megatron_trainer.yaml
- one_step_off_ppo_trainer.yaml
- dapo_7b_math_fsdp2_4_12.sh
- dapo_7b_math_fsdp2_colocate.sh
- dapo_7b_math_megatron_4_12.sh
- dapo_7b_math_megatron_colocate.sh
- fsdp_workers.py
- grpo_0.6b_gsm8k_fsdp2_2_6.sh
- grpo_3b_gsm8k_fsdp2_2_6.sh
- main_ppo.py
- megatron_workers.py
- ray_trainer.py
- README.md
- utils.py
- vllm_sharding_manager.py
- prime_trainer.yaml
- __init__.py
- main_prime.py
- prime_core_algos.py
- prime_dp_rm.py
- prime_fsdp_workers.py
- prime_ray_trainer.py
- run_prime_qwen.sh
- run_prime_qwen_code.sh
- evaluation.yaml
- __init__.py
- gpqa.py
- livecodebench.py
- math.py
- __init__.py
- data_process.py
- main_eval.py
- README.md
- reward_score.py
- run_r1_distill_qwen.sh
- README.md
- retool.py
- retool_multi_turn_sft_preprocess.py
- retool_sft_preprocess.py
- run_qwen2-32b_dapo.sh
- run_qwen2-32b_ppo.sh
- run_qwen2-32b_sft.sh
- run_qwen2.5_32b_sp8.sh
- run_qwen2.5_7b_sp4.sh
- run_qwen2_7b_dapo.sh
- run_qwen2_7b_sft.sh
- run_qwen2_7b_sft_npu.sh
- run_qwen3_4b_sp4.sh
- run_sft.sh
- sandbox_fusion_tool_config.yaml
- spin_trainer.yaml
- core_algos.py
- dp_actor.py
- fsdp_workers.py
- main_spin.py
- README.md
- run_spin.sh
- spin_trainer.py
- utils.py
- sppo_trainer.yaml
- __init__.py
- config.py
- dp_actor.py
- main_sppo.py
- README.md
- run_qwen2.5-7b_rm.sh
- sppo_ray_trainer.py
- sppo_worker.py
- README.md
- __init__.py
- converter_hf_to_mcore.py
- diagnose.py
- generate_trainer_config.sh
- init_random_model.py
- install_vllm_sglang_mcore.sh
- legacy_model_merger.py
- model_merger.py
- print_cfg.py
- rollout_viewer.py
- sft_weights_to_hf.py
- agent_utils.py
- qwen_vl_tool_chat_template.jinja2
- test_agent_loop_reward.py
- test_basic_agent_loop.py
- test_multi_modal.py
- __init__.py
- test_gsm8k_interaction.py
- test_interaction_registry.py
- test_megatron_engine.py
- test_transformer.py
- test_transformers_ulysses.py
- test_decorator.py
- main.py
- client.py
- README.md
- run.sh
- server.py
- __init__.py
- test_auto_padding_on_cpu.py
- test_colocated_workers.py
- test_colocated_workers_fused.py
- test_data_transfer.py
- test_decorator_on_cpu.py
- test_device_mesh_register.py
- test_driverfunc_to_worker.py
- test_fused_workers_on_cpu.py
- test_high_level_scheduling_api.py
- test_nested_worker.py
- test_ray_collectives.py
- test_ray_local_envs_on_cpu.py
- test_ray_utils_on_cpu.py
- test_rvdz.py
- test_worker_group_basics.py
- test_worker_group_torch.py
- README.md
- run_all.sh
- test_fsdp_ckpt.py
- test_mcore_config_converter.py
- test_tensor_dict.py
- __init__.py
- task.py
- tokenizer.py
- __init__.py
- run_gen_qwen05.sh
- qwen2moe_minimal.json
- run_function_reward.sh
- run_model_reward.sh
- run_single_gpu.sh
- run_single_gpu_with_engine.sh
- run_sft.sh
- test_sp_loss_match.py
- __init__.py
- check_custom_rwd_fn.py
- check_results.py
- README.md
- run_dapo.sh
- run_genrm_remote.sh
- run_geo3k_fsdp_sgl_multiturn_w_tool.sh
- run_grpo_lora_with_merge.sh
- run_gsm8k_fsdp_sgl_multiturn_sf_tool.sh
- run_gsm8k_fsdp_sgl_multiturn_w_tool.sh
- run_one_step_off_policy.sh
- run_ppo_trainer_megatron.sh
- run_prime.sh
- run_r1_distill_qwen_aime24_eval.sh
- run_spin.sh
- run_sppo.sh
- run_test.sh
- run_qwen2_5_05b_dapo.sh
- run_qwen2_5_05b_grpo.sh
- run_qwen2_5_05b_grpo_mindspeed.sh
- run_qwen2_5_05b_sft_peft_sp2.sh
- run_qwen2_5_vl_3b_npu.sh
- check_api_docs.py
- check_device_api_usage.py
- check_docs_time_info.py
- check_docstrings.py
- check_license.py
- check_pr_description.py
- check_pr_title.py
- test_config_docs.py
- test_import.py
- type_coverage_check.py
- validate_imported_docs.py
- validate_structure.py
- README.md
- test_memory_buffers.py
- test_base_tool_on_cpu.py
- __init__.py
- legacy_ppo_megatron_trainer.yaml
- legacy_ppo_trainer.yaml
- test_algo_config_on_cpu.py
- test_critic_config_on_cpu.py
- test_legacy_config_on_cpu.py
- __init__.py
- test_core_algos_on_cpu.py
- test_metric_utils_on_cpu.py
- __init__.py
- test_esi_save_ckpt_on_cpu.py
- test_create_rl_sampler_on_cpu.py
- test_multiturn_sft_dataset_on_cpu.py
- test_rl_collate_fn_on_cpu.py
- test_rl_dataset_on_cpu.py
- test_sft_dataset_on_cpu.py
- test_metrics.py
- test_pipeline_parallel.py
- test_sandbox_fusion_on_cpu.py
- test_sandbox_on_cpu.py
- _test_module.py
- test_activation_offload.py
- test_config_on_cpu.py
- test_flops_counter.py
- test_fs_on_cpu.py
- test_import_utils_on_cpu.py
- test_linear_cross_entropy.py
- test_linear_cross_entropy_tp.py
- test_model_on_cpu.py
- test_nvtx_profile.py
- test_rollout_skip_on_cpu.py
- test_rollout_trace_on_cpu.py
- test_seqlen_balancing.py
- test_special_linear_cross_entropy_tp.py
- test_special_mstx_profile.py
- test_temp_env_on_cpu.py
- test_timeout_decorator_cpu.py
- test_torch_functional.py
- test_special_dp_actor.py
- test_actor_config_on_cpu.py
- test_critic_config_on_cpu.py
- test_engine_config_on_cpu.py
- test_optim_config_on_cpu.py
- test_special_dp_critic.py
- test_special_fsdp_engine.py
- test_registry_on_cpu.py
- vllm_async_rollout.py
- mcp_server.json
- mcp_tool_config
- sandbox_fusion_tool_config
- search_tool_config
- test_http_server_engine.py
- run_fsdp_vllm.py
- test_vllm_chat_scheduler.py
- test_vllm_hf_loader.py
- test_vllm_model_rope_scaling.py
- test_vllm_spmd.py
- async_rollout_utils.py
- test_async_sglang_server.py
- test_async_sglang_server_on_cpu.py
- test_custom_completion_callback.py
- test_hf_rollout.py
- test_sglang_async_rollout_mcp_tools.py
- test_sglang_async_rollout_multimodal_delta.py
- test_sglang_async_rollout_search_tools.py
- test_sglang_async_rollout_sf_tools.py
- test_sglang_async_rollout_w_interaction.py
- test_sglang_async_rollout_w_tools.py
- test_sglang_multi_interaction.py
- test_sglang_rollout_sharding_manager.py
- test_sglang_spmd.py
- utils_sglang.py
- test_fsdp_workers.py
- __init__.py
- kill_github_tests.sh
- README.md
- test_base_config_on_cpu.py
- test_protocol_on_cpu.py
- __init__.py
- agent_loop.py
- single_turn_agent_loop.py
- tool_agent_loop.py
- tool_parser.py
- __init__.py
- sampler.py
- __init__.py
- dynamicgen_dataset.py
- __init__.py
- __init__.py
- interaction_registry.py
- __init__.py
- base.py
- gsm8k_interaction.py
- __init__.py
- __main__.py
- base_model_merger.py
- fsdp_model_merger.py
- megatron_model_merger.py
- __init__.py
- llama_loader.py
- llama_loader_depracated.py
- llama_saver.py
- __init__.py
- parallel_attention.py
- parallel_decoder.py
- parallel_linear.py
- parallel_mlp.py
- parallel_rmsnorm.py
- __init__.py
- modeling_llama_megatron.py
- __init__.py
- __init__.py
- attention.py
- model.py
- rope_utils.py
- vision_config.py
- vision_model.py
- vision_transformer_block.py
- __init__.py
- config_converter.py
- loader.py
- mbridge.py
- model_forward.py
- model_forward_fused.py
- model_initializer.py
- patch_v012.py
- readme.md
- registry.py
- saver.py
- util.py
- weight_converter.py
- __init__.py
- qwen2_loader.py
- qwen2_loader_depracated.py
- qwen2_saver.py
- __init__.py
- parallel_attention.py
- parallel_decoder.py
- parallel_linear.py
- parallel_mlp.py
- parallel_rmsnorm.py
- __init__.py
- modeling_qwen2_megatron.py
- __init__.py
- __init__.py
- dense_common.py
- kimi_vl.py
- llama.py
- monkey_patch.py
- npu_patch.py
- qwen2.py
- qwen2_5_vl.py
- qwen2_vl.py
- __init__.py
- README.md
- registry.py
- weight_loader_registry.py
- __init__.py
- worker.py
- worker_group.py
- __init__.py
- ray.py
- __init__.py
- decorator.py
- worker.py
- worker_group.py
- __init__.py
- base.py
- megatron.py
- __init__.py
- __init__.py
- parallel_state.py
- __init__.py
- state_dict.py
- __init__.py
- _state_dict_utils.py
- __init__.py
- __init__.py
- arg_utils.py
- config.py
- dtensor_weight_loaders.py
- hf_weight_loader.py
- llm.py
- llm_engine_sp.py
- megatron_weight_loaders.py
- model_loader.py
- model_runner.py
- parallel_state.py
- spmd_gpu_executor.py
- tokenizer.py
- worker.py
- __init__.py
- arg_utils.py
- config.py
- dtensor_weight_loaders.py
- hf_weight_loader.py
- llm.py
- llm_engine_sp.py
- megatron_weight_loaders.py
- model_loader.py
- model_runner.py
- parallel_state.py
- spmd_gpu_executor.py
- tokenizer.py
- worker.py
- __init__.py
- __init__.py
- McpClientManager.py
- utils.py
- __init__.py
- search_r1_like_utils.py
- tool_registry.py
- __init__.py
- base_tool.py
- geo3k_tool.py
- gsm8k_tool.py
- image_zoom_in_tool.py
- mcp_base_tool.py
- mcp_search_tool.py
- sandbox_fusion_tools.py
- schemas.py
- search_tool.py
- actor.yaml
- dp_actor.yaml
- megatron_actor.yaml
- critic.yaml
- dp_critic.yaml
- megatron_critic.yaml
- legacy_data.yaml
- dapo_32b.yaml
- gspo.yaml
- hcv_ipho_numeric_val.yaml
- hcv_numeric_val.yaml
- ipho_hcv_val.yaml
- ipho_numeric_val.yaml
- log_all_reward.yaml
- math_verify_reward.yaml
- pass_at_n.yaml
- simple_eval.yaml
- syn_data.yaml
- use_JEEBench.yaml
- use_kl.yaml
- use_legacy_prompt.yaml
- use_olympiad_bench.yaml
- use_PHYSICS.yaml
- 4.1-mini-or.yaml
- 4.1-mini.yaml
- 4o-mini.yaml
- 4o.yaml
- flash.yaml
- gemma2-9b.yaml
- gemma3-27b.yaml
- gpt-oss-or.yaml
- gpt4.1.yaml
- gpt5.yaml
- llama-70b-local.yaml
- llama-8b.yaml
- o1.yaml
- o3-mini.yaml
- o3.yaml
- o4-mini.yaml
- qwen2.5-14b-instruct.yaml
- qwen2.5-32b-instruct.yaml
- qwen2.5-32b-symbolic-translated-numeric-reverse_sleek-firefly-214_global-step-650.yaml
- qwen2.5-32b_numeric-symbolic_lucky-feather-119_global-step-1200.yaml
- qwen2.5-32b_olive-valley-63_global_step_550.yaml
- qwen2.5-3b-instruct.yaml
- qwen2.5-3b-rare-glade-1739-global_step_185.yaml
- qwen2.5-3b-rare-glade-1739.yaml
- qwen2.5-3b-rural-paper-1690.yaml
- qwen2.5-3b.yaml
- qwen2.5-72b-free_api.yaml
- qwen2.5-72b-instruct.yaml
- qwen2.5-72b_api.yaml
- qwen2.5-7b-instruct.yaml
- qwen235b-instruct-local.yaml
- qwen235b-instruct-or.yaml
- qwen235b-instruct.yaml
- qwen235b-thinking-local.yaml
- qwen235b-thinking-or.yaml
- qwen3-4b-base.yaml
- qwen3-4b-instruct.yaml
- qwen3-4b-thinking.yaml
- qwen3-4b.yaml
- qwen3-80b-instruct.yaml
- qwen3-80b-thinking.yaml
- qwen30b-instruct-local.yaml
- qwen30b-instruct-or.yaml
- qwen30b-thinking-local.yaml
- qwen30b-thinking-or.yaml
- r1-1.5b.yaml
- r1-1.5b_local.yaml
- r1-14b.yaml
- r1-14b_local.yaml
- r1-32b.yaml
- r1-70b.yaml
- r1-8b.yaml
- r1-8b_local.yaml
- r1.yaml
- v3.yaml
- npu_profile.yaml
- dp_ref.yaml
- megatron_ref.yaml
- ref.yaml
- dp_reward_model.yaml
- megatron_reward_model.yaml
- reward_model.yaml
- rollout.yaml
- __init__.py
- _generated_ppo_megatron_trainer.yaml
- _generated_ppo_trainer.yaml
- algorithm.py
- config.py
- evaluation.yaml
- generation.yaml
- ppo_megatron_trainer.bak
- ppo_megatron_trainer.yaml
- ppo_trainer.bak
- ppo_trainer.yaml
- sft_trainer.yaml
- __init__.py
- core_algos.py
- dynamic_token_budget_utils.py
- metric_utils.py
- ray_trainer.py
- reward.py
- utils.py
- __init__.py
- constants_ppo.py
- fsdp_sft_trainer.py
- main_eval.py
- main_generation.py
- main_ppo.py
- runtime_env.yaml
- runtime_env_nvidia.yaml
- __init__.py
- checkpoint_manager.py
- fsdp_checkpoint_manager.py
- megatron_checkpoint_manager.py
- __init__.py
- custom_dataloader.py
- multiturn_sft_dataset.py
- online_rl_dataset.py
- README.md
- real_dataset.py
- rl_dataset.py
- rm_dataset.py
- sft_dataset.py
- vision_utils.py
- __init__.py
- empty_annotations.py
- metrics.py
- nvtx_profile.py
- performance.py
- profile.py
- trajectory_tracker.py
- __init__.py
- torch_functional.py
- __init__.py
- kernels.py
- linear_cross_entropy.py
- __init__.py
- aggregate_logger.py
- __init__.py
- dist_checkpointing.py
- memory.py
- optimizer.py
- pipeline_parallel.py
- sequence_parallel.py
- tensor_parallel.py
- __init__.py
- utils.py
- __init__.py
- config.py
- empty_annotations.py
- mstx_profile.py
- nvtx_profile.py
- performance.py
- profile.py
- __init__.py
- ray_backend.py
- __init__.py
- README.md
- testing_util.py
- utils.py
- __init__.py
- grader.py
- math_normalize.py
- __init__.py
- utils.py
- Reward_Functions_Stats.txt
- qa_pairs_with_predicted_answers_sample.json
- run_reward_function_tests.sh
- test_reward_function.py
- __init__.py
- final_reward_function.py
- geo3k.py
- gsm8k.py
- math.py
- math_batch.py
- math_combined.py
- math_dapo.py
- math_p.py
- math_p_symbolic.py
- math_verify.py
- math_verify_numerical_reward.py
- numerical_reward.py
- reward_conf1.py
- reward_conf2.py
- reward_conf3.py
- search_r1_like_qa_em.py
- symbolic_reward.py
- val_reward_fn.py
- verl_reward_score_wrapper.py
- __init__.py
- patch.py
- utils.py
- __init__.py
- activation_offload.py
- config.py
- device.py
- distributed.py
- flops_counter.py
- fs.py
- fsdp_utils.py
- hdfs_io.py
- import_utils.py
- logging_utils.py
- megatron_utils.py
- memory_buffer.py
- memory_utils.py
- model.py
- net_utils.py
- py_functional.py
- ray_utils.py
- rollout_skip.py
- rollout_trace.py
- seqlen_balancing.py
- tokenizer.py
- torch_dtypes.py
- torch_functional.py
- tracking.py
- transformers_compat.py
- ulysses.py
- vllm_utils.py
- version
- __init__.py
- base.py
- dp_actor.py
- megatron_actor.py
- __init__.py
- actor.py
- critic.py
- engine.py
- model.py
- optimizer.py
- rollout.py
- __init__.py
- base.py
- dp_critic.py
- megatron_critic.py
- __init__.py
- engine_impl.py
- utils.py
- __init__.py
- engine_impl.py
- utils.py
- __init__.py
- base.py
- __init__.py
- abstract.py
- batch.py
- dapo.py
- naive.py
- prime.py
- registry.py
- __init__.py
- reward_model.py
- __init__.py
- base.py
- __init__.py
- losses.py
- __init__.py
- actor.py
- critic.py
- hybrid_engine.py
- __init__.py
- naive_rollout.py
- __init__.py
- async_sglang_server.py
- http_server_engine.py
- sglang_rollout.py
- utils.py
- __init__.py
- fire_vllm_rollout.py
- python_executor.py
- vllm_async_server.py
- vllm_rollout.py
- vllm_rollout_spmd.py
- __init__.py
- async_server.py
- base.py
- chat_scheduler.py
- hf_rollout.py
- rollout_worker.py
- schemas.py
- tokenizer.py
- __init__.py
- base.py
- fsdp_sglang.py
- fsdp_ulysses.py
- fsdp_vllm.py
- megatron_sglang.py
- megatron_vllm.py
- __init__.py
- fsdp_workers.py
- megatron_workers.py
- __init__.py
- base_config.py
- protocol.py
- py.typed
- .gitignore
- .pre-commit-config.yaml
- .readthedocs.yaml
- CONTRIBUTING.md
- LICENSE
- Notice.txt
- pyproject.toml
- README.md
- requirements-npu.txt
- requirements.txt
- requirements_sglang.txt
- setup.py
- .gitignore
- README.md
- setup_jeebench.sh
# 설치 가이드
git clone https://github.com/Sim2Reason/Sim2Reason
깃허브에서 프로젝트 코드 전체를 내 컴퓨터로 내려받습니다.
cd Sim2Reason
방금 내려받은 프로젝트 폴더 안으로 이동합니다.
2. 공식 설치 스크립트
쉬움 추천- Python 3 pip 명령어를 쓰려면 Python이 필요합니다.
pip install bpy
PyPI에 배포된 패키지를 바로 설치합니다. 소스 클론이 필요 없습니다.
pip install mujoco==3.3.4 ImageIO
PyPI에 배포된 패키지를 바로 설치합니다. 소스 클론이 필요 없습니다.
pip install ipdb scipy tabulate pandas matplotlib
PyPI에 배포된 패키지를 바로 설치합니다. 소스 클론이 필요 없습니다.
pip install hydra-core omegaconf
PyPI에 배포된 패키지를 바로 설치합니다. 소스 클론이 필요 없습니다.
이 레포의 README에 적힌 실제 명령어를 그대로 가져왔습니다.
3. Docker
쉬움- Git GitHub에서 프로젝트 코드를 내려받으려면 필요합니다.
- Docker Desktop 컨테이너를 빌드하고 실행하려면 필요합니다. 설치 후 실행해서 백그라운드에 켜두세요.
docker build -f verl_v4/docker/Dockerfile.awsefa -t sim2reason .
Dockerfile을 기반으로 실행 가능한 이미지를 빌드합니다.
docker run -p 8080:80 sim2reason
빌드된 이미지를 실제 컨테이너로 실행합니다.
4. Python
쉬움pip install bpy
PyPI에 배포된 패키지를 바로 설치합니다. 소스 클론이 필요 없습니다.
pip install mujoco==3.3.4 ImageIO
PyPI에 배포된 패키지를 바로 설치합니다. 소스 클론이 필요 없습니다.
pip install ipdb scipy tabulate pandas matplotlib
PyPI에 배포된 패키지를 바로 설치합니다. 소스 클론이 필요 없습니다.
pip install hydra-core omegaconf
PyPI에 배포된 패키지를 바로 설치합니다. 소스 클론이 필요 없습니다.
pip install tqdm wandb
PyPI에 배포된 패키지를 바로 설치합니다. 소스 클론이 필요 없습니다.
이 레포의 README에 적힌 실제 명령어를 그대로 가져왔습니다.
