SUPER-GIANT
Custom LLM that will soon (never) turn into god. ๐ฅTUESFEST
ํ์ผ ํ์๊ธฐ
์ต์ข ๋ฒ์ ๋ค์ด๋ก๋ (.zip)- Build_images.yml
- Dockerfile
- Dockerfile
- README.md
- gpu_job.py
- README.md
- runpod-gpu.sh
- s3.py
- sync-gpu.py
- wait-new-gpu.sh
- arxiv-src
- wait-new-gpu.sh
- a100-optimization-guide.md
- diffusers-h100.md
- diffusers-integration.md
- h100-optimization-guide.md
- huggingface-kernels-integration.md
- kernel-templates.md
- t4-optimization-guide.md
- transformers-integration.md
- troubleshooting.md
- benchmark_example.py
- benchmark_rmsnorm.py
- huggingface_kernels_example.py
- ltx_kernel_injection_example.py
- transformers_injection_example.py
- SKILL.md
- runpod-gpu.sh
- SKILL.md
- SKILL.md
- SKILL.md
- SKILL.md
- SKILL.md
- grep_in_session.ts
- README.md
- 00-cover.typ
- 01-intro.typ
- 02-ch1.typ
- 03-ch2.typ
- 03b-transformer-model.typ
- 03c-speculative-decoding.typ
- 04-ch3.typ
- 05-ch4.typ
- 06-conclusion.typ
- 07-bibliography.typ
- 08-appendix-resources.typ
- 09-experiments-embedded.typ
- _common.typ
- DOCS.md
- index.html
- main.pdf
- main.typ
- main_body.pdf
- main_final.pdf
- main_final_qpdf.pdf
- STANOVISTE.pdf
- ZADANIE.pdf
- ะะธะฟะปะพะผะฝะฐ ะ ะฐะฑะพัะฐ ะะฝัะพะฝ ะฅัะธััะพะฒ - black and white.pdf
- ะะธะฟะปะพะผะฝะฐ ะ ะฐะฑะพัะฐ ะะฝัะพะฝ ะฅัะธััะพะฒ - color.pdf
- ะะธะฟะปะพะผะฝะฐ ะ ะฐะฑะพัะฐ ะะฝัะพะฝ ะฅัะธััะพะฒ.pdf
- Anchor_TiDAR_forward_pass.png
- Anchor_TiDAR_KV_cache_diagram.png
- TiDAR_infernece_mask.png
- TiDAR_Single_Forward_pass.png
- Vanilla_speculative_decoding_with_smaller_model.png
- Attention_diagram.png
- Attention_diagram_bw_swapped copy.png
- Attention_diagram_bw_swapped.png
- DecoderTransformer.png
- Docker.png
- DR_Cover.pdf
- DR_cover.png
- Google_search_speed_matters.jpg
- Google_Search_Speed_Matters.png
- jax.png
- MHA_diagram_and_math_formula.png
- mha_img_original.png
- My Movie 1.iMovieMobile
- PDF-preview-greedy-results.png
- Prefill_and_Decode_diagram_wide.png
- RoPE_visualization.gif
- SUPER-GIANT-avatar-logo.png
- SUPER-GIANT-text-logo.png
- super_giant_artifacts.dot
- super_giant_artifacts.png
- super_giant_artifacts.svg
- super_giant_framework.dot
- super_giant_framework.png
- super_giant_framework.svg
- super_giant_ops.dot
- super_giant_ops.png
- super_giant_ops.svg
- Tensor_Core_Utilization_Diagram.png
- TiDAR_acceptance_best_small_big.png
- tidar_agenda.png
- TiDAR_decode_AND_attention_mask_and_agenda.png
- TiDAR_gpu_measured_k_sweep.png
- TiDAR_prefill_mask.png
- TiDAR_training_mask.png
- attn_res_core_torch.py
- block_attnres_forward_torch.py
- summary.md
- S3.mp4
- ARTIFACTS.md
- FRAMEWORK.md
- OPERATIONS.md
- README.md
- __init__.py
- Config.yml
- data.py
- evaluate_emission.py
- generate_greedy.py
- graph_rnn_llm.py
- README.md
- train.py
- __init__.py
- checkpoint_io.py
- checkpoint_manager.py
- optimizer_utils.py
- Config.yml
- dataset.py
- generate_dataset.py
- infer.py
- jit_inference.py
- Run_training.py
- Training_step.py
- Config.yml
- TRM.py
- TRM_block.py
- __init__.py
- Config.yml
- Config_easy.yml
- Config_medium.yml
- jit_inference.py
- Run_training.py
- sample_sudoku_pairs.py
- smoke_test.py
- sudoku_dataset.py
- Training_step.py
- visualize_sudoku_solving.py
- exp1_eval.png
- exp2_eval.png
- recursive_eval.png
- Config.yml
- Config_dataset_large.yml
- Config_experiment.yml
- Config_full30.yml
- Config_full30_large.yml
- Config_recursive_experiment.yml
- Config_smoke.yml
- dataset.py
- EXPERIMENT_RESULTS_INITIAL.md
- EXPERIMENT_RESULTS_MEDIUM_LONG.md
- EXPERIMENT_RESULTS_RECURSIVE_5090.md
- EXPERIMENTS_OVERVIEW.pdf
- EXPERIMENTS_OVERVIEW.typ
- prepare_dataset.py
- Run_training.py
- Training_step.py
- TRM-encoder-architecture.png
- TRMEncoderDecoder.py
- build_sudoku_trm_dataset.py
- Config.yml
- config_utils.py
- EXPERIMENTS_TRM_TOKEN_TOOL-1.png
- EXPERIMENTS_TRM_TOKEN_TOOL-2.png
- EXPERIMENTS_TRM_TOKEN_TOOL.pdf
- EXPERIMENTS_TRM_TOKEN_TOOL.typ
- LLM-token-TRM.png
- README.md
- Run_eval_general_loss.py
- Run_eval_ood_sudoku.py
- Run_inference_sudoku.py
- Run_training.py
- tokenizer_utils.py
- __init__.py
- cpu_smoke_test.py
- Global_Config.yml
- TRM_comparison_report.md
- TRM_documentaion.md
- README.md
- Config.py
- Config.yml
- Data_loader.py
- Evaluate.py
- Generate_text.py
- GiantGPT.py
- prepare_dataset.py
- Run_inference.py
- Run_training.py
- Save_params.py
- Training_step.py
- Transformer_block.py
- README.md
- synthetic_db_generator.py
- Generate_text_fast.py
- Generate_text_fast_old.py
- GiantGPT.py
- GiantGPT_old.py
- prepare_dataset.py
- Run_training.py
- Run_training_old.py
- Training_step.py
- Training_step_old.py
- Transformer_block.py
- Transformer_block_old.py
- config_rl.yml
- generate_math_dataset.py
- math_env.py
- math_tokenizer.py
- run_rl_training.py
- checkpoint_io.py
- checkpoint_manager.py
- Config.yml
- Create_custom_tokenizer.py
- Data_loader.py
- Evaluate.py
- Flash_att_transformer.py
- Generate_faster.py
- Generate_text.py
- Generate_text_fast.py
- GiantGPT.py
- inspect_shard.py
- jit_inference.py
- prepare_dataset.py
- prepare_dataset_new.py
- Run_training.py
- Save_params.py
- test_tokenizer.py
- Training_step.py
- Transformer_block.py
- use_hg_model.py
- checkpoint_io.py
- checkpoint_torch_io.py
- Config.yml
- Create_custom_tokenizer.py
- Generate_text_fast.py
- GiantGPT.py
- Transformer_block.py
- Model_Overview.md
- README.md
- __init__.py
- build_corpus.py
- cleaning.py
- Config.yml
- __init__.py
- arrow_data_loader.py
- Chat.py
- checkpoint_manager.py
- Config.yml
- Create_custom_tokenizer.py
- Evaluate.py
- Generate_faster.py
- GiantGPT.py
- jit_chat_step.py
- jit_inference.py
- optimizer_utils.py
- prepare_dataset.py
- Run_training.py
- test_batched_inference.py
- test_transformer_block_gqa.py
- Training_step.py
- Transformer_block.py
- Config.yml
- convert_from_hf.py
- Generate_faster.py
- Qwen_block.py
- Qwen_model.py
- __init__.py
- checkpoint_io.py
- checkpoint_manager.py
- Config.yml
- Config_135m.yml
- Config_360m.yml
- download_and_convert_smollm.py
- download_and_convert_smollm_135m.py
- Generate_faster.py
- GiantGPT.py
- jit_inference.py
- Transformer_block.py
- Config.yml
- Create_custom_chat_tokenizer.py
- DOCS.md
- peek_arrow_sample.py
- Save_tokenizer_locally.py
- Teacher_config.yml
- teacher_generate.py
- collect_used_tokens.py
- download_mixed_corpus.py
- latent_decoding.py
- latent_decoding_methods.py
- match_tokenizers.py
- match_tokenizers_on_data.py
- random_text.py
- test_chat_loss_mask_pipeline.py
- test_data_pipeline_smoke.py
- __init__.py
- device_utils.py
- Global_Config.yml
- kv_speed_test.py
- README.md
- vllm_speed_test.py
- curriculum_20b.yml
- giant_bg_quality_keep_bpe32k.yml
- giant_chat_curated_booster.yml
- giant_chat_curated_booster_bg_en_bpe32k.yml
- giant_chat_curated_booster_strong56.yml
- giant_chat_pretraining.yml
- giant_chat_pretraining_4b.yml
- giant_chat_pretraining_bg_en_900m_bpe32k.yml
- giant_chat_sft.yml
- giant_chat_sft_4x.yml
- giant_chat_sft_bg_en_smoltalk_bpe32k.yml
- giant_bg_en_bpe32k_quality.yml
- giant_chat_bg_en_bpe32k.yml
- giant_chat_bpe16k.yml
- giant_chat_bpe24k.yml
- attention_size_xsa_loss_walltime.png
- model_200m_mlp_x1_no_xsa.yml
- model_200m_mlp_x1_xsa.yml
- model_200m_mlp_x4_no_xsa.yml
- longdsl_30m_l1_ctx1024.yml
- longdsl_30m_l1_ctx128.yml
- longdsl_30m_l1_ctx128_transfer.yml
- longdsl_30m_l1_ctx16384.yml
- longdsl_30m_l1_ctx2048.yml
- longdsl_30m_l1_ctx256.yml
- longdsl_30m_l1_ctx256_transfer.yml
- longdsl_30m_l1_ctx32768.yml
- longdsl_30m_l1_ctx4096.yml
- longdsl_30m_l1_ctx512.yml
- longdsl_30m_l1_ctx64.yml
- longdsl_30m_l1_ctx65536.yml
- longdsl_30m_l1_ctx8192.yml
- longdsl_30m_l2_ctx1024.yml
- longdsl_30m_l2_ctx16384.yml
- longdsl_30m_l2_ctx2048.yml
- longdsl_30m_l2_ctx256.yml
- longdsl_30m_l2_ctx32768.yml
- longdsl_30m_l2_ctx4096.yml
- longdsl_30m_l2_ctx512.yml
- longdsl_30m_l2_ctx65536.yml
- longdsl_30m_l2_ctx8192.yml
- long_100m_l1_ctx128k_chi40x_decoder_xsa_5090.yml
- long_40m_continue_ctx2k_to_16k_lm_chi20x_a6000.yml
- long_40m_l1_ctx128.yml
- long_40m_l1_ctx2048_chinchilla20x_warm40_5090.yml
- long_40m_l1_ctx256.yml
- long_40m_l1_ctx512.yml
- long_40m_l1_ctx512_99klm_1kans_3ep_decoder_xsa.yml
- long_40m_l1_ctx512_ans_100k_3ep_decoder_xsa.yml
- long_40m_l1_ctx512_ans_10k_baseline.yml
- long_40m_l1_ctx512_ans_10k_encoder_xsa.yml
- long_40m_l1_ctx512_ans_1k_baseline.yml
- long_40m_l1_ctx512_ans_1k_encoder_xsa.yml
- long_40m_l1_ctx512_ans_20k_1ep_baseline.yml
- long_40m_l1_ctx512_ans_20k_1ep_encoder_xsa.yml
- long_40m_l1_ctx512_ans_20k_1ep_encoder_xsa_answer_hidden.yml
- long_40m_l1_ctx512_ans_2k_10ep_baseline.yml
- long_40m_l1_ctx512_ans_2k_10ep_encoder_xsa.yml
- long_40m_l1_grokking_probe_100k.yml
- long_40m_l1_sweep_ansheavy.yml
- long_40m_l1_sweep_lmheavy.yml
- long_40m_l2_ctx128.yml
- long_40m_l2_ctx256.yml
- long_40m_l2_ctx512.yml
- 1_pretraining_100m_bg_en_ctx256_32k_1p8b.yml
- 1_pretraining_100m_ctx256_4b.yml
- 1_pretraining_50m_ctx256.yml
- 2_curated_booster_100m_bg_en_ctx256_32k.yml
- 2_curated_booster_100m_ctx256.yml
- 2_curated_booster_50m_ctx256.yml
- 3_sft_100m_bg_en_ctx256_32k_smoltalk.yml
- 3_sft_100m_ctx256_4x.yml
- 3_sft_50m_ctx256.yml
- big_model.yml
- dev_100m_hq.yml
- giant_chat_pretrain_50m_ctx256.yml
- giant_chat_sft_50m_ctx256.yml
- README.md
- registry.yml
- teacher_kd_smoltalk_example.yml
- benchmark_bg_full_accessible.yml
- benchmark_bg_full_with_bpos.yml
- benchmark_bg_recommended.yml
- benchmark_bpos_probe.yml
- benchmark_tiny_local.yml
- train_fasttext_bg_full_accessible.yml
- train_fasttext_bg_full_with_bpos.yml
- train_fasttext_tiny_local.yml
- __init__.py
- build_quality_benchmark.py
- common.py
- Config.yml
- filter_lumees_pretrain.py
- filter_lumees_pretrain_parallel.py
- materialize_tier_views.py
- README.md
- score_lumees_sft_fasttext.py
- train_quality_filter.py
- __init__.py
- build_demo_target_pack.py
- build_prompt_bank.py
- curated_wikipedia_search.py
- export_top_candidates.py
- README.md
- sample_targets_100.jsonl
- samples.md
- teacher_distill.py
- TEACHER_KD.md
- translate_curated_booster.py
- translate_smoltalk_to_bg.py
- __init__.py
- build_corpus.py
- cleaning.py
- Config.yml
- README.md
- train_tokenizer.py
- Long_CTX2048_ANS_Tail_Failure.pdf
- Long_CTX2048_ANS_Tail_Failure.typ
- Long_Grokking_Probe_100k.pdf
- Long_Grokking_Probe_100k.typ
- Long_Grokking_Probe_10k.pdf
- Long_Grokking_Probe_10k.typ
- Long_LM_vs_ANS_Schedule_Sweep.pdf
- Long_LM_vs_ANS_Schedule_Sweep.typ
- Long_Natural_Record_Transition.pdf
- Long_Natural_Record_Transition.typ
- LongDSL_Night1_Bootstrap_Report.pdf
- LongDSL_Night1_Bootstrap_Report.typ
- throughtput_before
- __init__.py
- eval_longdsl.py
- eval_openai_long.py
- Lexicon.yml
- longdsl.py
- prepare_longdsl.py
- README.md
- write_training_configs.py
- __init__.py
- arrow_data_loader.py
- checkpoint_manager.py
- Config.yml
- Create_custom_tokenizer.py
- Evaluate.py
- Generate_chat.py
- Generate_faster.py
- GiantGPT.py
- jit_chat_step.py
- jit_inference.py
- model_mode.py
- optimizer_utils.py
- prepare_dataset.py
- README.md
- Run_training.py
- test_batched_inference.py
- test_transformer_block_gqa.py
- Training_step.py
- Transformer_block.py
- collect_used_tokens.py
- device_utils.py
- download_mixed_corpus.py
- kv_speed_test.py
- latent_decoding.py
- latent_decoding_methods.py
- match_tokenizers.py
- match_tokenizers_on_data.py
- random_text.py
- test_chat_loss_mask_pipeline.py
- test_data_pipeline_smoke.py
- test_longdsl.py
- test_teacher_kd_pipeline.py
- vllm_speed_test.py
- __init__.py
- device_utils.py
- Global_Config.yml
- README.md
- run_manifest.py
- __init__.py
- README.md
- __init__.py
- cli.py
- 1_8B_first_run_config.yml
- 1B_first_run_config.yml
- fineweb_edu_1p5b_cosmo2_ctx512.yml
- fineweb_edu_topup_64m_cosmo2_ctx512.yml
- Greedy_exp_500m.yml
- sweep_1_lr1e4_a1_l0.yml
- sweep_5_300m_ctxmix.yml
- tinystories_300m_512.yml
- TinyStories_full.yml
- __init__.py
- Config_docs.md
- Run_pipeline.py
- Smoke_config.yml
- giant_h18_kv6_gqa_logs.txt
- giant_h18_mha_logs.txt
- giant_h6_hd32_mha_logs.txt
- giant_h6_mha_logs.txt
- tidar_h18_kv6_gqa_logs.txt
- tidar_h18_mha_logs.txt
- tidar_h6_hd32_mha_logs.txt
- tidar_h6_mha_logs.txt
- greedy_135m_not_full.txt
- greedy_360m_not_full.txt
- sweep_1_logs_broken_masking.txt
- sweep_2_logs_broken_masking.txt
- sweep_3_logs_broken_masking.txt
- sweep_4_logs_broken_masking.txt
- sweep_5_draft_2_logs.txt
- sweep_5_draft_3_logs.txt
- sweep_5_logs_broken_masking.txt
- tidar_0p2_0_0p5_0_0p3_logs.txt
- tidar_1_1_0p1_0_0p3_logs.txt
- tidar_bigGamma_andDelta_logs.txt
- tidar_delta_masked_logs.txt
- tidar_eta0p04_t2_logs.txt
- tidar_later_stage_topk_logs.txt
- tidar_masked_delta_later_logs.txt
- tidar_noAB_greedy_eta_logs.txt
- tidar_smallAR_big_greedy_eta_logs.txt
- tidar_stable_bigger_beta_logs.txt
- tidar_stable_logs.txt
- Anchor_TiDAR_KV_cache_diagram.typ
- Bucket_Prefix0_1600_A40_Spot.pdf
- Bucket_Prefix0_1600_A40_Spot.typ
- Finding_Free_token_slots_A40.pdf
- Finding_Free_token_slots_A40.typ
- Implementation_Notes.md
- JAX_prod.md
- K+-2.md
- KV_Cache_Policy_A40.pdf
- KV_Cache_Policy_A40.typ
- README.md
- Results_Greedy_runs_135_360.pdf
- Results_Greedy_runs_135_360.typ
- Results_Sweep_4_runs_smollm135.typ
- TiDAR_105k_Inference_HeadToHead.typ
- TiDAR_AR_A40_H100_HeadToHead.pdf
- TiDAR_AR_A40_H100_HeadToHead.typ
- TiDAR_docs.md
- TiDAR_inference_experiments_summary.pdf
- TiDAR_inference_experiments_summary.typ
- TiDAR_vs_AR_SingleForward_Updated.pdf
- TiDAR_vs_AR_SingleForward_Updated.typ
- TinyStories_Attention_HeadSweep_Losses.pdf
- TinyStories_Attention_HeadSweep_Losses.typ
- TinyStories_TiDAR_losses_Comparison.pdf
- TinyStories_TiDAR_losses_Comparison.typ
- TinyStories_TiDAR_Losses_Stable_vs_KL.pdf
- Training_experiments.md
- check_tokenizer.py
- train_tokenizer.py
- tidar_smollm135_finewebedu_k8_ctx256_1b_normal.yml
- tidar_smollm135_finewebedu_k8_ctx256_probe.yml
- tidar_smollm135_finewebedu_k8_ctx256_timing.yml
- tidar_smollm135_finewebedu_k8_ctx256split_1b_normal.yml
- tidar_smollm135_finewebedu_k8_ctx256split_500m_continue_normal.yml
- tidar_smollm135_finewebedu_k8_ctx256split_500m_lossA_kl_only.yml
- tidar_smollm135_finewebedu_k8_ctx256split_500m_lossB_aggressive.yml
- Config.yml
- sweep_1_lr1e4_a1_l0.yml
- sweep_2_lr5e5_a0p7_l0p1.yml
- sweep_3_lr3e5_a0p3_l0p5.yml
- sweep_4_lr1e5_a0p5_l0p2.yml
- sweep_5_lr3e5_a0p2_l0p8_t2_ctxmix.yml
- tinystories_small_512.yml
- GIANT_tinystories_h18_kv6_gqa.yml
- GIANT_tinystories_h18_mha.yml
- GIANT_tinystories_h54_kv18_d1728.yml
- GIANT_tinystories_h6_hd32_mha.yml
- GIANT_tinystories_h6_mha_baseline.yml
- TiDAR_tinystories_h18_kv6_gqa_post.yml
- TiDAR_tinystories_h18_mha_post.yml
- TiDAR_tinystories_h54_kv18_d1728_post.yml
- TiDAR_tinystories_h6_hd32_mha_post.yml
- TiDAR_tinystories_h6_mha_post.yml
- GIANT_base_training.yml
- TiDAR_later_stage_aggressive_greedy.yml
- TiDAR_later_stage_agreement.yml
- TiDAR_later_stage_agreement_forward_kl.yml
- TiDAR_later_stage_bigGamma_andDelta.yml
- TiDAR_later_stage_distill.yml
- TiDAR_later_stage_topk.yml
- TiDAR_post_training.yml
- TiDAR_post_training_delta_masked.yml
- Greedy_exp_135m.yml
- Greedy_exp_360m.yml
- __init__.py
- analysis_inference_accepts.py
- Config.yml
- Config_135m.yml
- Config_360m.yml
- config_schema.py
- cudnn_attention_mask_example.py
- distributional_invariance_test.py
- GiantTiDAR.py
- inference.py
- Optimize_TiDAR_worst_case_decoding.md
- Prepare_mask_token.py
- Run_inference_logits_hist.py
- Run_training.py
- save_init_params.py
- test_inference_througthput.py
- tidar_core.py
- tidar_masks.py
- tidar_utils.py
- tokenizer_utils.py
- Training_step.py
- Transformer_block.py
- bench_cuDNN_vs_JAX_vs_Pallas_attention.py
- bench_generation_kvcache.py
- benchmark_decode_attention_speed.py
- create_mask.py
- find_batch_size.py
- incspecT_ultrachat_shard.py
- test_5term_loss_edge_cases.py
- test_agreement_loss.py
- test_nan_theory.py
- test_training_mask_alignment.py
- validate_training_batch.py
- visualize_training_config.py
- __init__.py
- Global_Config.yml
- README.md
- .gitignore
- index.html
- pyproject.toml
- README.md
- requirements.txt
// repository documentation
Was this content helpful?
(0 ratings)
