Qwen3.8-Flash-Next-DGX-Spark
Deploy the quantized version of the 180B hybrid-MoE model on a single NVIDIA DGX Spark (GB10, 128 GB unified memory, no multi-GPU, no cluster) and serve it stably
File Explorer
- benchmarks.md
- benchmarks.zh.md
- deploy_playbook.md
- deploy_playbook.zh.md
- deployment-matrix.md
- deployment-matrix.zh.md
- mtp-tracker.md
- mtp-tracker.zh.md
- speculative-analysis.md
- speculative-analysis.zh.md
- merge-mtp-shard.py
- qwen4exp-mtp-draft-head.patch
- baseline.json
- dl_q4_parallel.py
- monitor.py
- mtp_bench_safe.sh
- probe_mtp.py
- q4_bench.sh
- qwen38-q3.service
- run_qwen38_q3.sh
- system_watchdog.sh
- warm_table.py
- .gitignore
- LICENSE
- README.md
- README.zh.md
# Use via CDN
jsDelivrjsDelivr serves any public GitHub repository as a CDN with zero setup. Pick a version and a file to get a ready-to-paste link and snippet.
Link
Example
// repository documentation
Was this content helpful?
(0 ratings)
