awesome-agent-kernel-skills
No description available.
File Explorer
Showing a partial file list — download the zip above to see everything.
- generate_data.py
- udf_pipeline.py
- etl_pipeline.py
- generate_data.py
- generate_data.py
- groupby_analysis.py
- generate_data.py
- multi_join.py
- run_smoke.sh
- threaded_handoff.cu
- NOTICE.md
- generate_data.py
- null_pipeline.py
- generate_data.py
- parquet_pipeline.py
- generate_data.py
- reshape_analysis.py
- clean_contacts.py
- generate_data.py
- generate_data.py
- timeseries_analysis.py
- generate_data.py
- window_analysis.py
- train.py
- null_cleanup.py
- NOTICE.md
- evals.json
- api-patterns.md
- cudf-pandas-accelerator.md
- dask-cudf-patterns.md
- BENCHMARK.md
- skill-card.md
- SKILL.md
- skill.oms.sig
- SKILL.md
- SKILL.md
- SKILL.md
- SKILL.md
- advanced-techniques.md
- attention-kernels.md
- gemm-moe.md
- integration.md
- performance.md
- sampling.md
- SKILL.md
- SKILL.md
- benchmarking-and-profiling.md
- gemm-and-linear.md
- torch-compile-and-graphs.md
- triton-on-rocm.md
- SKILL.md
- dependency-debugging.md
- library-and-model-adaptation.md
- torch-compile-and-cudagraph.md
- verification-methodology.md
- SKILL.md
- config-example-rocm.yaml
- deepseek-r1-mi355x.yaml
- llama3.1-70b-mi355x.yaml
- qwen3-32b-mi355x.yaml
- SKILL.md
- skill.json
- SKILL.md
- SKILL.md
- skill.json
- SKILL.md
- skill.json
- SKILL.md
- structcu__dev__sm__resource__group__params.md
- structcuaccesspolicywindow__v1.md
- structcuarraymapinfo__v1.md
- structcuasyncnotificationinfo.md
- structcucheckpointcheckpointargs.md
- structcucheckpointgpupair.md
- structcucheckpointlockargs.md
- structcucheckpointrestoreargs.md
- structcucheckpointunlockargs.md
- structcuctxcigparam.md
- structcuctxcreateparams.md
- structcuda__array3d__descriptor__v2.md
- structcuda__array__descriptor__v2.md
- structcuda__array__memory__requirements__v1.md
- structcuda__array__sparse__properties__v1.md
- structcuda__batch__mem__op__node__params__v1.md
- structcuda__child__graph__node__params.md
- structcuda__conditional__node__params.md
- structcuda__event__record__node__params.md
- structcuda__event__wait__node__params.md
- structcuda__ext__sem__signal__node__params__v1.md
- structcuda__ext__sem__signal__node__params__v2.md
- structcuda__ext__sem__wait__node__params__v1.md
- structcuda__ext__sem__wait__node__params__v2.md
- structcuda__external__memory__buffer__desc__v1.md
- structcuda__external__memory__handle__desc__v1.md
- structcuda__external__memory__mipmapped__array__desc__v1.md
- structcuda__external__semaphore__handle__desc__v1.md
- structcuda__external__semaphore__signal__params__v1.md
- structcuda__external__semaphore__wait__params__v1.md
- structcuda__graph__instantiate__params.md
- structcuda__host__node__params__v1.md
- structcuda__host__node__params__v2.md
- structcuda__kernel__node__params__v1.md
- structcuda__kernel__node__params__v2.md
- structcuda__kernel__node__params__v3.md
- structcuda__launch__params__v1.md
- structcuda__mem__alloc__node__params__v1.md
- structcuda__mem__alloc__node__params__v2.md
- structcuda__mem__free__node__params.md
- structcuda__memcpy2d__v2.md
- structcuda__memcpy3d__peer__v1.md
- structcuda__memcpy3d__v2.md
- structcuda__memcpy__node__params.md
- structcuda__memset__node__params__v1.md
- structcuda__memset__node__params__v2.md
- structcuda__pointer__attribute__p2p__tokens__v1.md
- structcuda__resource__desc__v1.md
- structcuda__resource__view__desc__v1.md
- structcuda__texture__desc__v1.md
- structcudevprop__v1.md
- structcudevresource.md
- structcudevsmresource.md
- structcudevworkqueueconfigresource.md
- structcudevworkqueueresource.md
- structcueglframe__v1.md
- structcuexecaffinityparam__v1.md
- structcuexecaffinitysmcount__v1.md
- structcuextent3d__v1.md
- structcugraphedgedata.md
- structcugraphexecupdateresultinfo__v1.md
- structcugraphnodeparams.md
- structcuipceventhandle__v1.md
- structcuipcmemhandle__v1.md
- structculaunchattribute.md
- structculaunchconfig.md
- structculaunchmemsyncdomainmap.md
- structcumemaccessdesc__v1.md
- structcumemallocationprop__v1.md
- structcumemcpy3doperand__v1.md
- structcumemcpyattributes__v1.md
- structcumemdecompressparams.md
- structcumemfabrichandle__v1.md
- structcumemlocation__v1.md
- structcumempoolprops__v1.md
- structcumempoolptrexportdata__v1.md
- structcumulticastobjectprop__v1.md
- structcuoffset3d__v1.md
- structcutensormap.md
- group__cuda__checkpoint.md
- group__cuda__coredump.md
- group__cuda__ctx.md
- group__cuda__ctx__deprecated.md
- group__cuda__d3d10.md
- group__cuda__d3d10__deprecated.md
- group__cuda__d3d11.md
- group__cuda__d3d11__deprecated.md
- group__cuda__d3d9.md
- group__cuda__d3d9__deprecated.md
- group__cuda__device.md
- group__cuda__device__deprecated.md
- group__cuda__driver__entry__point.md
- group__cuda__egl.md
- group__cuda__error.md
- group__cuda__event.md
- group__cuda__exec.md
- group__cuda__exec__deprecated.md
- group__cuda__extres__interop.md
- group__cuda__gl.md
- group__cuda__gl__deprecated.md
- group__cuda__graph.md
- group__cuda__graphics.md
- group__cuda__green__contexts.md
- group__cuda__initialize.md
- group__cuda__library.md
- group__cuda__logs.md
- group__cuda__malloc__async.md
- group__cuda__mem.md
- group__cuda__memop.md
- group__cuda__module.md
- group__cuda__module__deprecated.md
- group__cuda__multicast.md
- group__cuda__occupancy.md
- group__cuda__peer__access.md
- group__cuda__primary__ctx.md
- group__cuda__profiler.md
- group__cuda__profiler__deprecated.md
- group__cuda__stream.md
- group__cuda__surfobject.md
- group__cuda__surfref__deprecated.md
- group__cuda__tensor__memory.md
- group__cuda__texobject.md
- group__cuda__texref__deprecated.md
- group__cuda__types.md
- group__cuda__unified.md
- group__cuda__va.md
- group__cuda__vdpau.md
- group__cuda__version.md
- INDEX.md
- structcudaaccesspolicywindow.md
- structcudaarraymemoryrequirements.md
- structcudaarraysparseproperties.md
- structcudaasyncnotificationinfo__t.md
- structcudachannelformatdesc.md
- structcudachildgraphnodeparams.md
- structcudaconditionalnodeparams.md
- structcudadeviceprop.md
- structcudadevresource.md
- structcudadevsmresource.md
- structcudadevsmresourcegroupparams.md
- structcudadevworkqueueconfigresource.md
- structcudadevworkqueueresource.md
- structcudaeglframe.md
- structcudaeglplanedesc.md
- structcudaeventrecordnodeparams.md
- structcudaeventwaitnodeparams.md
- structcudaextent.md
- structcudaexternalmemorybufferdesc.md
- structcudaexternalmemoryhandledesc.md
- structcudaexternalmemorymipmappedarraydesc.md
- structcudaexternalsemaphorehandledesc.md
- structcudaexternalsemaphoresignalnodeparams.md
- structcudaexternalsemaphoresignalnodeparamsv2.md
- structcudaexternalsemaphoresignalparams.md
- structcudaexternalsemaphorewaitnodeparams.md
- structcudaexternalsemaphorewaitnodeparamsv2.md
- structcudaexternalsemaphorewaitparams.md
- structcudafuncattributes.md
- structcudagraphedgedata.md
- structcudagraphexecupdateresultinfo.md
- structcudagraphinstantiateparams.md
- structcudagraphkernelnodeupdate.md
- structcudagraphnodeparams.md
- structcudahostnodeparams.md
- structcudahostnodeparamsv2.md
- structcudaipceventhandle__t.md
- structcudaipcmemhandle__t.md
- structcudakernelnodeparams.md
- structcudakernelnodeparamsv2.md
- structcudalaunchattribute.md
- structcudalaunchconfig__t.md
- structcudalaunchmemsyncdomainmap.md
- structcudamemaccessdesc.md
- structcudamemallocnodeparams.md
- structcudamemallocnodeparamsv2.md
- structcudamemcpy3doperand.md
- structcudamemcpy3dparms.md
- structcudamemcpy3dpeerparms.md
- structcudamemcpyattributes.md
- structcudamemcpynodeparams.md
- structcudamemfreenodeparams.md
- structcudamemlocation.md
- structcudamempoolprops.md
- structcudamempoolptrexportdata.md
- structcudamemsetparams.md
- structcudamemsetparamsv2.md
- structcudaoffset3d.md
- structcudapitchedptr.md
- structcudapointerattributes.md
- structcudapos.md
- structcudaresourcedesc.md
- structcudaresourceviewdesc.md
- structcudatexturedesc.md
- structcuuuid__st.md
- unioncudalaunchattributevalue.md
- group__cudart____version.md
- group__cudart__d3d10.md
- group__cudart__d3d10__deprecated.md
- group__cudart__d3d11.md
- group__cudart__d3d11__deprecated.md
- group__cudart__d3d9.md
- group__cudart__d3d9__deprecated.md
- group__cudart__device.md
- group__cudart__device__deprecated.md
- group__cudart__driver.md
- group__cudart__driver__entry__point.md
- group__cudart__egl.md
- group__cudart__error.md
- group__cudart__event.md
- group__cudart__execution.md
- group__cudart__execution__context.md
- group__cudart__execution__deprecated.md
- group__cudart__extres__interop.md
- group__cudart__graph.md
- group__cudart__highlevel.md
- group__cudart__interop.md
- group__cudart__library.md
- group__cudart__logs.md
- group__cudart__memory.md
- group__cudart__memory__deprecated.md
- group__cudart__memory__pools.md
- group__cudart__occupancy.md
- group__cudart__opengl.md
- group__cudart__opengl__deprecated.md
- group__cudart__peer.md
- group__cudart__profiler.md
- group__cudart__stream.md
- group__cudart__surface__object.md
- group__cudart__texture__object.md
- group__cudart__types.md
- group__cudart__unified.md
- group__cudart__vdpau.md
- INDEX.md
- 1.1-scalable-data-parallel-computing-using-gpus.md
- 1.2-goals-of-ptx.md
- 1.3-ptx-isa-version-91.md
- 1.4-document-structure.md
- 10.1-special-registerstid.md
- 10.10-special-registersgridid.md
- 10.11-special-registersis_explicit_cluster.md
- 10.12-special-registersclusterid.md
- 10.13-special-registersnclusterid.md
- 10.14-special-registerscluster_ctaid.md
- 10.15-special-registerscluster_nctaid.md
- 10.16-special-registerscluster_ctarank.md
- 10.17-special-registerscluster_nctarank.md
- 10.18-special-registerslanemask_eq.md
- 10.19-special-registerslanemask_le.md
- 10.2-special-registersntid.md
- 10.20-special-registerslanemask_lt.md
- 10.21-special-registerslanemask_ge.md
- 10.22-special-registerslanemask_gt.md
- 10.23-special-registersclockclock_hi.md
- 10.24-special-registersclock64.md
- 10.25-special-registerspm0pm7.md
- 10.26-special-registerspm0_64pm7_64.md
- 10.27-special-registersenvreg32.md
- 10.28-special-registersglobaltimerglobaltimer_loglobaltimer_hi.md
- 10.29-special-registersreserved_smem_offset_beginreserved_smem_offset_endreserved_smem_offset_capreserved_smem_offset_2.md
- 10.3-special-registerslaneid.md
- 10.30-special-registerstotal_smem_size.md
- 10.31-special-registersaggr_smem_size.md
- 10.32-special-registersdynamic_smem_size.md
- 10.33-special-registerscurrent_graph_exec.md
- 10.4-special-registerswarpid.md
- 10.5-special-registersnwarpid.md
- 10.6-special-registersctaid.md
- 10.7-special-registersnctaid.md
- 10.8-special-registerssmid.md
- 10.9-special-registersnsmid.md
- 11.1.1-ptx-module-directivesversion.md
- 11.1.2-ptx-module-directivestarget.md
- 11.1.3-ptx-module-directivesaddress_size.md
- 11.2.1-kernel-and-function-directivesentry.md
- 11.2.2-kernel-and-function-directivesfunc.md
- 11.2.3-kernel-and-function-directivesalias.md
- 11.3.1-control-flow-directivesbranchtargets.md
- 11.3.2-control-flow-directivescalltargets.md
- 11.3.3-control-flow-directivescallprototype.md
- 11.4.1-performance-tuning-directivesmaxnreg.md
- 11.4.2-performance-tuning-directivesmaxntid.md
- 11.4.3-performance-tuning-directivesreqntid.md
- 11.4.4-performance-tuning-directivesminnctapersm.md
- 11.4.5-performance-tuning-directivesmaxnctapersmdeprecated.md
- 11.4.6-performance-tuning-directivesnoreturn.md
- 11.4.7-performance-tuning-directivespragma.md
- 11.4.8-performance-tuning-directivesabi_preserve.md
- 11.4.9-performance-tuning-directivesabi_preserve_control.md
- 11.5.1-debugging-directivesdwarf.md
- 11.5.2-debugging-directivessection.md
- 11.5.3-debugging-directivesfile.md
- 11.5.4-debugging-directivesloc.md
- 11.6.1-linking-directivesextern.md
- 11.6.2-linking-directivesvisible.md
- 11.6.3-linking-directivesweak.md
- 11.6.4-linking-directivescommon.md
- 11.7.1-cluster-dimension-directivesreqnctapercluster.md
- 11.7.2-cluster-dimension-directivesexplicitcluster.md
- 11.7.3-cluster-dimension-directivesmaxclusterrank.md
- 11.8.1-miscellaneous-directivesblocksareclusters.md
- 12.1-pragma-stringsnounroll.md
- 12.2-pragma-stringsused_bytes_mask.md
- 12.3-pragma-stringsenable_smem_spilling.md
- 12.4-pragma-stringsfrequency.md
- 13.1-changes-in-ptx-isa-version-91.md
- 13.10-changes-in-ptx-isa-version-81.md
- 13.11-changes-in-ptx-isa-version-80.md
- 13.12-changes-in-ptx-isa-version-78.md
- 13.13-changes-in-ptx-isa-version-77.md
- 13.14-changes-in-ptx-isa-version-76.md
- 13.15-changes-in-ptx-isa-version-75.md
- 13.16-changes-in-ptx-isa-version-74.md
- 13.17-changes-in-ptx-isa-version-73.md
- 13.18-changes-in-ptx-isa-version-72.md
- 13.19-changes-in-ptx-isa-version-71.md
- 13.2-changes-in-ptx-isa-version-90.md
- 13.20-changes-in-ptx-isa-version-70.md
- 13.21-changes-in-ptx-isa-version-65.md
- 13.22-changes-in-ptx-isa-version-64.md
- 13.23-changes-in-ptx-isa-version-63.md
- 13.24-changes-in-ptx-isa-version-62.md
- 13.25-changes-in-ptx-isa-version-61.md
- 13.26-changes-in-ptx-isa-version-60.md
- 13.27-changes-in-ptx-isa-version-50.md
- 13.28-changes-in-ptx-isa-version-43.md
- 13.29-changes-in-ptx-isa-version-42.md
- 13.3-changes-in-ptx-isa-version-88.md
- 13.30-changes-in-ptx-isa-version-41.md
- 13.31-changes-in-ptx-isa-version-40.md
- 13.32-changes-in-ptx-isa-version-32.md
- 13.33-changes-in-ptx-isa-version-31.md
- 13.34-changes-in-ptx-isa-version-30.md
- 13.35-changes-in-ptx-isa-version-23.md
- 13.36-changes-in-ptx-isa-version-22.md
- 13.37-changes-in-ptx-isa-version-21.md
- 13.38-changes-in-ptx-isa-version-20.md
- 13.4-changes-in-ptx-isa-version-87.md
- 13.5-changes-in-ptx-isa-version-86.md
- 13.6-changes-in-ptx-isa-version-85.md
- 13.7-changes-in-ptx-isa-version-84.md
- 13.8-changes-in-ptx-isa-version-83.md
- 13.9-changes-in-ptx-isa-version-82.md
- 14.1-notice.md
- 14.2-opencl.md
- 14.3-trademarks.md
- 2.1-a-highly-multithreaded-coprocessor.md
- 2.2.1-cooperative-thread-arrays.md
- 2.2.2-cluster-of-cooperative-thread-arrays.md
- 2.2.3-grid-of-clusters.md
- 2.3-memory-hierarchy.md
- 3.1-a-set-of-simt-multiprocessors.md
- 3.2-independent-thread-scheduling.md
- 3.3-on-chip-shared-memory.md
- 4.1-source-format.md
- 4.2-comments.md
- 4.3.1-directive-statements.md
- 4.3.2-instruction-statements.md
- 4.4-identifiers.md
- 4.5.1-integer-constants.md
- 4.5.2-floating-point-constants.md
- 4.5.3-predicate-constants.md
- 4.5.4-constant-expressions.md
- 4.5.5-integer-constant-expression-evaluation.md
- 4.5.6-summary-of-constant-expression-evaluation-rules.md
- 5.1.1-register-state-space.md
- 5.1.2-special-register-state-space.md
- 5.1.3.1-banked-constant-state-space-deprecated.md
- 5.1.4-global-state-space.md
- 5.1.5-local-state-space.md
- 5.1.6.1-kernel-function-parameters.md
- 5.1.6.2-kernel-function-parameter-attributes.md
- 5.1.6.3-kernel-parameter-attributeptr.md
- 5.1.6.4-device-function-parameters.md
- 5.1.7-shared-state-space.md
- 5.1.8-texture-state-space-deprecated.md
- 5.2.1-fundamental-types.md
- 5.2.2-restricted-use-of-sub-word-sizes.md
- 5.2.3-alternate-floating-point-data-formats.md
- 5.2.4-fixed-point-data-format.md
- 5.2.5.1-packed-floating-point-data-types.md
- 5.2.5.2-packed-integer-data-types.md
- 5.2.5.3-packed-fixed-point-data-types.md
- 5.3.1-texture-and-surface-properties.md
- 5.3.2-sampler-properties.md
- 5.3.3-channel-data-type-and-channel-order-fields.md
- 5.4.1-variable-declarations.md
- 5.4.2-vectors.md
- 5.4.3-array-declarations.md
- 5.4.4-initializers.md
- 5.4.5-alignment.md
- 5.4.6-parameterized-variable-names.md
- 5.4.7-variable-attributes.md
- 5.4.8-variable-and-function-attribute-directiveattribute.md
- 5.5.1.1-sub-byte-types.md
- 5.5.2-tensor-access-modes.md
- 5.5.3.1-bounding-box.md
- 5.5.3.2-traversal-stride.md
- 5.5.3.3-out-of-boundary-access.md
- 5.5.3.4-tilescatter4andtilegather4modes.md
- 5.5.4.1-bounding-box.md
- 5.5.4.2-traversal-stride.md
- 5.5.4.3-out-of-boundary-access.md
- 5.5.5.1-bounding-box.md
- 5.5.5.2-traversal-stride.md
- 5.5.5.3-whalo.md
- 5.5.5.4-woffset.md
- 5.5.6-interleave-layout.md
- 5.5.7-swizzling-modes.md
- 5.5.8-tensor-map.md
- 6.1-operand-type-information.md
- 6.2-source-operands.md
- 6.3-destination-operands.md
- 6.4.1.1-generic-addressing.md
- 6.4.2-arrays-as-operands.md
- 6.4.3-vectors-as-operands.md
- 6.4.4-labels-and-function-names-as-operands.md
- 6.5.1-scalar-conversions.md
- 6.5.2-rounding-modifiers.md
- 6.6-operand-costs.md
- 7.1.1-changes-from-ptx-isa-version-1x.md
- 7.2-variadic-functions.md
- 7.3-alloca.md
- 8.1.1-limitations-on-atomicity-at-system-scope.md
- 8.10.1-coherence.md
- 8.10.2-fence-sc.md
- 8.10.3-atomicity.md
- 8.10.4-no-thin-air.md
- 8.10.5-sequential-consistency-per-location.md
- 8.10.6-causality.md
- 8.11.1-reductions-do-not-form-acquire-patterns.md
- 8.2.1-overlap.md
- 8.2.2-aliases.md
- 8.2.3-multimem-addresses.md
- 8.2.4-memory-operations-on-vector-data-types.md
- 8.2.5-memory-operations-on-packed-data-types.md
- 8.2.6-initialization.md
- 8.3-state-spaces.md
- 8.4.1-mmio-operation.md
- 8.4.2-volatile-operation.md
- 8.5-scope.md
- 8.6-proxies.md
- 8.7.1-conflict-and-data-races.md
- 8.7.2-limitations-on-mixed-size-data-races.md
- 8.8-release-and-acquire-patterns.md
- 8.9.1.1-asynchronous-operations.md
- 8.9.2-observation-order.md
- 8.9.3-fence-sc-order.md
- 8.9.4-memory-synchronization.md
- 8.9.5-causality-order.md
- 8.9.6-coherence-order.md
- 8.9.7-communication-order.md
- 9.1-format-and-semantics-of-instruction-descriptions.md
- 9.2-ptx-instructions.md
- 9.3.1.1-integer-and-bit-size-comparisons.md
- 9.3.1.2-floating-point-comparisons.md
- 9.3.2-manipulating-predicates.md
- 9.4.1-operand-size-exceeding-instruction-type-size.md
- 9.5-divergence-of-threads-in-control-constructs.md
- 9.6.1-machine-specific-semantics-of-16-bit-code.md
- 9.7.1.1-integer-arithmetic-instructionsadd.md
- 9.7.1.10-integer-arithmetic-instructionsabs.md
- 9.7.1.11-integer-arithmetic-instructionsneg.md
- 9.7.1.12-integer-arithmetic-instructionsmin.md
- 9.7.1.13-integer-arithmetic-instructionsmax.md
- 9.7.1.14-integer-arithmetic-instructionspopc.md
- 9.7.1.15-integer-arithmetic-instructionsclz.md
- 9.7.1.16-integer-arithmetic-instructionsbfind.md
- 9.7.1.17-integer-arithmetic-instructionsfns.md
- 9.7.1.18-integer-arithmetic-instructionsbrev.md
- 9.7.1.19-integer-arithmetic-instructionsbfe.md
- 9.7.1.2-integer-arithmetic-instructionssub.md
- 9.7.1.20-integer-arithmetic-instructionsbfi.md
- 9.7.1.21-integer-arithmetic-instructionsszext.md
- 9.7.1.22-integer-arithmetic-instructionsbmsk.md
- 9.7.1.23-integer-arithmetic-instructionsdp4a.md
- 9.7.1.24-integer-arithmetic-instructionsdp2a.md
- 9.7.1.3-integer-arithmetic-instructionsmul.md
- 9.7.1.4-integer-arithmetic-instructionsmad.md
- 9.7.1.5-integer-arithmetic-instructionsmul24.md
- 9.7.1.6-integer-arithmetic-instructionsmad24.md
- 9.7.1.7-integer-arithmetic-instructionssad.md
- 9.7.1.8-integer-arithmetic-instructionsdiv.md
- 9.7.1.9-integer-arithmetic-instructionsrem.md
- 9.7.10.1-texturing-modes.md
- 9.7.10.2-mipmaps.md
- 9.7.10.3-texture-instructionstex.md
- 9.7.10.4-texture-instructionstld4.md
- 9.7.10.5-texture-instructionstxq.md
- 9.7.10.6-texture-instructionsistypep.md
- 9.7.11.1-surface-instructionssuld.md
- 9.7.11.2-surface-instructionssust.md
- 9.7.11.3-surface-instructionssured.md
- 9.7.11.4-surface-instructionssuq.md
- 9.7.12.1-control-flow-instructions.md
- 9.7.12.2-control-flow-instructions.md
- 9.7.12.3-control-flow-instructionsbra.md
- 9.7.12.4-control-flow-instructionsbrxidx.md
- 9.7.12.5-control-flow-instructionscall.md
- 9.7.12.6-control-flow-instructionsret.md
- 9.7.12.7-control-flow-instructionsexit.md
- 9.7.13.1-parallel-synchronization-and-communication-instructionsbarbarrier.md
- 9.7.13.10-parallel-synchronization-and-communication-instructionsmatchsync.md
- 9.7.13.11-parallel-synchronization-and-communication-instructionsactivemask.md
- 9.7.13.12-parallel-synchronization-and-communication-instructionsreduxsync.md
- 9.7.13.13-parallel-synchronization-and-communication-instructionsgriddepcontrol.md
- 9.7.13.14-parallel-synchronization-and-communication-instructionselectsync.md
- 9.7.13.15-parallel-synchronization-and-communication-instructionsmbarrier.md
- 9.7.13.16-parallel-synchronization-and-communication-instructionstensormapcp_fenceproxy.md
- 9.7.13.17-parallel-synchronization-and-communication-instructionsclusterlaunchcontroltry_cancel.md
- 9.7.13.18-parallel-synchronization-and-communication-instructionsclusterlaunchcontrolquery_cancel.md
- 9.7.13.2-parallel-synchronization-and-communication-instructionsbarwarpsync.md
- 9.7.13.3-parallel-synchronization-and-communication-instructionsbarriercluster.md
- 9.7.13.4-parallel-synchronization-and-communication-instructionsmembarfence.md
- 9.7.13.5-parallel-synchronization-and-communication-instructionsatom.md
- 9.7.13.6-parallel-synchronization-and-communication-instructionsred.md
- 9.7.13.7-parallel-synchronization-and-communication-instructionsredasync.md
- 9.7.13.8-parallel-synchronization-and-communication-instructionsvotedeprecated.md
- 9.7.13.9-parallel-synchronization-and-communication-instructionsvotesync.md
- 9.7.14.1-matrix-shape.md
- 9.7.14.2-matrix-data-types.md
- 9.7.14.3-block-scaling-formmasync.md
- 9.7.14.4-matrix-multiply-accumulate-operation-usingwmmainstructions.md
- 9.7.14.5-matrix-multiply-accumulate-operation-usingmmainstruction.md
- 9.7.14.6-matrix-multiply-accumulate-operation-usingmmaspinstruction-with-sparse-matrix-a.md
- 9.7.15.1-warpgroup.md
- 9.7.15.2-matrix-shape.md
- 9.7.15.3-matrix-data-types.md
- 9.7.15.4-async-proxy.md
- 9.7.15.5-asynchronous-warpgroup-level-matrix-multiply-accumulate-operation-usingwgmmamma_asyncinstruction.md
- 9.7.15.6-asynchronous-warpgroup-level-multiply-and-accumulate-operation-usingwgmmamma_asyncspinstruction.md
- 9.7.15.7-asynchronouswgmmaproxy-operations.md
- 9.7.16.1-tensor-memory.md
- 9.7.16.10-tensorcore-5th-generation-matrix-multiply-and-accumulate-operations.md
- 9.7.16.11-tensorcore-5th-generation-specialized-synchronization-operations.md
- 9.7.16.12-tensorcore-5th-generation-async-synchronization-operations.md
- 9.7.16.2-matrix-and-data-movement-shape.md
- 9.7.16.3-major-ness-supported-by-strides.md
- 9.7.16.4-matrix-descriptors.md
- 9.7.16.5-issue-granularity.md
- 9.7.16.6-memory-consistency-model-for-5th-generation-of-tensorcore-operations.md
- 9.7.16.7-tensor-memory-allocation-and-management-instructions.md
- 9.7.16.8-tensor-memory-and-register-loadstore-instructions.md
- 9.7.16.9-tensor-memory-data-movement-instructions.md
- 9.7.17.1-stack-manipulation-instructionsstacksave.md
- 9.7.17.2-stack-manipulation-instructionsstackrestore.md
- 9.7.17.3-stack-manipulation-instructionsalloca.md
- 9.7.18.1-scalar-video-instructions.md
- 9.7.18.2-simd-video-instructions.md
- 9.7.19.1-miscellaneous-instructionsbrkpt.md
- 9.7.19.2-miscellaneous-instructionsnanosleep.md
- 9.7.19.3-miscellaneous-instructionspmevent.md
- 9.7.19.4-miscellaneous-instructionstrap.md
- 9.7.19.5-miscellaneous-instructionssetmaxnreg.md
- 9.7.2.1-extended-precision-arithmetic-instructionsaddcc.md
- 9.7.2.2-extended-precision-arithmetic-instructionsaddc.md
- 9.7.2.3-extended-precision-arithmetic-instructionssubcc.md
- 9.7.2.4-extended-precision-arithmetic-instructionssubc.md
- 9.7.2.5-extended-precision-arithmetic-instructionsmadcc.md
- 9.7.2.6-extended-precision-arithmetic-instructionsmadc.md
- 9.7.3.1-floating-point-instructionstestp.md
- 9.7.3.10-floating-point-instructionsneg.md
- 9.7.3.11-floating-point-instructionsmin.md
- 9.7.3.12-floating-point-instructionsmax.md
- 9.7.3.13-floating-point-instructionsrcp.md
- 9.7.3.14-floating-point-instructionsrcpapproxftzf64.md
- 9.7.3.15-floating-point-instructionssqrt.md
- 9.7.3.16-floating-point-instructionsrsqrt.md
- 9.7.3.17-floating-point-instructionsrsqrtapproxftzf64.md
- 9.7.3.18-floating-point-instructionssin.md
- 9.7.3.19-floating-point-instructionscos.md
- 9.7.3.2-floating-point-instructionscopysign.md
- 9.7.3.20-floating-point-instructionslg2.md
- 9.7.3.21-floating-point-instructionsex2.md
- 9.7.3.22-floating-point-instructionstanh.md
- 9.7.3.3-floating-point-instructionsadd.md
- 9.7.3.4-floating-point-instructionssub.md
- 9.7.3.5-floating-point-instructionsmul.md
- 9.7.3.6-floating-point-instructionsfma.md
- 9.7.3.7-floating-point-instructionsmad.md
- 9.7.3.8-floating-point-instructionsdiv.md
- 9.7.3.9-floating-point-instructionsabs.md
- 9.7.4.1-half-precision-floating-point-instructionsadd.md
- 9.7.4.10-half-precision-floating-point-instructionsex2.md
- 9.7.4.2-half-precision-floating-point-instructionssub.md
- 9.7.4.3-half-precision-floating-point-instructionsmul.md
- 9.7.4.4-half-precision-floating-point-instructionsfma.md
- 9.7.4.5-half-precision-floating-point-instructionsneg.md
- 9.7.4.6-half-precision-floating-point-instructionsabs.md
- 9.7.4.7-half-precision-floating-point-instructionsmin.md
- 9.7.4.8-half-precision-floating-point-instructionsmax.md
- 9.7.4.9-half-precision-floating-point-instructionstanh.md
- 9.7.5.1-mixed-precision-floating-point-instructionsadd.md
- 9.7.5.2-mixed-precision-floating-point-instructionssub.md
- 9.7.5.3-mixed-precision-floating-point-instructionsfma.md
- 9.7.6.1-comparison-and-selection-instructionsset.md
- 9.7.6.2-comparison-and-selection-instructionssetp.md
- 9.7.6.3-comparison-and-selection-instructionsselp.md
- 9.7.6.4-comparison-and-selection-instructionsslct.md
- 9.7.7.1-half-precision-comparison-instructionsset.md
- 9.7.7.2-half-precision-comparison-instructionssetp.md
- 9.7.8.1-logic-and-shift-instructionsand.md
- 9.7.8.2-logic-and-shift-instructionsor.md
- 9.7.8.3-logic-and-shift-instructionsxor.md
- 9.7.8.4-logic-and-shift-instructionsnot.md
- 9.7.8.5-logic-and-shift-instructionscnot.md
- 9.7.8.6-logic-and-shift-instructionslop3.md
- 9.7.8.7-logic-and-shift-instructionsshf.md
- 9.7.8.8-logic-and-shift-instructionsshl.md
- 9.7.8.9-logic-and-shift-instructionsshr.md
- 9.7.9.1-cache-operators.md
- 9.7.9.10-data-movement-and-conversion-instructionsldu.md
- 9.7.9.11-data-movement-and-conversion-instructionsst.md
- 9.7.9.12-data-movement-and-conversion-instructionsstasync.md
- 9.7.9.13-data-movement-and-conversion-instructionsstbulk.md
- 9.7.9.14-data-movement-and-conversion-instructionsmultimemld_reducemultimemstmultimemred.md
- 9.7.9.15-data-movement-and-conversion-instructionsprefetchprefetchu.md
- 9.7.9.16-data-movement-and-conversion-instructionsapplypriority.md
- 9.7.9.17-data-movement-and-conversion-instructionsdiscard.md
- 9.7.9.18-data-movement-and-conversion-instructionscreatepolicy.md
- 9.7.9.19-data-movement-and-conversion-instructionsisspacep.md
- 9.7.9.2-cache-eviction-priority-hints.md
- 9.7.9.20-data-movement-and-conversion-instructionscvta.md
- 9.7.9.21-data-movement-and-conversion-instructionscvt.md
- 9.7.9.22-data-movement-and-conversion-instructionscvtpack.md
- 9.7.9.23-data-movement-and-conversion-instructionsmapa.md
- 9.7.9.24-data-movement-and-conversion-instructionsgetctarank.md
- 9.7.9.25-data-movement-and-conversion-instructions-asynchronous-copy.md
- 9.7.9.26-data-movement-and-conversion-instructionsmultimemcpasyncbulk.md
- 9.7.9.27-data-movement-and-conversion-instructionsmultimemcpreduceasyncbulk.md
- 9.7.9.28-data-movement-and-conversion-instructionstensormapreplace.md
- 9.7.9.3-data-movement-and-conversion-instructionsmov.md
- 9.7.9.4-data-movement-and-conversion-instructionsmov.md
- 9.7.9.5-data-movement-and-conversion-instructionsshfldeprecated.md
- 9.7.9.6-data-movement-and-conversion-instructionsshflsync.md
- 9.7.9.7-data-movement-and-conversion-instructionsprmt.md
- 9.7.9.8-data-movement-and-conversion-instructionsld.md
- 9.7.9.9-data-movement-and-conversion-instructionsldglobalnc.md
- INDEX.md
- cuda-driver.md
- cuda-runtime.md
- debugging-tools.md
- ncu-guide.md
- nsys-guide.md
- nvtx-patterns.md
- performance-traps.md
- ptx-isa.md
- SKILL.md
- axpby.cu
- axpby_binding.cpp
- __init__.py
- compile.py
- compile.sh
- profiling.py
- verification.py
- binding.cpp
- binding_registry.h
- model.py
- model_new.py
- SKILL.md
- ncu_analyse.py
- ncu_profile.sh
- SKILL.md
- SKILL.md
- SKILL.md
- SKILL.md
- SKILL.md
- cuda-optimization-strategies.md
- SKILL.md
- a100-optimization-guide.md
- diffusers-h100.md
- diffusers-integration.md
- h100-optimization-guide.md
- huggingface-kernels-integration.md
- kernel-templates.md
- t4-optimization-guide.md
- transformers-integration.md
- troubleshooting.md
- benchmark_example.py
- benchmark_rmsnorm.py
- huggingface_kernels_example.py
- ltx_kernel_injection_example.py
- transformers_injection_example.py
- manifest.txt
- SKILL.md
- 1-introduction.md
- 1.1-data-layout.md
- 1.2-new-and-legacy-cublas-api.md
- 1.3-example-code.md
- 1.4-forward-compatibility.md
- 1.5-floating-point-emulation.md
- 1.5.1-bf16x9.md
- 1.5.2-fixed-point.md
- 1.5.2.1-dynamic-mantissa-control.md
- 1.5.2.2-fixed-mantissa-control.md
- 1.5.2.3-representation-and-mappings.md
- 1.5.2.4-fixed-point-workspace-requirements.md
- 1.5.2.5-fixed-point-performance-guide.md
- 1.5.3-default-library-configurations.md
- 1.5.4-support-for-floating-point-special-values.md
- 2-using-the-cublas-api.md
- 2.1-general-description.md
- 2.1.1-error-status.md
- 2.1.10-gemm-algorithms-numerical-behavior.md
- 2.1.11-tensor-core-usage.md
- 2.1.12-cuda-graphs-support.md
- 2.1.13-bit-integer-interface.md
- 2.1.2-cublas-context.md
- 2.1.3-thread-safety.md
- 2.1.4-results-reproducibility.md
- 2.1.5-scalar-parameters.md
- 2.1.6-parallelism-with-streams.md
- 2.1.7-batching-kernels.md
- 2.1.8-cache-configuration.md
- 2.1.9-static-library-support.md
- 2.2-cublas-datatypes-reference.md
- 2.2.1-cublashandle_t.md
- 2.2.10-cublasmath_t.md
- 2.2.11-cublascomputetype_t.md
- 2.2.12-cublasemulationstrategy_t.md
- 2.2.2-cublasstatus_t.md
- 2.2.3-cublasoperation_t.md
- 2.2.4-cublasfillmode_t.md
- 2.2.5-cublasdiagtype_t.md
- 2.2.6-cublassidemode_t.md
- 2.2.7-cublaspointermode_t.md
- 2.2.8-cublasatomicsmode_t.md
- 2.2.9-cublasgemmalgo_t.md
- 2.3-cuda-datatypes-reference.md
- 2.3.1-cudadatatype_t.md
- 2.3.2-cudaemulationstrategy_t.md
- 2.3.3-cudaemulationmantissacontrol_t.md
- 2.3.4-cudaemulationspecialvaluessupport_t.md
- 2.3.5-librarypropertytype_t.md
- 2.4-cublas-helper-function-reference.md
- 2.4.1-cublascreate.md
- 2.4.10-cublasgetpointermode.md
- 2.4.11-cublassetpointermode.md
- 2.4.12-cublassetvector.md
- 2.4.13-cublasgetvector.md
- 2.4.14-cublassetmatrix.md
- 2.4.15-cublasgetmatrix.md
- 2.4.16-cublassetvectorasync.md
- 2.4.17-cublasgetvectorasync.md
- 2.4.18-cublassetmatrixasync.md
- 2.4.19-cublasgetmatrixasync.md
- 2.4.2-cublasdestroy.md
- 2.4.20-cublassetatomicsmode.md
- 2.4.21-cublasgetatomicsmode.md
- 2.4.22-cublassetmathmode.md
- 2.4.23-cublasgetmathmode.md
- 2.4.24-cublassetsmcounttarget.md
- 2.4.25-cublasgetsmcounttarget.md
- 2.4.26-cublassetemulationstrategy.md
- 2.4.27-cublasgetemulationstrategy.md
- 2.4.28-cublasgetemulationspecialvaluessupport.md
- 2.4.29-cublassetemulationspecialvaluessupport.md
- 2.4.3-cublasgetversion.md
- 2.4.30-cublasgetfixedpointemulationmantissacontrol.md
- 2.4.31-cublassetfixedpointemulationmantissacontrol.md
- 2.4.32-cublasgetfixedpointemulationmaxmantissabitcount.md
- 2.4.33-cublassetfixedpointemulationmaxmantissabitcount.md
- 2.4.34-cublasgetfixedpointemulationmantissabitoffset.md
- 2.4.35-cublassetfixedpointemulationmantissabitoffset.md
- 2.4.36-cublasgetfixedpointemulationmantissabitcountpointer.md
- 2.4.37-cublassetfixedpointemulationmantissabitcountpointer.md
- 2.4.38-cublasloggerconfigure.md
- 2.4.39-cublasgetloggercallback.md
- 2.4.4-cublasgetproperty.md
- 2.4.40-cublassetloggercallback.md
- 2.4.5-cublasgetstatusname.md
- 2.4.6-cublasgetstatusstring.md
- 2.4.7-cublassetstream.md
- 2.4.8-cublassetworkspace.md
- 2.4.9-cublasgetstream.md
- 2.5-cublas-level-1-function-reference.md
- 2.5.1-cublasitamax.md
- 2.5.10-cublastrotm.md
- 2.5.11-cublastrotmg.md
- 2.5.12-cublastscal.md
- 2.5.13-cublastswap.md
- 2.5.2-cublasitamin.md
- 2.5.3-cublastasum.md
- 2.5.4-cublastaxpy.md
- 2.5.5-cublastcopy.md
- 2.5.6-cublastdot.md
- 2.5.7-cublastnrm2.md
- 2.5.8-cublastrot.md
- 2.5.9-cublastrotg.md
- 2.6-cublas-level-2-function-reference.md
- 2.6.1-cublastgbmv.md
- 2.6.10-cublastsyr2.md
- 2.6.11-cublasttbmv.md
- 2.6.12-cublasttbsv.md
- 2.6.13-cublasttpmv.md
- 2.6.14-cublasttpsv.md
- 2.6.15-cublasttrmv.md
- 2.6.16-cublasttrsv.md
- 2.6.17-cublasthemv.md
- 2.6.18-cublasthbmv.md
- 2.6.19-cublasthpmv.md
- 2.6.2-cublastgemv.md
- 2.6.20-cublasther.md
- 2.6.21-cublasther2.md
- 2.6.22-cublasthpr.md
- 2.6.23-cublasthpr2.md
- 2.6.24-cublastgemvbatched.md
- 2.6.25-cublastgemvstridedbatched.md
- 2.6.3-cublastger.md
- 2.6.4-cublastsbmv.md
- 2.6.5-cublastspmv.md
- 2.6.6-cublastspr.md
- 2.6.7-cublastspr2.md
- 2.6.8-cublastsymv.md
- 2.6.9-cublastsyr.md
- 2.7-cublas-level-3-function-reference.md
- 2.7.1-cublastgemm.md
- 2.7.10-cublasttrmm.md
- 2.7.11-cublasttrsm.md
- 2.7.12-cublasttrsmbatched.md
- 2.7.13-cublasthemm.md
- 2.7.14-cublastherk.md
- 2.7.15-cublasther2k.md
- 2.7.16-cublastherkx.md
- 2.7.2-cublastgemm3m.md
- 2.7.3-cublastgemmbatched.md
- 2.7.4-cublastgemmstridedbatched.md
- 2.7.5-cublastgemmgroupedbatched.md
- 2.7.6-cublastsymm.md
- 2.7.7-cublastsyrk.md
- 2.7.8-cublastsyr2k.md
- 2.7.9-cublastsyrkx.md
- 2.8-blas-like-extension.md
- 2.8.1-cublastgeam.md
- 2.8.10-cublasttrttp.md
- 2.8.11-cublastgemmex.md
- 2.8.12-cublasgemmex.md
- 2.8.13-cublasgemmbatchedex.md
- 2.8.14-cublasgemmstridedbatchedex.md
- 2.8.15-cublasgemmgroupedbatchedex.md
- 2.8.16-cublascsyrkex.md
- 2.8.17-cublascsyrk3mex.md
- 2.8.18-cublascherkex.md
- 2.8.19-cublascherk3mex.md
- 2.8.2-cublastdgmm.md
- 2.8.20-cublasnrm2ex.md
- 2.8.21-cublasaxpyex.md
- 2.8.22-cublasdotex.md
- 2.8.23-cublasrotex.md
- 2.8.24-cublasscalex.md
- 2.8.3-cublastgetrfbatched.md
- 2.8.4-cublastgetrsbatched.md
- 2.8.5-cublastgetribatched.md
- 2.8.6-cublastmatinvbatched.md
- 2.8.7-cublastgeqrfbatched.md
- 2.8.8-cublastgelsbatched.md
- 2.8.9-cublasttpttr.md
- 3-using-the-cublaslt-api.md
- 3.1-general-description.md
- 3.1.1-problem-size-limitations.md
- 3.1.2-heuristics-cache.md
- 3.1.3-cublaslt-logging.md
- 3.1.4-narrow-precision-data-types-usage.md
- 3.1.4.1-tensorwide-scaling-for-fp8-data-types.md
- 3.1.4.2-experimental-per-batch-tensorwide-scaling-for-fp8-data-types.md
- 3.1.4.3-outer-vector-scaling-for-fp8-data-types.md
- 3.1.4.4-32-element-1d-block-scaling-for-fp8-and-fp4-data-types.md
- 3.1.4.5-element-1d-and-128x128-2d-block-scaling-for-fp8-data-types.md
- 3.1.5-disabling-cpu-instructions.md
- 3.2-cublaslt-code-examples.md
- 3.3-cublaslt-datatypes-reference.md
- 3.3.1-cublasltclustershape_t.md
- 3.3.10-cublasltmatmulheuristicresult_t.md
- 3.3.11-cublasltmatmulinnershape_t.md
- 3.3.12-cublasltmatmulpreference_t.md
- 3.3.13-cublasltmatmulpreferenceattributes_t.md
- 3.3.14-cublasltmatmulsearch_t.md
- 3.3.15-cublasltmatmultile_t.md
- 3.3.16-cublasltmatmulstages_t.md
- 3.3.17-cublasltnumericalimplflags_t.md
- 3.3.18-cublasltmatrixlayout_t.md
- 3.3.19-cublasltmatrixlayoutattribute_t.md
- 3.3.2-cublasltepilogue_t.md
- 3.3.20-cublasltintegerwidth_t.md
- 3.3.21-cublasltmatrixtransformdesc_t.md
- 3.3.22-cublasltmatrixtransformdescattributes_t.md
- 3.3.23-cublasltorder_t.md
- 3.3.24-cublasltpointermode_t.md
- 3.3.25-cublasltpointermodemask_t.md
- 3.3.26-cublasltreductionscheme_t.md
- 3.3.27-cublasltmatmulmatrixscale_t.md
- 3.3.28-cublasltbatchmode_t.md
- 3.3.29-cublasltemulationdesc_t.md
- 3.3.3-cublaslthandle_t.md
- 3.3.30-cublasltemulationdescattributes_t.md
- 3.3.4-cublasltloggercallback_t.md
- 3.3.5-cublasltmatmulalgo_t.md
- 3.3.6-cublasltmatmulalgocapattributes_t.md
- 3.3.7-cublasltmatmulalgoconfigattributes_t.md
- 3.3.8-cublasltmatmuldesc_t.md
- 3.3.9-cublasltmatmuldescattributes_t.md
- 3.4-cublaslt-api-reference.md
- 3.4.1-cublasltcreate.md
- 3.4.10-cublasltgetversion.md
- 3.4.11-cublasltloggersetcallback.md
- 3.4.12-cublasltloggersetfile.md
- 3.4.13-cublasltloggeropenfile.md
- 3.4.14-cublasltloggersetlevel.md
- 3.4.15-cublasltloggersetmask.md
- 3.4.16-cublasltloggerforcedisable.md
- 3.4.17-cublasltmatmul.md
- 3.4.18-cublasltmatmulalgocapgetattribute.md
- 3.4.19-cublasltmatmulalgocheck.md
- 3.4.2-cublasltdestroy.md
- 3.4.20-cublasltmatmulalgoconfiggetattribute.md
- 3.4.21-cublasltmatmulalgoconfigsetattribute.md
- 3.4.22-cublasltmatmulalgogetheuristic.md
- 3.4.23-cublasltmatmulalgogetids.md
- 3.4.24-cublasltmatmulalgoinit.md
- 3.4.25-cublasltmatmuldesccreate.md
- 3.4.26-cublasltmatmuldescinit.md
- 3.4.27-cublasltmatmuldescdestroy.md
- 3.4.28-cublasltmatmuldescgetattribute.md
- 3.4.29-cublasltmatmuldescsetattribute.md
- 3.4.3-cublasltdisablecpuinstructionssetmask.md
- 3.4.30-cublasltmatmulpreferencecreate.md
- 3.4.31-cublasltmatmulpreferenceinit.md
- 3.4.32-cublasltmatmulpreferencedestroy.md
- 3.4.33-cublasltmatmulpreferencegetattribute.md
- 3.4.34-cublasltmatmulpreferencesetattribute.md
- 3.4.35-cublasltmatrixlayoutcreate.md
- 3.4.36-cublasltmatrixlayoutinit.md
- 3.4.37-cublasltgroupedmatrixlayoutcreate.md
- 3.4.38-cublasltgroupedmatrixlayoutinit.md
- 3.4.39-cublasltmatrixlayoutdestroy.md
- 3.4.4-cublasltgetcudartversion.md
- 3.4.40-cublasltmatrixlayoutgetattribute.md
- 3.4.41-cublasltmatrixlayoutsetattribute.md
- 3.4.42-cublasltmatrixtransform.md
- 3.4.43-cublasltmatrixtransformdesccreate.md
- 3.4.44-cublasltmatrixtransformdescinit.md
- 3.4.45-cublasltmatrixtransformdescdestroy.md
- 3.4.46-cublasltmatrixtransformdescgetattribute.md
- 3.4.47-cublasltmatrixtransformdescsetattribute.md
- 3.4.48-cublasltemulationdescinit.md
- 3.4.49-cublasltemulationdesccreate.md
- 3.4.5-cublasltgetproperty.md
- 3.4.50-cublasltemulationdescdestroy.md
- 3.4.51-cublasltemulationdescsetattribute.md
- 3.4.52-cublasltemulationdescgetattribute.md
- 3.4.6-cublasltgetstatusname.md
- 3.4.7-cublasltgetstatusstring.md
- 3.4.8-cublasltheuristicscachegetcapacity.md
- 3.4.9-cublasltheuristicscachesetcapacity.md
- 4-using-the-cublasxt-api.md
- 4.1-general-description.md
- 4.1.1-tiling-design-approach.md
- 4.1.2-hybrid-cpu-gpu-computation.md
- 4.1.3-results-reproducibility.md
- 4.2-cublasxt-api-datatypes-reference.md
- 4.2.1-cublasxthandle_t.md
- 4.2.2-cublasxtoptype_t.md
- 4.2.3-cublasxtblasop_t.md
- 4.2.4-cublasxtpinningmemmode_t.md
- 4.3-cublasxt-api-helper-function-reference.md
- 4.3.1-cublasxtcreate.md
- 4.3.2-cublasxtdestroy.md
- 4.3.3-cublasxtdeviceselect.md
- 4.3.4-cublasxtsetblockdim.md
- 4.3.5-cublasxtgetblockdim.md
- 4.3.6-cublasxtsetcpuroutine.md
- 4.3.7-cublasxtsetcpuratio.md
- 4.3.8-cublasxtsetpinningmemmode.md
- 4.3.9-cublasxtgetpinningmemmode.md
- 4.4-cublasxt-api-math-functions-reference.md
- 4.4.1-cublasxttgemm.md
- 4.4.10-cublasxtttrsm.md
- 4.4.11-cublasxtttrmm.md
- 4.4.12-cublasxttspmm.md
- 4.4.2-cublasxtthemm.md
- 4.4.3-cublasxttsymm.md
- 4.4.4-cublasxttsyrk.md
- 4.4.5-cublasxttsyr2k.md
- 4.4.6-cublasxttsyrkx.md
- 4.4.7-cublasxttherk.md
- 4.4.8-cublasxtther2k.md
- 4.4.9-cublasxttherkx.md
- 5-using-the-cublasdx-api.md
- 6-using-the-cublas-legacy-api.md
- 6.1-error-status.md
- 6.2-initialization-and-shutdown.md
- 6.3-thread-safety.md
- 6.4-memory-management.md
- 6.5-scalar-parameters.md
- 6.6-helper-functions.md
- 6.7-level-123-functions.md
- 6.8-converting-legacy-to-the-cublas-api.md
- 6.9-examples.md
- 7-cublas-fortran-bindings.md
- 10.1-notice.md
- 10.2-opencl.md
- 10.3-trademarks.md
- 8-interaction-with-other-libraries-and-tools.md
- 8.1-nvprune.md
- structcu__dev__sm__resource__group__params.md
- structcuaccesspolicywindow__v1.md
- structcuarraymapinfo__v1.md
- structcuasyncnotificationinfo.md
- structcucheckpointcheckpointargs.md
- structcucheckpointgpupair.md
- structcucheckpointlockargs.md
- structcucheckpointrestoreargs.md
- structcucheckpointunlockargs.md
- structcuctxcigparam.md
- structcuctxcreateparams.md
- structcuda__array3d__descriptor__v2.md
- structcuda__array__descriptor__v2.md
- structcuda__array__memory__requirements__v1.md
- structcuda__array__sparse__properties__v1.md
- structcuda__batch__mem__op__node__params__v1.md
- structcuda__child__graph__node__params.md
- structcuda__conditional__node__params.md
- structcuda__event__record__node__params.md
- structcuda__event__wait__node__params.md
- structcuda__ext__sem__signal__node__params__v1.md
- structcuda__ext__sem__signal__node__params__v2.md
- structcuda__ext__sem__wait__node__params__v1.md
- structcuda__ext__sem__wait__node__params__v2.md
- structcuda__external__memory__buffer__desc__v1.md
- structcuda__external__memory__handle__desc__v1.md
- structcuda__external__memory__mipmapped__array__desc__v1.md
- structcuda__external__semaphore__handle__desc__v1.md
- structcuda__external__semaphore__signal__params__v1.md
- structcuda__external__semaphore__wait__params__v1.md
- structcuda__graph__instantiate__params.md
- structcuda__host__node__params__v1.md
- structcuda__host__node__params__v2.md
- structcuda__kernel__node__params__v1.md
- structcuda__kernel__node__params__v2.md
- structcuda__kernel__node__params__v3.md
- structcuda__launch__params__v1.md
- structcuda__mem__alloc__node__params__v1.md
- structcuda__mem__alloc__node__params__v2.md
- structcuda__mem__free__node__params.md
- structcuda__memcpy2d__v2.md
- structcuda__memcpy3d__peer__v1.md
- structcuda__memcpy3d__v2.md
- structcuda__memcpy__node__params.md
- structcuda__memset__node__params__v1.md
- structcuda__memset__node__params__v2.md
- structcuda__pointer__attribute__p2p__tokens__v1.md
- structcuda__resource__desc__v1.md
- structcuda__resource__view__desc__v1.md
- structcuda__texture__desc__v1.md
- structcudevprop__v1.md
- structcudevresource.md
- structcudevsmresource.md
- structcudevworkqueueconfigresource.md
- structcudevworkqueueresource.md
- structcueglframe__v1.md
- structcuexecaffinityparam__v1.md
- structcuexecaffinitysmcount__v1.md
- structcuextent3d__v1.md
- structcugraphedgedata.md
- structcugraphexecupdateresultinfo__v1.md
- structcugraphnodeparams.md
- structcuipceventhandle__v1.md
- structcuipcmemhandle__v1.md
- structculaunchattribute.md
- structculaunchconfig.md
- structculaunchmemsyncdomainmap.md
- structcumemaccessdesc__v1.md
- structcumemallocationprop__v1.md
- structcumemcpy3doperand__v1.md
- structcumemcpyattributes__v1.md
- structcumemdecompressparams.md
- structcumemfabrichandle__v1.md
- structcumemlocation__v1.md
- structcumempoolprops__v1.md
- structcumempoolptrexportdata__v1.md
- structcumulticastobjectprop__v1.md
- structcuoffset3d__v1.md
- structcutensormap.md
- group__cuda__checkpoint.md
- group__cuda__coredump.md
- group__cuda__ctx.md
- group__cuda__ctx__deprecated.md
- group__cuda__d3d10.md
- group__cuda__d3d10__deprecated.md
- group__cuda__d3d11.md
- group__cuda__d3d11__deprecated.md
- group__cuda__d3d9.md
- group__cuda__d3d9__deprecated.md
- group__cuda__device.md
- group__cuda__device__deprecated.md
- group__cuda__driver__entry__point.md
- group__cuda__egl.md
- group__cuda__error.md
- group__cuda__event.md
- group__cuda__exec.md
- group__cuda__exec__deprecated.md
- group__cuda__extres__interop.md
- group__cuda__gl.md
- group__cuda__gl__deprecated.md
- group__cuda__graph.md
- group__cuda__graphics.md
- group__cuda__green__contexts.md
- group__cuda__initialize.md
- group__cuda__library.md
- group__cuda__logs.md
- group__cuda__malloc__async.md
- group__cuda__mem.md
- group__cuda__memop.md
- group__cuda__module.md
- group__cuda__module__deprecated.md
- group__cuda__multicast.md
- group__cuda__occupancy.md
- group__cuda__peer__access.md
- group__cuda__primary__ctx.md
- group__cuda__profiler.md
- group__cuda__profiler__deprecated.md
- group__cuda__stream.md
- group__cuda__surfobject.md
- group__cuda__surfref__deprecated.md
- group__cuda__tensor__memory.md
- group__cuda__texobject.md
- group__cuda__texref__deprecated.md
- group__cuda__types.md
- group__cuda__unified.md
- group__cuda__va.md
- group__cuda__vdpau.md
- group__cuda__version.md
- INDEX.md
- struct____half.md
- struct____half2.md
- struct____half2__raw.md
- struct____half__raw.md
- struct____nv__bfloat16.md
- struct____nv__bfloat162.md
- struct____nv__bfloat162__raw.md
- struct____nv__bfloat16__raw.md
- struct____nv__fp4__e2m1.md
- struct____nv__fp4x2__e2m1.md
- struct____nv__fp4x4__e2m1.md
- struct____nv__fp6__e2m3.md
- struct____nv__fp6__e3m2.md
- struct____nv__fp6x2__e2m3.md
- struct____nv__fp6x2__e3m2.md
- struct____nv__fp6x4__e2m3.md
- struct____nv__fp6x4__e3m2.md
- struct____nv__fp8__e4m3.md
- struct____nv__fp8__e5m2.md
- struct____nv__fp8__e8m0.md
- struct____nv__fp8x2__e4m3.md
- struct____nv__fp8x2__e5m2.md
- struct____nv__fp8x2__e8m0.md
- struct____nv__fp8x4__e4m3.md
- struct____nv__fp8x4__e5m2.md
- struct____nv__fp8x4__e8m0.md
- group__cuda__math__double.md
- group__cuda__math__int.md
- group__cuda__math__intrinsic__bfloat16.md
- group__cuda__math__intrinsic__cast.md
- group__cuda__math__intrinsic__double.md
- group__cuda__math__intrinsic__fp4.md
- group__cuda__math__intrinsic__fp6.md
- group__cuda__math__intrinsic__fp8.md
- group__cuda__math__intrinsic__half.md
- group__cuda__math__intrinsic__int.md
- group__cuda__math__intrinsic__simd.md
- group__cuda__math__intrinsic__single.md
- group__cuda__math__quad.md
- group__cuda__math__single.md
- INDEX.md
- structcudaaccesspolicywindow.md
- structcudaarraymemoryrequirements.md
- structcudaarraysparseproperties.md
- structcudaasyncnotificationinfo__t.md
- structcudachannelformatdesc.md
- structcudachildgraphnodeparams.md
- structcudaconditionalnodeparams.md
- structcudadeviceprop.md
- structcudadevresource.md
- structcudadevsmresource.md
- structcudadevsmresourcegroupparams.md
- structcudadevworkqueueconfigresource.md
- structcudadevworkqueueresource.md
- structcudaeglframe.md
- structcudaeglplanedesc.md
- structcudaeventrecordnodeparams.md
- structcudaeventwaitnodeparams.md
- structcudaextent.md
- structcudaexternalmemorybufferdesc.md
- structcudaexternalmemoryhandledesc.md
- structcudaexternalmemorymipmappedarraydesc.md
- structcudaexternalsemaphorehandledesc.md
- structcudaexternalsemaphoresignalnodeparams.md
- structcudaexternalsemaphoresignalnodeparamsv2.md
- structcudaexternalsemaphoresignalparams.md
- structcudaexternalsemaphorewaitnodeparams.md
- structcudaexternalsemaphorewaitnodeparamsv2.md
- structcudaexternalsemaphorewaitparams.md
- structcudafuncattributes.md
- structcudagraphedgedata.md
- structcudagraphexecupdateresultinfo.md
- structcudagraphinstantiateparams.md
- structcudagraphkernelnodeupdate.md
- structcudagraphnodeparams.md
- structcudahostnodeparams.md
- structcudahostnodeparamsv2.md
- structcudaipceventhandle__t.md
- structcudaipcmemhandle__t.md
- structcudakernelnodeparams.md
- structcudakernelnodeparamsv2.md
- structcudalaunchattribute.md
- structcudalaunchconfig__t.md
- structcudalaunchmemsyncdomainmap.md
- structcudamemaccessdesc.md
- structcudamemallocnodeparams.md
- structcudamemallocnodeparamsv2.md
- structcudamemcpy3doperand.md
- structcudamemcpy3dparms.md
- structcudamemcpy3dpeerparms.md
- structcudamemcpyattributes.md
- structcudamemcpynodeparams.md
- structcudamemfreenodeparams.md
- structcudamemlocation.md
- structcudamempoolprops.md
- structcudamempoolptrexportdata.md
- structcudamemsetparams.md
- structcudamemsetparamsv2.md
- structcudaoffset3d.md
- structcudapitchedptr.md
- structcudapointerattributes.md
- structcudapos.md
- structcudaresourcedesc.md
- structcudaresourceviewdesc.md
- structcudatexturedesc.md
- structcuuuid__st.md
- unioncudalaunchattributevalue.md
- group__cudart____version.md
- group__cudart__d3d10.md
- group__cudart__d3d10__deprecated.md
- group__cudart__d3d11.md
- group__cudart__d3d11__deprecated.md
- group__cudart__d3d9.md
- group__cudart__d3d9__deprecated.md
- group__cudart__device.md
- group__cudart__device__deprecated.md
- group__cudart__driver.md
- group__cudart__driver__entry__point.md
- group__cudart__egl.md
- group__cudart__error.md
- group__cudart__event.md
- group__cudart__execution.md
- group__cudart__execution__context.md
- group__cudart__execution__deprecated.md
- group__cudart__extres__interop.md
- group__cudart__graph.md
- group__cudart__highlevel.md
- group__cudart__interop.md
- group__cudart__library.md
- group__cudart__logs.md
- group__cudart__memory.md
- group__cudart__memory__deprecated.md
- group__cudart__memory__pools.md
- group__cudart__occupancy.md
- group__cudart__opengl.md
- group__cudart__opengl__deprecated.md
- group__cudart__peer.md
- group__cudart__profiler.md
- group__cudart__stream.md
- group__cudart__surface__object.md
- group__cudart__texture__object.md
- group__cudart__types.md
- group__cudart__unified.md
- group__cudart__vdpau.md
- INDEX.md
- colls.md
- comms.md
- device.md
- device_gin.md
- device_memory.md
- device_reducecopy.md
- device_setup.md
- flags.md
- group.md
- ops.md
- p2p.md
- types.md
- ras.md
- bufferreg.md
- collectives.md
- communicators.md
- cudagraph.md
- data.md
- deviceapi.md
- groups.md
- inplace.md
- p2p.md
- streams.md
- threadsafety.md
- api.md
- env.md
- examples.md
- INDEX.md
- mpi.md
- nccl1.md
- overview.md
- setup.md
- troubleshooting.md
- usage.md
- 1.1-scalable-data-parallel-computing-using-gpus.md
- 1.2-goals-of-ptx.md
- 1.3-ptx-isa-version-91.md
- 1.4-document-structure.md
- 10.1-special-registerstid.md
- 10.10-special-registersgridid.md
- 10.11-special-registersis_explicit_cluster.md
- 10.12-special-registersclusterid.md
- 10.13-special-registersnclusterid.md
- 10.14-special-registerscluster_ctaid.md
- 10.15-special-registerscluster_nctaid.md
- 10.16-special-registerscluster_ctarank.md
- 10.17-special-registerscluster_nctarank.md
- 10.18-special-registerslanemask_eq.md
- 10.19-special-registerslanemask_le.md
- 10.2-special-registersntid.md
- 10.20-special-registerslanemask_lt.md
- 10.21-special-registerslanemask_ge.md
- 10.22-special-registerslanemask_gt.md
- 10.23-special-registersclockclock_hi.md
- 10.24-special-registersclock64.md
- 10.25-special-registerspm0pm7.md
- 10.26-special-registerspm0_64pm7_64.md
- 10.27-special-registersenvreg32.md
- 10.28-special-registersglobaltimerglobaltimer_loglobaltimer_hi.md
- 10.29-special-registersreserved_smem_offset_beginreserved_smem_offset_endreserved_smem_offset_capreserved_smem_offset_2.md
- 10.3-special-registerslaneid.md
- 10.30-special-registerstotal_smem_size.md
- 10.31-special-registersaggr_smem_size.md
- 10.32-special-registersdynamic_smem_size.md
- 10.33-special-registerscurrent_graph_exec.md
- 10.4-special-registerswarpid.md
- 10.5-special-registersnwarpid.md
- 10.6-special-registersctaid.md
- 10.7-special-registersnctaid.md
- 10.8-special-registerssmid.md
- 10.9-special-registersnsmid.md
- 11.1.1-ptx-module-directivesversion.md
- 11.1.2-ptx-module-directivestarget.md
- 11.1.3-ptx-module-directivesaddress_size.md
- 11.2.1-kernel-and-function-directivesentry.md
- 11.2.2-kernel-and-function-directivesfunc.md
- 11.2.3-kernel-and-function-directivesalias.md
- 11.3.1-control-flow-directivesbranchtargets.md
- 11.3.2-control-flow-directivescalltargets.md
- 11.3.3-control-flow-directivescallprototype.md
- 11.4.1-performance-tuning-directivesmaxnreg.md
- 11.4.2-performance-tuning-directivesmaxntid.md
- 11.4.3-performance-tuning-directivesreqntid.md
- 11.4.4-performance-tuning-directivesminnctapersm.md
- 11.4.5-performance-tuning-directivesmaxnctapersmdeprecated.md
- 11.4.6-performance-tuning-directivesnoreturn.md
- 11.4.7-performance-tuning-directivespragma.md
- 11.4.8-performance-tuning-directivesabi_preserve.md
- 11.4.9-performance-tuning-directivesabi_preserve_control.md
- 11.5.1-debugging-directivesdwarf.md
- 11.5.2-debugging-directivessection.md
- 11.5.3-debugging-directivesfile.md
- 11.5.4-debugging-directivesloc.md
- 11.6.1-linking-directivesextern.md
- 11.6.2-linking-directivesvisible.md
- 11.6.3-linking-directivesweak.md
- 11.6.4-linking-directivescommon.md
- 11.7.1-cluster-dimension-directivesreqnctapercluster.md
- 11.7.2-cluster-dimension-directivesexplicitcluster.md
- 11.7.3-cluster-dimension-directivesmaxclusterrank.md
- 11.8.1-miscellaneous-directivesblocksareclusters.md
- 12.1-pragma-stringsnounroll.md
- 12.2-pragma-stringsused_bytes_mask.md
- 12.3-pragma-stringsenable_smem_spilling.md
- 12.4-pragma-stringsfrequency.md
- 13.1-changes-in-ptx-isa-version-91.md
- 13.10-changes-in-ptx-isa-version-81.md
- 13.11-changes-in-ptx-isa-version-80.md
- 13.12-changes-in-ptx-isa-version-78.md
- 13.13-changes-in-ptx-isa-version-77.md
- 13.14-changes-in-ptx-isa-version-76.md
- 13.15-changes-in-ptx-isa-version-75.md
- 13.16-changes-in-ptx-isa-version-74.md
- 13.17-changes-in-ptx-isa-version-73.md
- 13.18-changes-in-ptx-isa-version-72.md
- 13.19-changes-in-ptx-isa-version-71.md
- 13.2-changes-in-ptx-isa-version-90.md
- 13.20-changes-in-ptx-isa-version-70.md
- 13.21-changes-in-ptx-isa-version-65.md
- 13.22-changes-in-ptx-isa-version-64.md
- 13.23-changes-in-ptx-isa-version-63.md
- 13.24-changes-in-ptx-isa-version-62.md
- 13.25-changes-in-ptx-isa-version-61.md
- 13.26-changes-in-ptx-isa-version-60.md
- 13.27-changes-in-ptx-isa-version-50.md
- 13.28-changes-in-ptx-isa-version-43.md
- 13.29-changes-in-ptx-isa-version-42.md
- 13.3-changes-in-ptx-isa-version-88.md
- 13.30-changes-in-ptx-isa-version-41.md
- 13.31-changes-in-ptx-isa-version-40.md
- 13.32-changes-in-ptx-isa-version-32.md
- 13.33-changes-in-ptx-isa-version-31.md
- 13.34-changes-in-ptx-isa-version-30.md
- 13.35-changes-in-ptx-isa-version-23.md
- 13.36-changes-in-ptx-isa-version-22.md
- 13.37-changes-in-ptx-isa-version-21.md
- 13.38-changes-in-ptx-isa-version-20.md
- 13.4-changes-in-ptx-isa-version-87.md
- 13.5-changes-in-ptx-isa-version-86.md
- 13.6-changes-in-ptx-isa-version-85.md
- 13.7-changes-in-ptx-isa-version-84.md
- 13.8-changes-in-ptx-isa-version-83.md
- 13.9-changes-in-ptx-isa-version-82.md
- 14.1-notice.md
- 14.2-opencl.md
- 14.3-trademarks.md
- 2.1-a-highly-multithreaded-coprocessor.md
- 2.2.1-cooperative-thread-arrays.md
- 2.2.2-cluster-of-cooperative-thread-arrays.md
- 2.2.3-grid-of-clusters.md
- 2.3-memory-hierarchy.md
- 3.1-a-set-of-simt-multiprocessors.md
- 3.2-independent-thread-scheduling.md
- 3.3-on-chip-shared-memory.md
- 4.1-source-format.md
- 4.2-comments.md
- 4.3.1-directive-statements.md
- 4.3.2-instruction-statements.md
- 4.4-identifiers.md
- 4.5.1-integer-constants.md
- 4.5.2-floating-point-constants.md
- 4.5.3-predicate-constants.md
- 4.5.4-constant-expressions.md
- 4.5.5-integer-constant-expression-evaluation.md
- 4.5.6-summary-of-constant-expression-evaluation-rules.md
- 5.1.1-register-state-space.md
- 5.1.2-special-register-state-space.md
- 5.1.3.1-banked-constant-state-space-deprecated.md
- 5.1.4-global-state-space.md
- 5.1.5-local-state-space.md
- 5.1.6.1-kernel-function-parameters.md
- 5.1.6.2-kernel-function-parameter-attributes.md
- 5.1.6.3-kernel-parameter-attributeptr.md
- 5.1.6.4-device-function-parameters.md
- 5.1.7-shared-state-space.md
- 5.1.8-texture-state-space-deprecated.md
- 5.2.1-fundamental-types.md
- 5.2.2-restricted-use-of-sub-word-sizes.md
- 5.2.3-alternate-floating-point-data-formats.md
- 5.2.4-fixed-point-data-format.md
- 5.2.5.1-packed-floating-point-data-types.md
- 5.2.5.2-packed-integer-data-types.md
- 5.2.5.3-packed-fixed-point-data-types.md
- 5.3.1-texture-and-surface-properties.md
- 5.3.2-sampler-properties.md
- 5.3.3-channel-data-type-and-channel-order-fields.md
- 5.4.1-variable-declarations.md
- 5.4.2-vectors.md
- 5.4.3-array-declarations.md
- 5.4.4-initializers.md
- 5.4.5-alignment.md
- 5.4.6-parameterized-variable-names.md
- 5.4.7-variable-attributes.md
- 5.4.8-variable-and-function-attribute-directiveattribute.md
- 5.5.1.1-sub-byte-types.md
- 5.5.2-tensor-access-modes.md
- 5.5.3.1-bounding-box.md
- 5.5.3.2-traversal-stride.md
- 5.5.3.3-out-of-boundary-access.md
- 5.5.3.4-tilescatter4andtilegather4modes.md
- 5.5.4.1-bounding-box.md
- 5.5.4.2-traversal-stride.md
- 5.5.4.3-out-of-boundary-access.md
- 5.5.5.1-bounding-box.md
- 5.5.5.2-traversal-stride.md
- 5.5.5.3-whalo.md
- 5.5.5.4-woffset.md
- 5.5.6-interleave-layout.md
- 5.5.7-swizzling-modes.md
- 5.5.8-tensor-map.md
- 6.1-operand-type-information.md
- 6.2-source-operands.md
- 6.3-destination-operands.md
- 6.4.1.1-generic-addressing.md
- 6.4.2-arrays-as-operands.md
- 6.4.3-vectors-as-operands.md
- 6.4.4-labels-and-function-names-as-operands.md
- 6.5.1-scalar-conversions.md
- 6.5.2-rounding-modifiers.md
- 6.6-operand-costs.md
- 7.1.1-changes-from-ptx-isa-version-1x.md
- 7.2-variadic-functions.md
- 7.3-alloca.md
- 8.1.1-limitations-on-atomicity-at-system-scope.md
- 8.10.1-coherence.md
- 8.10.2-fence-sc.md
- 8.10.3-atomicity.md
- 8.10.4-no-thin-air.md
- 8.10.5-sequential-consistency-per-location.md
- 8.10.6-causality.md
- 8.11.1-reductions-do-not-form-acquire-patterns.md
- 8.2.1-overlap.md
- 8.2.2-aliases.md
- 8.2.3-multimem-addresses.md
- 8.2.4-memory-operations-on-vector-data-types.md
- 8.2.5-memory-operations-on-packed-data-types.md
- 8.2.6-initialization.md
- 8.3-state-spaces.md
- 8.4.1-mmio-operation.md
- 8.4.2-volatile-operation.md
- 8.5-scope.md
- 8.6-proxies.md
- 8.7.1-conflict-and-data-races.md
- 8.7.2-limitations-on-mixed-size-data-races.md
- 8.8-release-and-acquire-patterns.md
- 8.9.1.1-asynchronous-operations.md
- 8.9.2-observation-order.md
- 8.9.3-fence-sc-order.md
- 8.9.4-memory-synchronization.md
- 8.9.5-causality-order.md
- 8.9.6-coherence-order.md
- 8.9.7-communication-order.md
- 9.1-format-and-semantics-of-instruction-descriptions.md
- 9.2-ptx-instructions.md
- 9.3.1.1-integer-and-bit-size-comparisons.md
- 9.3.1.2-floating-point-comparisons.md
- 9.3.2-manipulating-predicates.md
- 9.4.1-operand-size-exceeding-instruction-type-size.md
- 9.5-divergence-of-threads-in-control-constructs.md
- 9.6.1-machine-specific-semantics-of-16-bit-code.md
- 9.7.1.1-integer-arithmetic-instructionsadd.md
- 9.7.1.10-integer-arithmetic-instructionsabs.md
- 9.7.1.11-integer-arithmetic-instructionsneg.md
- 9.7.1.12-integer-arithmetic-instructionsmin.md
- 9.7.1.13-integer-arithmetic-instructionsmax.md
- 9.7.1.14-integer-arithmetic-instructionspopc.md
- 9.7.1.15-integer-arithmetic-instructionsclz.md
- 9.7.1.16-integer-arithmetic-instructionsbfind.md
- 9.7.1.17-integer-arithmetic-instructionsfns.md
- 9.7.1.18-integer-arithmetic-instructionsbrev.md
- 9.7.1.19-integer-arithmetic-instructionsbfe.md
- 9.7.1.2-integer-arithmetic-instructionssub.md
- 9.7.1.20-integer-arithmetic-instructionsbfi.md
- 9.7.1.21-integer-arithmetic-instructionsszext.md
- 9.7.1.22-integer-arithmetic-instructionsbmsk.md
- 9.7.1.23-integer-arithmetic-instructionsdp4a.md
- 9.7.1.24-integer-arithmetic-instructionsdp2a.md
- 9.7.1.3-integer-arithmetic-instructionsmul.md
- 9.7.1.4-integer-arithmetic-instructionsmad.md
- 9.7.1.5-integer-arithmetic-instructionsmul24.md
- 9.7.1.6-integer-arithmetic-instructionsmad24.md
- 9.7.1.7-integer-arithmetic-instructionssad.md
- 9.7.1.8-integer-arithmetic-instructionsdiv.md
- 9.7.1.9-integer-arithmetic-instructionsrem.md
- 9.7.10.1-texturing-modes.md
- 9.7.10.2-mipmaps.md
- 9.7.10.3-texture-instructionstex.md
- 9.7.10.4-texture-instructionstld4.md
- 9.7.10.5-texture-instructionstxq.md
- 9.7.10.6-texture-instructionsistypep.md
- 9.7.11.1-surface-instructionssuld.md
- 9.7.11.2-surface-instructionssust.md
- 9.7.11.3-surface-instructionssured.md
- 9.7.11.4-surface-instructionssuq.md
- 9.7.12.1-control-flow-instructions.md
- 9.7.12.2-control-flow-instructions.md
- 9.7.12.3-control-flow-instructionsbra.md
- 9.7.12.4-control-flow-instructionsbrxidx.md
- 9.7.12.5-control-flow-instructionscall.md
- 9.7.12.6-control-flow-instructionsret.md
- 9.7.12.7-control-flow-instructionsexit.md
- 9.7.13.1-parallel-synchronization-and-communication-instructionsbarbarrier.md
- 9.7.13.10-parallel-synchronization-and-communication-instructionsmatchsync.md
- 9.7.13.11-parallel-synchronization-and-communication-instructionsactivemask.md
- 9.7.13.12-parallel-synchronization-and-communication-instructionsreduxsync.md
- 9.7.13.13-parallel-synchronization-and-communication-instructionsgriddepcontrol.md
- 9.7.13.14-parallel-synchronization-and-communication-instructionselectsync.md
- 9.7.13.15-parallel-synchronization-and-communication-instructionsmbarrier.md
- 9.7.13.16-parallel-synchronization-and-communication-instructionstensormapcp_fenceproxy.md
- 9.7.13.17-parallel-synchronization-and-communication-instructionsclusterlaunchcontroltry_cancel.md
- 9.7.13.18-parallel-synchronization-and-communication-instructionsclusterlaunchcontrolquery_cancel.md
- 9.7.13.2-parallel-synchronization-and-communication-instructionsbarwarpsync.md
- 9.7.13.3-parallel-synchronization-and-communication-instructionsbarriercluster.md
- 9.7.13.4-parallel-synchronization-and-communication-instructionsmembarfence.md
- 9.7.13.5-parallel-synchronization-and-communication-instructionsatom.md
- 9.7.13.6-parallel-synchronization-and-communication-instructionsred.md
- 9.7.13.7-parallel-synchronization-and-communication-instructionsredasync.md
- 9.7.13.8-parallel-synchronization-and-communication-instructionsvotedeprecated.md
- 9.7.13.9-parallel-synchronization-and-communication-instructionsvotesync.md
- 9.7.14.1-matrix-shape.md
- 9.7.14.2-matrix-data-types.md
- 9.7.14.3-block-scaling-formmasync.md
- 9.7.14.4-matrix-multiply-accumulate-operation-usingwmmainstructions.md
- 9.7.14.5-matrix-multiply-accumulate-operation-usingmmainstruction.md
- 9.7.14.6-matrix-multiply-accumulate-operation-usingmmaspinstruction-with-sparse-matrix-a.md
- 9.7.15.1-warpgroup.md
- 9.7.15.2-matrix-shape.md
- 9.7.15.3-matrix-data-types.md
- 9.7.15.4-async-proxy.md
- 9.7.15.5-asynchronous-warpgroup-level-matrix-multiply-accumulate-operation-usingwgmmamma_asyncinstruction.md
- 9.7.15.6-asynchronous-warpgroup-level-multiply-and-accumulate-operation-usingwgmmamma_asyncspinstruction.md
- 9.7.15.7-asynchronouswgmmaproxy-operations.md
- 9.7.16.1-tensor-memory.md
- 9.7.16.10-tensorcore-5th-generation-matrix-multiply-and-accumulate-operations.md
- 9.7.16.11-tensorcore-5th-generation-specialized-synchronization-operations.md
- 9.7.16.12-tensorcore-5th-generation-async-synchronization-operations.md
- 9.7.16.2-matrix-and-data-movement-shape.md
- 9.7.16.3-major-ness-supported-by-strides.md
- 9.7.16.4-matrix-descriptors.md
- 9.7.16.5-issue-granularity.md
- 9.7.16.6-memory-consistency-model-for-5th-generation-of-tensorcore-operations.md
- 9.7.16.7-tensor-memory-allocation-and-management-instructions.md
- 9.7.16.8-tensor-memory-and-register-loadstore-instructions.md
- 9.7.16.9-tensor-memory-data-movement-instructions.md
- 9.7.17.1-stack-manipulation-instructionsstacksave.md
- 9.7.17.2-stack-manipulation-instructionsstackrestore.md
- 9.7.17.3-stack-manipulation-instructionsalloca.md
- 9.7.18.1-scalar-video-instructions.md
- 9.7.18.2-simd-video-instructions.md
- 9.7.19.1-miscellaneous-instructionsbrkpt.md
- 9.7.19.2-miscellaneous-instructionsnanosleep.md
- 9.7.19.3-miscellaneous-instructionspmevent.md
- 9.7.19.4-miscellaneous-instructionstrap.md
- 9.7.19.5-miscellaneous-instructionssetmaxnreg.md
- 9.7.2.1-extended-precision-arithmetic-instructionsaddcc.md
- 9.7.2.2-extended-precision-arithmetic-instructionsaddc.md
- 9.7.2.3-extended-precision-arithmetic-instructionssubcc.md
- 9.7.2.4-extended-precision-arithmetic-instructionssubc.md
- 9.7.2.5-extended-precision-arithmetic-instructionsmadcc.md
- 9.7.2.6-extended-precision-arithmetic-instructionsmadc.md
- 9.7.3.1-floating-point-instructionstestp.md
- 9.7.3.10-floating-point-instructionsneg.md
- 9.7.3.11-floating-point-instructionsmin.md
- 9.7.3.12-floating-point-instructionsmax.md
- 9.7.3.13-floating-point-instructionsrcp.md
- 9.7.3.14-floating-point-instructionsrcpapproxftzf64.md
- 9.7.3.15-floating-point-instructionssqrt.md
- 9.7.3.16-floating-point-instructionsrsqrt.md
- 9.7.3.17-floating-point-instructionsrsqrtapproxftzf64.md
- 9.7.3.18-floating-point-instructionssin.md
- 9.7.3.19-floating-point-instructionscos.md
- 9.7.3.2-floating-point-instructionscopysign.md
- 9.7.3.20-floating-point-instructionslg2.md
- 9.7.3.21-floating-point-instructionsex2.md
- 9.7.3.22-floating-point-instructionstanh.md
- 9.7.3.3-floating-point-instructionsadd.md
- 9.7.3.4-floating-point-instructionssub.md
- 9.7.3.5-floating-point-instructionsmul.md
- 9.7.3.6-floating-point-instructionsfma.md
- 9.7.3.7-floating-point-instructionsmad.md
- 9.7.3.8-floating-point-instructionsdiv.md
- 9.7.3.9-floating-point-instructionsabs.md
- 9.7.4.1-half-precision-floating-point-instructionsadd.md
- 9.7.4.10-half-precision-floating-point-instructionsex2.md
- 9.7.4.2-half-precision-floating-point-instructionssub.md
- 9.7.4.3-half-precision-floating-point-instructionsmul.md
- 9.7.4.4-half-precision-floating-point-instructionsfma.md
- 9.7.4.5-half-precision-floating-point-instructionsneg.md
- 9.7.4.6-half-precision-floating-point-instructionsabs.md
- 9.7.4.7-half-precision-floating-point-instructionsmin.md
- 9.7.4.8-half-precision-floating-point-instructionsmax.md
- 9.7.4.9-half-precision-floating-point-instructionstanh.md
- 9.7.5.1-mixed-precision-floating-point-instructionsadd.md
- 9.7.5.2-mixed-precision-floating-point-instructionssub.md
- 9.7.5.3-mixed-precision-floating-point-instructionsfma.md
- 9.7.6.1-comparison-and-selection-instructionsset.md
- 9.7.6.2-comparison-and-selection-instructionssetp.md
- 9.7.6.3-comparison-and-selection-instructionsselp.md
- 9.7.6.4-comparison-and-selection-instructionsslct.md
- 9.7.7.1-half-precision-comparison-instructionsset.md
- 9.7.7.2-half-precision-comparison-instructionssetp.md
- 9.7.8.1-logic-and-shift-instructionsand.md
- 9.7.8.2-logic-and-shift-instructionsor.md
- 9.7.8.3-logic-and-shift-instructionsxor.md
- 9.7.8.4-logic-and-shift-instructionsnot.md
- 9.7.8.5-logic-and-shift-instructionscnot.md
- 9.7.8.6-logic-and-shift-instructionslop3.md
- 9.7.8.7-logic-and-shift-instructionsshf.md
- 9.7.8.8-logic-and-shift-instructionsshl.md
- 9.7.8.9-logic-and-shift-instructionsshr.md
- 9.7.9.1-cache-operators.md
- 9.7.9.10-data-movement-and-conversion-instructionsldu.md
- 9.7.9.11-data-movement-and-conversion-instructionsst.md
- 9.7.9.12-data-movement-and-conversion-instructionsstasync.md
- 9.7.9.13-data-movement-and-conversion-instructionsstbulk.md
- 9.7.9.14-data-movement-and-conversion-instructionsmultimemld_reducemultimemstmultimemred.md
- 9.7.9.15-data-movement-and-conversion-instructionsprefetchprefetchu.md
- 9.7.9.16-data-movement-and-conversion-instructionsapplypriority.md
- 9.7.9.17-data-movement-and-conversion-instructionsdiscard.md
- 9.7.9.18-data-movement-and-conversion-instructionscreatepolicy.md
- 9.7.9.19-data-movement-and-conversion-instructionsisspacep.md
- 9.7.9.2-cache-eviction-priority-hints.md
- 9.7.9.20-data-movement-and-conversion-instructionscvta.md
- 9.7.9.21-data-movement-and-conversion-instructionscvt.md
- 9.7.9.22-data-movement-and-conversion-instructionscvtpack.md
- 9.7.9.23-data-movement-and-conversion-instructionsmapa.md
- 9.7.9.24-data-movement-and-conversion-instructionsgetctarank.md
- 9.7.9.25-data-movement-and-conversion-instructions-asynchronous-copy.md
- 9.7.9.26-data-movement-and-conversion-instructionsmultimemcpasyncbulk.md
- 9.7.9.27-data-movement-and-conversion-instructionsmultimemcpreduceasyncbulk.md
- 9.7.9.28-data-movement-and-conversion-instructionstensormapreplace.md
- 9.7.9.3-data-movement-and-conversion-instructionsmov.md
- 9.7.9.4-data-movement-and-conversion-instructionsmov.md
- 9.7.9.5-data-movement-and-conversion-instructionsshfldeprecated.md
- 9.7.9.6-data-movement-and-conversion-instructionsshflsync.md
- 9.7.9.7-data-movement-and-conversion-instructionsprmt.md
- 9.7.9.8-data-movement-and-conversion-instructionsld.md
- 9.7.9.9-data-movement-and-conversion-instructionsldglobalnc.md
- INDEX.md
- cublas.md
- cuda-driver.md
- cuda-math.md
- cuda-runtime.md
- debugging-tools.md
- nccl.md
- ncu-guide.md
- nsys-guide.md
- nvtx-patterns.md
- performance-traps.md
- ptx-isa.md
- SKILL.md
- SKILL.md
- samples-advanced-topics.md
- samples-cuda-graphs.md
- samples-cuda-libraries.md
- samples-framework-interop.md
- samples-getting-started.md
- samples-multi-gpu.md
- samples-performance.md
- samples-reduction-scan-sort.md
- samples-streams-async.md
- samples-tensor-core-gemm.md
- SKILL.md
- config.yml
- EVAL.md
- evals.json
- BENCHMARK.md
- skill-card.md
- SKILL.md
- skill.oms.sig
- SKILL.md
- update-cutlass.sh
- skill.json
- SKILL.md
- skill.json
- SKILL.md
- debug-tools-doc.md
- error-catalog.md
- interpreting-debug-output.md
- SKILL.md
- benchmark-patterns.md
- naming-conventions.md
- registration-patterns.md
- review-rules-judgment.md.deprecated
- review-rules-structural.md
- test-patterns.md
- fetch_pr_diff.py
- post_review.py
- review_operator.py
- LICENSE.txt
- README.md
- SKILL.md
- common-issues.md
- naming.md
- pr-checklist.md
- pr-template.md
- workflow.md
- check_operator.py
- check_overload_consistency.py
- extract_from_worktree.py
- gen_pr_description.py
- operator_registry.py
- pr_gate_check.sh
- submit_operator.py
- LICENSE.txt
- README.md
- SKILL.md
- SKILL.md
- gemm-optimization.md
- kernel-authoring.md
- layout-algebra.md
- reduction-and-norms.md
- SKILL.md
- skill.json
- SKILL.md
- SKILL.md
- SKILL.md
- skill.json
- SKILL.md
- SKILL.md
- examples.md
- primer.md
- schema.md
- _wiki_root.py
- collect_contest_code.py
- compute_core_prs.py
- extract_blog_code.py
- fetch_pr_diff.py
- kbs.py
- kbs_checks.py
- kbs_db.py
- refresh_candidate_ledger.py
- scripts.md
- update_pr_corpus.py
- validate.py
- verify_core_prs.py
- verify_verbatim.py
- corpus.yaml
- layout.yaml
- pr-update.yaml
- MANIFEST.yaml
- MANIFEST.yaml
- 01-tmem-allocation-tcgen05-mma-single-thread-launch.cu
- 02-tmem-load-into-registers-for-epilogue.cu
- 03-cutlass-mma-atom-wrapping.cpp
- PROVENANCE.yaml
- MANIFEST.yaml
- 01-nc-128-cuda-core-promotion-hopper-sm90.cpp
- 02-sm100-path-tcgen05-mma-with-ue8m0-block-scaling.cpp
- 03-moe-grouped-gemm-launch.cpp
- PROVENANCE.yaml
- MANIFEST.yaml
- 01-software-exp-cody-waite-horner.cu
- 02-2-cta-cooperative-backward.cu
- PROVENANCE.yaml
- MANIFEST.yaml
- 01-mla-decode-inner-loop.cu
- 02-sparse-mla-kv-retrieval-kernel-v3-2.cu
- PROVENANCE.yaml
- MANIFEST.yaml
- 01-chunk-parallel-prefill-reference-pytorch.py
- 02-triton-decode-step-kernel-streaming.py
- PROVENANCE.yaml
- MANIFEST.yaml
- MANIFEST.yaml
- MANIFEST.yaml
- MANIFEST.yaml
- MANIFEST.yaml
- MANIFEST.yaml
- MANIFEST.yaml
- MANIFEST.yaml
- 01-reference-core-computation-part-1.cpp
- 02-strategy-1-k-parallel-grid-with-atomic-accumulation.cpp
- 03-strategy-3-atomic-free-shared-memory-reduction.cpp
- PROVENANCE.yaml
- MANIFEST.yaml
- 01-basic-tcgen05-mma-kernel-17-of-peak.cu
- 02-128b-swizzling-46-of-peak.cu
- 03-pipelining-mbarrier-phases-62-of-peak.cu
- PROVENANCE.yaml
- MANIFEST.yaml
- 01-2-blackwell-specific-cutlass-schedules-and-tma.cu
- PROVENANCE.yaml
- MANIFEST.yaml
- MANIFEST.yaml
- MANIFEST.yaml
- 01-step-1-cute-dsl-baseline-100us.cpp
- 02-step-2-coalesced-memory-access-443us-then-improved.cpp
- 03-step-3-hardware-intrinsics-39us.cpp
- 04-step-4-ptx-assembly-27us.ptx
- 05-step-5-ilp-optimization-22-9us.cpp
- PROVENANCE.yaml
- MANIFEST.yaml
- 01-reference-core-computation-part-1.cpp
- 02-strategy-1-k-parallel-grid-with-atomic-accumulation.cpp
- 03-strategy-3-atomic-free-shared-memory-reduction.cpp
- PROVENANCE.yaml
- 01-step-1-cute-dsl-baseline-100us.cpp
- 02-step-2-coalesced-memory-access-443us-then-improved.cpp
- 03-step-3-hardware-intrinsics-39us.cpp
- 04-step-4-ptx-assembly-27us.ptx
- 05-step-5-ilp-optimization-22-9us.cpp
- PROVENANCE.yaml
- PROVENANCE.yaml
- sglang-PR-21239-deepgemm-integration.patch
- sm100_fp8_gemm_1d1d.cuh
- sm90_fp8_gemm_1d1d.cuh
- 01-nc128-cuda-core-promotion.cpp
- PROVENANCE.yaml
- nvfp4_scaled_mm_kernels.cu
- PROVENANCE.yaml
- tmem-load-into-registers-for-epilogue.cu
- 01-double-buffered-tmem-epilogue-skeleton.cu
- PROVENANCE.yaml
- PROVENANCE.yaml
- sm100_fmha_bwd_mla_kernel_tma_warpspecialized.hpp
- 01-software-exp-skeleton.cu
- PROVENANCE.yaml
- 77_blackwell_mla_fwd.cu
- PROVENANCE.yaml
- sm100_fmha_mla_tma_warpspecialized.hpp
- 01-mla-decode-inner-loop.cu
- PROVENANCE.yaml
- flashinfer_cutedsl.py
- moe-grouped-gemm-launch.cpp
- PROVENANCE.yaml
- vllm-PR-23696-dual-gemm.patch
- 01-routing-plus-fusion-skeleton.py
- PROVENANCE.yaml
- gdn_fused_proj.py
- PROVENANCE.yaml
- 01-chunk-parallel-prefill-reference-pytorch.py
- 02-triton-decode-step-kernel-streaming.py
- PROVENANCE.yaml
- blackwell-cutlass-schedules-and-tma.cu
- PROVENANCE.yaml
- vllm-PR-23696-gated-dual-gemm.patch
- 01-fused-epilogue-swiglu-skeleton.cu
- PROVENANCE.yaml
- blackwell-cutlass-schedules-and-tma.cu
- PR-2139-blockwise-groupwise-gemm.patch
- PROVENANCE.yaml
- 01-nvfp4-gemm-skeleton.cu
- PROVENANCE.yaml
- nvfp4_scaled_mm_kernels.cu
- PROVENANCE.yaml
- 01-step-1-cute-dsl-baseline-100us.cpp
- 02-step-2-coalesced-memory-access-443us-then-improved.cpp
- 03-step-3-hardware-intrinsics-39us.cpp
- 04-step-4-ptx-assembly-27us.ptx
- 05-step-5-ilp-optimization-22-9us.cpp
- PROVENANCE.yaml
- dense_gemm_persistent_prefetch.py
- PR-2161-persistent-scheduler-clc.patch
- PROVENANCE.yaml
- 01-clc-persistent-loop-skeleton.cu
- PROVENANCE.yaml
- PROVENANCE.yaml
- sm100_fmha_bwd_kernel_tma_warpspecialized.hpp
- 01-two-tile-ping-pong-skeleton.cu
- PROVENANCE.yaml
- PROVENANCE.yaml
- sm100_fmha_fwd_mainloop_tma_warpspecialized.hpp
- 01-producer-consumer-skeleton.cu
- PROVENANCE.yaml
- 67_hopper_fp8_warp_specialized_gemm_with_blockwise_scaling.cu
- 67_hopper_fp8_warp_specialized_gemm_with_groupwise_scaling.cu
- hopper_fp8_commandline.hpp
- gemm_with_groupwise_scaling.h
- 68_hopper_fp8_warp_specialized_grouped_gemm_with_blockwise_scaling.cu
- hopper_fp8_commandline.hpp
- 81_blackwell_gemm_blockwise.cu
- 81_blackwell_gemm_groupwise.cu
- sm100_blockwise_scale_layout.hpp
- sm100_blockwise_umma_builder.inl
- sm90_gmma_builder.inl
- collective_builder.hpp
- collective_mma.hpp
- fp8_accumulation.hpp
- sm100_mma_warpspecialized_blockwise_scaling.hpp
- sm90_mma_array_tma_gmma_rs_warpspecialized_mixed_input.hpp
- sm90_mma_array_tma_gmma_ss_warpspecialized.hpp
- sm90_mma_array_tma_gmma_ss_warpspecialized_fp8.hpp
- sm90_mma_array_tma_gmma_ss_warpspecialized_fp8_blockwise_scaling.hpp
- sm90_mma_tma_gmma_ss_warpspecialized_fp8_blockwise_scaling.hpp
- gemm_universal.hpp
- sm100_gemm_tma_warpspecialized_mma_transform.hpp
- sm90_gemm_array_tma_warpspecialized_cooperative.hpp
- sm90_gemm_array_tma_warpspecialized_pingpong.hpp
- dispatch_policy.hpp
- diff.patch
- PROVENANCE.yaml
- grid_dependency_control.h
- sm100_mma_warpspecialized_blockwise_scaling.hpp
- diff.patch
- PROVENANCE.yaml
- mma_sm89.hpp
- mma_traits_sm89.hpp
- diff.patch
- PROVENANCE.yaml
- fmha_fusion.hpp
- fmha_device_bwd.hpp
- sm100_mla.hpp
- sm100_fmha_bwd_kernel_tma_warpspecialized.hpp
- sm100_fmha_bwd_mla_kernel_tma_warpspecialized.hpp
- sm100_fmha_mla_reduction.hpp
- sm100_fmha_mla_tma_warpspecialized.hpp
- sm100_mla_tile_scheduler.hpp
- 77_blackwell_fmha.cu
- 77_blackwell_fmha_bwd.cu
- 77_blackwell_mla.cu
- 77_blackwell_mla_fwd.cu
- diff.patch
- PROVENANCE.yaml
- fmha_fusion.hpp
- sm100_fmha_fwd_epilogue_tma_warpspecialized.hpp
- sm100_fmha_fwd_mainloop_tma_warpspecialized.hpp
- sm100_fmha_mla_fwd_mainloop_tma_warpspecialized.hpp
- sm100_fmha_mla_load_tma_warpspecialized.hpp
- pipeline_mla.hpp
- fmha_causal_tile_scheduler.hpp
- sm100_fmha_bwd_kernel_tma_warpspecialized.hpp
- sm100_fmha_fwd_kernel_tma_warpspecialized.hpp
- fmha_fwd_reference.hpp
- 77_blackwell_fmha.cu
- 77_blackwell_mla_fwd.cu
- diff.patch
- PROVENANCE.yaml
- 04_mma_tma_2sm_sm100.cu
- diff.patch
- PROVENANCE.yaml
- fmha.py
- diff.patch
- PROVENANCE.yaml
- fp16_gemm_1.py
- diff.patch
- PROVENANCE.yaml
- dense_blockscaled_gemm_persistent_prefetch.py
- dense_gemm_persistent_prefetch.py
- diff.patch
- PROVENANCE.yaml
- 01_mma_sm100.cu
- diff.patch
- PROVENANCE.yaml
- dense_blockscaled_gemm_persistent.py
- dense_blockscaled_gemm_persistent_prefetch.py
- diff.patch
- PROVENANCE.yaml
- clc.py
- diff.patch
- PROVENANCE.yaml
- grouped_gemm.py
- tensormap_manager.py
- diff.patch
- PROVENANCE.yaml
- fp16_gemm_0.py
- fp16_gemm_1.py
- fp16_gemm_2.py
- fp16_gemm_3.py
- fp16_gemm_3_1.py
- fp16_gemm_4.py
- fp16_gemm_5.py
- fp16_gemm_6.py
- diff.patch
- PROVENANCE.yaml
- fp16_gemm_2.py
- fp16_gemm_3.py
- fp16_gemm_3_1.py
- fp16_gemm_4.py
- fp16_gemm_5.py
- fp16_gemm_6.py
- dense_blockscaled_gemm_persistent_prefetch.py
- dense_gemm_persistent_prefetch.py
- distributed_all_gather_gemm_blackwell.py
- distributed_gemm_all_reduce_blackwell.py
- distributed_gemm_reduce_scatter_blackwell.py
- diff.patch
- PROVENANCE.yaml
- attention.hpp
- einsum.hpp
- gemm.hpp
- hyperconnection.hpp
- layout.hpp
- mega.hpp
- runtime.hpp
- main.cu
- cache.hpp
- compiler.hpp
- device_runtime.hpp
- handle.hpp
- include_parser.hpp
- kernel_runtime.hpp
- common.hpp
- config.hpp
- mega_moe.hpp
- runtime.hpp
- sm100.hpp
- sm90.hpp
- utils.hpp
- epilogue.hpp
- runtime_utils.hpp
- sm100_bf16_gemm.hpp
- sm100_bmk_bnk_mn.hpp
- sm100_fp8_fp4_gemm_1d1d.hpp
- sm100_fp8_fp4_mega_moe.hpp
- sm100_tf32_hc_prenorm_gemm.hpp
- sm90_bf16_gemm.hpp
- sm90_bmk_bnk_mn.hpp
- sm90_fp8_gemm_1d1d.hpp
- sm90_fp8_gemm_1d2d.hpp
- sm90_tf32_hc_prenorm_gemm.hpp
- smxx_clean_logits.hpp
- smxx_cublaslt.hpp
- smxx_fp8_fp4_mqa_logits.hpp
- smxx_fp8_fp4_paged_mqa_logits.hpp
- smxx_layout.hpp
- exception.hpp
- hash.hpp
- layout.hpp
- math.hpp
- system.hpp
- python_api.cpp
- barrier.cuh
- compile.cuh
- cute_tie.cuh
- exception.cuh
- math.cuh
- tma_copy.cuh
- types.cuh
- utils.cuh
- sm100_store_cd.cuh
- sm100_store_cd_swap_ab.cuh
- transform.cuh
- sm100_bf16_gemm.cuh
- sm100_bmk_bnk_mn.cuh
- sm100_fp4_mqa_logits.cuh
- sm100_fp4_paged_mqa_logits.cuh
- sm100_fp8_fp4_gemm_1d1d.cuh
- sm100_fp8_fp4_mega_moe.cuh
- sm100_fp8_mqa_logits.cuh
- sm100_fp8_paged_mqa_logits.cuh
- sm100_tf32_hc_prenorm_gemm.cuh
- sm90_bf16_gemm.cuh
- sm90_bmk_bnk_mn.cuh
- sm90_fp8_gemm_1d1d.cuh
- sm90_fp8_gemm_1d2d.cuh
- sm90_fp8_mqa_logits.cuh
- sm90_fp8_paged_mqa_logits.cuh
- sm90_tf32_hc_prenorm_gemm.cuh
- smxx_clean_logits.cuh
- smxx_layout.cuh
- mega_moe.cuh
- sym_buffer.cuh
- sm100.cuh
- sm90.cuh
- ld_st.cuh
- tcgen05.cuh
- tma.cuh
- utils.cuh
- wgmma.cuh
- gemm.cuh
- mega_moe.cuh
- paged_mqa_logits.cuh
- __init__.py
- swiglu_apply_weight_to_fp8.py
- utils.py
- diff.patch
- PROVENANCE.yaml
- __init__.py
- format_conversion.py
- diff.patch
- PROVENANCE.yaml
- fmha_cutlass_sm100.cu
- fmha_cutlass_sm100_pybind.cu
- __init__.py
- format_conversion.py
- fmha_common.hpp
- fmha_fusion.hpp
- sm100_fmha_fwd_epilogue_tma_warpspecialized.hpp
- sm100_fmha_fwd_mainloop_tma_warpspecialized.hpp
- sm100_fmha_gen_epilogue_warpspecialized.hpp
- sm100_fmha_gen_mainloop_warpspecialized.hpp
- sm100_fmha_load_cpasync_warpspecialized.hpp
- sm100_fmha_load_tma_warpspecialized.hpp
- pow_2.hpp
- fmha.hpp
- sm100_mla.hpp
- fmha_options.hpp
- fmha_tile_scheduler.hpp
- sm100_fmha_fwd_kernel_tma_warpspecialized.hpp
- sm100_fmha_gen_kernel_warpspecialized.hpp
- sm100_fmha_mla_reduction.hpp
- sm100_fmha_mla_tma_warpspecialized.hpp
- sm100_mla_tile_scheduler.hpp
- fmha_cutlass_sm100.cuh
- diff.patch
- PROVENANCE.yaml
- batch_decode.cu
- batch_decode_jit_pybind.cu
- batch_decode_mla_cute_sm80.cu
- batch_decode_mla_plan.cu
- batch_decode_mla_pybind.cu
- batch_decode_mla_run.cu
- batch_prefill.cu
- batch_prefill_fp8_sm90.cu
- batch_prefill_jit_pybind.cu
- batch_prefill_sm90.cu
- batch_prefill_sm90_jit_pybind.cu
- flashinfer_ops.cu
- flashinfer_ops_sm90.cu
- pod.cu
- pod_jit_pybind.cu
- prefill_sm90.cuh
- prefill_sm90.cuh
- cascade.cuh
- decode.cuh
- pod.cuh
- prefill.cuh
- batch_decode.cu
- batch_decode_jit_tvm_binding.cu
- batch_prefill.cu
- batch_prefill_jit_tvm_binding.cu
- batch_prefill_sm90.cu
- batch_prefill_sm90_jit_tvm_binding.cu
- diff.patch
- PROVENANCE.yaml
- gemm_allreduce_two_shot.py
- diff.patch
- PROVENANCE.yaml
- fmha_cutlass_sm100.cu
- sm100_fmha_fwd_mainloop_tma_warpspecialized.hpp
- sm100_fmha_gen_mainloop_warpspecialized.hpp
- diff.patch
- PROVENANCE.yaml
- selective_state_update.cuh
- diff.patch
- PROVENANCE.yaml
- sm_constraint_gemm.py
- __init__.py
- sm_constraint_gemm.py
- diff.patch
- PROVENANCE.yaml
- norm.py
- norm.py
- diff.patch
- PROVENANCE.yaml
- triton_heuristics.py
- diff.patch
- PROVENANCE.yaml
- triton_compat.py
- diff.patch
- PROVENANCE.yaml
- triton_combo_kernel.py
- diff.patch
- PROVENANCE.yaml
- triton_kernel_wrap.py
- diff.patch
- PROVENANCE.yaml
- triton.py
- triton_heuristics.py
- diff.patch
- PROVENANCE.yaml
- norm.py
- rmsnorm_onepass.py
- rotary.py
- scale_shift.py
- awq_marlin_repack.py
- debug_utils.py
- flash_attention_v4.py
- fused_store_index_cache.py
- gptq_marlin.py
- gptq_marlin_repack.py
- hicache.py
- moe_wna16_marlin.py
- ngram_embedding.py
- norm.py
- nvfp4.py
- per_tensor_quant_fp8.py
- per_token_group_quant_8bit.py
- rope.py
- timestep_embedding.py
- kernel_api_logging.py
- __init__.py
- debug_utils.py
- flash_attn.py
- diff.patch
- PROVENANCE.yaml
- gdn_fused_proj.py
- diff.patch
- PROVENANCE.yaml
- cutedsl_kda.py
- kda_cutedsl.py
- diff.patch
- PROVENANCE.yaml
- norm.py
- diff.patch
- PROVENANCE.yaml
- rmsnorm_onepass.py
- all_reduce.py
- awq_marlin_repack.py
- flash_attention_v4.py
- fused_store_index_cache.py
- gptq_marlin.py
- gptq_marlin_repack.py
- hicache.py
- kvcache.py
- moe_wna16_marlin.py
- ngram_embedding.py
- norm.py
- nvfp4.py
- per_tensor_quant_fp8.py
- per_token_group_quant_8bit.py
- rope.py
- timestep_embedding.py
- kernel_api_logging.py
- diff.patch
- PROVENANCE.yaml
- flashinfer_cutedsl.py
- diff.patch
- PROVENANCE.yaml
- killall.py
- common.py
- server_args.py
- check_env.py
- diff.patch
- PROVENANCE.yaml
- extend_attention.py
- diff.patch
- PROVENANCE.yaml
- triton_backend.py
- diff.patch
- PROVENANCE.yaml
- cutlass_mla_entry.cu
- cutlass_mla_kernels.cu
- nvfp4_scaled_mm_kernels.cu
- ops.h
- torch_bindings.cpp
- _custom_ops.py
- diff.patch
- PROVENANCE.yaml
- cuda.py
- diff.patch
- PROVENANCE.yaml
- kernel_warmup.py
- diff.patch
- PROVENANCE.yaml
- triton_mla.py
- triton_decode_attention.py
- diff.patch
- PROVENANCE.yaml
- flashinfer_cutedsl_moe.py
- diff.patch
- PROVENANCE.yaml
- cutlass_mla.py
- flashinfer_mla.py
- flashinfer_mla_sparse.py
- triton_mla.py
- flashinfer.py
- diff.patch
- PROVENANCE.yaml
- gpt_oss_triton_kernels_moe.py
- diff.patch
- PROVENANCE.yaml
- flashinfer_cutedsl_moe.py
- modular_kernel.py
- diff.patch
- PROVENANCE.yaml
- mamba_attn.py
- utils.py
- backend.py
- diff.patch
- PROVENANCE.yaml
- diff.patch
- PROVENANCE.yaml
- requirements.txt
- SKILL.md
- CONTRIBUTING.md
- README.md
# Use via CDN
jsDelivrjsDelivr serves any public GitHub repository as a CDN with zero setup. Pick a version and a file to get a ready-to-paste link and snippet.
Showing a partial file list.
Link
Example
// repository documentation
Was this content helpful?
(0 ratings)
