TensorPlay
Reference guides
latest (dev)
Copy
View Markdown

Latest development documentation · Updated 2026-10-08

tensorplay.backends

tensorplay.backends exposes per-library controls: which math libraries this build links against, whether they are available, and the library-specific knobs that change how kernels run. The submodules are cpu, cuda, cudnn, mkl, mkldnn, nnpack, and openmp.

tensorplay.backends.cpu

cpu.get_cpu_capability

Return the CPU instruction set selected for this build.

tensorplay.backends.cpu.get_cpu_capability() reports the highest SIMD capability the CPU dispatch layer selects for ("AVX2", "AVX512", …). Kernels compiled for several instruction sets dispatch on this value, so it is the first thing to check when CPU throughput looks wrong.

tensorplay.backends.cuda

Controls for the CUDA libraries the CUDA device layer links. cuBLASModule is the cuBLAS handle module; cuFFTPlanCache is the per-device plan cache for FFTs, with clear(), size(), and max_size for managing it — plans are expensive to build, and the cache is what makes repeated FFTs cheap.

cuda.allow_fp16_bf16_reduction_math_sdp

cuda.can_use_flash_attention

Check if FlashAttention can be utilized in scaled_dot_product_attention.

cuda.can_use_efficient_attention

Check if efficient_attention can be utilized in scaled_dot_product_attention.

cuda.can_use_cudnn_attention

Check if cudnn_attention can be utilized in scaled_dot_product_attention.

cuda.cuBLASModule

cuda.cuFFTPlanCache

Represent a specific plan cache for a specific device_index.

  • tensorplay.backends.cuda.allow_fp16_bf16_reduction_math_sdp() toggles whether the math implementation of scaled-dot-product attention may accumulate its reduction in fp16/bf16 (enabled or disabled as a plain call; pass False when numerical checks require full fp32 reductions).

  • The can_use_* predicates take an SDPAParams record and report whether the corresponding attention kernel would accept it — the same probes the dispatcher consults (see attention).

tensorplay.backends.cudnn and tensorplay.backends.mkldnn

Both are thin re-export modules (m) for the cuDNN and oneDNN-style dense-CPU library bindings compiled into this build. They exist so backend-specific code can be written against a stable import path.

tensorplay.backends.mkl

mkl.is_available

Return whether MKL kernels were included in this build.

tensorplay.backends.mkl.is_available() reports whether the CPU math library is linked and usable.

tensorplay.backends.nnpack

nnpack.is_available

Return whether NNPACK kernels were included in this build.

nnpack.flags

Temporarily set the process-wide NNPACK enable flag.

nnpack.set_flags

Set the process-wide NNPACK enable flag.

nnpack is the packing-based CPU convolution path. flags(enabled=...) is the context-manager form — the setting applies inside the with block and reverts on exit — and set_flags changes it without a context.

tensorplay.backends.openmp

openmp.is_available

Return whether OpenMP support was included in this build.

tensorplay.backends.openmp.is_available() reports whether the OpenMP thread pool backs CPU parallelism in this build; when it is absent, intra-op parallelism uses the built-in pool instead.

On this page

Ask DeepWiki