# tensorplay.backends Source: https://www.tensorplay.cn/docs/backends.html tensorplay.backends exposes per-library controls: which math libraries this build links against, whether they are available, and the library-specific knobs that change how kernels run. The submodules are cpu, cuda, cudnn, mkl, mkldnn, nnpack, and openmp. ## tensorplay.backends.cpu | cpu.get_cpu_capability |Return the CPU instruction set selected for this build. | | --- | --- | [tensorplay.backends.cpu.get_cpu_capability()](/docs/generated/tensorplay.backends.cpu.get_cpu_capability.html#tensorplay.backends.cpu.get_cpu_capability) reports the highest SIMD capability the CPU dispatch layer selects for ("AVX2", "AVX512", …). Kernels compiled for several instruction sets dispatch on this value, so it is the first thing to check when CPU throughput looks wrong. ## tensorplay.backends.cuda Controls for the CUDA libraries the CUDA device layer links. cuBLASModule is the cuBLAS handle module; cuFFTPlanCache is the per-device plan cache for FFTs, with clear(), size(), and max_size for managing it — plans are expensive to build, and the cache is what makes repeated FFTs cheap. | cuda.allow_fp16_bf16_reduction_math_sdp | | | --- | --- | | cuda.can_use_flash_attention | Check if FlashAttention can be utilized in scaled_dot_product_attention. | | cuda.can_use_efficient_attention | Check if efficient_attention can be utilized in scaled_dot_product_attention. | | cuda.can_use_cudnn_attention | Check if cudnn_attention can be utilized in scaled_dot_product_attention. | | cuda.cuBLASModule | | | cuda.cuFFTPlanCache | Represent a specific plan cache for a specific device_index. | - [tensorplay.backends.cuda.allow_fp16_bf16_reduction_math_sdp()](/docs/generated/tensorplay.backends.cuda.allow_fp16_bf16_reduction_math_sdp.html#tensorplay.backends.cuda.allow_fp16_bf16_reduction_math_sdp) toggles whether the math implementation of scaled-dot-product attention may accumulate its reduction in fp16/bf16 (enabled or disabled as a plain call; pass False when numerical checks require full fp32 reductions). - The can_use_* predicates take an SDPAParams record and report whether the corresponding attention kernel would accept it — the same probes the dispatcher consults (see [attention](/docs/nn.attention.html)). ## tensorplay.backends.cudnn and tensorplay.backends.mkldnn Both are thin re-export modules (m) for the cuDNN and oneDNN-style dense-CPU library bindings compiled into this build. They exist so backend-specific code can be written against a stable import path. ## tensorplay.backends.mkl | mkl.is_available |Return whether MKL kernels were included in this build. | | --- | --- | [tensorplay.backends.mkl.is_available()](/docs/generated/tensorplay.backends.mkl.is_available.html#tensorplay.backends.mkl.is_available) reports whether the CPU math library is linked and usable. ## tensorplay.backends.nnpack | nnpack.is_available |Return whether NNPACK kernels were included in this build. | | --- | --- | | nnpack.flags | Temporarily set the process-wide NNPACK enable flag. | | nnpack.set_flags | Set the process-wide NNPACK enable flag. | nnpack is the packing-based CPU convolution path. flags(enabled=...) is the context-manager form — the setting applies inside the with block and reverts on exit — and set_flags changes it without a context. ## tensorplay.backends.openmp | openmp.is_available |Return whether OpenMP support was included in this build. | | --- | --- | [tensorplay.backends.openmp.is_available()](/docs/generated/tensorplay.backends.openmp.is_available.html#tensorplay.backends.openmp.is_available) reports whether the OpenMP thread pool backs CPU parallelism in this build; when it is absent, intra-op parallelism uses the built-in pool instead.