latest (dev)
Copy
Latest development documentation · Updated 2026-10-08
tensorplay.backends.cuda API
Functions 4
allow_fp16_bf16_reduction_math_sdp
functionFull reference ↗can_use_cudnn_attention
functionFull reference ↗- tensorplay.backends.cuda.can_use_cudnn_attention(params: _SDPAParams, debug: bool = False) bool[source]
Check if cudnn_attention can be utilized in scaled_dot_product_attention.
- Parameters:
params – An instance of SDPAParams containing the tensors for query, key, value, an optional attention mask, dropout rate, and a flag indicating if the attention is causal.
debug – Whether to logging.warn debug information as to why cuDNN attention could not be run. Defaults to False.
- Returns:
True if cuDNN can be used with the given parameters; otherwise, False.
Note
This function is dependent on a CUDA-enabled build of TensorPlay. It will return False in non-CUDA environments.
can_use_efficient_attention
functionFull reference ↗- tensorplay.backends.cuda.can_use_efficient_attention(params: _SDPAParams, debug: bool = False) bool[source]
Check if efficient_attention can be utilized in scaled_dot_product_attention.
- Parameters:
params – An instance of SDPAParams containing the tensors for query, key, value, an optional attention mask, dropout rate, and a flag indicating if the attention is causal.
debug – Whether to logging.warn debug information as to why efficient_attention could not be run. Defaults to False.
- Returns:
True if efficient_attention can be used with the given parameters; otherwise, False.
Note
This function is dependent on a CUDA-enabled build of TensorPlay. It will return False in non-CUDA environments.
can_use_flash_attention
functionFull reference ↗- tensorplay.backends.cuda.can_use_flash_attention(params: _SDPAParams, debug: bool = False) bool[source]
Check if FlashAttention can be utilized in scaled_dot_product_attention.
- Parameters:
params – An instance of SDPAParams containing the tensors for query, key, value, an optional attention mask, dropout rate, and a flag indicating if the attention is causal.
debug – Whether to logging.warn debug information as to why FlashAttention could not be run. Defaults to False.
- Returns:
True if FlashAttention can be used with the given parameters; otherwise, False.
Note
This function is dependent on a CUDA-enabled build of TensorPlay. It will return False in non-CUDA environments.
Classes 2
cuBLASModule
classFull reference ↗- class tensorplay.backends.cuda.cuBLASModule[source]
cuFFTPlanCache
classFull reference ↗- class tensorplay.backends.cuda.cuFFTPlanCache(device_index)[source]
Represent a specific plan cache for a specific device_index.
The attributes size and max_size, and method clear, can fetch and/or change properties of the C++ cuFFT plan cache.
Help improve this page
Found an error, an unclear step, or a missing example?

