TensorPlay
API reference
latest (dev)
Copy
View Markdown

Latest development documentation · Updated 2026-10-08

tensorplay.backends.cuda API

Functions 4

#

allow_fp16_bf16_reduction_math_sdp

functionFull reference ↗
tensorplay.backends.cuda.allow_fp16_bf16_reduction_math_sdp(enabled: bool)[source]

Warning

This flag is beta and subject to change.

Enables or disables fp16/bf16 reduction in math scaled dot product attention.

#

can_use_cudnn_attention

functionFull reference ↗
tensorplay.backends.cuda.can_use_cudnn_attention(params: _SDPAParams, debug: bool = False) → bool[source]

Check if cudnn_attention can be utilized in scaled_dot_product_attention.

Parameters:
  • params – An instance of SDPAParams containing the tensors for query, key, value, an optional attention mask, dropout rate, and a flag indicating if the attention is causal.

  • debug – Whether to logging.warn debug information as to why cuDNN attention could not be run. Defaults to False.

Returns:

True if cuDNN can be used with the given parameters; otherwise, False.

Note

This function is dependent on a CUDA-enabled build of TensorPlay. It will return False in non-CUDA environments.

#

can_use_efficient_attention

functionFull reference ↗
tensorplay.backends.cuda.can_use_efficient_attention(params: _SDPAParams, debug: bool = False) → bool[source]

Check if efficient_attention can be utilized in scaled_dot_product_attention.

Parameters:
  • params – An instance of SDPAParams containing the tensors for query, key, value, an optional attention mask, dropout rate, and a flag indicating if the attention is causal.

  • debug – Whether to logging.warn debug information as to why efficient_attention could not be run. Defaults to False.

Returns:

True if efficient_attention can be used with the given parameters; otherwise, False.

Note

This function is dependent on a CUDA-enabled build of TensorPlay. It will return False in non-CUDA environments.

#

can_use_flash_attention

functionFull reference ↗
tensorplay.backends.cuda.can_use_flash_attention(params: _SDPAParams, debug: bool = False) → bool[source]

Check if FlashAttention can be utilized in scaled_dot_product_attention.

Parameters:
  • params – An instance of SDPAParams containing the tensors for query, key, value, an optional attention mask, dropout rate, and a flag indicating if the attention is causal.

  • debug – Whether to logging.warn debug information as to why FlashAttention could not be run. Defaults to False.

Returns:

True if FlashAttention can be used with the given parameters; otherwise, False.

Note

This function is dependent on a CUDA-enabled build of TensorPlay. It will return False in non-CUDA environments.

Classes 2

#

cuFFTPlanCache

classFull reference ↗
class tensorplay.backends.cuda.cuFFTPlanCache(device_index)[source]

Represent a specific plan cache for a specific device_index.

The attributes size and max_size, and method clear, can fetch and/or change properties of the C++ cuFFT plan cache.

On this page

Ask DeepWiki