# tensorplay.backends.cuda API Source: https://www.tensorplay.cn/docs/api/tensorplay.backends.cuda.html ## Functions 4 [#](#api-tensorplay.backends.cuda.allow_fp16_bf16_reduction_math_sdp) ### allow_fp16_bf16_reduction_math_sdp function[Full reference ↗](/docs/generated/tensorplay.backends.cuda.allow_fp16_bf16_reduction_math_sdp.html) ```python tensorplay.backends.cuda.allow_fp16_bf16_reduction_math_sdp(enabled: bool) ``` > **Warning** > > This flag is beta and subject to change. Enables or disables fp16/bf16 reduction in math scaled dot product attention. [#](#api-tensorplay.backends.cuda.can_use_cudnn_attention) ### can_use_cudnn_attention function[Full reference ↗](/docs/generated/tensorplay.backends.cuda.can_use_cudnn_attention.html) ```python tensorplay.backends.cuda.can_use_cudnn_attention(params: _SDPAParams, debug: bool = False) → bool ``` Check if cudnn_attention can be utilized in scaled_dot_product_attention. Parameters: - params – An instance of SDPAParams containing the tensors for query, key, value, an optional attention mask, dropout rate, and a flag indicating if the attention is causal. - debug – Whether to logging.warn debug information as to why cuDNN attention could not be run. Defaults to False. Returns: True if cuDNN can be used with the given parameters; otherwise, False. > **Note** > > This function is dependent on a CUDA-enabled build of TensorPlay. It will return False in non-CUDA environments. [#](#api-tensorplay.backends.cuda.can_use_efficient_attention) ### can_use_efficient_attention function[Full reference ↗](/docs/generated/tensorplay.backends.cuda.can_use_efficient_attention.html) ```python tensorplay.backends.cuda.can_use_efficient_attention(params: _SDPAParams, debug: bool = False) → bool ``` Check if efficient_attention can be utilized in scaled_dot_product_attention. Parameters: - params – An instance of SDPAParams containing the tensors for query, key, value, an optional attention mask, dropout rate, and a flag indicating if the attention is causal. - debug – Whether to logging.warn debug information as to why efficient_attention could not be run. Defaults to False. Returns: True if efficient_attention can be used with the given parameters; otherwise, False. > **Note** > > This function is dependent on a CUDA-enabled build of TensorPlay. It will return False in non-CUDA environments. [#](#api-tensorplay.backends.cuda.can_use_flash_attention) ### can_use_flash_attention function[Full reference ↗](/docs/generated/tensorplay.backends.cuda.can_use_flash_attention.html) ```python tensorplay.backends.cuda.can_use_flash_attention(params: _SDPAParams, debug: bool = False) → bool ``` Check if FlashAttention can be utilized in scaled_dot_product_attention. Parameters: - params – An instance of SDPAParams containing the tensors for query, key, value, an optional attention mask, dropout rate, and a flag indicating if the attention is causal. - debug – Whether to logging.warn debug information as to why FlashAttention could not be run. Defaults to False. Returns: True if FlashAttention can be used with the given parameters; otherwise, False. > **Note** > > This function is dependent on a CUDA-enabled build of TensorPlay. It will return False in non-CUDA environments. ## Classes 2 [#](#api-tensorplay.backends.cuda.cuBLASModule) ### cuBLASModule class[Full reference ↗](/docs/generated/tensorplay.backends.cuda.cuBLASModule.html) ```python class tensorplay.backends.cuda.cuBLASModule ``` [#](#api-tensorplay.backends.cuda.cuFFTPlanCache) ### cuFFTPlanCache class[Full reference ↗](/docs/generated/tensorplay.backends.cuda.cuFFTPlanCache.html) ```python class tensorplay.backends.cuda.cuFFTPlanCache(device_index) ``` Represent a specific plan cache for a specific device_index. The attributes size and max_size, and method clear, can fetch and/or change properties of the C++ cuFFT plan cache.