latest (dev)
Copy
Latest development documentation · Updated 2026-10-08
tensorplay.nn.attention.sdpa_kernel
- tensorplay.nn.attention.sdpa_kernel(backends: list[_SDPBackend] | _SDPBackend, set_priority: bool = False)[source]
Context manager to select which backend to use for scaled dot product attention.
Warning
This function is beta and subject to change.
- Parameters:
backends (Union[List[SDPBackend], SDPBackend]) – A backend or list of backends for scaled dot product attention.
set_priority (bool=False) – Whether the ordering of the backends is interpreted as their priority order.
Example:
from tensorplay.nn.functional import scaled_dot_product_attention from tensorplay.nn.attention import SDPBackend, sdpa_kernel # Only enable flash attention backend with sdpa_kernel(SDPBackend.FLASH_ATTENTION): scaled_dot_product_attention(...) # Enable the Math or Efficient attention backends with sdpa_kernel([SDPBackend.MATH, SDPBackend.EFFICIENT_ATTENTION]): scaled_dot_product_attention(...) # Enable the cuDNN or flash attention backends, and in that order with sdpa_kernel( [SDPBackend.CUDNN_ATTENTION, SDPBackend.FLASH_ATTENTION], set_priority=True ): scaled_dot_product_attention(...)This context manager can be used to select which backend to use for scaled dot product attention. Upon exiting the context manager, the previous state of the flags will be restored, enabling all backends.
Help improve this page
Found an error, an unclear step, or a missing example?
Was this page helpful?

