TensorPlay
latest (dev)
Copy
View Markdown

Latest development documentation · Updated 2026-10-08

tensorplay.nn.attention.varlen.varlen_attn_out

tensorplay.nn.attention.varlen.varlen_attn_out(out: Tensor, query: Tensor, key: Tensor, value: Tensor, cu_seq_q: Tensor, cu_seq_k: Tensor | None, max_q: int, max_k: int, *, return_aux: AuxRequest | None = None, scale: float | None = None, window_size: tuple[int, int] = (-1, -1), enable_gqa: bool = False, seqused_k: Tensor | None = None, block_table: Tensor | None = None, num_splits: int | None = None) → Tensor | tuple[Tensor, Tensor][source]

Compute variable-length attention using Flash Attention with a pre-allocated output tensor.

Same as varlen_attn() but writes the attention output into the provided out tensor instead of allocating a new one.

On this page

Ask DeepWiki