TensorPlay
latest (dev)
Copy
View Markdown

Latest development documentation · Updated 2026-10-08

tensorplay.distributed.algorithms.ddp_comm_hooks.ddp_zero_hook.hook_with_zero_step_interleaved

tensorplay.distributed.algorithms.ddp_comm_hooks.ddp_zero_hook.hook_with_zero_step_interleaved(hook: Callable[[Any, GradBucket], Any], ddp: DistributedDataParallel, zero: ZeroRedundancyOptimizer, shard_buckets: bool = False) → Callable[[Any, GradBucket], Any][source]

Modify hook to overlap ZeRO’s optimizer step with the DDP backward pass.

Once a bucket’s gradients have been computed, the optimizer computation using those gradients launches, yielding an interleaving of all-reduces and broadcasts in the communication stream. Preferred over hook_with_zero_step() when communication is relatively fast.

On this page

Ask DeepWiki