latest (dev)
Copy
Latest development documentation · Updated 2026-10-08
tensorplay.distributed.algorithms.ddp_comm_hooks.ddp_zero_hook.hook_with_zero_step_interleaved
- tensorplay.distributed.algorithms.ddp_comm_hooks.ddp_zero_hook.hook_with_zero_step_interleaved(hook: Callable[[Any, GradBucket], Any], ddp: DistributedDataParallel, zero: ZeroRedundancyOptimizer, shard_buckets: bool = False) Callable[[Any, GradBucket], Any][source]
Modify
hookto overlap ZeRO’s optimizer step with the DDP backward pass.Once a bucket’s gradients have been computed, the optimizer computation using those gradients launches, yielding an interleaving of all-reduces and broadcasts in the communication stream. Preferred over
hook_with_zero_step()when communication is relatively fast.
Help improve this page
Found an error, an unclear step, or a missing example?
Was this page helpful?

