latest (dev)
Copy
Latest development documentation · Updated 2026-10-08
tensorplay.distributed.algorithms.ddp_comm_hooks.ddp_zero_hook.hook_with_zero_step
- tensorplay.distributed.algorithms.ddp_comm_hooks.ddp_zero_hook.hook_with_zero_step(hook: Callable[[Any, GradBucket], Any], ddp: DistributedDataParallel, zero: ZeroRedundancyOptimizer, shard_buckets: bool = False) Callable[[Any, GradBucket], Any][source]
Modify
hookto overlap ZeRO’s optimizer step with the DDP backward pass.The optimizer computation follows the backward computation, overlapping with outstanding backward communication. May be preferred over
hook_with_zero_step_interleaved()when communication is relatively slow compared to computation.- Parameters:
hook – the hook to modify.
ddp – the DDP instance to use.
zero – the ZeRO instance to use.
shard_buckets (bool) – if
True, each DDP bucket assignment is partitioned across possibly multiple ranks.
- Raises:
ValueError – if
zerowas constructed withoverlap_with_ddp=False.
Warning
The first two or three training iterations do not perform parameter updates while DDP bucketing information is being collected.
Help improve this page
Found an error, an unclear step, or a missing example?

