TensorPlay
latest (dev)
Copy
View Markdown

Latest development documentation · Updated 2026-10-08

tensorplay.distributed.algorithms.ddp_comm_hooks.ddp_zero_hook.hook_with_zero_step

tensorplay.distributed.algorithms.ddp_comm_hooks.ddp_zero_hook.hook_with_zero_step(hook: Callable[[Any, GradBucket], Any], ddp: DistributedDataParallel, zero: ZeroRedundancyOptimizer, shard_buckets: bool = False) → Callable[[Any, GradBucket], Any][source]

Modify hook to overlap ZeRO’s optimizer step with the DDP backward pass.

The optimizer computation follows the backward computation, overlapping with outstanding backward communication. May be preferred over hook_with_zero_step_interleaved() when communication is relatively slow compared to computation.

Parameters:
  • hook – the hook to modify.

  • ddp – the DDP instance to use.

  • zero – the ZeRO instance to use.

  • shard_buckets (bool) – if True, each DDP bucket assignment is partitioned across possibly multiple ranks.

Raises:

ValueError – if zero was constructed with overlap_with_ddp=False.

Warning

The first two or three training iterations do not perform parameter updates while DDP bucketing information is being collected.

On this page

Ask DeepWiki