# tensorplay.distributed.algorithms.ddp_comm_hooks.ddp_zero_hook API Source: https://www.tensorplay.cn/docs/api/tensorplay.distributed.algorithms.ddp_comm_hooks.ddp_zero_hook.html ## Functions 2 [#](#api-tensorplay.distributed.algorithms.ddp_comm_hooks.ddp_zero_hook.hook_with_zero_step_interleaved) ### hook_with_zero_step_interleaved function[Full reference ↗](/docs/generated/tensorplay.distributed.algorithms.ddp_comm_hooks.ddp_zero_hook.hook_with_zero_step_interleaved.html) ```python tensorplay.distributed.algorithms.ddp_comm_hooks.ddp_zero_hook.hook_with_zero_step_interleaved(hook: Callable[[Any, GradBucket], Any], ddp: DistributedDataParallel, zero: ZeroRedundancyOptimizer, shard_buckets: bool = False) → Callable[[Any, GradBucket], Any] ``` Modify hook to overlap ZeRO’s optimizer step with the DDP backward pass. Once a bucket’s gradients have been computed, the optimizer computation using those gradients launches, yielding an interleaving of all-reduces and broadcasts in the communication stream. Preferred over [hook_with_zero_step()](/docs/generated/tensorplay.distributed.algorithms.ddp_comm_hooks.ddp_zero_hook.hook_with_zero_step.html#tensorplay.distributed.algorithms.ddp_comm_hooks.ddp_zero_hook.hook_with_zero_step) when communication is relatively fast. [#](#api-tensorplay.distributed.algorithms.ddp_comm_hooks.ddp_zero_hook.hook_with_zero_step) ### hook_with_zero_step function[Full reference ↗](/docs/generated/tensorplay.distributed.algorithms.ddp_comm_hooks.ddp_zero_hook.hook_with_zero_step.html) ```python tensorplay.distributed.algorithms.ddp_comm_hooks.ddp_zero_hook.hook_with_zero_step(hook: Callable[[Any, GradBucket], Any], ddp: DistributedDataParallel, zero: ZeroRedundancyOptimizer, shard_buckets: bool = False) → Callable[[Any, GradBucket], Any] ``` Modify hook to overlap ZeRO’s optimizer step with the DDP backward pass. The optimizer computation follows the backward computation, overlapping with outstanding backward communication. May be preferred over [hook_with_zero_step_interleaved()](/docs/generated/tensorplay.distributed.algorithms.ddp_comm_hooks.ddp_zero_hook.hook_with_zero_step_interleaved.html#tensorplay.distributed.algorithms.ddp_comm_hooks.ddp_zero_hook.hook_with_zero_step_interleaved) when communication is relatively slow compared to computation. Parameters: - hook – the hook to modify. - ddp – the DDP instance to use. - zero – the ZeRO instance to use. - shard_buckets ([bool](https://docs.python.org/3/builtins/functions.html#bool)) – if True, each DDP bucket assignment is partitioned across possibly multiple ranks. Raises: [ValueError](https://docs.python.org/3/builtins/exceptions.html#ValueError) – if zero was constructed with overlap_with_ddp=False. > **Warning** > > The first two or three training iterations do not perform parameter updates while DDP bucketing information is being collected.