# tensorplay.distributed.algorithms.ddp_comm_hooks.ddp_zero_hook.hook_with_zero_step Source: https://www.tensorplay.cn/docs/generated/tensorplay.distributed.algorithms.ddp_comm_hooks.ddp_zero_hook.hook_with_zero_step.html ```python tensorplay.distributed.algorithms.ddp_comm_hooks.ddp_zero_hook.hook_with_zero_step(hook: Callable[[Any, GradBucket], Any], ddp: DistributedDataParallel, zero: ZeroRedundancyOptimizer, shard_buckets: bool = False) → Callable[[Any, GradBucket], Any] ``` Modify hook to overlap ZeRO’s optimizer step with the DDP backward pass. The optimizer computation follows the backward computation, overlapping with outstanding backward communication. May be preferred over [hook_with_zero_step_interleaved()](/docs/generated/tensorplay.distributed.algorithms.ddp_comm_hooks.ddp_zero_hook.hook_with_zero_step_interleaved.html#tensorplay.distributed.algorithms.ddp_comm_hooks.ddp_zero_hook.hook_with_zero_step_interleaved) when communication is relatively slow compared to computation. Parameters: - hook – the hook to modify. - ddp – the DDP instance to use. - zero – the ZeRO instance to use. - shard_buckets ([bool](https://docs.python.org/3/builtins/functions.html#bool)) – if True, each DDP bucket assignment is partitioned across possibly multiple ranks. Raises: [ValueError](https://docs.python.org/3/builtins/exceptions.html#ValueError) – if zero was constructed with overlap_with_ddp=False. > **Warning** > > The first two or three training iterations do not perform parameter updates while DDP bucketing information is being collected.