# tensorplay.distributed.algorithms.ddp_comm_hooks.powerSGD_hook API Source: https://www.tensorplay.cn/docs/api/tensorplay.distributed.algorithms.ddp_comm_hooks.powerSGD_hook.html ## Functions 2 [#](#api-tensorplay.distributed.algorithms.ddp_comm_hooks.powerSGD_hook.batched_powerSGD_hook) ### batched_powerSGD_hook function[Full reference ↗](/docs/generated/tensorplay.distributed.algorithms.ddp_comm_hooks.powerSGD_hook.batched_powerSGD_hook.html) ```python tensorplay.distributed.algorithms.ddp_comm_hooks.powerSGD_hook.batched_powerSGD_hook(state: PowerSGDState, bucket: GradBucket) ``` Implement simplified PowerSGD algorithm. This DDP communication hook implements a simplified PowerSGD gradient compression algorithm described in the [paper](https://arxiv.org/abs/1905.13727). This variant does not compress the gradients layer by layer, but instead compresses the flattened input tensor that batches all the gradients. Therefore, it is faster than [powerSGD_hook()](/docs/generated/tensorplay.distributed.algorithms.ddp_comm_hooks.powerSGD_hook.powerSGD_hook.html#tensorplay.distributed.algorithms.ddp_comm_hooks.powerSGD_hook.powerSGD_hook), but usually results in a much lower accuracy, unless matrix_approximation_rank is 1. > **Warning** > > Increasing matrix_approximation_rank here may not necessarily increase the accuracy, because batching per-parameter tensors without column/row alignment can destroy low-rank structure. Parameters: - state ([PowerSGDState](/docs/generated/tensorplay.distributed.algorithms.ddp_comm_hooks.powerSGD_hook.PowerSGDState.html#tensorplay.distributed.algorithms.ddp_comm_hooks.powerSGD_hook.PowerSGDState)) – State information to configure the compression rate and support error feedback, warm start, etc. - bucket (dist.GradBucket) – Bucket that stores a 1D flattened gradient tensor that batches multiple per-variable tensors. Returns: Future handler of the communication, which updates the gradients in place. Example:: ``` >>> # xdoctest: +SKIP >>> state = PowerSGDState(process_group=process_group, matrix_approximation_rank=1) >>> ddp_model.register_comm_hook(state, batched_powerSGD_hook) ``` [#](#api-tensorplay.distributed.algorithms.ddp_comm_hooks.powerSGD_hook.powerSGD_hook) ### powerSGD_hook function[Full reference ↗](/docs/generated/tensorplay.distributed.algorithms.ddp_comm_hooks.powerSGD_hook.powerSGD_hook.html) ```python tensorplay.distributed.algorithms.ddp_comm_hooks.powerSGD_hook.powerSGD_hook(state: PowerSGDState, bucket: GradBucket) ``` Implement PowerSGD algorithm. This DDP communication hook implements PowerSGD gradient compression algorithm described in the [paper](https://arxiv.org/abs/1905.13727). Note that this communication hook enforces vanilla allreduce for the first state.start_powerSGD_iter iterations. Parameters: - state ([PowerSGDState](/docs/generated/tensorplay.distributed.algorithms.ddp_comm_hooks.powerSGD_hook.PowerSGDState.html#tensorplay.distributed.algorithms.ddp_comm_hooks.powerSGD_hook.PowerSGDState)) – State information to configure the compression rate and support error feedback, warm start, etc. - bucket (dist.GradBucket) – Bucket that stores a 1D flattened gradient tensor that batches multiple per-variable tensors. Returns: Future handler of the communication, which updates the gradients in place. Example:: ``` >>> # xdoctest: +SKIP >>> state = PowerSGDState(process_group=process_group, matrix_approximation_rank=1, start_powerSGD_iter=10, min_compression_rate=0.5) >>> ddp_model.register_comm_hook(state, powerSGD_hook) ``` ## Classes 1 [#](#api-tensorplay.distributed.algorithms.ddp_comm_hooks.powerSGD_hook.PowerSGDState) ### PowerSGDState class[Full reference ↗](/docs/generated/tensorplay.distributed.algorithms.ddp_comm_hooks.powerSGD_hook.PowerSGDState.html) ```python class tensorplay.distributed.algorithms.ddp_comm_hooks.powerSGD_hook.PowerSGDState(process_group, matrix_approximation_rank=1, start_powerSGD_iter=1000, min_compression_rate=2, use_error_feedback=True, warm_start=True, orthogonalization_epsilon=0, random_seed=0, compression_stats_logging_frequency=10000, batch_tensors_with_same_shape: bool = False) ``` Store both the algorithm’s hyperparameters and internal state for all gradients during training. Particularly, matrix_approximation_rank and start_powerSGD_iter are the main hyperparameters that should be tuned by the user. For performance, we suggest to keep binary hyperparameters use_error_feedback and warm_start on. - matrix_approximation_rank controls the size of compressed low-rank tensors, which determines the compression rate. The lower the rank, the stronger the compression. To tune matrix_approximation_rank, we suggest to start from 1 and increase by factors of 2 (like an exponential grid search, 1, 2, 4, …), until a satisfactory accuracy is reached. - start_powerSGD_iter defers PowerSGD compression until step start_powerSGD_iter, and vanilla allreduce runs prior to step start_powerSGD_iter. - min_compression_rate is the minimum compression rate required when a layer is compressed. Compression statistics are logged every compression_stats_logging_frequency iterations once PowerSGD compression starts. - orthogonalization_epsilon can be a very small value (e.g., 1e-8) added to every normalized matrix column in orthogonalization step, to prevent div-by-zero error if any column has all 0s. - batch_tensors_with_same_shape controls whether to compress and decompress tensors with same shape in a batched operation to achieve higher parallelism. > **Warning** > > If error feedback or warm-up is enabled, the minimum value of start_powerSGD_iter allowed in DDP is 2. ```python compression_stats() ``` Return latest compression statistics as tuple. Returns tuple of form (compress_rate, numel_before_compression, numel_after_compression). ```python maybe_increase_iter(bucket) ``` Track iterations and trigger log message at start of local SGD.