TensorPlay
API reference
latest (dev)
Copy
View Markdown

Latest development documentation · Updated 2026-10-08

tensorplay.distributed.algorithms.ddp_comm_hooks.powerSGD_hook API

Functions 2

#

batched_powerSGD_hook

functionFull reference ↗
tensorplay.distributed.algorithms.ddp_comm_hooks.powerSGD_hook.batched_powerSGD_hook(state: PowerSGDState, bucket: GradBucket)[source]

Implement simplified PowerSGD algorithm.

This DDP communication hook implements a simplified PowerSGD gradient compression algorithm described in the paper. This variant does not compress the gradients layer by layer, but instead compresses the flattened input tensor that batches all the gradients. Therefore, it is faster than powerSGD_hook(), but usually results in a much lower accuracy, unless matrix_approximation_rank is 1.

Warning

Increasing matrix_approximation_rank here may not necessarily increase the accuracy, because batching per-parameter tensors without column/row alignment can destroy low-rank structure.

Parameters:
  • state (PowerSGDState) – State information to configure the compression rate and support error feedback, warm start, etc.

  • bucket (dist.GradBucket) – Bucket that stores a 1D flattened gradient tensor that batches multiple per-variable tensors.

Returns:

Future handler of the communication, which updates the gradients in place.

Example::
>>> # xdoctest: +SKIP
>>> state = PowerSGDState(process_group=process_group, matrix_approximation_rank=1)
>>> ddp_model.register_comm_hook(state, batched_powerSGD_hook)
#

powerSGD_hook

functionFull reference ↗
tensorplay.distributed.algorithms.ddp_comm_hooks.powerSGD_hook.powerSGD_hook(state: PowerSGDState, bucket: GradBucket)[source]

Implement PowerSGD algorithm.

This DDP communication hook implements PowerSGD gradient compression algorithm described in the paper.

Note that this communication hook enforces vanilla allreduce for the first state.start_powerSGD_iter iterations.

Parameters:
  • state (PowerSGDState) – State information to configure the compression rate and support error feedback, warm start, etc.

  • bucket (dist.GradBucket) – Bucket that stores a 1D flattened gradient tensor that batches multiple per-variable tensors.

Returns:

Future handler of the communication, which updates the gradients in place.

Example::
>>> # xdoctest: +SKIP
>>> state = PowerSGDState(process_group=process_group, matrix_approximation_rank=1,
                          start_powerSGD_iter=10, min_compression_rate=0.5)
>>> ddp_model.register_comm_hook(state, powerSGD_hook)

Classes 1

#

PowerSGDState

classFull reference ↗
class tensorplay.distributed.algorithms.ddp_comm_hooks.powerSGD_hook.PowerSGDState(process_group, matrix_approximation_rank=1, start_powerSGD_iter=1000, min_compression_rate=2, use_error_feedback=True, warm_start=True, orthogonalization_epsilon=0, random_seed=0, compression_stats_logging_frequency=10000, batch_tensors_with_same_shape: bool = False)[source]

Store both the algorithm’s hyperparameters and internal state for all gradients during training.

Particularly, matrix_approximation_rank and start_powerSGD_iter are the main hyperparameters that should be tuned by the user. For performance, we suggest to keep binary hyperparameters use_error_feedback and warm_start on.

  1. matrix_approximation_rank controls the size of compressed low-rank tensors, which determines the compression rate. The lower the rank, the stronger the compression.

To tune matrix_approximation_rank, we suggest to start from 1 and increase by factors of 2 (like an exponential grid search, 1, 2, 4, …), until a satisfactory accuracy is reached.

  1. start_powerSGD_iter defers PowerSGD compression until step start_powerSGD_iter, and vanilla allreduce runs prior to step start_powerSGD_iter.

  2. min_compression_rate is the minimum compression rate required when a layer is compressed.

Compression statistics are logged every compression_stats_logging_frequency iterations once PowerSGD compression starts.

  1. orthogonalization_epsilon can be a very small value (e.g., 1e-8) added to every normalized matrix column in orthogonalization step, to prevent div-by-zero error if any column has all 0s.

  2. batch_tensors_with_same_shape controls whether to compress and decompress tensors with same shape in a batched operation to achieve higher parallelism.

Warning

If error feedback or warm-up is enabled, the minimum value of start_powerSGD_iter allowed in DDP is 2.

compression_stats()[source]

Return latest compression statistics as tuple.

Returns tuple of form (compress_rate, numel_before_compression, numel_after_compression).

maybe_increase_iter(bucket)[source]

Track iterations and trigger log message at start of local SGD.

On this page

Ask DeepWiki