latest (dev)
Copy
Latest development documentation · Updated 2026-10-08
tensorplay.distributed.algorithms.ddp_comm_hooks.powerSGD_hook.batched_powerSGD_hook
- tensorplay.distributed.algorithms.ddp_comm_hooks.powerSGD_hook.batched_powerSGD_hook(state: PowerSGDState, bucket: GradBucket)[source]
Implement simplified PowerSGD algorithm.
This DDP communication hook implements a simplified PowerSGD gradient compression algorithm described in the paper. This variant does not compress the gradients layer by layer, but instead compresses the flattened input tensor that batches all the gradients. Therefore, it is faster than
powerSGD_hook(), but usually results in a much lower accuracy, unlessmatrix_approximation_rankis 1.Warning
Increasing
matrix_approximation_rankhere may not necessarily increase the accuracy, because batching per-parameter tensors without column/row alignment can destroy low-rank structure.- Parameters:
state (PowerSGDState) – State information to configure the compression rate and support error feedback, warm start, etc.
bucket (dist.GradBucket) – Bucket that stores a 1D flattened gradient tensor that batches multiple per-variable tensors.
- Returns:
Future handler of the communication, which updates the gradients in place.
- Example::
>>> # xdoctest: +SKIP >>> state = PowerSGDState(process_group=process_group, matrix_approximation_rank=1) >>> ddp_model.register_comm_hook(state, batched_powerSGD_hook)
Help improve this page
Found an error, an unclear step, or a missing example?

