latest (dev)
Copy
Latest development documentation · Updated 2026-10-08
tensorplay.distributed.algorithms.ddp_comm_hooks.quantization_hooks.quantization_perchannel_hook
- tensorplay.distributed.algorithms.ddp_comm_hooks.quantization_hooks.quantization_perchannel_hook(process_group, bucket: GradBucket, bucket_size=512)[source]
Apply
quantize_per_channellogic to DDP usingallgatherprotocol.Compared to per-tensor, the main motivation of per-channel is for considerably large tensors such as a tensor that contains 6 million elements quantizing per a bucket size of 512 (or 128) elements may significantly increase the resolution.
It first splits
GradBuckettensor into multiple chunks (channels) ofbucket_sizeelements. Then, workers allgather the scales and zero points of their ownGradBucketprior to the quantization. After all workers have that information, the firstthencallback calledquantize_and_allgatherquantizes worker’s own gradient tensor, and usesallgatherto communicate these across all workers. The finalthencallback calleddequantize_and_aggregate, dequantizes, flattens, and aggregates each quantized gradient tensor locally and returns the mean.Warning
This is experimental, and uses
allgatherprotocol which is considerably slower thanallreduceprotocol. It works only with flattened grads.
Help improve this page
Found an error, an unclear step, or a missing example?

