latest (dev)
Copy
Latest development documentation · Updated 2026-10-08
tensorplay.distributed.algorithms.ddp_comm_hooks.quantization_hooks API
Functions 2
quantization_perchannel_hook
functionFull reference ↗- tensorplay.distributed.algorithms.ddp_comm_hooks.quantization_hooks.quantization_perchannel_hook(process_group, bucket: GradBucket, bucket_size=512)[source]
Apply
quantize_per_channellogic to DDP usingallgatherprotocol.Compared to per-tensor, the main motivation of per-channel is for considerably large tensors such as a tensor that contains 6 million elements quantizing per a bucket size of 512 (or 128) elements may significantly increase the resolution.
It first splits
GradBuckettensor into multiple chunks (channels) ofbucket_sizeelements. Then, workers allgather the scales and zero points of their ownGradBucketprior to the quantization. After all workers have that information, the firstthencallback calledquantize_and_allgatherquantizes worker’s own gradient tensor, and usesallgatherto communicate these across all workers. The finalthencallback calleddequantize_and_aggregate, dequantizes, flattens, and aggregates each quantized gradient tensor locally and returns the mean.Warning
This is experimental, and uses
allgatherprotocol which is considerably slower thanallreduceprotocol. It works only with flattened grads.
quantization_pertensor_hook
functionFull reference ↗- tensorplay.distributed.algorithms.ddp_comm_hooks.quantization_hooks.quantization_pertensor_hook(process_group, bucket: GradBucket)[source]
Apply
quantize_per_tensorlogic to DDP usingallgatherprotocol.Workers first allgather the scale and zero point of their own
GradBucketprior to the quantization. After all workers have that information, the firstthencallback calledquantize_and_allgatherquantizes worker’s own gradient tensor, and usesallgatherto communicate these across all workers. The finalthencallback calleddequantize_and_aggregate, dequantizes and aggregates each quantized gradient tensor locally and returns the mean.Warning
This is experimental, and uses
allgatherprotocol which is considerably slower thanallreduceprotocol. It works only with flattened grads.- Example::
>>> # xdoctest: +SKIP >>> ddp_model.register_comm_hook(process_group, quantization_pertensor_hook)
Help improve this page
Found an error, an unclear step, or a missing example?
tensorplay.distributed.algorithms.ddp_comm_hooks.powerSGD_hook API
Complete API reference for tensorplay.distributed.algorithms.ddp_comm_hooks.powerSGD_hook, including signatures, parameters, examples and members.
tensorplay.distributed.algorithms.model_averaging API
Complete API reference for tensorplay.distributed.algorithms.model_averaging, including signatures, parameters, examples and members.

