# tensorplay.distributed.algorithms.ddp_comm_hooks.quantization_hooks API Source: https://www.tensorplay.cn/docs/api/tensorplay.distributed.algorithms.ddp_comm_hooks.quantization_hooks.html ## Functions 2 [#](#api-tensorplay.distributed.algorithms.ddp_comm_hooks.quantization_hooks.quantization_perchannel_hook) ### quantization_perchannel_hook function[Full reference ↗](/docs/generated/tensorplay.distributed.algorithms.ddp_comm_hooks.quantization_hooks.quantization_perchannel_hook.html) ```python tensorplay.distributed.algorithms.ddp_comm_hooks.quantization_hooks.quantization_perchannel_hook(process_group, bucket: GradBucket, bucket_size=512) ``` Apply quantize_per_channel logic to DDP using allgather protocol. Compared to per-tensor, the main motivation of per-channel is for considerably large tensors such as a tensor that contains 6 million elements quantizing per a bucket size of 512 (or 128) elements may significantly increase the resolution. It first splits GradBucket tensor into multiple chunks (channels) of bucket_size elements. Then, workers allgather the scales and zero points of their own GradBucket prior to the quantization. After all workers have that information, the first then callback called quantize_and_allgather quantizes worker’s own gradient tensor, and uses allgather to communicate these across all workers. The final then callback called dequantize_and_aggregate, dequantizes, flattens, and aggregates each quantized gradient tensor locally and returns the mean. > **Warning** > > This is experimental, and uses allgather protocol which is considerably slower than allreduce protocol. It works only with flattened grads. [#](#api-tensorplay.distributed.algorithms.ddp_comm_hooks.quantization_hooks.quantization_pertensor_hook) ### quantization_pertensor_hook function[Full reference ↗](/docs/generated/tensorplay.distributed.algorithms.ddp_comm_hooks.quantization_hooks.quantization_pertensor_hook.html) ```python tensorplay.distributed.algorithms.ddp_comm_hooks.quantization_hooks.quantization_pertensor_hook(process_group, bucket: GradBucket) ``` Apply quantize_per_tensor logic to DDP using allgather protocol. Workers first allgather the scale and zero point of their own GradBucket prior to the quantization. After all workers have that information, the first then callback called quantize_and_allgather quantizes worker’s own gradient tensor, and uses allgather to communicate these across all workers. The final then callback called dequantize_and_aggregate, dequantizes and aggregates each quantized gradient tensor locally and returns the mean. > **Warning** > > This is experimental, and uses allgather protocol which is considerably slower than allreduce protocol. It works only with flattened grads. Example:: ``` >>> # xdoctest: +SKIP >>> ddp_model.register_comm_hook(process_group, quantization_pertensor_hook) ```