TensorPlay
API symbolsoptim
Copy
View MarkdownDownload .md

CosineAnnealingLR

class tensorplay.optim.lr_scheduler.CosineAnnealingLR(optimizer: Optimizer, T_max: int, eta_min: float = 0.0, last_epoch: int = -1)[source]

Set the learning rate of each parameter group using a cosine annealing schedule.

The learning rate is updated recursively using:

\[\eta_{t+1} = \eta_{\min} + (\eta_t - \eta_{\min}) \cdot \frac{1 + \cos\left(\frac{(T_{cur}+1) \pi}{T_{max}}\right)} {1 + \cos\left(\frac{T_{cur} \pi}{T_{max}}\right)}\]

This implements a recursive approximation of the closed-form schedule proposed in SGDR: Stochastic Gradient Descent with Warm Restarts:

\[\eta_t = \eta_{\min} + \frac{1}{2}(\eta_{\max} - \eta_{\min}) \left( 1 + \cos\left(\frac{T_{cur} \pi}{T_{max}}\right) \right)\]

where:

  • \(\eta_t\) is the learning rate at step \(t\)

  • \(T_{cur}\) is the number of epochs since the last restart

  • \(T_{max}\) is the maximum number of epochs in a cycle

Note

Although SGDR includes periodic restarts, this implementation performs cosine annealing without restarts, so \(T_{cur} = t\) and increases monotonically with each call to step().

Parameters:
  • optimizer (Optimizer) – Wrapped optimizer.

  • T_max (int) – Maximum number of iterations.

  • eta_min (float) – Minimum learning rate. Default: 0.

  • last_epoch (int) – The index of the last epoch. Default: -1.

Example

>>> # xdoctest: +SKIP
>>> num_epochs = 100
>>> scheduler = CosineAnnealingLR(optimizer, T_max=num_epochs)
>>> for epoch in range(num_epochs):
>>>     train(...)
>>>     validate(...)
>>>     scheduler.step()
get_last_lr() list[float | TensorBase]

Get the most recent learning rates computed by this scheduler.

Returns:

A list of learning rates with entries for each of the optimizer’s param_groups, with the same types as their group["lr"]s.

Return type:

list[float | Tensor]

Note

The returned Tensors are copies, and never alias the optimizer’s group["lr"]s.

get_lr() list[float | TensorBase][source]

Compute the next learning rate for each of the optimizer’s param_groups.

Scales the group["lr"]s in the optimizer’s param_groups such that their learning rates approximate

\[\texttt{eta\_min} + \frac{1}{2} (\texttt{base\_lr} - \texttt{eta\_min}) \left(1 + \cos\left(\pi \cdot \frac{\texttt{last\_epoch}}{\texttt{T\_max}}\right) \right)\]
Returns:

A list of learning rates for each of the optimizer’s param_groups with the same types as their current group["lr"]s.

Return type:

list[float | Tensor]

Note

If you’re trying to inspect the most recent learning rate, use get_last_lr() instead.

Note

The returned Tensors are copies, and never alias the optimizer’s group["lr"]s.

load_state_dict(state_dict: dict[str, Any]) None

Load the scheduler’s state.

Parameters:

state_dict (dict) – scheduler state. Should be an object returned from a call to state_dict().

state_dict() dict[str, Any]

Return the state of the scheduler as a dict.

It contains an entry for every variable in self.__dict__ which is not the optimizer.

step(epoch: int | None = None) None

Step the scheduler.

Parameters:

epoch (int, optional) –

Deprecated since version 1.4: If provided, sets last_epoch to epoch and uses _get_closed_form_lr() if it is available. This is not universally supported. Use step() without arguments instead.

Note

Call this method after calling the optimizer’s step().

Ask DeepWiki