Copy
CosineAnnealingLR
- class tensorplay.optim.lr_scheduler.CosineAnnealingLR(optimizer: Optimizer, T_max: int, eta_min: float = 0.0, last_epoch: int = -1)[source]
Set the learning rate of each parameter group using a cosine annealing schedule.
The learning rate is updated recursively using:
\[\eta_{t+1} = \eta_{\min} + (\eta_t - \eta_{\min}) \cdot \frac{1 + \cos\left(\frac{(T_{cur}+1) \pi}{T_{max}}\right)} {1 + \cos\left(\frac{T_{cur} \pi}{T_{max}}\right)}\]This implements a recursive approximation of the closed-form schedule proposed in SGDR: Stochastic Gradient Descent with Warm Restarts:
\[\eta_t = \eta_{\min} + \frac{1}{2}(\eta_{\max} - \eta_{\min}) \left( 1 + \cos\left(\frac{T_{cur} \pi}{T_{max}}\right) \right)\]where:
\(\eta_t\) is the learning rate at step \(t\)
\(T_{cur}\) is the number of epochs since the last restart
\(T_{max}\) is the maximum number of epochs in a cycle
Note
Although SGDR includes periodic restarts, this implementation performs cosine annealing without restarts, so \(T_{cur} = t\) and increases monotonically with each call to
step().- Parameters:
Example
>>> # xdoctest: +SKIP >>> num_epochs = 100 >>> scheduler = CosineAnnealingLR(optimizer, T_max=num_epochs) >>> for epoch in range(num_epochs): >>> train(...) >>> validate(...) >>> scheduler.step()
- get_last_lr() list[float | TensorBase]
Get the most recent learning rates computed by this scheduler.
- Returns:
A
listof learning rates with entries for each of the optimizer’sparam_groups, with the same types as theirgroup["lr"]s.- Return type:
Note
The returned
Tensors are copies, and never alias the optimizer’sgroup["lr"]s.
- get_lr() list[float | TensorBase][source]
Compute the next learning rate for each of the optimizer’s
param_groups.Scales the
group["lr"]s in the optimizer’sparam_groupssuch that their learning rates approximate\[\texttt{eta\_min} + \frac{1}{2} (\texttt{base\_lr} - \texttt{eta\_min}) \left(1 + \cos\left(\pi \cdot \frac{\texttt{last\_epoch}}{\texttt{T\_max}}\right) \right)\]- Returns:
A
listof learning rates for each of the optimizer’sparam_groupswith the same types as their currentgroup["lr"]s.- Return type:
Note
If you’re trying to inspect the most recent learning rate, use
get_last_lr()instead.Note
The returned
Tensors are copies, and never alias the optimizer’sgroup["lr"]s.
- load_state_dict(state_dict: dict[str, Any]) None
Load the scheduler’s state.
- Parameters:
state_dict (dict) – scheduler state. Should be an object returned from a call to
state_dict().
- state_dict() dict[str, Any]
Return the state of the scheduler as a
dict.It contains an entry for every variable in
self.__dict__which is not the optimizer.
- step(epoch: int | None = None) None
Step the scheduler.
- Parameters:
epoch (int, optional) –
Deprecated since version 1.4: If provided, sets
last_epochtoepochand uses_get_closed_form_lr()if it is available. This is not universally supported. Usestep()without arguments instead.
Note
Call this method after calling the optimizer’s
step().
Help improve this page
Found an error, an unclear step, or a missing example?
