latest (dev)
Copy
Latest development documentation · Updated 2026-10-08
Computation-graph visualization - tensorplay.utils.viz
Every result produced under autograd carries the recorded chain of
operations behind it. make_dot walks that chain and renders it as a
picture, so you can see exactly which operations produced a value before
calling backward on it.
Render the computation graph recorded for |
Rendering a graph
import tensorplay as tp
from tensorplay.utils.viz import make_dot
x = tp.randn(5, 5, requires_grad=True)
w = tp.randn(5, 5, requires_grad=True)
loss = ((x @ w).relu()).sum()
dot = make_dot(loss, params={"x": x, "w": w})
dot.render("graph", format="png") # writes graph.png in the working directory
Node colors carry the structure:
light blue — leaf tensors with
requires_grad=True; those listed inparamsare labeled by the name you passed,light green — the output tensor handed to
make_dot,white — the recorded operations, labeled by operation name.
Render backends
make_dot picks whichever backend is installed:
the
graphvizpackage returns a realgraphviz.Digraph— call.render(filename, format=...)to write an image, or.pipe()for the raw bytes. Rasterizing to a file also needs thedotprogram on yourPATH.otherwise
networkx+matplotlibdraw a hierarchical layout and return a wrapper with the same.render(filename, format="png")call.with neither installed, calling
make_dotraisesRuntimeError.
Reading the graph as text
The same chain is walkable without any rendering dependency: every tensor
with a gradient history exposes its grad_fn, and every recorded
operation exposes its name and the operations it consumed.
fn = loss.grad_fn
while fn is not None:
print(fn.name)
fn = fn.next_functions[0][0] if fn.next_functions else None
Help improve this page
Found an error, an unclear step, or a missing example?
Complex numbers
TensorPlay supports complex tensors — elements stored as a real and an imaginary part. Complex dtypes are tensorplay.complex64 and tensorplay.complex128 .
DDP Communication Hooks
By default, DistributedDataParallel averages each gradient bucket with an all-reduce as backward fills it: fp32 buckets use a native average collective, other dtypes are pre-divided by the world size and summed. A commun

