latest (dev)
Copy
Latest development documentation · Updated 2026-10-08
tensorplay.accelerator API
Functions 7
current_accelerator
functionFull reference ↗current_device_index
functionFull reference ↗device_count
functionFull reference ↗get_device_capability
functionFull reference ↗is_available
functionFull reference ↗set_device_index
functionFull reference ↗synchronize
functionFull reference ↗Classes 1
Graph
classFull reference ↗- class tensorplay.accelerator.Graph(keep_graph: bool = False, *, pool: Any = None, capture_error_mode: str = 'global')[source]
Capture/replay graph on the current accelerator device.
- Parameters:
keep_graph – accepted for generic code; the executable is compiled when capture ends in this backend.
pool –
Nonecaptures into a fresh private pool, otherwise a pool id fromgraph_pool_handle(), another graph, or another graph’s pool id shares that pool.capture_error_mode –
"global"fails the capture on unsafe calls anywhere in the process,"thread_local"only watches this thread,"relaxed"skips the guards.
- begin_capture_to_if_node(scalar_pred)
Inside an open capture, gate the following work on an
ifnode.scalar_predmust be a single-element CUDA Bool tensor; at replay time the driver samples it and runs the body captured between this call andend_capture_to_conditional_node()only when true.
- begin_capture_to_while_node(scalar_pred)
Like
begin_capture_to_if_node(), but the body loops while the predicate stays true (driver-level while node).
- capture_begin(pool: Any = '__unset__', capture_error_mode: Any = '__unset__', stream: Any = None) None[source]
Begin capture on the current stream with the stored settings.
- capture_end()
End capture and compile the executable (paid here, not on first replay).
- debug_dump(path)
Write a DOT rendering of the captured graph to
path.Call
enable_debug_mode()before capturing for a dump that includes full node attributes.
- end_capture_to_conditional_node()
Close the open conditional body; subsequent capture returns to the parent stream.
- instantiate()
No-op once instantiated; kept for late callers.
- pool()[source]
Opaque id of this graph’s memory pool, shareable with others.
- property pool_id
Allocator pool id this graph captured against.
- replay(stream=None)
Run the graph: launch the cached executable on the current stream.
- Parameters:
stream (Stream, optional) – launch on this explicit stream instead of querying the current one - shaves a TLS lookup off hot loops pinned to a single stream.
- reset()
Destroy the executable and release the pool reference.
All tensors allocated during the capture must be released first.
- set_conditional_handle_for_current_node(scalar_pred)
Refresh the predicate consumed by the innermost open conditional node (used for nested conditionals).
- stage_and_launch(static_inputs, inputs)
Stage every input onto its static buffer and replay in one call.
- Parameters:
static_inputs – buffers captured by the graph (kept alive by the caller).
inputs – fresh tensors whose contents overwrite the matching static buffer this iteration. Contiguous same-dtype/ same-device pairs take a raw async device-to-device copy; anything else falls back to full copy semantics.
This is the low-overhead bulk entry used by
tensorplay.compiler.backends.cudagraphs: one Python-to-native crossing for the whole replay instead of one dispatcher round trip per input.
Variables 3
graphs
dataFull reference ↗Device-agnostic capture/replay graphs for the current accelerator.
Graph records a sequence of operations on the active accelerator
backend and replays it with reduced launch overhead.
Classes
|
Capture/replay graph on the current accelerator device. |
memory
dataFull reference ↗Device-agnostic memory queries for the current accelerator.
Every function below targets the active accelerator backend (the GPU
backend in this build) and accepts a device object, a spelling such
as "cuda:1", an integer index, or None for the current device.
Functions
|
Release unoccupied cached device memory held by the allocator. |
|
Release unoccupied cached host memory. |
|
Free and total device memory in bytes for the given device. |
|
Peak bytes occupied by live tensors on the given device. |
|
Peak bytes managed by the allocator on the given device. |
|
Bytes currently occupied by live tensors on the given device. |
|
Bytes currently managed by the allocator on the given device. |
|
Allocator statistics for the given device. |
|
Reset historical accumulation counters on the given device. |
|
Reset peak counters on the given device. |
random
dataFull reference ↗Device-agnostic random-number helpers for the current accelerator.
Seeds and generator states are read and written through the active
accelerator backend (the GPU backend in this build). device
accepts a device object, a spelling such as "cuda:1", an integer
index, or None for the current device.
Functions
|
Generator state of the given accelerator device. |
|
Generator states of every accelerator device. |
|
Initial seed of the generator for the given accelerator device. |
|
Seed the generator of the current accelerator device. |
|
Seed the generators of every accelerator device. |
|
Reseed the generator of the current accelerator device. |
|
Reseed the generators of every accelerator device. |
|
Set the generator state of the given accelerator device. |
|
Set the generator states of every accelerator device. |
Help improve this page
Found an error, an unclear step, or a missing example?

