TensorPlay
API reference
latest (dev)
Copy
View Markdown

Latest development documentation · Updated 2026-10-08

tensorplay.accelerator API

Functions 7

#

current_accelerator

functionFull reference ↗
tensorplay.accelerator.current_accelerator(check_available: bool = False)[source]

The accelerator device selected at build time, if any.

Returns None on a host-only build. With check_available set, a runtime availability probe is also required before the device is reported.

#

current_device_index

functionFull reference ↗
tensorplay.accelerator.current_device_index() → int[source]

Index of the currently selected accelerator device.

#

device_count

functionFull reference ↗
tensorplay.accelerator.device_count() → int[source]

Number of devices for the current accelerator, or zero without one.

#

get_device_capability

functionFull reference ↗
tensorplay.accelerator.get_device_capability(device=None) → dict[str, Any][source]

Capability map for an accelerator device.

The map carries a supported_dtypes set listing the data types that can be allocated on the device.

#

is_available

functionFull reference ↗
tensorplay.accelerator.is_available() → bool[source]

Whether an accelerator was built and at least one device is visible.

#

set_device_index

functionFull reference ↗
tensorplay.accelerator.set_device_index(device) → None[source]

Select the accelerator device by index; negative indices are no-ops.

#

synchronize

functionFull reference ↗
tensorplay.accelerator.synchronize(device=None) → None[source]

Wait for all work on an accelerator device to complete.

Classes 1

#

Graph

classFull reference ↗
class tensorplay.accelerator.Graph(keep_graph: bool = False, *, pool: Any = None, capture_error_mode: str = 'global')[source]

Capture/replay graph on the current accelerator device.

Parameters:
  • keep_graph – accepted for generic code; the executable is compiled when capture ends in this backend.

  • pool – None captures into a fresh private pool, otherwise a pool id from graph_pool_handle(), another graph, or another graph’s pool id shares that pool.

  • capture_error_mode – "global" fails the capture on unsafe calls anywhere in the process, "thread_local" only watches this thread, "relaxed" skips the guards.

begin_capture_to_if_node(scalar_pred)

Inside an open capture, gate the following work on an if node.

scalar_pred must be a single-element CUDA Bool tensor; at replay time the driver samples it and runs the body captured between this call and end_capture_to_conditional_node() only when true.

begin_capture_to_while_node(scalar_pred)

Like begin_capture_to_if_node(), but the body loops while the predicate stays true (driver-level while node).

capture_begin(pool: Any = '__unset__', capture_error_mode: Any = '__unset__', stream: Any = None) → None[source]

Begin capture on the current stream with the stored settings.

capture_end()

End capture and compile the executable (paid here, not on first replay).

debug_dump(path)

Write a DOT rendering of the captured graph to path.

Call enable_debug_mode() before capturing for a dump that includes full node attributes.

end_capture_to_conditional_node()

Close the open conditional body; subsequent capture returns to the parent stream.

instantiate()

No-op once instantiated; kept for late callers.

pool()[source]

Opaque id of this graph’s memory pool, shareable with others.

property pool_id

Allocator pool id this graph captured against.

replay(stream=None)

Run the graph: launch the cached executable on the current stream.

Parameters:

stream (Stream, optional) – launch on this explicit stream instead of querying the current one - shaves a TLS lookup off hot loops pinned to a single stream.

reset()

Destroy the executable and release the pool reference.

All tensors allocated during the capture must be released first.

set_conditional_handle_for_current_node(scalar_pred)

Refresh the predicate consumed by the innermost open conditional node (used for nested conditionals).

stage_and_launch(static_inputs, inputs)

Stage every input onto its static buffer and replay in one call.

Parameters:
  • static_inputs – buffers captured by the graph (kept alive by the caller).

  • inputs – fresh tensors whose contents overwrite the matching static buffer this iteration. Contiguous same-dtype/ same-device pairs take a raw async device-to-device copy; anything else falls back to full copy semantics.

This is the low-overhead bulk entry used by tensorplay.compiler.backends.cudagraphs: one Python-to-native crossing for the whole replay instead of one dispatcher round trip per input.

Variables 3

#

graphs

dataFull reference ↗

Device-agnostic capture/replay graphs for the current accelerator.

Graph records a sequence of operations on the active accelerator backend and replays it with reduced launch overhead.

Classes

Graph([keep_graph, pool, capture_error_mode])

Capture/replay graph on the current accelerator device.

#

memory

dataFull reference ↗

Device-agnostic memory queries for the current accelerator.

Every function below targets the active accelerator backend (the GPU backend in this build) and accepts a device object, a spelling such as "cuda:1", an integer index, or None for the current device.

Functions

empty_cache()

Release unoccupied cached device memory held by the allocator.

empty_host_cache()

Release unoccupied cached host memory.

get_memory_info([device])

Free and total device memory in bytes for the given device.

max_memory_allocated([device])

Peak bytes occupied by live tensors on the given device.

max_memory_reserved([device])

Peak bytes managed by the allocator on the given device.

memory_allocated([device])

Bytes currently occupied by live tensors on the given device.

memory_reserved([device])

Bytes currently managed by the allocator on the given device.

memory_stats([device])

Allocator statistics for the given device.

reset_accumulated_memory_stats([device])

Reset historical accumulation counters on the given device.

reset_peak_memory_stats([device])

Reset peak counters on the given device.

#

random

dataFull reference ↗

Device-agnostic random-number helpers for the current accelerator.

Seeds and generator states are read and written through the active accelerator backend (the GPU backend in this build). device accepts a device object, a spelling such as "cuda:1", an integer index, or None for the current device.

Functions

get_rng_state([device])

Generator state of the given accelerator device.

get_rng_state_all()

Generator states of every accelerator device.

initial_seed([device])

Initial seed of the generator for the given accelerator device.

manual_seed(seed)

Seed the generator of the current accelerator device.

manual_seed_all(seed)

Seed the generators of every accelerator device.

seed()

Reseed the generator of the current accelerator device.

seed_all()

Reseed the generators of every accelerator device.

set_rng_state(new_state[, device])

Set the generator state of the given accelerator device.

set_rng_state_all(new_states)

Set the generator states of every accelerator device.

On this page

Ask DeepWiki