# tensorplay.accelerator API Source: https://www.tensorplay.cn/docs/api/tensorplay.accelerator.html ## Functions 7 [#](#api-tensorplay.accelerator.current_accelerator) ### current_accelerator function[Full reference ↗](/docs/generated/tensorplay.accelerator.current_accelerator.html) ```python tensorplay.accelerator.current_accelerator(check_available: bool = False) ``` The accelerator device selected at build time, if any. Returns None on a host-only build. With check_available set, a runtime availability probe is also required before the device is reported. [#](#api-tensorplay.accelerator.current_device_index) ### current_device_index function[Full reference ↗](/docs/generated/tensorplay.accelerator.current_device_index.html) ```python tensorplay.accelerator.current_device_index() → int ``` Index of the currently selected accelerator device. [#](#api-tensorplay.accelerator.device_count) ### device_count function[Full reference ↗](/docs/generated/tensorplay.accelerator.device_count.html) ```python tensorplay.accelerator.device_count() → int ``` Number of devices for the current accelerator, or zero without one. [#](#api-tensorplay.accelerator.get_device_capability) ### get_device_capability function[Full reference ↗](/docs/generated/tensorplay.accelerator.get_device_capability.html) ```python tensorplay.accelerator.get_device_capability(device=None) → dict[str, Any] ``` Capability map for an accelerator device. The map carries a supported_dtypes set listing the data types that can be allocated on the device. [#](#api-tensorplay.accelerator.is_available) ### is_available function[Full reference ↗](/docs/generated/tensorplay.accelerator.is_available.html) ```python tensorplay.accelerator.is_available() → bool ``` Whether an accelerator was built and at least one device is visible. [#](#api-tensorplay.accelerator.set_device_index) ### set_device_index function[Full reference ↗](/docs/generated/tensorplay.accelerator.set_device_index.html) ```python tensorplay.accelerator.set_device_index(device) → None ``` Select the accelerator device by index; negative indices are no-ops. [#](#api-tensorplay.accelerator.synchronize) ### synchronize function[Full reference ↗](/docs/generated/tensorplay.accelerator.synchronize.html) ```python tensorplay.accelerator.synchronize(device=None) → None ``` Wait for all work on an accelerator device to complete. ## Classes 1 [#](#api-tensorplay.accelerator.Graph) ### Graph class[Full reference ↗](/docs/generated/tensorplay.accelerator.Graph.html) ```python class tensorplay.accelerator.Graph(keep_graph: bool = False, *, pool: Any = None, capture_error_mode: str = 'global') ``` Capture/replay graph on the current accelerator device. Parameters: - keep_graph – accepted for generic code; the executable is compiled when capture ends in this backend. - pool – None captures into a fresh private pool, otherwise a pool id from graph_pool_handle(), another graph, or another graph’s pool id shares that pool. - capture_error_mode – "global" fails the capture on unsafe calls anywhere in the process, "thread_local" only watches this thread, "relaxed" skips the guards. ```python begin_capture_to_if_node(scalar_pred) ``` Inside an open capture, gate the following work on an if node. scalar_pred must be a single-element CUDA Bool tensor; at replay time the driver samples it and runs the body captured between this call and [end_capture_to_conditional_node()](#tensorplay.accelerator.Graph.end_capture_to_conditional_node) only when true. ```python begin_capture_to_while_node(scalar_pred) ``` Like [begin_capture_to_if_node()](#tensorplay.accelerator.Graph.begin_capture_to_if_node), but the body loops while the predicate stays true (driver-level while node). ```python capture_begin(pool: Any = '__unset__', capture_error_mode: Any = '__unset__', stream: Any = None) → None ``` Begin capture on the current stream with the stored settings. ```python capture_end() ``` End capture and compile the executable (paid here, not on first replay). ```python debug_dump(path) ``` Write a DOT rendering of the captured graph to path. Call enable_debug_mode() before capturing for a dump that includes full node attributes. ```python end_capture_to_conditional_node() ``` Close the open conditional body; subsequent capture returns to the parent stream. ```python instantiate() ``` No-op once instantiated; kept for late callers. ```python pool() ``` Opaque id of this graph’s memory pool, shareable with others. ```python property pool_id ``` Allocator pool id this graph captured against. ```python replay(stream=None) ``` Run the graph: launch the cached executable on the current stream. Parameters: stream ([Stream](/docs/generated/tensorplay.cuda.streams.Stream.html#tensorplay.cuda.streams.Stream), optional) – launch on this explicit stream instead of querying the current one - shaves a TLS lookup off hot loops pinned to a single stream. ```python reset() ``` Destroy the executable and release the pool reference. All tensors allocated during the capture must be released first. ```python set_conditional_handle_for_current_node(scalar_pred) ``` Refresh the predicate consumed by the innermost open conditional node (used for nested conditionals). ```python stage_and_launch(static_inputs, inputs) ``` Stage every input onto its static buffer and replay in one call. Parameters: - static_inputs – buffers captured by the graph (kept alive by the caller). - inputs – fresh tensors whose contents overwrite the matching static buffer this iteration. Contiguous same-dtype/ same-device pairs take a raw async device-to-device copy; anything else falls back to full copy semantics. This is the low-overhead bulk entry used by tensorplay.compiler.backends.cudagraphs: one Python-to-native crossing for the whole replay instead of one dispatcher round trip per input. ## Variables 3 [#](#api-tensorplay.accelerator.graphs) ### graphs data[Full reference ↗](/docs/generated/tensorplay.accelerator.graphs.html) Device-agnostic capture/replay graphs for the current accelerator. Graph records a sequence of operations on the active accelerator backend and replays it with reduced launch overhead. Classes | Graph([keep_graph, pool, capture_error_mode]) |Capture/replay graph on the current accelerator device. | | --- | --- | [#](#api-tensorplay.accelerator.memory) ### memory data[Full reference ↗](/docs/generated/tensorplay.accelerator.memory.html) Device-agnostic memory queries for the current accelerator. Every function below targets the active accelerator backend (the GPU backend in this build) and accepts a device object, a spelling such as "cuda:1", an integer index, or None for the current device. Functions | empty_cache() |Release unoccupied cached device memory held by the allocator. | | --- | --- | | empty_host_cache() | Release unoccupied cached host memory. | | get_memory_info([device]) | Free and total device memory in bytes for the given device. | | max_memory_allocated([device]) | Peak bytes occupied by live tensors on the given device. | | max_memory_reserved([device]) | Peak bytes managed by the allocator on the given device. | | memory_allocated([device]) | Bytes currently occupied by live tensors on the given device. | | memory_reserved([device]) | Bytes currently managed by the allocator on the given device. | | memory_stats([device]) | Allocator statistics for the given device. | | reset_accumulated_memory_stats([device]) | Reset historical accumulation counters on the given device. | | reset_peak_memory_stats([device]) | Reset peak counters on the given device. | [#](#api-tensorplay.accelerator.random) ### random data[Full reference ↗](/docs/generated/tensorplay.accelerator.random.html) Device-agnostic random-number helpers for the current accelerator. Seeds and generator states are read and written through the active accelerator backend (the GPU backend in this build). device accepts a device object, a spelling such as "cuda:1", an integer index, or None for the current device. Functions | get_rng_state([device]) |Generator state of the given accelerator device. | | --- | --- | | get_rng_state_all() | Generator states of every accelerator device. | | initial_seed([device]) | Initial seed of the generator for the given accelerator device. | | manual_seed(seed) | Seed the generator of the current accelerator device. | | manual_seed_all(seed) | Seed the generators of every accelerator device. | | seed() | Reseed the generator of the current accelerator device. | | seed_all() | Reseed the generators of every accelerator device. | | set_rng_state(new_state[, device]) | Set the generator state of the given accelerator device. | | set_rng_state_all(new_states) | Set the generator states of every accelerator device. |