# Graph Source: https://www.tensorplay.cn/docs/generated/tensorplay.accelerator.Graph.html ```python class tensorplay.accelerator.Graph(keep_graph: bool = False, *, pool: Any = None, capture_error_mode: str = 'global') ``` Capture/replay graph on the current accelerator device. Parameters: - keep_graph – accepted for generic code; the executable is compiled when capture ends in this backend. - pool – None captures into a fresh private pool, otherwise a pool id from graph_pool_handle(), another graph, or another graph’s pool id shares that pool. - capture_error_mode – "global" fails the capture on unsafe calls anywhere in the process, "thread_local" only watches this thread, "relaxed" skips the guards. ```python begin_capture_to_if_node(scalar_pred) ``` Inside an open capture, gate the following work on an if node. scalar_pred must be a single-element CUDA Bool tensor; at replay time the driver samples it and runs the body captured between this call and [end_capture_to_conditional_node()](#tensorplay.accelerator.Graph.end_capture_to_conditional_node) only when true. ```python begin_capture_to_while_node(scalar_pred) ``` Like [begin_capture_to_if_node()](#tensorplay.accelerator.Graph.begin_capture_to_if_node), but the body loops while the predicate stays true (driver-level while node). ```python capture_begin(pool: Any = '__unset__', capture_error_mode: Any = '__unset__', stream: Any = None) → None ``` Begin capture on the current stream with the stored settings. ```python capture_end() ``` End capture and compile the executable (paid here, not on first replay). ```python debug_dump(path) ``` Write a DOT rendering of the captured graph to path. Call enable_debug_mode() before capturing for a dump that includes full node attributes. ```python end_capture_to_conditional_node() ``` Close the open conditional body; subsequent capture returns to the parent stream. ```python instantiate() ``` No-op once instantiated; kept for late callers. ```python pool() ``` Opaque id of this graph’s memory pool, shareable with others. ```python property pool_id ``` Allocator pool id this graph captured against. ```python replay(stream=None) ``` Run the graph: launch the cached executable on the current stream. Parameters: stream ([Stream](/docs/generated/tensorplay.cuda.streams.Stream.html#tensorplay.cuda.streams.Stream), optional) – launch on this explicit stream instead of querying the current one - shaves a TLS lookup off hot loops pinned to a single stream. ```python reset() ``` Destroy the executable and release the pool reference. All tensors allocated during the capture must be released first. ```python set_conditional_handle_for_current_node(scalar_pred) ``` Refresh the predicate consumed by the innermost open conditional node (used for nested conditionals). ```python stage_and_launch(static_inputs, inputs) ``` Stage every input onto its static buffer and replay in one call. Parameters: - static_inputs – buffers captured by the graph (kept alive by the caller). - inputs – fresh tensors whose contents overwrite the matching static buffer this iteration. Contiguous same-dtype/ same-device pairs take a raw async device-to-device copy; anything else falls back to full copy semantics. This is the low-overhead bulk entry used by tensorplay.compiler.backends.cudagraphs: one Python-to-native crossing for the whole replay instead of one dispatcher round trip per input.