TensorPlay
API reference
latest (dev)
Copy
View Markdown

Latest development documentation · Updated 2026-10-08

tensorplay.cuda.tunable API

Functions 20

#

enable

functionFull reference ↗
tensorplay.cuda.tunable.enable(val: bool = True) → None[source]

Turn GEMM kernel tuning on or off.

While on, every GEMM on the tunable path reuses the algorithm recorded for its signature (or measures one, when tuning is also enabled and no winner is recorded yet). While off, dispatch is exactly the untuned behavior.

#

get_filename

functionFull reference ↗
tensorplay.cuda.tunable.get_filename() → str[source]

Return the configured results filename (empty until set or first use).

#

get_max_tuning_duration

functionFull reference ↗
tensorplay.cuda.tunable.get_max_tuning_duration() → int[source]

Return the per-candidate measurement time limit in milliseconds.

#

get_max_tuning_samples

functionFull reference ↗
tensorplay.cuda.tunable.get_max_tuning_samples() → int[source]

Return the per-candidate sample limit.

#

read_file

functionFull reference ↗
tensorplay.cuda.tunable.read_file(filename: str | None = None) → bool[source]

Merge a tuning results file into the in-memory database.

The file’s validator lines must match the current build and device; otherwise the file is rejected and False is returned. Entries already measured in this process win over the file’s. If filename is not given, the configured results file is used.

#

record_untuned_enable

functionFull reference ↗
tensorplay.cuda.tunable.record_untuned_enable(val: bool = True) → None[source]

Control logging of GEMMs that ran without a tuned choice.

When enabled, every unique signature that runs untuned is appended to the untuned file (tunableop_untuned<device>.csv in the working directory), one line per signature.

#

record_untuned_is_enabled

functionFull reference ↗
tensorplay.cuda.tunable.record_untuned_is_enabled() → bool[source]

Return whether untuned GEMMs are being logged.

#

set_filename

functionFull reference ↗
tensorplay.cuda.tunable.set_filename(filename: str, insert_device_ordinal: bool = False) → None[source]

Set the file used to persist tuning results.

If insert_device_ordinal is True, the current device ordinal is embedded in the name (a %d token is replaced in place, otherwise the ordinal lands before the extension). This keeps one-process-per-device runs from sharing a file. An empty filename turns file persistence off.

#

set_max_tuning_duration

functionFull reference ↗
tensorplay.cuda.tunable.set_max_tuning_duration(duration_ms: int) → None[source]

Bound the time spent measuring one candidate, in milliseconds.

Zero disables the limit. When both this and set_max_tuning_samples() are set, the smaller bound wins; a measurement always runs at least one timed sample.

#

set_max_tuning_samples

functionFull reference ↗
tensorplay.cuda.tunable.set_max_tuning_samples(samples: int) → None[source]

Bound the timed samples spent measuring one candidate.

Zero disables the limit. When both this and set_max_tuning_duration() are set, the smaller bound wins; a measurement always runs at least one timed sample.

#

set_verbose

functionFull reference ↗
tensorplay.cuda.tunable.set_verbose(val: bool) → None[source]

Turn diagnostic logging of the tuning context on or off.

Verbose output goes to stderr and reports state changes, measurement passes and file activity. It is meant for debugging.

#

tuning_enable

functionFull reference ↗
tensorplay.cuda.tunable.tuning_enable(val: bool = True) → None[source]

Control whether untuned GEMM shapes are measured.

When enabled, a GEMM without a recorded winner times every candidate algorithm and records the fastest; newly found winners are appended to the results file as they are measured. When disabled, such a GEMM runs the library heuristic’s top choice.

#

tuning_is_enabled

functionFull reference ↗
tensorplay.cuda.tunable.tuning_is_enabled() → bool[source]

Return whether untuned GEMM shapes are measured.

#

write_file

functionFull reference ↗
tensorplay.cuda.tunable.write_file() → None[source]

Rewrite the results file with every recorded winner.

The file is recreated with fresh validator lines followed by all in-memory results (both those read from files and those measured in this process).

On this page

Ask DeepWiki