latest (dev)
Copy
Latest development documentation · Updated 2026-10-08
tensorplay.cuda.tunable API
Functions 20
disable
functionFull reference ↗enable
functionFull reference ↗- tensorplay.cuda.tunable.enable(val: bool = True) None[source]
Turn GEMM kernel tuning on or off.
While on, every GEMM on the tunable path reuses the algorithm recorded for its signature (or measures one, when tuning is also enabled and no winner is recorded yet). While off, dispatch is exactly the untuned behavior.
get_filename
functionFull reference ↗get_max_tuning_duration
functionFull reference ↗get_max_tuning_samples
functionFull reference ↗get_results
functionFull reference ↗is_enabled
functionFull reference ↗is_verbose
functionFull reference ↗read_file
functionFull reference ↗- tensorplay.cuda.tunable.read_file(filename: str | None = None) bool[source]
Merge a tuning results file into the in-memory database.
The file’s validator lines must match the current build and device; otherwise the file is rejected and
Falseis returned. Entries already measured in this process win over the file’s. Iffilenameis not given, the configured results file is used.
record_untuned_disable
functionFull reference ↗- tensorplay.cuda.tunable.record_untuned_disable() None[source]
Stop logging untuned GEMMs. See
record_untuned_enable().
record_untuned_enable
functionFull reference ↗- tensorplay.cuda.tunable.record_untuned_enable(val: bool = True) None[source]
Control logging of GEMMs that ran without a tuned choice.
When enabled, every unique signature that runs untuned is appended to the untuned file (
tunableop_untuned<device>.csvin the working directory), one line per signature.
record_untuned_is_enabled
functionFull reference ↗set_filename
functionFull reference ↗- tensorplay.cuda.tunable.set_filename(filename: str, insert_device_ordinal: bool = False) None[source]
Set the file used to persist tuning results.
If
insert_device_ordinalisTrue, the current device ordinal is embedded in the name (a%dtoken is replaced in place, otherwise the ordinal lands before the extension). This keeps one-process-per-device runs from sharing a file. An empty filename turns file persistence off.
set_max_tuning_duration
functionFull reference ↗- tensorplay.cuda.tunable.set_max_tuning_duration(duration_ms: int) None[source]
Bound the time spent measuring one candidate, in milliseconds.
Zero disables the limit. When both this and
set_max_tuning_samples()are set, the smaller bound wins; a measurement always runs at least one timed sample.
set_max_tuning_samples
functionFull reference ↗- tensorplay.cuda.tunable.set_max_tuning_samples(samples: int) None[source]
Bound the timed samples spent measuring one candidate.
Zero disables the limit. When both this and
set_max_tuning_duration()are set, the smaller bound wins; a measurement always runs at least one timed sample.
set_verbose
functionFull reference ↗tuning_disable
functionFull reference ↗- tensorplay.cuda.tunable.tuning_disable() None[source]
Stop measuring untuned GEMM shapes. See
tuning_enable().
tuning_enable
functionFull reference ↗- tensorplay.cuda.tunable.tuning_enable(val: bool = True) None[source]
Control whether untuned GEMM shapes are measured.
When enabled, a GEMM without a recorded winner times every candidate algorithm and records the fastest; newly found winners are appended to the results file as they are measured. When disabled, such a GEMM runs the library heuristic’s top choice.
tuning_is_enabled
functionFull reference ↗write_file
functionFull reference ↗Help improve this page
Found an error, an unclear step, or a missing example?

