# tensorplay.distributed.checkpoint API Source: https://www.tensorplay.cn/docs/api/tensorplay.distributed.checkpoint.html ## Functions 12 [#](#api-tensorplay.distributed.checkpoint.async_save) ### async_save function[Full reference ↗](/docs/generated/tensorplay.distributed.checkpoint.async_save.html) ```python tensorplay.distributed.checkpoint.async_save(state_dict, *, checkpoint_id=None, storage_writer=None, planner=None, process_group=None, async_checkpointer_type: AsyncCheckpointerType = AsyncCheckpointerType.THREAD, async_stager: AsyncStager | None = None, no_dist=False, use_collectives=True) → Future[Any] | AsyncSaveResponse ``` Stage the input and execute the checkpoint write asynchronously. [#](#api-tensorplay.distributed.checkpoint.get_model_state_dict) ### get_model_state_dict function[Full reference ↗](/docs/generated/tensorplay.distributed.checkpoint.get_model_state_dict.html) ```python tensorplay.distributed.checkpoint.get_model_state_dict(model: Any, *, submodules: set[Any] | None = None, options: StateDictOptions | None = None) → dict[str, Any] ``` [#](#api-tensorplay.distributed.checkpoint.get_optimizer_state_dict) ### get_optimizer_state_dict function[Full reference ↗](/docs/generated/tensorplay.distributed.checkpoint.get_optimizer_state_dict.html) ```python tensorplay.distributed.checkpoint.get_optimizer_state_dict(model: Any, optimizers: Any, *, submodules: set[Any] | None = None, options: StateDictOptions | None = None) → dict[str, Any] ``` [#](#api-tensorplay.distributed.checkpoint.get_state_dict) ### get_state_dict function[Full reference ↗](/docs/generated/tensorplay.distributed.checkpoint.get_state_dict.html) ```python tensorplay.distributed.checkpoint.get_state_dict(model: Any, optimizers: Any, *, submodules: set[Any] | None = None, options: StateDictOptions | None = None) → tuple[dict[str, Any], dict[str, Any]] ``` [#](#api-tensorplay.distributed.checkpoint.load_sharded_optimizer_state_dict) ### load_sharded_optimizer_state_dict function[Full reference ↗](/docs/generated/tensorplay.distributed.checkpoint.load_sharded_optimizer_state_dict.html) ```python tensorplay.distributed.checkpoint.load_sharded_optimizer_state_dict(model_state_dict: dict[str, Any], optimizer_key: str, storage_reader: Any, planner: LoadPlanner | None = None) → dict[str, Any] ``` [#](#api-tensorplay.distributed.checkpoint.load_state_dict) ### load_state_dict function[Full reference ↗](/docs/generated/tensorplay.distributed.checkpoint.load_state_dict.html) ```python tensorplay.distributed.checkpoint.load_state_dict(state_dict: dict[str, Any], storage_reader: Any, process_group: Any = None, coordinator_rank: int = 0, no_dist: bool = False, planner: Any = None) → None ``` [#](#api-tensorplay.distributed.checkpoint.load) ### load function[Full reference ↗](/docs/generated/tensorplay.distributed.checkpoint.load.html) ```python tensorplay.distributed.checkpoint.load(state_dict, *, checkpoint_id=None, storage_reader=None, planner=None, process_group=None, no_dist=False) → None ``` Load checkpoint values into an existing state dictionary. [#](#api-tensorplay.distributed.checkpoint.save_state_dict) ### save_state_dict function[Full reference ↗](/docs/generated/tensorplay.distributed.checkpoint.save_state_dict.html) ```python tensorplay.distributed.checkpoint.save_state_dict(state_dict: dict[str, Any], storage_writer: Any, process_group: Any = None, coordinator_rank: int = 0, no_dist: bool = False, planner: Any = None) → Metadata | Any ``` [#](#api-tensorplay.distributed.checkpoint.save) ### save function[Full reference ↗](/docs/generated/tensorplay.distributed.checkpoint.save.html) ```python tensorplay.distributed.checkpoint.save(state_dict, *, checkpoint_id=None, storage_writer=None, planner=None, process_group=None, no_dist=False, use_collectives=True) → Any ``` Save a state dictionary with coordinated metadata commit. [#](#api-tensorplay.distributed.checkpoint.set_model_state_dict) ### set_model_state_dict function[Full reference ↗](/docs/generated/tensorplay.distributed.checkpoint.set_model_state_dict.html) ```python tensorplay.distributed.checkpoint.set_model_state_dict(model: Any, model_state_dict: dict[str, Any], *, options: StateDictOptions | None = None) → Any ``` [#](#api-tensorplay.distributed.checkpoint.set_optimizer_state_dict) ### set_optimizer_state_dict function[Full reference ↗](/docs/generated/tensorplay.distributed.checkpoint.set_optimizer_state_dict.html) ```python tensorplay.distributed.checkpoint.set_optimizer_state_dict(model: Any, optimizers: Any, optim_state_dict: dict[str, Any], *, options: StateDictOptions | None = None) → None ``` [#](#api-tensorplay.distributed.checkpoint.set_state_dict) ### set_state_dict function[Full reference ↗](/docs/generated/tensorplay.distributed.checkpoint.set_state_dict.html) ```python tensorplay.distributed.checkpoint.set_state_dict(model: Any, optimizers: Any, *, model_state_dict: dict[str, Any], optim_state_dict: dict[str, Any], options: StateDictOptions | None = None) → Any ``` ## Classes 36 [#](#api-tensorplay.distributed.checkpoint.AsyncCheckpointerType) ### AsyncCheckpointerType class[Full reference ↗](/docs/generated/tensorplay.distributed.checkpoint.AsyncCheckpointerType.html) ```python class tensorplay.distributed.checkpoint.AsyncCheckpointerType(*values) ``` [#](#api-tensorplay.distributed.checkpoint.AsyncSaveResponse) ### AsyncSaveResponse class[Full reference ↗](/docs/generated/tensorplay.distributed.checkpoint.AsyncSaveResponse.html) ```python class tensorplay.distributed.checkpoint.AsyncSaveResponse(staging_completion: 'Future[None]', upload_completion: 'Future[Any]') ``` [#](#api-tensorplay.distributed.checkpoint.BytesIOWriteData) ### BytesIOWriteData class[Full reference ↗](/docs/generated/tensorplay.distributed.checkpoint.BytesIOWriteData.html) ```python class tensorplay.distributed.checkpoint.BytesIOWriteData(nbytes: 'int') ``` [#](#api-tensorplay.distributed.checkpoint.BytesStorageMetadata) ### BytesStorageMetadata class[Full reference ↗](/docs/generated/tensorplay.distributed.checkpoint.BytesStorageMetadata.html) ```python class tensorplay.distributed.checkpoint.BytesStorageMetadata ``` [#](#api-tensorplay.distributed.checkpoint.CheckpointableTensor) ### CheckpointableTensor class[Full reference ↗](/docs/generated/tensorplay.distributed.checkpoint.CheckpointableTensor.html) ```python class tensorplay.distributed.checkpoint.CheckpointableTensor(*args, **kwargs) ``` [#](#api-tensorplay.distributed.checkpoint.ChunkStorageMetadata) ### ChunkStorageMetadata class[Full reference ↗](/docs/generated/tensorplay.distributed.checkpoint.ChunkStorageMetadata.html) ```python class tensorplay.distributed.checkpoint.ChunkStorageMetadata(offsets: 'tuple[int, ...]', sizes: 'tuple[int, ...]') ``` [#](#api-tensorplay.distributed.checkpoint.DefaultLoadPlanner) ### DefaultLoadPlanner class[Full reference ↗](/docs/generated/tensorplay.distributed.checkpoint.DefaultLoadPlanner.html) ```python class tensorplay.distributed.checkpoint.DefaultLoadPlanner(flatten_state_dict: bool = True, flatten_sharded_tensors: bool = True, allow_partial_load: bool = False) ``` [#](#api-tensorplay.distributed.checkpoint.DefaultSavePlanner) ### DefaultSavePlanner class[Full reference ↗](/docs/generated/tensorplay.distributed.checkpoint.DefaultSavePlanner.html) ```python class tensorplay.distributed.checkpoint.DefaultSavePlanner(flatten_state_dict: bool = True, flatten_sharded_tensors: bool = True, dedup_replicated_tensors: bool | None = None, dedup_save_to_lowest_rank: bool = False, enable_plan_caching: bool = False) ``` [#](#api-tensorplay.distributed.checkpoint.FileSystem) ### FileSystem class[Full reference ↗](/docs/generated/tensorplay.distributed.checkpoint.FileSystem.html) ```python class tensorplay.distributed.checkpoint.FileSystem ``` [#](#api-tensorplay.distributed.checkpoint.FileSystemBase) ### FileSystemBase class[Full reference ↗](/docs/generated/tensorplay.distributed.checkpoint.FileSystemBase.html) ```python class tensorplay.distributed.checkpoint.FileSystemBase ``` [#](#api-tensorplay.distributed.checkpoint.FileSystemReader) ### FileSystemReader class[Full reference ↗](/docs/generated/tensorplay.distributed.checkpoint.FileSystemReader.html) ```python class tensorplay.distributed.checkpoint.FileSystemReader(path: str | PathLike[str], _extension_registry: ExtensionRegistry | None = None) ``` Read checkpoint transactions written by [FileSystemWriter](/docs/generated/tensorplay.distributed.checkpoint.FileSystemWriter.html#tensorplay.distributed.checkpoint.FileSystemWriter). [#](#api-tensorplay.distributed.checkpoint.FileSystemWriter) ### FileSystemWriter class[Full reference ↗](/docs/generated/tensorplay.distributed.checkpoint.FileSystemWriter.html) ```python class tensorplay.distributed.checkpoint.FileSystemWriter(path: str | PathLike[str], single_file_per_rank: bool = True, sync_files: bool = True, thread_count: int = 1, per_thread_copy_ahead: int = 10000000, cache_staged_state_dict: bool = False, overwrite: bool = True, _extensions: Sequence[StreamTransformExtension] | None = None, serialization_format: SerializationFormat = SerializationFormat.TORCH_SAVE) ``` [#](#api-tensorplay.distributed.checkpoint.HuggingFaceStorageReader) ### HuggingFaceStorageReader class[Full reference ↗](/docs/generated/tensorplay.distributed.checkpoint.HuggingFaceStorageReader.html) ```python class tensorplay.distributed.checkpoint.HuggingFaceStorageReader(path: str | PathLike[str], thread_count: int = 1) ``` [#](#api-tensorplay.distributed.checkpoint.HuggingFaceStorageWriter) ### HuggingFaceStorageWriter class[Full reference ↗](/docs/generated/tensorplay.distributed.checkpoint.HuggingFaceStorageWriter.html) ```python class tensorplay.distributed.checkpoint.HuggingFaceStorageWriter(path: str | PathLike[str], fqn_to_index_mapping: dict[str, int] | None = None, thread_count: int = 1, save_distributed: bool = False, enable_consolidation: bool = False, thread_count_consolidation: int = 1) ``` [#](#api-tensorplay.distributed.checkpoint.LoadItemType) ### LoadItemType class[Full reference ↗](/docs/generated/tensorplay.distributed.checkpoint.LoadItemType.html) ```python class tensorplay.distributed.checkpoint.LoadItemType(*values) ``` [#](#api-tensorplay.distributed.checkpoint.LoadPlan) ### LoadPlan class[Full reference ↗](/docs/generated/tensorplay.distributed.checkpoint.LoadPlan.html) ```python class tensorplay.distributed.checkpoint.LoadPlan(items: 'list[ReadItem]', storage_data: 'Any' = None, planner_data: 'Any' = None) ``` [#](#api-tensorplay.distributed.checkpoint.LoadPlanner) ### LoadPlanner class[Full reference ↗](/docs/generated/tensorplay.distributed.checkpoint.LoadPlanner.html) ```python class tensorplay.distributed.checkpoint.LoadPlanner ``` [#](#api-tensorplay.distributed.checkpoint.MegaStorageReader) ### MegaStorageReader class[Full reference ↗](/docs/generated/tensorplay.distributed.checkpoint.MegaStorageReader.html) ```python class tensorplay.distributed.checkpoint.MegaStorageReader(path: str | PathLike) ``` Reads MEGA checkpoints produced by [MegaStorageWriter](/docs/generated/tensorplay.distributed.checkpoint.MegaStorageWriter.html#tensorplay.distributed.checkpoint.MegaStorageWriter) (.mega shards described by model.mega.index.json). [#](#api-tensorplay.distributed.checkpoint.MegaStorageWriter) ### MegaStorageWriter class[Full reference ↗](/docs/generated/tensorplay.distributed.checkpoint.MegaStorageWriter.html) ```python class tensorplay.distributed.checkpoint.MegaStorageWriter(path: str | PathLike, fqn_to_index_mapping: dict[str, int] | None = None) ``` Writes MEGA-format shards (model[-N-of-M].mega) plus model.mega.index.json; accepts plain directories and mega:// URIs. [#](#api-tensorplay.distributed.checkpoint.Metadata) ### Metadata class[Full reference ↗](/docs/generated/tensorplay.distributed.checkpoint.Metadata.html) ```python class tensorplay.distributed.checkpoint.Metadata(state_dict_metadata: 'dict[str, TensorStorageMetadata | BytesStorageMetadata]', planner_data: 'Any' = None, storage_data: 'Any' = None, storage_meta: 'StorageMeta | None' = None, version: 'str | None' = None) ``` [#](#api-tensorplay.distributed.checkpoint.MetadataIndex) ### MetadataIndex class[Full reference ↗](/docs/generated/tensorplay.distributed.checkpoint.MetadataIndex.html) ```python class tensorplay.distributed.checkpoint.MetadataIndex(fqn: 'str', offset: 'Sequence[int] | None' = None, index: 'int | None' = None) → 'None' ``` [#](#api-tensorplay.distributed.checkpoint.QuantizedHuggingFaceStorageReader) ### QuantizedHuggingFaceStorageReader class[Full reference ↗](/docs/generated/tensorplay.distributed.checkpoint.QuantizedHuggingFaceStorageReader.html) ```python class tensorplay.distributed.checkpoint.QuantizedHuggingFaceStorageReader(path: str, thread_count: int = 1, target_dtype: Any = None, block_size: int = 128) ``` [#](#api-tensorplay.distributed.checkpoint.ReadItem) ### ReadItem class[Full reference ↗](/docs/generated/tensorplay.distributed.checkpoint.ReadItem.html) ```python class tensorplay.distributed.checkpoint.ReadItem(type: 'LoadItemType', dest_index: 'MetadataIndex', dest_offsets: 'tuple[int, ...]', storage_index: 'MetadataIndex', storage_offsets: 'tuple[int, ...]', lengths: 'tuple[int, ...]') ``` [#](#api-tensorplay.distributed.checkpoint.SavePlan) ### SavePlan class[Full reference ↗](/docs/generated/tensorplay.distributed.checkpoint.SavePlan.html) ```python class tensorplay.distributed.checkpoint.SavePlan(items: 'list[WriteItem]', storage_data: 'Any' = None, planner_data: 'Any' = None, usable: 'bool' = True) ``` [#](#api-tensorplay.distributed.checkpoint.SavePlanner) ### SavePlanner class[Full reference ↗](/docs/generated/tensorplay.distributed.checkpoint.SavePlanner.html) ```python class tensorplay.distributed.checkpoint.SavePlanner ``` [#](#api-tensorplay.distributed.checkpoint.SerializationFormat) ### SerializationFormat class[Full reference ↗](/docs/generated/tensorplay.distributed.checkpoint.SerializationFormat.html) ```python class tensorplay.distributed.checkpoint.SerializationFormat(*values) ``` [#](#api-tensorplay.distributed.checkpoint.StateDictOptions) ### StateDictOptions class[Full reference ↗](/docs/generated/tensorplay.distributed.checkpoint.StateDictOptions.html) ```python class tensorplay.distributed.checkpoint.StateDictOptions(full_state_dict: 'bool' = False, cpu_offload: 'bool' = False, ignore_frozen_params: 'bool' = False, keep_submodule_prefixes: 'bool' = True, strict: 'bool' = True, broadcast_from_rank0: 'bool' = False, flatten_optimizer_state_dict: 'bool' = False, dsd_fqn_modifiers: 'str' = '_fqn_modifiers') ``` [#](#api-tensorplay.distributed.checkpoint.Stateful) ### Stateful class[Full reference ↗](/docs/generated/tensorplay.distributed.checkpoint.Stateful.html) ```python class tensorplay.distributed.checkpoint.Stateful(*args, **kwargs) ``` [#](#api-tensorplay.distributed.checkpoint.StorageMeta) ### StorageMeta class[Full reference ↗](/docs/generated/tensorplay.distributed.checkpoint.StorageMeta.html) ```python class tensorplay.distributed.checkpoint.StorageMeta(checkpoint_id: 'str | os.PathLike[str] | None' = None, save_id: 'str | None' = None, load_id: 'str | None' = None, modules: 'list[str]' = ) ``` [#](#api-tensorplay.distributed.checkpoint.StorageReader) ### StorageReader class[Full reference ↗](/docs/generated/tensorplay.distributed.checkpoint.StorageReader.html) ```python class tensorplay.distributed.checkpoint.StorageReader ``` [#](#api-tensorplay.distributed.checkpoint.StorageWriter) ### StorageWriter class[Full reference ↗](/docs/generated/tensorplay.distributed.checkpoint.StorageWriter.html) ```python class tensorplay.distributed.checkpoint.StorageWriter ``` [#](#api-tensorplay.distributed.checkpoint.TensorProperties) ### TensorProperties class[Full reference ↗](/docs/generated/tensorplay.distributed.checkpoint.TensorProperties.html) ```python class tensorplay.distributed.checkpoint.TensorProperties(dtype: 'Any' = , layout: 'Any' = , requires_grad: 'bool' = False, memory_format: 'Any' = , pin_memory: 'bool' = False) ``` [#](#api-tensorplay.distributed.checkpoint.TensorStorageMetadata) ### TensorStorageMetadata class[Full reference ↗](/docs/generated/tensorplay.distributed.checkpoint.TensorStorageMetadata.html) ```python class tensorplay.distributed.checkpoint.TensorStorageMetadata(properties: 'TensorProperties', size: 'tuple[int, ...]', chunks: 'list[ChunkStorageMetadata]') ``` [#](#api-tensorplay.distributed.checkpoint.TensorWriteData) ### TensorWriteData class[Full reference ↗](/docs/generated/tensorplay.distributed.checkpoint.TensorWriteData.html) ```python class tensorplay.distributed.checkpoint.TensorWriteData(chunk: 'ChunkStorageMetadata', properties: 'TensorProperties', size: 'tuple[int, ...]') ``` [#](#api-tensorplay.distributed.checkpoint.WriteItem) ### WriteItem class[Full reference ↗](/docs/generated/tensorplay.distributed.checkpoint.WriteItem.html) ```python class tensorplay.distributed.checkpoint.WriteItem(index: 'MetadataIndex', type: 'WriteItemType', bytes_io_data: 'BytesIOWriteData | None' = None, tensor_data: 'TensorWriteData | None' = None) ``` [#](#api-tensorplay.distributed.checkpoint.WriteItemType) ### WriteItemType class[Full reference ↗](/docs/generated/tensorplay.distributed.checkpoint.WriteItemType.html) ```python class tensorplay.distributed.checkpoint.WriteItemType(*values) ``` ## Exceptions 1 [#](#api-tensorplay.distributed.checkpoint.CheckpointException) ### CheckpointException exception[Full reference ↗](/docs/generated/tensorplay.distributed.checkpoint.CheckpointException.html) ```python exception tensorplay.distributed.checkpoint.CheckpointException(msg: str, failures: dict[int, tuple[BaseException, StackSummary]]) ```