Buckets:

HuggingFaceDocBuilder's picture
|
download
raw
7.61 kB
# DeepSpeed utilities
## DeepSpeedPlugin
## get_active_deepspeed_plugin[[accelerate.utils.get_active_deepspeed_plugin]]
#### accelerate.utils.get_active_deepspeed_plugin[[accelerate.utils.get_active_deepspeed_plugin]]
```python
accelerate.utils.get_active_deepspeed_plugin(state)
```
[Source](https://github.com/huggingface/accelerate/blob/vr_4157/src/accelerate/utils/deepspeed.py#L100)
**Raises:** ``ValueError``
- ``ValueError`` -- If DeepSpeed was not enabled and this function is called.
Returns the currently active DeepSpeedPlugin.
#### accelerate.DeepSpeedPlugin[[accelerate.DeepSpeedPlugin]]
```python
accelerate.DeepSpeedPlugin(hf_ds_config: typing.Any = None, gradient_accumulation_steps: int = None, gradient_clipping: float = None, zero_stage: int = None, is_train_batch_min: bool = True, offload_optimizer_device: str = None, offload_param_device: str = None, offload_optimizer_nvme_path: str = None, offload_param_nvme_path: str = None, zero3_init_flag: bool = None, zero3_save_16bit_model: bool = None, transformer_moe_cls_names: str = None, enable_msamp: bool = None, msamp_opt_level: typing.Optional[typing.Literal['O1', 'O2']] = None)
```
[Source](https://github.com/huggingface/accelerate/blob/vr_4157/src/accelerate/utils/dataclasses.py#L1122)
**Parameters:**
hf_ds_config (`Any`, defaults to `None`) : Path to DeepSpeed config file or dict or an object of class `accelerate.utils.deepspeed.HfDeepSpeedConfig`.
gradient_accumulation_steps (`int`, defaults to `None`) : Number of steps to accumulate gradients before updating optimizer states. If not set, will use the value from the `Accelerator` directly.
gradient_clipping (`float`, defaults to `None`) : Enable gradient clipping with value.
zero_stage (`int`, defaults to `None`) : Possible options are 0, 1, 2, 3. Default will be taken from environment variable.
is_train_batch_min (`bool`, defaults to `True`) : If both train & eval dataloaders are specified, this will decide the `train_batch_size`.
offload_optimizer_device (`str`, defaults to `None`) : Possible options are none|cpu|nvme. Only applicable with ZeRO Stages 2 and 3.
offload_param_device (`str`, defaults to `None`) : Possible options are none|cpu|nvme. Only applicable with ZeRO Stage 3.
offload_optimizer_nvme_path (`str`, defaults to `None`) : Possible options are /nvme|/local_nvme. Only applicable with ZeRO Stage 3.
offload_param_nvme_path (`str`, defaults to `None`) : Possible options are /nvme|/local_nvme. Only applicable with ZeRO Stage 3.
zero3_init_flag (`bool`, defaults to `None`) : Flag to indicate whether to save 16-bit model. Only applicable with ZeRO Stage-3.
zero3_save_16bit_model (`bool`, defaults to `None`) : Flag to indicate whether to save 16-bit model. Only applicable with ZeRO Stage-3.
transformer_moe_cls_names (`str`, defaults to `None`) : Comma-separated list of Transformers MoE layer class names (case-sensitive). For example, `MixtralSparseMoeBlock`, `Qwen2MoeSparseMoeBlock`, `JetMoEAttention`, `JetMoEBlock`, etc.
enable_msamp (`bool`, defaults to `None`) : Flag to indicate whether to enable MS-AMP backend for FP8 training.
msamp_opt_level (`Optional[Literal["O1", "O2"]]`, defaults to `None`) : Optimization level for MS-AMP (defaults to 'O1'). Only applicable if `enable_msamp` is True. Should be one of ['O1' or 'O2'].
This plugin is used to integrate DeepSpeed.
#### deepspeed_config_process[[accelerate.DeepSpeedPlugin.deepspeed_config_process]]
```python
deepspeed_config_process(prefix = '', mismatches = None, config = None, must_match = True, **kwargs)
```
[Source](https://github.com/huggingface/accelerate/blob/vr_4157/src/accelerate/utils/dataclasses.py#L1392)
Process the DeepSpeed config with the values from the kwargs.
#### select[[accelerate.DeepSpeedPlugin.select]]
```python
select(_from_accelerator_state: bool = False)
```
[Source](https://github.com/huggingface/accelerate/blob/vr_4157/src/accelerate/utils/dataclasses.py#L1554)
Sets the HfDeepSpeedWeakref to use the current deepspeed plugin configuration
#### accelerate.utils.DummyScheduler[[accelerate.utils.DummyScheduler]]
```python
accelerate.utils.DummyScheduler(optimizer, total_num_steps = None, warmup_num_steps = 0, lr_scheduler_callable = None, **kwargs)
```
[Source](https://github.com/huggingface/accelerate/blob/vr_4157/src/accelerate/utils/deepspeed.py#L362)
**Parameters:**
optimizer (`torch.optim.optimizer.Optimizer`) : The optimizer to wrap.
total_num_steps (int, *optional*) : Total number of steps.
warmup_num_steps (int, *optional*) : Number of steps for warmup.
lr_scheduler_callable (callable, *optional*) : A callable function that creates an LR Scheduler. It accepts only one argument `optimizer`.
- ****kwargs** (additional keyword arguments, *optional*) : Other arguments.
Dummy scheduler presents model parameters or param groups, this is primarily used to follow conventional training
loop when scheduler config is specified in the deepspeed config file.
## DeepSpeedEnginerWrapper[[accelerate.utils.DeepSpeedEngineWrapper]]
#### accelerate.utils.DeepSpeedEngineWrapper[[accelerate.utils.DeepSpeedEngineWrapper]]
```python
accelerate.utils.DeepSpeedEngineWrapper(engine)
```
[Source](https://github.com/huggingface/accelerate/blob/vr_4157/src/accelerate/utils/deepspeed.py#L253)
**Parameters:**
engine (deepspeed.runtime.engine.DeepSpeedEngine) : deepspeed engine to wrap
Internal wrapper for deepspeed.runtime.engine.DeepSpeedEngine. This is used to follow conventional training loop.
#### get_global_grad_norm[[accelerate.utils.DeepSpeedEngineWrapper.get_global_grad_norm]]
```python
get_global_grad_norm()
```
[Source](https://github.com/huggingface/accelerate/blob/vr_4157/src/accelerate/utils/deepspeed.py#L286)
Get the global gradient norm from DeepSpeed engine.
## DeepSpeedOptimizerWrapper[[accelerate.utils.DeepSpeedOptimizerWrapper]]
#### accelerate.utils.DeepSpeedOptimizerWrapper[[accelerate.utils.DeepSpeedOptimizerWrapper]]
```python
accelerate.utils.DeepSpeedOptimizerWrapper(optimizer)
```
[Source](https://github.com/huggingface/accelerate/blob/vr_4157/src/accelerate/utils/deepspeed.py#L295)
**Parameters:**
optimizer (`torch.optim.optimizer.Optimizer`) : The optimizer to wrap.
Internal wrapper around a deepspeed optimizer.
## DeepSpeedSchedulerWrapper[[accelerate.utils.DeepSpeedSchedulerWrapper]]
#### accelerate.utils.DeepSpeedSchedulerWrapper[[accelerate.utils.DeepSpeedSchedulerWrapper]]
```python
accelerate.utils.DeepSpeedSchedulerWrapper(scheduler, optimizers)
```
[Source](https://github.com/huggingface/accelerate/blob/vr_4157/src/accelerate/utils/deepspeed.py#L322)
**Parameters:**
scheduler (`torch.optim.lr_scheduler.LambdaLR`) : The scheduler to wrap.
optimizers (one or a list of `torch.optim.Optimizer`) --
Internal wrapper around a deepspeed scheduler.
## DummyOptim[[accelerate.utils.DummyOptim]]
#### accelerate.utils.DummyOptim[[accelerate.utils.DummyOptim]]
```python
accelerate.utils.DummyOptim(params, lr = 0.001, weight_decay = 0, **kwargs)
```
[Source](https://github.com/huggingface/accelerate/blob/vr_4157/src/accelerate/utils/deepspeed.py#L339)
**Parameters:**
lr (float) : Learning rate.
params (iterable) : iterable of parameters to optimize or dicts defining parameter groups
weight_decay (float) : Weight decay.
- ****kwargs** (additional keyword arguments, *optional*) : Other arguments.
Dummy optimizer presents model parameters or param groups, this is primarily used to follow conventional training
loop when optimizer config is specified in the deepspeed config file.
## DummyScheduler

Xet Storage Details

Size:
7.61 kB
·
Xet hash:
1841abbd2c81a0f8f9fab1e29c6d887129c5f442291d7cd468aca642e335b687

Xet efficiently stores files, intelligently splitting them into unique chunks and accelerating uploads and downloads. More info.