Buckets:
DeepSpeed utilities
DeepSpeedPlugin
get_active_deepspeed_plugin[[accelerate.utils.get_active_deepspeed_plugin]]
ValueError-- If DeepSpeed was not enabled and this function is called.ValueError
Returns the currently active DeepSpeedPlugin.
- hf_ds_config (
Any, defaults toNone) -- Path to DeepSpeed config file or dict or an object of classaccelerate.utils.deepspeed.HfDeepSpeedConfig. - gradient_accumulation_steps (
int, defaults toNone) -- Number of steps to accumulate gradients before updating optimizer states. If not set, will use the value from theAcceleratordirectly. - gradient_clipping (
float, defaults toNone) -- Enable gradient clipping with value. - zero_stage (
int, defaults toNone) -- Possible options are 0, 1, 2, 3. Default will be taken from environment variable. - is_train_batch_min (
bool, defaults toTrue) -- If both train & eval dataloaders are specified, this will decide thetrain_batch_size. - offload_optimizer_device (
str, defaults toNone) -- Possible options are none|cpu|nvme. Only applicable with ZeRO Stages 2 and 3. - offload_param_device (
str, defaults toNone) -- Possible options are none|cpu|nvme. Only applicable with ZeRO Stage 3. - offload_optimizer_nvme_path (
str, defaults toNone) -- Possible options are /nvme|/local_nvme. Only applicable with ZeRO Stage 3. - offload_param_nvme_path (
str, defaults toNone) -- Possible options are /nvme|/local_nvme. Only applicable with ZeRO Stage 3. - zero3_init_flag (
bool, defaults toNone) -- Flag to indicate whether to save 16-bit model. Only applicable with ZeRO Stage-3. - zero3_save_16bit_model (
bool, defaults toNone) -- Flag to indicate whether to save 16-bit model. Only applicable with ZeRO Stage-3. - transformer_moe_cls_names (
str, defaults toNone) -- Comma-separated list of Transformers MoE layer class names (case-sensitive). For example,MixtralSparseMoeBlock,Qwen2MoeSparseMoeBlock,JetMoEAttention,JetMoEBlock, etc. - enable_msamp (
bool, defaults toNone) -- Flag to indicate whether to enable MS-AMP backend for FP8 training. - msamp_opt_level (
Optional[Literal["O1", "O2"]], defaults toNone) -- Optimization level for MS-AMP (defaults to 'O1'). Only applicable ifenable_msampis True. Should be one of ['O1' or 'O2'].
This plugin is used to integrate DeepSpeed.
Process the DeepSpeed config with the values from the kwargs.
Sets the HfDeepSpeedWeakref to use the current deepspeed plugin configuration
- optimizer (
torch.optim.optimizer.Optimizer) -- The optimizer to wrap. - total_num_steps (int, optional) -- Total number of steps.
- warmup_num_steps (int, optional) -- Number of steps for warmup.
- lr_scheduler_callable (callable, optional) --
A callable function that creates an LR Scheduler. It accepts only one argument
optimizer. - **kwargs (additional keyword arguments, optional) -- Other arguments.
Dummy scheduler presents model parameters or param groups, this is primarily used to follow conventional training loop when scheduler config is specified in the deepspeed config file.
DeepSpeedEnginerWrapper[[accelerate.utils.DeepSpeedEngineWrapper]]
- engine (deepspeed.runtime.engine.DeepSpeedEngine) -- deepspeed engine to wrap
Internal wrapper for deepspeed.runtime.engine.DeepSpeedEngine. This is used to follow conventional training loop.
Get the global gradient norm from DeepSpeed engine.
DeepSpeedOptimizerWrapper[[accelerate.utils.DeepSpeedOptimizerWrapper]]
- optimizer (
torch.optim.optimizer.Optimizer) -- The optimizer to wrap.
Internal wrapper around a deepspeed optimizer.
DeepSpeedSchedulerWrapper[[accelerate.utils.DeepSpeedSchedulerWrapper]]
- scheduler (
torch.optim.lr_scheduler.LambdaLR) -- The scheduler to wrap. - optimizers (one or a list of
torch.optim.Optimizer) --
Internal wrapper around a deepspeed scheduler.
DummyOptim[[accelerate.utils.DummyOptim]]
- lr (float) -- Learning rate.
- params (iterable) -- iterable of parameters to optimize or dicts defining parameter groups
- weight_decay (float) -- Weight decay.
- **kwargs (additional keyword arguments, optional) -- Other arguments.
Dummy optimizer presents model parameters or param groups, this is primarily used to follow conventional training loop when optimizer config is specified in the deepspeed config file.
DummyScheduler
Xet Storage Details
- Size:
- 4.58 kB
- Xet hash:
- 6b10cf430c69f9ed52e2db3fb5c79cd2974be67dbb1df3ee4bfbbdc19adcc92b
Xet efficiently stores files, intelligently splitting them into unique chunks and accelerating uploads and downloads. More info.