Buckets:
CMStochasticIterativeScheduler
Consistency Models by Yang Song, Prafulla Dhariwal, Mark Chen, and Ilya Sutskever introduced a multistep and onestep scheduler (Algorithm 1) that is capable of generating good samples in one or a small number of steps.
The abstract from the paper is:
Diffusion models have significantly advanced the fields of image, audio, and video generation, but they depend on an iterative sampling process that causes slow generation. To overcome this limitation, we propose consistency models, a new family of models that generate high quality samples by directly mapping noise to data. They support fast one-step generation by design, while still allowing multistep sampling to trade compute for sample quality. They also support zero-shot data editing, such as image inpainting, colorization, and super-resolution, without requiring explicit training on these tasks. Consistency models can be trained either by distilling pre-trained diffusion models, or as standalone generative models altogether. Through extensive experiments, we demonstrate that they outperform existing distillation techniques for diffusion models in one- and few-step sampling, achieving the new state-of-the-art FID of 3.55 on CIFAR-10 and 6.20 on ImageNet 64x64 for one-step generation. When trained in isolation, consistency models become a new family of generative models that can outperform existing one-step, non-adversarial generative models on standard benchmarks such as CIFAR-10, ImageNet 64x64 and LSUN 256x256.
The original codebase can be found at openai/consistency_models.
CMStochasticIterativeScheduler[[diffusers.CMStochasticIterativeScheduler]]
- num_train_timesteps (
int, defaults to 40) -- The number of diffusion steps to train the model. - sigma_min (
float, defaults to 0.002) -- Minimum noise magnitude in the sigma schedule. Defaults to 0.002 from the original implementation. - sigma_max (
float, defaults to 80.0) -- Maximum noise magnitude in the sigma schedule. Defaults to 80.0 from the original implementation. - sigma_data (
float, defaults to 0.5) -- The standard deviation of the data distribution from the EDM paper. Defaults to 0.5 from the original implementation. - s_noise (
float, defaults to 1.0) -- The amount of additional noise to counteract loss of detail during sampling. A reasonable range is [1.000, 1.011]. Defaults to 1.0 from the original implementation. - rho (
float, defaults to 7.0) -- The parameter for calculating the Karras sigma schedule from the EDM paper. Defaults to 7.0 from the original implementation. - clip_denoised (
bool, defaults toTrue) -- Whether to clip the denoised outputs to(-1, 1). - timesteps (
listornp.ndarrayortorch.Tensor, optional) -- An explicit timestep schedule that can be optionally specified. The timesteps are expected to be in increasing order.
Multistep and onestep sampling for consistency models.
This model inherits from SchedulerMixin and ConfigMixin. Check the superclass documentation for the generic methods the library implements for all schedulers such as loading and saving.
- original_samples (
torch.Tensor) -- The original samples to which noise will be added. - noise (
torch.Tensor) -- The noise tensor to add to the original samples. - timesteps (
torch.Tensor) -- The timesteps at which to add noise, determining the noise level from the schedule.torch.TensorThe noisy samples with added noise scaled according to the timestep schedule.
Add noise to the original samples according to the noise schedule at the specified timesteps.
- sigma (
torch.Tensor) -- The current sigma value in the noise schedule.tuple[torch.Tensor, torch.Tensor]A tuple containingc_skip(scaling for the input sample) andc_out(scaling for the model output).
Computes the scaling factors for the consistency model output.
- sigma (
torch.Tensor) -- The current sigma in the Karras sigma schedule.tuple[torch.Tensor, torch.Tensor]A two-element tuple wherec_skip(which weights the current sample) is the first element andc_out(which weights the consistency model output) is the second element.
Gets the scalings used in the consistency model parameterization (from Appendix C of the paper) to enforce boundary condition.
>
epsilonin the equations forc_skipandc_outis set tosigma_min.
- timestep (
floatortorch.Tensor) -- The timestep value to find in the schedule. - schedule_timesteps (
torch.Tensor, optional) -- The timestep schedule to search in. IfNone, usesself.timesteps.intThe index of the timestep in the schedule. For the very first step, returns the second index if multiple matches exist to avoid skipping a sigma when starting mid-schedule (e.g., for image-to-image).
Find the index of a given timestep in the timestep schedule.
- sample (
torch.Tensor) -- The input sample. - timestep (
floatortorch.Tensor) -- The current timestep in the diffusion chain.torch.TensorA scaled input sample.
Scales the consistency model input by (sigma**2 + sigma_data**2) ** 0.5.
- begin_index (
int, defaults to0) -- The begin index for the scheduler.
Sets the begin index for the scheduler. This function should be run from pipeline before the inference.
- num_inference_steps (
int, optional) -- The number of diffusion steps used when generating samples with a pre-trained model. - device (
strortorch.device, optional) -- The device to which the timesteps should be moved to. IfNone, the timesteps are not moved. - timesteps (
list[int], optional) -- Custom timesteps used to support arbitrary spacing between timesteps. IfNone, then the default timestep spacing strategy of equal spacing between timesteps is used. Iftimestepsis passed,num_inference_stepsmust beNone.
Sets the timesteps used for the diffusion chain (to be run before inference).
- sigmas (
floatornp.ndarray) -- A single Karras sigma or an array of Karras sigmas.np.ndarrayA scaled input timestep array.
Gets scaled timesteps from the Karras sigmas for input to the consistency model.
- model_output (
torch.Tensor) -- The direct output from the learned diffusion model. - timestep (
floatortorch.Tensor) -- The current timestep in the diffusion chain. - sample (
torch.Tensor) -- A current instance of a sample created by the diffusion process. - generator (
torch.Generator, optional) -- A random number generator. - return_dict (
bool, defaults toTrue) -- Whether or not to return a CMStochasticIterativeSchedulerOutput ortuple.CMStochasticIterativeSchedulerOutput ortupleIf return_dict isTrue, CMStochasticIterativeSchedulerOutput is returned, otherwise a tuple is returned where the first element is the sample tensor.
Predict the sample from the previous timestep by reversing the SDE. This function propagates the diffusion process from the learned model outputs (most often the predicted noise).
CMStochasticIterativeSchedulerOutput[[diffusers.schedulers.scheduling_consistency_models.CMStochasticIterativeSchedulerOutput]]
- prev_sample (
torch.Tensorof shape(batch_size, num_channels, height, width)for images) -- Computed sample(x_{t-1})of previous timestep.prev_sampleshould be used as next model input in the denoising loop.
Output class for the scheduler's step function.
Xet Storage Details
- Size:
- 8.39 kB
- Xet hash:
- bb302b84699e1779d93c4cd582c368f27c3b340ef127e4be63bb387eed0f10d6
Xet efficiently stores files, intelligently splitting them into unique chunks and accelerating uploads and downloads. More info.