Audio-to-Audio
Transformers
Safetensors
mossformer-dns
feature-extraction
mossformer2
speech-enhancement
denoising
48khz
custom-code
custom_code
Instructions to use MigoXV/mossformer2-se-48k with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use MigoXV/mossformer2-se-48k with Transformers:
# Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("MigoXV/mossformer2-se-48k", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
| import torch.nn as nn | |
| from torch import Tensor | |
| class Transpose(nn.Module): | |
| """Wrapper class of torch.transpose() for Sequential module.""" | |
| def __init__(self, shape: tuple): | |
| super(Transpose, self).__init__() | |
| self.shape = shape | |
| def forward(self, x: Tensor) -> Tensor: | |
| return x.transpose(*self.shape) | |
| class DepthwiseConv1d(nn.Module): | |
| """ | |
| When groups == in_channels and out_channels == K * in_channels, where K is a positive integer, | |
| this operation is termed in literature as depthwise convolution. | |
| Args: | |
| in_channels (int): Number of channels in the input | |
| out_channels (int): Number of channels produced by the convolution | |
| kernel_size (int or tuple): Size of the convolving kernel | |
| stride (int, optional): Stride of the convolution. Default: 1 | |
| padding (int or tuple, optional): Zero-padding added to both sides of the input. Default: 0 | |
| bias (bool, optional): If True, adds a learnable bias to the output. Default: True | |
| Inputs: inputs | |
| - **inputs** (batch, in_channels, time): Tensor containing input vector | |
| Returns: outputs | |
| - **outputs** (batch, out_channels, time): Tensor produces by depthwise 1-D convolution. | |
| """ | |
| def __init__( | |
| self, | |
| in_channels: int, | |
| out_channels: int, | |
| kernel_size: int, | |
| stride: int = 1, | |
| padding: int = 0, | |
| bias: bool = False, | |
| ) -> None: | |
| super(DepthwiseConv1d, self).__init__() | |
| assert ( | |
| out_channels % in_channels == 0 | |
| ), "out_channels should be constant multiple of in_channels" | |
| self.conv = nn.Conv1d( | |
| in_channels=in_channels, | |
| out_channels=out_channels, | |
| kernel_size=kernel_size, | |
| groups=in_channels, | |
| stride=stride, | |
| padding=padding, | |
| bias=bias, | |
| ) | |
| def forward(self, inputs: Tensor) -> Tensor: | |
| return self.conv(inputs) | |
| class ConvModule(nn.Module): | |
| """ | |
| Conformer convolution module starts with a pointwise convolution and a gated linear unit (GLU). | |
| This is followed by a single 1-D depthwise convolution layer. Batchnorm is deployed just after the convolution | |
| to aid training deep models. | |
| Args: | |
| in_channels (int): Number of channels in the input | |
| kernel_size (int or tuple, optional): Size of the convolving kernel Default: 31 | |
| dropout_p (float, optional): probability of dropout | |
| Inputs: inputs | |
| inputs (batch, time, dim): Tensor contains input sequences | |
| Outputs: outputs | |
| outputs (batch, time, dim): Tensor produces by conformer convolution module. | |
| """ | |
| def __init__( | |
| self, | |
| in_channels: int, | |
| kernel_size: int = 17, | |
| expansion_factor: int = 2, | |
| dropout_p: float = 0.1, | |
| ) -> None: | |
| super(ConvModule, self).__init__() | |
| assert ( | |
| kernel_size - 1 | |
| ) % 2 == 0, "kernel_size should be a odd number for 'SAME' padding" | |
| assert expansion_factor == 2, "Currently, Only Supports expansion_factor 2" | |
| self.sequential = nn.Sequential( | |
| Transpose(shape=(1, 2)), | |
| DepthwiseConv1d( | |
| in_channels, | |
| in_channels, | |
| kernel_size, | |
| stride=1, | |
| padding=(kernel_size - 1) // 2, | |
| ), | |
| ) | |
| def forward(self, inputs: Tensor) -> Tensor: | |
| return inputs + self.sequential(inputs).transpose(1, 2) | |