VyvoUp EfficientHBR v2.0

Canlı Demo

VyvoUp is a compact causal speech-bandwidth-extension model. It accepts mono 24 kHz speech and returns exactly aligned mono 48 kHz speech. This release is the validation-promoted successor to the immutable v1.0-real-pcm tag. It uses a conservatively retained 3% of a synthetic-speech adaptation update; the other 97% of its trainable weights remain at the real-PCM v1 checkpoint.

The release includes the model, checksum sidecar, fixed-policy outputs, complete validation and test reports, failed adaptation audits, synthetic-data provenance, RTX 5070 Ti benchmarks, and all-utterance FlashSR/LSD and DNSMOS comparisons.

VoiceSR: Vyvo TTS ile 10 yeni karşılaştırma

Aşağıdaki 10 İngilizce giriş, sabitlenmiş Vyvo/Vyvo-Multilingual-EN-FT-v0.1 TTS checkpoint'i ile üretildi. Her metin için dört TTS adayı oluşturuldu; yalnızca 24 kHz girişin Whisper transkriptinde en az kelime hatası veren aday, iki SR modeli çalıştırılmadan önce seçildi. SR çıktıları seçimi etkilemedi.

VyvoUp yayımlanan 24 kHz PCM16 girişi doğrudan alır. FlashSR yalnızca 16 kHz kabul ettiği için aynı konuşmanın deterministik 24→16 kHz sürümünü alır. Her iki SR çıktısı mono 48 kHz PCM16'dır. Oynat düğmeleri doğrudan bu model deposundaki dosyaları açar; bu örnekler Space sayfasına bağlı değildir.

# Metin Vyvo TTS girişi, 24 kHz VyvoUp, 48 kHz FlashSR, 48 kHz
01 Clear speech should remain natural when the bandwidth is extended.
02 The quiet morning train arrived exactly at seven thirty.
03 Please place the bright blue folder beside the wooden chair.
04 A quick brown fox jumps over the lazy dog near the river.
05 Engineers compared every signal before choosing the final model.
06 Fresh coffee and warm bread filled the small kitchen.
07 Can you hear the soft rain tapping against the window?
08 Three curious children watched the silver rocket cross the sky.
09 Accurate audio restoration requires careful listening and fair tests.
10 Today we are testing two speech super resolution systems.

Toplam benzersiz konuşma süresi 37,20 saniye. Sabitlenmiş Whisper-large-v3-turbo anlaşılabilirlik kontrolünde 97 hedef kelime için TTS girişi ve VyvoUp çıktısı 3, FlashSR çıktısı 5 kelime hatası verdi. Bu kontrol algısal ses kalitesini veya üretilen yüksek frekansların doğruluğunu ölçmez. 48 kHz gerçek hedef bulunmadığı için bu bölüm LSD/DNSMOS üstünlük testi değil, doğrudan dinleme karşılaştırmasıdır.

Aynı konuşmaların gerçek FlashSR 16 kHz girdileri, VCTK referansı, 71 oynatıcı, bütün transkriptler, sürümler ve SHA-256 kayıtları bu model deposundaki ayrıntılı sayfadadır.

Hemen dinle: 5 giriş + 5 çıktı

Her satırda önce modelin aldığı mono 24 kHz giriş, sonra VyvoUp'ın ürettiği mono 48 kHz çıktı yer alır. Oynat düğmelerine basarak sayfadan ayrılmadan toplam 10 WAV dosyasını dinleyebilirsiniz. Kendi sesinizi yükleyip denemek için canlı demo Space'ini açın.

Örnek 24 kHz giriş 48 kHz VyvoUp çıktısı Süre
1 — %10 yüksek-frekans enerjisi 3.50 sn
2 — %30 yüksek-frekans enerjisi 2.97 sn
3 — %50 yüksek-frekans enerjisi 3.04 sn
4 — %70 yüksek-frekans enerjisi 3.16 sn
5 — %90 yüksek-frekans enerjisi 3.87 sn

Dosyalar tarayıcı uyumluluğu için PCM16 WAV olarak yayımlandı. Model çıktıları, yayımlanan giriş WAV'ları tekrar okunarak üretildi; dolayısıyla her çıktı tam olarak yanındaki girişe karşılık gelir. Seçim politikası, kimlikler ve tüm SHA-256 değerleri demo/metadata.json içindedir.

Model

Property Value
Architecture EfficientHBR, bias-free causal gated temporal convolutions
Trainable parameters 32,112
Input / output mono 24 kHz / mono 48 kHz
Receptive field 2,059 input samples / 85.79 ms
Algorithmic delay 308 output samples / 6.42 ms
Estimated compute 801.288 million MAC/s
Streaming state (FP32) 398,932 bytes
Selected checkpoint 3% v2-adapted + 97% v1-real trainable parameters
Checkpoint SHA-256 133a1d019d05c9a146f3ea618b868021268a58842f09841f85f900e5b05774c7

The neural network predicts two interleaved high-band residual phases. A fixed complementary FIR path preserves the observed band and adds only a high-pass residual. Exact digital silence maps to exact silence.

Data

The real-speech source is the official CSTR VCTK Corpus 0.92 archive, microphone 1, licensed CC BY 4.0. The speaker-disjoint real split contains 10.0001 hours for training, 0.5001 hours for validation, and 0.5008 hours for the held-out minimum-phase test.

Synthetic targets were generated as exact mono 48 kHz audio. All model and implementation revisions are pinned in data/synthetic-*-generation.json.

Round 48 kHz TTS sources Nominal train Unique usable train
1 dots.tts SOAR, Freya, Irodori v4.1 Small 10.0057 h 8.6692 h
2 dots.tts MF 2-step, Freya, Irodori v4.1 Small 5.0063 h 5.0063 h
Total three independently implemented model families 15.0120 h 13.6755 h

Round 1 contained 495 byte-identical repeated Freya targets totaling 1.3365 hours. The audit detected and removed them before combined training. Round 2 rejected duplicates globally during generation. The resulting combined train set has 23.6756 unique hours: 10.0001 real, 8.6692 round-1 synthetic, and 5.0063 round-2 synthetic.

Pinned round-2 sources, current at generation time (2026-08-13):

  • dots-studio/dots.tts-mf-2steps revision 159b33d33de0f9610d9ea73725a0820d27261fd7, native 48 kHz.
  • freyavoice/freya-tts revision d124e07493615208f58bdd21d432736849ee4230, native 48 kHz.
  • Aratako/Irodori-TTS-v4.1-Small revision 2b28324dc263ed5e6638b3cf3dd94c82ead07b4b, 48 kHz output. Its SilentCipher watermark path internally uses 44.1 kHz and restores 48 kHz; this is disclosed because it is not an untouched native-48 kHz path.

Train/validation/test voice controls and target text are disjoint. Training inputs use Kaiser-sinc and zero-phase Chebyshev-II degradations; tests use an unseen minimum-phase degradation. Generator source snapshots and SHA-256 hashes are under data/source-snapshots/.

Improvement history

The initial real-only model trained for 80,000 steps on one RTX 5070 Ti. Two subsequent gradient-training attempts were kept as evidence but not promoted:

  1. Synthetic-heavy v2: 60,000 steps, 18m20.9s. It greatly improved synthetic validation but regressed unseen-real LSD, so its test splits stayed sealed.
  2. Real-replay v3: 5,000 steps with an explicit 80% real / 12% round-1 / 8% round-2 crop mix, 98.42s. It passed all safety gates but still exceeded the real-domain regression cap, so it was rejected without opening tests.
  3. Validation soup v4: a predeclared 1–10% interpolation grid between v1 and the strongest v2 validation checkpoint. The 3% candidate was the best one satisfying every hard gate and was promoted before test data was opened.

All rejected and promoted selection reports are in training/.

Validation and held-out tests

Lower 12–24 kHz log-spectral distance (LSD) is better. The locked validation score weights unseen-real VCTK 0.60, round-1 synthetic 0.25, and round-2 synthetic 0.15. Every validation set also required zero clipping and worst-case protected-band error at or below -115 dB.

Validation set v1 LSD v2.0 LSD Change
Unseen-real VCTK 8.7003 8.7757 +0.0754 dB
Round-1 synthetic 41.7658 41.2764 -0.4894 dB
Round-2/latest synthetic 42.2015 41.7082 -0.4934 dB
Weighted score 21.9919 21.8408 -0.1511 dB

The synthetic tests were opened exactly once after checkpoint selection. The VCTK test had already been opened for v1 and is only a disclosed regression set for this release.

Test set v1 LSD v2.0 LSD Result
VCTK, unseen speakers + degradation (527 / 30.05 min) 8.3674 8.3903 +0.0229 dB
Round-1 synthetic (171 / 30.40 min) 41.5747 41.0840 0.4907 dB better
Round-2/latest synthetic (65 / 15.57 min) 42.8360 42.3354 0.5005 dB better

All final test sets had zero clipping. Worst protected-band errors were -126.25, -126.56, and -135.29 dB respectively. The complete per-item evidence and the test-opening declaration are in metrics/.

Synthetic-domain LSD remains much higher than VCTK LSD and the model over-produces synthetic high-band energy by roughly 7–8 dB on average. This is an important limitation, not hidden by the relative improvement.

FlashSR comparison

The pinned official repository at ysharma3501/FlashSR, implementation revision 2a69326250613c0a0f6c1c8d9f0c48cb779842b8, and checkpoint SHA-256 62c70874ac4efeb4dc9c8aa9dc0a611a951e1c36292abeb4c406d7fb91e0eefc were evaluated on all 527 VCTK test utterances (30.05 minutes).

Protocol VyvoUp LSD FlashSR LSD VyvoUp margin
Native inputs: 24 kHz vs 16 kHz 8.3320 11.7206 3.3886 dB
Same 16 kHz information 9.8090 11.7206 1.9116 dB
Native, RMS-gain aligned diagnostic 8.3300 12.9994 4.6695 dB
Same information, RMS-gain aligned 9.7950 12.9994 3.2044 dB

VyvoUp has 32,112 trainable parameters; the loaded FlashSR graph has 131,904 parameters (88,128 active in its forward path). Mean end-to-end RTF in this comparison was 0.004044 for VyvoUp and 0.004042 for FlashSR. FlashSR's official wrapper peak-normalizes every clip, so both raw wrapper output and a gain-aligned diagnostic are shown. The linked FlashSR repository implements a small HiFi-GAN/HierSpeech++-style upsampler, not the diffusion model from the similarly named paper. Full caveats, per-item metrics, and four paired listening sets are in comparison/.

DNSMOS P.835 diagnostic

The pinned local models from Microsoft DNS-Challenge were also run on the same 527 held-out utterances using the regular, non-personalized DNSMOS P.835 protocol. Higher scores are better. The main table below matches each output's RMS to its target with one scalar gain, preventing FlashSR's official per-clip peak normalization from dominating the result.

Protocol / system SIG BAK OVRL P.808
VyvoUp, native 24 kHz input 3.5174 4.0397 3.2149 3.6400
VyvoUp, same 16 kHz information 3.5167 4.0396 3.2142 3.6383
FlashSR, 16 kHz input 3.4883 3.9483 3.1272 3.6124

Paired VyvoUp-minus-FlashSR differences and 95% bootstrap intervals:

Protocol SIG BAK OVRL P.808
Native +0.0291 [0.0216, 0.0366] +0.0915 [0.0820, 0.1011] +0.0877 [0.0788, 0.0966] +0.0276 [0.0220, 0.0332]
Same 16 kHz information +0.0284 [0.0209, 0.0359] +0.0914 [0.0819, 0.1010] +0.0870 [0.0781, 0.0958] +0.0259 [0.0205, 0.0313]

The raw official-wrapper OVRL scores were 3.2151 for VyvoUp native, 3.2144 for VyvoUp equal-information, and 2.9000 for FlashSR. DNSMOS downsamples every output to 16 kHz, removing the synthesized 8–24 kHz band. It is therefore a low-band speech/noise-quality diagnostic, not a direct high-band reconstruction metric or a subjective listening result. The 12–24 kHz LSD comparison remains primary. Full distributions, caveats, pinned Microsoft revision/model hashes, and all 527 per-item records are in comparison/dnsmos/.

Aynı 10 konuşma: üç çıkışı doğrudan dinle

DNSMOS ses üretmez; puanlama modelidir. Aşağıdaki üç sütun iki bağımsız modelin üç çalışma yolunu gösterir: VyvoUp doğal 24 kHz giriş, aynı VyvoUp checkpoint'i ile eş-bilgi 16 kHz yolu ve FlashSR 16 kHz yolu. Eş-bilgi VyvoUp ile FlashSR aynı 16 kHz WAV'dan başlar.

Yayımlanan PCM16 dosyalar üzerindeki ortalamalar:

Çıkış SIG ↑ BAK ↑ OVRL ↑ P.808 ↑ 12–24 kHz LSD ↓
VyvoUp — doğal 24 kHz 3.5624 4.0752 3.2724 3.5796 8.3094
VyvoUp — aynı 16 kHz bilgi 3.5622 4.0742 3.2717 3.5779 9.7898
FlashSR — 16 kHz 3.4984 4.0020 3.1655 3.5756 11.4514

Karşılaştırma demosunu açın veya tam girişleri, hedefleri, 60 oynatıcıyı ve dosya bazlı ölçümleri bu model deposunda açın.

# VyvoUp doğal VyvoUp aynı bilgi FlashSR
01
02
03
04
05
06
07
08
09
10

Doğal/eş-bilgi VyvoUp, LSD'de sırasıyla 10/10 ve 9/10; RMS-hizalı DNSMOS OVRL'de iki yol da 8/10 örnekte FlashSR'den daha iyi çıktı. 70 WAV'ın tamamı mono PCM16, hedefle tam hizalı, benzersiz SHA-256'ya sahip ve clipping/saturation içermiyor. metadata.json bütün dosya hash'lerini, giriş sözleşmelerini ve örnek bazlı sonuçları içerir.

RTX 5070 Ti benchmark

One second of audio, BF16, batch one, 20 warmups and 200 measured iterations:

Path Mean p95 RTF Peak allocated
Eager 1.913 ms 1.943 ms 0.001913 26.0 MB
torch.compile 1.086 ms 1.205 ms 0.001086 14.5 MB

Compilation took 2.57 seconds and is excluded from measured latency. See benchmarks/ for environment and boundary details.

Example outputs

Examples follow a fixed policy: 10th, 50th, and 90th percentiles of target high-band energy, plus the hardest candidate LSD. Each directory contains the 24 kHz input, sinc baseline, model output, and 48 kHz target.

Selection Input Baseline VyvoUp v2.0 Target
10th percentile input sinc output target
Median input sinc output target
90th percentile input sinc output target
Hardest LSD input sinc output target

examples/metadata.json binds every WAV to its SHA-256 and evaluation report.

Inference

Install this VyvoUp source checkout and huggingface_hub, then download both the checkpoint and its mandatory checksum sidecar:

import numpy as np
import soundfile as sf
import torch
from huggingface_hub import hf_hub_download

from audio_upscaler.models import load_efficient_hbr

repo_id = "kadirnar/VyvoUp"
checkpoint = hf_hub_download(repo_id, "model.pt")
hf_hub_download(repo_id, "model.json")

audio, sample_rate = sf.read("speech-24khz.wav", dtype="float32")
if sample_rate != 24_000 or audio.ndim != 1:
    raise ValueError("Input must be mono 24 kHz audio")

model = load_efficient_hbr(checkpoint, device="cuda")
inputs = torch.from_numpy(np.asarray(audio)).view(1, 1, -1).cuda()
with torch.inference_mode(), torch.autocast("cuda", dtype=torch.bfloat16):
    output = model.forward_aligned(inputs)
sf.write("speech-48khz.wav", output.float().cpu().numpy()[0, 0], 48_000)

forward_aligned compensates FIR delay and returns exactly 2 * input_samples. For bounded-state deployment, use stream and flush.

CPU tabanlı demoda kullanılan dinamik-uzunluk TorchScript dosyası model.torchscript.pt olarak ayrıca yayımlanır. Dosya kimliği ve giriş/çıkış sözleşmesi model.torchscript.json içindedir.

Canlı demo, sunucu veya GPU kullanmadan model.onnx dosyasını ziyaretçinin tarayıcı CPU'sunda çalıştırır. ONNX grafiği dinamik ses uzunluğunu destekler; üç farklı uzunlukta ONNX Runtime CPU doğrulaması, PyTorch'a karşı en fazla 2.24e-7 mutlak hata ve tam 2:1 çıktı uzunluğu verdi. Ayrıntılar model.onnx.json içindedir.

Intended use, limits, and license

  • Research use for 24→48 kHz mono speech bandwidth extension.
  • English real-speech training is limited to VCTK. Other languages, music, environmental audio, stereo, telephony, and arbitrary rates are out of scope.
  • Generated high frequencies are plausible estimates, not recovered evidence; do not use them for forensic claims.
  • Objective metrics do not replace blinded listening or deployment-domain safety testing. No subjective-quality superiority claim is made.
  • Irodori-generated data inherits upstream ethical-use restrictions; do not use this work for impersonation, deception, misinformation, or rights violations.

Weights and bundled VCTK-derived listening examples are released under CC BY 4.0 to preserve VCTK attribution. Attribute CSTR VCTK Corpus 0.92 and VyvoUp when redistributing them. The source-code repository did not declare a software license at release time; this weight license grants no additional source-code rights.

VCTK citation: Christophe Veaux, Junichi Yamagishi, and Kirsten MacDonald, CSTR VCTK Corpus: English Multi-speaker Corpus for CSTR Voice Cloning Toolkit, version 0.92 (2019), DOI: 10.7488/ds/2645.

Downloads last month
-
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Space using kadirnar/VyvoUp 1

Paper for kadirnar/VyvoUp