af_heart_arm_tts / SAMPLES.md
tarunchy's picture
Upload 4 files
5620ed1 verified
|
Raw
History Blame Contribute Delete
1.21 kB

Samples

Six clips from samples/, synthesized by this model on the Arm CPU alone (ONNX Runtime, 9 threads) on a DGX Spark GB10. No GPU involved, no cherry picking — these are the first six sentences of the held-out evaluation set in order, the same set the WER/CER numbers in EVAL.md come from.

file text
samples/sample_00.wav The weather changed suddenly this afternoon.
samples/sample_01.wav Could you please pass me that book?
samples/sample_02.wav Wow, I did not expect to see you here!
samples/sample_03.wav Our meeting has been moved to three o'clock.
samples/sample_04.wav She walked quietly across the empty room.
samples/sample_05.wav How many people are coming to dinner tonight?

24 kHz mono, 16-bit PCM.

Listen before trusting any metric. The word and character error rates in EVAL.md measure whether an ASR model can recover the words. They say nothing about naturalness, pacing or timbre. This model is a 3.76M-parameter vocoder distilled from 3 hours of audio, and it sounds like it — the point is that it is 4.8× smaller and 10× faster on an Arm CPU than the teacher, not that it matches it.