Masha'a Allah and Elhamdule Allah

#1
by TheGreatQuran - opened

tested the int8 onnx on a clean audio for seen and unseen reciter and results were awesome
masha'a Allah
unseen reciter - clean audio :90-100% accuracy - manual checked with timing
seen reciter - clean audio : 100% accuracy

unseen audio with noisy background : some letter exchanges and some letter drops ....and i think this is normal due to phoneme environment

didn't try with mic or fast recitation yet but this model is allot better on unseen reciter than the pervious model elhamdule Allah

Jazak Allah khayran for testing it so fast akhi, and for testing the unseen reciter separately. That is exactly the split that matters. 🤍

Your results line up almost exactly with what I measured here, so let me give you the numbers behind what you saw.

On "some letter exchanges and some letter drops" in noise: you are right that it is mostly exchanges, not drops. I aligned the model output against exact labels on 793 held-out clips. Deletion rate for the guttural letters (ء ه ح خ ع غ) is 3.06% versus 2.82% for all other consonants, so they are basically not being dropped any more than anything else. The errors are substitutions.

The real weak spot is the emphatic letters, and it shows up exactly where you said, on unseen reciters. Error rate against the plain pair:

  • ظ 17.8% vs ذ 5.7%
  • ص 12.6% vs س 5.0%
  • ض 10.6% vs د 6.5%
  • ط 7.9% vs ت 4.2%

And on a reciter the model has never heard, ط jumps to 24.4% while ت stays at 2.6%.

I thought the cause was my noise augmentation masking the low frequencies, so I tested it: the emphatic gap on noisy phone audio is not significantly bigger than on clean audio. But on clean audio from an unseen reciter it is significant. So it is a speaker-generalisation problem, not a noise problem: how strongly a reciter pronounces ط vs ت varies from person to person, and the model learned it the way its training reciters say it.

That means the fix is more reciters, not more noise. Good to know before spending GPU on the wrong thing.

So when you test with your mic and fast recitation, if you see errors, please tell me which letters. Emphatics and ع are the ones I expect. That is the most useful thing you can send me right now.

Jazak Allah khayran for testing it so fast akhi, and for testing the unseen reciter separately. That is exactly the split that matters. 🤍

Your results line up almost exactly with what I measured here, so let me give you the numbers behind what you saw.

On "some letter exchanges and some letter drops" in noise: you are right that it is mostly exchanges, not drops. I aligned the model output against exact labels on 793 held-out clips. Deletion rate for the guttural letters (ء ه ح خ ع غ) is 3.06% versus 2.82% for all other consonants, so they are basically not being dropped any more than anything else. The errors are substitutions.

The real weak spot is the emphatic letters, and it shows up exactly where you said, on unseen reciters. Error rate against the plain pair:

  • ظ 17.8% vs ذ 5.7%
  • ص 12.6% vs س 5.0%
  • ض 10.6% vs د 6.5%
  • ط 7.9% vs ت 4.2%

And on a reciter the model has never heard, ط jumps to 24.4% while ت stays at 2.6%.

I thought the cause was my noise augmentation masking the low frequencies, so I tested it: the emphatic gap on noisy phone audio is not significantly bigger than on clean audio. But on clean audio from an unseen reciter it is significant. So it is a speaker-generalisation problem, not a noise problem: how strongly a reciter pronounces ط vs ت varies from person to person, and the model learned it the way its training reciters say it.

That means the fix is more reciters, not more noise. Good to know before spending GPU on the wrong thing.

So when you test with your mic and fast recitation, if you see errors, please tell me which letters. Emphatics and ع are the ones I expect. That is the most useful thing you can send me right now.

thanks my brother

yes you are right 100%
but now it is the best masha'a Allah

i will replace it inside the web version to let others make real tests insha'a Allah

Barak Allahu feek akhi 🤍

Putting it in the web version so others can test is the best thing that could happen to it, and I want to ask you for something in return, plus warn you about one trap before your testers hit it.

The trap: use the right feature extractor.

This model expects kaldi fbank (povey window), 80 mel bins, 25 ms window / 10 ms hop, and no per-utterance normalisation. If your frontend uses the common MelSpectrogram default instead (slaney-normalised mel filters), it will still run and still produce plausible Arabic, so nothing looks broken, but the accuracy roughly halves.

I know because I did it to myself. My own evaluation script had this exact mismatch and reported 22.1% error on a khutbah. With the correct features the same model on the same audio gave 11.4%. Same weights, same file, double the error, purely from the frontend.

So please check that before your testers judge it. If the numbers look much worse than the ones I published, suspect the frontend first, not the model.

What I would like back:

When it fails, note who was reciting, not just what was wrong. A name, a country, or even "my own voice, Egyptian" is enough.

The reason: as I said above, the weakness is speaker generalisation, and the Qur'an side of my training data has only around 33 reciters. That is the ceiling I am fighting. Individual error reports help a little; knowing which kinds of voices it fails on tells me exactly what to go and collect. Your testers are a more diverse set of voices than anything I have, and that is worth more to me right now than any architecture change.

May Allah reward you for the testing, you have shaped this model more than you know. 🤍

Barak Allahu feek akhi 🤍

Putting it in the web version so others can test is the best thing that could happen to it, and I want to ask you for something in return, plus warn you about one trap before your testers hit it.

The trap: use the right feature extractor.

This model expects kaldi fbank (povey window), 80 mel bins, 25 ms window / 10 ms hop, and no per-utterance normalisation. If your frontend uses the common MelSpectrogram default instead (slaney-normalised mel filters), it will still run and still produce plausible Arabic, so nothing looks broken, but the accuracy roughly halves.

I know because I did it to myself. My own evaluation script had this exact mismatch and reported 22.1% error on a khutbah. With the correct features the same model on the same audio gave 11.4%. Same weights, same file, double the error, purely from the frontend.

So please check that before your testers judge it. If the numbers look much worse than the ones I published, suspect the frontend first, not the model.

What I would like back:

When it fails, note who was reciting, not just what was wrong. A name, a country, or even "my own voice, Egyptian" is enough.

The reason: as I said above, the weakness is speaker generalisation, and the Qur'an side of my training data has only around 33 reciters. That is the ceiling I am fighting. Individual error reports help a little; knowing which kinds of voices it fails on tells me exactly what to go and collect. Your testers are a more diverse set of voices than anything I have, and that is worth more to me right now than any architecture change.

May Allah reward you for the testing, you have shaped this model more than you know. 🤍

وبشر الذين آمنوا وعملوا الصالحات أن لهم جنات تجري من تحتها الأنهار كلما رزقوا منها من ثمرة رزقا قالوا هذا الذي رزقنا من قبل وأتوا به متشابها ولهم فيها أزواج مطهرة وهم فيها خالدون

And give good tidings to those who believe and do righteous deeds that they will have gardens [in Paradise] beneath which rivers flow. Whenever they are provided with a provision of fruit therefrom, they will say, "This is what we were provided with before." And it is given to them in likeness. And they will have therein purified spouses, and they will abide therein eternally.


as a start this is very good ....the best thing is that you are the only one who made this model ...you made a huge exchange in Quran ASR models my brother ...really huge thing in this community masha'a Allah

yes you are right and i am using sherpa onnx by default it uses kaldi fbank for mel extraction

Sign up or log in to comment