DatasetWaxal

← Back to examples

How do the models score overall?

Two industry-standard measures, plus latency, across every model in the benchmark. Numbers come from the full evaluation set (1749 Shona clips from Waxal Shona).

< 15% — Excellent 15–30% — Good 30–50% — Rough > 50% — Struggling
� Model Size vs Performance

How does model size relate to WER?

� Word Error Rate (WER)

Percentage of words the model got wrong. Lower is better.

ModelWERNErrors
facebook/omniASR-LLM-7BOpen40.1%17490
aadel4/omniASR-CTC-1B-v2Open42.5%17490
facebook/omniASR-LLM-1BOpen43.4%17490
facebook/omniASR-CTC-7BOpen43.5%17490
facebook/omniASR-CTC-3BOpen44.7%17490
facebook/omniASR-CTC-1BOpen45.7%17490
facebook/omniASR-LLM-3BOpen46.2%17490
facebook/omniASR-LLM-300MOpen46.7%17490
aadel4/omniASR-CTC-300M-v2Open50.1%17490
facebook/omniASR-CTC-300MOpen57.7%17490
openai/whisper-large-v3Open114.6%17490
openai/whisper-smallOpen156.7%17490
Average across models61.0%
Best: facebook/omniASR-LLM-7B (40.1%)
🔤 Character Error Rate (CER)

Percentage of individual characters wrong. Lower is better.

ModelCERNErrors
facebook/omniASR-LLM-7BOpen8.9%17490
aadel4/omniASR-CTC-1B-v2Open9.2%17490
facebook/omniASR-CTC-7BOpen9.6%17490
facebook/omniASR-CTC-3BOpen9.8%17490
facebook/omniASR-LLM-1BOpen9.9%17490
facebook/omniASR-CTC-1BOpen10.0%17490
facebook/omniASR-LLM-3BOpen10.7%17490
aadel4/omniASR-CTC-300M-v2Open10.8%17490
facebook/omniASR-LLM-300MOpen10.9%17490
facebook/omniASR-CTC-300MOpen12.4%17490
openai/whisper-large-v3Open34.8%17490
openai/whisper-smallOpen65.1%17490
Average across models16.8%
Best: facebook/omniASR-LLM-7B (8.9%)
⚡ Latency

Average time it takes per sample.

ModelAvg latency (per sample)
facebook/omniASR-CTC-300MOpen0.02s
aadel4/omniASR-CTC-300M-v2Open0.03s
facebook/omniASR-CTC-1BOpen0.04s
facebook/omniASR-CTC-3BOpen0.08s
aadel4/omniASR-CTC-1B-v2Open0.09s
facebook/omniASR-CTC-7BOpen0.15s
openai/whisper-smallOpen0.30s
openai/whisper-large-v3Open0.41s
facebook/omniASR-LLM-300MOpen1.01s
facebook/omniASR-LLM-1BOpen1.03s
facebook/omniASR-LLM-3BOpen1.08s
facebook/omniASR-LLM-7BOpen1.13s
What do these numbers mean?

WER (Word Error Rate): out of every 100 spoken words, how many did the model get wrong?

CER (Character Error Rate): same idea but at the character level.

Lower is always better. 0% would mean a perfect transcription.

← Back to example gallery