arxiv:2605.20712
๐ In a Training Loop
Kavya Manohar
ยท
AI & ML interests
Speech Recognition, Low Resource Languages, Malayalam
Recent Activity
new activity 4 days ago
adalat-ai/fleurs-ro:v2.0: add Telugu config (466 rows, self-hosted Gemma 4 31B curation) + docs updated a dataset 4 days ago
adalat-ai/fleurs-ro posted an update about 1 month ago
Some bugs teach you more than they cost you. It's a small bug with a big lesson I keep running into: speech tools are built and benchmarked on English, so their limits are quietly calibrated for English. You only find the edges when you work in the languages they weren't tested on.
Learnt things the hard way, wrote it down so you don't have to.
https://huggingface.co/blog/adalat-ai/whisper-token-limit