Voice Quests: Teaching ORBII to Hear India
Speech models are trained mostly on American and British voices. ORBII is building its own distress-word model on Indian ones, and you can help train it, entirely by choice.
Ask most speech recognition systems to understand a Hindi-accented “help” shouted on a Delhi street at night, and they struggle. Not because the technology is bad, but because of what it was trained on.
The open speech models the world runs on were built largely from American and British English, recorded in reasonable conditions by people speaking calmly. Almost none of that describes the moment ORBII needs to work.
The gap we are actually trying to close
ORBII currently uses Vosk for on-device speech recognition. It is genuinely good, it runs offline, and it is the reason your audio never has to leave your phone. But it comes with a real cost: the English model alone is 68 MB sitting inside the app, because it is a general-purpose model built to transcribe any sentence you might say.
ORBII does not need to transcribe any sentence. It needs to recognise a handful of words, spoken by Indian voices, under bad conditions, with near-perfect reliability.
That is a much narrower problem, and a narrower problem can be solved by a much smaller model. The one we are building, OrbiiKWSCNN, is a keyword-spotting network with about 72,000 parameters. Stored at 8-bit precision that is roughly 70.6 KB. A model that fits in the space of a small photograph, doing one job properly instead of a thousand jobs adequately.
Smaller matters for real reasons. It uses less battery, so people leave Voice SOS switched on. It responds faster. And it makes room for more Indian languages rather than forcing us to choose two.
But a model is only as good as what it heard while learning. Which is where you come in, if you want to be.
What Voice Quests are
Voice Quests are short missions inside ORBII. Say “help” in a quiet room. Say “bachao” the way you would actually shout it. Say “madad” while there is traffic outside your window.
Each mission is a different word, a different setting, or a different tone, because a model that has only heard calm indoor voices is exactly the model that fails on the street. There are streaks, XP, and badges, because “help us build a dataset” is a bad ask and “finish today’s quest” is a good one.
Every clip you send makes the model better at hearing people who sound like you.
The rules we set ourselves
Asking people to send us audio, on a privacy-first safety app, is not something to do casually. So the constraints came first:
It is off by default. Voice donation does not exist for you until you go and switch it on.
It never runs in the background. There is no ambient collection, no “while you are here anyway”, no clever capture during an SOS. You tap record, you say the word, you see exactly what you are sending, and you send it.
Adults only. A child’s voice is biometric data and cannot be collected on the basis of the child’s own consent, so we do not collect it. Anyone under 18 does not see this feature.
You can withdraw at any time, and your clips are deleted with your account or on request.
It goes nowhere near advertising. Clips live in private storage that no other user can read, are used only to train distress-word detection, and are never sold. There is no second purpose hiding behind the first.
Your consent is recorded properly. Turning it on writes a row. Turning it off writes a row too. Withdrawals are logged exactly like grants, because a consent record that only tracks yeses is not a record, it is marketing.
What this does not change
Normal Voice SOS still works the way it always has. Your phone listens for your distress word on the device, recognises it on the device, and that audio never leaves your phone. That is the foundation of the product and nothing about Voice Quests touched it.
If you never open Voice Quests, ORBII never receives a single second of your audio. That is the honest default, and most people will stay on it.
Why bother building our own model at all
Because “works well enough on Indian voices” is not a standard anyone should accept from a safety app.
If a woman shouts the word she was told would summon help, and the phone does not react because it was trained on somebody else’s accent, everything else we built is worth nothing in that moment. The rest of the app is downstream of the microphone being right.
Buying our way out of this is not an option. There is no off-the-shelf dataset of Indian women shouting for help, and there should not be one collected without consent. So it gets built the slow, honest way: by people who chose to help, one clip at a time.
If that is you, it is in the app under Voice Quests. If it is not, that is a completely reasonable answer and the app works exactly the same.
Read why we keep your voice on your phone, or see how Voice SOS works.
Frequently asked questions
Does ORBII upload my voice?
Not during normal use. Voice SOS detection runs entirely on your phone and that audio never leaves the device. The only audio that ever reaches us is a Voice Quest clip you deliberately record and send, which is off by default.
Can I stop donating and delete what I sent?
Yes. Turn it off any time in the app, and everything you donated is deleted along with your account, or on request to privacy@orbii.in.
Who can listen to donated clips?
Only ORBII, and only to improve distress-word detection. Clips sit in private storage that no other user can read. They are never sold, never used for advertising, and never used to identify you.