Small Models, Big Reach

Small Models, Big Reach
Task-based precision that makes national-scale AI affordable
Every AI model reflects an architecture decision, and that decision has a cost. At Wadhwani AI, we build Small AI - models trained for a single task - instead of general-purpose language models adapted to fit one. It's the only architecture that gets us two bets we’re funding: technology good enough for a government partner to trust and a cost structure low enough that philanthropic dollars can fund it at national scale.
Here's how that architecture choice actually plays out.
Small AI has become its own line of research, distinct from the scaling race that dominates most AI headlines. Researchers at NVIDIA and Georgia Tech argued last year that small, task-specific models are the more suitable and more economical choice for narrow, repetitive tasks, not just a compromise on the way to something bigger. Model builders like Google DeepMind and Alibaba have started releasing models purpose-built for on-device inference instead of data-center deployment. We've been building this way since 2018.
The architecture difference comes down to task distribution. A large language model is trained across a broad distribution of tasks, which is what gives it general capability, and also what drives its compute footprint. Every inference call carries the overhead of everything else the model knows how to do, whether that capacity gets used or not. A small AI model is fine-tuned on a single task distribution. For example, classify a cough recording or score a child's reading fluency. There's no general-purpose overhead to carry, so the compute cost per inference drops sharply.
That lower compute cost is what makes offline deployment possible. A model with a narrow task distribution and a small parameter footprint can run inference locally on a basic smartphone, no data connection required. And because it's small enough to sit inside the tools people already carry, deployment doesn't require new hardware or a change in how a health worker or teacher does their job. Our reading assessment tool runs inside a teacher's existing classroom app. Our telehealth model runs inside the national platform a health worker already uses. The AI slots into an existing workflow instead of asking the system to build one around it, which is a large part of why adoption holds at scale instead of stalling after a pilot.
The instinct might be to hear "small" and assume "less capable." But we've learned the opposite is true for a fixed task like screening for TB or improving childhood literacy in the Global South. A model trained only on cough recordings, or only on children's oral reading, doesn't have to allocate any of its capacity to unrelated tasks, so all of it goes toward the problem in front of it. That's the core claim behind small AI research: precision on a narrow, well-defined task, not general-purpose reach, is what a lot of real-world deployment actually needs.
The clearest proof is in the numbers. Wadhwani AI’s reading assessment tool, built by Wadhwani AI India, runs at 94% precision in Hindi and 95% in Gujarati, and costs less than one penny per assessment. That per-unit economics is what makes the tool a national-scale, not a pilot-scale, solution. At a fraction of a cent per student, the marginal cost of assessing the next million children approaches zero, which is the only cost curve that lets a philanthropically funded program reach population scale without needing a proportional increase in budget. It's also a number a donor can build a defensible cost-per-child case around, something a per-token LLM pricing model can't offer at the same precision. And because the tool replaces roughly ten minutes of one-on-one teacher time with about one minute of automated scoring, the cost savings show up twice: once in the compute bill, and once in the teacher-hours it frees up for actual instruction.
That's the case for Small AI: not a smaller version of what the rest of the field is building, but a different bet on what AI is for. The next test isn't whether the model works in one state or one sector. Instead, it’'s whether the same logic, small enough to fit real infrastructure, cheap enough to fund at scale, holds as we take on new problems and new languages where the constraint has never been what AI could do, only whether it was built to actually reach the people who need it.
News
Insights from the frontline
The strategy, systems, and hard-won lessons behind AI that's built and deployed inside real government programs at population scale
