Twenty-six research and development projects have received LINGUA Africa grants from the Masakhane African Languages Hub, in what represents one of the most targeted funding pushes yet to close the gap between African languages and functional artificial intelligence systems, according to Tech In Africa.

Masakhane, a grassroots research organisation that has built its reputation on community-driven natural language processing for African languages, is channelling the LINGUA Africa programme specifically toward languages that global AI developers have largely bypassed. Of the roughly 2,000 languages spoken across Africa, only a handful — Swahili, Amharic, Hausa, Yoruba — appear in any meaningful volume in the training data that powers large language models from OpenAI, Google, and Meta.

The 26 grants signal an attempt to systematically expand that coverage. While the precise grant amounts per project have not been disclosed in early reporting, the number of awards — 26 simultaneous grants — is notable for a continent where language-focused AI funding has historically arrived in single-digit tranches, if at all. Each grantee is expected to produce datasets, tools, or applications that can feed back into the broader open-source African language AI ecosystem that Masakhane has been cultivating since its founding in 2018.

The practical stakes are high. AI-powered services — from agricultural advisories and health chatbots to financial literacy tools and customer support — are spreading rapidly across sub-Saharan Africa. But the overwhelming majority of these systems default to English, French, or Portuguese, effectively locking out the hundreds of millions of Africans who communicate primarily in indigenous languages. A smallholder farmer in rural Tanzania seeking crop disease guidance, or a microfinance borrower in northern Nigeria trying to understand a loan term, encounters a system built for someone else's language. The LINGUA Africa grants are, at their core, an attempt to fix that infrastructure deficit at the data layer — where the problem originates.

Masakhane's approach has always been to decentralise the work, recruiting researchers, students, and language speakers from across the continent rather than concentrating development in a single institution or country. That philosophy appears embedded in the 26-grant structure, which by design spreads resources and accountability across multiple teams and, presumably, multiple language communities. Distributed grant-making of this kind also hedges against the risk that a single flagship project stalls or loses momentum.

For operators building AI products in Africa, the LINGUA Africa output — if grantees deliver — should reduce one of the more stubborn cost inputs: the price of acquiring or commissioning African-language training data. Right now, companies that want a functional Twi or Tigrinya interface often must fund bespoke data collection themselves, a cost that is prohibitive for early-stage startups. Open datasets produced under Masakhane's open-source norms could change that calculation materially.

Investors backing African AI infrastructure should watch which of the 26 projects produce datasets at scale and which surface as candidates for commercialisation or further grant funding. Masakhane's track record — it has contributed to benchmark datasets such as MasakhaNER and MasakhaPOS, used by researchers worldwide — gives it credibility as a steward of this work, but the translation from academic dataset to product-grade AI capability still requires additional capital and engineering capacity that most grantees will not have internally.

Why it matters: With 26 simultaneous grants, Masakhane is making the most concentrated single push in the LINGUA Africa programme's history to populate the one layer — training data — without which no amount of downstream AI investment in Africa can produce tools that actually serve the continent's majority-language speakers. For any operator or investor building AI products for African consumers, the datasets these grantees produce in the next 12 to 24 months could be among the most valuable open-source resources to emerge from the continent.