BharatGen CEO On Why Models Alone Can’t Help India Gain In The AI Race
BharatGen, a non-profit consortium backed by the IndiaAI Mission, is building multilingual and multimodal AI models for all 22 scheduled Indian languages, focusing on research, data collection, and ecosystem development.
Intelligence analysis by Gemini 2.5 Flash

BharatGen, anchored at IIT Bombay with significant government funding, aims to define India's AI future by developing foundational models across diverse Indian languages. Unlike for-profit ventures, it prioritizes a robust research ecosystem and extensive data collection to address India's unique linguistic and cultural needs.
Imagine India wants to build its own super-smart computer brains (AI) that can understand all its many languages, not just English. BharatGen is like a special government-funded school and library that's teaching these computer brains and collecting all the unique stories and knowledge from India. They're making sure India has its own smart computers that truly understand Indian culture and languages, helping everyone from farmers to shopkeepers.
Analysis
BharatGen's Ecosystem-First Approach
BharatGen distinguishes itself from typical AI startups by operating as a non-profit consortium, anchored at IIT Bombay and backed by the Department of Science and Technology. With substantial funding of ₹988.6 Cr from the IndiaAI Mission, its mission extends beyond merely releasing AI models; it aims to cultivate the entire AI ecosystem in India. CEO Rishi Bal emphasizes that models and applications are just the visible tip of the iceberg, with deep research and talent development forming the much larger, invisible foundation. This approach is critical for India to become a global player in AI, fostering a robust research environment that can generate the next generation of AI innovators and technologies.
Bridging India's Linguistic Data Divide
A significant challenge for AI development in India is the massive scarcity of digital data for its numerous languages, especially beyond Hindi. BharatGen is tackling this through its "Bharat Data Sagar" initiative, a dataset repository designed to capture India's linguistic and cultural diversity at scale. The organization employs teams on the ground to engage with publishers, digitisation organizations, and heritage preservation groups, converting out-of-print books and old newspapers into native digital text. This labor-intensive process, while expensive, is crucial for training models that truly understand and serve India's diverse linguistic landscape, offering a quid-pro-quo arrangement where enterprises gain model access in exchange for data.
Navigating Talent and Compute Challenges
Despite its unique non-profit model and government backing, BharatGen faces considerable hurdles, primarily in talent acquisition and retention. Building large language models (LLMs) is fundamentally a "people business," and attracting and keeping top-tier AI talent is challenging, particularly for a non-profit entity competing with well-funded for-profit ventures. Additionally, the capital-intensive nature of training foundational models constitutes the bulk of BharatGen's expenses, with roughly 80% allocated to training and 20% to data collection. Despite these challenges, BharatGen continues to develop purpose-driven AI tools like Krishi Sathi for farmers and e-VikrAI for Indian sellers, demonstrating its commitment to real-world impact and continuous model progression.
Key points
- BharatGen is a non-profit consortium, anchored at IIT Bombay and backed by the IndiaAI Mission, with ₹988.6 Cr in funding.
- It is building multilingual and multimodal AI models across all 22 scheduled Indian languages, with over 1 lakh downloads of its open-source releases.
- The organization prioritizes a robust research ecosystem, having published over 30 papers and placed 100+ interns in 18 months.
- BharatGen is creating the "Bharat Data Sagar," a repository for India's linguistic and cultural data, using unique methods like digitizing out-of-print materials.
- Key challenges include talent acquisition and retention, and the high capital cost of training models (80% of expenses).
BharatGen's efforts could significantly advance India's position in the global AI landscape, fostering a self-reliant ecosystem capable of developing AI solutions tailored to its unique linguistic and cultural diversity. The focus on research and talent development promises a sustainable pipeline of innovation and skilled professionals.
The organization faces substantial challenges in attracting and retaining top AI talent, especially as a non-profit competing with private firms. The immense capital required for model training and the labor-intensive data collection for underserved languages could also strain resources and slow progress.


