top of page

How Karya Is Building Digital Work Infrastructure for Rural India Through AI Data Labeling

yanabijoor
11 minutes ago
5 min read

What is the problem?

All prominent AI algorithms worldwide run on human-labeled data. Someone has to collect voice recordings, transcribe them, annotate images by drawing bounding boxes around objects, and label content across multiple languages. This process, known as data annotation, employs millions of people across the developing world at exploitative rates. The average hourly wage for a data annotator in India is about 40 cents. 


By 2030, the data annotation industry in India is expected to employ nearly one million people. However, their salaries are growing nowhere near as fast as the AI industry's revenue. 


Meanwhile, the AI models these workers help train are getting worse at serving the people who built them. Most AI systems are trained on English text scraped from the web. India has roughly 700 dialects.


The voice recognition algorithm that works perfectly in American English cannot understand a farmer speaking in Odia, Kannada, or Marathi. Hence, rural India is excluded from the AI economy, both as underpaid labor and as underserved consumers.

Woman doing mobile data entry for AI models
Young woman who left her garment factory job due to health issues now works for Karya (Time Magazine)

What is the solution?

Karya hires rural Indians to perform microtasks for AI, such as voice recording, transcription, image annotation, and video labeling, using an app designed for people who don’t read and that works offline. The company pays workers 20 times the Indian minimum wage (around $5 per hour), which is well above the industry average. 


The company works exclusively with Indian languages, so workers create datasets in their native languages, improving the AI experience for people like them. Additionally, workers receive a share of all reselling profits from their data, which is unique in the annotation industry. Every worker who uses the app has a lifetime income cap of around $1,500 (about the Indian average annual income), after which Karya moves on to the next person in the queue.

data entry on phone
Young man using Karya after losing his right hand at age 8 (Time Magazine)

What is the business model?

Karya operates as a mission-driven social enterprise, sometimes categorized as a nonprofit and sometimes as a social impact firm because its business model blends elements of both. It generates funding by selling AI training datasets to major technology companies.


Karya covers its costs and shares the remaining revenue with workers. Key partners include Microsoft, Google, and the Bill and Melinda Gates Foundation. These grants help fund operations alongside commercial revenue.


In December 2024, Google.org granted Karya $1 million to design more accurate AI models for local languages and develop a multilingual chatbot for rural users. The Gates Foundation also helps fund research on bias mitigation using data Karya collects.


How is it structured and funded?

Manu Chopra, Vivek Seshadri, and Safiya Husain founded Karya in Bengaluru in 2021. Chopra was born in Shakur Basti, a slum in West Delhi, and earned a scholarship to an elite school, where he was bullied for “smelling of poverty.” He studied computer science at Stanford University, rejected Stanford's “make a billion dollars” mentality, and started Karya in India. 


Vivek Seshadri previously worked at Microsoft Research India. Safiya Husain previously worked in civil society organizations. Safal Kaul later joined Karya. 


Karya's investors include the Bill and Melinda Gates Foundation, Google.org, Rockefeller Foundation, ACT Grants, and others.


startup CEO from India
Manu Chopra, Co-Founder & CEO, Karya

Why is it innovative?

Karya rethinks two problems with AI right now. First, who benefits from the AI data work. Most annotation companies capture the margin between what they charge clients and what they pay workers. Karya inverts this ratio–its workers get the lion's share, and the company funds its operations from the rest.

Two, which languages are AI built to serve. Data-related companies worldwide harvest data either in English or in languages spoken in large countries because tech companies demand it. In contrast, Karya focuses on Indian minority languages, so workers produce data used to teach models those languages. It looks more like building digital infrastructure for their community rather than working in an exploitative way. 


The lifetime earnings limit adds the third dimension to Karya's approach: rather than creating a loyal group of contractors, Karya wants to distribute money among as many rural families as possible. According to Chopra, it is the quickest way to lift millions of people out of poverty.

mobile data entry
Karya workers all in their 20's and 30's using Karya to supplement their income (Time Magazine)

What is the impact?

Based on Karya's own reporting and press coverage:

  • Scale of Workers: Over 200,000 rural and low-income workers have been onboarded to the platform, with coverage across all 28 Indian states.

  • Total Wages Distributed: The platform has directly paid out over $4.42 million USD (~₹370 million) to its workforce, driving a 6.8x year-over-year increase in distributed wages.

  • Task Volume: Workers have completed more than 42 million paid AI microtasks (such as voice recordings and image annotations) via the mobile app.

  • Wage Premium & Income Boost: Workers continue to receive dignified pay at roughly 20 times the Indian minimum wage. This has resulted in an average 25.8% annual income increase for workers overall, with a 35% increase for female workers.

  • Language Breadth: The dataset focus has expanded dramatically to cover over 115 languages and local dialects, intentionally gathering training data for historically underrepresented communities.

  • Evolving 2030 Goal: The mission has shifted from general headcount to direct economic outcomes. The updated target is to distribute $100 million in wages across India and facilitate $1 billion in wages globally by 2030.

  • Global Footprint: While rooted in India, the technology has officially transitioned to a Platform-as-a-Service (PaaS) model and has been deployed across 7 countries, including pilots in Kenya and Ethiopia.

mobile data entry
Young woman using Karya to support her sick father (Time Magazine)

What needs to improve?

Karya faces three major obstacles. The first is the tension between scale and mission. The core mission of income distribution requires Karya to set an upper bound on each worker's lifetime income, but that bound also limits Karya, too. To reach its objective of 100 million people by 2030, Karya would need to scale its user base 1000-fold, which would require massive investment in recruitment, training, and quality assurance.


The second problem is Karya's heavy dependence on a handful of customers. Microsoft, Google, and the Gates Foundation are clients, but the AI industry is consolidating and becoming even more dependent on a few big players. If any of them changed their strategy and cut data purchases, Karya would feel the financial impact.


The third problem is the long-term sustainability of the business model. AI firms are increasingly developing synthetic data pipelines in which one AI generates training data for another, reducing the need for human labelers. If this trend accelerates faster than Karya reaches 100 million, its business model might stall.


Sources:



Comments


Join 9,300 Subscribers Today

Thanks for submitting!

Inventaid
bottom of page