The people behind the label: working conditions in India's data annotation industry
- Prof. Daniel Okonkwo · University of Cape TownORCID 0000-0002-8834-1129
- Aisha Khan · IIIT Hyderabad
- Dr. Hannah Weiss · Universität WienORCID 0000-0001-3345-7781
- Journal
- Journal of AI Ethics & Society
- Volume
- 7, issue 4
- Pages
- 401–428
- Licence
- CC BY 4.0
Abstract
Machine learning depends on annotation labour that its literature rarely describes. Drawing on 68 interviews with annotators and eleven with managers across nine firms in Bengaluru, Kochi and Bhubaneswar, we document a labour process organised around quota, surveillance and a quality regime that transfers the cost of ambiguous data to the worker. We argue that annotation guidelines function as an unacknowledged site of normative decision-making: annotators routinely resolve genuine moral ambiguity under time pressure, and their resolutions are laundered into training data as ground truth.
Keywords
Introduction
Every supervised model is an argument about what its labels mean, and every label was applied by someone. The conditions under which that someone worked are treated in the machine learning literature as a procurement detail.
This paper takes those conditions as its object. We report interviews with 68 annotators and 11 managers across nine firms, conducted between January and October 2025, and describe a labour process whose features are not incidental to the data it produces.
The quota and the queue
Annotators in our sample worked to hourly item quotas ranging from 180 to 900 depending on task complexity, with pay structured as a low base plus a quota-contingent component that most workers described as the part they actually lived on.
Quality was enforced through gold-standard items seeded into the queue at rates between two and six per cent. Falling below a quality threshold cost the quota bonus for the whole shift. The combination — a rate that requires speed and an audit that punishes error — was described by workers in near-identical terms across all nine firms, and it produces a specific behaviour we heard about repeatedly: when an item is genuinely ambiguous, annotators do not deliberate. They apply whatever rule they believe the gold standard encodes.
Ambiguity as unpaid normative labour
The tasks that generated the most reported difficulty were not the technically hard ones but the morally contested ones: whether a political post constitutes incitement, whether a slur is being used in reclamation, whether an image of a protest is violent content.
These are questions on which reasonable people disagree, and they were being settled at a rate of one every four to nine seconds by workers paid between ₹18,000 and ₹31,000 a month, with no channel to escalate an ambiguous item and an explicit instruction not to leave items unlabelled. The resolutions became ground truth. Every downstream fairness metric computed against that ground truth inherits decisions made under those conditions and reports them as facts about the world.
What follows
We resist the conclusion that annotation should simply be better paid, true though that is, because it leaves the epistemics untouched. A well-paid annotator resolving a genuinely contested normative question in four seconds still produces a label that a fairness audit will treat as ground truth.
Our recommendations are procedural: ambiguity channels that let a worker mark an item as contested without penalty, publication of annotation guidelines alongside datasets as a matter of course, and reporting of inter-annotator disagreement at the item level rather than as a single aggregate. Each is cheap. None is standard.
References
- 1.Gray, M. L., & Suri, S. (2019). Ghost Work. Houghton Mifflin Harcourt.
- 2.Miceli, M., Posada, J., & Yang, T. (2022). Studying up machine learning data. Proceedings of the ACM on HCI, 6(GROUP).
- 3.Roberts, S. T. (2019). Behind the Screen. Yale University Press.
- 4.Denton, E., et al. (2021). On the genealogy of machine learning datasets. Big Data & Society, 8(2).
Cite this article
Okonkwo, D., Khan, A., Weiss, H. (2025). The people behind the label: working conditions in India's data annotation industry. Journal of AI Ethics & Society, 7(4), 401–428. https://doi.org/10.53456/jaies.2025.7.4.401
Related in Computer Science
An audit of Hindi and Hinglish content moderation classifiers on three major platforms
Dr. Rohit Malhotra, Dr. Priya Venkataraman, Aisha Khan
Journal of AI Ethics & SocietyVolume 8, issue 12026pp. 22–5064 citations
Content moderation classifiers are evaluated overwhelmingly on English. We construct a 24,000-item benchmark of Hindi, Romanised Hindi and Hinglish social media text, annotated by nine trained annotators against a published codebook with substantial inter-annotator agreement, and audit the public moderation endpoints of three major platforms against it. All three perform far worse on Romanised Hindi than on Devanagari Hindi, and worst of all on code-mixed Hinglish, which is the register most actually used. False negative rates on caste-based slurs reach 71 per cent. We release the benchmark and the annotation codebook.