Skip to content

Peer-reviewed · Quarterly · Since 2019

Journal of AI Ethics & Society

The social, legal and normative dimensions of machine learning systems — fairness, accountability, governance and the labour that trains them.

Open accessComputer ScienceScopusDOAJCrossrefDBLPGoogle Scholar

Editor-in-Chief: Dr. Priya Venkataraman, IIIT Hyderabad

Articles

Research ArticleOpen access

An audit of Hindi and Hinglish content moderation classifiers on three major platforms

Dr. Rohit Malhotra, Dr. Priya Venkataraman, Aisha Khan

Volume 8, issue 12026pp. 22–5064 citations

Content moderation classifiers are evaluated overwhelmingly on English. We construct a 24,000-item benchmark of Hindi, Romanised Hindi and Hinglish social media text, annotated by nine trained annotators against a published codebook with substantial inter-annotator agreement, and audit the public moderation endpoints of three major platforms against it. All three perform far worse on Romanised Hindi than on Devanagari Hindi, and worst of all on code-mixed Hinglish, which is the register most actually used. False negative rates on caste-based slurs reach 71 per cent. We release the benchmark and the annotation codebook.

Research ArticleOpen access

The people behind the label: working conditions in India's data annotation industry

Prof. Daniel Okonkwo, Aisha Khan, Dr. Hannah Weiss

Volume 7, issue 42025pp. 401–42852 citations

Machine learning depends on annotation labour that its literature rarely describes. Drawing on 68 interviews with annotators and eleven with managers across nine firms in Bengaluru, Kochi and Bhubaneswar, we document a labour process organised around quota, surveillance and a quality regime that transfers the cost of ambiguous data to the worker. We argue that annotation guidelines function as an unacknowledged site of normative decision-making: annotators routinely resolve genuine moral ambiguity under time pressure, and their resolutions are laundered into training data as ground truth.