An audit of Hindi and Hinglish content moderation classifiers on three major platforms
Dr. Rohit Malhotra, Dr. Priya Venkataraman, Aisha Khan
Journal of AI Ethics & SocietyVolume 8, issue 12026pp. 22–5064 citations
Content moderation classifiers are evaluated overwhelmingly on English. We construct a 24,000-item benchmark of Hindi, Romanised Hindi and Hinglish social media text, annotated by nine trained annotators against a published codebook with substantial inter-annotator agreement, and audit the public moderation endpoints of three major platforms against it. All three perform far worse on Romanised Hindi than on Devanagari Hindi, and worst of all on code-mixed Hinglish, which is the register most actually used. False negative rates on caste-based slurs reach 71 per cent. We release the benchmark and the annotation codebook.