Undergraduate Research Mentorship
Status: Completed
A pair of undergraduate mentoring programs around one question: how well models hold up far from their training data. Eight teams took the language side, four students the cultural one, and four of the projects grew into published benchmarks.
Multilingual evaluation
A semester in 2023. Each of the eight teams took a different low-resource or non-English language, from Urdu and Bengali to Korean, Russian, and Indonesian, and a different capability: chain-of-thought reasoning, sentiment analysis, question answering, named-entity recognition, or standardized-exam performance. The recurring finding, language after language, was that models grow sharply weaker, and less culturally grounded, the further a language sits from their English-heavy training data. Two teams carried their work to published benchmarks:
- BEnQA: A Question Answering Benchmark for Bengali and English (Findings of ACL 2024)
- CLIcK: A Benchmark Dataset of Cultural and Linguistic Intelligence in Korean (LREC-COLING 2024)
Cultural evaluation
The second program moved from language to culture: four students evaluating whether text-to-image and multimodal models represent cultures they rarely saw in training. Their work became two publications:
- Diffusion Models Through a Global Lens: Are They Culturally Inclusive? (ACL 2025, Oral)
- When Tom Eats Kimchi: Evaluating Cultural Awareness of Multimodal Large Language Models in Cultural Mixture Contexts (C3NLP @ NAACL 2025, Outstanding Paper)