Insights into the Challenges and Opportunities of Large Multi-Modal Models for Blind and Low Vision Users: CLIP
PARIKSHA: A Scalable, Democratic, Transparent Evaluation Platform for Assessing Indic Large Language Models
Publication LLM Reasoning as Trajectories: Step-Specific Representation Geometry and Correctness Signals Lihao Sun, Hang Dong, Bo Qiao, Qingwei Lin, Dongmei Zhang, Saravan Rajmohan April 2026 arXiv | April 2026
Publication Human Values Matter: Investigating How Misalignment Shapes Collective Behaviors in LLM Agent Communities Xiangxu Zhang, Jiaming Wang, Qinlin Zhao, Hanze Guo, Linzhuo Li, Jing Yao, Xiao Zhou, Xiaoyuan Yi, Xing Xie April 2026 arXiv | April 2026
Publication The Tool Illusion: Rethinking Tool Use in Web Agents Renze Lou, Baolin Peng, Wenlin Yao, Qianhui Wu, Hao Cheng, Suman Nath, Wenpeng Yin, Jianfeng Gao April 2026 arXiv | April 2026
Publication Magic, Madness, Heaven, Sin: LLM Output Diversity is Everything, Everywhere, All at Once Harnoor Dhingra April 2026 arXiv | April 2026
Publication Not Another EHR: Reimagining Physician Information Needs with Generative AI Technology Ruican Zhong, Jicahen Li, Gary Hsieh, David W. McDonald, Selin S. Everett, Alyssa Unell, Jonathan M. Carlson, Katie Claveau, Noel Codella, Khalil Malik, Scott Mackie, Eduardo Olvera, Scott Saponas, Eric Horvitz, David Rhew, James Weinstein, Jacob Gross, Amanda K. Hall CHI 2026 Workshops | April 2026
Publication From Binary Groundedness to Support Relations: Towards a Reader-Centred Taxonomy for Comprehension of AI Output Advait Sarkar, Christian Poelitz, Viktor Kewenig ACM CHI 2026 Workshop on Science and Technology for Augmenting Reading (CHI ’26 STAR) | April 2026 Project
Publication The State and Fate of Multilingual, Contextual Evaluation in the NLP World Manan Uppadhyay, Himanshu Beniwal, Prashant Kodali, Sunayana Sitaram April 2026
Publication Identifying Harm in Personalized, Generative AI Systems Require User-Centered Auditing at the Interaction Level Hannah Cha HEAL @ CHI ’26 | April 2026
Publication Mimetic Alignment with ASPECT: Evaluation of AI-inferred Personal Profiles Ruoxi Shang, Dan Marshall, Edward Cutrell, Denae Ford March 2026 arXiv | March 2026
Publication A Decade-Scale Benchmark Evaluating LLMs’ Clinical Practice Guidelines Detection and Adherence in Multi-turn Conversations Andong Tan, Shuyun Dai, Jinglu Wang, Fengtao Zhou, Yan Lu, Xi (Ada) Wang, Ying-Che Chen, Can Yang, Shujie Liu, Hao Chen March 2026 arXiv | March 2026