📝 Publications

sym

Why Vision Fails as a Universal Bridge: Rectifying Modality Asynchrony in Multilingual MLLMs | Static Badge

Yihang Du†, Juhao Liang†, Zhengzhao Lai, Siyu Li, Yan Hu.

“We uncover a modality-asynchrony phenomenon in multilingual multimodal large language models and introduce ANCHOR, a framework that promotes earlier visual–linguistic alignment to improve cross-lingual visual reasoning.”

sym

Can Multimodal LLMs See Materials Clearly? A Multimodal Benchmark on Materials Characterization | Static Badge

Zhengzhao Lai, Youbin Zheng, Zhenyang Cai, Haonan Lyu, Jingpu Yang, Hongqing Liang, Yan Hu, Benyou Wang

“MatCha is the first benchmark for materials characterization image understanding featuring the evaluation of multimodal large language models in real-world scientific scenarios.”