Biography

I am a Researcher at RBC Borealis. I recently received my Ph.D. from University of Waterloo, where I was supervised by Prof. Krzysztof Czarnecki. I obtained my master’s degree under the supervision of Prof. Lijun Zhang in the LAMDA Group led by Prof. Zhihua Zhou at Nanjing University. I have also had research experience at Amazon, SONY, Borealis, Tsinghua University with Prof. Jingjing Liu and Prof. Yang Liu, Tencent Lightspeed & Quantum Studios, Alibaba, and Netease Games.

My research focuses on reliable and efficient multimodal learning, with a focus on vision-language models and multimodal retrieval.

I am always happy to connect about research opportunities and collaborations in multimodal learning (reliability, efficiency, and retrieval).

Yimu Wang
Researcher, RBC Borealis
Ph.D., University of Waterloo

        

News

Academic Services

  1. Workshop Organizer: Grounded and Faithful Vision-Language Models for Real-World Deployment (VLM4RWD), NeurIPS 2026
  2. Workshop Organizer: Grounding Language Models: Learning Faithfully and Efficiently, EMNLP 2026
  3. Workshop Organizer: Vision Language Models: Challenges of Real World Deployment (VLM4RWD), NeurIPS 2025

Publications [Google Scholar]

  1. Where Does the Answer Come From? Benchmarking View-Level Visual Evidence Identification in Multi-View MLLMs for Autonomous Driving
    Yimu Wang, Yee Man Choi, Barry Zhang, Mozhgan Nasr Azadani, Sean Sedwards, Krzysztof Czarnecki
    arXiv preprint, 2026.

    arXiv 2026 Arxiv

  2. VISTAQA: Benchmarking Joint Visual Question Answering and Pixel-Level Evidence
    Mozhgan Nasr Azadani, Yimu Wang, Yongpeng Zhu, Lihong Chen, Milan Ganai, Sean Sedwards, Marco Pavone, Krzysztof Czarnecki
    arXiv preprint, 2026.

    arXiv 2026 Arxiv

  3. UNIFORM: Unifying Knowledge from Large-scale and Diverse Pre-trained Models
    Yimu Wang, Weiming Zhuang, Chen Chen, Jiabo Huang, Jingtao Li, Lingjuan Lyu
    IEEE Conference on Computer Vision and Pattern Recognition (CVPR Findings), 2026.

    CVPR Findings 2026 Paper

  4. Mitigating the Modality Gap: Few-Shot Out-of-Distribution Detection with Multi-modal Prototypes and Image Bias Estimation
    Yimu Wang, Evelien Riddell, Adrian Chow, Sean Sedwards, Krzysztof Czarnecki
    IEEE Winter Conference on Applications of Computer Vision (WACV), 2026.

    WACV 2026 Paper

  5. Lexicographic Lipschitz Bandits: New Algorithms and a Lower Bound
    Bo Xue, Ji Cheng, Fei Liu, Yimu Wang, Lijun Zhang, and Qingfu Zhang
    Journal of Machine Learning Research (JMLR), 2025.

    JMLR 2025 Paper

  6. Hawaii: Hierarchical Visual Knowledge Transfer for Efficient Vision-Language Models
    Yimu Wang, Mozhgan Nasr Azadani, Sean Sedwards, Krzysztof Czarnecki
    Annual Conference on Neural Information Processing Systems (NeurIPS), 2025.

    NeurIPS 2025 Arxiv

  7. Survey of Video Diffusion Models: Foundations, Implementations, and Applications
    Yimu Wang, Xuye Liu, Wei Pang, Li Ma, Shuai Yuan, Paul Debevec, Ning Yu
    Transactions on Machine Learning Research (TMLR), 2025.

    TMLR 2025 Arxiv Paper

  8. LEO-MINI: An Efficient Multimodal Large Language Model using Conditional Token Reduction and Mixture of Multi-Modal Experts
    Yimu Wang, Mozhgan Nasr Azadani, Sean Sedwards, Krzysztof Czarnecki
    Empirical Methods in Natural Language Processing (EMNLP), 2025.

    EMNLP 2025 Arxiv

  9. Rethinking Spectral Augmentation for Contrast-based Graph Self-Supervised Learning
    Xiangru Jian, Xinjian Zhao, Wei Pang, Chaolong Ying, Yimu Wang, Yaoyao Xu, Tianshu Yu
    Transactions on Machine Learning Research (TMLR), 2025.

    TMLR 2025

  10. OV-SCAN: Semantically Consistent Alignment for Novel Object Discovery in Open-Vocabulary 3D Object Detection
    A. Chow, E. Riddell, Yimu Wang, S. Sedwards, K. Czarnecki
    International Conference on Computer Vision (ICCV), 2025.

    ICCV 2025 Arxiv

  11. NBDESCRIB: A Dataset for Text Description Generation from Tables and Code in Jupyter Notebooks with Guidelines
    Xuye Liu, Tengfei Ma, Yimu Wang, Fengjie Wang, Jian Zhao
    Annual Meeting of the Association for Computational Linguistics (Findings of ACL), 2025.

    Findings of ACL 2025

  12. ELIOT: Zero-Shot Video-Text Retrieval through Relevance-Boosted Captioning and Structural Information Extraction
    Xuye Liu, Yimu Wang, Jian Zhao
    NAACL Student Research Workshop (SRW of NAACL), 2025.

    SRW of NAACL 2025 Paper

  13. DREAM: Improving Video-Text Retrieval Through Relevance-Based Augmentation Using Large Foundation Models
    Yimu Wang, Shuai Yuan, Bo Xue, Xiangru Jian, Wei Pang, Mushi Wang, Ning Yu
    Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics (NAACL), 2025.

    NAACL 2025 Paper

  14. AIDE: Improving 3D Open-Vocabulary Semantic Segmentation by Aligned Vision-Language Learning
    Yimu Wang, Krzysztof Czarnecki
    IEEE Winter Conference on Applications of Computer Vision (WACV), 2025.

    WACV 2025 Paper

  15. Conditional Generative Adversarial Network-Assisted System for Radiation-Free Evaluation of Scoliosis Using a Single Smartphone Photograph: A Model Development and Validation Study
    Zhong He, Neng Lu, Yi Chen, Elvis Chun-Sing Chui, Zhen Liu, Xiaodong Qin, Jie Li, Shengru Wang, Junlin Yang, Zhiwei Wang, et al.
    eClinicalMedicine, 2024.

    eClinicalMedicine 2024

  16. NICE: CVPR 2023 Challenge on Zero-Shot Image Captioning
    Taehoon Kim, Pyunghwan Ahn, Sangyun Kim, Sihaeng Lee, Mark Marsden, Alessandra Sala, Seung Hwan Kim, Bohyung Han, Kyoung Mu Lee, Honglak Lee, et al.
    IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPR Workshops), 2024.

    CVPR Workshops 2024

  17. Self-Supervised Pretext Tasks for Event Sequence Data from Detecting Misalignment
    Yimu Wang, He Zhao, Ruizhi Deng, Frederick Tung, Greg Mori
    Conference on Neural Information Processing Systems Workshop (NeurIPS workshop), 2024.

    NeurIPS Workshop 2024 Paper

  18. Lost Domain Generalization Is a Natural Consequence of Lack of Training Domains
    Yimu Wang, Yihan Wu, Hongyang Zhang
    Proceedings of the AAAI Conference on Artificial Intelligence (AAAI), 2024.

    AAAI 2024 Paper

  19. Multiobjective Lipschitz Bandits under Lexicographic Ordering
    Bo Xue, Ji Cheng, Fei Liu, Yimu Wang, Qingfu Zhang
    Proceedings of the AAAI Conference on Artificial Intelligence (AAAI), 2024.

    AAAI 2024 Paper

  20. Efficient Algorithms for Generalized Linear Bandits with Heavy-tailed Rewards
    Bo Xue, Yimu Wang, Yuanyu Wan, Jinfeng Yi, and Lijun Zhang
    Conference on Neural Information Processing Systems (NeurIPS), 2023.

    NeurIPS 2023 Paper

  21. Balance Act: Mitigating Hubness in Cross-Modal Retrieval with Query and Gallery Banks
    Yimu Wang, Xiangru Jian, Bo Xue
    Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing (EMNLP Oral), 2023.

    EMNLP (Oral) 2023 Paper Code

  22. Video-Text Retrieval by Supervised Sparse Multi-Grained Learning
    Yimu Wang, Peng Shi
    Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing (Findings of EMNLP), 2023.

    Findings of EMNLP 2023 Paper Code

  23. InvGC: Robust Cross-Modal Retrieval by Inverse Graph Convolution
    Xiangru Jian, Yimu Wang
    Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing (Findings of EMNLP), 2023.

    Findings of EMNLP 2023 Paper Code

  24. Cooperation or Competition: Avoiding Player Domination for Multi-target Robustness by Adaptive Budgets
    Yimu Wang, Dinghuai Zhang, Yihan Wu, Heng Huang, Hongyang Zhang
    IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2023.

    CVPR 2023 Paper

  25. Investigating the Existence of “Secret Language” in Language Models
    Yimu Wang, Peng Shi, Hongyang Zhang
    arXiv preprint, 2023.

    arXiv 2023 Arxiv

  26. Multimodal Federated Learning via Contrastive Representation Ensemble
    Qiying Yu, Yang Liu, Yimu Wang, Ke Xu, Jingjing Liu
    International Conference on Learning Representations (ICLR), 2023.

    ICLR 2023 Paper Code

  27. Multi-View Fusion Transformer for Sensor-Based Human Activity Recognition
    Yimu Wang, Kun Yu, Yubo Wang, Hongwei Xue
    arXiv preprint, 2022.

    arXiv 2022 Arxiv

  28. Deep Unified Cross-Modality Hashing by Pairwise Data Alignment
    Yimu Wang, Bo Xue, Quan Cheng, Yuhui Chen, and Lijun Zhang
    International Joint Conference on Artificial Intelligence (IJCAI), 2021.

    IJCAI 2021 Paper

  29. Classification of Neurofibromatosis-Related Dystrophic or Nondystrophic Scoliosis Based on Image Features Using Bilateral CNN
    Zhengwang He, Yimu Wang, Xue Qin, Rui Yin, Yong Qiu, Ke He, Zezhang Zhu
    Medical Physics, 2021.

    Medical Physics 2021

  30. Piecewise Hashing: A Deep Hashing Method for Large-Scale Fine-Grained Search
    Yimu Wang, Xiu-Shen Wei, Bo Xue, Lijun Zhang
    Chinese Conference on Pattern Recognition and Computer Vision (PRCV), 2020.

    PRCV 2020

  31. Searching Privately by Imperceptible Lying: A Novel Private Hashing Method with Differential Privacy
    Yimu Wang, Shiyin Lu, and Lijun Zhang
    ACM International Conference on Multimedia (ACM MM), 2020.

    ACM MM 2020 Paper

  32. Nearly Optimal Regret for Stochastic Linear Bandits with Heavy-Tailed Payoffs
    Bo Xue, Guanghui Wang, Yimu Wang, Lijun Zhang
    International Joint Conference on Artificial Intelligence (IJCAI), 2020.

    IJCAI 2020 Paper

  33. An Adversarial Domain Adaptation Network for Cross-Domain Fine-Grained Recognition
    Yimu Wang, Ren-Jie Song, Xiu-Shen Wei, and Lijun Zhang
    IEEE Winter Conference on Applications of Computer Vision (WACV), 2020.

    WACV 2020 Paper

Education and Research Experience

Experience

Awards