VLDB 2026 Research / reviewers in the wild / expert
Weiyun Wang
dblp:311/3689
· DBLP profile ↗
18ranked-venue papers
6as first author
18since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 12 · 3 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-author · 4 since 2021Human-computer interaction and ubiquitous computing · 4 · 3 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | From Information to Experience: Exploring Users' Engagement with Different Stress DisplaysabstractWith stress-tracking technologies becoming pervasive, a range of feedback displays has been explored to support engagement with stress data. However, most displays are studied in isolation, leaving limited understanding of how different forms of feedback shape engagement differently. This study compares three real-time stress feedback displays: screen-based quantitative, screen-based expressive, and physical expressive. Twenty-one participants took part in stress induction and recovery tasks to gain direct experience with each display, with data collected through surveys and semi-structured interviews. We found that different displays became associated with different modes of engagement: quantitative displays supported mobile, analytical use; screen-based expressive displays encouraged active monitoring; and physical expressive displays enabled peripheral awareness. We highlight the importance of balancing awareness with emotional experience, shaped by data representation and materiality. We challenge the assumption in personal informatics that more information or a more intuitive understanding is always beneficial, and offer a design framework for experiencing sensitive bio-signal feedback. Weiyun Wang, Scot Gilmour, Naral Chalermchaikosol, Ilyena Hirskyj-Douglas, Xianghua Ding |
DIS | 1 |
| 2026 | EvoMoE: Expert Evolution in Mixture of Experts for Multimodal Large Language ModelsabstractRecent advancements have shown that the Mixture of Experts (MoE) approach significantly enhances the capacity of large language models (LLMs) and improves performance on downstream tasks. Building on these promising results, multi-modal large language models (MLLMs) have increasingly adopted MoE techniques. However, existing multi-modal MoE tuning methods typically face two key challenges: expert uniformity and router rigidity. Expert uniformity occurs because MoE experts are often initialized by simply replicating the FFN parameters from LLMs, leading to homogenized expert functions and weakening the intended diversification of the MoE architecture. Meanwhile, router rigidity stems from the prevalent use of static linear routers for expert selection, which fail to distinguish between visual and textual tokens, resulting in similar expert distributions for image and text. To address these limitations, we propose EvoMoE, an innovative MoE tuning framework. EvoMoE introduces a meticulously designed expert initialization strategy that progressively evolves multiple robust experts from a single trainable expert, a process termed expert evolution that specifically targets severe expert homogenization. Furthermore, we introduce the Dynamic Token-aware Router (DTR), a novel routing mechanism that allocates input tokens to appropriate experts based on their modality and intrinsic token values. This dynamic routing is facilitated by hypernetworks, which dynamically generate routing weights tailored for each individual token. Extensive experiments demonstrate that EvoMoE significantly outperforms other sparse MLLMs across a variety of multi-modal benchmarks, including MME, MMBench, TextVQA, and POPE. Our results highlight the effectiveness of EvoMoE in enhancing the performance of MLLMs by addressing the critical issues of expert uniformity and router rigidity. Linglin Jing, Zhigang Wang 0002, Wang Lan, Weiyun Wang, Wenhai Wang, Qingpei Guo |
AAAI | 6 |
| 2026 | "I Want to Keep My Phone Away From the Bed": Designing a Smart Pillow for Sleep OnsetabstractPre-sleep digital consumption is widespread. While it is a common concern for bedtime procrastination, recent research also highlights its importance in fulfilling various pre-sleep needs, such as claiming "me time". However, these benefits and underlying needs have largely been overlooked in the design of digital sleep interventions. In this paper, we present a co-design workshop with 16 participants, exploring a smart pillow that allows audio-based digital consumption through non-distracting interactions. We illustrate the smart pillow’s potential in resolving the tension between digital consumption and sleep transition to support sleep onset — the transition from wakefulness to sleep. Our work highlights how the pillow’s physical form affords audio consumption control with minimal effort, and positions sleep onset as a distinct design context that demands careful attention when designing for sleep. We offer design implications that leverage tangible and bodily interactions to accommodate the sensitive transitioning state during sleep onset. Weiyun Wang, Ilyena Hirskyj-Douglas, Kejin Yu, Xianghua Ding |
TEI | 1 |
| 2025 | ChemVLM: Exploring the Power of Multimodal Large Language Models in Chemistry AreaabstractLarge Language Models (LLMs) have achieved remarkable success and have been applied across various scientific fields, including chemistry. However, many chemical tasks require the processing of visual information, which cannot be successfully handled by existing chemical LLMs. This brings a growing need for models capable of integrating multimodal information in the chemical domain. In this paper, we introduce ChemVLM, an open-source chemical multimodal large language model specifically designed for chemical applications. ChemVLM is trained on a carefully curated bilingual multimodal dataset that enhances its ability to understand both textual and visual chemical information, including molecular structures, reactions, and chemistry examination questions. We develop three datasets for comprehensive evaluation, tailored to Chemical Optical Character Recognition (OCR), Multimodal Chemical Reasoning (MMCR), and Multimodal Molecule Understanding tasks. We benchmark ChemVLM against a range of open-source and proprietary multimodal large language models on various tasks. Experimental results demonstrate that ChemVLM achieves competitive performance across all evaluated tasks. Junxian Li 0001, Di Zhang 0026, Xunzhi Wang, Zeying Hao, Jingdi Lei, Cai Zhou, Wei Liu 0123, Yaotian Yang, Xinrui Xiong, Weiyun Wang, Zhe Chen 0013, Wenhai Wang, Wei Li 0076, Mao Su, Shufei Zhang, Wanli Ouyang, Dongzhan Zhou |
AAAI | 11 |
| 2025 | OmniAlign-V: Towards Enhanced Alignment of MLLMs with Human PreferenceabstractXiangyu Zhao, Shengyuan Ding, Zicheng Zhang, Haian Huang, Maosongcao Maosongcao, Jiaqi Wang, Weiyun Wang, Xinyu Fang, Wenhai Wang, Guangtao Zhai, Hua Yang, Haodong Duan, Kai Chen. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Shengyuan Ding, Haian Huang, Maosongcao, Jiaqi Wang 0003, Weiyun Wang, Xinyu Fang, Wenhai Wang, Guangtao Zhai, Hua Yang 0001, Haodong Duan, Kai Chen 0026 |
ACL (1) | 7 |
| 2025 | Docopilot: Improving Multimodal Models for Document-Level UnderstandingabstractDespite significant progress in multimodal large language models (MLLMs), their performance on complex, multi-page document comprehension remains inadequate, largely due to the lack of high-quality, document-level datasets. While current retrieval-augmented generation (RAG) methods offer partial solutions, they suffer from issues, such as fragmented retrieval contexts, multi-stage error accumulation, and extra time costs of retrieval. In this work, we present a high-quality document-level dataset, Doc-750K, designed to support in-depth understanding of multimodal documents. This dataset includes diverse document structures, extensive cross-page dependencies, and real question-answer pairs derived from the original documents. Building on the dataset, we develop a native multimodal model—Docopilot, which can accurately handle document-level dependencies without relying on RAG. Experiments demonstrate that Docopilot achieves superior coherence, accuracy, and efficiency in document understanding tasks and multi-turn interactions, setting a new baseline for document-level multimodal understanding. Data, code, and models are released at https://github.com/OpenGVLab/Docopilot. Yuchen Duan, Zhe Chen 0017, Yusong Hu, Weiyun Wang, Shenglong Ye, Botian Shi, Lewei Lu, Qibin Hou, Tong Lu 0002, Hongsheng Li 0001, Jifeng Dai, Wenhai Wang |
CVPR | 4 |
| 2025 | Vision-RWKV: Efficient and Scalable Visual Perception with RWKV-Like ArchitecturesabstractTransformers have revolutionized computer vision and natural language processing, but their high computational complexity limits their application in high-resolution image processing and long-context analysis. This paper introduces Vision-RWKV (VRWKV), a model that builds upon the RWKV architecture from the NLP field with key modifications tailored specifically for vision tasks. Similar to the Vision Transformer (ViT), our model demonstrates robust global processing capabilities, efficiently handles sparse inputs like masked images, and can scale up to accommodate both large-scale parameters and extensive datasets. Its distinctive advantage is its reduced spatial aggregation complexity, enabling seamless processing of high-resolution images without the need for window operations. Our evaluations demonstrate that VRWKV surpasses ViT's performance in image classification and has significantly faster speeds and lower memory usage processing high-resolution inputs. In dense prediction tasks, it outperforms window-based models, maintaining comparable speeds. These results highlight VRWKV's potential as a more efficient alternative for visual perception tasks. Code and models are available at~\url{https://github.com/OpenGVLab/Vision-RWKV}. Yuchen Duan, Weiyun Wang, Zhe Chen 0017, Xizhou Zhu, Lewei Lu, Tong Lu 0002, Yu Qiao 0001, Hongsheng Li 0001, Jifeng Dai, Wenhai Wang |
ICLR | 2 |
| 2025 | OmniCorpus: A Unified Multimodal Corpus of 10 Billion-Level Images Interleaved with TextabstractImage-text interleaved data, consisting of multiple images and texts arranged in a natural document format, aligns with the presentation paradigm of internet data and closely resembles human reading habits. Recent studies have shown that such data aids multimodal in-context learning and maintains the capabilities of large language models during multimodal fine-tuning. However, the limited scale and diversity of current image-text interleaved data restrict the development of multimodal large language models. In this paper, we introduce OmniCorpus, a 10 billion-scale image-text interleaved dataset. Using an efficient data engine, we filter and extract large-scale high-quality documents, which contain 8.6 billion images and 1,696 billion text tokens. Compared to counterparts (e.g., MMC4, OBELICS), our dataset 1) has 15 times larger scales while maintaining good data quality; 2) features more diverse sources, including both English and non-English websites as well as video-centric websites; 3) is more flexible, easily degradable from an image-text interleaved format to pure text corpus and image-text pairs. Through comprehensive analysis and experiments, we validate the quality, usability, and effectiveness of the proposed dataset. We hope this could provide a solid data foundation for future multimodal model research. Qingyun Li, Zhe Chen 0017, Weiyun Wang, Wenhai Wang, Shenglong Ye, Zhenjiang Jin, Guanzhou Chen 0004, Yinan He, Zhangwei Gao, Erfei Cui, Jiashuo Yu, Hao Tian 0006, Bin Wang 0065, Xingjian Wei, Wei Li 0320, Wenjian Zhang, Bo Zhang 0069, Pinlong Cai |
ICLR | 3 |
| 2025 | OWMM-Agent: Open World Mobile Manipulation With Multi-modal Agentic Data SynthesisabstractThe rapid progress of navigation, manipulation, and vision models has made mobile manipulators capable in many specialized tasks.
However, the open-world mobile manipulation (OWMM) task remains a challenge due to the need for generalization to open-ended instructions and environments, as well as the systematic complexity to integrate high-level decision making with low-level robot control based on both global scene understanding and current agent state. To address this complexity, we propose a novel multi-modal agent architecture that maintains multi-view scene frames and agent states for decision-making and controls the robot by function calling.
A second challenge is the hallucination from domain shift. To enhance the agent performance, we further introduce an agentic data synthesis pipeline for the OWMM task to adapt the VLM model to our task domain with instruction fine-tuning. We highlight our fine-tuned OWMM-VLM as the first dedicated foundation model for mobile manipulators with global scene understanding, robot state tracking, and multi-modal action generation in a unified model. Through experiments, we demonstrate that our model achieves SOTA performance compared to other foundation models including GPT-4o and strong zero-shot generalization in real world.
The project page is at https://hhyhrhy.github.io/owmm-agent-project. Haotian Liang, Lingxiao Du, Weiyun Wang, Mengkang Hu, Yao Mu 0001, Wenhai Wang, Jifeng Dai, Ping Luo 0002, Wenqi Shao, Lin Shao 0002 |
NeurIPS | 4 |
| 2025 | Visual Thoughts: A Unified Perspective of Understanding Multimodal Chain-of-ThoughtabstractLarge Vision-Language Models (LVLMs) have achieved significant success in multimodal tasks, with multimodal chain-of-thought (MCoT) further enhancing performance and interpretability. Recent MCoT methods fall into two categories: (i) Textual-MCoT (T-MCoT), which takes multimodal input and produces textual output; and (ii) Interleaved-MCoT (I-MCoT), which generates interleaved image-text outputs. Despite advances in both approaches, the mechanisms driving these improvements are not fully understood. To fill this gap, we first reveal that MCoT boosts LVLMs by incorporating $\textit{visual thoughts}$, which convey image information to the reasoning process regardless of the MCoT format, depending only on clarity and conciseness of expression. Furthermore, to explore visual thoughts systematically, we define four distinct forms of visual thought expressions and analyze them comprehensively. Our findings demonstrate that these forms differ in clarity and conciseness, yielding varying levels of MCoT improvement. Additionally, we explore the internal nature of visual thoughts, finding that visual thoughts serve as intermediaries between the input image and reasoning to deeper transformer layers, enabling more advanced visual information transmission. We hope that the visual thoughts can inspire further breakthroughs for future MCoT research. Zihui Cheng, Qiguang Chen, Xiao Xu 0005, Jiaqi Wang 0012, Weiyun Wang, Hao Fei 0003, Yidong Wang 0003, Alex Jinpeng Wang, Zhi Chen 0006, Wanxiang Che, Libo Qin 0001 |
NeurIPS | 5 |
| 2025 | Demystify Transformers & Convolutions in Modern Image Deep NetworksabstractVision transformers have gained popularity recently, leading to the development of new vision backbones with improved features and consistent performance gains. However, these advancements are not solely attributable to novel feature transformation designs; certain benefits also arise from advanced network-level and block-level architectures. This paper aims to identify the real gains of popular convolution and attention operators through a detailed study. We find that the key difference among these feature transformation modules, such as attention or convolution, lies in their spatial feature aggregation approach, known as the "spatial token mixer" (STM). To facilitate an impartial comparison, we introduce a unified architecture to neutralize the impact of divergent network-level and block-level designs. Subsequently, various STMs are integrated into this unified framework for comprehensive comparative analysis. Our experiments on various tasks and an analysis of inductive bias show a significant performance boost due to advanced network-level and block-level designs, but performance differences persist among different STMs. Our detailed analysis also reveals various findings about different STMs, including effective receptive fields, invariance, and adversarial robustness tests. Xiaowei Hu 0001, Min Shi 0004, Weiyun Wang, Sitong Wu, Linjie Xing, Wenhai Wang, Xizhou Zhou, Lewei Lu, Jie Zhou 0001, Xiaogang Wang 0005, Yu Qiao 0001, Jifeng Dai |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2024 | The All-Seeing Project V2: Towards General Relation Comprehension of the Open World
Weiyun Wang, Yiming Ren 0001, Haowen Luo, Tiantong Li, Chenxiang Yan, Zhe Chen 0017, Wenhai Wang, Qingyun Li, Lewei Lu, Xizhou Zhu, Yu Qiao 0001, Jifeng Dai |
ECCV (33) | 1 |
| 2024 | The All-Seeing Project: Towards Panoptic Visual Recognition and Understanding of the Open WorldabstractWe present the All-Seeing (AS) project: a large-scale dataset and model for recognizing and understanding everything in the open world.
Using a scalable data engine that incorporates human feedback and efficient models in the loop, we create a new dataset (AS-1B) with over 1.2 billion regions annotated with semantic tags, question-answering pairs, and detailed captions. It covers a wide range of 3.5 million common and rare concepts in the real world and has 132.2 billion tokens that describe the concepts and their attributes. Leveraging this new dataset, we develop the All-Seeing model (ASM), a unified framework for panoptic visual recognition and understanding. The model is trained with open-ended language prompts and locations, which allows it to generalize to various vision and language tasks with remarkable zero-shot performance, including both region- and image-level retrieval, region recognition, captioning, and question-answering. We hope that this project can serve as a foundation for vision-language artificial general intelligence research. Code is available at https://github.com/OpenGVLab/all-seeing. Weiyun Wang, Min Shi 0004, Qingyun Li, Wenhai Wang, Zhenhang Huang, Linjie Xing, Zhe Chen 0017, Hao Li 0069, Xizhou Zhu, Zhiguo Cao 0001, Tong Lu 0002, Jifeng Dai, Yu Qiao 0001 |
ICLR | 1 |
| 2024 | Needle In A Multimodal HaystackabstractWith the rapid advancement of multimodal large language models (MLLMs), their evaluation has become increasingly comprehensive. However, understanding long multimodal content, as a foundational ability for real-world applications, remains underexplored. In this work, we present Needle In A Multimodal Haystack (MM-NIAH), the first benchmark specifically designed to systematically evaluate the capability of existing MLLMs to comprehend long multimodal documents. Our benchmark includes three types of evaluation tasks: multimodal retrieval, counting, and reasoning. In each task, the model is required to answer the questions according to different key information scattered throughout the given multimodal document. Evaluating the leading MLLMs on MM-NIAH, we observe that existing models still have significant room for improvement on these tasks, especially on vision-centric evaluation. We hope this work can provide a platform for further research on long multimodal document comprehension and contribute to the advancement of MLLMs. Code and benchmark are released at https://github.com/OpenGVLab/MM-NIAH. Weiyun Wang, Shuibo Zhang, Yiming Ren 0001, Yuchen Duan, Tiantong Li, Mengkang Hu, Zhe Chen 0017, Kaipeng Zhang, Lewei Lu, Xizhou Zhu, Ping Luo 0002, Yu Qiao 0001, Jifeng Dai, Wenqi Shao, Wenhai Wang |
NeurIPS | 1 |
| 2024 | How far are we to GPT-4V? Closing the gap to commercial multimodal models with open-source suites
Zhe Chen 0017, Weiyun Wang, Hao Tian 0006, Shenglong Ye, Zhangwei Gao, Erfei Cui, Wenwen Tong, Kongzhi Hu, Jiapeng Luo, Zheng Ma 0012, Jiaqi Wang 0003, Xiaoyi Dong, Hang Yan 0001, Hewei Guo, Conghui He, Botian Shi, Zhenjiang Jin, Bin Wang 0065, Xingjian Wei, Wei Li 0320, Wenjian Zhang, Bo Zhang 0069, Pinlong Cai, Licheng Wen, Xiangchao Yan, Min Dou, Lewei Lu, Xizhou Zhu, Tong Lu 0002, Dahua Lin, Yu Qiao 0001, Jifeng Dai, Wenhai Wang |
Sci. China Inf. Sci. | 2 |
| 2024 | MMInstruct: a high-quality multi-modal instruction tuning dataset with extensive diversity
Yangzhou Liu, Zhangwei Gao, Weiyun Wang, Zhe Chen 0017, Wenhai Wang, Hao Tian 0006, Lewei Lu, Xizhou Zhu, Tong Lu 0002, Yu Qiao 0001, Jifeng Dai |
Sci. China Inf. Sci. | 4 |
| 2023 | Everyday Space as an Interface for Health Data Engagement: Designing Tangible Displays of Stress DataabstractHealth data user engagement, particularly with stress data, remains a challenge despite the widespread use of self-tracking products, like smartwatches and smart bracelets. Stress data engagement is crucial to the early detection and intervention of long-term stress which could cause harmful health effects. This paper explores the design of tangible displays to enhance engagement with self-tracked stress data. We conducted two co-design workshops in which participants were invited to design and draw sketches of stress displays for three different contexts. The workshops revealed many innovative ideas for using everyday spaces and materials as an interface to structure user interactions with the data, aimed at increasing awareness of stress data and management strategies while addressing various concerns associated with how the data is displayed. By focusing on stress data, this study highlights important opportunities to use everyday spaces as an interface for health data engagement. Weiyun Wang, Xianghua Ding, Ilyena Hirskyj-Douglas |
Conference on Designing Interactive Systems | 1 |
| 2023 | Digital Making for Inheritance and Enlivening Intangible Cultural Heritage: A Case of Hairy Monkey HandicraftsabstractDigital technologies can conduct an important role in preserving intangible cultural heritage (ICH). Nonetheless, existing work tends to be limited to digital storage, presentation, dissemination, and education, with comparatively little concentrating on production and reproduction of the craft, the key to revitalizing ICH. In this paper, we explore digital making as an approach for both the inheritance and innovation of ICH handicrafts. Taking Hairy Monkey craftsmanship as an instance, we conducted a workshop with 15 groups of makers, teaching them the traditional Hairy Monkey craft and subsequently enabling them to create innovative works with digital technologies in their own time. As revealed by our interviews with these participants, ICH brings cultural inspiration to digital making, and digital making rejuvenates ICH through innovative art forms and its positive influence on participants. As demonstrated in this paper, digital technologies can be deeply integrated with ICH through making to revitalize ICH from the core through living transmission. Guanhong Liu, Xianghua Ding, Jinghe Cai, Weiyun Wang, Yuting Diao, Tianyu Yu 0001, Haiqing Xu 0001, Haipeng Mi |
CHI | 4 |