VLDB 2026 Research / reviewers in the wild / expert
Wei Wang 0354
dblp:35/7092-354
· DBLP profile ↗
9ranked-venue papers
5as first author
8since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 8 · 4 first-author · 7 since 2021Artificial intelligence and machine learning · 4 · 1 first-author · 4 since 2021Databases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | IDEA-Bench: How Far are Generative Models from Professional Designing?abstractRecent advancements in image generation models enable the creation of high-quality images and targeted modifications based on textual instructions. Some models even support multimodal complex guidance and demonstrate robust task generalization capabilities. However, they still fall short of meeting the nuanced, professional demands of designers. To bridge this gap, we introduce IDEA-Bench, a comprehensive benchmark designed to advance image generation models toward applications with robust task generalization. IDEA-Bench comprises 100 professional image generation tasks and 275 specific cases, categorized into five major types based on the current capabilities of existing models. Furthermore, we provide a representative subset of 18 tasks with enhanced evaluation criteria to facilitate more nuanced and reliable evaluations using Multimodal Large Language Models (MLLMs). By assessing models’ ability to comprehend and execute novel, complex tasks, IDEA-Bench paves the way toward the development of generative models with autonomous and versatile visual generation capabilities. Lianghua Huang, Jingwu Fang, Huanzhang Dou, Wei Wang 0354, Zhi-Fan Wu, Yupeng Shi, Junge Zhang, Xin Zhao 0012, Yu Liu 0063 |
CVPR | 5 |
| 2024 | Chains of Diffusion Models
Yanheng Wei, Lianghua Huang, Zhi-Fan Wu, Wei Wang 0354, Mingda Jia, Shuailei Ma |
ECCV (72) | 4 |
| 2024 | MultiGen: Zero-Shot Image Generation from Multi-modal Prompts
Zhi-Fan Wu, Lianghua Huang, Wei Wang 0354, Yanheng Wei |
ECCV (8) | 3 |
| 2023 | Weakly-Supervised Video Object Grounding via Causal InterventionabstractWe target at the task of weakly-supervised video object grounding (WSVOG), where only video-sentence annotations are available during model learning. It aims to localize objects described in the sentence to visual regions in the video, which is a fundamental capability needed in pattern analysis and machine learning. Despite the recent progress, existing methods all suffer from the severe problem of spurious association, which will harm the grounding performance. In this paper, we start from the definition of WSVOG and pinpoint the spurious association from two aspects: (1) the association itself is not object-relevant but extremely ambiguous due to weak supervision; and (2) the association is unavoidably confounded by the observational bias when taking the statistics-based matching strategy in existing methods. With this in mind, we design a unified causal framework to learn the deconfounded object-relevant association for more accurate and robust video object grounding. Specifically, we learn the object-relevant association by causal intervention from the perspective of video data generation process. To overcome the problems of lacking fine-grained supervision in terms of intervention, we propose a novel spatial-temporal adversarial contrastive learning paradigm. To further remove the accompanying confounding effect within the object-relevant association, we pursue the true causality by conducting causal intervention via backdoor adjustment. Finally, the deconfounded object-relevant association is learned and optimized under a unified causal framework in an end-to-end manner. Extensive experiments on both IID and OOD testing sets of three benchmarks demonstrate its accurate and robust grounding performance against state-of-the-arts. Wei Wang 0354, Junyu Gao 0002, Changsheng Xu |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2023 | Weakly-Supervised Video Object Grounding via Learning Uni-Modal AssociationsabstractGrounding objects described in natural language to visual regions in the video is a crucial capability needed in vision-and-language fields. In this paper, we deal with the weakly-supervised video object grounding (WSVOG) task, where only video-sentence pairs are provided for learning. The essence of this task is to learn the cross-modal associations between words in textual modality and regions in visual modality. Despite the recent progress, we find that most existing methods focus on the association learning for cross-modal samples, while the rich and complementary information within uni-modal samples has not been fully exploited. To this end, we propose to explicitly learn uni-modal associations on both textual and visual sides, so as to fully exploit the useful uni-modal information for accurate video object grounding. Specifically, (1) we learn textual prototypes by considering rich contextual information of the same object in different sentences, and (2) we estimate visual prototypes in an adaptive manner so as to overcome the uncertainties in selecting object-relevant visual regions. Besides, a cross-modal correspondence is learned which not only bridges the visual and textual modalities for WSVOG task, but also tightly cooperates with the uni-modal association learning process. We conduct extensive experiments on three popular datasets, and the favorable results demonstrate the effectiveness of our method. Wei Wang 0354, Junyu Gao 0002, Changsheng Xu |
IEEE Trans. Multim. | 1 |
| 2023 | Many Hands Make Light Work: Transferring Knowledge From Auxiliary Tasks for Video-Text RetrievalabstractThe problem of video-text retrieval, which searches videos via natural language descriptions or vice versa, has attracted growing attention due to the explosive scale of videos produced every day. The dominant approaches for this problem follow the pipeline that firstly learns compact feature representations of videos and texts, and then jointly embeds them into a common feature space where matched video-text pairs are close and unmatched pairs are far away. However, most of them neither consider the structural similarities among cross-modal samples in a global view, nor leverage useful information from other relevant retrieval processes. We argue that both information has great potential for video-text retrieval. In this paper, we treat the relevant retrieval processes as auxiliary tasks and we extract useful knowledge from them by exploiting structural similarities via Graph Neural Networks (GNNs). We then progressively transfer the knowledge from auxiliary tasks in a general-to-specific manner to assist the main task of the current retrieval process. Specifically, for the retrieval of the given query, we first construct a sequence of query-graphs whose central queries are chosen from distant to close to the given query. Then we conduct knowledge-guided message passing in each query-graph to exploit regional structural similarities and gather knowledge of different levels from the updated query-graphs with a knowledge-based attention mechanism. Finally, we transfer the extracted useful knowledge from general to specific to assist the current retrieval process. Extensive experimental results show that our model outperforms the state-of-the-arts on four benchmarks. Wei Wang 0354, Junyu Gao 0002, Xiaoshan Yang, Changsheng Xu |
IEEE Trans. Multim. | 1 |
| 2021 | Weakly-Supervised Video Object Grounding via Stable Context LearningabstractWe investigate the problem of weakly-supervised video object grounding (WSVOG), where only the video-sentence annotations are provided for training. It aims at localizing the queried objects described in the sentence to visual regions in the video. Despite the recent progress, existing approaches have not fully exploited the potential of the description sentences for cross-modal alignment in two aspects: (1) Most of them extract objects from the description sentences and represent them with fixed textual representations. While achieving promising results, they do not make full use of the contextual information in the sentence. (2) A few works have attempted to utilize contextual information to learn object representations, but found a significant decrease in performance due to the unstable training in cross-modal alignment. To address the above issues, in this paper, we propose a Stable Context Learning (SCL) framework for WSVOG which jointly enjoys the merits of stable learning and rich contextual information. Specifically, we design two modules named Context-Aware Object Stabilizer module and Cross-Modal Alignment Knowledge Transfer module, which are cooperated together to inject contextual information to stable object concepts in text modality and transfer contextualized knowledge in cross-modal alignment. Our approach is finally optimized under a frame-level MIL paradigm. Extensive experiments on three popular benchmarks demonstrate its significant effectiveness. Wei Wang 0354, Junyu Gao 0002, Changsheng Xu |
ACM Multimedia | 1 |
| 2021 | Learning Coarse-to-Fine Graph Neural Networks for Video-Text RetrievalabstractWe address the problem of video-text retrieval that searches videos via natural language description or vice versa. Most state-of-the-art methods only consider cross-modal learning for two or three data points in isolation, ignoring to get benefit from the structural information of other data points from a global view. In this paper, we propose to exploit the comprehensive relationships among cross-modal samples via Graph Neural Networks (GNN). To improve the discriminative ability for accurately finding the positive sample, a Coarse-to-Fine GNN is constructed, which can progressively optimize the retrieval results via multi-step reasoning. Specifically, we first adopt heuristic edge features to represent relationships. Then we design a scoring module in each layer to rank the edges connected to the query node and drop the edges with lower scores. Finally, to alleviate the class imbalance issue, we propose a random-drop focal loss to optimize the whole framework. Extensive experimental results show that our method consistently outperforms the state-of-the-arts on four benchmarks. Wei Wang 0354, Junyu Gao 0002, Xiaoshan Yang, Changsheng Xu |
IEEE Trans. Multim. | 1 |
| 2019 | Multimodal Attribute and Feature Embedding for Activity RecognitionabstractHuman Activity Recognition (HAR) automatically recognizes human activities such as daily life and work based on digital records, which is of great significance to medical and health fields. Egocentric video and human acceleration data comprehensively describe human activity patterns from different aspects, which have laid a foundation for activity recognition based on multimodal behavior data. However, on the one hand, the low-level multimodal signal structures differ greatly and the mapping to high-level activities is complicated. On the other hand, the activity labeling based on multimodal behavior data has high cost and limited data amount, which limits the technical development in this field. In this paper, an activity recognition model MAFE based on multimodal attribute feature embedding is proposed. Before the activity recognition, the middle-level attribute features are extracted from the low-level signals of different modes. On the one hand, the mapping complexity from the low-level signals to the high-level activities is reduced, and on the other hand, a large number of middle-level attribute labeling data can be used to reduce the dependency on the activity labeling data. We conducted experiments on Stanford-ECM datasets to verify the effectiveness of the proposed MAFE method. Yi Huang 0037, Wanting Yu, Xiaoshan Yang, Wei Wang 0354, Jitao Sang 0001 |
MMAsia | 5 |