VLDB 2026 Research / reviewers in the wild / expert
Lei Zhao 0017
dblp:87/734-17
· DBLP profile ↗
9ranked-venue papers
3as first author
9since 2021 · last 2025
0000-0002-8838-164XORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 7 · 2 first-author · 7 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Temporal-Guided Mixture-of-Experts for Zero-Shot Video Question AnsweringabstractVideo Question Answering (VideoQA) is a challenging task in the vision-language field. Due to the time-consuming and labor-intensive labeling process of the question-answer pairs, fully supervised methods are no longer suitable for the current increasing demand for data. This has led to the rise of zero-shot VideoQA, and some works propose to adapt large language models (LLMs) to assist zero-shot learning. Despite recent progress, the inadequacy of LLMs in comprehending temporal information in videos and the neglect of temporal differences, e.g., the different dynamic changes between scenes or objects, remain insufficiently addressed by existing attempts in zero-shot VideoQA. In light of these challenges, a novel Temporal-guided Mixture-of-Experts Network (T-MoENet) for zero-shot video question answering is proposed in this paper. Specifically, we apply a temporal module to imbue language models with the capacity to perceive temporal information. Then a temporal-guided mixture-of-experts module is proposed to further learn the temporal differences presented in different videos. It enables the model to effectively improve the capacity of generalization. Our proposed method achieves state-of-the-art performance on multiple zero-shot VideoQA benchmarks, notably improving accuracy by 5.6% on TGIF-FrameQA and 2.3% on MSRVTT-QA while remaining competitive with other methods in the fully supervised setting. The codes and models developed in this study will be made publicly available athttps://github.com/qyx1121/T-MoENet. Yixin Qin, Lei Zhao 0017, Lianli Gao, Haonan Zhang 0003, Pengpeng Zeng, Heng Tao Shen |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2023 | Heterogeneous Knowledge Network for Visual DialogabstractVisual dialog requires an agent to answer successive questions considering an image and dialog history, which is a classic vision-language task. Despite progress, there are still two key challenges: 1) parsing long or complex questions and answers and 2) dealing with the visual scene containing complicated interactions among entities. These challenges bring about the unsatisfactory consequence of current visual dialog methods. In this paper, we propose a novel Heterogeneous Knowledge Network (HKNet), which leverages textual sequence knowledge and graph knowledge to address the above issues. Specifically, the textual sequence knowledge is derived from the sentences that are retrieved from the image captions of the visual dialog dataset. The textual sequence knowledge can supplement essential common sense for parsing long or complex questions and answers. The graph knowledge is constructed via scene graph, which provides complete visual relationships for understanding the complicated interactions. These two kinds of heterogeneous knowledge complement each other and jointly improve the logical reasoning ability of the visual dialog. Extensive experimental results on two benchmark datasets: VisDial v0.9 and v1.0 demonstrate the superiority of the proposed HKNet. Ablation studies and visualization results further verify the effectiveness of our method. Lei Zhao 0017, Lianli Gao, Yunbo Rao, Jingkuan Song, Heng Tao Shen |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2023 | From External to Internal: Structuring Image for Text-to-Image Attributes ManipulationabstractManipulating visual attributes of an image through a natural language description, known as text-to-image attributes manipulation (T2AM), is a challenging task. However, existing approaches tend to search the whole image to manipulate the target instance indicated by a description, thus they often fail to locate and manipulate the accurate text-relevant regions, and even disturb the text-irrelevant contents, e.g. texture and background. Meanwhile, the model efficiency needs to be improved. To tackle the above issues, we introduce a novel yet simple GAN-based approach, namelyStructuringImage forManipulating(SIMGAN), to narrow down the optimization areas from external to internal. It consists of two major components: 1)External Structuring(ExST), a pretrained segmentation network, for recognizing and separating the target instances and background from an image; and 2)Internal Structuring(InST) for seeking out and editing the text-relevant attributes of the target instances based on the given description and masked hierarchical image representations from ExST. Specifically, the InST structures target instances from outline to detail by firstly drawing the sketch and colors underpainting of instances with anOutline-Oriented Structuring(OuST), and then enhancing the text-relevant attributes and elaborating on details with aDetail-Oriented Structuring(DeST). Extensive experiments on benchmark datasets demonstrate that our framework significantly outperforms state-of-the-art both quantitatively and qualitatively. Compared with the state-of-the-art method ManiGAN, our approach reduces the training time by 88%, while the inferring time is three times faster. In addition, our approach is easily extended to solve the instance-level image-to-image translation problem, and the results exhibit the versatility and effectiveness of our approach. This code is released inhttps://github.com/qikizh/SIMGAN. Lianli Gao, Qike Zhao, Junchen Zhu, Sitong Su, Lechao Cheng, Lei Zhao 0017 |
IEEE Trans. Multim. | 6 |
| 2022 | Context Gating with Multi-Level Ranking Learning for Visual DialogabstractVisual dialog aims to answer several consecutive questions based on image and dialog history. Most works resolve all questions with ambiguous references (e.g., “she”) by dialog history, which generates redundant information and gets in-accurate results. Also, they regard this task as a classification task, which ignores the diversity of response answers and results in poor generalization capability. To tackle these problems, we propose a novel Context Gating with Multi-level Ranking Learning (CGMRL). Specifically, the proposed context gating considers both question and image to adaptively determine whether the history is needed for question answering, which reduces the redundant or even noisy information generated by history. To improve the generalization capability of the model, a new constrained multi-level ranking learning is proposed to encourage the model to consider the correct semantic options rather than only choose the ground truth answer. Experimental validations on the VisDial v1.0 show the superiority of the proposed method compared with other methods. Implementation code is published in anonymous Github: https://github.com/sy742/CGMRL_. Tangming Chen, Lianli Gao, Lei Zhao 0017, Jingkuan Song |
ICME | 4 |
| 2022 | X-HRNet: Towards Lightweight Human Pose Estimation with Spatially Unidimensional Self-AttentionabstractHigh-resolution representation is necessary for human pose estimation to achieve high performance, and the ensuing problem is high computational complexity. In particular, predominant pose estimation methods estimate human joints by 2D single-peak heatmaps. Each 2D heatmap can be hori-zontally and vertically projected to and reconstructed by a pair of 1D heat vectors. Inspired by this observation, we introduce a lightweight and powerful alternative, Spatially Unidimensional Self-Attention (SUSA), to the pointwise (1 x 1) convolution that is the main computational bottleneck in the depthwise separable 3 x 3 convolution. Our SUSA reduces the computational complexity of the pointwise (1 x 1) convolution by 96% without sacrificing accuracy. Furthermore, we use the SUSA as the main module to build our lightweight pose estimation backbone X-HRNet, where$X$represents the estimated cross-shape attention vectors. Extensive experiments on the COCO benchmark demonstrate the superiority of our X-HRNet, and comprehensive ablation studies show the effectiveness of the SUSA modules. The code is publicly available at https://github.com/cool-xuan/x-hrnet. Yixuan Zhou 0001, Xuanhan Wang, Xing Xu 0001, Lei Zhao 0017, Jingkuan Song |
ICME | 4 |
| 2021 | SKANet: Structured Knowledge-Aware Network for Visual DialogabstractVisual dialog aims to generate an answer to each question based on an image and dialog history. Despite recent progress, existing methods still undergo degradation on the condition of complex scenarios. Handling these scenarios depends on logical reasoning that requires common sense priors. In this paper, we propose a novel visual dialog pipeline, named Structured Knowledge-Aware Network (SKANet), consisting of a Multi-Modality Fusion Module, an Image Knowledge-Aware Module, and a Caption Knowledge-Aware Module. The Multi-Modality Fusion Module explores the textual context about the dialog history and visual content. To deal with the complex scenarios, the Image and Caption Knowledge-Aware Modules construct common sense knowledge graphs from ConceptNet. Experimental results on the VisDial v1.0 dataset show that our proposed method effectively outperforms comparative methods. Lei Zhao 0017, Lianli Gao, Yuyu Guo 0001, Jingkuan Song, Heng Tao Shen |
ICME | 1 |
| 2021 | Exploring Contextual-Aware Representation and Linguistic-Diverse Expression for Visual DialogabstractVisual dialog is a fundamental vision-language task where an AI agent holds a meaningful dialogue about visual content with humans in nature. However, this task remains challenging, since there is still no consensus way to capture rich visual contextual information contained in the environment rather than only focusing on visual objects. Furthermore, conventional methods suffer from the single-answer learning strategy, where it only accepts one correct answer without considering the diverse expressions of the language (i.e., one identical meaning but multiple expressions via rephrasing or adopting synonyms etc). In this paper, we introduce Contextual-Aware Representation and linguistic-diverse Expression (CARE), a novel plug-and-play framework with contextual-based graph embedding and curriculum contrastive learning to solve the above two issues. Specifically, the contextual-based graph embedding (CGE) module aims to integrate the environmental context information with visual objects to improve the answer quality. In addition, we propose a curriculum contrastive learning (CCL) paradigm to imitate the learning habits of humans when facing a question with multiple correct answers sharing the same meaning but with diverse expressions. To support CCL, a CCL loss is designed to progressively strengthen the model's ability in identifying the answers with correct semantics. Extensive experiments are conducted on two benchmark datasets, and our proposed method outperforms the state-of-the-arts by a considerable margin on VisDial V1.0 (4.63% NDCG) and VisDial V0.9 (1.27% MRR, 1.74% [email protected], 0.87% [email protected], 1.28% [email protected], 0.26 Mean. Lianli Gao, Lei Zhao 0017, Jingkuan Song |
ACM Multimedia | 3 |
| 2021 | Generalized pyramid co-attention with learnable aggregation net for video question answering
Lianli Gao, Tangming Chen, Pengpeng Zeng, Lei Zhao 0017, Yuan-Fang Li |
Pattern Recognit. | 5 |
| 2021 | GuessWhich? Visual dialog with attentive memory network
Lei Zhao 0017, Xinyu Lyu, Jingkuan Song, Lianli Gao |
Pattern Recognit. | 1 |