VLDB 2026 Research / reviewers in the wild / expert
Zhengkun Zhang
dblp:218/0107
· DBLP profile ↗
6ranked-venue papers
3as first author
6since 2021 · last 2024
0000-0001-8822-2107ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 5 · 3 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 3 first-author · 4 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
4 papers |
Question answering and dialogue systems · 42% Vision and language · 26% Knowledge representation and reasoning · 21% | |
| Computer graphics and multimedia
3 papers |
Multimedia analysis and retrieval · 100% | |
| Databases, data mining, and information retrieval
1 paper |
Information retrieval · 100% |
Topics — the 7 heaviest of 11, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Multimedia analysis and retrieval
multimodal summarization |
1.1 | 2 | 2022 | UniMS: A Unified Framework for Multimodal Summarization with Knowledge Distillation · AAAI 2022 LAMS: A Location-aware Approach for Multimodal Summarization (Student Abstract) · AAAI 2021 |
Multimedia analysis and retrieval › multimedia analysis › multimedia collection analysis › image collection analysis
photo selection |
0.7 | 2 | 2022 | UniMS: A Unified Framework for Multimodal Summarization with Knowledge Distillation · AAAI 2022 News Content Completion with Location-Aware Image Selection · AAAI 2021 |
Natural language and speech › Question answering and dialogue systems › dialogue generation
empathetic response generation |
0.6 | 1 | 2022 | Multi-Party Empathetic Dialogue Generation: A New Task for Dialog Systems · ACL (1) 2022 |
Computer vision › Vision and language › multimodal understanding
multimodal machine comprehension |
0.6 | 1 | 2022 | Modeling Temporal-Modal Entity Graph for Procedural Multimodal Machine Comprehension · ACL (1) 2022 |
Natural language and speech › Question answering and dialogue systems
multi-party dialogue |
0.6 | 1 | 2022 | Multi-Party Empathetic Dialogue Generation: A New Task for Dialog Systems · ACL (1) 2022 |
Computer vision › Vision and language
multimodal fusion |
0.1 | 1 | 2021 | LAMS: A Location-aware Approach for Multimodal Summarization (Student Abstract) · AAAI 2021 |
Machine learning › Deep learning architectures and training
transformer |
0.1 | 1 | 2021 | LAMS: A Location-aware Approach for Multimodal Summarization (Student Abstract) · AAAI 2021 |
Methods — techniques the papers use, named apart from their topics
knowledge distillation · 1.1BART · 1.1transformer · 1.0multimodal fusion · 1.0location-aware image selection · 1.0vision-language pretrained model · 0.6vision-language pre-trained model · 0.6static-dynamic modeling · 0.6graph neural network · 0.6emotion analysis · 0.6
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | M3sum: A Novel Unsupervised Language-Guided Video SummarizationabstractLanguage-guided video summarization empowers users to use natural language queries to effortlessly summarize lengthy videos into concise and relevant summaries that cater specifically to their information needs, which is more friendly to access and digest. However, most of the previous works rely on tremendous (also expensive) annotated videos and complex designs to align different modals at the feature level. In this paper, we first explore the combination of off-the-shelf models for each modal to solve the complex multi-modal problem by proposing a novel unsupervised language-guided video summarization method: Modular Multi-Modal Summarization (M3Sum), which does not require any training data or parameter updates. Specifically, instead of training an alignment module at the feature level, we convert all modal information (e.g. audio and frames) into textual descriptions and design a parameter-free alignment mechanism to fuse text descriptions from different modals. Benefiting from the remarkable long-context understanding capability of large language models (LLMs), our approach demonstrates comparable performance to most unsupervised methods and even outperforms certain supervised methods. Hongru Wang 0003, Baohang Zhou, Zhengkun Zhang, Yiming Du, David Ho, Kam-Fai Wong |
ICASSP | 3 |
| 2022 | UniMS: A Unified Framework for Multimodal Summarization with Knowledge DistillationabstractWith the rapid increase of multimedia data, a large body of literature has emerged to work on multimodal summarization, the majority of which target at refining salient information from textual and image modalities to output a pictorial summary with the most relevant images. Existing methods mostly focus on either extractive or abstractive summarization and rely on the presence and quality of image captions to build image references. We are the first to propose a Unified framework for Multimodal Summarization grounding on BART, UniMS, that integrates extractive and abstractive objectives, as well as selecting the image output. Specially, we adopt knowledge distillation from a vision-language pretrained model to improve image selection, which avoids any requirement on the existence and quality of image captions. Besides, we introduce a visual guided decoder to better integrate textual and visual modalities in guiding abstractive text generation. Results show that our best model achieves a new state-of-the-art result on a large-scale benchmark dataset. The newly involved extractive objective as well as the knowledge distillation technique are proven to bring a noticeable improvement to the multimodal summarization task. Zhengkun Zhang, Xiaojun Meng, Yasheng Wang, Xin Jiang 0002, Qun Liu 0001, Zhenglu Yang |
AAAI | 1 |
| 2022 | Multi-Party Empathetic Dialogue Generation: A New Task for Dialog SystemsabstractEmpathetic dialogue assembles emotion understanding, feeling projection, and appropriate response generation.Existing work for empathetic dialogue generation concentrates on the two-party conversation scenario.Multiparty dialogues, however, are pervasive in reality.Furthermore, emotion and sensibility are typically confused; a refined empathy analysis is needed for comprehending fragile and nuanced human feelings.We address these issues by proposing a novel task called Multi-Party Empathetic Dialogue Generation in this study.Additionally, a Static-Dynamic model for Multi-Party Empathetic Dialogue Generation, SDMPED, is introduced as a baseline by exploring the static sensibility and dynamic emotion for the multi-party empathetic dialogue learning, the aspects that help SDMPED achieve the state-of-the-art performance. Lingyu Zhu 0007, Zhengkun Zhang, Jun Wang 0023, Haiying Wu, Zhenglu Yang |
ACL (1) | 2 |
| 2022 | Modeling Temporal-Modal Entity Graph for Procedural Multimodal Machine ComprehensionabstractHuibin Zhang, Zhengkun Zhang, Yao Zhang, Jun Wang, Yufan Li, Ning Jiang, Xin Wei, Zhenglu Yang. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022. Huibin Zhang, Zhengkun Zhang, Jun Wang 0023, Zhenglu Yang |
ACL (1) | 2 |
| 2021 | News Content Completion with Location-Aware Image Selection
Zhengkun Zhang, Jun Wang 0023, Adam Jatowt, Zhe Sun 0009, Shao-Ping Lu, Zhenglu Yang |
AAAI | 1 |
| 2021 | LAMS: A Location-aware Approach for Multimodal Summarization (Student Abstract)abstractMultimodal summarization aims to refine salient information from multiple modalities, among which texts and images are two mostly discussed ones. In recent years, many fantastic works have emerged in this field by modeling image-text interactions; however, they neglect the fact that most of multimodal documents have been elaborately organized by their writers. This means that a critical organized factor has long been short of enough attention, that is, image locations, which may carry illuminating information and imply the key contents of a document. To address this issue, we propose a location-aware approach for multimodal summarization (LAMS) based on Transformer. We investigate image locations for multimodal summarization via a stack of multimodal fusion block, which can formulate the high-order interactions among images and texts. An extensive experimental study on an extended multimodal dataset validates the superior summarization performance of the proposed model. Zhengkun Zhang, Jun Wang 0023, Zhe Sun 0009, Zhenglu Yang |
AAAI | 1 |