Zhengkun Zhang

dblp:218/0107 · DBLP profile ↗
← Back
6ranked-venue papers
3as first author
6since 2021 · last 2024
0000-0001-8822-2107ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 3 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 3 first-author · 4 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
4 papers
Question answering and dialogue systems · 42% Vision and language · 26% Knowledge representation and reasoning · 21%
Computer graphics and multimedia
3 papers
Multimedia analysis and retrieval · 100%
Databases, data mining, and information retrieval
1 paper
Information retrieval · 100%

Topics — the 7 heaviest of 11, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Multimedia analysis and retrieval
multimodal summarization
1.122022
UniMS: A Unified Framework for Multimodal Summarization with Knowledge Distillation · AAAI 2022
LAMS: A Location-aware Approach for Multimodal Summarization (Student Abstract) · AAAI 2021
Multimedia analysis and retrieval › multimedia analysis › multimedia collection analysis › image collection analysis
photo selection
0.722022
UniMS: A Unified Framework for Multimodal Summarization with Knowledge Distillation · AAAI 2022
News Content Completion with Location-Aware Image Selection · AAAI 2021
Natural language and speech › Question answering and dialogue systems › dialogue generation
empathetic response generation
0.612022
Multi-Party Empathetic Dialogue Generation: A New Task for Dialog Systems · ACL (1) 2022
Computer vision › Vision and language › multimodal understanding
multimodal machine comprehension
0.612022
Modeling Temporal-Modal Entity Graph for Procedural Multimodal Machine Comprehension · ACL (1) 2022
Natural language and speech › Question answering and dialogue systems
multi-party dialogue
0.612022
Multi-Party Empathetic Dialogue Generation: A New Task for Dialog Systems · ACL (1) 2022
Computer vision › Vision and language
multimodal fusion
0.112021
LAMS: A Location-aware Approach for Multimodal Summarization (Student Abstract) · AAAI 2021
Machine learning › Deep learning architectures and training
transformer
0.112021
LAMS: A Location-aware Approach for Multimodal Summarization (Student Abstract) · AAAI 2021

Methods — techniques the papers use, named apart from their topics

knowledge distillation · 1.1BART · 1.1transformer · 1.0multimodal fusion · 1.0location-aware image selection · 1.0vision-language pretrained model · 0.6vision-language pre-trained model · 0.6static-dynamic modeling · 0.6graph neural network · 0.6emotion analysis · 0.6
YearPublicationVenuePosition
2024 M3sum: A Novel Unsupervised Language-Guided Video Summarization
abstract
Language-guided video summarization empowers users to use natural language queries to effortlessly summarize lengthy videos into concise and relevant summaries that cater specifically to their information needs, which is more friendly to access and digest. However, most of the previous works rely on tremendous (also expensive) annotated videos and complex designs to align different modals at the feature level. In this paper, we first explore the combination of off-the-shelf models for each modal to solve the complex multi-modal problem by proposing a novel unsupervised language-guided video summarization method: Modular Multi-Modal Summarization (M3Sum), which does not require any training data or parameter updates. Specifically, instead of training an alignment module at the feature level, we convert all modal information (e.g. audio and frames) into textual descriptions and design a parameter-free alignment mechanism to fuse text descriptions from different modals. Benefiting from the remarkable long-context understanding capability of large language models (LLMs), our approach demonstrates comparable performance to most unsupervised methods and even outperforms certain supervised methods.
Hongru Wang 0003, Baohang Zhou, Zhengkun Zhang, Yiming Du, David Ho, Kam-Fai Wong
ICASSP3
2022 UniMS: A Unified Framework for Multimodal Summarization with Knowledge Distillation
abstract
With the rapid increase of multimedia data, a large body of literature has emerged to work on multimodal summarization, the majority of which target at refining salient information from textual and image modalities to output a pictorial summary with the most relevant images. Existing methods mostly focus on either extractive or abstractive summarization and rely on the presence and quality of image captions to build image references. We are the first to propose a Unified framework for Multimodal Summarization grounding on BART, UniMS, that integrates extractive and abstractive objectives, as well as selecting the image output. Specially, we adopt knowledge distillation from a vision-language pretrained model to improve image selection, which avoids any requirement on the existence and quality of image captions. Besides, we introduce a visual guided decoder to better integrate textual and visual modalities in guiding abstractive text generation. Results show that our best model achieves a new state-of-the-art result on a large-scale benchmark dataset. The newly involved extractive objective as well as the knowledge distillation technique are proven to bring a noticeable improvement to the multimodal summarization task.
Zhengkun Zhang, Xiaojun Meng, Yasheng Wang, Xin Jiang 0002, Qun Liu 0001, Zhenglu Yang
AAAI1
2022 Multi-Party Empathetic Dialogue Generation: A New Task for Dialog Systems
abstract
Empathetic dialogue assembles emotion understanding, feeling projection, and appropriate response generation.Existing work for empathetic dialogue generation concentrates on the two-party conversation scenario.Multiparty dialogues, however, are pervasive in reality.Furthermore, emotion and sensibility are typically confused; a refined empathy analysis is needed for comprehending fragile and nuanced human feelings.We address these issues by proposing a novel task called Multi-Party Empathetic Dialogue Generation in this study.Additionally, a Static-Dynamic model for Multi-Party Empathetic Dialogue Generation, SDMPED, is introduced as a baseline by exploring the static sensibility and dynamic emotion for the multi-party empathetic dialogue learning, the aspects that help SDMPED achieve the state-of-the-art performance.
Lingyu Zhu 0007, Zhengkun Zhang, Jun Wang 0023, Haiying Wu, Zhenglu Yang
ACL (1)2
2022 Modeling Temporal-Modal Entity Graph for Procedural Multimodal Machine Comprehension
abstract
Huibin Zhang, Zhengkun Zhang, Yao Zhang, Jun Wang, Yufan Li, Ning Jiang, Xin Wei, Zhenglu Yang. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022.
Huibin Zhang, Zhengkun Zhang, Jun Wang 0023, Zhenglu Yang
ACL (1)2
2021 News Content Completion with Location-Aware Image Selection
Zhengkun Zhang, Jun Wang 0023, Adam Jatowt, Zhe Sun 0009, Shao-Ping Lu, Zhenglu Yang
AAAI1
2021 LAMS: A Location-aware Approach for Multimodal Summarization (Student Abstract)
abstract
Multimodal summarization aims to refine salient information from multiple modalities, among which texts and images are two mostly discussed ones. In recent years, many fantastic works have emerged in this field by modeling image-text interactions; however, they neglect the fact that most of multimodal documents have been elaborately organized by their writers. This means that a critical organized factor has long been short of enough attention, that is, image locations, which may carry illuminating information and imply the key contents of a document. To address this issue, we propose a location-aware approach for multimodal summarization (LAMS) based on Transformer. We investigate image locations for multimodal summarization via a stack of multimodal fusion block, which can formulate the high-order interactions among images and texts. An extensive experimental study on an extended multimodal dataset validates the superior summarization performance of the proposed model.
Zhengkun Zhang, Jun Wang 0023, Zhe Sun 0009, Zhenglu Yang
AAAI1