Junnan Zhu

dblp:205/8977 · DBLP profile ↗
← Back
34ranked-venue papers
8as first author
22since 2021 · last 2026
0000-0002-9856-2946ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 32 · 8 first-author · 21 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 MIRAGE: Scaling Test-Time Inference with Parallel Graph-Retrieval-Augmented Reasoning Chains
abstract
Large reasoning models (LRMs) have shown significant progress in test-time scaling through chain-of-thought prompting. Current approaches like search-o1 integrate retrieval augmented generation (RAG) into multi-step reasoning processes but rely on a single, linear reasoning path while incorporating unstructured textual information in a flat, context-agnostic manner. As a result, these approaches can lead to error accumulation throughout the reasoning chain, which significantly limits its effectiveness in medical question-answering (QA) tasks where both accuracy and traceability are critical requirements. To address these challenges, we propose MIRAGE (Multi-path Inference with Retrieval-Augmented Graph Exploration), a novel test-time scalable reasoning framework that performs dynamic multi-path inference over structured medical knowledge graphs. Specifically, MIRAGE 1) decomposes complex queries into entity-grounded sub-questions, 2) executes parallel inference paths, 3) retrieves evidence adaptively via neighbor expansion and multi-hop traversal, and 4) integrates answers using cross-path verification to resolve contradictions. Experiments on three medical QA benchmarks (GenMedGPT-5k, CMCQA, and ExplainCPE) show that MIRAGE consistently outperforms GPT-4o, Tree-of-Thought variants, and other retrieval-augmented baselines in both automatic and human evaluations. Additionally, MIRAGE improves interpretability by generating explicit reasoning chains that trace each factual claim to concrete paths within the knowledge graph, making it especially suitable for complex medical reasoning scenarios.
Kaiwen Wei, Rui Shan, Dongsheng Zou, Jianzhong Yang, Bi Zhao, Junnan Zhu
AAAI6
2026 FocalOrder: Focal Preference Optimization for Reading Order Detection
abstract
Fuyuan Liu, Dianyu Yu, He Ren, Nayu Liu, Xiaomian Kang, Delai Qiu, Fa Zhang, Genpeng Zhen, Shengping Liu, Liang Jiaen, Weihuang, Yining Wang, Junnan Zhu. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Fuyuan Liu, Dianyu Yu, Nayu Liu, Xiaomian Kang, Delai Qiu, Genpeng Zhen, Shengping Liu, Jiaen Liang, Junnan Zhu
ACL (1)13
2026 MentalSeek-Dx: Towards Progressive Hypothetico-Deductive Reasoning for Real-world Psychiatric Diagnosis
abstract
Xiao Sun, Ymyang, Xinyi Jiang, Yu Tian, Junnan Zhu, Jiang Zhong, Qin Lei, Jingwang Huang, Haoyang Zeng, Xinyu Zhou, Xin Xiao, Kaiwen Wei. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Junnan Zhu, Qin Lei, Jingwang Huang, Kaiwen Wei
ACL (1)5
2026 From Past To Path: Masked History Learning for Next-Item Prediction in Generative Recommendation
abstract
Kaiwen Wei, Kejun he, Xiaomian Kang, Jie Zhang, Ymyang, Li Jin, Zhenyang Li, Jiang Zhong, Richard He Bai, Junnan Zhu. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Kaiwen Wei, Kejun He, Xiaomian Kang, Ymyang, He Bai 0002, Junnan Zhu
ACL (1)10
2026 GenProve: Learning to Generate Text with Fine-Grained Provenance
abstract
Jingxuan Wei, Xingyue Wang, Yanghaoyu Liao, Jie Dong, Yuchen Liu, Caijun Jia, Bihui Yu, Junnan Zhu. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Jingxuan Wei, Yanghaoyu Liao, Caijun Jia, Bihui Yu, Junnan Zhu
ACL (1)8
2026 Deciphering spatial heterogeneity by multimodal spatial transcriptomics modelling with SpatialModal
abstract
MOTIVATION: Advances in spatial transcriptomics (ST) technologies have made it possible to jointly acquire gene expression and histological image information while preserving spatial coordinates. This breakthrough presents unprecedented opportunities for the precise dissection of spatial heterogeneity in complex tissues. However, existing computational methods remain limited in their capacity for effective integration and synergistic modelling of multimodal ST data. RESULTS: We propose SpatialModal, a multimodal graph learning framework that learns robust joint representations by combining a hierarchical representation strategy with a dual-level contrastive learning mechanism. We perform extensive validation of SpatialModal across diverse ST datasets spanning human and mouse tissues. The results demonstrate that SpatialModal effectively reveals intricate brain architectures in humans and mice, dissects tumour microenvironment heterogeneity in breast cancer, delineates Alzheimer's disease patterns, and characterizes spatiotemporal developmental trajectories within the embryonic heart, underscoring its capability to decipher the spatial heterogeneity of biological tissues. Furthermore, SpatialModal exhibits remarkable versatility and robustness, maintaining superior efficacy even on unimodal datasets devoid of histological images, thereby ensuring its broad applicability across diverse ST platforms. AVAILABILITY AND IMPLEMENTATION: SpatialModal is implemented in Python and is freely available at https://github.com/xingyili/SpatialModal. The source code used in this study has been archived on Zenodo at DOI: https://doi.org/10.5281/zenodo.21264356. All datasets used in this study are publicly available at https://doi.org/10.5281/zenodo.18220735.
Xingyi Li 0003, Dongmin Zhao, Xiangting Jia, Gaoyuan Du, Jialuo Xu, Yingfu Wu, Junnan Zhu, Xuequn Shang 0001
Bioinform.10
2025 SARA: Salience-Aware Reinforced Adaptive Decoding for Large Language Models in Abstractive Summarization
abstract
Nayu Liu, Junnan Zhu, Yiming Ma, Zhicong Lu, Wenlei Xu, Yong Yang, Jiang Zhong, Kaiwen Wei. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Nayu Liu, Junnan Zhu, Zhicong Lu, Wenlei Xu, Yong Yang 0001, Kaiwen Wei
ACL (1)2
2025 TROVE: A Challenge for Fine-Grained Text Provenance via Source Sentence Tracing and Relationship Classification
abstract
LLMs have achieved remarkable fluency and coherence in text generation, yet their widespread adoption has raised concerns about content reliability and accountability. In high-stakes domains, it is crucial to understand where and how the content is created. To address this, we introduce the Text pROVEnance (TROVE) challenge, designed to trace each sentence of a target text back to specific source sentences within potentially lengthy or multi-document inputs. Beyond identifying sources, TROVE annotates the fine-grained relationships (quotation, compression, inference, and others), providing a deep understanding of how each target sentence is formed.To benchmark TROVE, we construct our dataset by leveraging three public datasets covering 11 diverse scenarios (e.g., QA and summarization) in English and Chinese, spanning source texts of varying lengths (0–5k, 5–10k, 10k+), emphasizing the multi-document and long-document settings essential for provenance. To ensure high-quality data, we employ a three-stage annotation process: sentence retrieval, GPT-4o provenance, and human provenance. We evaluate 11 LLMs under direct prompting and retrieval-augmented paradigms, revealing that retrieval is essential for robust performance, larger models perform better in complex relationship classification, and closed-source models often lead, yet open-source models show significant promise, particularly with retrieval augmentation. We make our dataset available here: https://github.com/ZNLP/ZNLP-Dataset.
Junnan Zhu, Feifei Zhai, Chengqing Zong
ACL (1)1
2025 ChartMind: A Comprehensive Benchmark for Complex Real-world Multimodal Chart Question Answering
abstract
Chart question answering (CQA) has become a critical multimodal task for evaluating the reasoning capabilities of vision-language models.While early approaches have shown promising performance by focusing on visual features or leveraging large-scale pre-training, most existing evaluations rely on rigid output formats and objective metrics, thus ignoring the complex, real-world demands of practical chart analysis.In this paper, we introduce ChartMind, a new benchmark designed for complex CQA tasks in real-world settings.ChartMind covers seven task categories, incorporates multilingual contexts, supports open-domain textual outputs, and accommodates diverse chart formats, bridging the gap between real-world applications and traditional academic benchmarks.Furthermore, we propose a context-aware yet modelagnostic framework, ChartLLM, that focuses on extracting key contextual elements, reducing noise, and enhancing the reasoning accuracy of multimodal large language models.Extensive evaluations on ChartMind and three representative public benchmarks with 14 mainstream multimodal models show our framework significantly outperforms the previous three common CQA paradigms: instruction-following, OCRenhanced, and chain-of-thought, highlighting the importance of flexible chart understanding for real-world CQA.These findings suggest new directions for developing more robust chart reasoning in future research.
Jingxuan Wei, Junnan Zhu, Haoyanni, Bihui Yu
EMNLP3
2025 TableEval: A Real-World Benchmark for Complex, Multilingual, and Multi-Structured Table Question Answering
abstract
LLMs have shown impressive progress in natural language processing.However, they still face significant challenges in TableQA, where real-world complexities such as diverse table structures, multilingual data, and domain-specific reasoning are crucial.Existing TableQA benchmarks are often limited by their focus on simple flat tables and suffer from data leakage.Furthermore, most benchmarks are monolingual and fail to capture the crosslingual and cross-domain variability in practical applications.To address these limitations, we introduce TableEval, a new benchmark designed to evaluate LLMs on realistic TableQA tasks.Specifically, TableEval includes tables with various structures (such as concise, hierarchical, and nested tables) collected from four domains (including government, finance, academia, and industry reports).Additionally, TableEval features cross-lingual scenarios with tables in Simplified Chinese, Traditional Chinese, and English.To reduce potential data leakage, we curate data from recent real-world documents.Considering that existing TableQA metrics fail to capture semantic accuracy, we further propose SEAT, a new evaluation framework that assesses the alignment between model responses and reference answers at the sub-question level.Experimental results have shown that SEAT achieves high agreement with human judgment.Extensive experiments on TableEval reveal critical gaps in the ability of state-of-the-art LLMs to handle these complex, real-world TableQA tasks, offering insights for future improvements.We make our dataset available here:
Junnan Zhu, Bohan Yu
EMNLP1
2025 Pay More Attention to Images: Numerous Images-Oriented Multimodal Summarization
abstract
Min Xiao, Junnan Zhu, Feifei Zhai, Chengqing Zong, Yu Zhou. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Junnan Zhu, Feifei Zhai, Chengqing Zong, Yu Zhou 0001
NAACL (Long Papers)2
2024 DIUSum: Dynamic Image Utilization for Multimodal Summarization
abstract
Existing multimodal summarization approaches focus on fusing image features in the encoding process, ignoring the individualized needs for images when generating different summaries. However, whether intuitively or empirically, not all images can improve summary quality. Therefore, we propose a novel Dynamic Image Utilization framework for multimodal Summarization (DIUSum) to select and utilize valuable images for summarization. First, to predict whether an image helps produce a high-quality summary, we propose an image selector to score the usefulness of each image. Second, to dynamically utilize the multimodal information, we incorporate the hard and soft guidance from the image selector. Under the guidance, the image information is plugged into the decoder to generate a summary. Experimental results have shown that DIUSum outperforms multiple strong baselines and achieves SOTA on two public multimodal summarization datasets. Further analysis demonstrates that the image selector can reflect the improved level of summary quality brought by the images.
Junnan Zhu, Feifei Zhai, Yu Zhou 0001, Chengqing Zong
AAAI2
2024 A Hybrid Approach towards Chinese Spelling and Splitting Error Correction
abstract
Existing Chinese spelling check (CSC) methods have limitations in correcting variable-length error characters, requiring the input and output to be the same length. They mainly focus on modelling Chinese characters’ phonetic information and generating candidates for each position. In contrast, few approaches delve into the intricacies of splitting Chinese characters to address glyph errors and splitting variable-length corrections. We define the Chinese Splitting Error Correction (CSEC) task and develop CSEC datasets in news and social media domains to address this issue. We then propose Soft-Masked Multi-feature Error Correction (SoMu) model, which first generates semantic, phonetic, graphic, and unique Chinese Wubi embeddings, then integrates those features through selective gating fusion, followed by a soft-mask strategy to filter incorrect tokens and finally use transformer layers to predict the correct ones. This model effectively addresses both spelling and splitting errors. Extensive analysis shows that our model significantly improves character-splitting information modelling for CSEC. Our dataset is available at https://github.com/Skywalker-Harrison/SoMu.
Junhong Liang, Junnan Zhu, Feifei Zhai, Nanchang Cheng, Chengqing Zong, Yu Zhou 0001
ECAI2
2023 CFSum Coarse-to-Fine Contribution Network for Multimodal Summarization
abstract
Multimodal summarization usually suffers from the problem that the contribution of the visual modality is unclear.Existing multimodal summarization approaches focus on designing the fusion methods of different modalities, while ignoring the adaptive conditions under which visual modalities are useful.Therefore, we propose a novel Coarse-to-Fine contribution network for multimodal Summarization (CFSum) to consider different contributions of images for summarization.First, to eliminate the interference of useless images, we propose a pre-filter module to abandon useless images.Second, to make accurate use of useful images, we propose two levels of visual complement modules, word level and phrase level.Specifically, image contributions are calculated and are adopted to guide the attention of both textual and visual modalities.Experimental results have shown that CFSum significantly outperforms multiple strong baselines on the standard benchmark.Furthermore, the analysis verifies that useful images can even help generate nonvisual words which are implicitly represented in the image 1 .
Junnan Zhu, Haitao Lin 0001, Yu Zhou 0001, Chengqing Zong
ACL (1)2
2023 Zero-shot language extension for dialogue state tracking via pre-trained models and multi-auxiliary-tasks fine-tuning
Lu Xiang, Yang Zhao 0007, Junnan Zhu, Yu Zhou 0001, Chengqing Zong
Knowl. Based Syst.3
2023 Topic-Oriented Dialogue Summarization
abstract
A multi-turn dialogue often contains multiple discussion topics. In several scenarios (e.g., customer service dispute, public opinion monitoring), people are only interested in the gist of a specific topic in the dialogue. Therefore, we propose a novel summarization task, i.e., Topic-Oriented Dialogue Summarization (TODS). Given a dialogue with a topic label, TODS aims to produce a summary covering the main content of the given topic in the dialogue. To model the relationship between dialogues and topics, three key abilities are needed for TODS: (1) Learning the semantic information of different topics. (2) Locating the topic-related content in the dialogue. (3) Distinguishing summaries for different topics in the same dialogue. Thus, we propose three topic-related auxiliary tasks to make the summarization model learn the three abilities above. First, the topic identification task aims at generating all the topics in the dialogue. Second, the topic attention restriction task tries to constrain the attention distribution on topic-related utterances. Third, the topic summary distinguishing task focuses on increasing the difference of summaries for different topics in the same dialogue. Experimental results on two public TODS datasets show that all auxiliary tasks are critical for TODS and help generate high-quality summaries. We also point out the expansions and challenges in TODS for future research.
Haitao Lin 0001, Junnan Zhu, Lu Xiang, Feifei Zhai, Yu Zhou 0001, Jiajun Zhang 0001, Chengqing Zong
IEEE ACM Trans. Audio Speech Lang. Process.2
2022 Other Roles Matter! Enhancing Role-Oriented Dialogue Summarization via Role Interactions
abstract
Role-oriented dialogue summarization is to generate summaries for different roles in the dialogue, e.g., merchants and consumers.Existing methods handle this task by summarizing each role's content separately and thus are prone to ignore the information from other roles.However, we believe that other roles' content could benefit the quality of summaries, such as the omitted information mentioned by other roles.Therefore, we propose a novel role interaction enhanced method for role-oriented dialogue summarization.It adopts cross attention and decoder self-attention interactions to interactively acquire other roles' critical information.The cross attention interaction aims to select other roles' critical dialogue utterances, while the decoder self-attention interaction aims to obtain key information from other roles' summaries.Experimental results have shown that our proposed method significantly outperforms strong baselines on two public role-oriented dialogue summarization datasets.Extensive analyses have demonstrated that other roles' content could help generate summaries with more complete semantics and correct topic structures. 1
Haitao Lin 0001, Junnan Zhu, Lu Xiang, Yu Zhou 0001, Jiajun Zhang 0001, Chengqing Zong
ACL (1)2
2021 CSDS: A Fine-Grained Chinese Dataset for Customer Service Dialogue Summarization
abstract
Dialogue summarization has drawn much attention recently.Especially in the customer service domain, agents could use dialogue summaries to help boost their works by quickly knowing customer's issues and service progress.These applications require summaries to contain the perspective of a single speaker and have a clear topic flow structure, while neither are available in existing datasets.Therefore, in this paper, we introduce a novel Chinese dataset for Customer Service Dialogue Summarization (CSDS).CSDS improves the abstractive summaries in two aspects: (1) In addition to the overall summary for the whole dialogue, role-oriented summaries are also provided to acquire different speakers' viewpoints.(2) All the summaries sum up each topic separately, thus containing the topic-level structure of the dialogue.We define tasks in CSDS as generating the overall summary and different role-oriented summaries for a given dialogue.Next, we compare various summarization methods on CSDS, and experiment results show that existing methods are prone to generate redundant and incoherent summaries.Besides, the performance becomes much worse when analyzing the performance on role-oriented summaries and topic structures.We hope that this study could benchmark Chinese dialogue summarization and benefit further studies.
Haitao Lin 0001, Liqun Ma, Junnan Zhu, Lu Xiang, Yu Zhou 0001, Jiajun Zhang 0001, Chengqing Zong
EMNLP (1)3
2021 Learning N: M Fine-grained Structured Sparse Neural Networks From Scratch
Aojun Zhou, Junnan Zhu, Wenxiu Sun, Hongsheng Li 0001
ICLR3
2021 Zero-Shot Deployment for Cross-Lingual Dialogue System
Lu Xiang, Yang Zhao 0007, Junnan Zhu, Yu Zhou 0001, Chengqing Zong
NLPCC (2)3
2021 Robust Cross-lingual Task-oriented Dialogue
abstract
Cross-lingual dialogue systems are increasingly important in e-commerce and customer service due to the rapid progress of globalization. In real-world system deployment, machine translation (MT) services are often used before and after the dialogue system to bridge different languages. However, noises and errors introduced in the MT process will result in the dialogue system's low robustness, making the system's performance far from satisfactory. In this article, we propose a novel MT-oriented noise enhanced framework that exploits multi-granularity MT noises and injects such noises into the dialogue system to improve the dialogue system's robustness. Specifically, we first design a method to automatically construct multi-granularity MT-oriented noises and multi-granularity adversarial examples, which contain abundant noise knowledge oriented to MT. Then, we propose two strategies to incorporate the noise knowledge: (i) Utterance-level adversarial learning and (ii) Knowledge-level guided method. The former adopts adversarial learning to learn a perturbation-invariant encoder, guiding the dialogue system to learn noise-independent hidden representations. The latter explicitly incorporates the multi-granularity noises, which contain the noise tokens and their possible correct forms, into the training and inference process, thus improving the dialogue system's robustness. Experimental results on three dialogue models, two dialogue datasets, and two language pairs have shown that the proposed framework significantly improves the performance of the cross-lingual dialogue system.
Lu Xiang, Junnan Zhu, Yang Zhao 0007, Yu Zhou 0001, Chengqing Zong
ACM Trans. Asian Low Resour. Lang. Inf. Process.2
2021 Graph-based Multimodal Ranking Models for Multimodal Summarization
abstract
Multimodal summarization aims to extract the most important information from the multimedia input. It is becoming increasingly popular due to the rapid growth of multimedia data in recent years. There are various researches focusing on different multimodal summarization tasks. However, the existing methods can only generate single-modal output or multimodal output. In addition, most of them need a lot of annotated samples for training, which makes it difficult to be generalized to other tasks or domains. Motivated by this, we propose a unified framework for multimodal summarization that can cover both single-modal output summarization and multimodal output summarization. In our framework, we consider three different scenarios and propose the respective unsupervised graph-based multimodal summarization models without the requirement of any manually annotated document-summary pairs for training: (1) generic multimodal ranking, (2) modal-dominated multimodal ranking, and (3) non-redundant text-image multimodal ranking. Furthermore, an image-text similarity estimation model is introduced to measure the semantic similarity between image and text. Experiments show that our proposed models outperform the single-modal summarization methods on both automatic and human evaluation metrics. Besides, our models can also improve the single-modal summarization with the guidance of the multimedia information. This study can be applied as the benchmark for further study on multimodal summarization task.
Junnan Zhu, Lu Xiang, Yu Zhou 0001, Jiajun Zhang 0001, Chengqing Zong
ACM Trans. Asian Low Resour. Lang. Inf. Process.1
2020 Keywords-Guided Abstractive Sentence Summarization
abstract
We study the problem of generating a summary for a given sentence. Existing researches on abstractive sentence summarization ignore that keywords in the input sentence provide significant clues for valuable content, and humans tend to write summaries covering these keywords. In this paper, we propose an abstractive sentence summarization method by applying guidance signals of keywords to both the encoder and the decoder in the sequence-to-sequence model. A multi-task learning framework is adopted to jointly learn to extract keywords and generate a summary for the input sentence. We apply keywords-guided selective encoding strategies to filter source information by investigating the interactions between the input sentence and the keywords. We extend pointer-generator network by a dual-attention and a dual-copy mechanism, which can integrate the semantics of the input sentence and the keywords, and copy words from both the input sentence and the keywords. We demonstrate that multi-task learning and keywords-oriented guidance facilitate sentence summarization task, achieving better performance than the competitive models on the English Gigaword sentence summarization dataset.
Haoran Li 0001, Junnan Zhu, Jiajun Zhang 0001, Chengqing Zong, Xiaodong He 0001
AAAI2
2020 Multimodal Summarization with Guidance of Multimodal Reference
abstract
Multimodal summarization with multimodal output (MSMO) is to generate a multimodal summary for a multimodal news report, which has been proven to effectively improve users' satisfaction. The existing MSMO methods are trained by the target of text modality, leading to the modality-bias problem that ignores the quality of model-selected image during training. To alleviate this problem, we propose a multimodal objective function with the guidance of multimodal reference to use the loss from the summary generation and the image selection. Due to the lack of multimodal reference data, we present two strategies, i.e., ROUGE-ranking and Order-ranking, to construct the multimodal reference by extending the text reference. Meanwhile, to better evaluate multimodal outputs, we propose a novel evaluation metric based on joint multimodal representation, projecting the model output and multimodal reference into a joint semantic space during evaluation. Experimental results have shown that our proposed model achieves the new state-of-the-art on both automatic and manual evaluation metrics. Besides, our proposed evaluation method can effectively improve the correlation with human judgments.
Junnan Zhu, Yu Zhou 0001, Jiajun Zhang 0001, Haoran Li 0001, Chengqing Zong, Changliang Li
AAAI1
2020 Attend, Translate and Summarize: An Efficient Method for Neural Cross-Lingual Summarization
abstract
Cross-lingual summarization aims at summarizing a document in one language (e.g., Chinese) into another language (e.g., English).In this paper, we propose a novel method inspired by the translation pattern in the process of obtaining a cross-lingual summary.We first attend to some words in the source text, then translate them into the target language, and summarize to get the final summary.Specifically, we first employ the encoder-decoder attention distribution to attend to the source words.Second, we present three strategies to acquire the translation probability, which helps obtain the translation candidates for each source word.Finally, each summary word is generated either from the neural distribution or from the translation candidates of source words.Experimental results on Chinese-to-English and English-to-Chinese summarization tasks have shown that our proposed method can significantly outperform the baselines, achieving comparable performance with the state-of-the-art.
Junnan Zhu, Yu Zhou 0001, Jiajun Zhang 0001, Chengqing Zong
ACL1
2020 Multimodal Sentence Summarization via Multimodal Selective Encoding
abstract
This paper studies the problem of generating a summary for a given sentence-image pair.Existing multimodal sequence-to-sequence approaches mainly focus on enhancing the decoder by visual signals, while ignoring that the image can improve the ability of the encoder to identify highlights of a news event or a document.Thus, we propose a multimodal selective gate network that considers reciprocal relationships between textual and multi-level visual features, including global image descriptor, activation grids, and object proposals, to select highlights of the event when encoding the source sentence.In addition, we introduce a modality regularization to encourage the summary to capture the highlights embedded in the image more accurately.To verify the generalization of our model, we adopt the multimodal selective gate to the text-based decoder and multimodal-based decoder.Experimental results on a public multimodal sentence summarization dataset demonstrate the advantage of our models over baselines.Further analysis suggests that our proposed multimodal selective gate network can effectively select important information in the input sentence.
Haoran Li 0001, Junnan Zhu, Jiajun Zhang 0001, Xiaodong He 0001, Chengqing Zong
COLING2
2020 Knowledge Graph Enhanced Neural Machine Translation via Multi-task Learning on Sub-entity Granularity
abstract
Previous studies combining knowledge graph (KG) with neural machine translation (NMT) have two problems: i) Knowledge under-utilization: they only focus on the entities that appear in both KG and training sentence pairs, making much knowledge in KG unable to be fully utilized.ii) Granularity mismatch: the current KG methods utilize the entity as the basic granularity, while NMT utilizes the sub-word as the granularity, making the KG different to be utilized in NMT.To alleviate above problems, we propose a multi-task learning method on sub-entity granularity.Specifically, we first split the entities in KG and sentence pairs into sub-entity granularity by using joint BPE.Then we utilize the multi-task learning to combine the machine translation task and knowledge reasoning task.The extensive experiments on various translation tasks have demonstrated that our method significantly outperforms the baseline models in both translation quality and handling the entities.
Yang Zhao 0007, Lu Xiang, Junnan Zhu, Jiajun Zhang 0001, Yu Zhou 0001, Chengqing Zong
COLING3
2019 NCLS: Neural Cross-Lingual Summarization
abstract
Junnan Zhu, Qian Wang, Yining Wang, Yu Zhou, Jiajun Zhang, Shaonan Wang, Chengqing Zong. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Junnan Zhu, Qian Wang 0061, Yu Zhou 0001, Jiajun Zhang 0001, Shaonan Wang, Chengqing Zong
EMNLP/IJCNLP (1)1
2019 Read, Watch, Listen, and Summarize: Multi-Modal Summarization for Asynchronous Text, Image, Audio and Video
abstract
Automatic text summarization is a fundamental natural language processing (NLP) application that aims to condense a source text into a shorter version. The rapid increase in multimedia data transmission over the Internet necessitates multi-modal summarization (MMS) from asynchronous collections of text, image, audio, and video. In this work, we propose an extractive MMS method that unites the techniques of NLP, speech processing, and computer vision to explore the rich information contained in multi-modal data and to improve the quality of multimedia news summarization. The key idea is to bridge the semantic gaps between multi-modal content. Audio and visual are main modalities in the video. For audio information, we design an approach to selectively use its transcription and to infer the salience of the transcription with audio signals. For visual information, we learn the joint representations of text and images using a neural network. Then, we capture the coverage of the generated summary for important visual information through text-image matching or multi-modal topic modeling. Finally, all the multi-modal aspects are considered to generate a textual summary by maximizing the salience, non-redundancy, readability, and coverage through the budgeted optimization of submodular functions. We further introduce a publicly available MMS corpus in English and Chinese.1 The experimental results obtained on our dataset demonstrate that our methods based on image matching and image topic framework outperform other competitive baseline methods.
Haoran Li 0001, Junnan Zhu, Cong Ma 0002, Jiajun Zhang 0001, Chengqing Zong
IEEE Trans. Knowl. Data Eng.2
2018 Ensure the Correctness of the Summary: Incorporate Entailment Knowledge into Abstractive Sentence Summarization
abstract
In this paper, we investigate the sentence summarization task that produces a summary from a source sentence. Neural sequence-to-sequence models have gained considerable success for this task, while most existing approaches only focus on improving the informativeness of the summary, which ignore the correctness, i.e., the summary should not contain unrelated information with respect to the source sentence. We argue that correctness is an essential requirement for summarization systems. Considering a correct summary is semantically entailed by the source sentence, we incorporate entailment knowledge into abstractive summarization models. We propose an entailment-aware encoder under multi-task framework (i.e., summarization generation and entailment recognition) and an entailment-aware decoder by entailment Reward Augmented Maximum Likelihood (RAML) training. Experiment results demonstrate that our models significantly outperform baselines from the aspects of informativeness and correctness.
Haoran Li 0001, Junnan Zhu, Jiajun Zhang 0001, Chengqing Zong
COLING2
2018 MSMO: Multimodal Summarization with Multimodal Output
abstract
Multimodal summarization has drawn much attention due to the rapid growth of multimedia data.The output of the current multimodal summarization systems is usually represented in texts.However, we have found through experiments that multimodal output can significantly improve user satisfaction for informativeness of summaries.In this paper, we propose a novel task, multimodal summarization with multimodal output (MSMO).To handle this task, we first collect a large-scale dataset for MSMO research.We then propose a multimodal attention model to jointly generate text and select the most relevant image from the multimodal input.Finally, to evaluate multimodal outputs, we construct a novel multimodal automatic evaluation (MMAE) method which considers both intramodality salience and intermodality relevance.The experimental results show the effectiveness of MMAE.
Junnan Zhu, Haoran Li 0001, Tianshang Liu, Yu Zhou 0001, Jiajun Zhang 0001, Chengqing Zong
EMNLP1
2018 Multi-modal Sentence Summarization with Modality Attention and Image Filtering
abstract
In this paper, we introduce a multi-modal sentence summarization task that produces a short summary from a pair of sentence and image. This task is more challenging than sentence summarization. It not only needs to effectively incorporate visual features into standard text summarization framework, but also requires to avoid noise of image. To this end, we propose a modality-based attention mechanism to pay different attention to image patches and text units, and we design image filters to selectively use visual information to enhance the semantics of the input sentence. We construct a multimodal sentence summarization dataset and extensive experiments on this dataset demonstrate that our models significantly outperform conventional models which only employ text as input. Further analyses suggest that sentence summarization task can benefit from visually grounded representations from a variety of aspects.
Haoran Li 0001, Junnan Zhu, Tianshang Liu, Jiajun Zhang 0001, Chengqing Zong
IJCAI2
2017 Multi-modal Summarization for Asynchronous Collection of Text, Image, Audio and Video
abstract
The rapid increase in multimedia data transmission over the Internet necessitates the multi-modal summarization (MMS) from collections of text, image, audio and video.In this work, we propose an extractive multi-modal summarization method that can automatically generate a textual summary given a set of documents, images, audios and videos related to a specific topic.The key idea is to bridge the semantic gaps between multi-modal content.For audio information, we design an approach to selectively use its transcription.For visual information, we learn the joint representations of text and images using a neural network.Finally, all of the multimodal aspects are considered to generate the textual summary by maximizing the salience, non-redundancy, readability and coverage through the budgeted optimization of submodular functions.We further introduce an MMS corpus in English and Chinese, which is released to the public 1 .The experimental results obtained on this dataset demonstrate that our method outperforms other competitive baseline methods.
Haoran Li 0001, Junnan Zhu, Cong Ma 0002, Jiajun Zhang 0001, Chengqing Zong
EMNLP2
2017 Augmenting Neural Sentence Summarization Through Extractive Summarization
Junnan Zhu, Haoran Li 0001, Jiajun Zhang 0001, Yu Zhou 0001, Chengqing Zong
NLPCC1