Haoran Li 0001

dblp:50/10038-1 · DBLP profile ↗
← Back
23ranked-venue papers
11as first author
6since 2021 · last 2022
0000-0002-2368-7541ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 21 · 10 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 3 first-authorSystems, architecture and hardware · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
13 papers
Language models and text generation · 53% Vision and language · 13% Deep learning architectures and training · 8%
Computer graphics and multimedia
4 papers
Multimedia analysis and retrieval · 94% Audio and music processing · 6%

Topics — the 27 heaviest of 30, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Language models and text generation › text summarization
abstractive summarization
2.252021
Learn to Copy from the Copying History: Correlational Copy Network for Abstractive Summarization · EMNLP (1) 2021
Self-Attention Guided Copy Mechanism for Abstractive Summarization · ACL 2020
Keywords-Guided Abstractive Sentence Summarization · AAAI 2020
Natural language and speech › Language models and text generation
text summarization
2.162022
Multimodal Summarization with Guidance of Multimodal Reference · AAAI 2020
Read, Watch, Listen, and Summarize: Multi-Modal Summarization for Asynchronous Text, Image, Audio and Video · IEEE Trans. Knowl. Data Eng. 2019
Attention With Sparsity Regularization for Neural Machine Translation and Summarization · IEEE ACM Trans. Audio Speech Lang. Process. 2019
Computer vision › Vision and language › vision-language generation
multimodal summarization
1.232020
Multimodal Summarization with Guidance of Multimodal Reference · AAAI 2020
Aspect-Aware Multimodal Summarization for Chinese E-Commerce Products · AAAI 2020
Multi-modal Sentence Summarization with Modality Attention and Image Filtering · IJCAI 2018
Multimedia analysis and retrieval
multimodal summarization
1.032019
Read, Watch, Listen, and Summarize: Multi-Modal Summarization for Asynchronous Text, Image, Audio and Video · IEEE Trans. Knowl. Data Eng. 2019
MSMO: Multimodal Summarization with Multimodal Output · EMNLP 2018
Multi-modal Summarization for Asynchronous Collection of Text, Image, Audio and Video · EMNLP 2017
Natural language and speech › Language models and text generation › text generation › neural text generation
copy mechanism
0.922021
Learn to Copy from the Copying History: Correlational Copy Network for Abstractive Summarization · EMNLP (1) 2021
Self-Attention Guided Copy Mechanism for Abstractive Summarization · ACL 2020
Multimedia analysis and retrieval
multimodal evaluation
0.822020
Multimodal Summarization with Guidance of Multimodal Reference · AAAI 2020
MSMO: Multimodal Summarization with Multimodal Output · EMNLP 2018
Natural language and speech › Question answering and dialogue systems › dialogue generation
dialogue response generation
0.612022
JDDC 2.1: A Multimodal Chinese Dialogue Dataset with Joint Tasks of Query Rewriting, Response Generation, Discourse Parsing, and Summarization · EMNLP 2022
Natural language and speech › Information extraction and text analysis › discourse analysis
discourse parsing
0.612022
JDDC 2.1: A Multimodal Chinese Dialogue Dataset with Joint Tasks of Query Rewriting, Response Generation, Discourse Parsing, and Summarization · EMNLP 2022
Computer vision › Vision and language
multimodal dialogue
0.612022
JDDC 2.1: A Multimodal Chinese Dialogue Dataset with Joint Tasks of Query Rewriting, Response Generation, Discourse Parsing, and Summarization · EMNLP 2022
Natural language and speech › Language models and text generation
pre-trained language model
0.612022
PRINCE: Prefix-Masked Decoding for Knowledge Enhanced Sequence-to-Sequence Pre-Training · EMNLP 2022
Knowledge, reasoning and agents › Knowledge representation and reasoning › ontology-based query answering
query rewriting
0.612022
JDDC 2.1: A Multimodal Chinese Dialogue Dataset with Joint Tasks of Query Rewriting, Response Generation, Discourse Parsing, and Summarization · EMNLP 2022
Natural language and speech › Language models and text generation › large language model training
sequence-to-sequence pretraining
0.612022
PRINCE: Prefix-Masked Decoding for Knowledge Enhanced Sequence-to-Sequence Pre-Training · EMNLP 2022
Computer vision › Image recognition and object detection
attribute recognition
0.412020
Multimodal Joint Attribute Prediction and Value Extraction for E-commerce Product · EMNLP (1) 2020
Natural language and speech › Information extraction and text analysis › relation extraction
attribute value extraction
0.412020
Multimodal Joint Attribute Prediction and Value Extraction for E-commerce Product · EMNLP (1) 2020
Machine learning › Generative modeling › diffusion model › guided diffusion
self-attention guidance
0.412020
Self-Attention Guided Copy Mechanism for Abstractive Summarization · ACL 2020
Natural language and speech › Language models and text generation › text summarization
sentence compression
0.412020
Keywords-Guided Abstractive Sentence Summarization · AAAI 2020
Machine learning › Deep learning architectures and training
transformer
0.412020
Self-Attention Guided Copy Mechanism for Abstractive Summarization · ACL 2020
Machine learning › Deep learning architectures and training
attention mechanism
0.412019
Attention With Sparsity Regularization for Neural Machine Translation and Summarization · IEEE ACM Trans. Audio Speech Lang. Process. 2019
Natural language and speech › Language models and text generation › text summarization
extractive summarization
0.412019
Read, Watch, Listen, and Summarize: Multi-Modal Summarization for Asynchronous Text, Image, Audio and Video · IEEE Trans. Knowl. Data Eng. 2019
Natural language and speech › Machine translation
neural machine translation
0.412019
Attention With Sparsity Regularization for Neural Machine Translation and Summarization · IEEE ACM Trans. Audio Speech Lang. Process. 2019
Machine learning › Deep learning architectures and training › attention mechanism
sparse attention
0.412019
Attention With Sparsity Regularization for Neural Machine Translation and Summarization · IEEE ACM Trans. Audio Speech Lang. Process. 2019
Natural language and speech › Language models and text generation › text summarization
multimodal sentence summarization
0.312018
Multi-modal Sentence Summarization with Modality Attention and Image Filtering · IJCAI 2018
Machine learning › Representation and self-supervised learning › representation learning › embedding learning
entity representation learning
0.212022
PRINCE: Prefix-Masked Decoding for Knowledge Enhanced Sequence-to-Sequence Pre-Training · EMNLP 2022
Natural language and speech › Language models and text generation › text summarization
dialogue summarization
0.112021
Learn to Copy from the Copying History: Correlational Copy Network for Abstractive Summarization · EMNLP (1) 2021
Natural language and speech › Information extraction and text analysis › sentiment analysis › aspect-based sentiment analysis
aspect extraction
0.112020
Aspect-Aware Multimodal Summarization for Chinese E-Commerce Products · AAAI 2020
Multimedia analysis and retrieval › cross-modal retrieval
image-text retrieval
0.112019
Read, Watch, Listen, and Summarize: Multi-Modal Summarization for Asynchronous Text, Image, Audio and Video · IEEE Trans. Knowl. Data Eng. 2019
Computer vision › Image recognition and object detection
visual feature extraction
0.112018
Multi-modal Sentence Summarization with Modality Attention and Image Filtering · IJCAI 2018

Methods — techniques the papers use, named apart from their topics

submodular optimization · 1.3neural network · 1.3multimodal learning · 1.0prefix masking · 0.6knowledge-enhanced decoding · 0.6denoising autoencoder · 0.6pre-training · 0.5encoder-decoder attention · 0.5order-ranking · 0.4joint multimodal representation · 0.4dual attention · 0.4attention · 0.4ROUGE-ranking · 0.4topic modeling · 0.4multimodal attention model · 0.3
YearPublicationVenuePosition
2022 PRINCE: Prefix-Masked Decoding for Knowledge Enhanced Sequence-to-Sequence Pre-Training
abstract
Pre-trained Language Models (PLMs) have shown effectiveness in various Natural Language Processing (NLP) tasks.Denoising autoencoder is one of the most successful pretraining frameworks, learning to recompose the original text given a noise-corrupted one.The existing studies mainly focus on injecting noises into the input.This paper introduces a simple yet effective pre-training paradigm, equipped with a knowledge-enhanced decoder that predicts the next entity token with noises in the prefix, explicitly strengthening the representation learning of entities that span over multiple input tokens.Specifically, when predicting the next token within an entity, we feed masks into the prefix in place of some of the previous ground-truth tokens that constitute the entity.Our model achieves new state-of-the-art results on two knowledge-driven data-to-text generation tasks with up to 2% BLEU gains.
Song Xu 0002, Haoran Li 0001, Peng Yuan 0002, Youzheng Wu, Xiaodong He 0001
EMNLP2
2022 JDDC 2.1: A Multimodal Chinese Dialogue Dataset with Joint Tasks of Query Rewriting, Response Generation, Discourse Parsing, and Summarization
abstract
The popularity of multimodal dialogue has stimulated the need for a new generation of dialogue agents with multimodal interactivity.When users communicate with customer service, they may express their requirements by means of text, images, or even videos.Visual information usually acts as discriminators for product models, or indicators of product failures, which play an important role in the Ecommerce scenario.On the other hand, detailed information provided by the images is limited, and typically, customer service systems cannot understand the intent of users without the input text.Thus, bridging the gap between the image and text is crucial for communicating with customers.In this paper, we construct JDDC 2.1, a large-scale multimodal multi-turn dialogue dataset collected from a mainstream Chinese E-commerce platform 1 , containing about 246K dialogue sessions, 3M utterances, and 507K images, along with product knowledge bases and image category annotations.Over our dataset, we jointly define four tasks: the multimodal dialogue response generation task, the multimodal query rewriting task, the multimodal dialogue discourse parsing task, and the multimodal dialogue summarization task.JDDC 2.1 is the first corpus with annotations for all the above tasks over the same dialogue sessions, which facilitates the comprehensive research around the dialogue.In addition, we present several text-only and multimodal baselines and show the importance of visual information for these tasks.Our dataset and implements will be publicly available.
Haoran Li 0001, Youzheng Wu, Xiaodong He 0001
EMNLP2
2022 An Input-Output Regulated Adaptive Ramp for Fast Load Transition of PWM Buck Convertor
abstract
An input-output regulated adaptive ramp ($\mathrm{IOR}^{2})$ for fast load transition of pulse width modulation (PWM) buck convertor is presented. The scheme employs an adaptive ramp regulated by input and the voltage from the error amplifier to achieve fast load response and low line and load regulation rate. Simulation shows that an under/overshoot voltage of −32 mV and 35 mV, with $8 \mu$ and $8.3 \mu \mathrm{s}$ recovery time are respectively obtained for the load current stepping between 1 A and 1.7 A. The line/load regulation rate is respectively $12.4 \mu \mathrm{V} / \mathrm{V}$ and $1.04 \mathrm{mV} / \mathbf{A}$. Implemented in a $0.25 \mu \mathrm{m}$ BCD process, the proposed IOR2PWM regulator is capable of converting input voltage of 5 V to 65 V to output range of 3.3 V to 60 V with adjustable switching frequency up to 2.2 MHz, showing a peak efficiency of 94.5% at 1 A load current.
Bingbing He, Haoran Li 0001, Yongfu Li 0002, Yan Liu 0016, Yang Zhao 0007
ISCAS2
2022 Overview of the NLPCC 2022 Shared Task on Multimodal Product Summarization
Haoran Li 0001, Peng Yuan 0002, Haoning Zhang, Weikang Li, Song Xu 0002, Youzheng Wu, Xiaodong He 0001
NLPCC (2)1
2021 Learn to Copy from the Copying History: Correlational Copy Network for Abstractive Summarization
abstract
The copying mechanism has had considerable success in abstractive summarization, facilitating models to directly copy words from the input text to the output summary.Existing works mostly employ encoder-decoder attention, which applies copying at each time step independently of the former ones.However, this may sometimes lead to incomplete copying.In this paper, we propose a novel copying scheme named Correlational Copying Network (CoCoNet) that enhances the standard copying mechanism by keeping track of the copying history.It thereby takes advantage of prior copying distributions and, at each time step, explicitly encourages the model to copy the input word that is relevant to the previously copied one.In addition, we strengthen CoCoNet through pretraining with suitable corpora that simulate the copying behaviors.Experimental results show that CoCoNet can copy more accurately and achieves new state-of-the-art performances on summarization benchmarks, including CNN/DailyMail for news summarization and SAMSum for dialogue summarization.Our code is available at https:// github.com/hrlinlp/coconet.
Haoran Li 0001, Song Xu 0002, Peng Yuan 0002, Youzheng Wu, Xiaodong He 0001, Bowen Zhou 0001
EMNLP (1)1
2021 Semantically Constrained Document-Level Chinese-Mongolian Neural Machine Translation
abstract
By using document-level contextual information, document-level neural machine translation can achieve better results than ordinary machine translation, but traditional document-level machine translation is difficult to focus on the contextual sentence articulation relations and deep positional relations within the discourse while utilizing document-level vocabulary, and the model can concentrate only on relatively shallow inter-sentential relations or positional information. In this paper, we consider that most adjacent sentences are connected in document translation, and such links help improve the quality of translation. We propose a document translation model that focuses more on inter-sentential relations based on the previous work, and propose two methods to strengthen the model's positional information input, and combine these two methods to enhance the traditional Transformer positional information input. This paper also proposes a method for inserting paragraph information to allow inter-sentential relations to be learned by the model, and uses the improved Transformer model for Chinese-Mongolian document translation. Experiments show that in the improved Transformer system, the BLEU scores are enhanced on the Chinese-Mongolian machine translation task after fusing positional information and inter-sentential relation information, and the translation achieves better performance.
Haoran Li 0001, Hongxu Hou, Nier Wu, Xiaoning Jia
IJCNN1
2020 Aspect-Aware Multimodal Summarization for Chinese E-Commerce Products
abstract
We present an abstractive summarization system that produces summary for Chinese e-commerce products. This task is more challenging than general text summarization. First, the appearance of a product typically plays a significant role in customers' decisions to buy the product or not, which requires that the summarization model effectively use the visual information of the product. Furthermore, different products have remarkable features in various aspects, such as “energy efficiency” and “large capacity” for refrigerators. Meanwhile, different customers may care about different aspects. Thus, the summarizer needs to capture the most attractive aspects of a product that resonate with potential purchasers. We propose an aspect-aware multimodal summarization model that can effectively incorporate the visual information and also determine the most salient aspects of a product. We construct a large-scale Chinese e-commerce product summarization dataset that contains approximately 1.4 million manually created product summaries that are paired with detailed product information, including an image, a title, and other textual descriptions for each product. The experimental results on this dataset demonstrate that our models significantly outperform the comparative methods in terms of both the ROUGE score and manual evaluations.
Haoran Li 0001, Peng Yuan 0002, Song Xu 0002, Youzheng Wu, Xiaodong He 0001, Bowen Zhou 0001
AAAI1
2020 Keywords-Guided Abstractive Sentence Summarization
abstract
We study the problem of generating a summary for a given sentence. Existing researches on abstractive sentence summarization ignore that keywords in the input sentence provide significant clues for valuable content, and humans tend to write summaries covering these keywords. In this paper, we propose an abstractive sentence summarization method by applying guidance signals of keywords to both the encoder and the decoder in the sequence-to-sequence model. A multi-task learning framework is adopted to jointly learn to extract keywords and generate a summary for the input sentence. We apply keywords-guided selective encoding strategies to filter source information by investigating the interactions between the input sentence and the keywords. We extend pointer-generator network by a dual-attention and a dual-copy mechanism, which can integrate the semantics of the input sentence and the keywords, and copy words from both the input sentence and the keywords. We demonstrate that multi-task learning and keywords-oriented guidance facilitate sentence summarization task, achieving better performance than the competitive models on the English Gigaword sentence summarization dataset.
Haoran Li 0001, Junnan Zhu, Jiajun Zhang 0001, Chengqing Zong, Xiaodong He 0001
AAAI1
2020 Multimodal Summarization with Guidance of Multimodal Reference
abstract
Multimodal summarization with multimodal output (MSMO) is to generate a multimodal summary for a multimodal news report, which has been proven to effectively improve users' satisfaction. The existing MSMO methods are trained by the target of text modality, leading to the modality-bias problem that ignores the quality of model-selected image during training. To alleviate this problem, we propose a multimodal objective function with the guidance of multimodal reference to use the loss from the summary generation and the image selection. Due to the lack of multimodal reference data, we present two strategies, i.e., ROUGE-ranking and Order-ranking, to construct the multimodal reference by extending the text reference. Meanwhile, to better evaluate multimodal outputs, we propose a novel evaluation metric based on joint multimodal representation, projecting the model output and multimodal reference into a joint semantic space during evaluation. Experimental results have shown that our proposed model achieves the new state-of-the-art on both automatic and manual evaluation metrics. Besides, our proposed evaluation method can effectively improve the correlation with human judgments.
Junnan Zhu, Yu Zhou 0001, Jiajun Zhang 0001, Haoran Li 0001, Chengqing Zong, Changliang Li
AAAI4
2020 Self-Attention Guided Copy Mechanism for Abstractive Summarization
abstract
Copy module has been widely equipped in the recent abstractive summarization models, which facilitates the decoder to extract words from the source into the summary.Generally, the encoder-decoder attention is served as the copy distribution, while how to guarantee that important words in the source are copied remains a challenge.In this work, we propose a Transformer-based model to enhance the copy mechanism.Specifically, we identify the importance of each source word based on the degree centrality with a directed graph built by the self-attention layer in the Transformer.We use the centrality of each source word to guide the copy process explicitly.Experimental results show that the self-attention graph provides useful guidance for the copy distribution.Our proposed models significantly outperform the baseline methods on the CNN/Daily Mail dataset and the Gigaword dataset.
Song Xu 0002, Haoran Li 0001, Peng Yuan 0002, Youzheng Wu, Xiaodong He 0001, Bowen Zhou 0001
ACL2
2020 Multimodal Sentence Summarization via Multimodal Selective Encoding
abstract
This paper studies the problem of generating a summary for a given sentence-image pair.Existing multimodal sequence-to-sequence approaches mainly focus on enhancing the decoder by visual signals, while ignoring that the image can improve the ability of the encoder to identify highlights of a news event or a document.Thus, we propose a multimodal selective gate network that considers reciprocal relationships between textual and multi-level visual features, including global image descriptor, activation grids, and object proposals, to select highlights of the event when encoding the source sentence.In addition, we introduce a modality regularization to encourage the summary to capture the highlights embedded in the image more accurately.To verify the generalization of our model, we adopt the multimodal selective gate to the text-based decoder and multimodal-based decoder.Experimental results on a public multimodal sentence summarization dataset demonstrate the advantage of our models over baselines.Further analysis suggests that our proposed multimodal selective gate network can effectively select important information in the input sentence.
Haoran Li 0001, Junnan Zhu, Jiajun Zhang 0001, Xiaodong He 0001, Chengqing Zong
COLING1
2020 On the Faithfulness for E-commerce Product Summarization
abstract
In this work, we present a model to generate e-commerce product summaries.The consistency between the generated summary and the product attributes is an essential criterion for the ecommerce product summarization task.To enhance the consistency, first, we encode the product attribute table to guide the process of summary generation.Second, we identify the attribute words from the vocabulary, and we constrain these attribute words can be presented in the summaries only through copying from the source, i.e., the attribute words not in the source cannot be generated.We construct a Chinese e-commerce product summarization dataset, and the experimental results on this dataset demonstrate that our models significantly improve the faithfulness.
Peng Yuan 0002, Haoran Li 0001, Song Xu 0002, Youzheng Wu, Xiaodong He 0001, Bowen Zhou 0001
COLING2
2020 Multimodal Joint Attribute Prediction and Value Extraction for E-commerce Product
abstract
Product attribute values are essential in many e-commerce scenarios, such as customer service robots, product recommendations, and product retrieval.While in the real world, the attribute values of a product are usually incomplete and vary over time, which greatly hinders the practical applications.In this paper, we propose a multimodal method to jointly predict product attributes and extract values from textual product descriptions with the help of the product images.We argue that product attributes and values are highly correlated, e.g., it will be easier to extract the values on condition that the product attributes are given.Thus, we jointly model the attribute prediction and value extraction tasks from multiple aspects towards the interactions between attributes and values.Moreover, product images have distinct effects on our tasks for different product attributes and values.Thus, we selectively draw useful visual information from product images to enhance our model.We annotate a multimodal product attribute value dataset that contains 87,194 instances, and the experimental results on this dataset demonstrate that explicitly modeling the relationship between attributes and values facilitates our method to establish the correspondence between them, and selectively utilizing visual product information is necessary for the task.Our code and dataset are available 1 .
Tiangang Zhu, Haoran Li 0001, Youzheng Wu, Xiaodong He 0001, Bowen Zhou 0001
EMNLP (1)3
2019 Towards Personalized Review Summarization via User-Aware Sequence Network
abstract
We address personalized review summarization, which generates a condensed summary for a user’s review, accounting for his preference on different aspects or his writing style. We propose a novel personalized review summarization model named User-aware Sequence Network (USN) to consider the aforementioned users’ characteristics when generating summaries, which contains a user-aware encoder and a useraware decoder. Specifically, the user-aware encoder adopts a user-based selective mechanism to select the important information of a review, and the user-aware decoder incorporates user characteristic and user-specific word-using habits into word prediction process to generate personalized summaries. To validate our model, we collected a new dataset Trip, comprising 536,255 reviews from 19,400 users. With quantitative and human evaluation, we show that USN achieves state-ofthe-art performance on personalized review summarization.
Haoran Li 0001, Chengqing Zong
AAAI2
2019 Incorporating Multi-Level User Preference into Document-Level Sentiment Classification
abstract
Document-level sentiment classification aims to predict a user’s sentiment polarity in a document about a product. Most existing methods only focus on review contents and ignore users who post reviews. In fact, when reviewing a product, different users have different word-using habits to express opinions (i.e., word-level user preference), care about different attributes of the product (i.e., aspect-level user preference), and have different characteristics to score the review (i.e., polarity-level user preference). These preferences have great influence on interpreting the sentiment of text. To address this issue, we propose a model called Hierarchical User Attention Network (HUAN), which incorporates multi-level user preference into a hierarchical neural network to perform document-level sentiment classification. Specifically, HUAN encodes different kinds of information (word, sentence, aspect, and document) in a hierarchical structure and imports user embedding and user attention mechanism to model these preferences. Empirical results on two real-world datasets show that HUAN achieves state-of-the-art performance. Furthermore, HUAN can also mine important attributes of products for different users.
Haoran Li 0001, Xiaomian Kang, Haitong Yang, Chengqing Zong
ACM Trans. Asian Low Resour. Lang. Inf. Process.2
2019 Attention With Sparsity Regularization for Neural Machine Translation and Summarization
abstract
The attention mechanism has become thede factostandard component in neural sequence to sequence tasks, such as machine translation and abstractive summarization. It dynamically determines which parts in the input sentence should be focused on when generating each word in the output sequence. Ideally, only few relevant input words should be attended to at each decoding time step and the attention weight distribution should be sparse and sharp. However, previous methods have no good mechanism to control this attention weight distribution. In this paper, we propose a sparse attention model in which a sparsity regularization term is designed to augment the objective function. We explore two kinds of regularizations:$L_{\infty }$-norm regularization and minimum entropy regularization, both of which aim to sharpen the attention weight distribution. Extensive experiments on both neural machine translation and abstractive summarization demonstrate that our proposed sparse attention model can substantially outperform the strong baselines. And the detailed analyses reveal that the final attention distribution indeed becomes sparse and sharp.
Jiajun Zhang 0001, Yang Zhao 0007, Haoran Li 0001, Chengqing Zong
IEEE ACM Trans. Audio Speech Lang. Process.3
2019 Read, Watch, Listen, and Summarize: Multi-Modal Summarization for Asynchronous Text, Image, Audio and Video
abstract
Automatic text summarization is a fundamental natural language processing (NLP) application that aims to condense a source text into a shorter version. The rapid increase in multimedia data transmission over the Internet necessitates multi-modal summarization (MMS) from asynchronous collections of text, image, audio, and video. In this work, we propose an extractive MMS method that unites the techniques of NLP, speech processing, and computer vision to explore the rich information contained in multi-modal data and to improve the quality of multimedia news summarization. The key idea is to bridge the semantic gaps between multi-modal content. Audio and visual are main modalities in the video. For audio information, we design an approach to selectively use its transcription and to infer the salience of the transcription with audio signals. For visual information, we learn the joint representations of text and images using a neural network. Then, we capture the coverage of the generated summary for important visual information through text-image matching or multi-modal topic modeling. Finally, all the multi-modal aspects are considered to generate a textual summary by maximizing the salience, non-redundancy, readability, and coverage through the budgeted optimization of submodular functions. We further introduce a publicly available MMS corpus in English and Chinese.1 The experimental results obtained on our dataset demonstrate that our methods based on image matching and image topic framework outperform other competitive baseline methods.
Haoran Li 0001, Junnan Zhu, Cong Ma 0002, Jiajun Zhang 0001, Chengqing Zong
IEEE Trans. Knowl. Data Eng.1
2018 Ensure the Correctness of the Summary: Incorporate Entailment Knowledge into Abstractive Sentence Summarization
abstract
In this paper, we investigate the sentence summarization task that produces a summary from a source sentence. Neural sequence-to-sequence models have gained considerable success for this task, while most existing approaches only focus on improving the informativeness of the summary, which ignore the correctness, i.e., the summary should not contain unrelated information with respect to the source sentence. We argue that correctness is an essential requirement for summarization systems. Considering a correct summary is semantically entailed by the source sentence, we incorporate entailment knowledge into abstractive summarization models. We propose an entailment-aware encoder under multi-task framework (i.e., summarization generation and entailment recognition) and an entailment-aware decoder by entailment Reward Augmented Maximum Likelihood (RAML) training. Experiment results demonstrate that our models significantly outperform baselines from the aspects of informativeness and correctness.
Haoran Li 0001, Junnan Zhu, Jiajun Zhang 0001, Chengqing Zong
COLING1
2018 MSMO: Multimodal Summarization with Multimodal Output
abstract
Multimodal summarization has drawn much attention due to the rapid growth of multimedia data.The output of the current multimodal summarization systems is usually represented in texts.However, we have found through experiments that multimodal output can significantly improve user satisfaction for informativeness of summaries.In this paper, we propose a novel task, multimodal summarization with multimodal output (MSMO).To handle this task, we first collect a large-scale dataset for MSMO research.We then propose a multimodal attention model to jointly generate text and select the most relevant image from the multimodal input.Finally, to evaluate multimodal outputs, we construct a novel multimodal automatic evaluation (MMAE) method which considers both intramodality salience and intermodality relevance.The experimental results show the effectiveness of MMAE.
Junnan Zhu, Haoran Li 0001, Tianshang Liu, Yu Zhou 0001, Jiajun Zhang 0001, Chengqing Zong
EMNLP2
2018 Multi-modal Sentence Summarization with Modality Attention and Image Filtering
abstract
In this paper, we introduce a multi-modal sentence summarization task that produces a short summary from a pair of sentence and image. This task is more challenging than sentence summarization. It not only needs to effectively incorporate visual features into standard text summarization framework, but also requires to avoid noise of image. To this end, we propose a modality-based attention mechanism to pay different attention to image patches and text units, and we design image filters to selectively use visual information to enhance the semantics of the input sentence. We construct a multimodal sentence summarization dataset and extensive experiments on this dataset demonstrate that our models significantly outperform conventional models which only employ text as input. Further analyses suggest that sentence summarization task can benefit from visually grounded representations from a variety of aspects.
Haoran Li 0001, Junnan Zhu, Tianshang Liu, Jiajun Zhang 0001, Chengqing Zong
IJCAI1
2017 Multi-modal Summarization for Asynchronous Collection of Text, Image, Audio and Video
abstract
The rapid increase in multimedia data transmission over the Internet necessitates the multi-modal summarization (MMS) from collections of text, image, audio and video.In this work, we propose an extractive multi-modal summarization method that can automatically generate a textual summary given a set of documents, images, audios and videos related to a specific topic.The key idea is to bridge the semantic gaps between multi-modal content.For audio information, we design an approach to selectively use its transcription.For visual information, we learn the joint representations of text and images using a neural network.Finally, all of the multimodal aspects are considered to generate the textual summary by maximizing the salience, non-redundancy, readability and coverage through the budgeted optimization of submodular functions.We further introduce an MMS corpus in English and Chinese, which is released to the public 1 .The experimental results obtained on this dataset demonstrate that our method outperforms other competitive baseline methods.
Haoran Li 0001, Junnan Zhu, Cong Ma 0002, Jiajun Zhang 0001, Chengqing Zong
EMNLP1
2017 Augmenting Neural Sentence Summarization Through Extractive Summarization
Junnan Zhu, Haoran Li 0001, Jiajun Zhang 0001, Yu Zhou 0001, Chengqing Zong
NLPCC3
2017 Implicit Discourse Relation Recognition for English and Chinese with Multiview Modeling and Effective Representation Learning
abstract
Discourse relations between two text segments play an important role in many Natural Language Processing (NLP) tasks. The connectives strongly indicate the sense of discourse relations, while in fact, there are no connectives in a large proportion of discourse relations, that is, implicit discourse relations. Compared with explicit relations, implicit relations are much harder to detect and have drawn significant attention. Until now, there have been many studies focusing on English implicit discourse relations, and few studies address implicit relation recognition in Chinese even though the implicit discourse relations in Chinese are more common than those in English. In our work, both the English and Chinese languages are our focus. The key to implicit relation prediction is to properly model the semantics of the two discourse arguments, as well as the contextual interaction between them. To achieve this goal, we propose a neural network based framework that consists of two hierarchies. The first one is the model hierarchy, in which we propose a max-margin learning method to explore the implicit discourse relation from multiple views. The second one is the feature hierarchy, in which we learn multilevel distributed representations from words, arguments, and syntactic structures to sentences. We have conducted experiments on the standard benchmarks of English and Chinese, and the results show that compared with several methods our proposed method can achieve the best performance in most cases.
Haoran Li 0001, Jiajun Zhang 0001, Chengqing Zong
ACM Trans. Asian Low Resour. Lang. Inf. Process.1