Zhenglu Yang

dblp:43/5146 · DBLP profile ↗
← Back
73ranked-venue papers
9as first author
28since 2021 · last 2026
0000-0001-9528-965XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 43 · 4 first-author · 19 since 2021Graphics, computer vision, multimedia, augmented reality and games · 22 · 4 first-author · 7 since 2021Databases, data management, data science and information retrieval · 21 · 5 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 1 first-author · 2 since 2021Computer networks · 2 · 2 since 2021Systems, architecture and hardware · 1
YearPublicationVenuePosition
2026 Multi-grained multi-label feature selection: Jointly towards semantic-aware and instance-specific features
Hengpeng Xu, Zhenglu Yang, Jun Wang 0023
Knowl. Based Syst.3
2026 Adaptive knowledge selection in dialogue systems: Accommodating diverse knowledge types, requirements, and generation models
Zhongtian Bao, Hongru Liang, Jun Wang 0023, Zhenglu Yang, Zhe Sun 0009, Andrzej Cichocki
Neural Networks6
2025 Enhancing Cross-Lingual Dialogue Summarization Through Interpretable Chain-of-Thought
Zhongtian Bao, Jun Wang 0023, Adam Jatowt, Zhenglu Yang
DASFAA (2)5
2024 Can We Learn Question, Answer, and Distractors All from an Image? A New Task for Multiple-choice Visual Question Answering
abstract
Multiple-choice visual question answering (MC VQA) requires an answer picked from a list of distractors, based on a question and an image. This research has attracted wide interest from the fields of visual question answering, visual question generation, and visual distractor generation. However, these fields still stay in their own territories, and how to jointly generate meaningful questions, correct answers, and challenging distractors remains unexplored. In this paper, we introduce a novel task, Visual Question-Answer-Distractors Generation (VQADG), which can bridge this research gap as well as take as a cornerstone to promote existing VQA models. Specific to the VQADG task, we present a novel framework consisting of a vision-and-language model to encode the given image and generate QADs jointly, and contrastive learning to ensure the consistency of the generated question, answer, and distractors. Empirical evaluations on the benchmark dataset validate the performance of our model in the VQADG task.
Wenjian Ding, Jun Wang 0023, Adam Jatowt, Zhenglu Yang
LREC/COLING5
2024 Exploring Union and Intersection of Visual Regions for Generating Questions, Answers, and Distractors
abstract
Multiple-choice visual question answering (VQA) is to automatically choose a correct answer from a set of choices after reading an image.Existing efforts have been devoted to a separate generation of an image-related question, a correct answer, or challenge distractors.By contrast, we turn to a holistic generation and optimization of questions, answers, and distractors (QADs) in this study.This integrated generation strategy eliminates the need for human curation and guarantees information consistency.Furthermore, we first propose to put the spotlight on different image regions to diversify QADs.Accordingly, a novel framework ReBo is formulated in this paper.ReBo cyclically generates each QAD based on a recurrent multimodal encoder, and each generation is focusing on a different area of the image compared to those already concerned by the previously generated QADs.In addition to traditional VQA comparisons with state-of-the-art approaches, we also validate the capability of ReBo in generating augmented data to benefit VQA models.
Wenjian Ding, Jun Wang 0023, Adam Jatowt, Zhenglu Yang
EMNLP5
2023 HaPPy: Harnessing the Wisdom from Multi-Perspective Graphs for Protein-Ligand Binding Affinity Prediction (Student Abstract)
abstract
Gathering information from multi-perspective graphs is an essential issue for many applications especially for proteinligand binding affinity prediction. Most of traditional approaches obtained such information individually with low interpretability. In this paper, we harness the rich information from multi-perspective graphs with a general model, which abstractly represents protein-ligand complexes with better interpretability while achieving excellent predictive performance. In addition, we specially analyze the protein-ligand binding affinity problem, taking into account the heterogeneity of proteins and ligands. Experimental evaluations demonstrate the effectiveness of our data representation strategy on public datasets by fusing information from different perspectives.
Xianfeng Zhang, Yanhui Gu, Guandong Xu, Jinlan Wang, Zhenglu Yang
AAAI6
2023 Well Begun is Half Done: Generator-agnostic Knowledge Pre-Selection for Knowledge-Grounded Dialogue
abstract
Accurate knowledge selection is critical in knowledge-grounded dialogue systems.Towards a closer look at it, we offer a novel perspective to organize existing literature, i.e., knowledge selection coupled with, after, and before generation.We focus on the third underexplored category of study, which can not only select knowledge accurately in advance, but has the advantage to reduce the learning, adjustment, and interpretation burden of subsequent response generation models, especially LLMs.We propose GATE, a generator-agnostic knowledge selection method, to prepare knowledge for subsequent response generation models by selecting context-related knowledge among different knowledge structures and variable knowledge requirements.Experimental results demonstrate the superiority of GATE, and indicate that knowledge selection before generation is a lightweight yet effective way to facilitate LLMs (e.g., ChatGPT) to generate more informative responses.
Hongru Liang, Jun Wang 0023, Zhenglu Yang
EMNLP5
2023 Multi-path Based Self-adaptive Cross-lingual Summarization
Zhongtian Bao, Jun Wang 0023, Zhenglu Yang
KSEM (3)3
2022 UniMS: A Unified Framework for Multimodal Summarization with Knowledge Distillation
abstract
With the rapid increase of multimedia data, a large body of literature has emerged to work on multimodal summarization, the majority of which target at refining salient information from textual and image modalities to output a pictorial summary with the most relevant images. Existing methods mostly focus on either extractive or abstractive summarization and rely on the presence and quality of image captions to build image references. We are the first to propose a Unified framework for Multimodal Summarization grounding on BART, UniMS, that integrates extractive and abstractive objectives, as well as selecting the image output. Specially, we adopt knowledge distillation from a vision-language pretrained model to improve image selection, which avoids any requirement on the existence and quality of image captions. Besides, we introduce a visual guided decoder to better integrate textual and visual modalities in guiding abstractive text generation. Results show that our best model achieves a new state-of-the-art result on a large-scale benchmark dataset. The newly involved extractive objective as well as the knowledge distillation technique are proven to bring a noticeable improvement to the multimodal summarization task.
Zhengkun Zhang, Xiaojun Meng, Yasheng Wang, Xin Jiang 0002, Qun Liu 0001, Zhenglu Yang
AAAI6
2022 Multi-Party Empathetic Dialogue Generation: A New Task for Dialog Systems
abstract
Empathetic dialogue assembles emotion understanding, feeling projection, and appropriate response generation.Existing work for empathetic dialogue generation concentrates on the two-party conversation scenario.Multiparty dialogues, however, are pervasive in reality.Furthermore, emotion and sensibility are typically confused; a refined empathy analysis is needed for comprehending fragile and nuanced human feelings.We address these issues by proposing a novel task called Multi-Party Empathetic Dialogue Generation in this study.Additionally, a Static-Dynamic model for Multi-Party Empathetic Dialogue Generation, SDMPED, is introduced as a baseline by exploring the static sensibility and dynamic emotion for the multi-party empathetic dialogue learning, the aspects that help SDMPED achieve the state-of-the-art performance.
Lingyu Zhu 0007, Zhengkun Zhang, Jun Wang 0023, Haiying Wu, Zhenglu Yang
ACL (1)6
2022 Modeling Temporal-Modal Entity Graph for Procedural Multimodal Machine Comprehension
abstract
Huibin Zhang, Zhengkun Zhang, Yao Zhang, Jun Wang, Yufan Li, Ning Jiang, Xin Wei, Zhenglu Yang. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022.
Huibin Zhang, Zhengkun Zhang, Jun Wang 0023, Zhenglu Yang
ACL (1)8
2022 Improving Self-Supervised Learning for Speech Recognition with Intermediate Layer Supervision
abstract
Recently, pioneer work finds that self-supervised pre-training methods can improve multiple downstream speech tasks, because the model utilizes bottom layers to learn speaker-related information and top layers to encode content-related information. Since the network capacity is limited, we believe the speech recognition performance could be further improved if the model is dedicated to audio content information learning. To this end, we propose Intermediate Layer Supervision for Self-Supervised Learning (ILS-SSL), which forces the model to concentrate on content information as much as possible by adding an additional SSL loss on the intermediate layers. Experiments on LibriSpeech test-other set show that our method outperforms HuBERT significantly, which achieves a 23.5%/11.6% relative word error rate reduction in the w/o language model setting for Base/Large models. Detailed analysis shows the bottom layers of our model have a better correlation with phonetic units, which is consistent with our intuition and explains the success of our method for ASR. We will release our code and model at https://github.com/microsoft/UniSpeech.
Chengyi Wang 0002, Yu Wu 0012, Sanyuan Chen, Shujie Liu 0001, Jinyu Li 0001, Yao Qian, Zhenglu Yang
ICASSP7
2022 Interacting with Non-Cooperative User: A New Paradigm for Proactive Dialogue Policy
abstract
Proactive dialogue system is able to lead the conversation to a goal topic and has advantaged potential in bargain, persuasion, and negotiation. Current corpus-based learning manner limits its practical application in real-world scenarios. To this end, we contribute to advancing the study of the proactive dialogue policy to a more natural and challenging setting, i.e., interacting dynamically with users. Further, we call attention to the non-cooperative user behavior - the user talks about off-path topics when he/she is not satisfied with the previous topics introduced by the agent. We argue that the targets of reaching the goal topic quickly and maintaining a high user satisfaction are not always converged, because the topics close to the goal and the topics user preferred may not be the same. Towards this issue, we propose a new solution named I-Pro that can learn Proactive policy in the Interactive setting. Specifically, we learn the trade-off via a learned goal weight, which consists of four factors (dialogue turn, goal completion difficulty, user satisfaction estimation, and cooperative degree). The experimental results demonstrate I-Pro significantly outperforms baselines in terms of effectiveness and interpretability.
Wenqiang Lei, Feifan Song 0001, Hongru Liang, Jiaxin Mao, Jiancheng Lv 0001, Zhenglu Yang, Tat-Seng Chua
SIGIR7
2022 Deep convolutional recurrent model for region recommendation with spatial and temporal contexts
Hengpeng Xu, Wenjian Ding, Wei Shen 0004, Jun Wang 0023, Zhenglu Yang
Ad Hoc Networks5
2022 Domain classifier-based transfer learning for visual attention prediction
Zhiwen Zhang 0004, Feng Duan 0006, Cesar F. Caiafa, Jordi Solé i Casals, Zhenglu Yang, Zhe Sun 0009
World Wide Web5
2021 News Content Completion with Location-Aware Image Selection
Zhengkun Zhang, Jun Wang 0023, Adam Jatowt, Zhe Sun 0009, Shao-Ping Lu, Zhenglu Yang
AAAI6
2021 LAMS: A Location-aware Approach for Multimodal Summarization (Student Abstract)
abstract
Multimodal summarization aims to refine salient information from multiple modalities, among which texts and images are two mostly discussed ones. In recent years, many fantastic works have emerged in this field by modeling image-text interactions; however, they neglect the fact that most of multimodal documents have been elaborately organized by their writers. This means that a critical organized factor has long been short of enough attention, that is, image locations, which may carry illuminating information and imply the key contents of a document. To address this issue, we propose a location-aware approach for multimodal summarization (LAMS) based on Transformer. We investigate image locations for multimodal summarization via a stack of multimodal fusion block, which can formulate the high-order interactions among images and texts. An extensive experimental study on an extended multimodal dataset validates the superior summarization performance of the proposed model.
Zhengkun Zhang, Jun Wang 0023, Zhe Sun 0009, Zhenglu Yang
AAAI4
2021 Generalized Relation Learning with Semantic Correlation Awareness for Link Prediction
abstract
Developing link prediction models to automatically complete knowledge graphs has recently been the focus of significant research interest. The current methods for the link prediction task have two natural problems: 1) the relation distributions in KGs are usually unbalanced, and 2) there are many unseen relations that occur in practical situations. These two problems limit the training effectiveness and practical applications of the existing link prediction models. We advocate a holistic understanding of KGs and we propose in this work a unified Generalized Relation Learning framework GRL to address the above two problems, which can be plugged into existing link prediction models. GRL conducts a generalized relation learning, which is aware of semantic correlations between relations that serve as a bridge to connect semantically similar relations. After training with GRL, the closeness of semantically similar relations in vector space and the discrimination of dissimilar relations are improved. We perform comprehensive experiments on six benchmarks to demonstrate the superior capability of GRL in the link prediction task. In particular, GRL is found to enhance the existing link prediction models making them insensitive to unbalanced relation distributions and capable of learning unseen relations.
Jun Wang 0023, Hongru Liang, Wenqiang Lei, Zhe Sun 0009, Adam Jatowt, Zhenglu Yang
AAAI8
2021 RepSum: Unsupervised Dialogue Summarization based on Replacement Strategy
abstract
Xiyan Fu, Yating Zhang, Tianyi Wang, Xiaozhong Liu, Changlong Sun, Zhenglu Yang. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Xiyan Fu, Xiaozhong Liu 0001, Changlong Sun, Zhenglu Yang
ACL/IJCNLP (1)6
2021 GMH: A General Multi-hop Reasoning Model for KG Completion
abstract
Knowledge graphs are essential for numerous downstream natural language processing applications, but are typically incomplete with many facts missing.This results in research efforts on multi-hop reasoning task, which can be formulated as a search process and current models typically perform short distance reasoning.However, the long-distance reasoning is also vital with the ability to connect the superficially unrelated entities.To the best of our knowledge, there lacks a general framework that approaches multi-hop reasoning in mixed long-short distance reasoning scenarios.We argue that there are two key issues for a general multi-hop reasoning model: i) where to go, and ii) when to stop.Therefore, we propose a general model which resolves the issues with three modules: 1) the local-global knowledge module to estimate the possible paths, 2) the differentiated action dropout module to explore a diverse set of paths, and 3) the adaptive stopping search module to avoid over searching.The comprehensive results on three datasets demonstrate the superiority of our model with significant improvements against baselines in both short and long distance reasoning scenarios.
Hongru Liang, Adam Jatowt, Wenqiang Lei, Zhenglu Yang
EMNLP (1)7
2021 Deep Symmetric Network for Underexposed Image Enhancement with Recurrent Attentional Learning
abstract
Underexposed image enhancement is of importance in many research domains. In this paper, we take this problem as image feature transformation between the underexposed image and its paired enhanced version, and we propose a deep symmetric network for the issue. Our symmetric network adapts invertible neural networks (INN) for bidirectional feature learning between images, and to ensure the mutual propagation invertible we specifically construct two pairs of encoder-decoder with the same pretrained parameters. This invertible mechanism with bidirectional feature transformations enable us to both avoid colour bias and recover the content effectively for image enhancement. In addition, we propose a new recurrent residual-attention module (RRAM), where the recurrent learning network is designed to gradually perform the desired colour adjustments. Ablation experiments are executed to show the role of each component of our new architecture. We conduct a large number of experiments on two datasets to demonstrate that our method achieves the state-of-the-art effect in underexposed image enhancement. Code is available at https://www.shaopinglu.net/proj-iccv21/ImageEnhancement.html.
Shao-Ping Lu, Tao Chen 0015, Zhenglu Yang, Ariel Shamir
ICCV4
2021 MM-AVS: A Full-Scale Dataset for Multi-modal Summarization
abstract
Multimodal summarization becomes increasingly significant as it is the basis for question answering, Web search, and many other downstream tasks.However, its learning materials have been lacking a holistic organization by integrating resources from various modalities, thereby lagging behind the research progress of this field.In this study, we present a full-scale multimodal dataset comprehensively gathering documents, summaries, images, captions, videos, audios, transcripts, and titles in English from CNN and Daily Mail.To our best knowledge, this is the first collection that spans all modalities and nearly comprises all types of materials available in this community.In addition, we devise a baseline model based on the novel dataset, which employs a newly proposed Jump-Attention mechanism based on transcripts.The experimental results validate the important assistance role of the external information for multimodal summarization.
Xiyan Fu, Jun Wang 0023, Zhenglu Yang
NAACL-HLT3
2021 Joint Open Knowledge Base Canonicalization and Linking
abstract
Open Information Extraction (OIE) methods extract a large number of OIE triples (noun phrase, relation phrase, noun phrase) from text, which compose large Open Knowledge Bases (OKBs). However, noun phrases (NPs) and relation phrases (RPs) in OKBs are not canonicalized and often appear in different paraphrased textual variants, which leads to redundant and ambiguous facts. To address this problem, there are two related tasks: OKB canonicalization (i.e., convert NPs and RPs to canonicalized form) and OKB linking (i.e., link NPs and RPs with their corresponding entities and relations in a curated Knowledge Base (e.g., DBPedia). These two tasks are tightly coupled, and one task can benefit significantly from the other. However, they have been studied in isolation so far. In this paper, we explore the task of joint OKB canonicalization and linking for the first time, and propose a novel framework JOCL based on factor graph model to make them reinforce each other. JOCL is flexible enough to combine different signals from both tasks, and able to extend to fit any new signals. A thorough experimental study over two large scale OIE triple data sets shows that our framework outperforms all the baseline methods for the task of OKB canonicalization (OKB linking) in terms of average F1 (accuracy).
Yinan Liu 0001, Wei Shen 0004, Yuanfei Wang, Jianyong Wang 0001, Zhenglu Yang, Xiaojie Yuan
SIGMOD Conference5
2021 TRFR: A ternary relation link prediction framework on Knowledge graphs
Hengpeng Xu, Xingxing Wu, Zhenglu Yang
Ad Hoc Networks5
2021 Component-mixing strategy: A decomposition-based data augmentation algorithm for motor imagery signals
Binghua Li 0001, Zhiwen Zhang 0004, Feng Duan 0006, Zhenglu Yang, Qibin Zhao, Zhe Sun 0009, Jordi Solé i Casals
Neurocomputing4
2021 From text to graph: a general transition-based AMR parsing using neural network
Yanhui Gu, Weilan Luo, Guandong Xu, Zhenglu Yang, Junsheng Zhou, Weiguang Qu
Neural Comput. Appl.5
2021 Named Entity Location Prediction Combining Twitter and Web
abstract
Knowledge bases are critical to many applications. However, they are greatly incomplete. Enriching knowledge bases with new entities and new location attributes becomes increasingly important. Given a named entity with tweets and Web documents where the entity appears, we aim to predict the entity city-level location combining the geographical location knowledge embedded in both Twitter and Web. This task is helpful for knowledge base enrichment and tweet location prediction. In this paper we propose NELPTW, the first unsupervised framework forNamedEntityLocationPrediction by leveraging the knowledge fromTwitter andWeb. Based on each data source, NELPTW utilizes a linear function ranking model to generate several rankings to the candidate location set for each entity. To combine the knowledge from two sources which have different reliability and importance for the location prediction, an unsupervised rank aggregation algorithm is developed to aggregate multiple rankings for each entity to obtain a better ranking. A learning algorithm based on the EM method is proposed to automatically learn the parameters of the ranking model without requiring any training labels. The experimental results over a real world Twitter and Web data set show that our framework significantly outperforms the baselines in terms of accuracy.
Yinan Liu 0001, Wei Shen 0004, Zonghai Yao, Jianyong Wang 0001, Zhenglu Yang, Xiaojie Yuan
IEEE Trans. Knowl. Data Eng.5
2021 Multiple interleaving interests modeling of sequential user behaviors in e-commerce platform
Yuqiang Han, Qian Li 0016, Hucheng Zhou, Zhenglu Yang, Jian Wu 0001
World Wide Web5
2020 Document Summarization with VHTM: Variational Hierarchical Topic-Aware Mechanism
abstract
Automatic text summarization focuses on distilling summary information from texts. This research field has been considerably explored over the past decades because of its significant role in many natural language processing tasks; however, two challenging issues block its further development: (1) how to yield a summarization model embedding topic inference rather than extending with a pre-trained one and (2) how to merge the latent topics into diverse granularity levels. In this study, we propose a variational hierarchical model to holistically address both issues, dubbed VHTM. Different from the previous work assisted by a pre-trained single-grained topic model, VHTM is the first attempt to jointly accomplish summarization with topic inference via variational encoder-decoder and merge topics into multi-grained levels through topic embedding and attention. Comprehensive experiments validate the superior performance of VHTM compared with the baselines, accompanying with semantically consistent topics.
Xiyan Fu, Jun Wang 0023, Jinghan Zhang 0004, Jinmao Wei 0001, Zhenglu Yang
AAAI5
2020 Bridging the Gap between Pre-Training and Fine-Tuning for End-to-End Speech Translation
abstract
End-to-end speech translation, a hot topic in recent years, aims to translate a segment of audio into a specific language with an end-to-end model. Conventional approaches employ multi-task learning and pre-training methods for this task, but they suffer from the huge gap between pre-training and fine-tuning. To address these issues, we propose a Tandem Connectionist Encoding Network (TCEN) which bridges the gap by reusing all subnets in fine-tuning, keeping the roles of subnets consistent, and pre-training the attention module. Furthermore, we propose two simple but effective methods to guarantee the speech encoder outputs and the MT encoder inputs are consistent in terms of semantic representation and sequence length. Experimental results show that our model leads to significant improvements in En-De and En-Fr translation irrespective of the backbones.
Chengyi Wang 0002, Yu Wu 0012, Shujie Liu 0001, Zhenglu Yang, Ming Zhou 0001
AAAI4
2020 Multi-Point Semantic Representation for Intent Classification
abstract
Detecting user intents from utterances is the basis of natural language understanding (NLU) task. To understand the meaning of utterances, some work focuses on fully representing utterances via semantic parsing in which annotation cost is labor-intentsive. While some researchers simply view this as intent classification or frequently asked questions (FAQs) retrieval, they do not leverage the shared utterances among different intents. We propose a simple and novel multi-point semantic representation framework with relatively low annotation cost to leverage the fine-grained factor information, decomposing queries into four factors, i.e., topic, predicate, object/condition, query type. Besides, we propose a compositional intent bi-attention model under multi-task learning with three kinds of attention mechanisms among queries, labels and factors, which jointly combines coarse-grained intent and fine-grained factor information. Extensive experiments show that our framework and model significantly outperform several state-of-the-art approaches with an improvement of 1.35%-2.47% in terms of accuracy.
Jinghan Zhang 0004, Yuxiao Ye, Yue Zhang 0004, Likun Qiu, Yang Li 0218, Zhenglu Yang, Jian Sun 0021
AAAI7
2020 Curriculum Pre-training for End-to-End Speech Translation
abstract
End-to-end speech translation poses a heavy burden on the encoder because it has to transcribe, understand, and learn cross-lingual semantics simultaneously.To obtain a powerful encoder, traditional methods pre-train it on ASR data to capture speech features.However, we argue that pre-training the encoder only through simple speech recognition is not enough, and high-level linguistic knowledge should be considered.Inspired by this, we propose a curriculum pre-training method that includes an elementary course for transcription learning and two advanced courses for understanding the utterance and mapping words in two languages.The difficulty of these courses is gradually increasing.Experiments show that our curriculum pre-training method leads to significant improvements on En-De and En-Fr speech translation benchmarks.
Chengyi Wang 0002, Yu Wu 0012, Shujie Liu 0001, Ming Zhou 0001, Zhenglu Yang
ACL5
2020 PiRhDy: Learning Pitch-, Rhythm-, and Dynamics-aware Embeddings for Symbolic Music
abstract
Definitive embeddings remain a fundamental challenge of computational musicology for symbolic music in deep learning today. Analogous to natural language, music can be modeled as a sequence of tokens. This motivates the majority of existing solutions to explore the utilization of word embedding models to build music embeddings. However, music differs from natural languages in two key aspects: (1) musical token is multi-faceted -- it comprises of pitch, rhythm and dynamics information; and (2) musical context is two-dimensional -- each musical token is dependent on both melodic and harmonic contexts. In this work, we provide a comprehensive solution by proposing a novel framework named PiRhDy that integrates pitch, rhythm, and dynamics information seamlessly. PiRhDy adopts a hierarchical strategy which can be decomposed into two steps: (1) token (i.e., note event) modeling, which separately represents pitch, rhythm, and dynamics and integrates them into a single token embedding; and (2) context modeling, which utilizes melodic and harmonic knowledge to train the token embedding. A thorough study was made on each component and sub-strategy of PiRhDy.We further validate our embeddings in three downstream tasks -- melody completion, accompaniment suggestion, and genre classification. Results indicate a significant advancement of the neural approach towards symbolic music as well as PiRhDy's potential as a pretrained tool for a broad range of symbolic music applications.
Hongru Liang, Wenqiang Lei, Paul Y. Chan, Zhenglu Yang, Maosong Sun 0001, Tat-Seng Chua
ACM Multimedia4
2019 Attention Optimization for Abstractive Document Summarization
abstract
Min Gui, Junfeng Tian, Rui Wang, Zhenglu Yang. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Min Gui, Rui Wang 0005, Zhenglu Yang
EMNLP/IJCNLP (1)4
2019 A general framework for learning prosodic-enhanced representation of rap lyrics
Hongru Liang, Haozheng Wang, Qian Li 0016, Jun Wang 0023, Guandong Xu, Jinmao Wei 0001, Zhenglu Yang
World Wide Web8
2018 A Multi-Attention based Neural Network with External Knowledge for Story Ending Predicting Task
abstract
Enabling a mechanism to understand a temporal story and predict its ending is an interesting issue that has attracted considerable attention, as in case of the ROC Story Cloze Task (SCT). In this paper, we develop a multi-attention-based neural network (MANN) with well-designed optimizations, like Highway Network, and concatenated features with embedding representations into the hierarchical neural network model. Considering the particulars of the specific task, we thoughtfully extend MANN with external knowledge resources, exceeding state-of-the-art results obviously. Furthermore, we develop a thorough understanding of our model through a careful hand analysis on a subset of the stories. We identify what traits of MANN contribute to its outperformance and how external knowledge is obtained in such an ending prediction task.
Qian Li 0016, Jinmao Wei 0001, Yanhui Gu, Adam Jatowt, Zhenglu Yang
COLING6
2018 JTAV: Jointly Learning Social Media Content Representation by Fusing Textual, Acoustic, and Visual Features
abstract
Learning social media content is the basis of many real-world applications, including information retrieval and recommendation systems, among others. In contrast with previous works that focus mainly on single modal or bi-modal learning, we propose to learn social media content by fusing jointly textual, acoustic, and visual information (JTAV). Effective strategies are proposed to extract fine-grained features of each modality, that is, attBiGRU and DCRNN. We also introduce cross-modal fusion and attentive pooling techniques to integrate multi-modal information comprehensively. Extensive experimental evaluation conducted on real-world datasets demonstrate our proposed model outperforms the state-of-the-art approaches by a large margin.
Hongru Liang, Haozheng Wang, Jun Wang 0023, Shaodi You, Zhe Sun 0009, Jinmao Wei 0001, Zhenglu Yang
COLING7
2018 Neural Framework for Joint Evolution Modeling of User Feedback and Social Links in Dynamic Social Networks
abstract
Modeling the evolution of user feedback and social links in dynamic social networks is of considerable significance, because it is the basis of many applications, including recommendation systems and user behavior analyses. Most of the existing methods in this area model user behaviors separately and consider only certain aspects of this problem, such as dynamic preferences of users, dynamic attributes of items, evolutions of social networks, and their partial integration. This work proposes a comprehensive general neural framework with several optimal strategies to jointly model the evolution of user feedback and social links. The framework considers the dynamic user preferences, dynamic item attributes, and time-dependent social links in time evolving social networks. Experimental results conducted on two real-world datasets demonstrate that our proposed model performs remarkably better than state-of-the-art methods.
Peizhi Wu, Xiaojie Yuan, Adam Jatowt, Zhenglu Yang
IJCAI5
2018 Probabilistic Topic and Role Model for Information Diffusion in Social Network
Hengpeng Xu, Jinmao Wei 0001, Zhenglu Yang, Jianhua Ruan, Jun Wang 0023
PAKDD (2)3
2018 HAVAE: Learning Prosodic-Enhanced Representations of Rap Lyrics
Hongru Liang, Qian Li 0016, Haozheng Wang, Jun Wang 0023, Zhe Sun 0009, Jinmao Wei 0001, Zhenglu Yang
PRICAI (1)8
2018 Unsupervised learning of semantic representation for documents with the law of total probability
abstract
Abstract The semantic information of documents needs to be represented because it is the basis for many applications, such as document summarization, web search, and text analysis. Although many studies have explored this problem by enriching document vectors with the relatedness of the words involved, the performance remains far from satisfactory because the physical boundaries of documents hinder the evaluation of the relatedness between words. To address this problem, we propose an effective approach to further infer the implicit relatedness between words via their common related words. To avoid overestimation of the implicit relatedness, we restrict the inference in terms of the marginal probabilities of the words based on the law of total probability. The proposed method measures the relatedness between words, which is confirmed theoretically and experimentally. Thorough evaluation on real datasets illustrates that significant improvement on document clustering has been achieved with the proposed method compared with state-of-the-art methods.
Jinmao Wei 0001, Zhenglu Yang
Nat. Lang. Eng.3
2018 SHINE+: A General Framework for Domain-Specific Entity Linking with Heterogeneous Information Networks
abstract
Heterogeneous information networks that consist of multi-type, interconnected objects are becoming increasingly popular, such as social media networks and bibliographic networks. The task of linking named entity mentions detected from unstructured Web text with their corresponding entities in a heterogeneous information network is of practical importance for the problem of information network population. This task is challenging due to name ambiguity and limited knowledge existing in the network. Most existing entity linking methods focus on linking entities with Wikipedia and cannot be applied to our task. In this paper, we present SHINE+, a general framework for linking named entitieS in Web free text with a Heterogeneous I nformation NEtwork. We propose a probabilistic linking model, which unifies an entity popularity model with an entity object model. As the entity knowledge contained in the information network is insufficient, we propose a knowledge population algorithm to iteratively enrich the network entity knowledge by leveraging the context information of mentions mapped by the linking model with high confidence, which subsequently boosts the linking performance. Experimental results over two real heterogeneous information networks (i.e., DBLP and IMDb) demonstrate the effectiveness and efficiency of our proposed framework in comparison with the baselines.
Wei Shen 0004, Jiawei Han 0001, Jianyong Wang 0001, Xiaojie Yuan, Zhenglu Yang
IEEE Trans. Knowl. Data Eng.5
2018 An enhanced short text categorization model with deep abundant representation
Yanhui Gu, Guandong Xu, Zhenglu Yang, Junsheng Zhou, Weiguang Qu
World Wide Web5
2017 Unsupervised Feature Selection with Joint Clustering Analysis
abstract
Unsupervised feature selection has raised considerable interests in the past decade, due to its remarkable performance in reducing dimensionality without any prior class information. Preserving reliable locality information and achieving excellent cluster separation are two critical issues for unsupervised feature selection. However, existing methods cannot tackle two issues simultaneously. To address the problems, we propose a novel unsupervised approach that integrates sparse feature selection and robust joint clustering analysis. The joint clustering analysis seamlessly unifies the spectral clustering and the orthogonal basis clustering. Specifically, a probabilistic neighborhood graph is utilized to preserve reliable locality information in the spectral clustering, and an orthogonal basis matrix is incorporated to achieve excellent cluster separation in the orthogonal basis clustering. A compact and effective iterative algorithm is designed to optimize the proposed selection framework. Extensive experiments on both synthetic data and real-world data validate the effectiveness of our approach under various evaluation indices.
Shuai An 0003, Jun Wang 0023, Jinmao Wei 0001, Zhenglu Yang
CIKM4
2017 Effective Representing of Information Network by Variational Autoencoder
abstract
Network representation is the basis of many applications and of extensive interest in various fields, such as information retrieval, social network analysis, and recommendation systems. Most previous methods for network representation only consider the incomplete aspects of a problem, including link structure, node information, and partial integration. The present study proposes a deep network representation model that seamlessly integrates the text information and structure of a network. Our model captures highly non-linear relationships between nodes and complex features of a network by exploiting the variational autoencoder (VAE), which is a deep unsupervised generation algorithm. We also merge the representation learned with a paragraph vector model and that learned with the VAE to obtain the network representation that preserves both structure and text information. We conduct comprehensive empirical experiments on benchmark datasets and find our model performs better than state-of-the-art techniques by a large margin.
Haozheng Wang, Zhenglu Yang
IJCAI3
2017 Automatic Generation of Grounded Visual Questions
abstract
In this paper, we propose the first model to be able to generate visually grounded questions with diverse types for a single image. Visual question generation is an emerging topic which aims to ask questions in natural language based on visual input. To the best of our knowledge, it lacks automatic methods to generate meaningful questions with various types for the same visual input. To circumvent the problem, we propose a model that automatically generates visually grounded questions with varying types. Our model takes as input both images and the captions generated by a dense caption model, samples the most probable question types, and generates the questions in sequel. The experimental results on two real world datasets show that our model outperforms the strongest baseline in terms of both correctness and diversity with a wide margin.
Lizhen Qu, Shaodi You, Zhenglu Yang, Jiawan Zhang
IJCAI4
2017 Discrimination Structure Complementarity-Based Feature Selection
abstract
Feature selection is crucial, particularly for processing high‐dimensional data. Existing selection methods generally compute a discriminant value for a feature with respect to class variable to indicate its classification ability. However, a scalar value can hardly reveal the multifaceted classification abilities of a feature for different subproblems of a complicated multiclass problem. In view of this, we propose to select features based on discrimination structure complementarity. To this end, the classification abilities of a feature for different subproblems are evaluated individually. Consequently, a discrimination structure vector can be obtained to indicate if the feature is discriminative respectively for different subproblems. Based on discrimination structure, indispensable and dispensable features (ID‐features for short) are defined. In selection process, the ID‐features, which are complementary in discrimination structure to the selected ones, are selected. The proposed method tries to equally treat all subproblems and hence can avoid falling into the pitfall that the discriminative features for difficult subproblems are prone to be covered by the features for easy ones in multi‐class classification. Two algorithms are developed and compared with several feature selection methods using some open data sets. Experimental results demonstrate the effectiveness of the proposed method.
Jinmao Wei 0001, Zhenglu Yang
Comput. Intell.3
2017 Feature Selection by Maximizing Independent Classification Information
abstract
Feature selection approaches based on mutual information can be roughly categorized into two groups. The first group minimizes the redundancy of features between each other. The second group maximizes the new classification information of features providing for the selected subset. A critical issue is that large new information does not signify little redundancy, and vice versa. Features with large new information but with high redundancy may be selected by the second group, and features with low redundancy but with little relevance with classes may be highly scored by the first group. Existing approaches fail to balance the importance of both terms. As such, a new information term denoted as Independent Classification Information is proposed in this paper. It assembles the newly provided information and the preserved information negatively correlated with the redundant information. Redundancy and new information are properly unified and equally treated in the new term. This strategy helps find the predictive features providing large new information and little redundancy. Moreover, independent classification information is proved as a loose upper bound of the total classification information of feature subset. Its maximization is conducive to achieve a high global discriminative performance. Comprehensive experiments demonstrate the effectiveness of the new approach.
Jun Wang 0023, Jinmao Wei 0001, Zhenglu Yang, Shu-Qin Wang
IEEE Trans. Knowl. Data Eng.3
2016 Supervised Feature Selection by Preserving Class Correlation
abstract
Feature selection is an effective technique for dimension reduction, which assesses the importance of features and constructs an optimal feature subspace suitable for recognition task. Two recognition scenarios, i.e., single-label learning and multi-label learning, pose different challenges for feature selection. For the single-label task, how to accurately measure and reduce feature redundancy is crucial. For the multi-label task, how to effectively exploit class correlation information during selection is critical. However, both issues cannot be simultaneously resolved by any existing selection methods. In this paper, we propose effective supervised feature selection techniques to address the problems. The original class correlation information in the reduced feature space is preserved, and meanwhile the feature redundancy for classification is alleviated. To the best of our knowledge, this study is the first attempt to accomplish both recognition tasks in a unified framework. Comprehensive experimental evaluations on artificial, single-label, and multi-label data sets demonstrate the effectiveness of the new approach.
Jun Wang 0023, Jinmao Wei 0001, Zhenglu Yang
CIKM3
2016 Joint Probability Consistent Relation Analysis for Document Representation
Jinmao Wei 0001, Zhenglu Yang
DASFAA (1)3
2015 Extended Strategies for Document Clustering with Word Co-occurrences
Jinmao Wei 0001, Zhenglu Yang
APWeb3
2015 Enriching Document Representation with the Deviations of Word Co-occurrence Frequencies
Jinmao Wei 0001, Zhenglu Yang
ICA3PP (2)3
2014 Identifying Domain-Dependent Influential Microblog Users: A Post-Feature Based Approach
abstract
Users of a social network like to follow the posts published by influential users. Such posts usually are delivered quickly and thus will produce a strong influence on public opinions. In this paper, we focus on the problem of identifying domain-dependent influential users(or topic experts). Some of traditional approaches are based on the post contents of users user’s to identify influential users, which may be biased by spammers who try to make posts related to some topics through a simple copy and paste. Others make use of user authentication information given by a service platform or user self description (introduction or label) in finding influential users. However, what users have published is not necessarily related to what they have registed and described. In addition, if there is no comments from other users, it’s less objective to assess a user’s post quality. To improve effectiveness of recognizing influential users in a topic of microblogs, we propose a post-feature based approach which is supplementary to post-content based approaches. Our experimental results show that the post-feature based approach produces relatively higher precision than that of the content based approach.
Lin Li 0001, Guandong Xu, Zhenglu Yang
AAAI4
2014 Exploration on efficient similar sentences extraction
Yanhui Gu, Zhenglu Yang, Guandong Xu, Miyuki Nakano, Masashi Toyoda, Masaru Kitsuregawa
World Wide Web2
2013 QUBiC: An adaptive approach to query-based recommendation
Lin Li 0001, Luo Zhong, Zhenglu Yang, Masaru Kitsuregawa
J. Intell. Inf. Syst.3
2013 An efficient approach to suggesting topically related web queries using hidden topic model
Lin Li 0001, Guandong Xu, Zhenglu Yang, Peter Dolog, Yanchun Zhang, Masaru Kitsuregawa
World Wide Web3
2012 Recommending Related Microblogs: A Comparison Between Topic and WordNet based Approaches
abstract
Computing similarity between short microblogs is an important step in microblog recommendation. In this paper, we investigate a topic based approach and a WordNet based approach to estimate similarity scores between microblogs and recommend top related ones to users. Empirical study is conducted to compare their recommendation effectiveness using two evaluation measures. The results show that the WordNet based approach has relatively higher precision than that of the topic based approach using 548 tweets as dataset. In addition, the Kendall tau distance between two lists recommended by WordNet and topic approaches is calculated. Its average of all the 548 pair lists tells us the two approaches have the relative high disaccord in the ranking of related tweets.
Lin Li 0001, Guandong Xu, Zhenglu Yang, Masaru Kitsuregawa
AAAI4
2012 Towards Efficient Similar Sentences Extraction
Yanhui Gu, Zhenglu Yang, Miyuki Nakano, Masaru Kitsuregawa
IDEAL2
2011 Efficient Searching Top-k Semantic Similar Words
abstract
Measuring the semantic meaning between words is an important issue because it is the basis for many applications, such as word sense disambiguation, document summarization, and so forth. Although it has been explored for several decades, most of the studies focus on improving the effectiveness of the problem, i.e., precision and recall. In this paper, we propose to address the efficiency issue, that given a collection of words, how to efficiently discover the top-k most semantic similar words to the query. This issue is very important for real applications yet the existing state-of-the-art strategies cannot satisfy users with reasonable performance. Efficient strategies on searching top-k semantic similar words are proposed. We provide an extensive comparative experimental evaluation demonstrating the advantages of the introduced strategies over the state-of-the-art approaches.
Zhenglu Yang, Masaru Kitsuregawa
IJCAI1
2011 TOAST: A Topic-Oriented Tag-Based Recommender System
Guandong Xu, Yanhui Gu, Yanchun Zhang, Zhenglu Yang, Masaru Kitsuregawa
WISE4
2010 Fast Algorithms for Top-k Approximate String Matching
abstract
Top-k approximate querying on string collections is an important data analysis tool for many applications, and it has been exhaustively studied. However, the scale of the problem has increased dramatically because of the prevalence of the Web. In this paper, we aim to explore the efficient top-k similar string matching problem. Several efficient strategies are introduced, such as length aware and adaptive q-gram selection. We present a general q-gram based framework and propose two efficient algorithms based on the strategies introduced. Our techniques are experimentally evaluated on three real data sets and show a superior performance.
Zhenglu Yang, Jianjun Yu, Masaru Kitsuregawa
AAAI1
2010 Multi-View Clustering with Web and Linguistic Features for Relation Extraction
abstract
Binary semantic relation extraction is particularly useful for various NLP and Web applications. Currently Web-based methods and Linguistic-based methods are two types of leading methods for semantic relation extraction task. With a novel view on integrating linguistic analysis on local text with Web frequent information, we propose a multi-view co-clustering approach for semantic relation extraction. One is feature clustering by automatically learning clustering functions for Web features, linguistic features simultaneously based on a subset of entity pairs. The other is relation clustering, using the feature clustering functions to define learning function for relation extraction. Our experiments demonstrate the superiority of our clustering approach comparing with several state-of-the-art clustering methods.
Yulan Yan, Haibo Li 0002, Yutaka Matsuo, Zhenglu Yang, Mitsuru Ishizuka
APWeb4
2010 Fires on the Web: Towards Efficient Exploring Historical Web Graphs
Zhenglu Yang, Jeffrey Xu Yu, Zheng Liu 0001, Masaru Kitsuregawa
DASFAA (1)1
2009 Unsupervised Relation Extraction by Mining Wikipedia Texts Using Information from the Web
Yulan Yan, Naoaki Okazaki, Yutaka Matsuo, Zhenglu Yang, Mitsuru Ishizuka
ACL/IJCNLP4
2008 Query-URL Bipartite Based Approach to Personalized Query Recommendation
Lin Li 0001, Zhenglu Yang, Ling Liu 0001, Masaru Kitsuregawa
AAAI2
2008 Efficient Querying Relaxed Dominant Relationship between Product Items Based on Rank Aggregation
Zhenglu Yang, Lin Li 0001, Masaru Kitsuregawa
AAAI1
2008 Using Ontology-Based User Preferences to Aggregate Rank Lists in Web Search
Lin Li 0001, Zhenglu Yang, Masaru Kitsuregawa
PAKDD2
2007 Aggregating User-Centered Rankings to Improve Web Search
Lin Li 0001, Zhenglu Yang, Masaru Kitsuregawa
AAAI2
2007 Towards Efficient Dominant Relationship Exploration of the Product Items on the Web
Zhenglu Yang, Lin Li 0001, Masaru Kitsuregawa
AAAI1
2007 LAPIN: Effective Sequential Pattern Mining Algorithms by Last Position Induction for Dense Databases
Zhenglu Yang, Masaru Kitsuregawa
DASFAA1
2007 Towards efficient dominant relationship exploration of the product items on the web
abstract
In recent years, there has been a prevalence of search engines being employed to find useful information in the Web as they efficiently explore hyperlinks between web pages which define a natural graph structure that yields a good ranking. Unfortunately, current search engines cannot effectively rank those relational data, which exists on dynamic websites supported by online databases. In this study, to rank such structured data (i.e., find the "best" items), we propose an integrated online system consisting of compressed data structure to encode the dominant relationship of the relational data. Efficient querying strategies and updating scheme are devised to facilitate the ranking process. Extensive experiments illustrate the effectiveness and efficiency of our methods. As such, we believe the work in this poster can be complementary to traditional search engines.
Zhenglu Yang, Lin Li 0001, Masaru Kitsuregawa
WWW1
2006 An Effective System for Mining Web Log
Zhenglu Yang, Masaru Kitsuregawa
APWeb1
2006 PAID: Mining Sequential Patterns by Passed Item Deduction in Large Databases
abstract
Sequential pattern mining is very important because it is the basis of many applications. Yet how to efficiently implement the mining is difficult due to the inherent characteristic of the problem - the large size of the dataset. Although there has been a great deal of effort on sequential pattern mining in recent years, its performance is still far from satisfactory. In this paper, we have proposed a new algorithm called passed item deduced sequential pattern mining (abbreviated as PAID), which can efficiently get all the frequent sequential patterns from a large database. The main difference between our strategy and the existing works is that other algorithms accumulate the candidate support in each iteration from scratch, in contrast, PAID makes good use of the temporary results (support value) of k-length frequent patterns on discovering (k+1)-length patterns, which can reduce the search space greatly in mining sequential patterns. Our experimental results and performance studies show that PAID outperforms the previous works by meaningful margins on large datasets
Zhenglu Yang, Masaru Kitsuregawa
IDEAS1