VLDB 2026 Research / reviewers in the wild / expert
Lujun Zhao
dblp:222/7936
· DBLP profile ↗
12ranked-venue papers
4as first author
3since 2021 · last 2023
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 10 · 2 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 2 · 2 first-author · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
9 papers |
Information extraction and text analysis · 43% Language models and text generation · 31% Deep learning architectures and training · 11% | |
| Databases, data mining, and information retrieval
1 paper |
Data mining · 100% |
Topics — the 23 heaviest of 24, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Language models and text generation › text summarization
dialogue summarization |
1.0 | 2 | 2021 | Topic-Oriented Spoken Dialogue Summarization for Customer Service with Saliency-Aware Topic Modeling · AAAI 2021 Unsupervised Summarization for Chat Logs with Topic-Oriented Ranking and Context-Aware Auto-Encoders · AAAI 2021 |
Natural language and speech › Information extraction and text analysis › topic model
neural topic model |
0.5 | 1 | 2021 | Topic-Oriented Spoken Dialogue Summarization for Customer Service with Saliency-Aware Topic Modeling · AAAI 2021 |
Natural language and speech › Language models and text generation
text summarization |
0.5 | 1 | 2021 | Unsupervised Summarization for Chat Logs with Topic-Oriented Ranking and Context-Aware Auto-Encoders · AAAI 2021 |
Natural language and speech › Language models and text generation › text summarization › controllable summarization
topic-aware summarization |
0.5 | 1 | 2021 | Topic-Oriented Spoken Dialogue Summarization for Customer Service with Saliency-Aware Topic Modeling · AAAI 2021 |
Natural language and speech › Information extraction and text analysis
topic model |
0.5 | 1 | 2021 | Topic-Oriented Spoken Dialogue Summarization for Customer Service with Saliency-Aware Topic Modeling · AAAI 2021 |
Natural language and speech › Language models and text generation › text summarization
unsupervised summarization |
0.5 | 1 | 2021 | Unsupervised Summarization for Chat Logs with Topic-Oriented Ranking and Context-Aware Auto-Encoders · AAAI 2021 |
Natural language and speech › Information extraction and text analysis › sentiment analysis
aspect-based sentiment analysis |
0.4 | 1 | 2019 | Cold-Start Aware Deep Memory Network for Multi-Entity Aspect-Based Sentiment Analysis · IJCAI 2019 |
Natural language and speech › Information extraction and text analysis › named entity recognition
chinese named entity recognition |
0.4 | 1 | 2019 | CNN-Based Chinese NER with Lexicon Rethinking · IJCAI 2019 |
Natural language and speech › Information extraction and text analysis
customer satisfaction analysis |
0.4 | 1 | 2019 | Using Customer Service Dialogues for Satisfaction Analysis with Context-Assisted Multiple Instance Learning · EMNLP/IJCNLP (1) 2019 |
Machine learning › Graph learning › relation modeling
dependency modeling |
0.4 | 1 | 2019 | Long Short-Term Memory with Dynamic Skip Connections · AAAI 2019 |
Natural language and speech › Information extraction and text analysis
dialogue analysis |
0.4 | 1 | 2019 | Using Customer Service Dialogues for Satisfaction Analysis with Context-Assisted Multiple Instance Learning · EMNLP/IJCNLP (1) 2019 |
Natural language and speech › Question answering and dialogue systems › dialogue generation
dialogue response generation |
0.4 | 1 | 2019 | Review Response Generation in E-Commerce Platforms with External Product Information · WWW 2019 |
Machine learning › Deep learning architectures and training › recurrent neural network
LSTM |
0.4 | 1 | 2019 | Long Short-Term Memory with Dynamic Skip Connections · AAAI 2019 |
Natural language and speech › Information extraction and text analysis
named entity recognition |
0.4 | 1 | 2019 | CNN-Based Chinese NER with Lexicon Rethinking · IJCAI 2019 |
Machine learning › Deep learning architectures and training
recurrent neural network |
0.4 | 1 | 2019 | Long Short-Term Memory with Dynamic Skip Connections · AAAI 2019 |
Natural language and speech › Information extraction and text analysis
sentiment analysis |
0.4 | 1 | 2019 | Cold-Start Aware Deep Memory Network for Multi-Entity Aspect-Based Sentiment Analysis · IJCAI 2019 |
Natural language and speech › Information extraction and text analysis
sequence labeling |
0.4 | 1 | 2019 | Sequence Labeling With Deep Gated Dual Path CNN · IEEE ACM Trans. Audio Speech Lang. Process. 2019 |
Natural language and speech › Language models and text generation › language modeling › language model architecture
sequence-to-sequence model |
0.4 | 1 | 2019 | Review Response Generation in E-Commerce Platforms with External Product Information · WWW 2019 |
Natural language and speech › Information extraction and text analysis › word segmentation
chinese word segmentation |
0.3 | 1 | 2018 | Neural Networks Incorporating Unlabeled and Partially-labeled Data for Cross-domain Chinese Word Segmentation · IJCAI 2018 |
Natural language and speech › Question answering and dialogue systems › task-oriented dialogue
customer service dialogue |
0.3 | 2 | 2021 | Topic-Oriented Spoken Dialogue Summarization for Customer Service with Saliency-Aware Topic Modeling · AAAI 2021 Using Customer Service Dialogues for Satisfaction Analysis with Context-Assisted Multiple Instance Learning · EMNLP/IJCNLP (1) 2019 |
Data mining › text mining
topic modeling |
0.1 | 1 | 2021 | Unsupervised Summarization for Chat Logs with Topic-Oriented Ranking and Context-Aware Auto-Encoders · AAAI 2021 |
Machine learning › Deep learning architectures and training › memory-augmented neural networks
memory network |
0.1 | 1 | 2019 | Cold-Start Aware Deep Memory Network for Multi-Entity Aspect-Based Sentiment Analysis · IJCAI 2019 |
Machine learning › Deep learning architectures and training
sequence modeling |
0.1 | 1 | 2019 | Long Short-Term Memory with Dynamic Skip Connections · AAAI 2019 |
Methods — techniques the papers use, named apart from their topics
topic-oriented ranking · 1.0denoising autoencoder · 1.0reinforcement learning · 0.8convolutional neural network · 0.8two-stage summarization · 0.5saliency-aware topic modeling · 0.5multiple instance learning · 0.4lexicon rethinking · 0.4context-assisted learning · 0.4LSTM · 0.4
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | A novel model for assessing the degree of intelligent manufacturing readiness in the process industry: process-industry intelligent manufacturing readiness index (PIMRI)abstractRecently, the implementation of Industry 4.0 has become a new tendency, and it brings both opportunities and challenges to worldwide manufacturing companies. Thus, many manufacturing companies are attempting to find advanced technologies to launch intelligent manufacturing transformation. In this study, we propose a new model to measure the intelligent manufacturing readiness for the process industry, which aims to guide companies in recognizing their current stage and short slabs when carrying out intelligent manufacturing transformation. Although some models have already been reported to measure Industry 4.0 readiness and maturity, there are no models that are aimed at the process industry. This newly proposed model has six levels to describe different development stages for intelligent manufacturing. In addition, the model consists of four races, nine species, and 25 domains that are relevant to the essential businesses of companies’ daily operation and capability requirements of intelligent manufacturing. Furthermore, these 25 domains are divided into 249 characteristic items to evaluate the manufacturing readiness in detail. A questionnaire is also designed based on the proposed model to help process-industry companies easily carry out self-diagnosis. Using the new method, a case including 196 real-world process-industry companies is evaluated to introduce the method of how to use the proposed model. Overall, the proposed model provides a new way to assess the degree of intelligent manufacturing readiness for process-industry companies. Lujun Zhao, Jiaming Shao, Yuqi Qi, Jian Chu, Yiping Feng |
Frontiers Inf. Technol. Electron. Eng. | 1 |
| 2021 | Unsupervised Summarization for Chat Logs with Topic-Oriented Ranking and Context-Aware Auto-EncodersabstractAutomatic chat summarization can help people quickly grasp important information from numerous chat messages. Unlike conventional documents, chat logs usually have fragmented and evolving topics. In addition, these logs contain a quantity of elliptical and interrogative sentences, which make the chat summarization highly context dependent. In this work, we propose a novel unsupervised framework called RankAE to perform chat summarization without employing manually labeled data. RankAE consists of a topic-oriented ranking strategy that selects topic utterances according to centrality and diversity simultaneously, as well as a denoising auto-encoder that is carefully designed to generate succinct but context-informative summaries based on the selected utterances. To evaluate the proposed method, we collect a large-scale dataset of chat logs from a customer service environment and build an annotated set only for model evaluation. Experimental results show that RankAE significantly outperforms other unsupervised methods and is able to generate high-quality summaries in terms of relevance and topic coverage. Yicheng Zou, Lujun Zhao, Yangyang Kang, Zhuoren Jiang, Changlong Sun, Qi Zhang 0001, Xuanjing Huang 0001, Xiaozhong Liu 0001 |
AAAI | 3 |
| 2021 | Topic-Oriented Spoken Dialogue Summarization for Customer Service with Saliency-Aware Topic ModelingabstractIn a customer service system, dialogue summarization can boost service efficiency by automatically creating summaries for long spoken dialogues in which customers and agents try to address issues about specific topics. In this work, we focus on topic-oriented dialogue summarization, which generates highly abstractive summaries that preserve the main ideas from dialogues. In spoken dialogues, abundant dialogue noise and common semantics could obscure the underlying informative content, making the general topic modeling approaches difficult to apply. In addition, for customer service, role-specific information matters and is an indispensable part of a summary. To effectively perform topic modeling on dialogues and capture multi-role information, in this work we propose a novel topic-augmented two-stage dialogue summarizer (TDS) jointly with a saliency-aware neural topic model (SATM) for topic-oriented summarization of customer service dialogues. Comprehensive studies on a real-world Chinese customer service dataset demonstrated the superiority of our method against several strong baselines. Yicheng Zou, Lujun Zhao, Yangyang Kang, Minlong Peng, Zhuoren Jiang, Changlong Sun, Qi Zhang 0001, Xuanjing Huang 0001, Xiaozhong Liu 0001 |
AAAI | 2 |
| 2020 | Behavior Based Dynamic Summarization on Product Aspects via Reinforcement Neighbour SelectionabstractDynamic summarization on product aspects, as a newly proposed topic, is an important task in E-commerce for tracking and understanding the nature of products. This can benefit both customers and sellers in different downstream tasks, such as explainable recommendations. However, most existing research works focus on analyzing product static reviews but miss dynamic sentiment changes. In this paper, we propose an innovative multi-task model to sample neighbour products whose information is simultaneously utilized to generate product summarization. In detail, a reinforcement learning approach selects neighbour products from a group of seed products by considering their pairwise similarities calculated from user behaviors. Meanwhile, a generative model helps to summarize product aspects via product descriptive phrases and selected neighbour products' sentimental phrases. To the best of our knowledge, this is the first work that studies dynamic product summarization leveraging user behaviors instead of self-reviews. It means that the proposed approach can naturally address the cold-start scenario where few recent product reviews are available. Extensive experiments are conducted with real-world reviews plus behavior data to validate the proposed method against several strong alternatives. Zheng Gao 0001, Lujun Zhao, Hongsong Li, Changlong Sun, Luo Si, Xiaozhong Liu 0001 |
ECAI | 2 |
| 2019 | Long Short-Term Memory with Dynamic Skip ConnectionsabstractIn recent years, long short-term memory (LSTM) has been successfully used to model sequential data of variable length. However, LSTM can still experience difficulty in capturing long-term dependencies. In this work, we tried to alleviate this problem by introducing a dynamic skip connection, which can learn to directly connect two dependent words. Since there is no dependency information in the training data, we propose a novel reinforcement learning-based method to model the dependency relationship and connect dependent words. The proposed model computes the recurrent transition functions based on the skip connections, which provides a dynamic skipping advantage over RNNs that always tackle entire sentences sequentially. Our experimental results on three natural language processing tasks demonstrate that the proposed method can achieve better performance than existing methods. In the number prediction experiment, the proposed model outperformed LSTM with respect to accuracy by nearly 20%. Tao Gui, Qi Zhang 0001, Lujun Zhao, Yaosong Lin, Minlong Peng, Jingjing Gong, Xuanjing Huang 0001 |
AAAI | 3 |
| 2019 | Cross-domain Aspect Category Transfer and Detection via Traceable Heterogeneous Graph Representation LearningabstractAspect category detection is an essential task for sentiment analysis and opinion mining. However, the cost of categorical data labeling, e.g., label the review aspect information for a large number of product domains, can be inevitable but unaffordable. In this study, we propose a novel problem, cross-domain aspect category transfer and detection, which faces three challenges: various feature spaces, different data distributions, and diverse output spaces. To address these problems, we propose an innovative solution, Traceable Heterogeneous Graph Representation Learning (THGRL). Unlike prior text-based aspect detection works, THGRL explores latent domain aspect category connections via massive user behavior information on a heterogeneous graph. Moreover, an innovative latent variable "Walker Tracer" is introduced to characterize the global semantic/aspect dependencies and capture the informative vertexes on the random walk paths. By using THGRL, we project different domains' feature spaces into a common one, while allowing data distributions and output spaces stay differently. Experiment results show that the proposed method outperforms a series of state-of-the-art baseline models. Zhuoren Jiang, Lujun Zhao, Changlong Sun, Yao Lu 0007, Xiaozhong Liu 0001 |
CIKM | 3 |
| 2019 | Using Customer Service Dialogues for Satisfaction Analysis with Context-Assisted Multiple Instance LearningabstractKaisong Song, Lidong Bing, Wei Gao, Jun Lin, Lujun Zhao, Jiancheng Wang, Changlong Sun, Xiaozhong Liu, Qiong Zhang. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Kaisong Song, Lidong Bing, Wei Gao 0001, Lujun Zhao, Changlong Sun, Xiaozhong Liu 0001, Qi Zhang 0001 |
EMNLP/IJCNLP (1) | 5 |
| 2019 | CNN-Based Chinese NER with Lexicon RethinkingabstractCharacter-level Chinese named entity recognition (NER) that applies long short-term memory (LSTM) to incorporate lexicons has achieved great success. However, this method fails to fully exploit GPU parallelism and candidate lexicons can conflict. In this work, we propose a faster alternative to Chinese NER: a convolutional neural network (CNN)-based method that incorporates lexicons using a rethinking mechanism. The proposed method can model all the characters and potential words that match the sentence in parallel. In addition, the rethinking mechanism can address the word conflict by feeding back the high-level features to refine the networks. Experimental results on four datasets show that the proposed method can achieve better performance than both word-level and character-level baseline methods. In addition, the proposed method performs up to 3.21 times faster than state-of-the-art methods, while realizing better performance. Tao Gui, Ruotian Ma, Qi Zhang 0001, Lujun Zhao, Yu-Gang Jiang 0001, Xuanjing Huang 0001 |
IJCAI | 4 |
| 2019 | Cold-Start Aware Deep Memory Network for Multi-Entity Aspect-Based Sentiment AnalysisabstractVarious types of target information have been considered in aspect-based sentiment analysis, such as entities and aspects. Existing research has realized the importance of targets and developed methods with the goal of precisely modeling their contexts via generating target-specific representations. However, all these methods ignore that these representations cannot be learned well due to the lack of sufficient human-annotated target-related reviews, which leads to the data sparsity challenge, a.k.a. cold-start problem here. In this paper, we focus on a more general multiple entity aspect-based sentiment analysis (ME-ABSA) task which aims at identifying the sentiment polarity of different aspects of multiple entities in their context. Faced with severe cold-start scenario, we develop a novel and extensible deep memory network framework with cold-start aware computational layers which use frequency-guided attention mechanism to accentuate on the most related targets, and then compose their representations into a complementary vector for enhancing the representations of cold-start entities and aspects. To verify the effectiveness of the framework, we instantiate it with a concrete context encoding method and then apply the model to the ME-ABSA task. Experimental results conducted on two public datasets demonstrate that the proposed approach outperforms state-of-the-art baselines on ME-ABSA task. Kaisong Song, Wei Gao 0001, Lujun Zhao, Changlong Sun, Xiaozhong Liu 0001 |
IJCAI | 3 |
| 2019 | Review Response Generation in E-Commerce Platforms with External Product Informationabstract''User reviews” are becoming an essential component of e-commerce. When buyers write a negative or doubting review, ideally, the sellers need to quickly give a response to minimize the potential impact. When the number of reviews is growing at a frightening speed, there is an urgent need to build a response writing assistant for customer service providers. In order to generate high-quality responses, the algorithm needs to consume and understand the information from both the original review and the target product. The classical sequence-to-sequence (Seq2Seq) methods can hardly satisfy this requirement. In this study, we propose a novel deep neural network model based on the Seq2Seq framework for the review response generation task in e-commerce platforms, which can incorporate product information by a gated multi-source attention mechanism and a copy mechanism. Moreover, we employ a reinforcement learning technique to reduce the exposure bias problem. To evaluate the proposed model, we constructed a large-scale dataset from a popular e-commerce website, which contains product information. Empirical studies on both automatic evaluation metrics and human annotations show that the proposed model can generate informative and diverse responses, significantly outperforming state-of-the-art text generation models. Lujun Zhao, Kaisong Song, Changlong Sun, Qi Zhang 0001, Xuanjing Huang 0001, Xiaozhong Liu 0001 |
WWW | 1 |
| 2019 | Sequence Labeling With Deep Gated Dual Path CNNabstractSequence labeling, such as part-of-speech (POS) tagging, named entity recognition (NER), text chunking, is a classic task in natural language processing. Most existing neural networks models for sequence labeling are based on recurrent neural networks. Recently, convolutional neural networks have been proposed to replace the recurrent components for sequence labeling. However, they are usually shallow compared to deep convolutional networks that achieve start-of-the-art performance in other fields. Due to the vanishing gradient problem, these models usually can not work well when simply increasing the number of layers. In this paper, we propose using deep CNN architecture in sequence labeling, which can capture a large context through stacked convolutions. To reduce the vanishing gradient problem, the proposed method incorporates gated linear units, residual connections, and dense connections. Experimental results on three sequence labeling tasks show that the proposed model can achieve competitive performance to the RNN-based state-of-the-art method while maintaining 2.41 × faster speed, even with up to 10 convolutional layers. Lujun Zhao, Xipeng Qiu, Qi Zhang 0001, Xuanjing Huang 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2018 | Neural Networks Incorporating Unlabeled and Partially-labeled Data for Cross-domain Chinese Word SegmentationabstractMost existing Chinese word segmentation (CWS) methods are usually supervised. Hence, large-scale annotated domain-specific datasets are needed for training. In this paper, we seek to address the problem of CWS for the resource-poor domains that lack annotated data. A novel neural network model is proposed to incorporate unlabeled and partially-labeled data. To make use of unlabeled data, we combine a bidirectional LSTM segmentation model with two character-level language models using a gate mechanism. These language models can capture co-occurrence information. To make use of partially-labeled data, we modify the original cross entropy loss function of RNN. Experimental results demonstrate that the method performs well on CWS tasks in a series of domains. Lujun Zhao, Qi Zhang 0001, Peng Wang 0095 |
IJCAI | 1 |