VLDB 2026 Research / reviewers in the wild / expert
Wei Zhao 0033
dblp:181/2852-33
· DBLP profile ↗
26ranked-venue papers
11as first author
9since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 23 · 11 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 3 first-authorDatabases, data management, data science and information retrieval · 4 · 2 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | APODICTUS: Automatic Processing of DICTionary Update candidateS
Felix Blessing, Johannes S. Sax, Julian Kaufmann, Wei Zhao 0033, Nikolay Arefyev, Dominik Schlechtweg |
LREC | 4 |
| 2024 | Graph-based Clustering for Detecting Semantic Change Across Time and LanguagesabstractDespite the predominance of contextualized embeddings in NLP, approaches to detect semantic change relying on these embeddings and clustering methods underperform simpler counterparts based on static word embeddings.This stems from the poor quality of the clustering methods to produce sense clusters-which struggle to capture word senses, especially those with low frequency.This issue hinders the next step in examining how changes in word senses in one language influence another.To address this issue, we propose a graph-based clustering approach to capture nuanced changes in both high-and low-frequency word senses across time and languages, including the acquisition and loss of these senses over time.Our experimental results show that our approach substantially surpasses previous approaches in the SemEval2020 binary classification task across four languages.Moreover, we showcase the ability of our approach as a versatile visualization tool to detect semantic changes in both intra-language and inter-language setups.We make our code and data available 1 . Xianghe Ma, Michael Strube 0001, Wei Zhao 0033 |
EACL (1) | 3 |
| 2024 | Towards Explainable Evaluation Metrics for Machine TranslationabstractUnlike classical lexical overlap metrics such as BLEU, most current evaluation metrics for machine translation (for example, COMET or BERTScore) are based on black-box large language models. They often achieve strong correlations with human judgments, but recent research indicates that the lower-quality classical metrics remain dominant, one of the potential reasons being that their decision processes are more transparent. To foster more widespread acceptance of novel high-quality metrics, explainability thus becomes crucial. In this concept paper, we identify key properties as well as key goals of explainable machine translation metrics and provide a comprehensive synthesis of recent techniques, relating them to our established goals and properties. In this context, we also discuss the latest state-of-the-art approaches to explainable metrics based on generative models such as ChatGPT and GPT4. Finally, we contribute a vision of next-generation approaches, including natural language explanations. We hope that our work can help catalyze and guide future research on explainable evaluation metrics and, mediately, also contribute to better and more transparent machine translation systems. Christoph Leiter, Piyawat Lertvittayakumjorn, Marina Fomicheva, Wei Zhao 0033, Yang Gao 0021, Steffen Eger |
J. Mach. Learn. Res. | 4 |
| 2023 | DiscoScore: Evaluating Text Generation with BERT and Discourse CoherenceabstractRecently, there has been a growing interest in designing text generation systems from a discourse coherence perspective, e.g., modeling the interdependence between sentences.Still, recent BERT-based evaluation metrics are weak in recognizing coherence, and thus are not reliable in a way to spot the discourselevel improvements of those text generation systems.In this work, we introduce DiscoScore, a parametrized discourse metric, which uses BERT to model discourse coherence from different perspectives, driven by Centering theory.Our experiments encompass 16 non-discourse and discourse metrics, including DiscoScore and popular coherence models, evaluated on summarization and document-level machine translation (MT).We find that (i) the majority of BERT-based metrics correlate much worse with human rated coherence than early discourse metrics, invented a decade ago; (ii) the recent state-of-the-art BARTScore is weak when operated at system level-which is particularly problematic as systems are typically compared in this manner.DiscoScore, in contrast, achieves strong system-level correlation with human ratings, not only in coherence but also in factual consistency and other aspects, and surpasses BARTScore by over 10 correlation points on average.Further, aiming to understand DiscoScore, we provide justifications to the importance of discourse coherence for evaluation metrics, and explain the superiority of one variant over another.Our code is available at https://github.com/AIPHES/ DiscoScore. Wei Zhao 0033, Michael Strube 0001, Steffen Eger |
EACL | 1 |
| 2023 | Modeling Graphs Beyond Hyperbolic: Graph Neural Networks in Symmetric Positive Definite Matrices
Wei Zhao 0033, Federico López 0002, J. Maxwell Riestenberg, Michael Strube 0001, Diaaeldin Taha, Steve Trettel |
ECML/PKDD (3) | 1 |
| 2022 | Constrained Density Matching and Modeling for Cross-lingual Alignment of Contextualized Representations
Wei Zhao 0033, Steffen Eger |
ACML | 1 |
| 2021 | Better than Average: Paired Evaluation of NLP systemsabstractMaxime Peyrard, Wei Zhao, Steffen Eger, Robert West. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Maxime Peyrard, Wei Zhao 0033, Steffen Eger, Robert West 0001 |
ACL/IJCNLP (1) | 2 |
| 2021 | Global Explainability of BERT-Based Evaluation Metrics by Disentangling along Linguistic FactorsabstractEvaluation metrics are a key ingredient for progress of text generation systems.In recent years, several BERT-based evaluation metrics have been proposed (including BERTScore, MoverScore, BLEURT, etc.) which correlate much better with human assessment of text generation quality than BLEU or ROUGE, invented two decades ago.However, little is known what these metrics, which are based on black-box language model representations, actually capture (it is typically assumed they model semantic similarity).In this work, we use a simple regression based global explainability technique to disentangle metric scores along linguistic factors, including semantics, syntax, morphology, and lexical overlap.We show that the different metrics capture all aspects to some degree, but that they are all substantially sensitive to lexical overlap, just like BLEU and ROUGE.This exposes limitations of these novelly proposed metrics, which we also highlight in an adversarial test scenario. Marvin Kaster, Wei Zhao 0033, Steffen Eger |
EMNLP (1) | 2 |
| 2021 | Graph routing between capsules
Yang Li 0055, Wei Zhao 0033, Erik Cambria, Suhang Wang, Steffen Eger |
Neural Networks | 2 |
| 2020 | SUPERT: Towards New Frontiers in Unsupervised Evaluation Metrics for Multi-Document SummarizationabstractWe study unsupervised multi-document summarization evaluation metrics, which require neither human-written reference summaries nor human annotations (e.g.preferences, ratings, etc.).We propose SUPERT, which rates the quality of a summary by measuring its semantic similarity with a pseudo reference summary, i.e. selected salient sentences from the source documents, using contextualized embeddings and soft token alignment techniques.Compared to the state-of-theart unsupervised evaluation metrics, SUPERT correlates better with human ratings by 18-39%.Furthermore, we use SUPERT as rewards to guide a neural-based reinforcement learning summarizer, yielding favorable performance compared to the state-of-the-art unsupervised summarizers. Yang Gao 0021, Wei Zhao 0033, Steffen Eger |
ACL | 2 |
| 2020 | On the Limitations of Cross-lingual Encoders as Exposed by Reference-Free Machine Translation EvaluationabstractEvaluation of cross-lingual encoders is usually performed either via zero-shot cross-lingual transfer in supervised downstream tasks or via unsupervised cross-lingual textual similarity.In this paper, we concern ourselves with reference-free machine translation (MT) evaluation where we directly compare source texts to (sometimes low-quality) system translations, which represents a natural adversarial setup for multilingual encoders.Referencefree evaluation holds the promise of web-scale comparison of MT systems.We systematically investigate a range of metrics based on state-of-the-art cross-lingual semantic representations obtained with pretrained M-BERT and LASER.We find that they perform poorly as semantic encoders for reference-free MT evaluation and identify their two key limitations, namely, (a) a semantic mismatch between representations of mutual translations and, more prominently, (b) the inability to punish "translationese", i.e., low-quality literal translations.We propose two partial remedies: (1) post-hoc re-alignment of the vector spaces and (2) coupling of semantic-similarity based metrics with target-side language modeling.In segment-level MT evaluation, our best metric surpasses reference-based BLEU by 5.7 correlation points.We make our MT evaluation code available.1 Wei Zhao 0033, Goran Glavas, Maxime Peyrard, Yang Gao 0021, Robert West 0001, Steffen Eger |
ACL | 1 |
| 2020 | Leveraging Long and Short-Term Information in Content-Aware Movie Recommendation via Adversarial TrainingabstractMovie recommendation systems provide users with ranked lists of movies based on individual's preferences and constraints. Two types of models are commonly used to generate ranking results: 1) long-term models and 2) session-based models. The long-term-based models represent the interactions between users and movies that are supposed to change slowly across time, while the session-based models encode the information of users' interests and changing dynamics of movies' attributes in short terms. In this paper, we propose the LSIC model, leveraging long and short-term information for content-aware movie recommendation using adversarial training. In the adversarial process, we train a generator as an agent of reinforcement learning which recommends the next movie to a user sequentially. We also train a discriminator which attempts to distinguish the generated list of movies from the real records. The poster information of movies is integrated to further improve the performance of movie recommendation, which is specifically essential when few ratings are available. The experiments demonstrate that the proposed model has robust superiority over competitors and achieves the state-of-the-art results. Wei Zhao 0033, Benyou Wang, Min Yang 0007, Jianbo Ye, Zhou Zhao 0001, Xiaojun Chen 0006, Ying Shen 0001 |
IEEE Trans. Cybern. | 1 |
| 2019 | Towards Scalable and Reliable Capsule Networks for Challenging NLP ApplicationsabstractObstacles hindering the development of capsule networks for challenging NLP applications include poor scalability to large output spaces and less reliable routing processes.In this paper, we introduce (i) an agreement score to evaluate the performance of routing processes at instance level; (ii) an adaptive optimizer to enhance the reliability of routing; (iii) capsule compression and partial routing to improve the scalability of capsule networks.We validate our approach on two NLP tasks, namely: multi-label text classification and question answering.Experimental results show that our approach considerably improves over strong competitors on both tasks.In addition, we gain the best results in low-resource settings with few training instances.1 Wei Zhao 0033, Haiyun Peng, Steffen Eger, Erik Cambria, Min Yang 0007 |
ACL (1) | 1 |
| 2019 | A Unified Generation-Retrieval Framework for Image CaptioningabstractRecent image captioning approaches are typically trained on generation-based or retrieval-based approaches. Both methods have their advantages but limited by the disadvantages. In this paper, we propose a Unified Generation-Retrieval framework for Image Captioning (UGRIC) by using adversarial learning. Different from previous methods, the proposed UGRIC model leverages the informative contents of N-best response candidates provided by the retrieval-based model to enhance the generation-based method. In addition, to further improve the informativeness of the generated caption, we employ copying mechanism to choose words from the retrieved candidate captions and put them into proper positions of the output sequence. Experiments on MSCOCO dataset demonstrate the effectiveness of the UGRIC model through various evaluation metrics.\footnoteCode and data are available at: \urlhttp://tinyurl.com/y6z2x6ho. Chunpu Xu, Wei Zhao 0033, Min Yang 0007, Xiang Ao 0001, Wangrong Cheng, Jinwen Tian |
CIKM | 2 |
| 2019 | MoverScore: Text Generation Evaluating with Contextualized Embeddings and Earth Mover DistanceabstractWei Zhao, Maxime Peyrard, Fei Liu, Yang Gao, Christian M. Meyer, Steffen Eger. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Wei Zhao 0033, Maxime Peyrard, Fei Liu 0004, Yang Gao 0021, Christian M. Meyer, Steffen Eger |
EMNLP/IJCNLP (1) | 1 |
| 2019 | Investigating the transferring capability of capsule networks for text classification
Min Yang 0007, Wei Zhao 0033, Lei Chen 0072, Qiang Qu 0001, Zhou Zhao 0001, Ying Shen 0001 |
Neural Networks | 2 |
| 2019 | Multitask Learning for Cross-Domain Image CaptioningabstractRecent artificial intelligence research has witnessed great interest in automatically generating text descriptions of images, which are known as the image captioning task. Remarkable success has been achieved on domains where a large number of paired data in multimedia are available. Nevertheless, annotating sufficient data is labor-intensive and time-consuming, establishing significant barriers for adapting the image captioning systems to new domains. In this study, we introduc a novel Multitask Learning Algorithm for cross-Domain Image Captioning (MLADIC). MLADIC is a multitask system that simultaneously optimizes two coupled objectives via a dual learning mechanism: image captioning and text-to-image synthesis, with the hope that by leveraging the correlation of the two dual tasks, we are able to enhance the image captioning performance in the target domain. Concretely, the image captioning task is trained with an encoder-decoder model (i.e., CNN-LSTM) to generate textual descriptions of the input images. The image synthesis task employs the conditional generative adversarial network (C-GAN) to synthesize plausible images based on text descriptions. In C-GAN, a generative model $G$ synthesizes plausible images given text descriptions, and a discriminative model $D$ tries to distinguish the images in training data from the generated images by $G$. The adversarial process can eventually guide $G$ to generate plausible and high-quality images. To bridge the gap between different domains, a two-step strategy is adopted in order to transfer knowledge from the source domains to the target domains. First, we pre-train the model to learn the alignment between the neural representations of images and that of text data with the sufficient labeled source domain data. Second, we fine-tune the learned model by leveraging the limited image-text pairs and unpaired data in the target domain. We conduct extensive experiments to evaluate the performance of MLADIC by using the MSCOCO as the source domain data, and using Flickr30k and Oxford-102 as the target domain data. The results demonstrate that MLADIC achieves substantially better performance than the strong competitors for the cross-domain image captioning task. Min Yang 0007, Wei Zhao 0033, Yabing Feng, Zhou Zhao 0001, Xiaojun Chen 0006, Kai Lei |
IEEE Trans. Multim. | 2 |
| 2018 | A Semi-Supervised Network Embedding Model for Protein Complexes DetectionabstractProtein complex is a group of associated polypeptide chains which plays essential roles in biological process. Given a graph representing protein-protein interactions (PPI) network, it is critical but non-trivial to detect protein complexes.In this paper, we propose a semi-supervised network embedding model by adopting graph convolutional networks to effectively detect densely connected subgraphs. We conduct extensive experiment on two popular PPI networks with various data sizes and densities. The experimental results show our approach achieves state-of-the-art performance. Wei Zhao 0033, Jia Zhu 0003, Min Yang 0007, Danyang Xiao, Gabriel Pui Cheong Fung, Xiaojun Chen 0006 |
AAAI | 1 |
| 2018 | Aspect and Sentiment Aware Abstractive Review SummarizationabstractReview text has been widely studied in traditional tasks such as sentiment analysis and aspect extraction. However, to date, no work is towards the abstractive review summarization that is essential for business organizations and individual consumers to make informed decisions. This work takes the lead to study the aspect/sentiment-aware abstractive review summarization by exploring multi-factor attentions. Specifically, we propose an interactive attention mechanism to interactively learns the representations of context words, sentiment words and aspect words within the reviews, acted as an encoder. The learned sentiment and aspect representations are incorporated into the decoder to generate aspect/sentiment-aware review summaries via an attention fusion network. In addition, the abstractive summarizer is jointly trained with the text categorization task, which helps learn a category-specific text encoder, locating salient aspect information and exploring the variations of style and wording of content with respect to different text categories. The experimental results on a real-life dataset demonstrate that our model achieves impressive results compared to other strong competitors. Min Yang 0007, Qiang Qu 0001, Ying Shen 0001, Qiao Liu 0003, Wei Zhao 0033, Jia Zhu 0003 |
COLING | 5 |
| 2018 | Investigating Capsule Networks with Dynamic Routing for Text ClassificationabstractIn this study, we explore capsule networks with dynamic routing for text classification.We propose three strategies to stabilize the dynamic routing process to alleviate the disturbance of some noise capsules which may contain "background" information or have not been successfully trained.A series of experiments are conducted with capsule networks on six text classification benchmarks.Capsule networks achieve competitive results over the compared baseline methods on 4 out of 6 datasets, which shows the effectiveness of capsule networks for text classification.We additionally show that capsule networks exhibit significant improvement when transfer single-label to multi-label text classification over the competitors.To the best of our knowledge, this is the first work that capsule networks have been empirically investigated for text modeling 1 . Min Yang 0007, Wei Zhao 0033, Jianbo Ye, Zeyang Lei, Zhou Zhao 0001, Soufei Zhang |
EMNLP | 2 |
| 2018 | PLASTIC: Prioritize Long and Short-term Information in Top-n Recommendation using Adversarial TrainingabstractRecommender systems provide users with ranked lists of items based on individual's preferences and constraints. Two types of models are commonly used to generate ranking results: long-term models and session-based models. While long-term models represent the interactions between users and items that are supposed to change slowly across time, session-based models encode the information of users' interests and changing dynamics of items' attributes in short terms. In this paper, we propose a PLASTIC model, Prioritizing Long And Short-Term Information in top-n reCommendation using adversarial training. In the adversarial process, we train a generator as an agent of reinforcement learning which recommends the next item to a user sequentially. We also train a discriminator which attempts to distinguish the generated list of items from the real list recorded. Extensive experiments show that our model exhibits significantly better performances on two widely used real-world datasets. Wei Zhao 0033, Benyou Wang, Jianbo Ye, Yongqiang Gao, Min Yang 0007, Xiaojun Chen 0006 |
IJCAI | 1 |
| 2018 | A Multi-task Learning Approach for Image CaptioningabstractIn this paper, we propose a Multi-task Learning Approach for Image Captioning (MLAIC ), motivated by the fact that humans have no difficulty performing such task because they possess capabilities of multiple domains. Specifically, MLAIC consists of three key components: (i) A multi-object classification model that learns rich category-aware image representations using a CNN image encoder; (ii) A syntax generation model that learns better syntax-aware LSTM based decoder; (iii) An image captioning model that generates image descriptions in text, sharing its CNN encoder and LSTM decoder with the object classification task and the syntax generation task, respectively. In particular, the image captioning model can benefit from the additional object categorization and syntax knowledge. To verify the effectiveness of our approach, we conduct extensive experiments on MS-COCO dataset. The experimental results demonstrate that our model achieves impressive results compared to other strong competitors. Wei Zhao 0033, Benyou Wang, Jianbo Ye, Min Yang 0007, Zhou Zhao 0001, Ruotian Luo, Yu Qiao 0001 |
IJCAI | 1 |
| 2018 | Task-oriented keyphrase extraction from social media
Min Yang 0007, Yuzhi Liang, Wei Zhao 0033, Jia Zhu 0003, Qiang Qu 0001 |
Multim. Tools Appl. | 3 |
| 2017 | Dual Learning for Cross-domain Image CaptioningabstractRecent AI research has witnessed increasing interests in automatically generating image descriptions in text, which is coined as theimage captioning problem. Significant progresses have been made in domains where plenty of labeled training data (i.e. image-text pairs) are readily available or collected. However, obtaining rich annotated data is a time-consuming and expensive process, creating a substantial barrier for applying image captioning methods to a new domain. In this paper, we propose a cross-domain image captioning approach that uses a novel dual learning mechanism to overcome this barrier. First, we model the alignment between the neural representations of images and that of natural languages in the source domain where one can access sufficient labeled data. Second, we adjust the pre-trained model based on examining limited data (or unpaired data) in the target domain. In particular, we introduce a dual learning mechanism with a policy gradient method that generates highly rewarded captions. The mechanism simultaneously optimizes two coupled objectives: generating image descriptions in text and generating plausible images from text descriptions, with the hope that by explicitly exploiting their coupled relation, one can safeguard the performance of image captioning in the target domain. To verify the effectiveness of our model, we use MSCOCO dataset as the source domain and two other datasets (Oxford-102 and Flickr30k) as the target domains. The experimental results show that our model consistently outperforms previous methods for cross-domain image captioning. Wei Zhao 0033, Min Yang 0007, Jianbo Ye, Zhou Zhao 0001, Yabing Feng, Yu Qiao 0001 |
CIKM | 1 |
| 2017 | Identifying and Tracking Sentiments and Topics from Social Media Texts during Natural DisastersabstractWe study the problem of identifying the topics and sentiments and tracking their shifts from social media texts in different geographical regions during emergencies and disasters.We propose a location-based dynamic sentiment-topic model (LDST) which can jointly model topic, sentiment, time and Geolocation information.The experimental results demonstrate that LDST performs very well at discovering topics and sentiments from social media and tracking their shifts in different geographical regions during emergencies and disasters 1 . Min Yang 0007, Jincheng Mei, Heng Ji 0001, Wei Zhao 0033, Zhou Zhao 0001, Xiaojun Chen 0006 |
EMNLP | 4 |
| 2017 | Personalized Response Generation via Domain adaptationabstractIn this paper, we propose a novel personalized response generation model via domain adaptation (PRG-DM). First, we learn the human responding style from large general data (without user-specific information). Second, we fine tune the model on a small size of personalized data to generate personalized responses with a dual learning mechanism. Moreover, we propose three new rewards to characterize good conversations that are personalized, informative and grammatical. We employ the policy gradient method to generate highly rewarded responses. Experimental results show that our model can generate better personalized responses for different users. Min Yang 0007, Zhou Zhao 0001, Wei Zhao 0033, Xiaojun Chen 0006, Jia Zhu 0003, Lianqiang Zhou, Zigang Cao |
SIGIR | 3 |