VLDB 2026 Research / reviewers in the wild / expert
Zhibo Man
dblp:283/7862
· DBLP profile ↗
7ranked-venue papers
5as first author
7since 2021 · last 2026
0009-0002-3336-5703ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 6 · 5 first-author · 6 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
2 papers |
Generative modeling · 21% Information extraction and text analysis · 21% Language models and text generation · 21% | |
| Databases, data mining, and information retrieval
1 paper |
Recommender systems · 100% |
Topics — the 9 heaviest of 10, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Language models and text generation
large language model evaluation |
0.9 | 1 | 2025 | DMDTEval: An Evaluation and Analysis of LLMs on Disambiguation in Multi-domain Translation · EMNLP 2025 |
Machine learning › Generative modeling › generative adversarial network › image-to-image translation
multi-domain image translation |
0.9 | 1 | 2025 | DMDTEval: An Evaluation and Analysis of LLMs on Disambiguation in Multi-domain Translation · EMNLP 2025 |
Natural language and speech › Information extraction and text analysis
word sense disambiguation |
0.9 | 1 | 2025 | DMDTEval: An Evaluation and Analysis of LLMs on Disambiguation in Multi-domain Translation · EMNLP 2025 |
Recommender systems
fairness and bias |
0.9 | 1 | 2025 | Dual Debiasing in LLM-based Recommendation · SIGIR 2025 |
Recommender systems › recommender system evaluation › off-policy evaluation
inverse propensity scoring |
0.9 | 1 | 2025 | Dual Debiasing in LLM-based Recommendation · SIGIR 2025 |
Recommender systems
large language model-based recommendation |
0.9 | 1 | 2025 | Dual Debiasing in LLM-based Recommendation · SIGIR 2025 |
Recommender systems › debiased recommendation
popularity bias mitigation |
0.9 | 1 | 2025 | Dual Debiasing in LLM-based Recommendation · SIGIR 2025 |
Machine learning › Representation and self-supervised learning › representation learning
domain-aware representation learning |
0.8 | 1 | 2024 | WDSRL: Multi-Domain Neural Machine Translation With Word-Level Domain-Sensitive Representation Learning · IEEE ACM Trans. Audio Speech Lang. Process. 2024 |
Natural language and speech › Machine translation
multi-domain neural machine translation |
0.8 | 1 | 2024 | WDSRL: Multi-Domain Neural Machine Translation With Word-Level Domain-Sensitive Representation Learning · IEEE ACM Trans. Audio Speech Lang. Process. 2024 |
Methods — techniques the papers use, named apart from their topics
prompt engineering · 0.9large language model · 0.9inverse propensity score weighting · 0.9disambiguation metrics · 0.9topic knowledge representation · 0.8domain discriminator · 0.8convolutional neural network · 0.8
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | DKF: Domain knowledge fusion in progressive incremental learning for multi-domain machine translation
Zhibo Man, Yuanmeng Chen, Yufeng Chen 0005, Jin An Xu |
Expert Syst. Appl. | 1 |
| 2025 | DMDTEval: An Evaluation and Analysis of LLMs on Disambiguation in Multi-domain TranslationabstractCurrently, Large Language Models (LLMs) have achieved remarkable results in machine translation.However, their performance in multi-domain translation (MDT) is less satisfactory, the meanings of words can vary across different domains, highlighting the significant ambiguity inherent in MDT.Therefore, evaluating the disambiguation ability of LLMs in MDT remains an open problem.To this end, we present an evaluation and analysis of LLMs on disambiguation in multi-domain translation (DMDTEval), our systematic evaluation framework consisting of three aspects: (1) we construct a translation test set with multi-domain ambiguous word annotation, (2) we curate a diverse set of disambiguation prompt strategies, and (3) we design precise disambiguation metrics, and study the efficacy of various prompt strategies on multiple state-of-the-art LLMs.We conduct comprehensive experiments across 4 language pairs and 13 domains, our extensive experiments reveal a number of crucial findings that we believe will pave the way and also facilitate further research in the critical area of improving the disambiguation of LLMs. Zhibo Man, Yuanmeng Chen, Jin An Xu |
EMNLP | 1 |
| 2025 | Dual Debiasing in LLM-based RecommendationabstractLarge language models (LLMs) have been widely applied in recommender systems, achieving remarkable success. However, LLM-based recommendation (LR) suffers from more severe popularity bias than conventional recommendation (CR), stemming from both training and inference stages. In this paper, we propose a novel debiasing method for LR, which performs debiasing in such two stages, so termed as Dual Debiasing in LR (D²LR). Concretely, in the training stage, we conduct token-wise inverse propensity score weighting to force the LLM to pay more attention on unpopular tokens. In the inference stage, we train a more biased CR model by increasing the weights of popular items, which adjusts the generation probability of corresponding tokens according to its scores for items, hoping to suppress the excessive generation of popular tokens. Experiments conducted on three real-world datasets validate the effectiveness of our D²LR in mitigating popularity bias in LR. Sijin Lu, Zhibo Man, Fangyuan Luo, Jun Wu 0007 |
SIGIR | 2 |
| 2024 | An Ensemble Strategy with Gradient Conflict for Multi-Domain Neural Machine TranslationabstractMulti-domain neural machine translation aims to construct a unified neural machine translation model to translate sentences across various domains. Nevertheless, previous studies have one limitation is the incapacity to acquire both domain-general and domain-specific representations concurrently. To this end, we propose an ensemble strategy with gradient conflict for multi-domain neural machine translation that automatically learns model parameters by identifying both domain-shared and domain-specific features. Specifically, our approach consists of (1) a parameter-sharing framework, where the parameters of all the layers are originally shared and equivalent to each domain, and (2) ensemble strategy, in which we design an Extra Ensemble strategy via a piecewise condition function to learn direction and distance-based gradient conflict. In addition, we give a detailed theoretical analysis of the gradient conflict to further validate the effectiveness of our approach. Experimental results on two multi-domain datasets show the superior performance of our proposed model compared to previous work. Zhibo Man, Yu Li 0025, Yuanmeng Chen, Yufeng Chen 0005, Jin An Xu |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 1 |
| 2024 | WDSRL: Multi-Domain Neural Machine Translation With Word-Level Domain-Sensitive Representation LearningabstractDue to the strong reliance on domain-specific knowledge, the joint learning manner of domain discrimination and translation has been widely considered in the Multi-Domain Neural Machine Translation (MDNMT) task. However, the word ambiguity problem still inevitably exists in MDNMT, especially when mixed multi-domain data is brought into the model training phase. Although word-level MDNMT can mitigate this problem to some extent, poor domain discrimination yet remains and severely hinders performance. Based on the above limitation, we observed that coarser granularity strings may provide more specific semantics, which is more conducive to domain discrimination. Thus, we propose a Word-level Domain-Sensitive Representation Learning (WDSRL) method. Specifically, we focus on two aspects of our approach: domain representation and domain discrimination. To extend the scope of domain representation, we adopt Convolution Neural Networks (CNN) to encode Local Domain Representation at different granularities, and then integrate Topic Knowledge Representation into each word. By doing so, context features related to the domain could be comprehensively enriched. Regarding domain discrimination, we design a Domain-Sensitive Discriminator, which could not only generate domain features for each word but also enhance domain representation learning. Experimental results demonstrate our substantial improvements over several representative baselines on multiple language pairs. Furthermore, the extensive analysis also indicates the superiority of our proposed domain-sensitive feature encoding strategy and domain-sensitive discriminator for word-level representation learning. Zhibo Man, Zengcheng Huang, Yu Li 0025, Yuanmeng Chen, Yufeng Chen 0005, Jin An Xu |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2023 | Exploring Domain-shared and Domain-specific Knowledge in Multi-Domain Neural Machine TranslationabstractCurrently, multi-domain neural machine translation (NMT) has become a significant research topic in domain adaptation machine translation, which trains a single model by mixing data from multiple domains. Multi-domain NMT aims to improve the performance of the low-resources domain through data augmentation. However, mixed domain data brings more translation ambiguity. Previous work focused on domain-general or domain-context knowledge learning, respectively. Therefore, there is a challenge for acquiring domain-general or domain-context knowledge simultaneously. To this end, we propose a unified framework for learning simultaneously domain-general and domain-specific knowledge, we are the first to apply parameter differentiation in multi-domain NMT. Specifically, we design the differentiation criterion and differentiation granularity to obtain domain-specific parameters. Experimental results on multi-domain UM-corpus English-to-Chinese and OPUS German-to-English datasets show that the average BLEU scores of the proposed method exceed the strong baseline by 1.22 and 1.87, respectively. In addition, we investigate the case study to illustrate the effectiveness of the proposed method in acquiring domain knowledge. Zhibo Man, Yuanmeng Chen, Yufeng Chen 0005, Jin An Xu |
MTSummit (1) | 1 |
| 2021 | A Neural Joint Model with BERT for Burmese Syllable Segmentation, Word Segmentation, and POS TaggingabstractThe smallest semantic unit of the Burmese language is called the syllable. In the present study, it is intended to propose the first neural joint learning model for Burmese syllable segmentation, word segmentation, and part-of-speech ( POS ) tagging with the BERT. The proposed model alleviates the error propagation problem of the syllable segmentation. More specifically, it extends the neural joint model for Vietnamese word segmentation, POS tagging, and dependency parsing [28] with the pre-training method of the Burmese character, syllable, and word embedding with BiLSTM-CRF-based neural layers. In order to evaluate the performance of the proposed model, experiments are carried out on Burmese benchmark datasets, and we fine-tune the model of multilingual BERT. Obtained results show that the proposed joint model can result in an excellent performance. Cunli Mao, Zhibo Man, Zhengtao Yu 0001, Shengxiang Gao, Zhenhan Wang, Hongbin Wang 0002 |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 2 |