EDBT 2026 Demo / reviewers in the wild / expert
Xiaobo Liang
dblp:224/6053
· DBLP profile ↗
17ranked-venue papers
8as first author
16since 2021 · last 2026
0009-0001-1550-2877ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 16 · 8 first-author · 15 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Escaping the Echo Trap: On Credit Assignment Failure in Multi-turn LLM Self-ReflectionabstractLinxuan Du, Guangquan Xue, Xiaobo Liang, Qipeng Huang, Yuyang Ding, Xinyu Shi, Zhang Yijun, Ji Qi, Wenpeng Zhu, Juntao Li, Min Zhang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Linxuan Du, Guangquan Xue, Xiaobo Liang, Qipeng Huang, Yuyang Ding, Xinyu Shi 0005, Zhang Yijun, Wenpeng Zhu, Juntao Li 0005, Min Zhang 0005 |
ACL (1) | 3 |
| 2026 | DUAL RM: Beyond Rule-based Preference Reward Modeling via Meta-RewardabstractXiaobo Liang, Wanfu Wang, Qipeng Huang, Yuyang Ding, Zecheng Tang, Yixin Ji, Qianben Chen, Zhe Zhao, Kehai Chen, Juntao Li, Min Zhang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Xiaobo Liang, Wanfu Wang, Qipeng Huang, Yuyang Ding, Zecheng Tang, Yixin Ji, Qianben Chen, Kehai Chen, Juntao Li 0005, Min Zhang 0005 |
ACL (1) | 1 |
| 2026 | Adaptive learning guided dual-function scheduling for delay-bounded quality of service in energy-constrained fifth-generation network slicesabstractThe rapid growth of latency-sensitive applications such as live video streaming, telemedicine, and emergency communications in fifth-generation mobile networks has intensified the need for time-bounded Quality of Service guarantees while maintaining energy efficiency. Within the context of network slicing, heterogeneous traffic patterns, rapid variations in user density, and stringent energy constraints pose major challenges to conventional scheduling approaches, which typically rely on fixed priority rules or non-adaptive parameters and therefore suffer from degraded performance under high network load. This paper proposes an Adaptive Learning Guided Dual-Function Scheduling framework to deliver delay-bounded Quality of Service in energy-constrained fifth-generation network slices. The proposed scheduler integrates a bi-functional prioritization mechanism, combining exponential and logarithmic components, with a lightweight Artificial Intelligence–based online learning agent. This learning agent adaptively tunes scheduling parameters based on real-time observations of energy consumption trends, traffic intensity, queue dynamics, and delay stability, enabling context-aware and energy-efficient resource allocation. From a computational perspective, the proposed approach is designed to maintain low processing overhead, making it suitable for deployment in radio access network nodes with limited computational capabilities. Extensive system-level simulations conducted on a fifth-generation network slicing platform demonstrate that the proposed method significantly reduces average latency, lowers packet loss, and decreases overall energy consumption, while preserving a high level of fairness among network slices. Compared with existing benchmark scheduling schemes, the proposed Artificial Intelligence–enabled framework consistently achieves superior performance under both moderate and heavy traffic conditions. Peigang Wei, Xiaofeng Nong, Xiaobo Liang |
Eng. Appl. Artif. Intell. | 4 |
| 2025 | Unleashing LLM Reasoning Capability via Scalable Question Synthesis from ScratchabstractImproving the mathematical reasoning capabilities of Large Language Models (LLMs) is critical for advancing artificial intelligence. However, access to extensive, diverse, and high-quality reasoning datasets remains a significant challenge, particularly for the open-source community. In this paper, we propose ScaleQuest, a novel, scalable, and cost-effective data synthesis method that enables the generation of large-scale mathematical reasoning datasets using lightweight 7B-scale models. ScaleQuest introduces a two-stage question-tuning process comprising Question Fine-Tuning (QFT) and Question Preference Optimization (QPO) to unlock the question generation capabilities of problem-solving models. By generating diverse questions from scratch – without relying on powerful proprietary models or seed data – we produce a dataset of 1 million problem-solution pairs. Our experiments demonstrate that models trained on our data outperform existing open-source datasets in both in-domain and out-of-domain evaluations. Furthermore, our approach shows continued performance improvement as the volume of training data increases, highlighting its potential for ongoing data scaling. The extensive improvements observed in code reasoning tasks demonstrate the generalization capabilities of our proposed method. Our work provides the open-source community with a practical solution to enhance the mathematical reasoning abilities of LLMs. Yuyang Ding, Xinyu Shi 0005, Xiaobo Liang, Juntao Li 0005, Zhaopeng Tu, Qiaoming Zhu, Min Zhang 0005 |
ACL (1) | 3 |
| 2025 | Generative Reward Modeling via Synthetic Criteria Preference LearningabstractGenerative Reward Models (GenRMs) leverage synthesized Chains of Thought (CoT) to reduce the need for massive labeled data, but this approach introduces risks of overoptimization due to the inability to guarantee the correctness of the CoTs.Identifying and optimizing unexpected behaviors within these synthesized CoT remains a challenge, as it heavily depends on precise annotations of intermediate behavior, similar to process supervision.In this work, we introduce a criteria-based preference tree for reward modeling, where each path in the tree represents a reasoning trajectory based on synthesized criteria.Crucially, each reasoning trajectory can be independently optimized through RL algorithm.These fine-grained process reward signals are derived from the inferencetime computations and predefined rules, eliminating the need for human supervision.In experiments, SyncPL 1 showed significant improvements over baselines on multiple human preference benchmarks.We further demonstrate that synthesized data can be learned using a long CoT format, analogous to an o1-like model, further enhancing performance while keeping stability and efficiency during training. Xiaobo Liang, Haoke Zhang, Juntao Li 0005, Kehai Chen, Qiaoming Zhu, Min Zhang 0005 |
ACL (1) | 1 |
| 2025 | SCAN: Self-Denoising Monte Carlo Annotation for Robust Process Reward LearningabstractProcess reward models (PRMs) offer fine-grained, step-level evaluations that facilitate deeper reasoning processes in large language models (LLMs), proving effective in complex tasks like mathematical reasoning.
However, developing PRMs is challenging due to the high cost and limited scalability of human-annotated data.
Synthetic data from Monte Carlo (MC) estimation is a promising alternative but suffers from a high noise ratio, which can cause overfitting and hinder large-scale training.
In this work, we conduct a preliminary study on the noise distribution in synthetic data from MC estimation, identifying that annotation models tend to both underestimate and overestimate step correctness due to limitations in their annotation capabilities.
Building on these insights, we propose {\bf S}elf-Denoising Monte {\bf C}arlo {\bf An}notation (\textsc{Scan}), an efficient data synthesis and noise-tolerant learning framework.
Our key findings indicate that:
(1) Even lightweight models (e.g., 1.5B parameters) can produce high-quality annotations through self-denoising strategy, enabling PRMs to achieve superior performance with only 6\% the inference cost required by vanilla MC estimation.
(2) With our robust learning strategy, PRMs can effectively learn from this weak supervision, achieving a 39.2 F1 score improvement (from 19.9 to 59.1) in ProcessBench.
Despite using only a compact synthetic dataset, our models surpass strong baselines, including those trained on large-scale human-annotated datasets such as PRM800K.
Furthermore, performance continues to improve as we scale up the synthetic data, highlighting the potential of \textsc{Scan} for scalable, cost-efficient, and robust PRM training. Yuyang Ding, Xinyu Shi 0005, Juntao Li 0005, Xiaobo Liang, Zhaopeng Tu, Min Zhang 0005 |
NeurIPS | 4 |
| 2024 | Randomness Regularization With Simple Consistency Training for Neural NetworksabstractRandomness is widely introduced in neural network training to simplify model optimization or avoid the over-fitting problem. Among them, dropout and its variations in different aspects (e.g., data, model structure) are prevalent in regularizing the training of deep neural networks. Though effective and performing well, the randomness introduced by these dropout-based methods causes nonnegligible inconsistency between training and inference. In this paper, we introduce a simple consistency training strategy to regularize such randomness, namely R-Drop, which forces two output distributions sampled by each type of randomness to be consistent. Specifically, R-Drop minimizes the bidirectional KL-divergence between two output distributions produced by dropout-based randomness for each training sample. Theoretical analysis reveals that R-Drop can reduce the above inconsistency by reducing the inconsistency among the sampled sub structures and bridging the gap between the loss calculated by the full model and sub structures. Experiments on 7 widely-used deep learning tasks ( 23 datasets in total) demonstrate that R-Drop is universally effective for different types of neural networks (i.e., feed-forward, recurrent, and graph neural networks) and different learning paradigms (supervised, parameter-efficient, and semi-supervised). In particular, it achieves state-of-the-art performances with the vanilla Transformer model on WMT14 English → German translation ( 30.91 BLEU) and WMT14 English → French translation ( 43.95 BLEU), even surpassing models trained with extra large-scale data and expert-designed advanced variants of Transformer models. Juntao Li 0005, Xiaobo Liang, Lijun Wu 0003, Yue Wang 0039, Tao Qin 0001, Min Zhang 0005, Tie-Yan Liu |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2024 | Enhancing Low-Resource NLP by Consistency Training With Data and Model PerturbationsabstractNatural language processing (NLP) has recently shown significant progress in rich-resource scenarios. However, it is much less effective for low-resource scenarios due to the model easily overfitting to limited training data and generalizing poorly on testing data. In recent years, consistency training has been widely adopted and shown great promise in deep learning, but still remains unexplored in low-resource settings. In this work, we propose DM-CT, a framework that incorporates both data-level and model-level consistency training as well as advanced data augmentation techniques for low-resource scenarios. Concretely, the input data is first augmented, and the output distributions of different sub-models generated by model variance are forced to be consistent (model-level consistency). Meanwhile, the predictions of the original input and the augmented one are also constrained to be consistent (data-level consistency). Experiments on different low-resource NLP tasks, including neural machine translation (4 IWSLT14 translation tasks, multilingual translation task, and WMT16 Romanian$\to$English translation), natural language understanding tasks (GLUE benchmark), and named entity recognition (Conll2003 and WikiGold), well demonstrate the superiority of DM-CT by obtaining significant and consistent performance improvements. Xiaobo Liang, Runze Mao, Lijun Wu 0003, Juntao Li 0005, Min Zhang 0005, Qing Li 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2023 | Dynamic and Efficient Inference for Text Generation via BERT FamilyabstractDespite the excellent performance of Pretrained Language Models on many text generation tasks, they suffer from inefficient inference on computation and memory due to their largescale parameters and the universal autoregressive decoding paradigm.In this work, we propose a novel fine-tuning method DEER, which can make a single pre-trained model support Dynamic and Efficient infERence and achieve an adaptive trade-off between model performance and latency.In particular, our critical insight is to jointly utilize the non-autoregressive (NAR) generation and dynamic parameter pruning techniques, which can flexibly control the decoding iteration steps and model sizes according to memory and latency limitations.Besides, we also explore the effectiveness of the pre-trained MLMs (i.e., the BERT family) for text generation tasks since their bidirectional attention nature is more suitable for the NAR training objective.Extensive experiments on both monolingual and multilingual pre-trained MLMs demonstrate the effectiveness of our proposed DEER method by consistently achieving (1) higher BLEU scores than the strong autoregressive Transformer model on three neural machine translation tasks with 3 → 12 times speedup, (2) competitive performance (but with much faster inference speed) compared with the BART model on four GLGE benchmark tasks.Our code will be publicly available at GitHub 1 . Xiaobo Liang, Juntao Li 0005, Lijun Wu 0003, Ziqiang Cao, Min Zhang 0005 |
ACL (1) | 1 |
| 2023 | Open-ended Long Text Generation via Masked Language ModelingabstractPre-trained autoregressive (AR) language models such as BART and GPTs have dominated Open-ended Long Text Generation (Open-LTG).However, the AR nature will decrease the inference efficiency along with the increase of generation length, which hinder their application in Open-LTG.To improve inference efficiency, we alternatively explore the potential of the pre-trained masked language models (MLMs) along with a representative iterative non-autoregressive (NAR) decoding strategy for Open-LTG.Our preliminary study shows that pre-trained MLMs can merely generate short text and will collapse for long text modeling.To enhance the long text generation capability of MLMs, we introduce two simple yet effective strategies for the iterative NAR model: dynamic sliding window attention (DSWA) and linear temperature decay (LTD).It can alleviate long-distance collapse problems and achieve longer text generation with a flexible trade-off between performance and inference speedup.Experiments on the storytelling and multi-paragraph opinionated article writing tasks show that pre-trained MLMs can achieve more than 3 × → 13 × speedup with better performance than strong AR models.Our code is available at GitHub * . Xiaobo Liang, Zecheng Tang, Juntao Li 0005, Min Zhang 0005 |
ACL (1) | 1 |
| 2023 | CT4Rec: Simple yet Effective Consistency Training for Sequential RecommendationabstractSequential recommendation methods are increasingly important in cutting-edge recommender systems. Through leveraging historical records, the systems can capture user interests and perform recommendations accordingly. State-of-the-art sequential recommendation models proposed very recently combine contrastive learning techniques for obtaining high-quality user representations. Though effective and performing well, the models based on contrastive learning require careful selection of data augmentation methods and pretext tasks, efficient negative sampling strategies, and massive hyper-parameters validation. In this paper, we propose an ultra-simple alternative for obtaining better user representations and improving sequential recommendation performance. Specifically, we present a simple yet effective Consistency T braining method for sequential Recommendation (CT4Rec) in which only two extra training objectives are utilized without any structural modifications and data augmentation. Experiments on three benchmark datasets and one large newly crawled industrial corpus demonstrate that our proposed method outperforms SOTA models by a large margin and with much less training time than these based on contrastive learning. Online evaluation on real-world content recommendation system also achieves 2.717% improvement on the click-through rate and 3.679% increase on the average click number per capita. Further exploration reveals that such a simple method has great potential for CTR prediction. Our code is available at https://github.com/ct4rec/CT4Rec.git. Xiaoyang Liu 0012, Rongqin Zheng, Xiaobo Liang, Juntao Li 0005, Lijun Wu 0003, Min Zhang 0005, Leyu Lin |
KDD | 5 |
| 2023 | Are the BERT family zero-shot learners? A study on their potential and limitations
Yue Wang 0039, Lijun Wu 0003, Juntao Li 0005, Xiaobo Liang, Min Zhang 0005 |
Artif. Intell. | 4 |
| 2023 | R2-DDI: relation-aware feature refinement for drug-drug interaction predictionabstractPrecisely predicting the drug-drug interaction (DDI) is an important application and host research topic in drug discovery, especially for avoiding the adverse effect when using drug combination treatment for patients. Nowadays, machine learning and deep learning methods have achieved great success in DDI prediction. However, we notice that most of the works ignore the importance of the relation type when building the DDI prediction models. In this work, we propose a novel R$^2$-DDI framework, which introduces a relation-aware feature refinement module for drug representation learning. The relation feature is integrated into drug representation and refined in the framework. With the refinement features, we also incorporate the consistency training method to regularize the multi-branch predictions for better generalization. Through extensive experiments and studies, we demonstrate our R$^2$-DDI approach can significantly improve the DDI prediction performance over multiple real-world datasets and settings, and our method shows better generalization ability with the help of the feature refinement design. Jiacheng Lin, Lijun Wu 0003, Jinhua Zhu 0001, Xiaobo Liang, Yingce Xia, Shufang Xie 0003, Tao Qin 0001, Tie-Yan Liu |
Briefings Bioinform. | 4 |
| 2022 | JANUS: Joint Autoregressive and Non-autoregressive Training with Auxiliary Loss for Sequence GenerationabstractTransformer-based autoregressive and nonautoregressive models have played an essential role in sequence generation tasks.The autoregressive model can obtain excellent performance, while the non-autoregressive model brings fast decoding speed for inference.In this paper, we propose JANUS, a Joint Autoregressive and Non-autoregressive training method using aUxiliary losS to enhance the model performance in both AR and NAR manner simultaneously and effectively alleviate the problem of distribution discrepancy.Further, we pre-train BART with JANUS on a large corpus with minimal cost (16 GPU days) and make the BART-JANUS capable of nonautoregressive generation, demonstrating that our approach can transfer the AR knowledge to NAR.Empirically, we show our approach and BART-JANUS can achieve significant improvement on multiple generation tasks, including machine translation and GLGE benchmarks.Our code is available at Github 1 . Xiaobo Liang, Lijun Wu 0003, Juntao Li 0005, Min Zhang 0005 |
EMNLP | 1 |
| 2022 | Multi-Teacher Distillation With Single Model for Neural Machine TranslationabstractKnowledge distillation (KD) is an effective strategy for neural machine translation (NMT) to improve the performance of a student model. Usually, the teacher can guide the student to be better by distilling the soft label or data knowledge from the teacher itself. However, the data diversity and teacher knowledge are limited with only one teacher model. Though a natural solution is to adopt multiple randomized teacher models, one big shortcoming is that the model parameters and training costs are largely increased with the number of teacher models. In this work, we explore to mimic multiple teacher distillation from the sub-network space and permuted variants of one single teacher model. Specifically, we train a teacher by multiple sub-network extraction paradigms: sub-layer reordering, layer-drop, and dropout variants. In doing so, one teacher model can provide multiple outputs variants and causes neither additional parameters nor much extra training cost. Experiments on $8$ IWSLT datasets: (IWSLT14 En $\leftrightarrow$ De, En $\leftrightarrow$ Es, and IWSLT17 En $\leftrightarrow$ Fr, En $\leftrightarrow$ Zh) and the large WMT14 EN $\to$ DE translation tasks show that our method even achieves nearly comparable performance with multiple teacher models with different randomized parameters, both word-level, and sequence-level knowledge distillation. Our code is available at GitHub\footnote{https://github.com/dropreg/RLD} Xiaobo Liang, Lijun Wu 0003, Juntao Li 0005, Tao Qin 0001, Min Zhang 0005, Tie-Yan Liu |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2021 | R-Drop: Regularized Dropout for Neural NetworksabstractDropout is a powerful and widely used technique to regularize the training of deep neural networks. Though effective and performing well, the randomness introduced by dropout causes unnegligible inconsistency between training and inference. In this paper, we introduce a simple consistency training strategy to regularize dropout, namely R-Drop, which forces the output distributions of different sub models generated by dropout to be consistent with each other. Specifically, for each training sample, R-Drop minimizes the bidirectional KL-divergence between the output distributions of two sub models sampled by dropout. Theoretical analysis reveals that R-Drop reduces the above inconsistency. Experiments on $\bf{5}$ widely used deep learning tasks ($\bf{18}$ datasets in total), including neural machine translation, abstractive summarization, language understanding, language modeling, and image classification, show that R-Drop is universally effective. In particular, it yields substantial improvements when applied to fine-tune large-scale pre-trained models, e.g., ViT, RoBERTa-large, and BART, and achieves state-of-the-art (SOTA) performances with the vanilla Transformer model on WMT14 English$\to$German translation ($\bf{30.91}$ BLEU) and WMT14 English$\to$French translation ($\bf{43.95}$ BLEU), even surpassing models trained with extra large-scale data and expert-designed advanced variants of Transformer models. Our code is available at GitHub\footnote{\url{https://github.com/dropreg/R-Drop}}. Xiaobo Liang, Lijun Wu 0003, Juntao Li 0005, Yue Wang 0039, Tao Qin 0001, Wei Chen 0034, Min Zhang 0005, Tie-Yan Liu |
NeurIPS | 1 |
| 2018 | Neural Relation Classification with Text DescriptionsabstractRelation classification is an important task in natural language processing fields. State-of-the-art methods usually concentrate on building deep neural networks based classification models on the training data in which the relations of the labeled entity pairs are given. However, these methods usually suffer from the data sparsity issue greatly. On the other hand, we notice that it is very easily to obtain some concise text descriptions for almost all of the entities in a relation classification task. The text descriptions can provide helpful supplementary information for relation classification. But they are ignored by most of existing methods. In this paper, we propose DesRC, a new neural relation classification method which integrates entities’ text descriptions into deep neural networks models. We design a two-level attention mechanism to select the most useful information from the “intra-sentence” aspect and the “cross-sentence” aspect. Besides, the adversarial training method is also used to further improve the classification per-formance. Finally, we evaluate the proposed method on the SemEval 2010 dataset. Extensive experiments show that our method achieves much better experimental results than other state-of-the-art relation classification methods. Feiliang Ren, Rongsheng Zhao, Yongkang Liu 0002, Xiaobo Liang |
COLING | 7 |