EDBT 2026 Demo / reviewers in the wild / expert
Chunliang Zhang
dblp:54/8637
· DBLP profile ↗
30ranked-venue papers
2as first author
22since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 25 · 2 first-author · 17 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 6 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Probing Preference Representations: A Multi-Dimensional Evaluation and Analysis Method for Reward ModelsabstractPrevious methods evaluate reward models by testing them on a fixed pairwise ranking test set, but they typically do not provide performance information on each preference dimension. In this work, we address the evaluation challenge of reward models by probing preference representations. To confirm the effectiveness of this evaluation method, we construct a Multi-dimensional Reward Model Benchmark (MRMBench), a collection of six probing tasks for different preference dimensions. We design it to favor and encourage reward models that better capture preferences across different dimensions. Furthermore, we introduce an analysis method, inference-time probing, which identifies the dimensions used during the reward prediction and enhances its interpretability. Through extensive experiments, we find that MRMBench strongly correlates with LLM alignment performance, supporting it as a reliable reference for developing advanced reward models. By analyzing the evaluation results on MRMBench, we reveal that reward models struggle to simultaneously capture preferences across multiple dimensions, highlighting the potential of multi-objective optimization in reward modeling. Furthermore, our results demonstrate that the proposed inference-time probing method provides a reliable metric for assessing the confidence of reward predictions, leading to improved alignment of large language models. Chenglong Wang 0002, Yifu Huo, Yang Gan, Yongyu Mu, Qiaozhi He, Murun Yang, Chunliang Zhang, Tongran Liu, Anxiang Ma, Zhengtao Yu 0001, Tong Xiao 0001 |
AAAI | 8 |
| 2026 | GRAM-R²: Self-Training Generative Foundation Reward Models for Reward ReasoningabstractMajor progress in reward modeling over recent years has been driven by a paradigm shift from task-specific designs to generalist reward models. Despite this trend, developing effective reward models remains a fundamental challenge: the heavy reliance on large-scale labeled preference data. Pre-training on abundant unlabeled data offers a promising direction, but existing approaches fall short in instilling explicit reasoning capabilities into reward models. To bridge this gap, we propose a self-training approach that can leverage unlabeled data to scale up reward reasoning in reward models. Based on this approach, we develop GRAM-R² a generative reward model trained to produce not only preference labels but also accompanying reward rationales. GRAM-R² can serve as a foundation model for reward reasoning and can be applied to a wide range of tasks with minimal or no additional fine-tuning. It can support downstream applications such as policy optimization and task-specific reward tuning. Experiments on response ranking, task adaptation, and reinforcement learning from human feedback demonstrate that GRAM-R² consistently delivers strong performance, outperforming several strong discriminative and generative baselines. Chenglong Wang 0002, Yongyu Mu, Yifu Huo, Jiali Zeng, Murun Yang, Xiaoyang Hao, Chunliang Zhang, Fandong Meng, Tong Xiao 0001 |
AAAI | 10 |
| 2025 | RoVRM: A Robust Visual Reward Model Optimized via Auxiliary Textual Preference DataabstractLarge vision-language models (LVLMs) often fail to align with human preferences, leading to issues like generating misleading content without proper visual context (also known as hallucination). A promising solution to this problem is using human-preference alignment techniques, such as best-of-n sampling and reinforcement learning. However, these techniques face the difficulty arising from the scarcity of visual preference data, which is required to train a visual reward model (VRM). In this work, we continue the line of research. We present a Robust Visual Reward Model (RoVRM) which improves human-preference alignment for LVLMs. RoVRM leverages auxiliary textual preference data through a three-phase progressive training and optimal transport-based preference data selection to effectively mitigate the scarcity of visual preference data. We experiment with RoVRM on the commonly used vision-language tasks based on the LLaVA-1.5-7B and -13B models. Experimental results demonstrate that RoVRM consistently outperforms traditional VRMs. Furthermore, our three-phase progressive training and preference data selection approaches can yield consistent performance gains over ranking-based alignment techniques, such as direct preference optimization. Chenglong Wang 0002, Yang Gan, Yifu Huo, Yongyu Mu, Murun Yang, Qiaozhi He, Tong Xiao 0001, Chunliang Zhang, Tongran Liu |
AAAI | 8 |
| 2025 | GRAM: A Generative Foundation Reward Model for Reward GeneralizationabstractIn aligning large language models (LLMs), reward models have played an important role, but are standardly trained as discriminative models and rely only on labeled human preference data. In this paper, we explore methods that train reward models using both unlabeled and labeled data. Building on the generative models in LLMs, we develop a generative reward model that is first trained via large-scale unsupervised learning and then fine-tuned via supervised learning. We also show that by using label smoothing, we are in fact optimizing a regularized pairwise ranking loss. This result, in turn, provides a new view of training reward models, which links generative models and discriminative models under the same class of training objectives. The outcome of these techniques is a foundation reward model, which can be applied to a wide range of tasks with little or no further fine-tuning effort. Extensive experiments show that this model generalizes well across several tasks, including response ranking, reinforcement learning from human feedback, and task adaptation with fine-tuning, achieving significant performance improvements over several strong baseline models. Chenglong Wang 0002, Yang Gan, Yifu Huo, Yongyu Mu, Qiaozhi He, Murun Yang, Tong Xiao 0001, Chunliang Zhang, Tongran Liu |
ICML | 9 |
| 2025 | MRO: Enhancing Reasoning in Diffusion Language Models via Multi-Reward OptimizationabstractRecent advances in diffusion language models (DLMs) have presented a promising alternative to traditional autoregressive large language models (LLMs). However, DLMs still lag behind LLMs in reasoning performance, especially as the number of denoising steps decreases. Our analysis reveals that this shortcoming arises primarily from the independent generation of masked tokens across denoising steps, which fails to capture the token correlation. In this paper, we define two types of token correlation: intra-sequence correlation and inter-sequence correlation, and demonstrate that enhancing these correlations improves reasoning performance. To this end, we propose a Multi-Reward Optimization (MRO) approach, which encourages DLMs to consider the token correlation during the denoising process. More specifically, our MRO approach leverages test-time scaling, reject sampling, and reinforcement learning to directly optimize the token correlation with multiple elaborate rewards. Additionally, we introduce group step and importance sampling strategies to mitigate reward variance and enhance sampling efficiency. Through extensive experiments, we demonstrate that MRO not only improves reasoning performance but also achieves significant sampling speedups while maintaining high performance on reasoning benchmarks. Chenglong Wang 0002, Yang Gan, Chi Hu, Yongyu Mu, Murun Yang, Chunliang Zhang, Tongran Liu, Zhengtao Yu 0001, Tong Xiao 0001 |
NeurIPS | 9 |
| 2024 | Revealing the Parallel Multilingual Learning within Large Language ModelsabstractYongyu Mu, Peinan Feng, Zhiquan Cao, Yuzhang Wu, Bei Li, Chenglong Wang, Tong Xiao, Kai Song, Tongran Liu, Chunliang Zhang, JingBo Zhu. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024. Yongyu Mu, Peinan Feng, Zhiquan Cao, Yuzhang Wu, Chenglong Wang 0002, Tong Xiao 0001, Tongran Liu, Chunliang Zhang |
EMNLP | 10 |
| 2024 | Soft Alignment of Modality Space for End-to-End Speech TranslationabstractEnd-to-end Speech Translation (ST) aims to convert speech into target text within a unified model. The inherent differences between speech and text modalities often impede effective cross-modal and cross-lingual transfer. Existing methods typically employ hard alignment (H-Align) of individual speech and text segments, which can degrade textual representations. To address this, we introduce Soft Alignment (S-Align), using adversarial training to align the representation spaces of both modalities. S-Align creates a modality-invariant space while preserving individual modality quality. Experiments on three languages from the MuST-C dataset show S-Align outperforms H-Align across multiple tasks and offers translation capabilities on par with specialized translation models. Kaiqi Kou, Chen Xu 0008, Chunliang Zhang, Tong Xiao 0001 |
ICASSP | 5 |
| 2024 | Practical Fixed-Time Adaptive ERBFNNs Event-Triggered Control for Uncertain Nonlinear Systems With Dead-Zone ConstraintabstractThe issue of practical fixed-time control is investigated for a category of uncertain nonlinear systems with input dead-zone constraint. Many practical control systems are subject to the constraint of communication resources and input dead zone, which affects the system’s performance and even results in system instability. To handle the above problems, an extended radial basis function neural networks (ERBFNNs) adaptive event-triggered control method is developed to enable the online compensation of input dead zone and schedule the update of control signals. On this foundation, based on the fixed-time stability theorem, a practical fixed-time event-triggered controller is established by the backstepping technique. Technically, the controller can guarantee that the tracking error converges into a small and adjustable set in a fixed time under different initial states, and the boundary of convergence time is dependent on the adjustable design parameters. Meanwhile, all the closed-loop signals are bounded, the communication resources are saved, and the Zeno behavior is also avoided. Finally, some simulation examples are given to illustrate the validity of the presented strategy. Jianhui Wang 0003, Chen Wang 0116, Zhi Liu 0001, C. L. Philip Chen, Chunliang Zhang |
IEEE Trans. Syst. Man Cybern. Syst. | 5 |
| 2023 | Improving End-to-End Speech Translation by Leveraging Auxiliary Speech and Text DataabstractWe present a method for introducing a text encoder into pre-trained end-to-end speech translation systems. It enhances the ability of adapting one modality (i.e., source-language speech) to another (i.e., source-language text). Thus, the speech translation model can learn from both unlabeled and labeled data, especially when the source-language text data is abundant. Beyond this, we present a denoising method to build a robust text encoder that can deal with both normal and noisy text data. Our system sets new state-of-the-arts on the MuST-C En-De, En-Fr, and LibriSpeech En-Fr tasks. Chen Xu 0008, Bojie Hu, Chunliang Zhang, Tong Xiao 0001 |
AAAI | 4 |
| 2023 | Rethinking and Improving Multi-task Learning for End-to-end Speech TranslationabstractSignificant improvements in end-to-end speech translation (ST) have been achieved through the application of multi-task learning.However, the extent to which auxiliary tasks are highly consistent with the ST task, and how much this approach truly helps, have not been thoroughly studied.In this paper, we investigate the consistency between different tasks, considering different times and modules.We find that the textual encoder primarily facilitates cross-modal conversion, but the presence of noise in speech impedes the consistency between text and speech representations.Furthermore, we propose an improved multi-task learning (IMTL) approach for the ST task, which bridges the modal gap by mitigating the difference in length and representation.We conduct experiments on the MuST-C dataset.The results demonstrate that our method attains stateof-the-art results.Moreover, when additional data is used, we achieve the new SOTA result on MuST-C English to Spanish task with 20.8% of the training time required by the current SOTA method. Chen Xu 0008, Tong Xiao 0001, Chunliang Zhang |
EMNLP | 6 |
| 2023 | Large-scale mobile users deployment optimization based on a two-stage hybrid global HS-DE algorithm in multi-UAV-enabled mobile edge computing
Haibin Ouyang, Chunliang Zhang, Steven Li, Liqun Gao |
Eng. Appl. Artif. Intell. | 3 |
| 2023 | Fixed-time adaptive fuzzy event-triggered control for uncertain nonlinear systems with output constraint and actuator failures
Jianhui Wang 0003, Chen Wang 0116, Chunliang Zhang, Zhi Liu 0001, C. L. Philip Chen |
Fuzzy Sets Syst. | 3 |
| 2023 | Fixed-time event-triggered fuzzy adaptive control for uncertain nonlinear systems with full-state constraints
Chen Wang 0116, Jianhui Wang 0003, Yongping Du, Chunliang Zhang, Zhi Liu 0001, C. L. Philip Chen |
Inf. Sci. | 4 |
| 2023 | Finite-time consensus control for multi-agent systems with full-state constraints and actuator failures
Jianhui Wang 0003, Yancheng Yan, Zhi Liu 0001, C. L. Philip Chen, Chunliang Zhang, Kairui Chen |
Neural Networks | 5 |
| 2023 | Fast Finite-Time Event-Triggered Consensus Control for Uncertain Nonlinear Multiagent Systems With Full-State ConstraintsabstractThe fast finite-time event-triggered consensus control is investigated for a category of uncertain nonlinear multiagent systems (MASs) with full-state constraints. The uncertainty of the system is estimated by the radial basis function neural networks (RBFNNs). Furthermore, to achieve the fast finite-time stability and not violate the full-state constraints, a fast finite-time event-triggered consensus control method is proposed. The proposed control method can achieve the fast finite-time stability of the system, and all the followers can track the output signal of the leader. Meanwhile, the system states do not exceed the boundaries of the full-state constraints, and the communication resources of the system can be saved. Finally, some simulation examples are provided to verify the feasibility of the proposed approach. Jianhui Wang 0003, Chen Wang 0116, C. L. Philip Chen, Zhi Liu 0001, Chunliang Zhang |
IEEE Trans. Circuits Syst. I Regul. Pap. | 5 |
| 2022 | Adaptive Robust Control for a Spatial Flexible Timoshenko Manipulator Subject to Input Dead-ZoneabstractThis article investigates the adaptive robust spatial vibration control for a flexible Timoshenko manipulator subject to input dead-zone nonlinearity characteristic. The “disturbance-like” terms and dead-zone nonlinearity are first incorporated into the context of control design, and the new boundary robust adaptive control laws are constructed to reduce the shear deformation and elastic oscillation, ensure the expected angle orientation, handle the input dead-zone, and estimate the upper bound of compound disturbances. The convergence of states and the stability of the system are analyzed and proven without simplifying the infinite dimensional dynamics. In the end, the effectiveness of the presented scheme is demonstrated by the result of simulation research. Shouyan Chen, Zhijia Zhao 0002, Dachang Zhu, Chunliang Zhang, Han-Xiong Li |
IEEE Trans. Syst. Man Cybern. Syst. | 4 |
| 2021 | Learning Light-Weight Translation Models from Deep TransformerabstractRecently, deep models have shown tremendous improvements in neural machine translation (NMT). However, systems of this kind are computationally expensive and memory intensive. In this paper, we take a natural step towards learning strong but light-weight NMT systems. We proposed a novel group-permutation based knowledge distillation approach to compressing the deep Transformer model into a shallow model. The experimental results on several benchmarks validate the effectiveness of our method. Our compressed model is 8 times shallower than the deep model, with almost no loss in BLEU. To further enhance the teacher model, we present a Skipping Sub-Layer method to randomly omit sub-layers to introduce perturbation into training, which achieves a BLEU score of 30.63 on English-German newstest2014. The code is publicly available at https://github.com/libeineu/GPKD. Quan Du, Tong Xiao 0001, Chunliang Zhang |
AAAI | 6 |
| 2021 | An Improved GA with Matrix-Coding for Optimizing a Complex Disassembly Sequence Problem on ELVabstractOptimal disassembly sequencing is an NP-hard problem and has always been an ambition for industry production. In the context of increasing public concerns over environmental impacts, in addition to the feasibility of a disassembly sequence, dismantling enterprises have to consider the relationship between potential profits and the impacts. Thus, an ideal disassembly sequence should weight these three factors comprehensively. Up to now, an appropriate ELV disassembly sequence still mainly relies on people’s intuitive experience and seeking an optimal disassembly sequencing method assumes enormous importance. This paper aims to address the optimal disassembly sequencing problem of ELVs by means of an improved genetic algorithm, in which a matrix coding mechanism and an elite strategy are employed. The weight of different factors can be adjusted according to the actual conditions of factories. The paper gives a case and a series of Pareto fronts are obtained. The effects of population size and maximum evolutionary time on the Pareto solutions were investigated. Ultimately, the optimal Pareto disassembly sequence corresponding to balanced profit and environmental impact is achieved, thereby providing an appropriate disassembly depth defined by the aforementioned disassembly sequence. This can contribute to timely disassembly decisions for end-of-life vehicle (ELV) dismantling enterprises, achieving a cost-effective disassembly process for survival in the context of growing environmental concerns. This paper seeks to offer a viable decision-making approach prior to real disassembly of ELVs by detailing a Pareto disassembly depth and sequence. Chunliang Zhang |
Int. J. Pattern Recognit. Artif. Intell. | 1 |
| 2021 | Novel fuzzy event-triggered adaptive control for nonlinear systems with input hysteresis
Zicong Chen, Jianhui Wang 0003, Kemao Ma, Peisen Zhu, Biaotao He, Chunliang Zhang |
Soft Comput. | 6 |
| 2021 | Self-adaptively commensal learning-based Jaya algorithm with multi-populations and its application
Zuanjia Xie, Chunliang Zhang, Haibin Ouyang, Steven Li, Liqun Gao |
Soft Comput. | 2 |
| 2021 | Adaptive Neural-Network Boundary Control for a Flexible Manipulator With Input Constraints and Model UncertaintiesabstractThis article develops an adaptive neural-network (NN) boundary control scheme for a flexible manipulator subject to input constraints, model uncertainties, and external disturbances. First, a radial basis function NN method is utilized to tackle the unknown input saturations, dead zones, and model uncertainties. Then, based on the backstepping approach, two adaptive NN boundary controllers with update laws are employed to stabilize the like-position loop subsystem and like-posture loop subsystem, respectively. With the introduced control laws, the uniform ultimate boundedness of the deflection and angle tracking errors for the flexible manipulator are guaranteed. Finally, the control performance of the developed control technique is examined by a numerical example. Yong Ren 0003, Zhijia Zhao 0002, Chunliang Zhang, Qinmin Yang, Keum Shik Hong |
IEEE Trans. Cybern. | 3 |
| 2021 | Fuzzy Adaptive Two-Bit-Triggered Control for a Class of Uncertain Nonlinear Systems With Actuator Failures and Dead-Zone ConstraintabstractThis article investigates a fuzzy adaptive two-bit-triggered control for uncertain nonlinear systems with actuator failures and dead-zone constraint. Actuator failures and dead-zone constraint exist frequently in practical systems, which will affect the system performance greatly. Based on the improved fuzzy-logic systems (FLSs), a fuzzy adaptive compensation control is established to address these issues. The approximation error is introduced to the control design as a time-varying function. In addition, for the limited transmission resources of the practical system, a two-bit-triggered control mechanism is proposed to further save system transmission resources. It is proved that the proposed method can guarantee the system tracking performance and all the signals are bounded. Its effectiveness is verified by the simulation examples. Chunliang Zhang, Zicong Chen, Jianhui Wang 0003, Zhi Liu 0001, C. L. Philip Chen |
IEEE Trans. Cybern. | 1 |
| 2020 | Enhanced harmony search algorithm with circular region perturbation for global optimization problems
Wenqiang Wu, Haibin Ouyang, Ali Wagdy Mohamed, Chunliang Zhang, Steven Li |
Appl. Intell. | 4 |
| 2019 | Improved Differentiable Architecture Search for Language Modeling and Named Entity RecognitionabstractYufan Jiang, Chi Hu, Tong Xiao, Chunliang Zhang, Jingbo Zhu. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Yufan Jiang, Chi Hu, Tong Xiao 0001, Chunliang Zhang |
EMNLP/IJCNLP (1) | 4 |
| 2019 | Improved harmony search with general iteration models for engineering design optimization problems
Haibin Ouyang, Wenqiang Wu, Chunliang Zhang, Steven Li, Dexuan Zou, Guiyun Liu |
Soft Comput. | 3 |
| 2018 | Neural network based boundary control of a vibrating string system with input deadzone
Zhijia Zhao 0002, Xiaogang Wang 0011, Chunliang Zhang, Zhijie Liu 0001 |
Neurocomputing | 3 |
| 2017 | Fast Parallel Training of Neural Language ModelsabstractTraining neural language models (NLMs) is very time consuming and we need parallelization for system speedup. However, standard training methods have poor scalability across multiple devices (e.g., GPUs) due to the huge time cost required to transmit data for gradient sharing in the back-propagation process. In this paper we present a sampling-based approach to reducing data transmission for better scaling of NLMs. As a ''bonus'', the resulting model also improves the training speed on a single device. Our approach yields significant speed improvements on a recurrent neural network-based language model. On four NVIDIA GTX1080 GPUs, it achieves a speedup of 2.1+ times over the standard asynchronous stochastic gradient descent baseline, yet with no increase in perplexity. This is even 4.2 times faster than the naive single GPU counterpart. Tong Xiao 0001, Tongran Liu, Chunliang Zhang |
IJCAI | 4 |
| 2016 | Syntactic Skeleton-Based TranslationabstractIn this paper we propose an approach to modeling syntactically-motivated skeletal structure of source sentence for machine translation. This model allows for application of high-level syntactic transfer rules and low-level non-syntactic rules. It thus involves fully syntactic, non-syntactic, and partially syntactic derivations via a single grammar and decoding paradigm. On large-scale Chinese-English and English-Chinese translation tasks, we obtain an average improvement of +0.9 BLEU across the newswire and web genres. Tong Xiao 0001, Chunliang Zhang, Tongran Liu |
AAAI | 3 |
| 2012 | Multi-Aspect Rating Inference with Aspect-Based SegmentationabstractThis paper explores the problem of content-based rating inference from online opinion-based texts, which often expresses differing opinions on multiple aspects. To sufficiently capture information from various aspects, we propose an aspect-based segmentation algorithm to first segment a user review into multiple single-aspect textual parts, and an aspect-augmentation approach to generate the aspect-specific feature vector of each aspect for aspect-based rating inference. To tackle the problem of inconsistent rating annotation, we present a tolerance-based criterion to optimize training sample selection for parameter updating during the model training process. Finally, we present a collaborative rating inference model which explores meaningful correlations between ratings across a set of aspects of user opinions for multi-aspect rating inference. We compared our proposed methods with several other approaches, and experiments on real Chinese restaurant reviews demonstrated that our approaches achieve significant improvements over others. Chunliang Zhang, Matthew Y. Ma |
IEEE Trans. Affect. Comput. | 2 |
| 2011 | Unsupervised Discovery of Domain-Specific Knowledge from Text
Dirk Hovy, Chunliang Zhang, Eduard H. Hovy, Anselmo Peñas |
ACL | 2 |