EDBT 2026 Demo / reviewers in the wild / expert
Changlong Yu
dblp:76/238
· DBLP profile ↗
24ranked-venue papers
5as first author
19since 2021 · last 2026
0000-0002-4758-9014ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 18 · 4 first-author · 14 since 2021Databases, data management, data science and information retrieval · 5 · 3 since 2021Security and privacy · 2 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Low coupling and high interaction dual-branch contrastive pseudo supervision for semi-supervised medical image segmentation
Changlong Yu, Yunfeng Zhang 0001, Rui Zhang 0072, Fangxun Bao, Huijian Han |
Neurocomputing | 1 |
| 2026 | Learning to Optimize Multi-Objective Alignment Through Dynamic Reward WeightingabstractAbstract Prior works in multi-objective reinforcement learning typically use linear reward scalarization with fixed weights, which provably fail to capture non-convex Pareto fronts and thus yield suboptimal results. This limitation becomes especially critical in online preference alignment for large language models. Here, stochastic trajectories generated by parameterized policies create highly non-linear and non-convex mappings from parameters to objectives that no single static weighting scheme can find optimal trade-offs. We address this limitation by introducing dynamic reward weighting, which adaptively adjusts reward weights during the online reinforcement learning process. Unlike existing approaches that rely on fixed-weight interpolation, our dynamic weighting continuously balances and prioritizes objectives in training, facilitating effective exploration of Pareto fronts in objective space. We introduce two approaches of increasing sophistication and generalizability: hypervolume-guided weight adaptation and gradient-based weight optimization, offering a versatile toolkit for online multi-objective alignment. Our extensive experiments demonstrate their compatibility with commonly used online reinforcement learning algorithms, effectiveness across multiple datasets, and applicability to different model families, consistently achieving Pareto dominant solutions with fewer training steps than fixed-weight linear scalarization baselines. Yining Lu, Changlong Yu, Qingyu Yin, Meng Jiang 0001 |
Trans. Assoc. Comput. Linguistics | 5 |
| 2025 | EcomScriptBench: A Multi-task Benchmark for E-commerce Script Planning via Step-wise Intention-Driven Product AssociationabstractWeiqi Wang, Limeng Cui, Xin Liu, Sreyashi Nag, Wenju Xu, Chen Luo, Sheikh Muhammad Sarwar, Yang Li, Hansu Gu, Hui Liu, Changlong Yu, Jiaxin Bai, Yifan Gao, Haiyang Zhang, Qi He, Shuiwang Ji, Yangqiu Song. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Weiqi Wang 0001, Limeng Cui, Xin Liu 0039, Sreyashi Nag, Wenju Xu, Chen Luo 0003, Sheikh Muhammad Sarwar, Yang Li 0055, Hansu Gu, Hui Liu 0033, Changlong Yu, Jiaxin Bai, Yifan Gao 0001, Qi He 0002, Shuiwang Ji, Yangqiu Song |
ACL (1) | 11 |
| 2025 | WebAgent-R1: Training Web Agents via End-to-End Multi-Turn Reinforcement LearningabstractZhepei Wei, Wenlin Yao, Yao Liu, Weizhi Zhang, Qin Lu, Liang Qiu, Changlong Yu, Puyang Xu, Chao Zhang, Bing Yin, Hyokun Yun, Lihong Li. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Zhepei Wei, Wenlin Yao, Changlong Yu, Puyang Xu, Chao Zhang 0014, Hyokun Yun, Lihong Li 0001 |
EMNLP | 7 |
| 2025 | Discriminative Finetuning of Generative Large Language Models without Reward Models and Human Preference DataabstractSupervised fine-tuning (SFT) has become a crucial step for aligning pretrained large language models (LLMs) using supervised datasets of input-output pairs. However, despite being supervised, SFT is inherently limited by its generative training objective. To address its limitations, the existing common strategy is to follow SFT with a separate phase of preference optimization (PO), which relies on either human-labeled preference data or a strong reward model to guide the learning process. In this paper, we address the limitations of SFT by exploring one of the most successful techniques in conventional supervised learning: discriminative learning. We introduce Discriminative Fine-Tuning (DFT), an improved variant of SFT, which mitigates the burden of collecting human-labeled preference data or training strong reward models. Unlike SFT that employs a generative approach and overlooks negative data, DFT adopts a discriminative paradigm that increases the probability of positive answers while suppressing potentially negative ones, aiming for data prediction instead of token prediction. Our contributions include: (i) a discriminative probabilistic framework for fine-tuning LLMs by explicitly modeling the discriminative likelihood of an answer among all possible outputs given an input; (ii) efficient algorithms to optimize this discriminative likelihood; and (iii) extensive experiments demonstrating DFT’s effectiveness, achieving performance better than SFT and comparable to if not better than SFT$\rightarrow$PO. The code can be found at https://github.com/Optimization-AI/DFT. Siqi Guo 0003, Ilgee Hong, Vicente Balmaseda, Changlong Yu, Haoming Jiang, Tuo Zhao, Tianbao Yang |
ICML | 4 |
| 2025 | Think-RM: Enabling Long-Horizon Reasoning in Generative Reward ModelsabstractReinforcement learning from human feedback (RLHF) has become a powerful post-training paradigm for aligning large language models with human preferences. A core challenge in RLHF is constructing accurate reward signals, where the conventional Bradley-Terry reward models (BT RMs) often suffer from sensitivity to data size and coverage, as well as vulnerability to reward hacking. Generative reward models (GenRMs) offer a more robust alternative by generating chain-of-thought (CoT) rationales followed by a final verdict. However, existing GenRMs rely on shallow, vertically scaled reasoning, limiting their capacity to handle nuanced or complex tasks. Moreover, their pairwise preference outputs are incompatible with standard RLHF algorithms that require pointwise reward signals. In this work, we introduce Think-RM, a training framework that enables long-horizon reasoning in GenRMs by modeling an internal thinking process. Rather than producing structured, externally provided rationales, Think-RM generates flexible, self-guided reasoning traces that support advanced capabilities such as self-reflection, hypothetical reasoning, and divergent reasoning. To elicit these reasoning abilities, we first warm-up the models by supervised fine-tuning (SFT) over long CoT data. We then further improve the model's long-horizon abilities by rule-based reinforcement learning (RL). In addition, we propose a novel pairwise RLHF pipeline that directly optimizes policies from pairwise comparisons, eliminating the need for pointwise reward conversion. Experiments show that Think-RM outperforms baselines on both in-distribution and out-of-distribution tasks, with particularly strong gains on reasoning-heavy benchmarks: more than 10\% and 5\% on RewardBench's Chat Hard and Reasoning, and 12\% on RM-Bench's Math domain. When combined with our pairwise RLHF pipeline, it demonstrates superior end-policy performance compared to traditional approaches. This depth-oriented approach not only broadens the GenRM design space but also establishes a new paradigm for preference-based policy optimization in RLHF. Ilgee Hong, Changlong Yu, Weixiang Yan, Zhenghao Xu, Haoming Jiang, Qingru Zhang, Xin Liu 0039, Chao Zhang 0014, Tuo Zhao |
NeurIPS | 2 |
| 2025 | Ask a Strong LLM Judge when Your Reward Model is UncertainabstractReward model (RM) plays a pivotal role in reinforcement learning with human feedback (RLHF) for aligning large language models (LLMs). However, classical RMs trained on human preferences are vulnerable to reward hacking and generalize poorly to out-of-distribution (OOD) inputs.
By contrast, strong LLM judges equipped with reasoning capabilities demonstrate superior generalization, even without additional training, but incur significantly higher inference costs, limiting their applicability in online RLHF.
In this work, we propose an uncertainty-based routing framework that efficiently complements a fast RM with a strong but costly LLM judge. Our approach formulates advantage estimation in policy gradient (PG) methods as pairwise preference classification, enabling principled uncertainty quantification to guide routing. Uncertain pairs are forwarded to the LLM judge, while confident ones are evaluated by the RM. Experiments on RM benchmarks demonstrate that our uncertainty-based routing strategy significantly outperforms random judge calling at the same cost, and downstream alignment results showcase its effectiveness in improving online RLHF. Zhenghao Xu, Qingru Zhang, Ilgee Hong, Changlong Yu, Wenlin Yao, Haoming Jiang, Lihong Li 0001, Hyokun Yun, Tuo Zhao |
NeurIPS | 6 |
| 2025 | TAPAS: An Efficient Online APT Detection with Task-guided Process Provenance Graph Segmentation and Analysis
Bo Zhang 0150, Yansong Gao 0001, Changlong Yu, Boyu Kuang, Zhi Zhang 0001, Hyoungshick Kim, Anmin Fu |
USENIX Security Symposium | 3 |
| 2025 | Event AutoAugment: Word-level data augmentation for textual event detection
Huan Zhao 0002, Wei Wang 0138, Changlong Yu, Ruifeng Xu 0001 |
Expert Syst. Appl. | 4 |
| 2025 | Decider: A Dual-System Rule-Controllable Decoding Framework for Language GenerationabstractConstrained decoding approaches aim to control the meaning or style of text generated by a Pre-trained Language Model (PLM) for various task-specific objectives at inference time. However, these methods often guide plausible continuations by greedily and explicitly selecting targets, which, while fulfilling the task requirements, may overlook the natural patterns of human language generation. In this work, we propose a novel decoding framework,Decider, which enables us to program high-level rules on how we might effectively complete tasks to control a PLM. Differing from previous works, our framework transforms the encouragement of concrete target words into the encouragement of all words that satisfy the high-level rules. Specifically,Decideris a dual system in which a PLM is equipped and controlled by a First-Order Logic (FOL) reasoner to express and evaluate the rules, along with a decision function that merges the outputs from both systems to guide the generation. Experiments on CommonGen and PersonaChat demonstrate thatDecidercan effectively follow given rules to guide a PLM in achieving generation tasks in a more human-like manner. Tian Lan 0003, Changlong Yu, Wei Wang 0138, Qunxi Dong, Kun Qian 0003, Piji Li, Wei Bi, Bin Hu 0001 |
IEEE Trans. Knowl. Data Eng. | 4 |
| 2024 | MIND: Multimodal Shopping Intention Distillation from Large Vision-language Models for E-commerce Purchase UnderstandingabstractBaixuan Xu, Weiqi Wang, Haochen Shi, Wenxuan Ding, Huihao Jing, Tianqing Fang, Jiaxin Bai, Xin Liu, Changlong Yu, Zheng Li, Chen Luo, Qingyu Yin, Bing Yin, Long Chen, Yangqiu Song. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024. Baixuan Xu, Weiqi Wang 0001, Wenxuan Ding 0001, Huihao Jing, Tianqing Fang, Jiaxin Bai, Xin Liu 0039, Changlong Yu, Zheng Li 0018, Chen Luo 0003, Qingyu Yin, Long Chen 0016, Yangqiu Song |
EMNLP | 9 |
| 2024 | SNIPER: Detect Complex Attacks Accurately from Traffic
Changlong Yu, Bo Zhang 0150, Boyu Kuang, Anmin Fu |
ISPEC | 1 |
| 2023 | A Generative Approach for Script Event Prediction via Contrastive Fine-TuningabstractScript event prediction aims to predict the subsequent event given the context. This requires the capability to infer the correlations between events. Recent works have attempted to improve event correlation reasoning by using pretrained language models and incorporating external knowledge (e.g., discourse relations). Though promising results have been achieved, some challenges still remain. First, the pretrained language models adopted by current works ignore event-level knowledge, resulting in an inability to capture the correlations between events well. Second, modeling correlations between events with discourse relations is limited because it can only capture explicit correlations between events with discourse markers, and cannot capture many implicit correlations. To this end, we propose a novel generative approach for this task, in which a pretrained language model is fine-tuned with an event-centric pretraining objective and predicts the next event within a generative paradigm. Specifically, we first introduce a novel event-level blank infilling strategy as the learning objective to inject event-level knowledge into the pretrained language model, and then design a likelihood-based contrastive loss for fine-tuning the generative model. Instead of using an additional prediction layer, we perform prediction by using sequence likelihoods generated by the generative model. Our approach models correlations between events in a soft way without any external knowledge. The likelihood-based prediction eliminates the need to use additional networks to make predictions and is somewhat interpretable since it scores each word in the event. Experimental results on the multi-choice narrative cloze (MCNC) task demonstrate that our approach achieves better results than other state-of-the-art baselines. Our code will be available at https://github.com/zhufq00/mcnc. Fangqi Zhu, Changlong Yu, Wei Wang 0138, Xin Mu, Min Yang 0007, Ruifeng Xu 0001 |
AAAI | 3 |
| 2023 | Controllable Contrastive Generation for Multilingual Biomedical Entity LinkingabstractMultilingual biomedical entity linking (MBEL) aims to map language-specific mentions in the biomedical text to standardized concepts in a multilingual knowledge base (KB) such as Unified Medical Language System (UMLS).In this paper, we propose Con2GEN, a prompt-based controllable contrastive generation framework for MBEL, which summarizes multidimensional information of the UMLS concept mentioned in biomedical text into a natural sentence following a predefined template.Instead of tackling the MBEL problem with a discriminative classifier, we formulate it as a sequence-tosequence generation task, which better exploits the shared dependencies between source mentions and target entities.Moreover, Con2GEN matches against UMLS concepts in as many languages and types as possible, hence facilitating cross-information disambiguation.Extensive experiments show that our model achieves promising performance improvements compared with several state-of-the-art techniques on the XL-BEL and the Mantra GSC datasets spanning 12 typologically diverse languages. Tiantian Zhu 0002, Yang Qin 0001, Qingcai Chen, Xin Mu, Changlong Yu, Yang Xiang 0003 |
EMNLP | 5 |
| 2022 | Improving Event Representation via Simultaneous Weakly Supervised Contrastive Learning and ClusteringabstractRepresentations of events described in text are important for various tasks.In this work, we present SWCC: a Simultaneous Weakly supervised Contrastive learning and Clustering framework for event representation learning.SWCC learns event representations by making better use of co-occurrence information of events.Specifically, we introduce a weakly supervised contrastive learning method that allows us to consider multiple positives and multiple negatives, and a prototype-based clustering method that avoids semantically related events being pulled apart.For model training, SWCC learns representations by simultaneously performing weakly supervised contrastive learning and prototypebased clustering.Experimental results show that SWCC outperforms other baselines on Hard Similarity and Transitive Sentence Similarity tasks.In addition, a thorough analysis of the prototypebased clustering method demonstrates that the learned prototype vectors are able to implicitly capture various relations between events.Our code will be available at https://github. com/gaojun4ever/SWCC4Event. Wei Wang 0138, Changlong Yu, Huan Zhao 0002, Wilfred Ng, Ruifeng Xu 0001 |
ACL (1) | 3 |
| 2022 | XDM: Improving Sequential Deep Matching with Unclicked User Behaviors for Recommender System
Fuyu Lv, Mengxue Li, Tonglei Guo, Changlong Yu, Fei Sun 0001, Taiwei Jin, Wilfred Ng |
DASFAA (3) | 4 |
| 2022 | Title2Event: Benchmarking Open Event Extraction with a Large-scale Chinese Title DatasetabstractHaolin Deng, Yanan Zhang, Yangfan Zhang, Wangyang Ying, Changlong Yu, Jun Gao, Wei Wang, Xiaoling Bai, Nan Yang, Jin Ma, Xiang Chen, Tianhua Zhou. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. 2022. Haolin Deng, Yangfan Zhang, Wangyang Ying, Changlong Yu, Wei Wang 0138, Xiaoling Bai, Jin Ma 0003, Tianhua Zhou |
EMNLP | 5 |
| 2022 | An Empirical Revisiting of Linguistic Knowledge Fusion in Language Understanding TasksabstractThough linguistic knowledge emerges during large-scale language model pretraining, recent work attempt to explicitly incorporate humandefined linguistic priors into task-specific finetuning.Infusing language models with syntactic or semantic knowledge from parsers has shown improvements on many language understanding tasks.To further investigate the effectiveness of structural linguistic priors, we conduct empirical study of replacing parsed graphs or trees with trivial ones (rarely carrying linguistic knowledge e.g., balanced tree) for tasks in the GLUE benchmark.Encoding with trivial graphs achieves competitive or even better performance in fully-supervised and few-shot settings.It reveals that the gains might not be significantly attributed to explicit linguistic priors but rather to more feature interactions brought by fusion layers.Hence we call for attention to using trivial graphs as necessary baselines to design advanced knowledge fusion methods in the future. Changlong Yu, Tianyi Xiao, Lingpeng Kong, Yangqiu Song, Wilfred Ng |
EMNLP | 1 |
| 2022 | Translation-Based Implicit Annotation Projection for Zero-Shot Cross-Lingual Event Argument ExtractionabstractZero-shot cross-lingual event argument extraction (EAE) is a challenging yet practical problem in Information Extraction. Most previous works heavily rely on external structured linguistic features, which are not easily accessible in real-world scenarios. This paper investigates a translation-based method to implicitly project annotations from the source language to the target language. With the use of translation-based parallel corpora, no additional linguistic features are required during training and inference. As a result, the proposed approach is more cost effective than previous works on zero-shot cross-lingual EAE. Moreover, our implicit annotation projection approach introduces less noises and hence is more effective and robust than explicit ones. Experimental results show that our model achieves the best performance, outperforming a number of competitive baselines. The thorough analysis further demonstrates the effectiveness of our model compared to explicit annotation projection approaches. Chenwei Lou, Changlong Yu, Wei Wang 0138, Huan Zhao 0002, Wei-Wei Tu, Ruifeng Xu 0001 |
SIGIR | 3 |
| 2020 | Hypernymy Detection for Low-Resource Languages via Meta LearningabstractHypernymy detection, a.k.a.lexical entailment, is a fundamental sub-task of many natural language understanding tasks.Previous explorations mostly focus on monolingual hypernymy detection on high-resource languages, e.g., English, but few investigate the lowresource scenarios.This paper addresses the problem of low-resource hypernymy detection by combining high-resource languages.We extensively compare three joint training paradigms and for the first time propose applying meta learning to relieve the low-resource issue.Experiments demonstrate the superiority of our method among the three settings, which substantially improves the performance of extremely low-resource languages by preventing over-fitting on small datasets.* Work done when C. Yu and J. Han were with Tencent AI Lab. Changlong Yu, Jialong Han, Haisong Zhang, Wilfred Ng |
ACL | 1 |
| 2020 | When Hearst Is not Enough: Improving Hypernymy Detection from Corpus with Distributional ModelsabstractWe address hypernymy detection, i.e., whether an is-a relationship exists between words (x, y), with the help of large textual corpora.Most conventional approaches to this task have been categorized to be either pattern-based or distributional.Recent studies suggest that pattern-based ones are superior, if large-scale Hearst pairs are extracted and fed, with the sparsity of unseen (x, y) pairs relieved.However, they become invalid in some specific sparsity cases, where x or y is not involved in any pattern.For the first time, this paper quantifies the non-negligible existence of those specific cases.We also demonstrate that distributional methods are ideal to make up for patternbased ones in such cases.We devise a complementary framework, under which a patternbased and a distributional model collaborate seamlessly in cases which they each prefer.On several benchmark datasets, our framework achieves competitive improvements and the case study shows its better interpretability. Changlong Yu, Jialong Han, Peifeng Wang, Yangqiu Song, Hongming Zhang 0009, Wilfred Ng, Shuming Shi 0001 |
EMNLP (1) | 1 |
| 2019 | SDM: Sequential Deep Matching Model for Online Large-scale Recommender SystemabstractCapturing users' precise preferences is a fundamental problem in large-scale recommender system. Currently, item-based Collaborative Filtering (CF) methods are common matching approaches in industry. However, they are not effective to model dynamic and evolving preferences of users. In this paper, we propose a new sequential deep matching (SDM) model to capture users' dynamic preferences by combining short-term sessions and long-term behaviors. Compared with existing sequence-aware recommendation methods, we tackle the following two inherent problems in real-world applications: (1) there could exist multiple interest tendencies in one session. (2) long-term preferences may not be effectively fused with current session interests. Long-term behaviors are various and complex, hence those highly related to the short-term session should be kept for fusion. We propose to encode behavior sequences with two corresponding components: multi-head self-attention module to capture multiple types of interests and long-short term gated fusion module to incorporate long-term preferences. Successive items are recommended after matching between sequential user behavior vector and item embedding vectors. Offline experiments on real-world datasets show the superior performance of the proposed SDM. Moreover, SDM has been successfully deployed on online large-scale recommender system at Taobao and achieves improvements in terms of a range of commercial metrics. Fuyu Lv, Taiwei Jin, Changlong Yu, Fei Sun 0001, Quan Lin, Keping Yang, Wilfred Ng |
CIKM | 3 |
| 2019 | Multiplex Word Embeddings for Selectional Preference AcquisitionabstractHongming Zhang, Jiaxin Bai, Yan Song, Kun Xu, Changlong Yu, Yangqiu Song, Wilfred Ng, Dong Yu. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Hongming Zhang 0009, Jiaxin Bai, Yan Song 0003, Kun Xu 0005, Changlong Yu, Yangqiu Song, Wilfred Ng, Dong Yu 0001 |
EMNLP/IJCNLP (1) | 5 |
| 2017 | A Neural Network Model for Semi-supervised Review Aspect Identification
Ying Ding 0005, Changlong Yu, Jing Jiang 0001 |
PAKDD (2) | 2 |