EDBT 2026 Demo / reviewers in the wild / expert
Chenhe Dong
dblp:254/8252
· DBLP profile ↗
9ranked-venue papers
2as first author
8since 2021 · last 2026
0000-0002-2211-5138ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 4 · 1 first-author · 4 since 2021Databases, data management, data science and information retrieval · 4 · 1 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Learning to Trust: Dynamic Utilization of Retrieval-Augmented Generation for E-commerce Search RelevanceabstractAccurately estimating query-item relevance is vital for e-commerce ranking and conversion. While Large Language Models (LLMs) excel at reasoning, they often lack specialized knowledge required for long-tail or fast-evolving queries, necessitating Retrieval-Augmented Generation (RAG). However, production environments face three critical challenges: (1) external context is inherently noisy and inconsistent; (2) extreme latency budgets prohibit multi-stage processing or refinement; and (3) the model must simultaneously assess relevance and context-trust within a unified inference pass. We propose DyKnow-RAG, a reinforcement learning framework that teaches LLMs to learn to trust through dynamic utilization of external knowledge. Built on Group Relative Policy Optimization (GRPO), DyKnow-RAG utilizes a dual-group rollout strategy (parametric-only vs. with-context) and a posterior-driven inter-group advantage scaling mechanism. This enables the model to optimize context utilization without human process labels or extra inference overhead. Our pipeline further integrates structured Chain-of-Thought (CoT) and an uncertainty-prioritized RL pool to stabilize training. Offline evaluations show significant Macro-F1 and Accuracy gains, particularly on noise-sensitive query slices. Importantly, DyKnow-RAG has been deployed in Taobao's production system, serving hundreds of millions of active users and billions of daily search requests. Controlled A/B tests demonstrate consistent lifts in key business metrics, including GSB and Item Goodrate, while maintaining a p99 latency under 400ms. This work provides a scalable and deployable paradigm for operationalizing noisy RAG under extreme efficiency constraints of large-scale industrial search. Tingqiao Xu, Shaowei Yao, Chenhe Dong, Zerui Huang, Dan Ou, Haihong Tang, Bo Zheng 0007 |
SIGIR | 3 |
| 2026 | TaoSR-AGRL: Adaptive Guided Reinforcement Learning Framework for E-commerce Search RelevanceabstractQuery-product relevance prediction is fundamental to e-commerce search and has become even more critical in the era of AI-powered shopping, where semantic understanding and complex reasoning directly shape the user experience and business conversion. Large Language Models (LLMs) enable generative, reasoning-based approaches, typically aligned via supervised fine-tuning (SFT) or preference optimization methods like Direct Preference Optimization (DPO). However, the increasing complexity of business rules and user queries exposes the inability of existing methods to endow models with robust reasoning capacity for long-tail and challenging cases. Efforts to address this via reinforcement learning strategies like Group Relative Policy Optimization (GRPO) often suffer from sparse terminal rewards, offering insufficient guidance for multi-step reasoning and slowing convergence. To address these challenges, we propose TaoSR-AGRL, an Adaptive Guided Reinforcement Learning framework for LLM-based relevance prediction in Taobao Search Relevance. TaoSR-AGRL introduces two key innovations: (1) Rule-aware Reward Shaping, which decomposes the final relevance judgment into dense, structured rewards aligned with domain-specific relevance criteria; and (2) Adaptive Guided Replay, which identifies low-accuracy rollouts during training and injects targeted ground-truth guidance to steer the policy away from stagnant, rule-violating reasoning patterns toward compliant trajectories. TaoSR-AGRL was evaluated on large-scale real-world datasets and through online side-by-side human evaluations on Taobao Search. It consistently outperforms DPO and standard GRPO baselines in offline experiments, improving relevance accuracy, rule adherence, and training stability. The model trained with TaoSR-AGRL has been successfully deployed in the main search scenario on Taobao, serving hundreds of millions of users. Jianhui Yang 0001, Pengkun Jiao, Chenhe Dong, Zerui Huang, Shaowei Yao, Xiaojiang Zhou, Dan Ou, Haihong Tang |
WWW | 4 |
| 2025 | Enhancing Factual Consistency in Text Summarization via Counterfactual DebiasingabstractDespite significant progress in abstractive text summarization aimed at generating fluent and informative outputs, how to ensure the factual consistency of generated summaries remains a crucial and challenging issue. In this study, drawing inspiration from advancements in causal inference, we construct causal graphs to analyze the process of abstractive text summarization methods and identify intrinsic causes of factual inconsistency, specifically language bias and irrelevancy bias, and we propose CoFactSum, a novel framework that mitigates the causal effects of these biases through counterfactual estimation for enhancing the factual consistency of the generated content. CoFactSum provides two counterfactual estimation strategies, including Explicit Counterfactual Masking, which employs a dynamic masking approach, and Implicit Counterfactual Training, which utilizes a discriminative cross-attention mechanism. Besides, we propose a Debiasing Degree Adjustment mechanism to dynamically calibrate the level of debiasing at each decoding step. Extensive experiments conducted on two widely used summarization datasets demonstrate the effectiveness and advantages of the proposed CoFactSum in enhancing the factual consistency of generated summaries, outperforming several baseline methods. Zhenqing Ling, Yuexiang Xie, Chenhe Dong, Ying Shen 0001 |
COLING | 3 |
| 2024 | A Unified Framework for Contextual and Factoid Question GenerationabstractQuestion generation (QG) aims to automatically generate fluent and relevant questions, where the two most mainstream directions are generating questions from unstructured contextual texts (CQG), such as news articles, and generating questions from structured factoid texts (FQG), such as knowledge graphs or tables. Existing methods for these two tasks mainly face challenges of limited internal structural information as well as scarce background information, while these two tasks can benefit each other for alleviating these issues. For example, when meeting the entity mention “United Kingdom” in CQG, it can be inferred that it is a country in European continent based on the structural knowledge “(Europe, countries_within, United Kingdom)” in FQG. And when meeting the entity “Houston Rockets” in FQG, more background information, such as “an American professional basketball team based in Houston since 1971”, can be found in the related passages of CQG. To this end, we propose a unified framework for the tasks of CQG and FQG, where: (i) two types of task-sharing modules are developed to learn shared contextual and structural knowledge, where the task format is unified with a pseudo passage reformulation strategy; (ii) for the CQG task, a task-specific knowledge module with a knowledge selection and aggregation mechanism is introduced, so as to incorporate more factoid knowledge from external knowledge graphs and alleviate the word ambiguity problem; and (iii) for the FQG task, a task-specific passage module with a multi-level passage fusion mechanism is designed to extract fine-grained word-level knowledge. Experimental results in both automatic and human evaluation show the effectiveness of our proposed method. Chenhe Dong, Ying Shen 0001, Shiyang Lin, Zhenzhou Lin, Yang Deng 0002 |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2023 | Unbiased curriculum learning enhanced global-local graph neural network for protein thermodynamic stability predictionabstractMOTIVATION: Proteins play crucial roles in biological processes, with their functions being closely tied to thermodynamic stability. However, measuring stability changes upon point mutations of amino acid residues using physical methods can be time-consuming. In recent years, several computational methods for protein thermodynamic stability prediction (PTSP) based on deep learning have emerged. Nevertheless, these approaches either overlook the natural topology of protein structures or neglect the inherent noisy samples resulting from theoretical calculation or experimental errors. RESULTS: We propose a novel Global-Local Graph Neural Network powered by Unbiased Curriculum Learning for the PTSP task. Our method first builds a Siamese graph neural network to extract protein features before and after mutation. Since the graph's topological changes stem from local node mutations, we design a local feature transformation module to make the model focus on the mutated site. To address model bias caused by noisy samples, which represent unavoidable errors from physical experiments, we introduce an unbiased curriculum learning method. This approach effectively identifies and re-weights noisy samples during the training process. Extensive experiments demonstrate that our proposed method outperforms advanced protein stability prediction methods, and surpasses state-of-the-art learning methods for regression prediction tasks. AVAILABILITY AND IMPLEMENTATION: All code and data is available at https://github.com/haifangong/UCL-GLGNN. Haifan Gong, Chenhe Dong, Yue Wang 0101, Guanqi Chen, Bilin Liang, Haofeng Li, Lanxuan Liu, Jie Xu 0068, Guanbin Li |
Bioinform. | 3 |
| 2022 | Cross-perspective Graph Contrastive Learning
Shiyang Lin, Chenhe Dong, Ying Shen 0001 |
KSEM (1) | 2 |
| 2021 | HRKD: Hierarchical Relational Knowledge Distillation for Cross-domain Language Model CompressionabstractOn many natural language processing tasks, large pre-trained language models (PLMs) have shown overwhelming performances compared with traditional neural network methods.Nevertheless, their huge model size and low inference speed have hindered the deployment on resource-limited devices in practice.In this paper, we target to compress PLMs with knowledge distillation, and propose a hierarchical relational knowledge distillation (HRKD) method to capture both hierarchical and domain relational information.Specifically, to enhance the model capability and transferability, we leverage the idea of metalearning and set up domain-relational graphs to capture the relational information across different domains.And to dynamically select the most representative prototypes for each domain, we propose a hierarchical compareaggregate mechanism to capture hierarchical relationships.Extensive experiments on public multi-domain datasets demonstrate the superior performance of our HRKD method as well as its strong few-shot learning ability.For reproducibility, we release the code at https: //github.com/cheneydon/hrkd. Chenhe Dong, Yaliang Li, Ying Shen 0001, Minghui Qiu |
EMNLP (1) | 1 |
| 2021 | Wasserstein Distance-Based Auto-Encoder Tracking
Ying Wei 0007, Chenhe Dong, Chuaqiao Xu, Zhaofu Diao |
Neural Process. Lett. | 3 |
| 2019 | Combination of Appearance and License Plate Features for Vehicle Re-IdentificationabstractIn this work, we propose a two-module framework that combines appearance and corresponding license plate features for vehicle re-identification (Re-ID). In the appearance module, we design a Two-Branch Network to extract comprehensive global features. To obtain more discriminative feature representations, we propose an enhanced triplet loss (ETL) and combine ETL with softmax loss to optimize the parameter of Two-Branch Network. In the license plate module, we present a license plate Re-ID network that incorporates the bidirectional LSTMs into CNNs, which is effective for capturing the contexts in license plate images and significantly improves the performance of license plate Re-ID. We validate our method on both VeRi-776 dataset [1] and VehicleID dataset [2]. The experimental results show that our method outperforms most state-of-the-art approaches for vehicle Re-ID, even if only the appearance module is used. Yanguang He, Chenhe Dong, Ying Wei 0007 |
ICIP | 2 |