EDBT 2026 Demo / reviewers in the wild / expert
Yingyou Wen
dblp:40/1624
· DBLP profile ↗
20ranked-venue papers
1as first author
12since 2021 · last 2027
0000-0002-6659-1785ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 7 · 6 since 2021Databases, data management, data science and information retrieval · 6 · 5 since 2021Security and privacy · 3Computer networks · 2 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2Systems, architecture and hardware · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2027 | Stage-aware LLM-SLM collaboration for agentic tasks via joint routing and verification
Qingwen Yang, Xuejing Li, Tiezheng Guo, Yanyi Liu, Feiyu Qu, Yingyou Wen |
Inf. Process. Manag. | 7 |
| 2026 | From Detection to Diagnosis: Advancing Hallucination Analysis with Automated Data SynthesisabstractHallucinations in Large Language Models (LLMs), defined as the generation of content inconsistent with facts or context, represent a core obstacle to their reliable deployment in critical domains. Current research primarily focuses on binary "detection" approaches that, while capable of identifying hallucinations, fail to provide interpretable and actionable feedback for model improvement, thus limiting practical utility. To address this limitation, a new research paradigm is proposed, shifting from "detection" to "diagnosis". The Hallucination Diagnosis Task is introduced, a task which requires models to not only detect hallucinations, but also perform error localization, causal explanation, and content correction. We develop the Hallucination Diagnosis Generator (HDG), an automated pipeline that systematically generates high-quality training samples with rich diagnostic metadata from raw corpora through multi-dimensional augmentation strategies including controlled fact fabrication and reasoning chain perturbation. Using HDG-generated data, we train HDM-4B-RL, a 4-billion-parameter hallucination diagnosis model, employing Group Relative Policy Optimization (GRPO) with a comprehensive reward function incorporating structural, accuracy, and localization signals. Experimental results demonstrate that our model surpasses previous state-of-the-art detection models on the HaluEval benchmark while achieving comparable performance to advanced general-purpose models. In comprehensive diagnosis tasks, HDM-4B-RL matches the capabilities of larger general models while maintaining a smaller size. This work validates the feasibility and value of hallucination diagnosis, providing an effective methodology for building more trustworthy and reliable generative AI systems. Yanyi Liu, Qingwen Yang, Tiezheng Guo, Feiyu Qu, Yingyou Wen |
AAAI | 6 |
| 2026 | Augmented Runtime Collaboration for Self-Organizing Multi-Agent Systems: A Hybrid Bi-Criteria Routing ApproachabstractLLM-based multi-agent systems have demonstrated significant capabilities across diverse domains. However, the task performance and efficiency are fundamentally constrained by their collaboration strategies. Prevailing approaches rely on static topologies and centralized global planning, a paradigm that limits their scalability and adaptability in open, decentralized networks. Effective collaboration planning in distributed systems using only local information thus remains a formidable challenge. To address this, we propose BiRouter, a novel dual-criteria routing method for Self-Organizing Multi-Agent Systems (SO-MAS). This method enables each agent to autonomously execute "next-hop" task routing at runtime, relying solely on local information. Its core decision-making mechanism is predicated on balancing two metrics: (1) the ImpScore, which evaluates a candidate agent's long-term importance to the overall goal, and (2) the GapScore, which assesses its contextual continuity for the current task state. Furthermore, we introduce a dynamically updated reputation mechanism to bolster system robustness in untrustworthy environments and have developed a large-scale, cross-domain dataset, comprising thousands of annotated task-routing paths, to enhance the model's generalization. Extensive experiments demonstrate that BiRouter achieves superior performance and token efficiency over existing baselines, while maintaining strong robustness and effectiveness in information-limited, decentralized, and untrustworthy settings. Qingwen Yang, Feiyu Qu, Tiezheng Guo, Yanyi Liu, Yingyou Wen |
AAAI | 5 |
| 2026 | WADSeg: Exploiting weak attention associations for enhanced knowledge segmentation in RAG
Tiezheng Guo, Chen Wang 0150, Qingwen Yang, Yanyi Liu, Yingyou Wen |
Expert Syst. Appl. | 6 |
| 2026 | LOOM: Weaving high-quality long-texts through hierarchical planning and reflective feedback
Tiezheng Guo, Qingwen Yang, Yanyi Liu, Feiyu Qu, Yingyou Wen |
Expert Syst. Appl. | 6 |
| 2026 | From answering to discussing: Advancing human-AI cognitive collaboration in dialogue agents
Junchi Wang, Qingwen Yang, Tiezheng Guo, Yanyi Liu, Yingyou Wen |
Inf. Process. Manag. | 6 |
| 2025 | Leveraging inter-chunk interactions for enhanced retrieval in large language model-based question answering
Tiezheng Guo, Yanyi Liu, Sai Xu, Qingwen Yang, Xianlin Gao, Yingyou Wen |
Neurocomputing | 10 |
| 2025 | Adaptive-TOD: An LLM-driven and adaptive agent for diverse interaction modes
Qingwen Yang, Sai Xu, Xuejing Li, Yanyi Liu, Tiezheng Guo, Yingyou Wen |
Neurocomputing | 10 |
| 2025 | NeutronHeter: Optimizing Distributed Graph Neural Network Training for Heterogeneous Clusters
Chunyu Cao, Xin Ai 0006, Qiange Wang, Yanfeng Zhang 0001, Zhenbo Fu, Mingyi Cao, Chaoyi Chen, Yingyou Wen, Yu Gu 0002, Ge Yu 0001 |
Proc. ACM Manag. Data | 9 |
| 2025 | DepCache: A KV Cache Management Framework for GraphRAG with Dependency AttentionabstractGraph-based Retrieval-Augmented Generation (GraphRAG) has emerged as a promising paradigm for enhancing LLM reliability by enabling multi-hop reasoning over graph-structured knowledge. However, existing LLMs struggle to efficiently process graph-structured inputs, as traditional attention mechanisms are sequence-based and introduce significant redundancy when serializing graphs into prompt sequences, leading to excessive computation and memory overhead. To address this, we introduce dependency attention, a novel graph-aware attention mechanism that restricts attention computation to token pairs with structural dependencies in the retrieved subgraph. Unlike standard self-attention that computes fully connected interactions, dependency attention prunes irrelevant token pairs and reuses computations along shared relational paths, substantially reducing inference overhead. Building on this idea, we develop DepCache, a KV cache management framework tailored for dependency attention. DepCache enables efficient KV cache reuse through (i) a graph-based KV cache reuse strategy that aligns KV caches across varying prompt contexts, enabling efficient cross-request reuse in GraphRAG, and (ii) a locality-aware replacement policy that leverages spatial and temporal access patterns to improve KV cache hit rate. Evaluations across diverse models and datasets show that DepCache improves LLM inference throughput by 1.5×-5.0× and reduces time-to-first-token latency by up to 3.2×, without compromising generation accuracy. Xin Ai 0006, Qiange Wang, Peizheng Li, Jiayang Yu, Chaoyi Chen, Xinbo Yang, Yanfeng Zhang 0001, Zhenbo Fu, Yingyou Wen, Ge Yu 0001 |
Proc. ACM Manag. Data | 10 |
| 2025 | NeutronTask: Scalable and Efficient Multi-GPU GNN Training with Task ParallelismabstractGraph neural networks (GNNs) have emerged as a promising method for learning from graph data, but large-scale GNN training requires extensive memory and computation resources. To address this, researchers have proposed using multi-GPU processing, which partitions graph data across GPUs for parallel training. However, vertex dependencies in multi-GPU GNN training lead to significant neighbor replications across GPUs, increasing memory consumption. The substantial intermediate data generated during training further exacerbates this issue. Neighbor replication and intermediate data constitute the primary memory consumption in GNN training (i.e., typically accounting for over 80%). In this work, we propose GNN task parallelism for multi-GPU GNN training, which reduces neighbor replication by partitioning training tasks in each layer across different GPUs rather than partitioning the graph structure. This approach only partitions the graph data within individual GPUs, reducing the memory requirements of single tasks while overlapping subgraph computation across different GPUs. Shared neighbor embeddings among different subgraphs can be efficiently reused within a single GPU. Additionally, we employ a task-decoupled GNN training framework, which decouples different training tasks to manage their associated intermediate data independently and release it as early as possible to reduce memory usage. By integrating these techniques, we propose a multi-GPU GNN training system, NeutronTask. Experimental results on a 4×A5000 GPU server show that NeutronTask effectively supports billion-scale full-graph GNN training. For small graphs where the training data fits into the GPUs, NeutronTask achieves 1.27× - 5.47× speedup compared to state-of-the-art GNN systems including NeutronStar and Sancus. Zhenbo Fu, Xin Ai 0006, Qiange Wang, Yanfeng Zhang 0001, Shizhan Lu, Chaoyi Chen, Chunyu Cao, Zhewei Wei, Yu Gu 0002, Yingyou Wen, Ge Yu 0001 |
Proc. VLDB Endow. | 11 |
| 2022 | Multi-context unsupervised domain adaption for HEp-2 cell classification using maximum partial classifier discrepancy
Haoran Zhao 0001, Tao Ren 0002, Xiaotao Yang, Yingyou Wen |
J. Supercomput. | 5 |
| 2019 | Optimal Utility of Vehicles in LTE-V Scenario: An Immune Clone-Based Spectrum Allocation ApproachabstractWith the surge service requirements from vehicular users, especially for automated driving, providing real-time high-rate wireless connections to fast-moving vehicles is ever demanding. This motivates the development of the emerging LTE-V network, a 5G cellular-based vehicular technology. However, note that the vehicular user group is typically of a very large scale, whereas the bandwidth spectrum available for vehicular communications is very limited. To efficiently allocate and utilize the slim spectrum resource to vehicle users are therefore important. This paper develops a service priority oriented spectrum allocation scheme in an LTE-V network, which explores the features of vehicular networks toward economic yet QoS guaranteed spectrum allocation. Specifically, the work exploits two features of the vehicular networks. First, vehicles in the proximity typically have similar information requirements, e.g., road conditions. As a result, the location-based multicast (i.e., geocast) could be applied to save the spectrum. Second, different types of vehicles, e.g., ambulances, buses, and private cars, are of different bandwidth and service requirements. Therefore, differential services and spectrum allocations should be applied. By jointly considering the above features, we develop a 2-D service importance oriented framework for LTE-V network spectrum allocations. The spectrum allocation issue is finally modeled as a mixed integer programming problem to maximize the system utility, and solved using an immune clonal based algorithm. The convergence of the proposed algorithm is proved, and using numerical results, we show that our proposal can outperform the typical heuristics-based spectrum resource allocation in terms of convergence and average delay. Quyuan Luo, Changle Li, Tom H. Luan, Yingyou Wen |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2017 | 3D visual correlation model for wireless visual sensor networksabstractWireless visual sensor networks comprise a large number of camera-equipped sensor devices and obtain visual information from field of interest. In a Wireless visual sensor network, there exist visual correlation characteristics among images observed by cameras with overlapped field of views. To describe those characteristics, the conventional method is based on image processing. However, it is too complex to be applied to cameras in resource-constrained wireless visual sensor networks. In this paper, based on 3D sensing model and coordinate transformation theory, a novel 3D visual correlation model is designed to exploit the correlation characteristics among cameras for spatial wireless visual sensor networks. The designed model, of which a visual correlation function was proposed, and then a 3D visual correlation coefficient algorithm is derived. Experimental results demonstrated that the designed model can accurately model the visual correlation characteristics and the proposed 3D visual correlation coefficient algorithm outperforms the state-of-the-art algorithms. Xiaotao Yang, Yingyou Wen, Mingyang Zhang 0009 |
ICIS | 2 |
| 2013 | Design of an OSGi-Based WSN GatewayabstractA common application scenario for WSNs is that the sensor nodes sense the physical environment, process the sensed data and transmit them to remote user periodically. However, in most scenarios, client user is not interested in all the sensed data pushed by the gateway. Too much unnecessary sensed data may make a burden on the communication. In this paper, based on the OSGi (Open Service Gateway Initiative) platform, we propose an extensible and configurable WSN gateway with Data Process Engine. The sensed data can be examined, analyzed and filtered within the Data Process Engine, according to the queries that deployed by the client user. Furthermore, a gateway access method via web browser using XMLSocket is introduced in this paper. Finally we implement and deploy the gateway on the TI OMAP 3530 platform and an implementation scenario has been conducted to confirm the feasibility of this solution in an indoor environment. Dazhe Zhao, Yingyou Wen, Yongzhong Mu |
MSN | 3 |
| 2013 | A Determination Algorithm for Probability Density about Uncertain Data Streams Based on GMMabstractIn research of uncertain data stream in sensor network, the probability density is usually used to describe the uncertainty of continuous uncertain object values. Current researches are mainly based on the simply assumption that the values of uncertain object meets some conventional distribution, such as Gaussian distribution. However, the statistical distribution of sensor data streams cannot be described accurately in many cases. In this paper, we propose an algorithm for probability density fitting about continuous uncertain sensor data streams, which based on GMM. Experiments show that this algorithm can meet the time requirement of actual sensing application and improve the fitting effect on the probability density of the value uncertain object. Yingyou Wen, Shaopeng Wang, Tongjie Zhang |
MSN | 1 |
| 2011 | The Four Corners DV-Hop Localization Algorithm for Wireless Sensor NetworkabstractLocalization is crucial for wireless sensor networks. In the original DV-Hop algorithm, beacon nodes need to broadcast messages to compute minimum hop-count between each node pair, and beacon nodes also need to broadcast messages to make each unknown node get the average one-hop distance. Twice broadcasts cause a lot of communication. The communication causes a lot of energy consumption. In order to address this problem, an improved algorithm, named Four Corners DV-Hop algorithm, is proposed. The Four Corners DV-Hop algorithm utilizes the beacon placement to divide the whole area into some small regions. We analyze the communication model and set a threshold to make a message just broadcasted in a small region instead of in the whole area to reduce communication. Compared to the original DV-Hop algorithm, the Four Corners DV-Hop localization algorithm reduces a lot of communication and is more accurate through the simulation results. Yinghui Meng, Yingyou Wen |
TrustCom | 3 |
| 2008 | SPIT Detection and Prevention Method in VoIP EnvironmentabstractIn order to resolve the SPIT (Spam over Internet telephony) security risk in the ALL-IP convergence network, a detection and prevention method based on feedback judgment is proposed in this paper. Introducing the participation of the users and combining the mechanism of trust and reputation can make the direct and indirect use of feedback simple effective and lossless. Improved inference algorithm embodies the distribution characteristic of SPIT behavior and reflects the weight relation of factors which influence the result. In addition, the incremental learning algorithm gives the Real-time trust and reputation. The algorithm dynamically integrates the trust with the reputation and makes a Comprehensive Evaluation of the SPIT. Experiment and analysis results show the better accuracy and sensitivity of the method, indicate the efficiency of the detection and prevention. Guangyu He, Yingyou Wen |
ARES | 2 |
| 2007 | Formal Analysis of Secure Bootstrap in Trusted Computing
Yingyou Wen |
ATC | 2 |
| 2004 | Optimizing Sensor Node Distribution with Genetic Algorithm in Wireless Sensor Network
Yingyou Wen, Ruiqiang Shang |
ISNN (2) | 2 |