EDBT 2026 Demo / reviewers in the wild / expert
Yuan Yuan 0034
dblp:64/5845-34
· DBLP profile ↗
22ranked-venue papers
3as first author
15since 2021 · last 2026
0009-0004-5072-6093ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 10 · 1 first-author · 7 since 2021Software engineering, systems software and programming languages · 4 · 1 first-author · 4 since 2021Artificial intelligence and machine learning · 2 · 1 since 2021Computer networks · 2 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2Databases, data management, data science and information retrieval · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | CML-PowF: Data Clustering Matching Based Low-overhead Multiple CPU Real-time Power ForecastingabstractEfficient CPU power capping is essential for energy saving and fault tolerance in parallel computing clusters, but its effectiveness depends on accurate and timely processor power forecasting with minimal sampling overhead. Existing methods often struggle to balance these factors under scalability constraints, as hardware limitations tightly bound the available sampling resources. This article focuses on the issue of high-precision real-time processor power forecasting while maintaining (or minimally increasing) the total overhead of multiprocessor power forecasting, particularly when the parallelism scale ranges from P to 2P processors or when the problem size scales from M to 2M . We propose CML-PowF , a low-overhead multiprocessor real-time power forecasting approach based on data clustering. CML-PowF integrates two key algorithms: Alg-CEF , which conducts cluster matching on the runtime characteristics of the program at the P / M scale, and models the tradeoff among forecasting error, time span, and sampling overhead. Alg-MSF , which leverages execution patterns from smaller-scale runs to determine the optimal sampling overhead and forecasting time span at the 2P / 2M scale. We evaluate CML-PowF on x86 and ARM platforms with up to 32 computing nodes (2,048 cores). Results show that it achieves 3–6% forecasting error at large scales with only 0.2–0.5% degradation compared to the P / M scale, without increasing total sampling overhead. Integrated with the PowC control system, CML-PowF effectively maintains real-time processor power below target thresholds. Rongyu Deng, Juan Chen 0001, Yuan Yuan 0034, Yong Dong, Aolin Cao, Yida Gu, Dingwen Tao |
ACM Trans. Archit. Code Optim. | 4 |
| 2025 | Efficient and Accurate Anomaly Detection in HPC Systems via Coarse-Grained Clustering and Fine-Grained Model Sharing
Shaoyu Hu, Sibo Xia, Yongqian Sun, Xijie Pan, Yuan Yuan 0034, Shenglin Zhang |
APNet | 5 |
| 2025 | SPEA: Large-Scale Entity Alignment via Self-PartitioningabstractThe task of entity alignment (EA) seeks to identify corresponding entities across different knowledge graphs (KGs). However, in large-scale KG alignment tasks, the complexity of the problem renders traditional entity structure representation methods, designed for small-scale KGs, ineffective. Partition-based approaches address this challenge by breaking large KGs into smaller subgraphs, but this inevitably results in a loss of structural information. Although existing methods have sought to mitigate this issue, they have largely overlooked the interplay between partitioning and entity structure representation learning. To address this, we propose a Self-Partitioning Entity Alignment (SPEA) pipeline for large-scale EA, in which partitioning and entity structure representation learning are mutually optimized. Within this framework, we introduce the Inter-Subgraph Neighbor Interaction (ISNI) for enhanced entity structure representation, the Bidirectional Margin-based Confidence (BMC) for pseudo-pairing, the Seed-oriented Cross Graph Partitioner (SCGP) for dynamic repartitioning, and the Historical Confidence Ensembling (HCE) strategy for consistent training. Extensive experiments demonstrate that SPEA significantly outperforms existing methods for large-scale EA tasks. Kele Xu, Yuan Yuan 0034, Zixuan Dong |
ICASSP | 4 |
| 2025 | To Split or to Merge? Exploring Multi-modal Data Flexibly for Failure Classification in MicroservicesabstractClassifying failures for automated diagnosis is critical for maintaining the reliability of ever-increasing microservice systems.Existing failure classification methods focus solely on analyzing one specific data modality (e.g., metrics) or feeding multi-modal data (e.g., logs and traces) in a one-fit-all model.This would lead to biased inference as different modalities have differentiated proportions and manifest distinct sensitivities and patterns in failures of varied runtime contexts.This work proposes a novel failure classification framework for microservices, named AgileMFC, with the basic idea of flexibly learning and assembling the expertise of multi-modal monitoring data.The flexibility is attained with the split-and-merge design of data representation, where it assigns each modality a different gate that routes it to dedicated expert networks and uses a sharing gate to facilitate controllable knowledge interaction among modalities.Meanwhile, AgileMFC trains the representation networks and a classifier in consecutive stages with concatenated and weighted feature fusion, thus decoupling task-agnostic feature extraction with task-specific classification.Experiments on three datasets show that AgileMFC outperforms four baselines w.r.t., precision, recall, and F1-score on failure classification. Xiuhong Tan, Yuan Yuan 0034, Tongqing Zhou, Shiming He |
Internetware | 2 |
| 2025 | ClusterRCA: An End-to-End Approach for Network Fault Localization and Classification for HPC SystemabstractNetwork failure diagnosis is challenging yet critical for high-performance computing (HPC) systems. Existing methods cannot be directly applied to HPC scenarios due to data heterogeneity and lack of accuracy. This paper proposes a novel framework, called ClusterRCA, to localize culprit nodes and determine failure types by leveraging multimodal data. ClusterRCA extracts features from topologically connected network interface controller (NIC) pairs to analyze the diverse, multimodal data in HPC systems. To accurately localize culprit nodes and determine failure types, ClusterRCA combines classifier-based and graph-based approaches. A failure graph is constructed based on the output of the state classifier, and then it performs a customized random walk on the graph to localize the root cause. Experiments on datasets collected by a top-tier global HPC device vendor show ClusterRCA achieves high accuracy in diagnosing network failure for HPC systems. ClusterRCA also maintains robust performance across different application scenarios. Yongqian Sun, Xijie Pan, Xiao Xiong, Jiaju Wang, Shenglin Zhang, Yuan Yuan 0034, Kunlin Jian |
ISSRE | 7 |
| 2025 | TrioXpert: An Automated Incident Management Framework for Microservice SystemabstractAutomated incident management plays a pivotal role in large-scale microservice systems. However, many existing methods rely solely on single-modal data (e.g., metrics, logs, and traces) and struggle to simultaneously address multiple downstream tasks, including anomaly detection (AD), failure triage (FT), and root cause localization (RCL). Moreover, the lack of clear reasoning evidence in current techniques often leads to insufficient interpretability. To address these limitations, we propose TrioXpert, an end-to-end incident management framework capable of fully leveraging multimodal data. TrioXpert designs three independent data processing pipelines based on the inherent characteristics of different modalities, comprehensively characterizing the operational status of microservice systems from both numerical and textual dimensions. It employs a collaborative reasoning mechanism using large language models (LLMs) to simultaneously handle multiple tasks while providing clear reasoning evidence to ensure strong interpretability. We conducted extensive evaluations on two microservice system datasets, and the experimental results demonstrate that TrioXpert achieves outstanding performance in AD (improving by 4.7% to 57.7%), FT (improving by 2.1% to 40.6%), and RCL (improving by 1.6% to 163.1%) tasks. TrioXpert has also been deployed in Lenovo’s production environment, demonstrating substantial gains in diagnostic efficiency and accuracy. Yongqian Sun, Yu Luo 0011, Xidao Wen, Yuan Yuan 0034, Xiaohui Nie, Shenglin Zhang |
ASE | 4 |
| 2025 | Effective Node-Level Anomaly Detection in HPC Systems via Coarse-Grained Clustering and Fine-Grained Model SharingabstractHigh-performance computing (HPC) systems are crucial for scientific advancement and engineering breakthroughs. Unexpected performance degradation or system failures can severely impact these endeavors. This paper introduces NodeSentry, a novel unsupervised anomaly detection framework tailored for compute nodes of large-scale HPC systems. NodeSentry leverages a combined approach of coarse-grained clustering and fine-grained model sharing to effectively address the challenges posed by the massive node scales, frequent job transitions, and complex patterns characteristic of modern HPC deployments. Evaluation on two real-world HPC datasets demonstrates NodeSentry’s superior performance, achieving an F1-score exceeding 0.876. This represents a 0.560 average improvement over existing best baseline methods, while simultaneously reducing training overhead by an average of 45.69%. Furthermore, to promote reproducibility and contribute to the broader research community, we open-source NodeSentry’s codebase and introduce a novel clustering adjustment and anomaly labeling tool specifically designed for HPC systems. Sibo Xia, Yongqian Sun, Xijie Pan, Yuan Yuan 0034, Shenglin Zhang, Shaoyu Hu, Jinghua Feng |
SC | 4 |
| 2024 | Exploring Hierarchical Patterns for Alert Aggregation in SupercomputersabstractAlongside the high performance built on massive hardware, ever-larger computer systems bear tons of hardware alerts every day during reliability maintenance. Based on an exploratory study on a representative supercomputer system, this work first characterizes supercomputer alerts as an overload of continuous bursts for the operators. Yet, existing similarity-based aggregation solutions, tuned for in-band textual alerts, are myopic by finding dissimilar representatives instead of looking into the semantics in the supercomputer context. To fill the void of supercomputer alert aggregation, we propose the SuperAgg framework to extract the hierarchical patterns of real-world alerts and use them for online alert management. SuperAgg jointly integrates unsupervised state detection of time series and expert analysis to successfully discover 4 categories of sensor-tier alert patterns and exploits primary-and-secondary statistics between sensors for system-tier correlation patterns. With such extracted knowledge, SuperAgg then identifies the formulated patterns online and uses spatiotemporal combined strategies to reduce the alert influx. Evaluations on alerts generated from a production supercomputer show that SuperAgg provides over 98% aggregation rate and significantly higher aggregation accuracy (over 83.8% and 43.2% on different datasets) than 3 baselines. Production deployment further demonstrates its effectiveness from the perspective of system operators. The source code is available at: https://github.com/Txh-User/SuperAgg. Yuan Yuan 0034, Tongqing Zhou, Xiuhong Tan, Yongqian Sun, Zhiping Cai |
ISSRE | 1 |
| 2023 | Optimizing the Parallelism of Communication and Computation in Distributed Training Platform
Xiang Hou, Yuan Yuan 0034, Sheng Ma, Lizhou Wu |
ICA3PP (1) | 2 |
| 2023 | PowerDis: Fine-Grained Power Monitoring Through Power Disaggregation Model
Xinxin Qi, Juan Chen 0001, Rongyu Deng, Yuan Yuan 0034, Yonggang Che |
ICA3PP (4) | 6 |
| 2023 | HighRPM: Combining Integrated Measurement and Sofware Power Modeling for High-Resolution Power MonitoringabstractIn an era where power and energy are the first-class constraints of computing systems, accurate power information is crucial for energy efficiency optimization in parallel computing systems. Existing power monitoring techniques rely on either software-centric power models that suffer from poor accuracy or integrated hardware measurement schemes that have a low reading update frequency and coarse granularity. These result in a low spatiotemporal resolution for power monitoring. This paper introduces HighRPM, a new method for accurately measuring power consumption on parallel computing systems. HighRPM combines coarse-grained power sensor readings and software power modeling techniques to improve temporal and spatial resolutions. To provide high-frequent power readings in the temporal domain, HighRPM employs statistical modeling and machine learning techniques to predict the long-term power trend and the short-term fluctuations in power consumption. To improve spatial coverage, HighRPM takes low-time resolution node-level power consumption and uses a neural network to distribute the power readings to lower-level computing components like CPUs and memory components. We evaluate HighRPM by applying it to both ARM-based and X86-based platforms. Experimental results show that HighRPM improves time resolution by 10 times, provides accurate readings for CPUs and memory, and reduces error by 7-24% compared to other power modeling methods. Xinxin Qi, Juan Chen 0001, Yong Dong, Yuan Yuan 0034, Tao Xu 0052, Rongyu Deng, Kexing Zhou, Zheng Wang 0001 |
ICPP | 4 |
| 2023 | Turning backdoors for efficient privacy protection against image retrieval violations
Qiang Liu 0004, Tongqing Zhou, Zhiping Cai, Yuan Yuan 0034, Ming Xu 0002, Jiaohua Qin, Wentao Ma 0003 |
Inf. Process. Manag. | 4 |
| 2022 | SparG: A Sparse GEMM Accelerator for Deep Learning Applications
Sheng Ma, Yuan Yuan 0034, Xiang Hou, Xiao Yi |
ICA3PP | 3 |
| 2022 | SADD: A Novel Systolic Array Accelerator with Dynamic Dataflow for Sparse GEMM in Deep Learning
Sheng Ma, Zhong Liu 0003, Libo Huang 0002, Yuan Yuan 0034 |
NPC | 5 |
| 2021 | USET : A network based on Utterance hidden State transfEr for Task-oriented dialogueabstractMulti-turn dialogue is challenging because semantic information is not only contained in the current utterance, but also in the dialogue context. In fact, understanding multiturn dialogue is a dynamic process. With the increase of dialogue turn, users' understanding is also changing. In this case, we propose a network based on utterance hidden state transfer for task-oriented dialogue (USET). In our model, we first extract the hidden state of previous utterance as the previous comprehension. Then, this comprehension is passed to the next turn. We take the comprehension as a prior knowledge to understand the semantic information in the dialogue context. Finally, we put the previous comprehension and current utterance together to understand current utterance. In order to realize the transfer of comprehension in dialogue, we propose a continuous sample training method: CST, which takes a multi-turn dialogue as a whole to understand. All the sentences in a dialogue are put into a batch for training. Our method makes use of previous comprehension and achieves the information exchange among dialogue. Experimental results on Stanford Multi-Domain dataset demonstrate that our model is superior to existing models. Code is available at https://github.com/season1blue/USET Shezheng Song, Dongsong Zhang, Yuxing Peng 0001, Yuan Yuan 0034 |
IJCNN | 5 |
| 2020 | PMC-Based Dynamic Adaptive CPU and DRAM Power Modeling
Yunfang Zhang, Yong Dong, Juan Chen 0001, Zhixin Ou, Yuan Yuan 0034 |
ICA3PP (1) | 5 |
| 2020 | A Distributed Computing Framework Based on Variance Reduction Method to Accelerate Training Machine Learning ModelsabstractTo support large-scale intelligent applications, distributed machine learning based on JointCloud is an intuitive solution scheme. However, the distributed machine learning is difficult to train due to that the corresponding optimization solver algorithms converge slowly, which highly demand on computing and memory resources. To overcome the challenges, we propose a computing framework for L-BFGS optimization algorithm based on variance reduction method, which can utilize a fixed big learning rate to linearly accelerate the convergence speed. To validate our claims, we have conducted several experiments on multiple classical datasets. Experimental results show that the computing framework accelerate the training process of solver and obtain accurate results for machine learning algorithms. Zhen Huang 0006, Mingxing Tang, Jinyan Qiu, Hangjun Zhou, Yuan Yuan 0034, Dongsheng Li 0001, Yuxing Peng 0001 |
JCC | 6 |
| 2019 | TFPN: Twin Feature Pyramid Networks for Object DetectionabstractFPN (Feature Pyramid Networks) is one of the most popular object detection networks, which can improve small object detection by enhancing shallow features. However, limited attention has been paid to the improvement of large object detection via deeper feature enhancement. One existing approach merges the feature maps of different layers into a new feature map for object detection, but can lead to increased noise and loss of information. The other approach adds a bottom-up structure after the feature pyramid of FPN, which superimposes the information from shallow layers into the deep feature map but weakens the strength of FPN in detecting small objects. To address these challenges, this paper proposes TFPN (Twin Feature Pyramid Networks), which consists of (1) FPN+, a bottom-up structure that improves large object detection; (2) TPS, a Twin Pyramid Structure that improves medium object detection; and (3) innovative integration of these two with FPN, which can significantly improve the detection accuracy of large and medium objects while maintaining the advantage of FPN in small object detection. Extensive experiments using the MSCOCO object detection datasets and the BDD100K automatic driving dataset demonstrate that TFPN significantly improves over existing models, achieving up to 2.2 improvement in detection accuracy (e.g., 36.3 for FPN vs. 38.5 for TFPN on COCO Val-17). Our method can obtain the same accuracy as FPN with ResNet-101 based on ResNet-50 and needs fewer parameters. Fangzhao Li, Yuxing Peng 0001, Qin Lv, Yuan Yuan 0034, Zhen Huang 0006 |
ICTAI | 6 |
| 2018 | Moving from exascale to zettascale computing: challenges and techniquesabstractHigh-performance computing (HPC) is essential for both traditional and emerging scientific fields, enabling scientific activities to make progress. With the development of high-performance computing, it is foreseeable that exascale computing will be put into practice around 2020. As Moore’s law approaches its limit, high-performance computing will face severe challenges when moving from exascale to zettascale, making the next 10 years after 2020 a vital period to develop key HPC techniques. In this study, we discuss the challenges of enabling zettascale computing with respect to both hardware and software. We then present a perspective of future HPC technology evolution and revolution, leading to our main recommendations in support of zettascale computing in the coming future. Xiangke Liao, Kai Lu 0001, Canqun Yang, Jin-wen Li, Yuan Yuan 0034, Libo Huang 0002, Pingjing Lu, Jianbin Fang, Jie Shen 0003 |
Frontiers Inf. Technol. Electron. Eng. | 5 |
| 2014 | Hybrid hierarchy storage system in MilkyWay-2 supercomputer
Yutong Lu, Enqiang Zhou, Zhenlong Song, Yong Dong, Dengping Wei, Jianying Xing, Yuan Yuan 0034 |
Frontiers Comput. Sci. | 12 |
| 2011 | Performance of Acyclic Stochastic Networks with Network CodingabstractNetwork coding allows a network node to code the information flows before forwarding them. While it has been theoretically proved that network coding can achieve maximum network throughput, the theoretical results usually do not consider the burstiness of data traffic, delays, and the stochastic nature in information processing and transmission. There is currently no theory to systematically model and evaluate the performance of network coding, especially when node's capacity (i.e., coding and transmission) becomes stochastic. Without such a theory, the performance of network coding under various system settings is far from clear. To fill the vacancy, we develop an analytical approach by extending the stochastic network calculus theory to tackle the special difficulties in the evaluation of network coding. We prove the new properties of the stochastic network calculus and design an algorithm to obtain the performance bounds for acyclic stochastic networks with network coding. The tightness of theoretical bounds is validated with simulation. Yuan Yuan 0034, Kui Wu 0001, Weijia Jia 0001, Yuming Jiang 0001 |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2009 | FOCUS: A Cost-Effective Approach for Large-Scale Crop Monitoring with Sensor NetworksabstractCurrent investment in crop monitoring consumes a large amount of financial cost, and how to reduce this cost has been a long-standing problem in agriculture. Traditional crop monitoring approaches are not cost-effective, because they rely on either heavy human labor or intensive computation with expensive instruments. In this paper, we explore the possibility of deploying networked sensor nodes for low-cost crop monitoring. As an example, we compute an important agricultural metric called global leaf area index (LAI) to illustrate the benefit of using sensor networks. We propose an approach called FOCUS that incrementally deploys sensor nodes into farmland to improve the accuracy of global LAI measurements. We design and implement a novel algorithm that calculates the total size of crop leaves with light intensity readings captured by the sensors under the crop canopies. FOCUS not only lowers the deployment cost considerably but also reduces the number of sensors for the long-term monitoring. Through a small-scale field test and large-scale simulations, we validate our design and show its effectiveness in crop monitoring. Yuan Yuan 0034, Shanshan Li 0001, Kui Wu 0001, Weijia Jia 0001, Yuxing Peng 0001 |
MASS | 1 |