Yaru Zhao 0001

dblp:233/4833-1 · DBLP profile ↗
← Back
10ranked-venue papers
6as first author
10since 2021 · last 2026
0000-0002-2151-5420ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 3 first-author · 5 since 2021Computer networks · 3 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 3 · 3 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Cost-sensitive fuzzy rough learning for multi-class imbalanced data classification
Yuan-Ting Yan, Yaru Zhao 0001, Peng Zhou 0008, Shu Zhao 0005
Int. J. Approx. Reason.3
2026 MBNAD: A dual-role memory bridge network for multivariate time series anomaly detection
Yaru Zhao 0001, Peiheng Li, Pietro Liò, Pan Hui 0001
Neurocomputing3
2026 Position Encoding-Enhanced Adaptive Expansion Graph Attention Network for Anomaly Detection in Multivariate Time Series
abstract
As cyber–physical systems continue to increase in complexity, multivariate time series exhibit not only intricate temporal patterns within individual variables but also complex inter-variable dependencies, including both synchronous and asynchronous propagation. Although existing Graph Neural Network (GNN)-based methods perform well in modeling variable dependencies, they often decouple temporal pattern processing from the graph construction. This separation often results in graph structures that lack temporal semantics and struggle to capture delayed responses among variables. To address this, we propose a Position Encoding–Enhanced Adaptive Expansion Graph Attention Network (PEE-AEGAT) for multivariate time series anomaly detection and root cause analysis. First, to embed temporal semantics directly into the graph construction process, we design a Multi-Frequency Positional Encoding Projection (MF-PEP) module that explicitly injects multi-scale periodic patterns and positional information into node embeddings. Second, we introduce a Hybrid Adaptive Graph Structure Learning (HAGSL) strategy that constructs a dual-mode graph by integrating static feature similarity with dynamic delay-aware similarity, together with a density-aware adaptive expansion mechanism for flexibly capturing dependencies. Finally, we develop a delay-aware graph attention network for information aggregation and propose an influence-based scoring mechanism to accurately identify anomaly root causes. Extensive experiments on four public real-world datasets demonstrate that PEE-AEGAT significantly outperforms state-of-the-art baselines in anomaly detection accuracy and root cause localization performance.
Yaru Zhao 0001, Pietro Liò, Pan Hui 0001
IEEE Internet Things J.3
2026 Intent-Driven Cognitive XR Networks: Multi-Agent Orchestration for Immersive Communication
abstract
Extended Reality (XR) is emerging as a central use case in the 6G era, requiring intelligent, adaptive, and low-latency communication to support immersive user experiences. However, existing network architectures remain reactive and lack the capability to interpret or act upon users’ fine-grained multimodal intents, limiting their responsiveness and efficiency. To address this gap, this paper presents theCognitive XR Network (CXN)framework, an intent-driven architecture that integrates perception, reasoning, and control through a hierarchy of Artificial Intelligence (AI) agents. CXN comprises three cooperating agents: the User Intent Agent, which infers and predicts user intents from multimodal sensory inputs; the Network Orchestration Agent, which performs global coordination through multi-agent reinforcement learning; and distributed Resource Management Agents which execute localized decisions under strategic guidance. Together, these agents form a closed cognitive loop that enables proactive, intent-aware orchestration across radio, compute, and caching domains. Simulations under realistic XR collaboration scenarios demonstrate that CXN sustains an intent satisfaction rate above 85% and an average interaction latency below 60 ms with 90 concurrent users, outperforming both reactive and centralized learning baselines.
Yakun Huang, Yaru Zhao 0001, Zhenguo Chen, Jing Lv, Xiuquan Qiao
IEEE J. Sel. Areas Commun.2
2026 EcoPath: Energy-Efficient Multi-Path Data Aggregation for Ubiquitous Connectivity Services
abstract
Ubiquitous connectivity is a key 6G usage scenario, in which large-scale sensing systems deployed in remote and underserved regions must deliver heterogeneous sensing data under stringent energy budgets and deadline constraints. This paper presents EcoPath, a two-tier data aggregation framework for clustered large-scale sensor networks. EcoPath separates low-power intra-cluster collection from a high-rate multi-interface backhaul operated by cluster heads, where Multipath QUIC (MPQUIC) can be practically deployed to exploit path diversity. At the cluster head, EcoPath jointly integrates (i) a deadline-aware bundling controller that aggregates sensor frames into MTU-bounded bundles to amortize protocol overhead while bounding additional waiting time, and (ii) a robust multi-path scheduler that prioritizes packets using Weighted Earliest- Deadline-First (W-EDF) with fairness protection and selects backhaul paths via a stability-aware quality metric with hysteresis to avoid flapping under time-varying links. We further formulate an explicit energy–timeliness optimization and show how its outputs parameterize the online bundling and scheduling policies. Extensive simulations with realistic wireless effects, together with baselines and ablations, demonstrate that EcoPath improves energy efficiency and deadline satisfaction for large-scale aggregation.
Yaru Zhao 0001, Yuan-Ting Yan, Man He, Yuanwei Zhu, Yi Yue 0001, Yakun Huang
IEEE Trans. Netw. Serv. Manag.1
2025 P2S-XR: Predictive and Scalable Scheduling for Concurrent Multi-User XR Services
Yaru Zhao 0001, Shou-lu Hou, Binyang Li, Yakun Huang
IEEE Big Data1
2024 SemDA: Communication-Efficient Data Aggregation Through Distributed Semantic Transmission
abstract
This paper introduces SemDA, a communication-efficient data aggregation method that uses distributed semantic communication for improved transmission and analysis. SemDA utilizes an end-to-end trainable network structure that reduces data transmission volume and deepens semantic feature aggregation. Key advances include an attention-based aggregation method for holistic semantic feature integration and a dual-attention decoding network that emphasizes viewpoint and content dimensions. Performance evaluations on CIFAR-10 and ImageNet datasets show that SemDA offers significant improvements in accuracy and system overhead compared to traditional and distributed semantic communication methods. Notable contributions include the proposal of a novel decoding structure, the introduction of a dual-attention decoding mechanism, and extensive evaluations against benchmark methods.
Yaru Zhao 0001, Yakun Huang
ICASSP1
2024 KiProL: A Knowledge-Injected Prompt Learning Framework for Language Generation
Yaru Zhao 0001, Yakun Huang, Bo Cheng 0001
PAKDD (6)1
2024 FluGCF: A Fluent Dialogue Generation Model With Coherent Concept Entity Flow
abstract
The integration of external knowledge graphs into dialogue systems effectively mitigates the generation of generic and uninteresting responses. This approach, particularly the explicit modeling of conversation flows from related concept entities, facilitates the generation of semantically rich and informative responses. However, recent models guided by concept entity flows present two primary limitations: (1) a limited semantic understanding of the post message, which complicates the selection of highly relevant 1-hop concept entities, and (2) an inability to extract dynamic and diverse semantic relations between the post message and 2-hop concept entities. To address these issues, we introduce FluGCF, a novel model that fluently generates dialogues with coherent guidance from concept entity flows. FluGCF employs a ternary fusion to explicitly model multi-hop concept entity flows using a post-aware knowledge encoding mechanism. This mechanism learns semantic concept entity features from both word and sentence-level text features. Additionally, we design a corresponding ternary decoding mechanism that dynamically selects concept entities or words from the vocabulary to enhance fluency and diversity in dialogue generation. FluGCF, implemented in PyTorch, was extensively evaluated on a large-scale dataset, revealing that it surpasses baseline models, including the state-of-the-art knowledge-aware model ConceptFlow, by nearly 15% in terms of fluency. Furthermore, it demonstrated notable enhancements in coherence, diversity and informativeness.
Yaru Zhao 0001, Bo Cheng 0001, Yakun Huang, Zhiguo Wan
IEEE ACM Trans. Audio Speech Lang. Process.1
2023 Beyond Words: An Intelligent Human-Machine Dialogue System with Multimodal Generation and Emotional Comprehension
abstract
Intelligent service robots have become an indispensable aspect of modern‐day society, playing a crucial role in various domains ranging from healthcare to hospitality. Among these robotic systems, human‐machine dialogue systems are particularly noteworthy as they deliver both auditory and visual services to users, effectively bridging the communication gap between humans and machines. Despite their utility, the majority of existing approaches to these systems primarily concentrate on augmenting the logical coherence of the system’s responses, inadvertently neglecting the significance of user emotions in shaping a comprehensive communication experience. To tackle this shortcoming, we propose the development of an innovative human‐machine dialogue system that is both intelligent and emotionally sensitive, employing multimodal generation techniques. This system is architecturally comprised of three components: (1) data collection and processing, responsible for gathering and preparing relevant information, (2) a dialogue engine, which generates contextually appropriate responses, and (3) an interaction module, responsible for facilitating the communication interface between users and the system. To validate our proposed approach, we have constructed a prototype system and conducted an evaluation of the performance of the core dialogue engine by utilizing an open dataset. The results of our study indicate that our system demonstrates a remarkable level of multimodal generation response, ultimately offering a more human‐like dialogue experience.
Yaru Zhao 0001, Bo Cheng 0001, Yakun Huang, Zhiguo Wan
Int. J. Intell. Syst.1