VLDB 2026 Research / reviewers in the wild / expert
Xinye Cao
dblp:322/9121
· DBLP profile ↗
10ranked-venue papers
4as first author
10since 2021 · last 2026
0009-0007-0060-6693ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 4 · 1 first-author · 4 since 2021Computer networks · 3 · 3 first-author · 3 since 2021Systems, architecture and hardware · 1 · 1 since 2021Security and privacy · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
4 papers |
Efficient and distributed learning · 35% Video understanding and tracking · 17% Deep learning architectures and training · 17% | |
| Computer architecture, parallel and distributed computing, and storage systems
3 papers |
Hardware accelerators and domain-specific architectures · 46% Cloud and datacenter computing · 28% Reconfigurable computing and FPGAs · 19% | |
| Computer networks
3 papers |
Network management and operations · 62% Cellular and mobile networks · 29% Edge and fog computing · 9% | |
| Network and information security
3 papers |
Network security · 100% |
Topics — the 22 heaviest of 26, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Network security › intrusion detection and prevention
intrusion detection |
1.3 | 2 | 2026 | Enhancing Cloud Network Resilience via a Robust LLM-Empowered Multi-Agent Reinforcement Learning Framework · IEEE Trans. Dependable Secur. Comput. 2026 Advancing LLM-Based Security Automation With Customized Group Relative Policy Optimization for Zero-Touch Networks · IEEE J. Sel. Areas Commun. 2026 |
Network security › attack resilience › attack mitigation
reinforcement-learning-based defense |
1.0 | 1 | 2026 | Enhancing Cloud Network Resilience via a Robust LLM-Empowered Multi-Agent Reinforcement Learning Framework · IEEE Trans. Dependable Secur. Comput. 2026 |
Cloud and datacenter computing
cloud security |
1.0 | 1 | 2026 | Enhancing Cloud Network Resilience via a Robust LLM-Empowered Multi-Agent Reinforcement Learning Framework · IEEE Trans. Dependable Secur. Comput. 2026 |
Machine learning › Efficient and distributed learning › model compression › sparsity
activation sparsity |
0.9 | 1 | 2025 | Advancing Expert Specialization for Better MoE · NeurIPS 2025 |
Computer vision › Vision and language › multimodal reasoning
compositional reasoning |
0.9 | 1 | 2025 | Advancing Compositional LLM Reasoning With Structured Task Relations in Interactive Multimodal Communications · IEEE J. Sel. Areas Commun. 2025 |
Machine learning › Deep learning architectures and training › mixture of experts
expert specialization |
0.9 | 1 | 2025 | Advancing Expert Specialization for Better MoE · NeurIPS 2025 |
Computer vision › Video understanding and tracking
long video understanding |
0.9 | 1 | 2025 | VideoMiner: Iteratively Grounding Key Frames of Hour-Long Videos via Tree-Based Group Relative Policy Optimization · ICCV 2025 |
Machine learning › Efficient and distributed learning › parameter-efficient fine-tuning
LoRA composition |
0.9 | 1 | 2025 | Advancing Compositional LLM Reasoning With Structured Task Relations in Interactive Multimodal Communications · IEEE J. Sel. Areas Commun. 2025 |
Machine learning › Deep learning architectures and training
mixture of experts |
0.9 | 1 | 2025 | Advancing Expert Specialization for Better MoE · NeurIPS 2025 |
Machine learning › Efficient and distributed learning › large-scale learning
model scaling |
0.9 | 1 | 2025 | Advancing Expert Specialization for Better MoE · NeurIPS 2025 |
Machine learning › Efficient and distributed learning
parameter-efficient fine-tuning |
0.9 | 1 | 2025 | Advancing Compositional LLM Reasoning With Structured Task Relations in Interactive Multimodal Communications · IEEE J. Sel. Areas Commun. 2025 |
Machine learning › Learning paradigms › multi-task learning
task relationship modeling |
0.9 | 1 | 2025 | Advancing Compositional LLM Reasoning With Structured Task Relations in Interactive Multimodal Communications · IEEE J. Sel. Areas Commun. 2025 |
Network management and operations
situation awareness |
0.9 | 1 | 2025 | Exploring LLM-Based Multi-Agent Situation Awareness for Zero-Trust Space-Air-Ground Integrated Network · IEEE J. Sel. Areas Commun. 2025 |
Cellular and mobile networks › 6g
space-air-ground integrated network |
0.9 | 1 | 2025 | Exploring LLM-Based Multi-Agent Situation Awareness for Zero-Trust Space-Air-Ground Integrated Network · IEEE J. Sel. Areas Commun. 2025 |
Computer vision › Image recognition and object detection
object detection |
0.8 | 1 | 2024 | A comprehensive analysis of DAC-SDC FPGA low power object detection challenge · Sci. China Inf. Sci. 2024 |
Hardware accelerators and domain-specific architectures › machine learning accelerator
DNN accelerator |
0.7 | 1 | 2023 | Uint-Packing: Multiply Your DNN Accelerator Performance via Unsigned Integer DSP Packing · DAC 2023 |
Reconfigurable computing and FPGAs › FPGA resource optimization
DSP block utilization |
0.7 | 1 | 2023 | Uint-Packing: Multiply Your DNN Accelerator Performance via Unsigned Integer DSP Packing · DAC 2023 |
Natural language and speech › Language models and text generation
large language model training |
0.3 | 1 | 2025 | Advancing Expert Specialization for Better MoE · NeurIPS 2025 |
Machine learning › Reinforcement learning
policy optimization |
0.3 | 1 | 2025 | VideoMiner: Iteratively Grounding Key Frames of Hour-Long Videos via Tree-Based Group Relative Policy Optimization · ICCV 2025 |
Edge and fog computing
resource-constrained inference |
0.3 | 1 | 2025 | Advancing Compositional LLM Reasoning With Structured Task Relations in Interactive Multimodal Communications · IEEE J. Sel. Areas Commun. 2025 |
Energy-efficient computing
low-power design |
0.2 | 1 | 2024 | A comprehensive analysis of DAC-SDC FPGA low power object detection challenge · Sci. China Inf. Sci. 2024 |
Hardware accelerators and domain-specific architectures › machine learning accelerator
DNN inference accelerator |
0.2 | 1 | 2023 | Uint-Packing: Multiply Your DNN Accelerator Performance via Unsigned Integer DSP Packing · DAC 2023 |
Methods — techniques the papers use, named apart from their topics
large language model · 5.7multi-agent reinforcement learning · 2.0human-in-the-loop · 2.0group relative policy optimization · 2.0reinforcement learning · 1.7multi-agent system · 1.7mixture of experts · 1.7variance loss · 0.9tree-based group relative policy optimization · 0.9orthogonality loss · 0.9multimodal large language model · 0.9load-balancing loss · 0.9hierarchical clustering · 0.9chain-of-thought · 0.9chain of thought · 0.9unsigned integer packing · 0.7MAC operation packing · 0.7
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Advancing LLM-Based Security Automation With Customized Group Relative Policy Optimization for Zero-Touch NetworksabstractZero-Touch Networks (ZTNs) represent a transformative paradigm toward fully automated and intelligent network management, providing the scalability and adaptability required for the complexity of sixth-generation (6G) networks. However, the distributed architecture, high openness, and deep heterogeneity of 6G networks expand the attack surface and pose unprecedented security challenges. To address this, security automation aims to enable intelligent security management across dynamic and complex environments, serving as a key capability for securing 6G ZTNs. Despite its promise, implementing security automation in 6G ZTNs presents two primary challenges: 1) automating the lifecycle from security strategy generation to validation and update under real-world, parallel, and adversarial conditions, and 2) adapting security strategies to evolving threats and dynamic environments. This motivates us to propose SecLoop and SA-GRPO. SecLoop constitutes the first fully automated framework that integrates large language models (LLMs) across the entire lifecycle of security strategy generation, orchestration, response, and feedback, enabling intelligent and adaptive defenses in dynamic network environments, thus tackling the first challenge. Furthermore, we propose SA-GRPO, a novel security-aware group relative policy optimization algorithm that iteratively refines security strategies by contrasting group feedback collected from parallel SecLoop executions, thereby addressing the second challenge. Extensive real-world experiments on five benchmarks, including 11 MITRE ATT&CK processes and over 20 types of attacks, demonstrate the superiority of the proposed SecLoop and SA-GRPO. We will release our platform to the community, facilitating the advancement of security automation towards next generation communications. Xinye Cao, Yihan Lin 0001, Guoshun Nan, Qinchuan Zhou, Yuhang Luo, Yurui Gao, Haolang Lu, Qimei Cui, Yan-Zhao Hou, Xiaofeng Tao 0001, Tony Q. S. Quek |
IEEE J. Sel. Areas Commun. | 1 |
| 2026 | Enhancing Cloud Network Resilience via a Robust LLM-Empowered Multi-Agent Reinforcement Learning FrameworkabstractWhile virtualization and resource pooling empower cloud networks with structural flexibility and elastic scalability, they inevitably expand the attack surface and challenge cyber resilience. Reinforcement Learning (RL)-based defense strategies have been developed to optimize resource deployment and isolation policies under adversarial conditions, aiming to enhance system resilience by maintaining and restoring network availability. However, existing approaches lack robustness as they require retraining to adapt to dynamic changes in network structure, node scale, attack strategies, and attack intensity. Furthermore, the lack of Human-in-the-Loop (HITL) support limits interpretability and flexibility. To address these limitations, we propose CyberOps-Bots, a hierarchical multi agent reinforcement learning framework empowered by Large Language Models (LLMs). Inspired by MITRE ATT&CK's “Tactics-Techniques” model, CyberOps-Bots features a two-layer architecture: (1) An upper-level LLM agent with four mod ules—ReAct planning, IPDRR-based perception, long-short term memory, and action/tool integration—performs global awareness, human intent recognition, and tactical planning; (2) Lower-level RL agents, developed via heterogeneous separated pre-training, execute atomic defense actions within localized network regions. This synergy preserves LLM adaptability and interpretability while ensuring reliable RL execution. Experiments on real cloud datasets show that, compared to state-of-the-art algorithms, CyberOps-Bots maintains network availability 68.5% higher and achieves a 34.7% jumpstart performance gain when shifting the scenarios without retraining. To our knowledge, this is the first study to establish a robust LLM-RL framework with HITL support for cloud defense. Yixiao Peng, Hao Hu 0005, Feiyang Li, Xinye Cao, Yingchang Jiang, Jipeng Tang, Guoshun Nan |
IEEE Trans. Dependable Secur. Comput. | 4 |
| 2025 | VideoMiner: Iteratively Grounding Key Frames of Hour-Long Videos via Tree-Based Group Relative Policy OptimizationabstractUnderstanding hour-long videos with multi-modal large language models (MM-LLMs) enriches the landscape of human-centered AI applications. However, for end-to-end video understanding with LLMs, uniformly sampling video frames results in LLMs being overwhelmed by a vast amount of irrelevant information as video length increases. Existing hierarchical key frame extraction methods improve the accuracy of video understanding but still face two critical challenges. 1) How can the interference of extensive redundant information in long videos be mitigated? 2) How can a model dynamically adapt to complex hierarchical structures while accurately identifying key frames? To address these issues, we propose VideoMiner, which iteratively segments, captions, and clusters long videos, forming a hierarchical tree structure. The proposed VideoMiner progresses from long videos to events to frames while preserving temporal coherence, effectively addressing the first challenge. To precisely locate key frames, we introduce T-GRPO, a tree-based group relative policy optimization in reinforcement learning method that guides the exploration of the VideoMiner. The proposed T-GRPO is specifically designed for tree structures, integrating spatiotemporal information at the event level while being guided by the question, thus solving the second challenge. We achieve superior performance in all long-video understanding tasks and uncover several interesting insights. Our proposed T-GRPO surprisingly incentivizes the model to spontaneously generate a reasoning chain. Additionally, the designed tree growth auxin dynamically adjusts the expansion depth, obtaining accuracy and efficiency gains. The code is publicly available at https://github.com/caoxinye/VideoMiner. Xinye Cao, Hongcan Guo, Jiawen Qian, Guoshun Nan, Yuqi Pan, Tianhao Hou, Yutong Gao 0001 |
ICCV | 1 |
| 2025 | Advancing Expert Specialization for Better MoEabstractMixture-of-Experts (MoE) models enable efficient scaling of large language models (LLMs) by activating only a subset of experts per input.
However, we observe that the commonly used auxiliary load balancing loss often leads to expert overlap and overly uniform routing, which hinders expert specialization and degrades overall performance during post-training.
To address this, we propose a simple yet effective solution that introduces two complementary objectives: (1) an orthogonality loss to encourage experts to process distinct types of tokens, and (2) a variance loss to encourage more discriminative routing decisions.
Gradient-level analysis demonstrates that these objectives are compatible with the existing auxiliary loss and contribute to optimizing the training process.
Experimental results over various model architectures and across multiple benchmarks show that our method significantly enhances expert specialization.
Notably, our method improves classic MoE baselines with auxiliary loss by up to 23.79\%, while also maintaining load balancing in downstream tasks, without any architectural modifications or additional components. We will release our code to contribute to the community. Hongcan Guo, Haolang Lu, Guoshun Nan, Bolun Chu, Jialin Zhuang, Wenhao Che, Xinye Cao, Sicong Leng, Qimei Cui |
NeurIPS | 8 |
| 2025 | PCSViT: Efficient and hardware friendly Pyramid Vision Transformer with channel and spatial self-attentions
Xiaofeng Zou, Yuanxi Peng, Xinye Cao |
Neurocomputing | 4 |
| 2025 | Advancing Compositional LLM Reasoning With Structured Task Relations in Interactive Multimodal CommunicationsabstractInteractive multimodal applications (IMAs), such as route planning in the Internet of Vehicles, enrich users’ personalized experiences by integrating various forms of data over wireless networks. Recent advances in large language models (LLMs) utilize mixture-of-experts (MoE) mechanisms to empower multiple IMAs, with each LLM trained individually for a specific task that presents different business workflows. In contrast to existing approaches that rely on multiple LLMs for IMAs, this paper presents a novel paradigm that accomplishes various IMAs using a single compositional LLM over wireless networks. The two primary challenges include 1) guiding a single LLM to adapt to diverse IMA objectives and 2) ensuring the flexibility and efficiency of the LLM in resource-constrained mobile environments. To tackle the first challenge, we propose ContextLoRA, a novel method that guides an LLM to learn the rich structured context among IMAs by constructing a task dependency graph. We partition the learnable parameter matrix of neural layers for each IMA to facilitate LLM composition. Then, we develop a step-by-step fine-tuning procedure guided by task relations, including training, freezing, and masking phases. This allows the LLM to learn to reason among tasks for better adaptation, capturing the latent dependencies between tasks. For the second challenge, we introduce ContextGear, a scheduling strategy to optimize the training procedure of ContextLoRA, aiming to minimize computational and communication costs through a strategic grouping mechanism. Experiments on three benchmarks show the superiority of the proposed ContextLoRA and ContextGear. Furthermore, we prototype our proposed paradigm on a real-world wireless testbed, demonstrating its practical applicability for various IMAs. We will release our code to the community. Xinye Cao, Hongcan Guo, Guoshun Nan, Jiaoyang Cui, Haoting Qian, Yihan Lin 0001, Yilin Peng, Diyang Zhang, Yan-Zhao Hou, Huici Wu, Xiaofeng Tao 0001, Tony Q. S. Quek |
IEEE J. Sel. Areas Commun. | 1 |
| 2025 | Exploring LLM-Based Multi-Agent Situation Awareness for Zero-Trust Space-Air-Ground Integrated NetworkabstractSpace-air-ground integrated network (SAGIN), which integrates satellite systems, aerial networks, and terrestrial communications, offers ubiquitous coverage for a multitude of applications. Nevertheless, the highly dynamic and open nature of SAGIN increases the network’s vulnerability. Hence, zero-trust security, operating on the principle of “never trust, always verify”, holds the significant potential of securing SAGIN. However, implementing zero-trust SAGIN in practice presents three primary challenges: 1) understanding massive unstructured threat information across diverse domains, 2) performing adaptive security assessments, and 3) making in-depth security decisions. This motivates us to propose SAG-Attack and LLM-SA to enhance zero-trust SAGIN. SAG-Attack serves as a simulator that aims to mimic various attacks in SAGIN. Our LLM-SA is a novel situation awareness method that explores the multiple agents of large language model (LLM). Specifically, the output logs of SAG-Attack will be fed into LLM-SA, and LLM-SA fuses vast amounts of heterogeneous threat information from various domains, thus tackling the first challenge. Then, our LLM-SA relies on multiple LLM-based agents to perform adaptive security assessments, utilizing the chain-of-thought capabilities of LLMs to automatically generate in-depth defense strategies, thereby addressing the second and third challenges. Experiments on five benchmarks demonstrate the superiority of the proposed SAG-Attack and LLM-SA. Notably, our method based on open-sourced Llama3-8B even outperforms ChatGPT-4 under the same setting, despite involving significantly fewer parameters. To foster further research in this area, we will release our platform to the community, facilitating the advancement of zero-trust SAGIN. Xinye Cao, Guoshun Nan, Hongcan Guo, Hanqing Mu, Yihan Lin 0001, Qinchuan Zhou, Baohua Qin, Qimei Cui, Xiaofeng Tao 0001, He Fang, Haitao Du, Tony Q. S. Quek |
IEEE J. Sel. Areas Commun. | 1 |
| 2024 | A comprehensive analysis of DAC-SDC FPGA low power object detection challenge
Xinye Cao |
Sci. China Inf. Sci. | 4 |
| 2023 | Uint-Packing: Multiply Your DNN Accelerator Performance via Unsigned Integer DSP PackingabstractDSP blocks are undoubtedly efficient solutions for implementing multiply-accumulate (MAC) operations on FPGA. Since DSP resources are scarce in FPGA, the advanced solution is to pack parallel multiplication operations into a single DSP. However, available methods are based on signed-type multiplication, leading to both loss of accuracy and increased area. To solve these issues simultaneously, we propose an unsigned integer DSP packing generalization model called uint-packing. Guided by this generalization model, we design the novel computational structure of the DNN accelerator. Our system design is state-of-the-art, with 2.8× throughput and 4× energy efficiency compared to the third-place DAC-SDC’22 design. Meng Zhang 0010, Xinye Cao |
DAC | 3 |
| 2022 | Efficient depthwise separable convolution accelerator for classification and UAV object detection
Meng Zhang 0010, Ruixia Wu, Xinye Cao, Wenzhao Liu |
Neurocomputing | 5 |