VLDB 2026 Research / reviewers in the wild / expert
Xiaojie Cai
dblp:398/8422
· DBLP profile ↗
4ranked-venue papers
1as first author
4since 2021 · last 2026
0009-0007-0215-6753ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 3 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
3 papers |
Video understanding and tracking · 34% Language models and text generation · 23% Multi-agent systems · 20% | |
| Computer architecture, parallel and distributed computing, and storage systems
1 paper |
Emerging computing paradigms · 100% |
Topics — the 8 heaviest of 9, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Knowledge, reasoning and agents › Multi-agent systems
autonomous agents |
1.0 | 1 | 2026 | AgencyBench: Benchmarking the Frontiers of Autonomous Agents in 1M-Token Real-World Contexts · ACL (1) 2026 |
Natural language and speech › Language models and text generation
LLM agents |
0.9 | 1 | 2025 | DeepResearcher: Scaling Deep Research via Reinforcement Learning in Real-world Environments · EMNLP 2025 |
Computer vision › Video understanding and tracking
video anomaly detection |
0.9 | 1 | 2025 | UCF-Crime-DVS: A Novel Event-Based Dataset for Video Anomaly Detection with Spiking Neural Networks · AAAI 2025 |
Computer vision › Video understanding and tracking › video anomaly detection
weakly supervised video anomaly detection |
0.9 | 1 | 2025 | UCF-Crime-DVS: A Novel Event-Based Dataset for Video Anomaly Detection with Spiking Neural Networks · AAAI 2025 |
Emerging computing paradigms
neuromorphic computing |
0.9 | 1 | 2025 | UCF-Crime-DVS: A Novel Event-Based Dataset for Video Anomaly Detection with Spiking Neural Networks · AAAI 2025 |
Emerging computing paradigms › neuromorphic computing
spiking neural network |
0.9 | 1 | 2025 | UCF-Crime-DVS: A Novel Event-Based Dataset for Video Anomaly Detection with Spiking Neural Networks · AAAI 2025 |
Natural language and speech › Language models and text generation
agent benchmarking |
0.3 | 1 | 2026 | AgencyBench: Benchmarking the Frontiers of Autonomous Agents in 1M-Token Real-World Contexts · ACL (1) 2026 |
Natural language and speech › Information extraction and text analysis
web information extraction |
0.3 | 1 | 2025 | DeepResearcher: Scaling Deep Research via Reinforcement Learning in Real-world Environments · EMNLP 2025 |
Methods — techniques the papers use, named apart from their topics
spiking neural network · 1.7multi-scale fusion · 1.7benchmark construction · 1.0retrieval-augmented generation · 0.9reinforcement learning · 0.9multi-agent architecture · 0.9
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | AgencyBench: Benchmarking the Frontiers of Autonomous Agents in 1M-Token Real-World ContextsabstractKeyu Li, Junhao Shi, Yang Xiao, Mohan Jiang, Jie Sun, Yunze Wu, Dayuan Fu, Shijie Xia, Xiaojie Cai, Tianze Xu, Weiye Si, Wenjie Li, Dequan Wang, Pengfei Liu. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Junhao Shi, Mohan Jiang, Jie Sun 0030, Yunze Wu, Dayuan Fu, Shijie Xia, Xiaojie Cai, Weiye Si, Wenjie Li 0002, Dequan Wang, Pengfei Liu 0003 |
ACL (1) | 9 |
| 2026 | Semantic Boosting via Knowledge Sharing and Feedback for Video Anomaly DetectionabstractVision-language models have the potential to enrich purely visual tasks by utilizing the combined representation of images/videos and corresponding textual descriptions. Recent advances in video anomaly detection have also integrated textual information to enhance the understanding of abnormal events. However, existing approaches often merge visual and textual modalities in a straightforward, bottom-up manner, failing to fully explore their interconnections. Moreover, textual captions themselves do not inherently convey “abnormal” attributes. Consequently, these joint representations tend to highlight all salient input features without adequately focusing on high-level tasks such as video anomaly detection. To direct the model’s attention towards anomalies more effectively, we propose incorporating a top-down mechanism into weakly supervised video anomaly detection tasks. A new Knowledge Sharing and Feedback (KSF) framework is designed to unify the representation of anomalies across both video and text. Specifically, we develop a category pattern sharing module that performs knowledge matching, acting as an alignment bridge between abnormal events and their corresponding descriptions. This ensures consistent representations for identical anomalies while maintaining distinct representations for different ones. Following this alignment process, matched high-level semantic priors are fed back into the forward path to enhance differentiation between abnormal and normal patterns. Comprehensive experiments on three benchmark datasets demonstrate the superiority of our proposed method in learning the implicit definition of anomaly patterns. The code is available at https://github.com/XJ-Cai/KSF. Xiaojie Cai, Yucheng Qian, Chong Wang 0001, Xiaohao Peng, Yuanbin Qian, Jiafei Wu |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2025 | UCF-Crime-DVS: A Novel Event-Based Dataset for Video Anomaly Detection with Spiking Neural NetworksabstractVideo anomaly detection plays a significant role in intelligent surveillance systems. To enhance model's anomaly recognition ability, previous works have typically involved RGB, optical flow, and text features. Recently, dynamic vision sensors (DVS) have emerged as a promising technology, which capture visual information as discrete events with a very high dynamic range and temporal resolution. It reduces data redundancy and enhances the capture capacity of moving objects compared to conventional camera. To introduce this rich dynamic information into the surveillance field, we created the first DVS video anomaly detection benchmark, namely UCF-Crime-DVS. To fully utilize this new data modality, a multi-scale spiking fusion network (MSF) is designed based on spiking neural networks (SNNs). This work explores the potential application of dynamic information from event data in video anomaly detection. Our experiments demonstrate the effectiveness of our framework on UCF-Crime-DVS and its superior performance compared to other models, establishing a new baseline for SNN-based weakly supervised video anomaly detection. Yuanbin Qian, Shuhan Ye, Chong Wang 0001, Xiaojie Cai, Jiangbo Qian, Jiafei Wu |
AAAI | 4 |
| 2025 | DeepResearcher: Scaling Deep Research via Reinforcement Learning in Real-world EnvironmentsabstractLarge Language Models (LLMs) with web search capabilities show significant potential for deep research, yet current methods-brittle prompt engineering or RAG-based reinforcement learning in controlled environments-fail to capture real-world complexities.In this paper, we introduce DeepResearcher, the first comprehensive framework for end-to-end training of LLM-based deep research agents through scaling reinforcement learning (RL) in real-world environments with authentic web search interactions.Unlike RAG approaches reliant on fixed corpora, DeepResearcher trains agents to navigate the noisy, dynamic open web.We implement a specialized multi-agent architecture where browsing agents extract relevant information from various webpage structures and overcoming significant technical challenges.Extensive experiments on open-domain research tasks demonstrate that DeepResearcher achieves substantial improvements of up to 28.9 points over prompt engineering-based baselines and up to 7.2 points over RAG-based RL agents.Our qualitative analysis reveals emergent cognitive behaviors from end-to-end RL training, such as planning, cross-validation, self-reflection for research redirection, and maintain honesty when unable to find definitive answers.Our results highlight that end-to-end training in realworld web environments is fundamental for developing robust research capabilities aligned with real-world applications.The source code for DeepResearcher is released at: https:// github.com/GAIR-NLP/DeepResearcher. Dayuan Fu, Xiangkun Hu, Xiaojie Cai, Lyumanshan Ye, Pengrui Lu, Pengfei Liu 0003 |
EMNLP | 4 |