Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Xiaojie Cai

dblp:398/8422 · DBLP profile ↗
← Back
4ranked-venue papers
1as first author
4since 2021 · last 2026
0009-0007-0215-6753ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 3 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Video understanding and tracking · 34% Language models and text generation · 23% Multi-agent systems · 20%
Computer architecture, parallel and distributed computing, and storage systems
1 paper
Emerging computing paradigms · 100%

Topics — the 8 heaviest of 9, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Knowledge, reasoning and agents › Multi-agent systems
autonomous agents
1.012026
AgencyBench: Benchmarking the Frontiers of Autonomous Agents in 1M-Token Real-World Contexts · ACL (1) 2026
Natural language and speech › Language models and text generation
LLM agents
0.912025
DeepResearcher: Scaling Deep Research via Reinforcement Learning in Real-world Environments · EMNLP 2025
Computer vision › Video understanding and tracking
video anomaly detection
0.912025
UCF-Crime-DVS: A Novel Event-Based Dataset for Video Anomaly Detection with Spiking Neural Networks · AAAI 2025
Computer vision › Video understanding and tracking › video anomaly detection
weakly supervised video anomaly detection
0.912025
UCF-Crime-DVS: A Novel Event-Based Dataset for Video Anomaly Detection with Spiking Neural Networks · AAAI 2025
Emerging computing paradigms
neuromorphic computing
0.912025
UCF-Crime-DVS: A Novel Event-Based Dataset for Video Anomaly Detection with Spiking Neural Networks · AAAI 2025
Emerging computing paradigms › neuromorphic computing
spiking neural network
0.912025
UCF-Crime-DVS: A Novel Event-Based Dataset for Video Anomaly Detection with Spiking Neural Networks · AAAI 2025
Natural language and speech › Language models and text generation
agent benchmarking
0.312026
AgencyBench: Benchmarking the Frontiers of Autonomous Agents in 1M-Token Real-World Contexts · ACL (1) 2026
Natural language and speech › Information extraction and text analysis
web information extraction
0.312025
DeepResearcher: Scaling Deep Research via Reinforcement Learning in Real-world Environments · EMNLP 2025

Methods — techniques the papers use, named apart from their topics

spiking neural network · 1.7multi-scale fusion · 1.7benchmark construction · 1.0retrieval-augmented generation · 0.9reinforcement learning · 0.9multi-agent architecture · 0.9
YearPublicationVenuePosition
2026 AgencyBench: Benchmarking the Frontiers of Autonomous Agents in 1M-Token Real-World Contexts
abstract
Keyu Li, Junhao Shi, Yang Xiao, Mohan Jiang, Jie Sun, Yunze Wu, Dayuan Fu, Shijie Xia, Xiaojie Cai, Tianze Xu, Weiye Si, Wenjie Li, Dequan Wang, Pengfei Liu. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Junhao Shi, Mohan Jiang, Jie Sun 0030, Yunze Wu, Dayuan Fu, Shijie Xia, Xiaojie Cai, Weiye Si, Wenjie Li 0002, Dequan Wang, Pengfei Liu 0003
ACL (1)9
2026 Semantic Boosting via Knowledge Sharing and Feedback for Video Anomaly Detection
abstract
Vision-language models have the potential to enrich purely visual tasks by utilizing the combined representation of images/videos and corresponding textual descriptions. Recent advances in video anomaly detection have also integrated textual information to enhance the understanding of abnormal events. However, existing approaches often merge visual and textual modalities in a straightforward, bottom-up manner, failing to fully explore their interconnections. Moreover, textual captions themselves do not inherently convey “abnormal” attributes. Consequently, these joint representations tend to highlight all salient input features without adequately focusing on high-level tasks such as video anomaly detection. To direct the model’s attention towards anomalies more effectively, we propose incorporating a top-down mechanism into weakly supervised video anomaly detection tasks. A new Knowledge Sharing and Feedback (KSF) framework is designed to unify the representation of anomalies across both video and text. Specifically, we develop a category pattern sharing module that performs knowledge matching, acting as an alignment bridge between abnormal events and their corresponding descriptions. This ensures consistent representations for identical anomalies while maintaining distinct representations for different ones. Following this alignment process, matched high-level semantic priors are fed back into the forward path to enhance differentiation between abnormal and normal patterns. Comprehensive experiments on three benchmark datasets demonstrate the superiority of our proposed method in learning the implicit definition of anomaly patterns. The code is available at https://github.com/XJ-Cai/KSF.
Xiaojie Cai, Yucheng Qian, Chong Wang 0001, Xiaohao Peng, Yuanbin Qian, Jiafei Wu
IEEE Trans. Circuits Syst. Video Technol.1
2025 UCF-Crime-DVS: A Novel Event-Based Dataset for Video Anomaly Detection with Spiking Neural Networks
abstract
Video anomaly detection plays a significant role in intelligent surveillance systems. To enhance model's anomaly recognition ability, previous works have typically involved RGB, optical flow, and text features. Recently, dynamic vision sensors (DVS) have emerged as a promising technology, which capture visual information as discrete events with a very high dynamic range and temporal resolution. It reduces data redundancy and enhances the capture capacity of moving objects compared to conventional camera. To introduce this rich dynamic information into the surveillance field, we created the first DVS video anomaly detection benchmark, namely UCF-Crime-DVS. To fully utilize this new data modality, a multi-scale spiking fusion network (MSF) is designed based on spiking neural networks (SNNs). This work explores the potential application of dynamic information from event data in video anomaly detection. Our experiments demonstrate the effectiveness of our framework on UCF-Crime-DVS and its superior performance compared to other models, establishing a new baseline for SNN-based weakly supervised video anomaly detection.
Yuanbin Qian, Shuhan Ye, Chong Wang 0001, Xiaojie Cai, Jiangbo Qian, Jiafei Wu
AAAI4
2025 DeepResearcher: Scaling Deep Research via Reinforcement Learning in Real-world Environments
abstract
Large Language Models (LLMs) with web search capabilities show significant potential for deep research, yet current methods-brittle prompt engineering or RAG-based reinforcement learning in controlled environments-fail to capture real-world complexities.In this paper, we introduce DeepResearcher, the first comprehensive framework for end-to-end training of LLM-based deep research agents through scaling reinforcement learning (RL) in real-world environments with authentic web search interactions.Unlike RAG approaches reliant on fixed corpora, DeepResearcher trains agents to navigate the noisy, dynamic open web.We implement a specialized multi-agent architecture where browsing agents extract relevant information from various webpage structures and overcoming significant technical challenges.Extensive experiments on open-domain research tasks demonstrate that DeepResearcher achieves substantial improvements of up to 28.9 points over prompt engineering-based baselines and up to 7.2 points over RAG-based RL agents.Our qualitative analysis reveals emergent cognitive behaviors from end-to-end RL training, such as planning, cross-validation, self-reflection for research redirection, and maintain honesty when unable to find definitive answers.Our results highlight that end-to-end training in realworld web environments is fundamental for developing robust research capabilities aligned with real-world applications.The source code for DeepResearcher is released at: https:// github.com/GAIR-NLP/DeepResearcher.
Dayuan Fu, Xiangkun Hu, Xiaojie Cai, Lyumanshan Ye, Pengrui Lu, Pengfei Liu 0003
EMNLP4