Jirong Zha

dblp:374/9277 · DBLP profile ↗
← Back
12ranked-venue papers
4as first author
12since 2021 · last 2026
0009-0006-4824-6976ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 7 · 1 first-author · 7 since 2021Artificial intelligence and machine learning · 4 · 3 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 first-author · 3 since 2021Security and privacy · 1 · 1 since 2021
YearPublicationVenuePosition
2026 DIMM: Decoupled Multi-hierarchy Kalman Filter via Reinforcement Learning
abstract
State estimation is challenging for target tracking with high maneuverability, as the target's state transition function changes rapidly, irregularly, and is unknown to the estimator. Existing work based on interacting multiple model (IMM) achieves more accurate estimation than single-filter approaches through model combination, aligning appropriate models for different motion modes of the target over time. However, two limitations of conventional IMM remain unsolved. First, the solution space of the model combination is constrained as the target's diverse kinematic properties in different directions are ignored. Second, the model combination weights calculated by the observation likelihood are not accurate enough due to the measurement uncertainty. In this paper, we propose a novel framework, DIMM, to effectively combine estimates from different motion models in each direction, thus increasing the target tracking accuracy. First, DIMM extends the model combination solution space of conventional IMM from a hyperplane to a hypercube by designing a 3D-decoupled multi-hierarchy filter bank, which describes the target's motion with various-order linear models. Second, DIMM generates more reliable combination weight matrices through a differentiable adaptive fusion network for importance allocation rather than solely relying on the observation likelihood; it contains an attention-based twin delayed deep deterministic policy gradient (TD3) method with a hierarchical reward. Experiments demonstrate that DIMM significantly improves the tracking accuracy of existing state estimation methods by 31.61%~99.23%.
Jirong Zha, Yuxuan Fan, Chen Gao 0001, Xinlei Chen
AAAI1
2026 AirCopBench: A Benchmark for Multi-drone Collaborative Embodied Perception and Reasoning
abstract
Multimodal Large Language Models (MLLMs) have shown promise in single-agent vision tasks, yet benchmarks for evaluating multi-agent collaborative perception remain scarce. This gap is critical, as multi-drone systems provide enhanced coverage, robustness, and collaboration compared to single-sensor setups. Existing multi-image benchmarks mainly target basic perception tasks using high-quality single-agent images, thus failing to evaluate MLLMs in more complex, egocentric collaborative scenarios, especially under real-world degraded perception conditions. To address these challenges, we introduce AirCopBench, the first comprehensive benchmark designed to evaluate MLLMs in embodied aerial collaborative perception under challenging perceptual conditions. AirCopBench includes 14.6k+ questions derived from both simulator and real-world data, spanning four key task dimensions: Scene Understanding, Object Understanding, Perception Assessment, and Collaborative Decision, across 14 task types. We construct the benchmark using data from challenging degraded-perception scenarios with annotated collaborative events, generating large-scale questions through model-, rule-, and human-based methods under rigorous quality control. Evaluations on 40 MLLMs show significant performance gaps in collaborative perception tasks, with the best model trailing humans by 24.38% on average and exhibiting inconsistent results across tasks. Fine-tuning experiments further confirm the feasibility of sim-to-real transfer in aerial collaborative perception.
Jirong Zha, Yuxuan Fan, Chen Gao 0001, Xinlei Chen
AAAI1
2026 Scalable UAV Multi-Hop Networking via Multi-Agent Reinforcement Learning With Large Language Models
abstract
In disaster scenarios, establishing robust emergency communication networks is critical, and unmanned aerial vehicles (UAVs) offer a promising solution to rapidly restore connectivity. However, organizing UAVs to form multi-hop networks in large-scale dynamic environments presents significant challenges, including limitations in algorithmic scalability and the vast exploration space required for coordinated decision-making. To address these issues, we propose MRLMN, a novel framework that integrates multi-agent reinforcement learning (MARL) and large language models (LLMs) to jointly optimize UAV agents toward achieving optimal networking performance. The framework incorporates a grouping strategy with reward decomposition to enhance algorithmic scalability and balance decision-making across UAVs. In addition, behavioral constraints are applied to selected key UAVs to improve the robustness of the network. Furthermore, the framework integrates LLM agents, leveraging knowledge distillation to transfer their high-level decision-making capabilities to MARL agents. This enhances both the efficiency of exploration and the overall training process. In the distillation module, a Hungarian algorithm-based matching scheme is applied to align the decision outputs of the LLM and MARL agents and define the distillation loss. Extensive simulation results validate the effectiveness of our approach, demonstrating significant improvements in network performance over the MAPPO baseline and other comparison methods, including enhanced coverage and communication quality.
Yanggang Xu, Jirong Zha, Weijie Hong, Xiangmin Yi, Chen-Chun Hsia, Xinlei Chen
IEEE Trans. Mob. Comput.2
2026 Breaking the Communication-Accuracy Trade-Off: A Sparsified Information Diffusion Framework for Multi-Agent Collaborative Perception
abstract
The growing relevance of multi-agent systems has drawn increasing focus on communication-efficient filters for collaborative perception to alleviate the system's communication burden. While the event-triggered (ET) mechanism can improve communication efficiency in collaborative state estimation, an inevitable trade-off exists between estimation accuracy and communication cost in ET filters. This paper proposes a fast and accurate ET diffusion-based filter for real-time multi-agent collaborative target tracking, aiming to reduce the system's data transmission without compromise in tracking performance. The proposed filter achieves improved tracking accuracy, reduced data transmission, and accelerated convergence using an error-minimized ET cubature information filter (CIF) for local estimation, and a correlation-aware diffusion strategy for global fusion. The experimental results confirm the scalability of the proposed EDC-CIF algorithm and demonstrate its efficacy in simultaneously reducing estimation error and computation time while significantly enhancing communication efficiency.
Jirong Zha, Chenyu Zhao 0002, Zhenyu Liu 0003, Tao Sun 0013, Xinlei Chen
IEEE Trans. Mob. Comput.1
2025 UrbanVideo-Bench: Benchmarking Vision-Language Models on Embodied Intelligence with Video Data in Urban Spaces
abstract
Baining Zhao, Jianjie Fang, Zichao Dai, Ziyou Wang, Jirong Zha, Weichen Zhang, Chen Gao, Yue Wang, Jinqiang Cui, Xinlei Chen, Yong Li. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Baining Zhao, Jianjie Fang, Zichao Dai, Ziyou Wang, Jirong Zha, Chen Gao 0001, Yue Wang 0007, Jinqiang Cui, Xinlei Chen, Yong Li 0008
ACL (1)5
2025 How to Enable LLM with 3D Capacity? A Survey of Spatial Reasoning in LLM
abstract
3D spatial understanding is essential in real-world applications such as robotics, autonomous vehicles, virtual reality, and medical imaging. Recently, Large Language Models (LLMs), having demonstrated remarkable success across various domains, have been leveraged to enhance 3D understanding tasks, showing potential to surpass traditional computer vision methods. In this survey, we present a comprehensive review of methods integrating LLMs with 3D spatial understanding. We propose a taxonomy that categorizes existing methods into three branches: image-based methods deriving 3D understanding from 2D visual data, point cloud-based methods working directly with 3D representations, and hybrid modality-based methods combining multiple data streams. We systematically review representative methods along these categories, covering data representations, architectural modifications, and training strategies that bridge textual and 3D modalities. Finally, we discuss current limitations, including dataset scarcity and computational challenges, while highlighting promising research directions in spatial perception, multi-modal fusion, and real-world applications.
Jirong Zha, Yuxuan Fan, Xinlei Chen
IJCAI1
2024 A Unified Membership Inference Method for Visual Self-supervised Encoder via Part-aware Capability
abstract
Self-supervised learning shows promise in harnessing extensive unlabeled data, but it also confronts significant privacy concerns, especially in vision. In this paper, we aim to perform membership inference on visual self-supervised models in a more realistic setting: self-supervised training method and details are unknown for an adversary when attacking as he usually faces a black-box system in practice. In this setting, considering that self-supervised model could be trained by completely different self-supervised paradigms, e.g., masked image modeling and contrastive learning, with complex training details, we propose a unified membership inference method called PartCrop. It is motivated by the shared part-aware capability among models and stronger part response on the training data. Specifically, PartCrop crops parts of objects in an image to query responses with the image in representation space. We conduct extensive attacks on self-supervised models with different training protocols and structures using three widely used image datasets. The results verify the effectiveness and generalization of PartCrop. Moreover, to defend against PartCrop, we evaluate two common approaches, i.e., early stop and differential privacy, and propose a tailored method called shrinking crop scale range. The defense experiments indicate that all of them are effective. Our code is available at https://github.com/JiePKU/PartCrop.
Jirong Zha, Ding Li 0001, Leye Wang
CCS2
2024 Poster Abstract: Emergency Networking Using UAVs: A Reinforcement Learning Approach with Large Language Model
abstract
Utilizing unmanned aerial vehicles (UAVs) as mobile access points can assist urban communication systems in establishing emergency networks in disaster scenarios. In this paper, to organize UAVs in large-scale environments for networking purposes, we propose a multi-agent reinforcement learning (MARL) model, in which the design of a selective parameter sharing mechanism and a grouping strategy enhances the model’s scalability. Furthermore, the model adopts a reward mechanism based on intrinsic motivation, using the Large Language Model (LLM), to accelerate the optimization process. Numerical results demonstrate that this algorithm outperforms existing alternatives.
Yanggang Xu, Zhuozhu Jian, Jirong Zha, Xinlei Chen
IPSN3
2024 Demo Abstract: Bio-inspired Tactile Sensing for MAV Landing with Extreme Low-cost Sensors
abstract
MAV (Micro Aerial Vehicle) requires landing on a docking platform for recharging during or after missions due to their limited energy capacity. Inspired by biological tactile sensing, we propose a proprioceptive sensing system that allows MAV to "touch", recognize, and locate the landing platform even when visual or other positioning systems are not functioning properly. We leverage a physical phenomenon: as the MAV approaches a beneath obstacle, it experiences attitude disturbances caused by the airflow generated by the rotor’s reflections from the ground. By employing traditional signal processing and learning-based techniques to analyze signals from the IMU (Inertial Measurement Unit) and motors, the MAV can sense the edges of the platform and further calculate the precise landing coordinates. With a power consumption of less than 40 mW, our system achieves an edge detection error of less than 2 cm and a landing success rate exceeding 90%.CCS CONCEPTS• Applied computing → Aerospace; • Computing methodologies → Machine learning approaches; • Computer systems organization → Sensors and actuators.
Chenyu Zhao 0002, Ciyu Ruan, Jirong Zha, Haoyang Wang 0012, Jiaqi Li 0028, Yuxuan Liu 0010, Xuzhe Wang, Xinlei Chen
IPSN4
2024 Multi-Agent Target Pursuit Using Perception Uncertainty-Aware Reinforcement Learning
abstract
Existing target pursuit systems are able to coordinate a team of mobile agents to capture or intercept unauthorized targets. Multi-agent reinforcement learning (MARL) further empowers pursuit strategies with the potential to emerge complex behaviors. However, existing solutions lack the ability to handle the perception uncertainty caused by relative position measurement noises, which blurs the understanding of the target's state and complicates the pursuit strategy learning process. This study proposes PUARL, which enhances the learning under the perception uncertainty process by guiding exploration with probabilistic estimation and adapting the policy based on awareness of perception uncertainty. We validate its performance in terms of both accuracy and efficiency. PUARL achieves a success rate increase of 12.3%+ and a reduction in total steps by 58.3%+, outperforming both state-of-the-art heuristic and learning-based solutions.
Yuhan Cheng, Jirong Zha, Renjue Yang, Susu Xu, Xinlei Chen
MobiCom2
2024 Scalable Multi-Agent Reinforcement Learning for Effective UAV Scheduling in Multi-Hop Emergency Networks
abstract
Utilizing unmanned aerial vehicles (UAVs) as mobile access points can assist urban communication systems in establishing emergency networks in disaster scenarios. However, in large-scale dynamic environments, the extensive exploration space makes effective collaboration among a large number of UAVs challenging. In this paper, to schedule the deployment of UAVs for networking purposes, we propose a novel approach, MAEN, using multi-agent reinforcement learning. The grouping and information sharing mechanisms in MAEN enable the algorithm to easily scale up the number of UAVs to dozens and address the issue of strategy equilibrium. Additionally, a reward decomposition module is designed to handle coordination and task allocation among UAVs. Experimental results demonstrate that the algorithm outperforms existing algorithms in terms of ground device coverage and communication quality.
Yanggang Xu, Jirong Zha, Jiyuan Ren, Xintao Jiang, Xinlei Chen
MobiCom2
2024 Foes or Friends: Embracing Ground Effect for Edge Detection on Lightweight Drones
abstract
Drone-based rapid and accurate environmental edge detection is highly advantageous for tasks such as disaster relief and autonomous navigation. Current methods, using radar or cameras, raise deployment costs and burden lightweight drones with high computational demands. In this paper, we propose AirTouch, a system that transforms the ground effect from a stability "foe" in traditional flight control views, into a "friend" for accurate and efficient edge detection. Our key insight is that analyzing drone sensor readings and flight commands allows us to detect ground effect changes. Such changes typically indicate the drone flying over an edge, making this information valuable for edge detection. We approach this insight through theoretical analysis, algorithm design, and implementation, fully leveraging the ground effect as a new sensing modality without compromising drone flight stability, thereby achieving accurate and efficient scene edge detection. Extensive evaluations demonstrate that our system achieves a high detection accuracy with mean detection distance errors of 0.051m, outperforming the baseline performance by 86%.
Chenyu Zhao 0002, Ciyu Ruan, Jingao Xu, Haoyang Wang 0012, Jiaqi Li 0028, Jirong Zha, Zheng Yang 0002, Yunhao Liu 0001, Xiao-Ping Zhang 0002, Xinlei Chen
MobiCom7