VLDB 2026 Research / reviewers in the wild / expert
Jianwu Fang
dblp:142/0412
· DBLP profile ↗
54ranked-venue papers
10as first author
33since 2021 · last 2026
0000-0002-0300-6892ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 30 · 2 first-author · 17 since 2021Graphics, computer vision, multimedia, augmented reality and games · 21 · 5 first-author · 12 since 2021Applied, interdisciplinary, general and emerging computing · 9 · 4 first-author · 8 since 2021Systems, architecture and hardware · 8 · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Guided Distillation and Risk Adaptive Evolution for Multi-Robot NavigationabstractRecent advancements in multi-robot navigation have explored methods that combine Large Language Models (LLMs) for tasks like scene understanding or high-level decision-making. However, these approaches face challenges with high inference latency and potential hallucinations. To address these challenges, we propose a knowledge-driven Reinforcement Learning (RL) framework, GUIDER, that utilizes an LLM in two different offline roles. First, we leverage the LLM as an offline knowledge source. Its expertise is distilled into a compact model, which is applied only when the RL agent is uncertain about its own value estimates and the model itself is confident in its prediction. Additionally, we utilize the LLM as an offline semantic engine. This process translates the LLM's high-level understanding of situational risk into a dynamic adjustment of the RL agent's behavioral style, evolving a function that optimally balances conservative and aggressive actions. We conduct extensive experiments in both terrestrial and maritime settings. Across all maritime scenarios (3–12 robots), GUIDER improves the task success rate and reduces the collision rate significantly compared to the state-of-the-art RL-based multi-robot navigation methods. Jianwu Fang, Lin Li 0085, Guangliang Li, Jianru Xue |
AAAI | 2 |
| 2026 | ADVersa: Abductive Driving Accident Video UnderstandingabstractUnderstanding traffic accident scenes is a long-standing research for vision-based safe driving. It seeks to answer why accidents occur, how near-crash scenes develop, and what the key elements of an accident are. This research is challenging due to the scarcity and fragmentation of accident data, as well as the complex accident environments. To study this, we present a framework of Abductive Driving accident Video understanding (ADVersa), which infers a plausible visual and textual explanation for the absent near-crash scenes. ADVersa underscores three groups of tasks: 1) visual past recovery of near-crash scenes, 2) visual prediction of near-crash scenes, and 3) accident cause involved video synthesis. To support the study, we first contribute MM-AU, a novel dataset for Multi-Modal Accident video Understanding. MM-AU contains 11,727 in-the-wild driving accident videos with temporally aligned text descriptions, 2.23 million well-annotated object boxes, and 58,650 pairs of video-based accident cause texts. We then propose an Abductive CLIP model and a Contrastive Graph Video Pre-training (CGVP) model, which exploit relation-aware cross-modal semantic learning to drive spatially abductive and temporally abductive accident video diffusion. Extensive experiments verify the superiority of ADVersa to the state-of-the-art approaches on different tasks, i.e., historical near-crash video frame recovering, crashing video frame prediction, textual accident cause and category reasoning, normal-to-accident video synthesis, and accident video editing. With these efforts, we hope this research can advance the progress on multimodal accident video understanding. Lei-Lei Li, Jianwu Fang, Junbin Xiao, Hongkai Yu, Chen Lv 0001, Jianru Xue, Zhengguo Li, Tat-Seng Chua |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2026 | A Human-Oriented Cooperative Driving Approach: Integrating Driving Intention, State, and ConflictabstractHuman-vehicle cooperative driving serves as a vital bridge to fully autonomous driving by improving driving flexibility and gradually building driver trust and acceptance of autonomous technology. To establish more natural and effective human-vehicle interaction, we propose a Human-Oriented Cooperative Driving (HOCD) approach that primarily minimizes human-machine conflict by prioritizing driver intention and state. In implementation, we take both tactical and operational levels into account to ensure seamless human-vehicle cooperation. At the tactical level, we design an intention-aware trajectory planning method, using intention consistency cost as the core metric to evaluate the trajectory and align it with driver intention. At the operational level, we develop a control authority allocation strategy based on reinforcement learning, optimizing the policy through a designed reward function to achieve consistency between driver state and authority allocation. The results of simulation and human-in-the-loop experiments demonstrate that our proposed approach not only aligns with driver intention in trajectory planning but also ensures a reasonable authority allocation. Compared to other cooperative driving approaches, the proposed HOCD approach significantly enhances driving performance and mitigates human-machine conflict. Shanmin Pang, Jianwu Fang, Shengye Dong, Fuhao Liu, Jianru Xue, Chen Lv 0001 |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2026 | Generative Adversarial Self-Imitation Learning With Large Language Model Feedback for Robot Control and Navigation
Enqi Zhao, Zicheng Sun, Jianwu Fang, Eric Nichols, Randy Gomez, Bo He 0002, Jianru Xue, Guangliang Li |
IEEE Trans. Robotics | 5 |
| 2025 | Causal-Entity Reflected Egocentric Traffic Accident Video Synthesis
Lei-Lei Li, Jianwu Fang, Junbin Xiao, Shanmin Pang, Hongkai Yu, Chen Lv 0001, Jianru Xue, Tat-Seng Chua |
ICCV | 2 |
| 2025 | Accelerate 3D Object Detection Models via Zero-Shot Attention Key PruningabstractQuery-based methods with dense features have demonstrated remarkable success in 3D object detection tasks. However, the computational demands of these models, particularly with large image sizes and multiple transformer layers, pose significant challenges for efficient running on edge devices. Existing pruning and distillation methods either need retraining or are designed for ViT models, which are hard to migrate to 3D detectors. To address this issue, we propose a zero-shot runtime pruning method for transformer decoders in 3D object detection models. The method, termed tgGBC (trim keys gradually Guided By Classification scores), systematically trims keys in transformer modules based on their importance. We expand the classification score to multiply it with the attention map to get the importance score of each key and then prune certain keys after each transformer layer according to their importance scores. Our method achieves a 1.99x speedup in the transformer decoder of the latest ToC3D model, with only a minimal performance loss of less than 1%. Interestingly, for certain models, our method even enhances their performance. Moreover, we deploy 3D detectors with tgGBC on an edge device, further validating the effectiveness of our method. The code can be found at https://github.com/iseri27/tg_gbc. Lizhen Xu, Xiuxiu Bai, Xiaojun Jia, Jianwu Fang, Shanmin Pang |
ICCV | 4 |
| 2025 | V2X-DG: Domain Generalization for Vehicle-to-Everything Cooperative PerceptionabstractLiDAR-based Vehicle-to-Everything (V2X) cooperative perception has demonstrated its impact on the safety and effectiveness of autonomous driving. Since current cooperative perception algorithms are trained and tested on the same dataset, the generalization ability of cooperative perception systems remains underexplored. This paper is the first work to study the Domain Generalization problem of LiDAR-based V2X cooperative perception (V2X-DG) for 3D detection based on four widely-used open source datasets: OPV2V, V2XSet, V2V4Real and DAIR-V2X. Our research seeks to sustain high performance not only within the source domain but also across other unseen domains, achieved solely through training on source domain. To this end, we propose Cooperative Mixup Augmentation based Generalization (CMAG) to improve the model generalization capability by simulating the unseen cooperation, which is designed compactly for the domain gaps in cooperative perception. Furthermore, we propose a constraint for the regularization of the robust generalized feature representation learning: Cooperation Feature Consistency (CFC), which aligns the intermediately fused features of the generalized cooperation by CMAG and the early fused features of the original cooperation in source domain. Extensive experiments demonstrate that our approach achieves significant performance gains when generalizing to other unseen datasets while it also maintains strong performance on the source dataset. Zongzhe Xu, Xinyu Liu 0009, Jianwu Fang, Xiaopeng Li 0020, Hongkai Yu |
ICRA | 5 |
| 2025 | Causal-Planner: Causal Interaction Disentangling with Episodic Memory Gating for Autonomous PlanningabstractAutonomous vehicle trajectory planning faces significant challenges in dynamic traffic environments due to the complex and mixed causal relationships between critical scene elements (e.g., pedestrians, vehicles, road markings) and safe decision-making. To identify the causal factors influencing planning outcomes, we propose Causal-Planner, which disentangles the scene interaction graph into causal and confounding components via attention-based adversarial graph learning. Additionally, we introduce a long-short-term episodic memory gating (LSTEM) module that enhances causal interaction disentangling by adaptively capturing evolving causal relationships in dynamic scenarios through bidirectional gated memory fusion. Extensive experiments on the nuPlan dataset suggest that Causal-Planner achieves competitive performance, performing well in both Test-random and Test-hard scenarios under open-loop and closed-loop evaluations. The code will be publicly available at https://github.com/Yyb-XJTU/Causal-Planner. Yibo Yuan, Jianwu Fang, Chen Lv 0001, Jianru Xue |
IROS | 2 |
| 2025 | HeightMapNet: Explicit Height Modeling for End-to-End HD Map Learning
Wenzhao Qiu, Shanmin Pang, Jianwu Fang, Jianru Xue |
WACV | 4 |
| 2025 | LLM-augmented hierarchical reinforcement learning for human-like decision-making of autonomous driving
Lin Li 0085, Runjia Tan, Jianwu Fang, Jianru Xue, Chen Lv 0001 |
Expert Syst. Appl. | 3 |
| 2025 | Gating Syn-to-Real Knowledge for Pedestrian Crossing Prediction in Safe DrivingabstractPedestrian crossing prediction (PCP) in driving scenes plays a critical role in ensuring the safe decision of intelligent vehicles. Due to the limited observations and annotations of pedestrian crossing behaviors in real situations, recent studies have begun to leverage synthetic data with flexible variation to boost prediction performance, employing domain adaptation frameworks. However, different domain knowledge has distinct cross-domain distribution gaps, which necessitates suitable domain knowledge adaption ways for PCP tasks. In this work, we propose a gated syn-to-real knowledge transfer approach for PCP (Gated-S2R-PCP), which has two aims: 1) designing the suitable domain adaptation ways for different kinds of crossing-domain knowledge, and 2) transferring suitable knowledge for specific situations with gated knowledge fusion. Specifically, we design a framework that contains three domain adaption methods including style transfer, distribution approximation, and knowledge distillation for various information, such as visual, semantic, depth, bounding boxes, etc. A learnable gated unit (LGU) is employed to fuse suitable cross-domain knowledge to boost pedestrian crossing prediction. We construct a new synthetic benchmark S2R-PCP-3181 with 3181 sequences (489,740 frames) which contains the pedestrian bounding boxes, RGB frames, semantic segmentation maps, and depth maps. With the synthetic S2R-PCP-3181, we transfer the knowledge to two real challenging datasets of PIE and JAAD, and superior PCP performance is obtained to the state-of-the-art methods. Jianwu Fang, Chen Lv 0001, Jianru Xue, Zhengguo Li |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2025 | EQ-TAA: Equivariant Traffic Accident Anticipation via Diffusion-Based Accident Video SynthesisabstractTraffic Accident Anticipation (TAA) in traffic scenes is a challenging problem for achieving zero fatalities in the future. Current approaches typically treat TAA as a supervised learning task needing the laborious annotation of accident occurrence duration. However, the inherent long-tailed, uncertain, and fast-evolving nature of traffic scenes has the problem that real causal parts of accidents are difficult to identify and are easily dominated by data bias, resulting in a background confounding issue. Thus, we propose an Attentive Video Diffusion (AVD) model that synthesizes additional accident video clips by generating the causal part in dashcam videos, i.e., from normal clips to accident clips. AVD aims to generate causal video frames based on accident or accident-free text prompts while preserving the style and content of frames for TAA after video generation. This approach can be trained using datasets collected from various driving scenes without any extra annotations. Additionally, AVD facilitates an Equivariant TAA (EQ-TAA) with an equivariant triple loss for an anchor accident-free video clip, along with the generated pair of contrastivepseudo-normalandpseudo-accidentclips. Extensive experiments have been conducted to evaluate the performance of AVD and EQ-TAA, and competitive performance compared to state-of-the-art methods has been obtained. Jianwu Fang, Lei-Lei Li, Zhedong Zheng, Hongkai Yu, Jianru Xue, Zhengguo Li, Tat-Seng Chua |
IEEE Trans. Multim. | 1 |
| 2024 | S³-Match: Common-View Aligned Image Matching via Self-Supervised Keypoint Selection
Shizhen Li, Jianwu Fang, Dezheng Gao, Jianru Xue |
BMVC | 3 |
| 2024 | Abductive Ego-View Accident Video Understanding for Safe Driving PerceptionabstractWe present MM-AU, a novel dataset for Multi-Modal Accident video Understanding. MM-AU contains 11,727 in-the-wild ego-view accident videos, each with temporally aligned text descriptions. We annotate over 2.23 mil-lion object boxes and 58,650 pairs of video-based accident reasons, covering 58 accident categories. MM-AU supports various accident understanding tasks, particularly multimodal video diffusion to understand accident cause-effect chains for safe driving. With MM-AU, we present an Abductive accident Video unders tanding framework for Safe Driving perception (AdVersa-SD). AdVersa-SD performs video diffusion via an Object-Centric Video Diffusion (OAVD) method which is driven by an abductive CLIP model. This model involves a contrastive interaction loss to learn the pair co-occurrence of normal, near-accident, accident frames with the corresponding text descriptions, such as accident reasons, prevention advice, and accident categories. OAVD enforces the object region learning while fixing the content of the original frame background in video generation, to find the dominant objects for certain accidents. Extensive experiments verify the abductive ability of AdVersa-SD and the superiority of OAVD against the state-of-the-art diffusion models. Additionally, we provide care-ful benchmark evaluations for object detection and accident reason answering since AdVersa-SD relies on precise object and accident reason information. Jianwu Fang, Lei-Lei Li, Junfei Zhou, Junbin Xiao, Hongkai Yu, Chen Lv 0001, Jianru Xue, Tat-Seng Chua |
CVPR | 1 |
| 2024 | AdvGPS: Adversarial GPS for Multi-Agent Perception AttackabstractThe multi-agent perception system collects visual data from sensors located on various agents and leverages their relative poses determined by GPS signals to effectively fuse information, mitigating the limitations of single-agent sensing, such as occlusion. However, the precision of GPS signals can be influenced by a range of factors, including wireless transmission and obstructions like buildings. Given the pivotal role of GPS signals in perception fusion and the potential for various interference, it becomes imperative to investigate whether specific GPS signals can easily mislead the multi-agent perception system. To address this concern, we frame the task as an adversarial attack challenge and introduce ADVGPS, a method capable of generating adversarial GPS signals which are also stealthy for individual agents within the system, significantly reducing object detection accuracy. To enhance the success rates of these attacks in a black-box scenario, we introduce three types of statistically sensitive natural discrepancies: appearance-based discrepancy, distribution-based discrepancy, and task-aware discrepancy. Our extensive experiments on the OPV2V dataset demonstrate that these attacks substantially undermine the performance of state-of-the-art methods, showcasing remarkable transferability across different point cloud based 3D detection systems. This alarming revelation underscores the pressing need to address security implications within multi-agent perception systems, thereby underscoring a critical area of research. The code is available at https://github.com/jinlong17/AdvGPS. Xinyu Liu 0009, Jianwu Fang, Felix Juefei-Xu, Qing Guo 0005, Hongkai Yu |
ICRA | 4 |
| 2024 | Vehicle Behavior Prediction by Episodic-Memory Implanted NDTabstractIn autonomous driving, predicting the behavior (turning left, stopping, etc.) of target vehicles is crucial for the self-driving vehicle to make safe decisions and avoid accidents. Existing deep learning-based methods have shown excellent and accurate performance, but the black-box nature makes it untrustworthy to apply them in practical use. In this work, we explore the interpretability of behavior prediction of target vehicles by an Episodic Memory implanted Neural Decision Tree (abbrev. eMem-NDT). The structure of eMem-NDT is constructed by hierarchically clustering the text embedding of vehicle behavior descriptions. eMem-NDT is a neural-backed part of a pre-trained deep learning model by changing the soft-max layer of the deep model to eMem-NDT, for grouping and aligning the memory prototypes of the historical vehicle behavior features in training data on a neural decision tree. Each leaf node of eMem-NDT is modeled by a neural network for aligning the behavior memory prototypes. By eMem-NDT, we infer each instance in behavior prediction of vehicles by bottom-up Memory Prototype Matching (MPM) (searching the appropriate leaf node and the links to the root node) and top-down Leaf Link Aggregation (LLA) (obtaining the probability of future behaviors of vehicles for certain instances). We validate eMem-NDT on BLVD and LOKI datasets, and the results show that our model can obtain a superior performance to other methods with clear explainability. The code is available in https://github.com/JWFangit/eMem-NDT. Peining Shen, Jianwu Fang, Hongkai Yu, Jianru Xue |
ICRA | 2 |
| 2024 | VehicleGAN: Pair-flexible Pose Guided Image Synthesis for Vehicle Re-identificationabstractVehicle Re-identification (Re-ID) has been broadly studied in the last decade; however, the different camera view angles leading to confused discrimination in the feature subspace for the vehicles of various poses, is still challenging for the Vehicle Re-ID models in the real world. To promote the Vehicle Re-ID models, this paper proposes to synthesize a large number of vehicle images in the target pose, whose idea is to project the vehicles of diverse poses into the unified target pose so as to enhance feature discrimination. Considering that the paired data of the same vehicles in different traffic surveillance cameras might be not available in the real world, we propose the first Pair-flexible Pose Guided Image Synthesis method for Vehicle Re-ID, named as VehicleGAN in this paper, which works for both supervised and unsupervised settings without the knowledge of geometric 3D models. Because of the feature distribution difference between real and synthetic data, simply training a traditional metric learning based Re-ID model with data-level fusion (i.e., data augmentation) is not satisfactory, therefore we propose a new Joint Metric Learning (JML) via effective feature-level fusion from both real and synthetic data. Intensive experimental results on the public VeRi-776 and VehicleID datasets prove the accuracy and effectiveness of our proposed VehicleGAN and JML. Ping Liu 0004, Lan Fu, Jianwu Fang, Zhigang Xu 0001, Hongkai Yu |
IV | 5 |
| 2024 | Scalable Traffic Simulation for Autonomous Driving via Multi-Agent Goal Assignment and Autoregressive Goal-Directed PlanningabstractSimulation provides a fast, cost-effective, and secure environment for developing autonomous driving systems. However, mitigating the gap between simulation and reality is a challenging task as it demands a behavior simulation method that is human-like, diverse, controllable, socially consistent, and scalable. This work proposes a data-driven traffic agent simulation method to address the aforementioned challenges. Our approach centers around a graph-based scene representation and an encoding method, dividing the simulation into two stages: Multi-Agent Goal assignment (MAG) and Goal-Directed Planning (GDP). Firstly, we create joint goal sets for all agents involved in the scenario. Subsequently, we assign target centerlines (TCLs) to each agent based on their predicted goals. To account for any potential mismatch between the predicted joint goal sets and the road structure, we further align the goals of each agent with their respective assigned TCLs. These on-TCL goals serve as inputs for our interactive autoregressive Goal-Directed Planner (AR-GDP), constituting the second stage of our method that generates roll-outs for simulations. Evaluation results on the leaderboard of the Waymo Open Sim Agents Challenge (WOSAC) 2023 show the competitiveness of the proposed method. Xiaoyu Mo, Zhiyu Huang, Jianwu Fang, Jianru Xue, Chen Lv 0001 |
IV | 4 |
| 2024 | EVD4UAV: An Altitude-Sensitive Benchmark to Evade Vehicle Detection in UAVabstractVehicle detection in Unmanned Aerial Vehicle (UAV) captured images has wide applications in aerial photography and remote sensing. There are many public benchmark datasets proposed for the vehicle detection and tracking in UAV images. Recent studies show that adding an adversarial patch on objects can fool the well-trained deep neural networks based object detectors, posing security concerns to the downstream tasks. However, the current public UAV datasets might ignore the diverse altitudes, vehicle attributes, fine-grained instance-level annotation in mostly side view with blurred vehicle roof, so none of them is good to study the adversarial patch based vehicle detection attack problem. In this paper, we propose a new dataset named EVD4UAV as an altitude-sensitive benchmark to evade vehicle detection in UAV with 6,284 images and 90,886 fine-grained annotated vehicles. The EVD4UAV dataset has diverse altitudes (50m, 70m, 90m), vehicle attributes (color, type), fine-grained annotation (horizontal and rotated bounding boxes, instance-level mask) in top view with clear vehicle roof. One white-box and two black-box patch based attack methods are implemented to attack three classic deep neural networks based object detectors on EVD4UAV. The experimental results show that these representative attack methods could not achieve the robust altitude-insensitive attack performance. Huiming Sun, Jiacheng Guo, Zibo Meng, Tianyun Zhang, Jianwu Fang, Yuewei Lin, Hongkai Yu |
IV | 5 |
| 2024 | GAFB-Mapper: Ground aware Forward-Backward View Transformation for Monocular BEV Semantic MappingabstractMonocular online map segmentation is of great significance to mapless autonomous driving, and the core step is the View Transformation Module (VTM), which is used to transfer feature from the image perspective to the Bird-Eye-View (BEV). Most existing methods directly draw from the field of 3D object perception, either projecting 2D features into 3D space based on depth estimation, or projecting 3D coordinates into 2D images to query corresponding features, while ignoring the geometry and semantics from the ground surface. In this paper, we proposed a ground aware forward-backward view transformation module. The forward projection is used to generate the initial sparse BEV features and the geometric and semantic prior information of the ground surface. The backward module refines the BEV features based on the geometric and semantic priors, thereby improving the accuracy of map segmentation. In addition, the data partitioning of most previous related works has the problem of data leakage, so we repartitioned and experimented on the nuScense data set to conduct a fair evaluation. Experimental results demonstrate that our method achieves the highest accuracy on the test set. Code will be released at https://github.com/Brickzhuantou/MonoBEVseg. Jiangtong Zhu, Yibo Yuan, Zhuo Yin, Shizhen Li, Jianwu Fang, Jianru Xue |
IV | 6 |
| 2024 | IC-Mapper: Instance-Centric Spatio-Temporal Modeling for Online Vectorized Map Construction
Jiangtong Zhu, Yinan Shi, Jianwu Fang, Jianru Xue |
ACM Multimedia | 4 |
| 2024 | Vision-Based Traffic Accident Detection and Anticipation: A SurveyabstractTraffic accident detection and anticipation is an obstinate road safety problem and painstaking efforts have been devoted. With the rapid growth of video data, Vision-based Traffic Accident Detection and Anticipation (named Vision-TAD and Vision-TAA) become the last one-mile problem for safe driving and surveillance safety. However, the long-tailed, unbalanced, highly dynamic, complex, and uncertain properties of traffic accidents form the Out-of-Distribution (OOD) feature for Vision-TAD and Vision-TAA. Current AI development may focus on these OOD but important problems. What has been done for Vision-TAD and Vision-TAA? What direction we should focus on in the future for this problem? A comprehensive survey is important. We present the first survey on Vision-TAD in the deep learning era and the first-ever survey for Vision-TAA. The pros and cons of each research prototype are discussed in detail during the investigation. In addition, we also provide a critical review of 31 publicly available benchmarks and related evaluation metrics. Through this survey, we want to spawn new insights and open possible trends for Vision-TAD and Vision-TAA tasks. Jianwu Fang, Jiahuan Qiao, Jianru Xue, Zhengguo Li |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2024 | Behavioral Intention Prediction in Driving Scenes: A SurveyabstractIn driving scenes, road agents often engage in frequent interaction and strive to understand their surroundings. Ego-agent (each road agent itself) predicts what behavior will be engaged by other road users all the time and expects a shared and consistent understanding for safe movement. To achieve this, Behavioral Intention Prediction (BIP) simulates such a human consideration process to anticipate specific behaviors, and the rapid development of BIP inevitably leads to new issues and challenges. To catalyze future research, this work provides a comprehensive review of BIP from the available datasets, key factors, challenges, pedestrian-centric and vehicle-centric BIP approaches, and BIP-aware applications. The investigation reveals that data-driven deep learning approaches have become the primary pipelines, while the behavioral intention types are still limited in most current datasets and methods (e.g., Crossing (C) and Not Crossing (NC) for pedestrians and Lane Changing (LC) for vehicles) in this field. In addition, current research on BIP in safe-critical scenarios (e.g., near-crashing situations) is limited. Through this investigation, we identify open issues in behavioral intention prediction and suggest possible insights for future research. Jianwu Fang, Jianru Xue, Tat-Seng Chua |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2024 | TADS: a novel dataset for road traffic accident detection from a surveillance perspective
Yachuang Chai, Jianwu Fang, Haoquan Liang, Wushour Slamu |
J. Supercomput. | 2 |
| 2023 | UAVM '23: 2023 Workshop on UAVs in Multimedia: Capturing the World from a New PerspectiveabstractUnmanned Aerial Vehicles (UAVs), also known as drones, have become increasingly popular in recent years due to their ability to capture high-quality multimedia data from the sky. With the rise of multimedia applications, such as aerial photography, cinematography, and mapping, UAVs have emerged as a powerful tool for gathering rich and diverse multimedia content. This workshop aims to bring together researchers, practitioners, and enthusiasts interested in UAV multimedia to explore the latest advancements, challenges, and opportunities in this exciting field. The workshop covers various topics related to UAV multimedia, including aerial image and video processing, machine learning for UAV data analysis, UAV swarm technology, and UAV-based multimedia applications. In the context of the ACM Multimedia conference, this workshop is highly relevant as multimedia data from UAVs is becoming an increasingly important source of content for many multimedia applications. The workshop provides a platform for researchers to share their work and discuss potential collaborations, as well as an opportunity for practitioners to learn about the latest developments in UAV multimedia technology. Overall, this workshop provides a unique opportunity to explore the exciting and rapidly evolving field of UAV multimedia and its potential impact on the wider multimedia community. Zhedong Zheng, Yujiao Shi 0002, Tingyu Wang 0002, Jun Liu 0036, Jianwu Fang, Yunchao Wei, Tat-Seng Chua |
ACM Multimedia | 5 |
| 2023 | Towards Trajectory Forecasting From DetectionabstractTrajectory forecasting for traffic participants (e.g., vehicles) is critical for autonomous platforms to make safe plans. Currently, most trajectory forecasting methods assume that object trajectories have been extracted and directly develop trajectory predictors based on the ground truth trajectories. However, this assumption does not hold in practical situations. Trajectories obtained from object detection and tracking are inevitably noisy, which could cause serious forecasting errors to predictors built on ground truth trajectories. In this paper, we propose to predict trajectories directly based on detection results without relying on explicitly formed trajectories. Different from traditional methods which encode the motion cues of an agent based on its clearly defined trajectory, we extract the motion information only based on the affinity cues among detection results, in which an affinity-aware state update mechanism is designed to manage the state information. In addition, considering that there could be multiple plausible matching candidates, we aggregate the states of them. These designs take the uncertainty of association into account which relax the undesirable effect of noisy trajectory obtained from data association and improve the robustness of the predictor. Extensive experiments validate the effectiveness of our method and its generalization ability to different detectors or forecasting schemes. Pu Zhang 0001, Lei Bai 0001, Jianwu Fang, Jianru Xue, Nanning Zheng 0001, Wanli Ouyang |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2023 | Heterogeneous Trajectory Forecasting via Risk and Scene Graph LearningabstractHeterogeneous trajectory forecasting is critical for intelligent transportation systems, but it is challenging because of the difficulty of modeling the complex interaction relations among the heterogeneous road agents as well as their agent-environment constraints. In this work, we propose a risk and scene graph learning method for trajectory forecasting of heterogeneous road agents, which consists of a Heterogeneous Risk Graph (HRG) and a Hierarchical Scene Graph (HSG) from the aspects of agent category and their movable semantic regions. HRG groups each kind of road agent and calculates their interaction adjacency matrix based on an effective collision risk metric. HSG of the driving scene is modeled by inferring the relationship between road agents and road semantic layout aligned by the road scene grammar. Based on this formulation, we can obtain effective trajectory forecasting in driving situations, and comparable performance to other state-of-the-art approaches is presented by extensive experiments on the nuScenes, ApolloScape, and Argoverse datasets. Jianwu Fang, Pu Zhang 0001, Hongkai Yu, Jianru Xue |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2023 | Parallel Multistage Wide Neural NetworkabstractDeep learning networks have achieved great success in many areas, such as in large-scale image processing. They usually need large computing resources and time and process easy and hard samples inefficiently in the same way. Another undesirable problem is that the network generally needs to be retrained to learn new incoming data. Efforts have been made to reduce the computing resources and realize incremental learning by adjusting architectures, such as scalable effort classifiers, multi-grained cascade forest (gcForest), conditional deep learning (CDL), tree CNN, decision tree structure with knowledge transfer (ERDK), forest of decision trees with radial basis function (RBF) networks, and knowledge transfer (FDRK). In this article, a parallel multistage wide neural network (PMWNN) is presented. It is composed of multiple stages to classify different parts of data. First, a wide radial basis function (WRBF) network is designed to learn features efficiently in the wide direction. It can work on both vector and image instances and can be trained in one epoch using subsampling and least squares (LS). Second, successive stages of WRBF networks are combined to make up the PMWNN. Each stage focuses on the misclassified samples of the previous stage. It can stop growing at an early stage, and a stage can be added incrementally when new training data are acquired. Finally, the stages of the PMWNN can be tested in parallel, thus speeding up the testing process. To sum up, the proposed PMWNN network has the advantages of: 1) optimized computing resources; 2) incremental learning; and 3) parallel testing with stages. The experimental results with the MNIST data, a number of large hyperspectral remote sensing data, and different types of data in different application areas, including many image and nonimage datasets, show that the WRBF and PMWNN can work well on both image and nonimage data and have very competitive accuracy compared to learning models, such as stacked autoencoders, deep belief nets, support vector machine (SVM), multilayer perceptron (MLP), LeNet-5, RBF network, recently proposed CDL, broad learning, gcForest, ERDK, and FDRK. Jiangbo Xi, Okan K. Ersoy, Jianwu Fang, Tianjun Wu, Chaoying Zhao |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2022 | Traffic Accident Detection via Self-Supervised Consistency Learning in Driving ScenariosabstractWith the rapid progress of autonomous driving and advanced driver assistance systems, there are growing efforts to promote their safety in natural driving scenarios, especially for the detection of the traffic accidents. However, because of the dynamic camera motion and complex scene in driving situations, traffic accident detection is still challenging. In this work, we aim to give the ability of Traffic Accident Detection for driving systems by proposing a Self-Supervised Consistency learning framework, termed as SSC-TAD, that involves the appearance, motion, and context consistency learning. The key formulation is to find the inconsistency of video frames, object locations and the spatial relation structure of scene temporally between different frames captured by the dashcam videos. Within this field, different from the previous works which concentrate on predicting the future object locations or frames, we further focus on predicting the visual scene context in driving scenarios and detecting the traffic accident by considering the temporal frame consistency, temporal object location consistency, and the spatial-temporal relation consistency of road participants. In this work, this formulation is fulfilled by a collaborative multi-task consistency learning network and the visual scene context feature is represented by a graph convolution network. The superiority to the state-of-the-art is verified by exhaustive evaluations on two large scale datasets, i.e., the AnAn Accident Detection (A3D) dataset and DADA-2000 dataset collected recently. Jianwu Fang, Jiahuan Qiao, Hongkai Yu, Jianru Xue |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2022 | DADA: Driver Attention Prediction in Driving Accident ScenariosabstractDriver attention prediction is becoming an essential research problem in human-like driving systems. This work makes an attempt to predict thedriverattention indrivingaccident scenarios (DADA). However, challenges tread on the heels of that because of the dynamic traffic scene, intricate and imbalanced accident categories. In this work, we design a semantic context induced attentive fusion network (SCAFNet). We first segment the RGB video frames into the images with different semantic regions (i.e., semantic images), where each region denotes one semantic category of the scene (e.g., road, trees, etc.), and learn the spatio-temporal features of RGB frames and semantic images in two parallel paths simultaneously. Then, the learned features are fused by an attentive fusion network to find the semantic-induced scene variation in driver attention prediction. The contributions are three folds. 1) With the semantic images, we introduce their semantic context features and verify the manifest promotion effect for helping the driver attention prediction, where the semantic context features are modeled by a graph convolution network (GCN) on semantic images; 2) We fuse the semantic context features of semantic images and the features of RGB frames in an attentive strategy, and the fused details are transferred over frames by a convolutional LSTM module to obtain the attention map of each video frame with the consideration of historical scene variation in driving situations; 3) The superiority of the proposed method is evaluated on our previously collected dataset (named as DADA-2000) and two other challenging datasets with state-of-the-art methods. Jianwu Fang, Dingxin Yan, Jiahuan Qiao, Jianru Xue, Hongkai Yu |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2022 | An Adaptive Invariant EKF for Map-Aided Localization Using 3D Point CloudabstractIn map-aided localization using 3D point cloud sets, the poses estimated by 3D Registration Algorithms (3DRAs) are typically fused with other sensor data via the Extended Kalman Filter (EKF) to obtain reliable and smooth results. However, the challenges of this combined method are: 1) The linearization process of EKF may cause errors and singularities due to the state defined by a 6D pose. 2) The results of 3DRA as the measurements of EKF may cause errors in residual calculation. 3) The approach relies heavily on 3DRA to overcome the effects of dynamic scenes. This paper proposes an adaptive localization framework based on Invariant Extended Kalman Filter (Invariant EKF), in which the Lie Group is introduced to define the state. In this framework, the points of the raw point cloud set are the measurements of the filter, and the 3DRA is only employed for data association between the raw 3D point cloud set and the 3D point cloud map. Then, a Concentric Ring Model (CRM) is proposed to reduce the influence of dynamic objects, which can adaptively estimate the covariance of each observed point via Gaussian Process Regression (GPR). Besides, the CRM considers Sensor Measurement Noise (SMN) and Sensor Vibration Noise (SVN). The performance of the proposed framework is evaluated on the KITTI dataset and our dataset. The experimental results show that the proposed method is superior to other state-of-the-art methods, and the CRM can achieve more accurate measurement than ever before, especially in high-dynamic scenes. Zhongxing Tao, Jianru Xue, Di Wang 0028, Gengxin Li, Jianwu Fang |
IEEE Trans. Intell. Transp. Syst. | 5 |
| 2022 | Deep Domain Adaptation Based Multi-Spectral Salient Object DetectionabstractSalient Object Detection (SOD) plays an important role in many image-related multimedia applications. Although there are many existing research works about the salient object detection in traditional RGB (visible-light spectrum) images, there are still many complex situations that regular RGB images cannot provide enough cues for the accurate SOD, such as the shadow effect, similar appearance between background and foreground, strong or insufficient illumination, etc. Because of the success of near-infrared spectrum in many computer vision tasks, we explore the multi-spectral SOD in the synchronized RGB images and near-infrared (NIR) images for the both simple and complex situations. We assume that the RGB SOD in the existing RGB image datasets could provide references for the multi-spectral SOD problem. In this paper, we mainly model this research problem as a deep learning based domain adaptation from the traditional RGB image data (source domain) to the multi-spectral data (target domain), and an adversarial deep domain adaptation model is proposed. We first collect and will publicize a large multi-spectral dataset, RGBN-SOD dataset, including 780 synchronized RGB and NIR image pairs for the multi-spectral SOD problem in the simple and complex situations. Intensive experimental results show the effectiveness and accuracy of the proposed deep domain adaptation for the multi-spectral SOD. Besides, due to the absence of research on the field of multi-spectral co-saliency detection, we also collect 200 synchronized RGB and NIR image pairs in addition to explore the multi-spectral co-saliency detection. Shaoyue Song, Zhenjiang Miao, Hongkai Yu, Jianwu Fang, Cong Ma 0004, Song Wang 0002 |
IEEE Trans. Multim. | 4 |
| 2021 | Video Frame Prediction by Deep Multi-Branch Mask NetworkabstractFuture frame prediction in video is one of the most important problem in computer vision, and useful for a range of practical applications, such as intention prediction or video anomaly detection. However, this task is challenging because of the complex and dynamic evolution of scene. The difficulty of video frame prediction is to model the inherent spatio-temporal correlation between frames and pose an adaptive and flexible framework for large motion change or appearance variation. In this paper, we construct a deep multi-branch mask network (DMMNet) which adaptively fuses the advantages of optical flow warping and RGB pixel synthesizing methods, i.e., the common two kinds of approaches in this task. In the procedure of DMMNet, we add mask layer in each branch to adaptively adjust the magnitude range of estimated optical flow and the weight of predicted frames by optical flow warping and RGB pixel synthesizing, respectively. In other words, we provide a more flexible masking network for motion and appearance fusion on video frame prediction. Exhaustive experiments on Caltech pedestrian and UCF101 datasets show that the proposed model can obtain favorable video frame prediction performance compared with the state-of-the-art methods. In addition, we also put our model into the video anomaly detection problem, and the superiority is verified by the experiments on UCSD dataset. Jianwu Fang, Hongke Xu, Jianru Xue |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2020 | Multi-Spectral Salient Object Detection by Adversarial Domain AdaptationabstractAlthough there are many existing research works about the salient object detection (SOD) in RGB images, there are still many complex situations that regular RGB images cannot provide enough cues for the accurate SOD, such as the shadow effect, similar appearance between background and foreground, strong or insufficient illumination, etc. Because of the success of near-infrared spectrum in many computer vision tasks, we explore the multi-spectral SOD in the synchronized RGB images and near-infrared (NIR) images for the both simple and complex situations. We assume that the RGB SOD in the existing RGB image datasets could provide references for the multi-spectral SOD problem. In this paper, we first collect and will publicize a large multi-spectral dataset including 780 synchronized RGB and NIR image pairs for the multi-spectral SOD problem in the simple and complex situations. We model this research problem as an adversarial domain adaptation from the existing RGB image dataset (source domain) to the collected multi-spectral dataset (target domain). Experimental results show the effectiveness and accuracy of the proposed adversarial domain adaptation for the multi-spectral SOD. Shaoyue Song, Hongkai Yu, Zhenjiang Miao, Jianwu Fang, Cong Ma 0004, Song Wang 0002 |
AAAI | 4 |
| 2020 | Navigation Command Matching for Vision-based Autonomous DrivingabstractLearning an optimal policy for autonomous driving task to confront with complex environment is a long- studied challenge. Imitative reinforcement learning is accepted as a promising approach to learn a robust driving policy through expert demonstrations and interactions with environments. However, this model utilizes non-smooth rewards, which have a negative impact on matching between navigation commands and trajectory (state-action pairs), and degrade the generalizability of an agent. Smooth rewards are crucial to discriminate actions generated from sub-optimal policy. In this paper, we propose a navigation command matching (NCM) model to address this issue. There are two key components in NCM, 1) a matching measurer produces smooth navigation rewards that measure matching between navigation commands and trajectory; 2) attention-guided agent performs actions given states where salient regions in RGB images (i.e. roadsides, lane markings and dynamic obstacles) are highlighted to amplify their influence on the final model. We obtain navigation rewards and store transitions to replay buffer after an episode, so NCM is able to discriminate actions generated from suboptimal policy. Experiments on CARLA driving benchmark show our proposed NCM outperforms previous state-of-the- art models on various tasks in terms of the percentage of successfully completed episodes. Moreover, our model improves generalizability of the agent and obtains good performance even in unseen scenarios. Yuxin Pan, Jianru Xue, Pengfei Zhang 0005, Wanli Ouyang, Jianwu Fang, Xingyu Chen 0001 |
ICRA | 5 |
| 2020 | Using Detection, Tracking and Prediction in Visual SLAM to Achieve Real-time Semantic Mapping of Dynamic ScenariosabstractIn this paper, we propose a lightweight system, RDS-SLAM, based on ORB-SLAM2, which can accurately estimate poses and build semantic maps at object level for dynamic scenarios in real time using only one commonly used Intel Core i7 CPU. In RDS-SLAM, three major improvements, as well as major architectural modifications, are proposed to overcome the limitations of ORB-SLAM2. Firstly, it adopts a lightweight object detection neural network in key frames. Secondly, an efficient tracking and prediction mechanism is embedded into the system to remove the feature points belonging to movable objects in all incoming frames. Thirdly, a semantic octree map is built by probabilistic fusion of detection and tracking results, which enables a robot to maintain a semantic description at object level for potential interactions in dynamic scenarios. We evaluate RDS-SLAM in TUM RGB-D dataset, and experimental results show that RDS-SLAM can run with 30.3 ms per frame in dynamic scenarios using only an Intel Core i7 CPU, and achieves comparable accuracy compared with the state-of-the-art SLAM systems which heavily rely on both Intel Core i7 CPUs and powerful GPUs. Xingyu Chen 0001, Jianru Xue, Jianwu Fang, Yuxin Pan, Nanning Zheng 0001 |
IV | 3 |
| 2020 | Improving 3D Object Detection via Joint Attribute-oriented 3D Lossabstract3D object detection has become a hot topic in intelligent vehicle applications in recent years. Generally, deep learning has been the primary framework used in 3D object detection, and regression of the object location and classification of the objectness are the two indispensable components. In the process of training, the ℓn(n=1,2) and the focal loss are considered as the frequent solutions to minimize the regression and classification loss, respectively. However, there are two problems to be solved in the existing methods. For regression component, there is a gap between evaluation metrics, e.g., 3D Intersection over Union (IoU), and the traditional regression loss. As for the classification component, confidence score exists ambiguous due to the binary label assignment of target. To solve these problems, we propose a loss by jointing 3D IoU and other geometric attributes (named as jointed attribute-oriented 3D loss), which can be directly used in optimizing the regression component. In addition, the jointed attribute-oriented 3D loss can assign a soft label for supervising the training of the classification. By incorporating the proposed loss function into several state-of-the-art 3D object detection methods, the significant performance improvement has been achieved on the KITTI benchmark. Jianru Xue, Jian Dou, Yuxin Pan, Jianwu Fang, Di Wang 0028, Nanning Zheng 0001 |
IV | 5 |
| 2020 | Vehicle re-identification in tunnel scenes via synergistically cascade forests
Rixing Zhu, Jianwu Fang, Qi Wang 0009, Hongke Xu, Jianru Xue, Hongkai Yu |
Neurocomputing | 2 |
| 2020 | A New Method and Benchmark for Detecting Co-Saliency Within a Single ImageabstractRecently, saliency detection in a single image and co-saliency detection in multiple images have drawn extensive research interest in the vision and multimedia communities. In this paper, we investigate a new problem of co-saliency detection within a single image, i.e., detecting within-image co-saliency. By identifying common saliency within an image, e.g., highlighting multiple occurrences of an object class with similar appearance, this work can benefit many important applications, such as the detection of objects of interest, more robust object recognition, reduction of information redundancy, and animation synthesis. We propose a new bottom-up method to address this problem. Specifically, a large number of object proposals are first detected from the image. Then we develop an optimization algorithm to derive a set of proposal groups, each of which contains multiple proposals showing good common saliency in the image. For each proposal group, we calculate a co-saliency map and then use a low-rank based algorithm to fuse the maps calculated from all the proposal groups for the final co-saliency map in the image. In the experiment, we collect a new benchmark dataset of 664 color images (two subsets) for within-image co-saliency detection. Experiment results show that the proposed method can better detect the within-image co-saliency than existing algorithms. The experimental results also show that the proposed method can be applied to detect the repetitive patterns in a single image and detect the co-saliency in multiple images. Hongkai Yu, Jianwu Fang, Hao Guo 0002, Song Wang 0002 |
IEEE Trans. Multim. | 3 |
| 2019 | Small Object Detection on Road by Embedding Focal-Area Loss
Jianwu Fang, Jian Dou, Jianru Xue |
ICIG (1) | 2 |
| 2019 | SEG-VoxelNet for 3D Vehicle Detection from RGB and LiDAR DataabstractThis paper proposes a SEG-VoxelNet that takes RGB images and LiDAR point clouds as inputs for accurately detecting 3D vehicles in autonomous driving scenarios, which for the first time introduces semantic segmentation technique to assist the 3D LiDAR point cloud based detection. Specifically, SEG-VoxelNet is composed of two sub-networks: an image semantic segmentation network (SEG-Net) and an improved-VoxelNet. The SEG-Net generates the semantic segmentation map which represents the probability of the category for each pixel. The improved-VoxelNet is capable of effectively fusing point cloud data with image semantic feature and generating accurate 3D bounding boxes of vehicles. Experiments on the KITTI 3D vehicle detection benchmark show that our approach outperforms the methods of state-of-the-art. Jian Dou, Jianru Xue, Jianwu Fang |
ICRA | 3 |
| 2019 | BLVD: Building A Large-scale 5D Semantics Benchmark for Autonomous DrivingabstractIn autonomous driving community, numerous benchmarks have been established to assist the tasks of 3D/2D object detection, stereo vision, semantic/instance segmentation. However, the more meaningful dynamic evolution of the surrounding objects of ego-vehicle is rarely exploited, and lacks a large-scale dataset platform. To address this, we introduce BLVD, a large-scale 5D semantics benchmark which does not concentrate on the static detection or semantic/instance segmentation tasks tackled adequately before. Instead, BLVD aims to provide a platform for the tasks of dynamic 4D (3D+temporal) tracking, 5D (4D+interactive) interactive event recognition and intention prediction. This benchmark will boost the deeper understanding of traffic scenes than ever before. We totally yield 249, 129 3D annotations, 4, 902 independent individuals for tracking with the length of overall 214, 922 points, 6, 004 valid fragments for 5D interactive event recognition, and 4, 900 individuals for 5D intention prediction. These tasks are contained in four kinds of scenarios depending on the object density (low and high) and light conditions (daytime and nighttime). The benchmark can be downloaded from our project site https://github.com/VCCIV/BLVD/. Jianru Xue, Jianwu Fang, Bohua Zhang, Pu Zhang 0001, Jian Dou |
ICRA | 2 |
| 2018 | Co-Saliency Detection Within a Single ImageabstractRecently, saliency detection in a single image and co-saliency detection in multiple images have drawn extensive research interest in the vision community. In this paper, we investigate a new problem of co-saliency detection within a single image, i.e., detecting within-image co-saliency. By identifying common saliency within an image, e.g., highlighting multiple occurrences of an object class with similar appearance, this work can benefit many important applications, such as the detection of objects of interest, more robust object recognition, reduction of information redundancy, and animation synthesis. We propose a new bottom-up method to address this problem. Specifically, a large number of object proposals are first detected from the image. Then we develop an optimization algorithm to derive a set of proposal groups, each of which contains multiple proposals showing good common saliency in the original image. For each proposal group, we calculate a co-saliency map and then use a low-rank based algorithm to fuse the maps calculated from all the proposal groups for the final co-saliency map in the image. In the experiment, we collect a new dataset of 364 color images with within-image cosaliency. Experiment results show that the proposed method can better detect the within-image co-saliency than existing algorithms. Hongkai Yu, Jianwu Fang, Hao Guo 0002, Wei Feng 0005, Song Wang 0002 |
AAAI | 3 |
| 2018 | Incrementally perceiving hazards in driving
Yuan Yuan 0001, Jianwu Fang, Qi Wang 0009 |
Neurocomputing | 2 |
| 2018 | Unsupervised Object-Based Change Detection via a Weibull Mixture Model-Based Binarization for High-Resolution Remote Sensing ImagesabstractObject-based change detection (CD) is an effective method of identifying detailed changes in land features by contrastively observing the same areas of high-resolution remote sensing images at different times. Binarization is the important step in partitioning changed and unchanged classes in the unsupervised domain. We formulate a novel binarization technique based on the Weibull mixture model, where generated similarity measure images are modeled using a mixture of nonnormal Weibull distributions. The parameters in the model are further globally estimated by employing a genetic algorithm. Two data sets with high-resolution remote sensing images are used to evaluate the effectiveness of the proposed method. Experimental results demonstrate that the method allows better and more robust unsupervised object-based CD than do state-of-the-art threshold-based and clustering-based methods. Advantages of the proposed method are embodied in the modeling of relatively few data of the changed class with a skewed and long tail distribution. Tianjun Wu, Jiancheng Luo, Jianwu Fang, Jianghong Ma, Xueli Song |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2017 | Boosting CNN-Based Pedestrian Detection via 3D LiDAR Fusion in Autonomous Driving
Jian Dou, Jianwu Fang, Jianru Xue |
ICIG (2) | 2 |
| 2017 | Online High-Accurate Calibration of RGB+3D-LiDAR for Autonomous Driving
Jianwu Fang, Di Wang 0028, Jianru Xue |
ICIG (3) | 2 |
| 2017 | Spatial-sequential-spectral context awareness trackingabstractVisual context has formed a robust stimulation for visual perception. Spatio-temporal context in existing trackers sometimes shows weak reliability in visible light videos with poor quality. Supplemented by the infrared perception, this work exploits the role of visual context in tracking in a spatial-sequential-spectral view, by which to excavate dominance of different contexts in various scenarios. Specifically, we infer it in the Fourier domain with a real-time speed, and incorporate a fully-occlusion handling and scale adaptation with a trajectory regression filter and object contour closure, respectively. Extensive experiments on 50 video clips simultaneously containing registered RGB and thermal bands demonstrate that our tracker shows a state-of-the-art performance. Jianwu Fang, Jianru Xue |
ICIP | 1 |
| 2017 | Online hash tracking with spatio-temporal saliency auxiliary
Jianwu Fang, Hongke Xu, Qi Wang 0009, Tianjun Wu |
Comput. Vis. Image Underst. | 1 |
| 2015 | Adaptive road detection via context-aware label transfer
Qi Wang 0009, Jianwu Fang, Yuan Yuan 0001 |
Neurocomputing | 2 |
| 2015 | Online Anomaly Detection in Crowd Scenes via Structure AnalysisabstractAbnormal behavior detection in crowd scenes is continuously a challenge in the field of computer vision. For tackling this problem, this paper starts from a novel structure modeling of crowd behavior. We first propose an informative structural context descriptor (SCD) for describing the crowd individual, which originally introduces the potential energy function of particle's interforce in solid-state physics to intuitively conduct vision contextual cueing. For computing the crowd SCD variation effectively, we then design a robust multi-object tracker to associate the targets in different frames, which employs the incremental analytical ability of the 3-D discrete cosine transform (DCT). By online spatial-temporal analyzing the SCD variation of the crowd, the abnormality is finally localized. Our contribution mainly lies on three aspects: 1) the new exploration of abnormal detection from structure modeling where the motion difference between individuals is computed by a novel selective histogram of optical flow that makes the proposed method can deal with more kinds of anomalies; 2) the SCD description that can effectively represent the relationship among the individuals; and 3) the 3-D DCT multi-object tracker that can robustly associate the limited number of (instead of all) targets which makes the tracking analysis in high density crowd situation feasible. Experimental results on several publicly available crowd video datasets verify the effectiveness of the proposed method. Yuan Yuan 0001, Jianwu Fang, Qi Wang 0009 |
IEEE Trans. Cybern. | 2 |
| 2014 | Multi-cue based tracking
Qi Wang 0009, Jianwu Fang, Yuan Yuan 0001 |
Neurocomputing | 2 |
| 2014 | Part-Based Online Tracking With Geometry Constraint and Attention SelectionabstractVisual tracking in condition of occlusion, appearance or illumination change has been a challenging task over decades. Recently, some online trackers, based on the detection by classification framework, have achieved good performance. However, problems are still embodied in at least one of the three aspects: 1) tracking the target with a single region has poor adaptability for occlusion, appearance or illumination change; 2) lack of sample weight estimation, which may cause overfitting issue; and 3) inadequate motion model to prevent target from drifting. For tackling the above problems, this paper presents the contributions as follows: 1) a novel part-based structure is utilized in the online AdaBoost tracking; 2) attentional sample weighting and selection is tackled by introducing a weight relaxation factor, instead of treating the samples equally as traditional trackers do; and 3) a two-stage motion model, multiple parts constraint, is proposed and incorporated into the part-based structure to ensure a stable tracking. The effectiveness and efficiency of the proposed tracker is validated upon several complex video sequences, compared with seven popular online trackers. The experimental results show that the proposed tracker can achieve increased accuracy with comparable computational cost. Jianwu Fang, Qi Wang 0009, Yuan Yuan 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2014 | Robust Superpixel Tracking via Depth FusionabstractAlthough numerous trackers have been designed to adapt to the nonstationary image streams that change over time, it remains a challenging task to facilitate a tracker to accurately distinguish the target from the background in every frame. This paper proposes a robust superpixel-based tracker via depth fusion, which exploits the adequate structural information and great flexibility of mid-level features captured by superpixels, as well as the depth-map's discriminative ability for the target and background separation. By introducing graph-regularized sparse coding into the appearance model, the local geometrical structure of data is considered, and the resulting appearance model has a more powerful discriminative ability. Meanwhile, the similarity of the target superpixels' neighborhoods in two adjacent frames is also incorporated into the refinement of the target estimation, which helps a more accurate localization. Most importantly, the depth cue is fused into the superpixel-based target estimation so as to tackle the cluttered background with similar appearance to the target. To evaluate the effectiveness of the proposed tracker, four video sequences of different challenging situations are contributed by the authors. The comparison results demonstrate that the proposed tracker has more robust and accurate performance than seven ones representing the state-of-the-art. Yuan Yuan 0001, Jianwu Fang, Qi Wang 0009 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |