Jianru Xue

dblp:16/6902 · DBLP profile ↗
← Back
149ranked-venue papers
10as first author
44since 2021 · last 2026
0000-0002-4994-9343ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 76 · 3 first-author · 17 since 2021Artificial intelligence and machine learning · 65 · 4 first-author · 21 since 2021Applied, interdisciplinary, general and emerging computing · 20 · 4 first-author · 11 since 2021Systems, architecture and hardware · 11 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 4Security and privacy · 2 · 1 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 first-author
YearPublicationVenuePosition
2026 Guided Distillation and Risk Adaptive Evolution for Multi-Robot Navigation
abstract
Recent advancements in multi-robot navigation have explored methods that combine Large Language Models (LLMs) for tasks like scene understanding or high-level decision-making. However, these approaches face challenges with high inference latency and potential hallucinations. To address these challenges, we propose a knowledge-driven Reinforcement Learning (RL) framework, GUIDER, that utilizes an LLM in two different offline roles. First, we leverage the LLM as an offline knowledge source. Its expertise is distilled into a compact model, which is applied only when the RL agent is uncertain about its own value estimates and the model itself is confident in its prediction. Additionally, we utilize the LLM as an offline semantic engine. This process translates the LLM's high-level understanding of situational risk into a dynamic adjustment of the RL agent's behavioral style, evolving a function that optimally balances conservative and aggressive actions. We conduct extensive experiments in both terrestrial and maritime settings. Across all maritime scenarios (3–12 robots), GUIDER improves the task success rate and reduces the collision rate significantly compared to the state-of-the-art RL-based multi-robot navigation methods.
Jianwu Fang, Lin Li 0085, Guangliang Li, Jianru Xue
AAAI6
2026 Redundant Queries in DETR-Based 3D Detection Methods: Unnecessary and Prunable
abstract
Query-based models are extensively used in 3D object detection tasks, with a wide range of pre-trained checkpoints readily available online. However, despite their popularity, these models often require an excessive number of object queries, far surpassing the actual number of objects to detect. The redundant queries result in unnecessary computational and memory costs. In this paper, we find that not all queries contribute equally -- a significant portion of queries have a much smaller impact compared to others. Based on this observation, we propose an embarrassingly simple approach called Gradually Pruning Queries (GPQ), which prunes queries incrementally based on their classification scores. A key advantage of GPQ is that it requires no additional learnable parameters. It is straightforward to implement in any query-based method, as it can be seamlessly integrated as a fine-tuning step using an existing checkpoint after training. With GPQ, users can easily generate multiple models with fewer queries, starting from a checkpoint with an excessive number of queries. Experiments on various advanced 3D detectors show that GPQ effectively reduces redundant queries while maintaining performance. Using our method, model inference on desktop GPUs can be accelerated by up to 1.35x. Moreover, after deployment on edge devices, it achieves up to a 67.86% reduction in FLOPs and a 65.16% decrease in inference time.
Lizhen Xu, Wenzhao Qiu, Shanmin Pang, Xiuxiu Bai, Jianru Xue
AAAI7
2026 ADVersa: Abductive Driving Accident Video Understanding
abstract
Understanding traffic accident scenes is a long-standing research for vision-based safe driving. It seeks to answer why accidents occur, how near-crash scenes develop, and what the key elements of an accident are. This research is challenging due to the scarcity and fragmentation of accident data, as well as the complex accident environments. To study this, we present a framework of Abductive Driving accident Video understanding (ADVersa), which infers a plausible visual and textual explanation for the absent near-crash scenes. ADVersa underscores three groups of tasks: 1) visual past recovery of near-crash scenes, 2) visual prediction of near-crash scenes, and 3) accident cause involved video synthesis. To support the study, we first contribute MM-AU, a novel dataset for Multi-Modal Accident video Understanding. MM-AU contains 11,727 in-the-wild driving accident videos with temporally aligned text descriptions, 2.23 million well-annotated object boxes, and 58,650 pairs of video-based accident cause texts. We then propose an Abductive CLIP model and a Contrastive Graph Video Pre-training (CGVP) model, which exploit relation-aware cross-modal semantic learning to drive spatially abductive and temporally abductive accident video diffusion. Extensive experiments verify the superiority of ADVersa to the state-of-the-art approaches on different tasks, i.e., historical near-crash video frame recovering, crashing video frame prediction, textual accident cause and category reasoning, normal-to-accident video synthesis, and accident video editing. With these efforts, we hope this research can advance the progress on multimodal accident video understanding.
Lei-Lei Li, Jianwu Fang, Junbin Xiao, Hongkai Yu, Chen Lv 0001, Jianru Xue, Zhengguo Li, Tat-Seng Chua
IEEE Trans. Pattern Anal. Mach. Intell.6
2026 A Human-Oriented Cooperative Driving Approach: Integrating Driving Intention, State, and Conflict
abstract
Human-vehicle cooperative driving serves as a vital bridge to fully autonomous driving by improving driving flexibility and gradually building driver trust and acceptance of autonomous technology. To establish more natural and effective human-vehicle interaction, we propose a Human-Oriented Cooperative Driving (HOCD) approach that primarily minimizes human-machine conflict by prioritizing driver intention and state. In implementation, we take both tactical and operational levels into account to ensure seamless human-vehicle cooperation. At the tactical level, we design an intention-aware trajectory planning method, using intention consistency cost as the core metric to evaluate the trajectory and align it with driver intention. At the operational level, we develop a control authority allocation strategy based on reinforcement learning, optimizing the policy through a designed reward function to achieve consistency between driver state and authority allocation. The results of simulation and human-in-the-loop experiments demonstrate that our proposed approach not only aligns with driver intention in trajectory planning but also ensures a reasonable authority allocation. Compared to other cooperative driving approaches, the proposed HOCD approach significantly enhances driving performance and mitigates human-machine conflict.
Shanmin Pang, Jianwu Fang, Shengye Dong, Fuhao Liu, Jianru Xue, Chen Lv 0001
IEEE Trans. Intell. Transp. Syst.6
2026 Generative Adversarial Self-Imitation Learning With Large Language Model Feedback for Robot Control and Navigation
Enqi Zhao, Zicheng Sun, Jianwu Fang, Eric Nichols, Randy Gomez, Bo He 0002, Jianru Xue, Guangliang Li
IEEE Trans. Robotics10
2025 Causal-Entity Reflected Egocentric Traffic Accident Video Synthesis
Lei-Lei Li, Jianwu Fang, Junbin Xiao, Shanmin Pang, Hongkai Yu, Chen Lv 0001, Jianru Xue, Tat-Seng Chua
ICCV7
2025 Causal-Planner: Causal Interaction Disentangling with Episodic Memory Gating for Autonomous Planning
abstract
Autonomous vehicle trajectory planning faces significant challenges in dynamic traffic environments due to the complex and mixed causal relationships between critical scene elements (e.g., pedestrians, vehicles, road markings) and safe decision-making. To identify the causal factors influencing planning outcomes, we propose Causal-Planner, which disentangles the scene interaction graph into causal and confounding components via attention-based adversarial graph learning. Additionally, we introduce a long-short-term episodic memory gating (LSTEM) module that enhances causal interaction disentangling by adaptively capturing evolving causal relationships in dynamic scenarios through bidirectional gated memory fusion. Extensive experiments on the nuPlan dataset suggest that Causal-Planner achieves competitive performance, performing well in both Test-random and Test-hard scenarios under open-loop and closed-loop evaluations. The code will be publicly available at https://github.com/Yyb-XJTU/Causal-Planner.
Yibo Yuan, Jianwu Fang, Chen Lv 0001, Jianru Xue
IROS6
2025 HeightMapNet: Explicit Height Modeling for End-to-End HD Map Learning
Wenzhao Qiu, Shanmin Pang, Jianwu Fang, Jianru Xue
WACV5
2025 LLM-augmented hierarchical reinforcement learning for human-like decision-making of autonomous driving
Lin Li 0085, Runjia Tan, Jianwu Fang, Jianru Xue, Chen Lv 0001
Expert Syst. Appl.4
2025 Gating Syn-to-Real Knowledge for Pedestrian Crossing Prediction in Safe Driving
abstract
Pedestrian crossing prediction (PCP) in driving scenes plays a critical role in ensuring the safe decision of intelligent vehicles. Due to the limited observations and annotations of pedestrian crossing behaviors in real situations, recent studies have begun to leverage synthetic data with flexible variation to boost prediction performance, employing domain adaptation frameworks. However, different domain knowledge has distinct cross-domain distribution gaps, which necessitates suitable domain knowledge adaption ways for PCP tasks. In this work, we propose a gated syn-to-real knowledge transfer approach for PCP (Gated-S2R-PCP), which has two aims: 1) designing the suitable domain adaptation ways for different kinds of crossing-domain knowledge, and 2) transferring suitable knowledge for specific situations with gated knowledge fusion. Specifically, we design a framework that contains three domain adaption methods including style transfer, distribution approximation, and knowledge distillation for various information, such as visual, semantic, depth, bounding boxes, etc. A learnable gated unit (LGU) is employed to fuse suitable cross-domain knowledge to boost pedestrian crossing prediction. We construct a new synthetic benchmark S2R-PCP-3181 with 3181 sequences (489,740 frames) which contains the pedestrian bounding boxes, RGB frames, semantic segmentation maps, and depth maps. With the synthetic S2R-PCP-3181, we transfer the knowledge to two real challenging datasets of PIE and JAAD, and superior PCP performance is obtained to the state-of-the-art methods.
Jianwu Fang, Chen Lv 0001, Jianru Xue, Zhengguo Li
IEEE Trans. Intell. Transp. Syst.5
2025 Enhancing Mapless Trajectory Prediction Through Knowledge Distillation
abstract
Scene information plays a crucial role in trajectory forecasting systems for autonomous driving by providing semantic clues and constraints on potential future paths of traffic agents. Prevalent trajectory prediction techniques often take High-Definition maps (HD maps) as part of the inputs to provide scene knowledge. Although HD maps offer accurate road information, they may suffer from the high cost of annotation or restrictions of law, which limits their widespread use. Therefore, it is crucial for trajectory prediction methods to generate reliable prediction results in mapless scenarios. In this paper, we tackle the problem of improving the consistency of predicted trajectories and the scene road topology when map information is unavailable during the test phase. To achieve this, we propose a universal knowledge distillation (KD) framework. This KD framework trains a map-based teacher network on samples with annotated HD maps and subsequently transfers the knowledge to a student mapless predictor through a two-fold knowledge distillation process. Experimental results show that our method stably improves prediction performance in test-time mapless situations on many widely used trajectory prediction baselines, and achieves state-of-the-art mapless prediction performances. Qualitative visualization results demonstrate that our approach helps infer unseen map information. Our solution is generalizable for common trajectory prediction networks and datasets.
Pu Zhang 0001, Lei Bai 0001, Jianru Xue
IEEE Trans. Intell. Transp. Syst.5
2025 Continual Multi-Agent Trajectory Prediction via Dual Knowledge Consolidation and Pseudo Data Generation
abstract
Multi-agent trajectory prediction (MATP) is pivotal in both robotics and autonomous driving. Current multi-agent trajectory prediction methodologies face significant performance degradation due to distributional discrepancies between pre-training data and real-world deployment environments. In practical scenarios, data arrive as a continuous data stream with distribution shifts and potential recurrence of older patterns, constituting a continual learning trajectory prediction problem. Training on a continuous data stream is difficult due to the evolving domain knowledge regarding both intra-sample interactions and trajectory statistical patterns, leading to catastrophic forgetting. We tackle this challenge by proposing a rehearsal-based dual knowledge consolidation framework, which systematically consolidates interaction patterns via intra-sample relational distillation, and retains trajectory statistics via an expandable pattern memory bank and knowledge distillation between pattern centers and features. A pseudo-data generation strategy is also newly proposed to enhance the rehearsal diversity. We design challenging yet practical continual prediction experiments on widely used datasets to evaluate our method. Results demonstrate that our approach significantly improves the prediction performance on continuous data, surpassing existing continuous prediction methods.
Pu Zhang 0001, Jianru Xue
IEEE Trans. Intell. Transp. Syst.5
2025 EQ-TAA: Equivariant Traffic Accident Anticipation via Diffusion-Based Accident Video Synthesis
abstract
Traffic Accident Anticipation (TAA) in traffic scenes is a challenging problem for achieving zero fatalities in the future. Current approaches typically treat TAA as a supervised learning task needing the laborious annotation of accident occurrence duration. However, the inherent long-tailed, uncertain, and fast-evolving nature of traffic scenes has the problem that real causal parts of accidents are difficult to identify and are easily dominated by data bias, resulting in a background confounding issue. Thus, we propose an Attentive Video Diffusion (AVD) model that synthesizes additional accident video clips by generating the causal part in dashcam videos, i.e., from normal clips to accident clips. AVD aims to generate causal video frames based on accident or accident-free text prompts while preserving the style and content of frames for TAA after video generation. This approach can be trained using datasets collected from various driving scenes without any extra annotations. Additionally, AVD facilitates an Equivariant TAA (EQ-TAA) with an equivariant triple loss for an anchor accident-free video clip, along with the generated pair of contrastivepseudo-normalandpseudo-accidentclips. Extensive experiments have been conducted to evaluate the performance of AVD and EQ-TAA, and competitive performance compared to state-of-the-art methods has been obtained.
Jianwu Fang, Lei-Lei Li, Zhedong Zheng, Hongkai Yu, Jianru Xue, Zhengguo Li, Tat-Seng Chua
IEEE Trans. Multim.5
2024 S³-Match: Common-View Aligned Image Matching via Self-Supervised Keypoint Selection
Shizhen Li, Jianwu Fang, Dezheng Gao, Jianru Xue
BMVC5
2024 Abductive Ego-View Accident Video Understanding for Safe Driving Perception
abstract
We present MM-AU, a novel dataset for Multi-Modal Accident video Understanding. MM-AU contains 11,727 in-the-wild ego-view accident videos, each with temporally aligned text descriptions. We annotate over 2.23 mil-lion object boxes and 58,650 pairs of video-based accident reasons, covering 58 accident categories. MM-AU supports various accident understanding tasks, particularly multimodal video diffusion to understand accident cause-effect chains for safe driving. With MM-AU, we present an Abductive accident Video unders tanding framework for Safe Driving perception (AdVersa-SD). AdVersa-SD performs video diffusion via an Object-Centric Video Diffusion (OAVD) method which is driven by an abductive CLIP model. This model involves a contrastive interaction loss to learn the pair co-occurrence of normal, near-accident, accident frames with the corresponding text descriptions, such as accident reasons, prevention advice, and accident categories. OAVD enforces the object region learning while fixing the content of the original frame background in video generation, to find the dominant objects for certain accidents. Extensive experiments verify the abductive ability of AdVersa-SD and the superiority of OAVD against the state-of-the-art diffusion models. Additionally, we provide care-ful benchmark evaluations for object detection and accident reason answering since AdVersa-SD relies on precise object and accident reason information.
Jianwu Fang, Lei-Lei Li, Junfei Zhou, Junbin Xiao, Hongkai Yu, Chen Lv 0001, Jianru Xue, Tat-Seng Chua
CVPR7
2024 Vehicle Behavior Prediction by Episodic-Memory Implanted NDT
abstract
In autonomous driving, predicting the behavior (turning left, stopping, etc.) of target vehicles is crucial for the self-driving vehicle to make safe decisions and avoid accidents. Existing deep learning-based methods have shown excellent and accurate performance, but the black-box nature makes it untrustworthy to apply them in practical use. In this work, we explore the interpretability of behavior prediction of target vehicles by an Episodic Memory implanted Neural Decision Tree (abbrev. eMem-NDT). The structure of eMem-NDT is constructed by hierarchically clustering the text embedding of vehicle behavior descriptions. eMem-NDT is a neural-backed part of a pre-trained deep learning model by changing the soft-max layer of the deep model to eMem-NDT, for grouping and aligning the memory prototypes of the historical vehicle behavior features in training data on a neural decision tree. Each leaf node of eMem-NDT is modeled by a neural network for aligning the behavior memory prototypes. By eMem-NDT, we infer each instance in behavior prediction of vehicles by bottom-up Memory Prototype Matching (MPM) (searching the appropriate leaf node and the links to the root node) and top-down Leaf Link Aggregation (LLA) (obtaining the probability of future behaviors of vehicles for certain instances). We validate eMem-NDT on BLVD and LOKI datasets, and the results show that our model can obtain a superior performance to other methods with clear explainability. The code is available in https://github.com/JWFangit/eMem-NDT.
Peining Shen, Jianwu Fang, Hongkai Yu, Jianru Xue
ICRA4
2024 Scalable Traffic Simulation for Autonomous Driving via Multi-Agent Goal Assignment and Autoregressive Goal-Directed Planning
abstract
Simulation provides a fast, cost-effective, and secure environment for developing autonomous driving systems. However, mitigating the gap between simulation and reality is a challenging task as it demands a behavior simulation method that is human-like, diverse, controllable, socially consistent, and scalable. This work proposes a data-driven traffic agent simulation method to address the aforementioned challenges. Our approach centers around a graph-based scene representation and an encoding method, dividing the simulation into two stages: Multi-Agent Goal assignment (MAG) and Goal-Directed Planning (GDP). Firstly, we create joint goal sets for all agents involved in the scenario. Subsequently, we assign target centerlines (TCLs) to each agent based on their predicted goals. To account for any potential mismatch between the predicted joint goal sets and the road structure, we further align the goals of each agent with their respective assigned TCLs. These on-TCL goals serve as inputs for our interactive autoregressive Goal-Directed Planner (AR-GDP), constituting the second stage of our method that generates roll-outs for simulations. Evaluation results on the leaderboard of the Waymo Open Sim Agents Challenge (WOSAC) 2023 show the competitiveness of the proposed method.
Xiaoyu Mo, Zhiyu Huang, Jianwu Fang, Jianru Xue, Chen Lv 0001
IV5
2024 GAFB-Mapper: Ground aware Forward-Backward View Transformation for Monocular BEV Semantic Mapping
abstract
Monocular online map segmentation is of great significance to mapless autonomous driving, and the core step is the View Transformation Module (VTM), which is used to transfer feature from the image perspective to the Bird-Eye-View (BEV). Most existing methods directly draw from the field of 3D object perception, either projecting 2D features into 3D space based on depth estimation, or projecting 3D coordinates into 2D images to query corresponding features, while ignoring the geometry and semantics from the ground surface. In this paper, we proposed a ground aware forward-backward view transformation module. The forward projection is used to generate the initial sparse BEV features and the geometric and semantic prior information of the ground surface. The backward module refines the BEV features based on the geometric and semantic priors, thereby improving the accuracy of map segmentation. In addition, the data partitioning of most previous related works has the problem of data leakage, so we repartitioned and experimented on the nuScense data set to conduct a fair evaluation. Experimental results demonstrate that our method achieves the highest accuracy on the test set. Code will be released at https://github.com/Brickzhuantou/MonoBEVseg.
Jiangtong Zhu, Yibo Yuan, Zhuo Yin, Shizhen Li, Jianwu Fang, Jianru Xue
IV7
2024 IC-Mapper: Instance-Centric Spatio-Temporal Modeling for Online Vectorized Map Construction
Jiangtong Zhu, Yinan Shi, Jianwu Fang, Jianru Xue
ACM Multimedia5
2024 Fusing crops representation into snippet via mutual learning for weakly supervised surveillance anomaly detection
abstract
Abstract In recent years, the challenge of detecting anomalies in real‐world surveillance videos using weakly supervised data has emerged. Traditional methods, utilising multi‐instance learning (MIL) with video snippets, struggle with background noise and tend to overlook subtle anomalies. To tackle this, the authors propose a novel approach that crops snippets to create multiple instances with less noise, separately evaluates them and then fuses these evaluations for more precise anomaly detection. This method, however, leads to higher computational demands, especially during inference. Addressing this, our solution employs mutual learning to guide snippet feature training using these low‐noise crops. The authors integrate multiple instance learning (MIL) for the primary task with snippets as inputs and multiple‐multiple instance learning (MMIL) for an auxiliary task with crops during training. The authors’ approach ensures consistent multi‐instance results in both tasks and incorporates a temporal activation mutual learning module (TAML) for aligning temporal anomaly activations between snippets and crops, improving the overall quality of snippet representations. Additionally, a snippet feature discrimination enhancement module (SFDE) refines the snippet features further. Tested across various datasets, the authors’ method shows remarkable performance, notably achieving a frame‐level AUC of 85.78% on the UCF‐Crime dataset, while reducing computational costs.
Bohua Zhang, Jianru Xue
IET Comput. Vis.2
2024 HDF-Net: Capturing Homogeny Difference Features to Localize the Tampered Image
abstract
Modern image editing software enables anyone to alter the content of an image to deceive the public, which can pose a security hazard to personal privacy and public safety. The detection and localization of image tampering is becoming an urgent issue to be addressed. We have revealed that the tampered region exhibits homogenous differences (the changes in metadata organization form and organization structure of the image) from the real region after manipulations such as splicing, copy-move, and removal. Therefore, we propose a novel end-to-end network named HDF-Net to extract these homogeny difference features for precise localization of tampering artifacts. The HDF-Net is composed of RGB and SRM dual-stream networks, including three complementary modules, namely the suspicious tampering-artifact prominent (STP) module, the fine tampering-artifact salient (FTS) module, and the tampering-artifact edge refined (TER) module. We utilize the fully attentional block (FLA) to enhance the characterization ability of homogeny difference features extracted by each module and preserve the specifics of tampering artifacts. These modules are gradually merged according to the strategy of "coarse-fine-finer", which significantly improves the localization accuracy and edge refinement. Extensive experiments demonstrate that HDF-Net performs better than state-of-the-art tampering localization models on five benchmarks, achieving satisfactory generalization and robustness.
Ruidong Han, Ningning Bai, Jianpeng Hou, Jianru Xue
IEEE Trans. Pattern Anal. Mach. Intell.6
2024 Vision-Based Traffic Accident Detection and Anticipation: A Survey
abstract
Traffic accident detection and anticipation is an obstinate road safety problem and painstaking efforts have been devoted. With the rapid growth of video data, Vision-based Traffic Accident Detection and Anticipation (named Vision-TAD and Vision-TAA) become the last one-mile problem for safe driving and surveillance safety. However, the long-tailed, unbalanced, highly dynamic, complex, and uncertain properties of traffic accidents form the Out-of-Distribution (OOD) feature for Vision-TAD and Vision-TAA. Current AI development may focus on these OOD but important problems. What has been done for Vision-TAD and Vision-TAA? What direction we should focus on in the future for this problem? A comprehensive survey is important. We present the first survey on Vision-TAD in the deep learning era and the first-ever survey for Vision-TAA. The pros and cons of each research prototype are discussed in detail during the investigation. In addition, we also provide a critical review of 31 publicly available benchmarks and related evaluation metrics. Through this survey, we want to spawn new insights and open possible trends for Vision-TAD and Vision-TAA tasks.
Jianwu Fang, Jiahuan Qiao, Jianru Xue, Zhengguo Li
IEEE Trans. Circuits Syst. Video Technol.3
2024 Behavioral Intention Prediction in Driving Scenes: A Survey
abstract
In driving scenes, road agents often engage in frequent interaction and strive to understand their surroundings. Ego-agent (each road agent itself) predicts what behavior will be engaged by other road users all the time and expects a shared and consistent understanding for safe movement. To achieve this, Behavioral Intention Prediction (BIP) simulates such a human consideration process to anticipate specific behaviors, and the rapid development of BIP inevitably leads to new issues and challenges. To catalyze future research, this work provides a comprehensive review of BIP from the available datasets, key factors, challenges, pedestrian-centric and vehicle-centric BIP approaches, and BIP-aware applications. The investigation reveals that data-driven deep learning approaches have become the primary pipelines, while the behavioral intention types are still limited in most current datasets and methods (e.g., Crossing (C) and Not Crossing (NC) for pedestrians and Lane Changing (LC) for vehicles) in this field. In addition, current research on BIP in safe-critical scenarios (e.g., near-crashing situations) is limited. Through this investigation, we identify open issues in behavioral intention prediction and suggest possible insights for future research.
Jianwu Fang, Jianru Xue, Tat-Seng Chua
IEEE Trans. Intell. Transp. Syst.3
2024 A Mutually Textual and Visual Refinement Network for Image-Text Matching
abstract
Image-text matching is vital important in the field of multi-modal intelligence. Recently, it is advocated in a way that decomposes images and texts into local fragments and followed by region-word aligning. As a result, the image-text relevance score is given by aggregating semantic similarities between matched region-word pairs. Despite effectiveness, this strategy fails to express data relations exactly. From the perspective of the text side, text words decomposed from a concise language sentence usually have limited contextual information, which can result in semantic identical but actually false text-region alignments. From the perspective of the image side, semantic ambiguity that multiple objects share the same semantic meaning can further exacerbate this problem. In this manuscript, we introduce a mutually Textual and Visual Refinement Network (TVRN), to tackle the inaccurate cross-modal alignment problem. In a nutshell, TVRN improves inter-modal matching by improving contextual information in sentences meanwhile reduces semantic ambiguity in images to capture the maximized relevant relations. More specifically, we develop a new module that integrates visual contextual clues into the text modality to generate informational text features with richer geometric contexts. Mutually, we further design a semantic alignment enhancement module that leverages consensus affinity of local image and text features to guide deeper semantic image embedding with the supervision of global image vectors. At the image-text matching stage, similarities at the local and global levels are integrated to capture coarse-grained and fine-grained interactions between vision and language. A large number of experiments on Flickr30K and MS-COCO benchmarks demonstrate that TVRN is superior to existing methods.
Shanmin Pang, Yueyang Zeng, Jianru Xue
IEEE Trans. Multim.4
2023 FEND: A Future Enhanced Distribution-Aware Contrastive Learning Framework for Long-Tail Trajectory Prediction
abstract
Predicting the future trajectories of the traffic agents is a gordian technique in autonomous driving. However, trajectory prediction suffers from data imbalance in the prevalent datasets, and the tailed data is often more complicated and safety-critical. In this paper, we focus on dealing with the long-tail phenomenon in trajectory prediction. Previous methods dealing with long-tail data did not take into account the variety of motion patterns in the tailed data. In this paper, we put forward a future enhanced contrastive learning framework to recognize tail trajectory patterns and form a feature space with separate pattern clusters. Furthermore, a distribution aware hyper predictor is brought up to better utilize the shaped feature space. Our method is a model-agnostic framework and can be plugged into many well-known baselines. Experimental results show that our framework outperforms the state-of-the-art long-tail prediction method on tailed samples by 9.5% on ADE and 8.5% on FDE, while maintaining or slightly improving the averaged performance. Our method also surpasses many long-tail techniques on trajectory prediction task.
Pu Zhang 0001, Lei Bai 0001, Jianru Xue
CVPR4
2023 CalibDepth: Unifying Depth Map Representation for Iterative LiDAR-Camera Online Calibration
abstract
LiDAR-Camera online calibration is of great significance for building a stable autonomous driving perception system. For online calibration, a key challenge lies in constructing a unified and robust representation between multi-modal sensor data. Most methods extract features manually or implicitly with an end-to-end deep learning method. The former suffers poor robustness, while the latter has poor interpretability. In this paper, we propose CalibDepth, which uses depth maps as the unified representation for image and LiDAR point cloud. CalibDepth introduces a sub-network for monocular depth estimation to assist online calibration tasks. To further improve the performance, we regard online calibration as a sequence prediction problem, and introduce global and local losses to optimize the calibration results. CalibDepth shows excellent performance in different experimental setups. Code is open-sourced at https://github.com/Brickzhuantou/CalibDepth.
Jiangtong Zhu, Jianru Xue, Pu Zhang 0001
ICRA2
2023 Weakly-supervised anomaly detection with a Sub-Max strategy
Bohua Zhang, Jianru Xue
Neurocomputing2
2023 Towards Trajectory Forecasting From Detection
abstract
Trajectory forecasting for traffic participants (e.g., vehicles) is critical for autonomous platforms to make safe plans. Currently, most trajectory forecasting methods assume that object trajectories have been extracted and directly develop trajectory predictors based on the ground truth trajectories. However, this assumption does not hold in practical situations. Trajectories obtained from object detection and tracking are inevitably noisy, which could cause serious forecasting errors to predictors built on ground truth trajectories. In this paper, we propose to predict trajectories directly based on detection results without relying on explicitly formed trajectories. Different from traditional methods which encode the motion cues of an agent based on its clearly defined trajectory, we extract the motion information only based on the affinity cues among detection results, in which an affinity-aware state update mechanism is designed to manage the state information. In addition, considering that there could be multiple plausible matching candidates, we aggregate the states of them. These designs take the uncertainty of association into account which relax the undesirable effect of noisy trajectory obtained from data association and improve the robustness of the predictor. Extensive experiments validate the effectiveness of our method and its generalization ability to different detectors or forecasting schemes.
Pu Zhang 0001, Lei Bai 0001, Jianwu Fang, Jianru Xue, Nanning Zheng 0001, Wanli Ouyang
IEEE Trans. Pattern Anal. Mach. Intell.5
2023 Tensor-Based Incomplete Multi-View Clustering With Low-Rank Data Reconstruction and Consistency Guidance
abstract
We propose a new approach, called Tensor-based Incomplete Multi-view Clustering with Low-rank data Reconstruction and Consistency guidance (TIMC-RC), to perform clustering on multi-view data with missing views. Existing methods usually leverage original incomplete data to explore the partial correlations among multiple views, and do not make sufficient use of both consistent and complementary information across views. To explore the full information of missing and available views, TIMC-RC introduces low-rank data reconstruction and consistency view establishment. Specifically, 1) it adopts a low-rank constraint to reconstruct data representations so as to reduce the negative effect of missing data and obtain more reasonable data representations. 2) It builds a new consistency view by self-representation matrices and therefore explores the consistent correlation of different views. 3) It formalizes view-specific self-representation matrices and the consistent matrix as a tensor and utilizes the tensor singular value decomposition-based nuclear norm to enhance the consistency and complementarity of multi-view representations. Experiments conducted on eight benchmarks verify the effectiveness and advancement of the proposed TIMC-RC.
Wenyu Hao, Shanmin Pang, Xiuxiu Bai, Jianru Xue
IEEE Trans. Circuits Syst. Video Technol.4
2023 FCD-Net: Learning to Detect Multiple Types of Homologous Deepfake Face Images
abstract
With the rapid development of artificial intelligence technology, a variety of GAN generated deepfake face images/videos have emerged endlessly. The abuse of deepfake has brought serious negative effects to many industries. Therefore, there is an urgent need to develop advanced methods to combat the abuse of deepfake. As far as we know, there are almost no techniques that can distinguish multiple types of homologous deepfake face images. In this study, we propose a method based on the multi-classification task to address this issue. The proposed method relies on a novel network framework named FCD-Net that consists of the facial synaptic saliency module (FSS), the contour detail feature extraction module (CDFE), and the distinguishing feature fusion module (DFF). Utilizing this method, the imperceptible features introduced by deepfake can be exposed, and the differences caused by different types of deepfake can be distinguished, even if deepfake images are homologous. To test the proposed method and compare it with other SOTA methods, we establish a new homologous dataset named HDFD that contains real face images, entire face synthesis images, face swap images, and facial attribute manipulation images. Among them, the three types of deepfake images are all generated from the same real face images through different deepfake techniques. Abundant experiment results demonstrate that the proposed method has a high-level detection accuracy and relatively strong robustness against content-preserving manipulations. Moreover, the generalization of our method is superior to other SOTA methods.
Ruidong Han, Ningning Bai, Zinian Liu, Jianru Xue
IEEE Trans. Inf. Forensics Secur.6
2023 Heterogeneous Trajectory Forecasting via Risk and Scene Graph Learning
abstract
Heterogeneous trajectory forecasting is critical for intelligent transportation systems, but it is challenging because of the difficulty of modeling the complex interaction relations among the heterogeneous road agents as well as their agent-environment constraints. In this work, we propose a risk and scene graph learning method for trajectory forecasting of heterogeneous road agents, which consists of a Heterogeneous Risk Graph (HRG) and a Hierarchical Scene Graph (HSG) from the aspects of agent category and their movable semantic regions. HRG groups each kind of road agent and calculates their interaction adjacency matrix based on an effective collision risk metric. HSG of the driving scene is modeled by inferring the relationship between road agents and road semantic layout aligned by the road scene grammar. Based on this formulation, we can obtain effective trajectory forecasting in driving situations, and comparable performance to other state-of-the-art approaches is presented by extensive experiments on the nuScenes, ApolloScape, and Argoverse datasets.
Jianwu Fang, Pu Zhang 0001, Hongkai Yu, Jianru Xue
IEEE Trans. Intell. Transp. Syst.5
2022 Stochastic Navigation Command Matching for Imitation Learning of a Driving Policy
Xiangning Meng, Jianru Xue, Gengxin Li, Mengsen Wu
PRCV (3)2
2022 Human - machine augmented intelligence: research and applications
Jianru Xue, Lingxi Li 0001, Junping Zhang
Frontiers Inf. Technol. Electron. Eng.1
2022 Tensor-based multi-view clustering with consistency exploration and diversity regularization
Wenyu Hao, Shanmin Pang, Bo Yang 0041, Jianru Xue
Knowl. Based Syst.4
2022 Social-Aware Pedestrian Trajectory Prediction via States Refinement LSTM
abstract
In the task of pedestrian trajectory prediction, social interaction could be one of the most complicated factors since it is difficult to be interpreted through simple rules. Recent studies have shown a great ability of LSTM networks in learning social behaviors from datasets, e.g., introducing LSTM hidden states of the neighbors at the last time step into LSTM recursion. However, those methods depend on previous neighboring features which lead to a delayed observation. In this paper, we propose a data-driven states refinement LSTM network (SR-LSTM) to enable the utilization of the current intention of neighbors through a message passing framework. Moreover, the model performs in the form of self-updating by jointly refining the current states of all participants, rather than an input-output mechanism served by feature concatenation. In the process of states refinement, a social-aware information selection module consisting of an element-wise motion gate and a pedestrian-wise attention is designed to serve as the guidance of the message passing process. Considering the pedestrian walking space as a graph where each pedestrian is a node and each pedestrian pair with an edge, spatial-edge LSTMs are further exploited to enhance the model capacity, where two kinds of LSTMs interact with each other so that states of them are interactively refined. Experimental results on four widely used pedestrian trajectory datasets, ETH, UCY, PWPD, and NYGC demonstrate the effectiveness of the proposed model.
Pu Zhang 0001, Jianru Xue, Pengfei Zhang 0005, Nanning Zheng 0001, Wanli Ouyang
IEEE Trans. Pattern Anal. Mach. Intell.2
2022 Coarse-to-fine-grained method for image splicing region detection
Jinjin Lei, Jianru Xue
Pattern Recognit.6
2022 Traffic Accident Detection via Self-Supervised Consistency Learning in Driving Scenarios
abstract
With the rapid progress of autonomous driving and advanced driver assistance systems, there are growing efforts to promote their safety in natural driving scenarios, especially for the detection of the traffic accidents. However, because of the dynamic camera motion and complex scene in driving situations, traffic accident detection is still challenging. In this work, we aim to give the ability of Traffic Accident Detection for driving systems by proposing a Self-Supervised Consistency learning framework, termed as SSC-TAD, that involves the appearance, motion, and context consistency learning. The key formulation is to find the inconsistency of video frames, object locations and the spatial relation structure of scene temporally between different frames captured by the dashcam videos. Within this field, different from the previous works which concentrate on predicting the future object locations or frames, we further focus on predicting the visual scene context in driving scenarios and detecting the traffic accident by considering the temporal frame consistency, temporal object location consistency, and the spatial-temporal relation consistency of road participants. In this work, this formulation is fulfilled by a collaborative multi-task consistency learning network and the visual scene context feature is represented by a graph convolution network. The superiority to the state-of-the-art is verified by exhaustive evaluations on two large scale datasets, i.e., the AnAn Accident Detection (A3D) dataset and DADA-2000 dataset collected recently.
Jianwu Fang, Jiahuan Qiao, Hongkai Yu, Jianru Xue
IEEE Trans. Intell. Transp. Syst.5
2022 DADA: Driver Attention Prediction in Driving Accident Scenarios
abstract
Driver attention prediction is becoming an essential research problem in human-like driving systems. This work makes an attempt to predict thedriverattention indrivingaccident scenarios (DADA). However, challenges tread on the heels of that because of the dynamic traffic scene, intricate and imbalanced accident categories. In this work, we design a semantic context induced attentive fusion network (SCAFNet). We first segment the RGB video frames into the images with different semantic regions (i.e., semantic images), where each region denotes one semantic category of the scene (e.g., road, trees, etc.), and learn the spatio-temporal features of RGB frames and semantic images in two parallel paths simultaneously. Then, the learned features are fused by an attentive fusion network to find the semantic-induced scene variation in driver attention prediction. The contributions are three folds. 1) With the semantic images, we introduce their semantic context features and verify the manifest promotion effect for helping the driver attention prediction, where the semantic context features are modeled by a graph convolution network (GCN) on semantic images; 2) We fuse the semantic context features of semantic images and the features of RGB frames in an attentive strategy, and the fused details are transferred over frames by a convolutional LSTM module to obtain the attention map of each video frame with the consideration of historical scene variation in driving situations; 3) The superiority of the proposed method is evaluated on our previously collected dataset (named as DADA-2000) and two other challenging datasets with state-of-the-art methods.
Jianwu Fang, Dingxin Yan, Jiahuan Qiao, Jianru Xue, Hongkai Yu
IEEE Trans. Intell. Transp. Syst.4
2022 An Adaptive Invariant EKF for Map-Aided Localization Using 3D Point Cloud
abstract
In map-aided localization using 3D point cloud sets, the poses estimated by 3D Registration Algorithms (3DRAs) are typically fused with other sensor data via the Extended Kalman Filter (EKF) to obtain reliable and smooth results. However, the challenges of this combined method are: 1) The linearization process of EKF may cause errors and singularities due to the state defined by a 6D pose. 2) The results of 3DRA as the measurements of EKF may cause errors in residual calculation. 3) The approach relies heavily on 3DRA to overcome the effects of dynamic scenes. This paper proposes an adaptive localization framework based on Invariant Extended Kalman Filter (Invariant EKF), in which the Lie Group is introduced to define the state. In this framework, the points of the raw point cloud set are the measurements of the filter, and the 3DRA is only employed for data association between the raw 3D point cloud set and the 3D point cloud map. Then, a Concentric Ring Model (CRM) is proposed to reduce the influence of dynamic objects, which can adaptively estimate the covariance of each observed point via Gaussian Process Regression (GPR). Besides, the CRM considers Sensor Measurement Noise (SMN) and Sensor Vibration Noise (SVN). The performance of the proposed framework is evaluated on the KITTI dataset and our dataset. The experimental results show that the proposed method is superior to other state-of-the-art methods, and the CRM can achieve more accurate measurement than ever before, especially in high-dynamic scenes.
Zhongxing Tao, Jianru Xue, Di Wang 0028, Gengxin Li, Jianwu Fang
IEEE Trans. Intell. Transp. Syst.2
2021 PCRLaneNet: Lane Marking Detection via Point Coordinate Regression
abstract
Lane detection is one of the most important task in autonomous driving. While the semantic segmentation based method is widely explored and recognized in recent decade, some post-processing are required to estimate the exact location of the predicted lane markings and can be easily failed in complex scenarios. To tackle these limitations, this paper proposes a novel lane detection network named PCRLaneNet. Firstly, we use a fully convolutional network to predict the coordinates of lane marking points directly, which can better meet with the requirements of autonomous driving. Secondly, to take the fully advantage of the correlation of these lane marking points, a point feature fusion strategy is designed to fuse feature maps of the points on the same lane marking, which makes our method capable of handling challenging scenarios. Lastly, the robustness, accuracy and latency of the proposed method are extensively verified in two datasets (CULane and TuSimple).
Jianru Xue, Jian Dou, Di Wang 0028
IV2
2021 Perceptual hash-based coarse-to-fine grained image tampering forensics method
Chuntao Jiang, Jianru Xue
J. Vis. Commun. Image Represent.4
2021 An efficient batch images encryption method based on DNA encoding and PWLCM
Jinjin Lei, Jianru Xue
Multim. Tools Appl.5
2021 Video Frame Prediction by Deep Multi-Branch Mask Network
abstract
Future frame prediction in video is one of the most important problem in computer vision, and useful for a range of practical applications, such as intention prediction or video anomaly detection. However, this task is challenging because of the complex and dynamic evolution of scene. The difficulty of video frame prediction is to model the inherent spatio-temporal correlation between frames and pose an adaptive and flexible framework for large motion change or appearance variation. In this paper, we construct a deep multi-branch mask network (DMMNet) which adaptively fuses the advantages of optical flow warping and RGB pixel synthesizing methods, i.e., the common two kinds of approaches in this task. In the procedure of DMMNet, we add mask layer in each branch to adaptively adjust the magnitude range of estimated optical flow and the weight of predicted frames by optical flow warping and RGB pixel synthesizing, respectively. In other words, we provide a more flexible masking network for motion and appearance fusion on video frame prediction. Exhaustive experiments on Caltech pedestrian and UCF101 datasets show that the proposed model can obtain favorable video frame prediction performance compared with the state-of-the-art methods. In addition, we also put our model into the video anomaly detection problem, and the superiority is verified by the experiments on UCSD dataset.
Jianwu Fang, Hongke Xu, Jianru Xue
IEEE Trans. Circuits Syst. Video Technol.4
2021 Improve Regression Network on Depth Hand Pose Estimation With Auxiliary Variable
abstract
The regression based deep neural networks have achieved state-of-the-arts performance on depth 3D hand pose estimation task. This paper focuses on improving the regression mapping between features and pose joints. Inspired by the distribution modeling ability of Variational Autoencoders, we introduce an auxiliary variable into the regression network. During training, the auxiliary variable is modeled by an inference distribution that learns the underlying structural kinematics of human hand. Different with other regression methods on hand poses, our network estimates the pose joints from input depth features and the learned auxiliary variable as well. We show that by introducing the auxiliary variable, the regression is benefited from 1) regularization modeled by inference distribution; and 2) prior information carried by the auxiliary model. The effectiveness of the proposed regression method is evaluated with extensively self-comparative experiments and in comparison with other regression methods on hand pose datasets. The proposed network is easy to train in an end-to-end manner and can work with various feature extraction methods. We apply the proposed regression method to an existing hand pose estimation system, and improves the estimation accuracy by 18.35% and 16.65% on public hand pose datasets.
Ji'an Tao, Jianru Xue
IEEE Trans. Circuits Syst. Video Technol.4
2020 Semantics-Guided Neural Networks for Efficient Skeleton-Based Human Action Recognition
abstract
Skeleton-based human action recognition has attracted great interest thanks to the easy accessibility of the human skeleton data. Recently, there is a trend of using very deep feedforward neural networks to model the 3D coordinates of joints without considering the computational efficiency. In this paper, we propose a simple yet effective semantics-guided neural network (SGN) for skeleton-based action recognition. We explicitly introduce the high level semantics of joints (joint type and frame index) into the network to enhance the feature representation capability. In addition, we exploit the relationship of joints hierarchically through two modules, i.e., a joint-level module for modeling the correlations of joints in the same frame and a framelevel module for modeling the dependencies of frames by taking the joints in the same frame as a whole. A strong baseline is proposed to facilitate the study of this field. With an order of magnitude smaller model size than most previous works, SGN achieves the state-of-the-art performance on the NTU60, NTU120, and SYSU datasets.
Pengfei Zhang 0005, Cuiling Lan, Wenjun Zeng 0001, Junliang Xing, Jianru Xue, Nanning Zheng 0001
CVPR5
2020 Navigation Command Matching for Vision-based Autonomous Driving
abstract
Learning an optimal policy for autonomous driving task to confront with complex environment is a long- studied challenge. Imitative reinforcement learning is accepted as a promising approach to learn a robust driving policy through expert demonstrations and interactions with environments. However, this model utilizes non-smooth rewards, which have a negative impact on matching between navigation commands and trajectory (state-action pairs), and degrade the generalizability of an agent. Smooth rewards are crucial to discriminate actions generated from sub-optimal policy. In this paper, we propose a navigation command matching (NCM) model to address this issue. There are two key components in NCM, 1) a matching measurer produces smooth navigation rewards that measure matching between navigation commands and trajectory; 2) attention-guided agent performs actions given states where salient regions in RGB images (i.e. roadsides, lane markings and dynamic obstacles) are highlighted to amplify their influence on the final model. We obtain navigation rewards and store transitions to replay buffer after an episode, so NCM is able to discriminate actions generated from suboptimal policy. Experiments on CARLA driving benchmark show our proposed NCM outperforms previous state-of-the- art models on various tasks in terms of the percentage of successfully completed episodes. Moreover, our model improves generalizability of the agent and obtains good performance even in unseen scenarios.
Yuxin Pan, Jianru Xue, Pengfei Zhang 0005, Wanli Ouyang, Jianwu Fang, Xingyu Chen 0001
ICRA2
2020 Using Detection, Tracking and Prediction in Visual SLAM to Achieve Real-time Semantic Mapping of Dynamic Scenarios
abstract
In this paper, we propose a lightweight system, RDS-SLAM, based on ORB-SLAM2, which can accurately estimate poses and build semantic maps at object level for dynamic scenarios in real time using only one commonly used Intel Core i7 CPU. In RDS-SLAM, three major improvements, as well as major architectural modifications, are proposed to overcome the limitations of ORB-SLAM2. Firstly, it adopts a lightweight object detection neural network in key frames. Secondly, an efficient tracking and prediction mechanism is embedded into the system to remove the feature points belonging to movable objects in all incoming frames. Thirdly, a semantic octree map is built by probabilistic fusion of detection and tracking results, which enables a robot to maintain a semantic description at object level for potential interactions in dynamic scenarios. We evaluate RDS-SLAM in TUM RGB-D dataset, and experimental results show that RDS-SLAM can run with 30.3 ms per frame in dynamic scenarios using only an Intel Core i7 CPU, and achieves comparable accuracy compared with the state-of-the-art SLAM systems which heavily rely on both Intel Core i7 CPUs and powerful GPUs.
Xingyu Chen 0001, Jianru Xue, Jianwu Fang, Yuxin Pan, Nanning Zheng 0001
IV2
2020 An Efficient Sampling-Based Hybrid A* Algorithm for Intelligent Vehicles
abstract
In this paper, we propose an improved sampling-based hybrid A* (SBA*) algorithm for path planning of intelligent vehicles, which works efficiently in complex urban environments. Two main modifications are introduced into the traditional hybrid A* algorithm to improve its adaptivity in both structured and unstructured traffic scenes. Firstly, a hybrid potential field (HPF) model considering both traffic regulation and obstacle configuration is proposed to represent the vehicle's workspace, which is utilized as a heuristic function. Secondly, a set of directional motion primitives is generated by taking the prior topological structure of the workspace into account. The path planner using SBA* not only obeys traffic regulations in structured scenes but also is capable of exploring complex unstructured scenes rapidly. Finally, a post-optimization step is adopted to increase the feasibility of the path. The efficacy of the proposed algorithm is extensively validated and tested with an autonomous vehicle in real traffic scenes. The experimental results show that SBA* works well in complex urban environments.
Gengxin Li, Jianru Xue, Di Wang 0028, Zhongxing Tao, Nanning Zheng 0001
IV2
2020 Improving 3D Object Detection via Joint Attribute-oriented 3D Loss
abstract
3D object detection has become a hot topic in intelligent vehicle applications in recent years. Generally, deep learning has been the primary framework used in 3D object detection, and regression of the object location and classification of the objectness are the two indispensable components. In the process of training, the ℓn(n=1,2) and the focal loss are considered as the frequent solutions to minimize the regression and classification loss, respectively. However, there are two problems to be solved in the existing methods. For regression component, there is a gap between evaluation metrics, e.g., 3D Intersection over Union (IoU), and the traditional regression loss. As for the classification component, confidence score exists ambiguous due to the binary label assignment of target. To solve these problems, we propose a loss by jointing 3D IoU and other geometric attributes (named as jointed attribute-oriented 3D loss), which can be directly used in optimizing the regression component. In addition, the jointed attribute-oriented 3D loss can assign a soft label for supervising the training of the classification. By incorporating the proposed loss function into several state-of-the-art 3D object detection methods, the significant performance improvement has been achieved on the KITTI benchmark.
Jianru Xue, Jian Dou, Yuxin Pan, Jianwu Fang, Di Wang 0028, Nanning Zheng 0001
IV2
2020 Look, Listen and Infer
abstract
Inspired by the ability of human beings on recognizing the relations between visual scenes and sounds, many cross-modal learning methods have been developed for modeling images or videos and associated sounds. In this work, for the first time, a Look, Listen and Infer Network (LLINet) is proposed to learn a zero-shot model that can infer the relations of visual scenes and sounds from novel categories never appeared before. LLINet is mainly desired to qualify for two tasks, i.e., image-audio cross-modal retrieval and sound localization in each image. Towards this end, it is designed as a two-branch encoding network that builds a common space for images and audios. Besides, a cross-modal attention mechanism is proposed in LLINet to localize sound objects. To evaluate LLINet, a new data set, named INSTRUMENT-32CLASS, is collected in this work. Besides zero-shot cross-modal retrieval and sound localization, a zero-shot image recognition task based on sounds is also conducted on this database. All experimental results on these tasks demonstrate the effectiveness of LLINet, indicating that zero-shot learning for visual scenes and sounds is feasible. The project page for LLINet is available at https://llinet.github.io/.
Ruijian Jia, Shanmin Pang, Jihua Zhu, Jianru Xue
ACM Multimedia5
2020 Vehicle re-identification in tunnel scenes via synergistically cascade forests
Rixing Zhu, Jianwu Fang, Qi Wang 0009, Hongke Xu, Jianru Xue, Hongkai Yu
Neurocomputing6
2020 Time Series Prediction Method Based on Variant LSTM Recurrent Neural Network
Jiaojiao Hu, Depeng Zhang, Jianru Xue
Neural Process. Lett.6
2020 Image alignment based perceptual image hash for content authentication
Xiaorui Zhou, Bingchao Xu, Jianru Xue
Signal Process. Image Commun.5
2020 EleAtt-RNN: Adding Attentiveness to Neurons in Recurrent Neural Networks
abstract
Recurrent neural networks (RNNs) are capable of modeling temporal dependencies of complex sequential data. In general, current available structures of RNNs tend to concentrate on controlling the contributions of current and previous information. However, the exploration of different importance levels of different elements within an input vector is always ignored. We propose a simple yet effective Element-wise-Attention Gate (EleAttG), which can be easily added to an RNN block (e.g. all RNN neurons in an RNN layer), to empower the RNN neurons to have attentiveness capability. For an RNN block, an EleAttG is used for adaptively modulating the input by assigning different levels of importance, i.e., attention, to each element/dimension of the input. We refer to an RNN block equipped with an EleAttG as an EleAtt-RNN block. Instead of modulating the input as a whole, the EleAttG modulates the input at fine granularity, i.e., element-wise, and the modulation is content adaptive. The proposed EleAttG, as an additional fundamental unit, is general and can be applied to any RNN structures, e.g., standard RNN, Long Short-Term Memory (LSTM), or Gated Recurrent Unit (GRU). We demonstrate the effectiveness of the proposed EleAtt-RNN by applying it to different tasks including the action recognition, from both skeleton-based data and RGB videos, gesture recognition, and sequential MNIST classification. Experiments show that adding attentiveness through EleAttGs to RNN blocks significantly improves the power of RNNs.
Pengfei Zhang 0005, Jianru Xue, Cuiling Lan, Wenjun Zeng 0001, Zhanning Gao, Nanning Zheng 0001
IEEE Trans. Image Process.2
2019 A Novel Hierarchical Collaborative Method Based on Multi-objective Optimization for Modularization of Product Platform
abstract
The modular design of product platform becomes an indispensable aspect in manufacturing industry with increasingly tailored requirements from customers. Amongst all components in this platform, the modularization mechanism addresses key issue in the design process of manufacturing product. However, the traditional modular methods only solve the modularization with single one or several objectives. In this paper, a novel hierarchical collaborative method based on multi-objective optimization is developed for modularization of product platform, which is integrated various of modular standards, including the product structure, product function, design of products, assembly of products and so on, and trades off the contradiction of these standards to achieve the optimal modularization with multiple hierarchical objectives. Moreover, for solutions to this method, a hierarchical collaborative multi-objective optimization algorithm, the Hierarchical Non-dominated Sorting Genetic Algorithm (H-NSGA), is developed to optimize the modularization with multilevel objectives. Furthermore, an industrial case illustrates the details of the proposed model and algorithm. Finally, the experimental comparison demonstrates that the hierarchical collaborative method and the hierarchical multi-objective optimization algorithm can achieve efficient modular solutions with multiple hierarchical objectives.
Qihao Wan, Jianru Xue, Chiyuan Zhang, Heming Zhang 0001
CSCWD3
2019 SR-LSTM: State Refinement for LSTM Towards Pedestrian Trajectory Prediction
abstract
In crowd scenarios, reliable trajectory prediction of pedestrians requires insightful understanding of their social behaviors. These behaviors have been well investigated by plenty of studies, while it is hard to be fully expressed by hand-craft rules. Recent studies based on LSTM networks have shown great ability to learn social behaviors. However, many of these methods rely on previous neighboring hidden states but ignore the important current intention of the neighbors. In order to address this issue, we propose a data-driven state refinement module for LSTM network (SR-LSTM), which activates the utilization of the current intention of neighbors, and jointly and iteratively refines the current states of all participants in the crowd through a message passing mechanism. To effectively extract the social effect of neighbors, we further introduce a social-aware information selection mechanism consisting of an element-wise motion gate and a pedestrian-wise attention to select useful message from neighboring pedestrians. Experimental results on two public datasets, i.e. ETH and UCY, demonstrate the effectiveness of our proposed SR-LSTM and we achieve state-of-the-art results.
Pu Zhang 0001, Wanli Ouyang, Pengfei Zhang 0005, Jianru Xue, Nanning Zheng 0001
CVPR4
2019 Deep Conditional Variational Estimation for Depth-Based Hand Poses
abstract
We propose a novel and effective approach for 3D hand pose estimation on single depth image. Instead of doing deterministic regression from depth images, our model focuses on learning a latent distribution to model the high dimensional space of pose joints, which can also be interpreted as a kinematics model for human hands. Specifically, the proposed network combines the framework of conditional variational autoencoder which learns an encoder and a decoder with standard convolutional network. The encoder models the latent variable as a prior or a regularization for the pose joints. Then probabilistic inference is performed by the decoder to generate the output prediction conditioned on input depth images. In addition, we introduce a pool-convolution module to improve the localization regression of the network. The architecture can be trained end-to-end. In experiments, we demonstrate the effectiveness of our proposed approach in comparison to various state-of-art holistic regression approaches.
Yinqi Li 0001, Ji'an Tao, Jianru Xue
FG5
2019 Small Object Detection on Road by Embedding Focal-Area Loss
Jianwu Fang, Jian Dou, Jianru Xue
ICIG (1)4
2019 SEG-VoxelNet for 3D Vehicle Detection from RGB and LiDAR Data
abstract
This paper proposes a SEG-VoxelNet that takes RGB images and LiDAR point clouds as inputs for accurately detecting 3D vehicles in autonomous driving scenarios, which for the first time introduces semantic segmentation technique to assist the 3D LiDAR point cloud based detection. Specifically, SEG-VoxelNet is composed of two sub-networks: an image semantic segmentation network (SEG-Net) and an improved-VoxelNet. The SEG-Net generates the semantic segmentation map which represents the probability of the category for each pixel. The improved-VoxelNet is capable of effectively fusing point cloud data with image semantic feature and generating accurate 3D bounding boxes of vehicles. Experiments on the KITTI 3D vehicle detection benchmark show that our approach outperforms the methods of state-of-the-art.
Jian Dou, Jianru Xue, Jianwu Fang
ICRA2
2019 BLVD: Building A Large-scale 5D Semantics Benchmark for Autonomous Driving
abstract
In autonomous driving community, numerous benchmarks have been established to assist the tasks of 3D/2D object detection, stereo vision, semantic/instance segmentation. However, the more meaningful dynamic evolution of the surrounding objects of ego-vehicle is rarely exploited, and lacks a large-scale dataset platform. To address this, we introduce BLVD, a large-scale 5D semantics benchmark which does not concentrate on the static detection or semantic/instance segmentation tasks tackled adequately before. Instead, BLVD aims to provide a platform for the tasks of dynamic 4D (3D+temporal) tracking, 5D (4D+interactive) interactive event recognition and intention prediction. This benchmark will boost the deeper understanding of traffic scenes than ever before. We totally yield 249, 129 3D annotations, 4, 902 independent individuals for tracking with the length of overall 214, 922 points, 6, 004 valid fragments for 5D interactive event recognition, and 4, 900 individuals for 5D intention prediction. These tasks are contained in four kinds of scenarios depending on the object density (low and high) and light conditions (daytime and nighttime). The benchmark can be downloaded from our project site https://github.com/VCCIV/BLVD/.
Jianru Xue, Jianwu Fang, Bohua Zhang, Pu Zhang 0001, Jian Dou
ICRA1
2019 Precise Correntropy-based 3D Object Modelling With Geometrical Traffic Prior
abstract
Robust 3D perception using LiDAR is of prime importance for robotics, and its fundamental core lies in precise object modelling resisting to noise and outliers. In this paper, a precise 3D object modelling algorithm is designed especially for the intelligent vehicles. The proposed algorithm is advantageous by leveraging the crucial traffic geometrical prior of road surface profile, and both the noise and outliers are elegantly handled by robust correntropy-based metric. More specifically, the road surface correction (RSC) method transforms each individual LiDAR measurement from its locally planar road surface to a globally ideal plane. This procedure essentially guarantees the reduction of vehicle's motion from arbitrary 3D motion to physically feasible 2D motion. To deal with the noise and outliers, a correntropy-based multi-frame matching (CorrMM) algorithm is proposed which has a robust objective function with respect to point-to-plane residual error. An efficient solver inspired by M-estimator and retraction technique on Lie group is developed, which elegantly converts the optimization of highly non-linear objective function into a simple quadratic programming (QP) problem. Extensive experimental results validate that the proposed algorithm attains more crisper 3D object models than several state-of-the-art algorithms on a challenging real traffic dataset.
Di Wang 0028, Jianru Xue, Yinghan Jin, Nanning Zheng 0001, Masayoshi Tomizuka
IROS2
2019 View Adaptive Neural Networks for High Performance Skeleton-Based Human Action Recognition
abstract
Skeleton-based human action recognition has recently attracted increasing attention thanks to the accessibility and the popularity of 3D skeleton data. One of the key challenges in action recognition lies in the large variations of action representations when they are captured from different viewpoints. In order to alleviate the effects of view variations, this paper introduces a novel view adaptation scheme, which automatically determines the virtual observation viewpoints over the course of an action in a learning based data driven manner. Instead of re-positioning the skeletons using a fixed human-defined prior criterion, we design two view adaptive neural networks, i.e., VA-RNN and VA-CNN, which are respectively built based on the recurrent neural network (RNN) with the Long Short-term Memory (LSTM) and the convolutional neural network (CNN). For each network, a novel view adaptation module learns and determines the most suitable observation viewpoints, and transforms the skeletons to those viewpoints for the end-to-end recognition with a main classification network. Ablation studies find that the proposed view adaptive models are capable of transforming the skeletons of various views to much more consistent virtual viewpoints. Therefore, the models largely eliminate the influence of the viewpoints, enabling the networks to focus on the learning of action-specific features and thus resulting in superior performance. In addition, we design a two-stream scheme (referred to as VA-fusion) that fuses the scores of the two networks to provide the final prediction, obtaining enhanced performance. Moreover, random rotation of skeleton sequences is employed to improve the robustness of view adaptation models and alleviate overfitting during training. Extensive experimental evaluations on five challenging benchmarks demonstrate the effectiveness of the proposed view-adaptive networks and superior performance over state-of-the-art approaches.
Pengfei Zhang 0005, Cuiling Lan, Junliang Xing, Wenjun Zeng 0001, Jianru Xue, Nanning Zheng 0001
IEEE Trans. Pattern Anal. Mach. Intell.5
2019 Unifying Sum and Weighted Aggregations for Efficient Yet Effective Image Representation Computation
abstract
Embedding and aggregating a set of local descriptors (e.g. SIFT) into a single vector is normally used to represent images in image search. Standard aggregation operations include sum and weighted aggregations. While showing high efficiency, sum aggregation lacks discriminative power. In contrast, weighted aggregation shows promising retrieval performance but suffers extremely high time cost. In this work, we present a general mixed aggregation method that unifies sum and weighted aggregation methods. Owing to its general formulation, our method is able to balance the trade-off between retrieval quality and image representation efficiency. Additionally, to improve query performance, we propose computing multiple weighting coefficients rather than one for each to be aggregated vector by partitioning them into several components with negligible computational cost. Extensive experimental results on standard public image retrieval benchmarks demonstrate that our aggregation method achieves state-of-the-art performance while showing over ten times speedup over baselines.
Shanmin Pang, Jianru Xue, Jihua Zhu, Li Zhu 0003, Qi Tian 0001
IEEE Trans. Image Process.2
2019 Deep Feature Aggregation and Image Re-Ranking With Heat Diffusion for Image Retrieval
abstract
Image retrieval based on deep convolutional features has demonstrated state-of-the-art performance in popular benchmarks. In this paper, we present a unified solution to address deep convolutional feature aggregation and image re-ranking by simulating the dynamics of heat diffusion. A distinctive problem in image retrieval is that repetitive or bursty features tend to dominate final image representations, resulting in representations less distinguishable. We show that by considering each deep feature as a heat source, our unsupervised aggregation method is able to avoid over-representation of bursty features. We additionally provide a practical solution for the proposed aggregation method and further show the efficiency of our method in experimental evaluation. Inspired by the aforementioned deep feature aggregation method, we also propose a method to re-rank a number of top ranked images for a given query image by considering the query as the heat source. Finally, we extensively evaluate the proposed approach with pre-trained and fine-tuned deep networks on common public benchmarks and show superior performance compared to previous work.
Shanmin Pang, Jianru Xue, Jihua Zhu, Vicente Ordonez
IEEE Trans. Multim.3
2019 Improving Object Retrieval Quality by Integration of Similarity Propagation and Query Expansion
abstract
Re-ranking is an essential step for accurate image retrieval, due to its well-known power in performance improvement. Although numerous works have been proposed for re-ranking, many of them are only customized for a certain image representation model. In contrast to most existing techniques, we develop generalized re-ranking algorithms that are applicable to different kinds of image encodings in this paper. We first employ a quite successful theory of similarity propagation to reconstruct vectors of a query and its top ranked images and, subsequently, get a re-ranked list by comparing the new image vectors. Furthermore, considering that the just mentioned strategy is directly compatible with query expansion and, thus, in order to leverage advantages of this milestone, we then propose integrating them into a unified framework for maximizing re-ranking benefits. Our re-ranking algorithms are memory and computation efficient, and experimental results on benchmark datasets demonstrate that they compare favorably with the state of the art. Our code is available at https://github.com/MaJinWakeUp/rerank.
Shanmin Pang, Jihua Zhu, Jianru Xue, Qi Tian 0001
IEEE Trans. Multim.4
2018 Adding Attentiveness to the Neurons in Recurrent Neural Networks
Pengfei Zhang 0005, Jianru Xue, Cuiling Lan, Wenjun Zeng 0001, Zhanning Gao, Nanning Zheng 0001
ECCV (9)2
2018 Accurate Mix-Norm-Based Scan Matching
abstract
Highly accurate mapping and localization is of prime importance for mobile robotics, and its core lies in efficient scan matching. Previous research are focusing on designing a robust objective function and the residual error distribution is often ignored or simply assumed as unitary or mixture of simple distributions. In this paper, a mixture of exponential power (MoEP) distributions is proposed to approximate the residual error distribution. The objective function induced by MoEP-based residual error modelling ensembles a mix-norm-based scan matching (MiNoM), which enhances the matching accuracy and convergence characteristic. Both the parameters of transformation (rotation and translation) and residual error distribution are estimated efficiently via an EM-like algorithm. The optimization of MiNoM is iteratively achieved via two phases: An on-line parameter learning (OPL) phase to learn residual error distribution for better representation according to the likelihood field model (LFM), and an iteratively reweighted least squares (IRLS) phase to attain transformation for accuracy and efficiency. Extensive experimental results validate that the proposed MiNoM out-performs several state-of-the-art scan matching algorithms in both convergence characteristic and matching accuracy.
Di Wang 0028, Jianru Xue, Zhongxing Tao, Dixiao Cui, Shaoyi Du, Nanning Zheng 0001
IROS2
2018 Accurate Localization in Underground Garages via Cylinder Feature based Map Matching
abstract
Autonomous driving in underground garages usually utilizes a 2D/3D occupancy map for localization. However, the real scene is changing, and may not be consistent with the map. Vehicles and other objects not contained in the map are considered as obstacles, which increase the difficulty of localization and affect the accuracy of result. In this paper, we propose a cylinder rotational projection statistics (Cy-RoPS) feature descriptor, which is a local surface feature descriptor to improve the accuracy of localization. The local surface feature motivated by RoPS feature is invariant to rotation of point set enclosed in a cylinder. We also propose to employ the local surface feature for localization in a real underground garage. The experimental results show that the proposed method is robust to dynamic obstacles in the underground garage, and has a higher accuracy in localization, compared with the state-of-the-art methods.
Zhongxing Tao, Jianru Xue, Di Wang 0028, Dixiao Cui, Shaoyi Du
Intelligent Vehicles Symposium2
2018 Precise Point Set Registration Using Point-to-Plane Distance and Correntropy for LiDAR Based Localization
abstract
In this paper, we propose a robust point set registration algorithm which combines correntropy and point-to-plane distance, which can register rigid point sets with noises and outliers. Firstly, as correntropy performs well in handling data with non-Gaussian noises, we introduce it to model rigid point set registration problem based on point-to-plane distance; Secondly, we propose an iterative algorithm to solve this problem, which repeats to compute correspondence and transformation parameters respectively in closed form solutions. Simulated experimental results demonstrate the high precision and robustness of the proposed algorithm. In addition, LiDAR based localization experiments on automated vehicle performs satisfactory for localization accuracy and time consumption.
Guanglin Xu, Shaoyi Du, Dixiao Cui, Sirui Zhang, Badong Chen, Xuetao Zhang 0001, Jianru Xue, Yue Gao 0002
Intelligent Vehicles Symposium7
2018 Temporality-enhanced knowledgememory network for factoid question answering
abstract
Question answering is an important problem that aims to deliver specific answers to questions posed by humans in natural language. How to efficiently identify the exact answer with respect to a given question has become an active line of research. Previous approaches in factoid question answering tasks typically focus on modeling the semantic relevance or syntactic relationship between a given question and its corresponding answer. Most of these models suffer when a question contains very little content that is indicative of the answer. In this paper, we devise an architecture named the temporality-enhanced knowledge memory network (TE-KMN) and apply the model to a factoid question answering dataset from a trivia competition called quiz bowl. Unlike most of the existing approaches, our model encodes not only the content of questions and answers, but also the temporal cues in a sequence of ordered sentences which gradually remark the answer. Moreover, our model collaboratively uses external knowledge for a better understanding of a given question. The experimental results demonstrate that our method achieves better performance than several state-of-the-art methods.
Xinyu Duan, Siliang Tang, Shengyu Zhang 0001, Yin Zhang 0006, Zhou Zhao 0001, Jianru Xue, Yueting Zhuang, Fei Wu 0001
Frontiers Inf. Technol. Electron. Eng.6
2018 Building discriminative CNN image representations for object retrieval using the replicator equation
Shanmin Pang, Jihua Zhu, Vicente Ordonez, Jianru Xue
Pattern Recognit.5
2018 Large-scale vocabularies with local graph diffusion and mode seeking
Shanmin Pang, Jianru Xue, Zhanning Gao, Lihong Zheng, Li Zhu 0003
Signal Process. Image Commun.2
2018 Data-Driven State-Increment Statistical Model and Its Application in Autonomous Driving
abstract
The aim of trajectory planning is to generate a feasible, collision-free trajectory to guide an autonomous vehicle from the initial state to the goal state safely. However, it is difficult to guarantee that the trajectory is feasible for the vehicle and the real path of the vehicle is collision-free when the vehicle follows the trajectory. In this paper, a state-increment statistical model (SISM) is proposed to describe the kinodynamic constraints of a vehicle by modeling the controller, the actuator, and the vehicle model jointly. The SISM consists of Gaussian distributions of lateral error increments in all state subspaces which are composed of the curvature radius, the velocity, and the lateral error. It is a data-driven modeling approach that can improve the SISM via increasing the number of samples of the increment-state, which is composed of the state and its corresponding increment of the lateral error. According to the SISM, the experience cost functions are designed to evaluate the trajectories for searching the best one with the lowest cost, and the real path can be predicted directly according to the planned trajectory and the vehicle state. The predicted path can be utilized effectually to evaluate the safety of the vehicle motion.
Chao Ma 0024, Jianru Xue, Yuehu Liu, Jing Yang 0014, Nanning Zheng 0001
IEEE Trans. Intell. Transp. Syst.2
2017 ER3: A Unified Framework for Event Retrieval, Recognition and Recounting
abstract
We develop a unified framework for complex event retrieval, recognition and recounting. The framework is based on a compact video representation that exploits the temporal correlations in image features. Our feature alignment procedure identifies and removes the feature redundancies across frames and outputs an intermediate tensor representation we call video imprint. The video imprint is then fed into a reasoning network, whose attention mechanism parallels that of memory networks used in language modeling. The reasoning network simultaneously recognizes the event category and locates the key pieces of evidence for event recounting. In event retrieval tasks, we show that the compact video representation aggregated from the video imprint achieves significantly better retrieval accuracy compared with existing methods. We also set new state of the art results in event recognition tasks with an additional benefit: The latent structure in our reasoning network highlights the areas of the video imprint and can be directly used for event recounting. As video imprint maps back to locations in the video frames, the network allows not only the identification of key frames but also specific areas inside each frame which are most influential to the decision process.
Zhanning Gao, Gang Hua 0001, Dongqing Zhang, Nebojsa Jojic, Le Wang 0003, Jianru Xue, Nanning Zheng 0001
CVPR6
2017 View Adaptive Recurrent Neural Networks for High Performance Human Action Recognition from Skeleton Data
abstract
Skeleton-based human action recognition has recently attracted increasing attention due to the popularity of 3D skeleton data. One main challenge lies in the large view variations in captured human actions. We propose a novel view adaptation scheme to automatically regulate observation viewpoints during the occurrence of an action. Rather than re-positioning the skeletons based on a human defined prior criterion, we design a view adaptive recurrent neural network (RNN) with LSTM architecture, which enables the network itself to adapt to the most suitable observation viewpoints from end to end. Extensive experiment analyses show that the proposed view adaptive RNN model strives to (1) transform the skeletons of various views to much more consistent viewpoints and (2) maintain the continuity of the action rather than transforming every frame to the same position with the same body orientation. Our model achieves significant improvement over the state-of-the-art approaches on three benchmark datasets.
Pengfei Zhang 0005, Cuiling Lan, Junliang Xing, Wenjun Zeng 0001, Jianru Xue, Nanning Zheng 0001
ICCV5
2017 Boosting CNN-Based Pedestrian Detection via 3D LiDAR Fusion in Autonomous Driving
Jian Dou, Jianwu Fang, Jianru Xue
ICIG (2)4
2017 Online High-Accurate Calibration of RGB+3D-LiDAR for Autonomous Driving
Jianwu Fang, Di Wang 0028, Jianru Xue
ICIG (3)5
2017 Spatial-sequential-spectral context awareness tracking
abstract
Visual context has formed a robust stimulation for visual perception. Spatio-temporal context in existing trackers sometimes shows weak reliability in visible light videos with poor quality. Supplemented by the infrared perception, this work exploits the role of visual context in tracking in a spatial-sequential-spectral view, by which to excavate dominance of different contexts in various scenarios. Specifically, we infer it in the Fourier domain with a real-time speed, and incorporate a fully-occlusion handling and scale adaptation with a trajectory regression filter and object contour closure, respectively. Extensive experiments on 50 video clips simultaneously containing registered RGB and thermal bands demonstrate that our tracker shows a state-of-the-art performance.
Jianwu Fang, Jianru Xue
ICIP3
2017 A robust submap-based road shape estimation via iterative Gaussian process regression
abstract
Road shape estimation is important for the safe driving of intelligent vehicles. The common road shape models such as line/parabola, spline and clothoid are lacking of flexibility in various urban traffic scenes. In this paper, a robust road shape model which consists of multiple overlapped submaps is proposed. Each individual submap is represented by a smooth curve generated through Gaussian process(GP). To estimate parameters of a GP submap, a framework involving pre-processing, pose correction, road shape regression and map updating/creating is proposed. Pose correction is achieved by fusion of vehicle motion model and simplified GP-based observation model. Road shape regression is used to extract a coarse road shape. Map updating/creating is used to adapt to the new coming data and generates refined road shape. A robust iterative Gaussian process regression(iGPR) is utilized in both road shape regression and map updating/creating. Extensive experimental results show the efficiency of the proposed method.
Di Wang 0028, Jianru Xue, Dixiao Cui
Intelligent Vehicles Symposium2
2017 A vision-centered multi-sensor fusing approach to self-localization and obstacle perception for robotic cars
abstract
Most state-of-the-art robotic cars’ perception systems are quite different from the way a human driver understands traffic environments. First, humans assimilate information from the traffic scene mainly through visual perception, while the machine perception of traffic environments needs to fuse information from several different kinds of sensors to meet safety-critical requirements. Second, a robotic car requires nearly 100% correct perception results for its autonomous driving, while an experienced human driver works well with dynamic traffic environments, in which machine perception could easily produce noisy perception results. In this paper, we propose a vision-centered multi-sensor fusing framework for a traffic environment perception approach to autonomous driving, which fuses camera, LIDAR, and GIS information consistently via both geometrical and semantic constraints for efficient self-localization and obstacle perception. We also discuss robust machine vision algorithms that have been successfully integrated with the framework and address multiple levels of machine vision techniques, from collecting training data, efficiently processing sensor data, and extracting low-level features, to higher-level object and environment mapping. The proposed framework has been tested extensively in actual urban scenes with our self-developed robotic cars for eight years. The empirical results validate its robustness and efficiency.
Jianru Xue, Di Wang 0028, Shaoyi Du, Dixiao Cui, Nanning Zheng 0001
Frontiers Inf. Technol. Electron. Eng.1
2017 Hybrid-augmented intelligence: collaboration and cognition
abstract
The long-term goal of artificial intelligence (AI) is to make machines learn and think like human beings. Due to the high levels of uncertainty and vulnerability in human life and the open-ended nature of problems that humans are facing, no matter how intelligent machines are, they are unable to completely replace humans. Therefore, it is necessary to introduce human cognitive capabilities or human-like cognitive models into AI systems to develop a new form of AI, that is, hybrid-augmented intelligence. This form of AI or machine intelligence is a feasible and important developing model. Hybrid-augmented intelligence can be divided into two basic models: one is human-in-the-loop augmented intelligence with human-computer collaboration, and the other is cognitive computing based augmented intelligence, in which a cognitive model is embedded in the machine learning system. This survey describes a basic framework for human-computer collaborative hybrid-augmented intelligence, and the basic elements of hybrid-augmented intelligence based on cognitive computing. These elements include intuitive reasoning, causal models, evolution of memory and knowledge, especially the role and basic principles of intuitive reasoning for complex problem solving, and the cognitive learning framework for visual scene understanding based on memory and reasoning. Several typical applications of hybrid-augmented intelligence in related fields are given.
Nanning Zheng 0001, Ziyi Liu 0001, Pengju Ren, Shi-tao Chen, Si-yu Yu, Jianru Xue, Badong Chen, Fei-Yue Wang 0001
Frontiers Inf. Technol. Electron. Eng.7
2017 Precise glasses detection algorithm for face with in-plane rotation
Shaoyi Du, Yuehu Liu, Xuetao Zhang 0001, Jianru Xue
Multim. Syst.5
2017 Video Object Discovery and Co-Segmentation with Extremely Weak Supervision
abstract
We present a spatio-temporal energy minimization formulation for simultaneous video object discovery and co-segmentation across multiple videos containing irrelevant frames. Our approach overcomes a limitation that most existing video co-segmentation methods possess, i.e., they perform poorly when dealing with practical videos in which the target objects are not present in many frames. Our formulation incorporates a spatio-temporal auto-context model, which is combined with appearance modeling for superpixel labeling. The superpixel-level labels are propagated to the frame level through a multiple instance boosting algorithm with spatial reasoning, based on which frames containing the target object are identified. Our method only needs to be bootstrapped with the frame-level labels for a few video frames (e.g., usually 1 to 3) to indicate if they contain the target objects or not. Extensive experiments on four datasets validate the efficacy of our proposed method: 1) object segmentation from a single video on the SegTrack dataset, 2) object co-segmentation from multiple videos on a video co-segmentation dataset, and 3) joint object discovery and co-segmentation from multiple videos containing irrelevant frames on the MOViCS dataset and XJTU-Stevens, a new dataset that we introduce in this paper. The proposed method compares favorably with the state-of-the-art in all of these experiments.
Le Wang 0003, Gang Hua 0001, Rahul Sukthankar, Jianru Xue, Zhenxing Niu, Nanning Zheng 0001
IEEE Trans. Pattern Anal. Mach. Intell.4
2017 A new compressive sensing video coding framework based on Gaussian mixture model
Xiangwei Li, Xuguang Lan, Meng Yang 0002, Jianru Xue, Nanning Zheng 0001
Signal Process. Image Commun.4
2016 Precise 2D point set registration using iterative closest algorithm and correntropy
abstract
The iterative closest point (ICP) algorithm is fast and accurate for rigid point set registration, but it works badly when there are many outliers and noises in the point sets. This paper instead proposes a novel method based on the ICP algorithm to deal with this problem. Firstly, correntropy is introduced into the rigid registration problem and then a new energy function based on maximum correntropy criterion is proposed. After that, a new ICP algorithm based on correntropy is proposed, which performs well in dealing with rigid registration with noises and outliers. This new algorithm converges moronically from any given parameters, which is similar to the ICP algorithm. Experimental results demonstrate its accuracy and efficiency compared with the traditional ICP algorithm.
Guanglin Xu, Shaoyi Du, Jianru Xue
IJCNN3
2016 Efficient compressive sensing video compression method based on Gaussian mixture models
abstract
In this paper, we propose an efficient lossy compression method for the compressive sensing video that utilizes Gaussian mixture models (GMM). The GMM is used to model the compressive sensing video (CSV) frames. Then we design an efficient lossy compression method based on the GMM. Each CSV frame can be efficiently compressed by the proposed method. The proposed method is for better compromise of compression efficiency and computational complexity. And it achieves a significant Bjontegaard-Delta (BD)-PSNR improvement about 8.84~11.81dB in average compared with existing low complexity compression solutions for compressing the CSV sequence.
Xiangwei Li, Xuguang Lan, Meng Yang 0002, Jianru Xue, Nanning Zheng 0001
VCIP4
2016 Robust isotropic scaling ICP algorithm with bidirectional distance and bounded rotation angle
Shaoyi Du, Chunjia Zhang, Zongze Wu 0001, Jianru Xue
Neurocomputing5
2016 Robust iterative closest point algorithm with bounded rotation angle for 2D registration
Chunjia Zhang, Shaoyi Du, Jianru Xue, Yuehu Liu
Neurocomputing5
2016 New iterative closest point algorithm for isotropic scaling registration of point sets with noise
Shaoyi Du, Bo Bi, Jihua Zhu, Jianru Xue
J. Vis. Commun. Image Represent.5
2016 Robust 3D Point Set Registration Using Iterative Closest Point Algorithm with Bounded Rotation Angle
Chunjia Zhang, Shaoyi Du, Jianru Xue
Signal Process.4
2016 Real-Time Global Localization of Robotic Cars in Lane Level via Lane Marking Detection and Shape Registration
abstract
In this paper, we propose an accurate and real-time positioning method for robotic cars in urban environments. The proposed method uses a robust lane marking detection algorithm, as well as an efficient shape registration algorithm between the detected lane markings and a GPS-based road shape prior, to improve the robustness and accuracy of the global localization of a robotic car. We show that, by formulating the positioning problem in a relative sense, we can estimate the global localization of a car in real time and bound its absolute error in the centimeter level by a cross-validation scheme. The cross-validation scheme integrates the vision-based lane marking detection with the shape registration, and it improves the accuracy and robustness of the overall localization system. The GPS localization can be refined by using lane marking detection when the GPS suffers from frequent satellite signal masking or blockage, whereas lane marking detection is validated and completed by the GPS-based road shape prior when it does not work well in adverse weather conditions or with poor lane signatures. We extensively evaluate the proposed method with a single forward-looking camera mounted on an autonomous vehicle that travels at 60 km/h through several urban street scenes.
Dixiao Cui, Jianru Xue, Nanning Zheng 0001
IEEE Trans. Intell. Transp. Syst.2
2016 Democratic Diffusion Aggregation for Image Retrieval
abstract
Content-based image retrieval is an important research topic in the multimedia field. In large-scale image search using local features, image features are encoded and aggregated into a compact vector to avoid indexing each feature individually. In the aggregation step, sum-aggregation is wildly used in many existing works and demonstrates promising performance. However, it is based on a strong and implicit assumption that the local descriptors of an image are identically and independently distributed in descriptor space and image plane. To address this problem, we propose a new aggregation method named democratic diffusion aggregation (DDA) with weak spatial context embedded. The main idea of our aggregation method is to re-weight the embedded vectors before sum-aggregation by considering the relevance among local descriptors. Different from previous work, by conducting a diffusion process on the improved kernel matrix, we calculate the weighting coefficients more efficiently without any iterative optimization. Besides considering the relevance of local descriptors from different images, we also discuss an efficient query fusion strategy which uses the initial top-ranked image vectors to enhance the retrieval performance. Experimental results show that our aggregation method exhibits much higher efficiency (about × 14 faster) and better retrieval accuracy compared with previous methods, and the query fusion strategy consistently improves the retrieval quality.
Zhanning Gao, Jianru Xue, Wengang Zhou 0001, Shanmin Pang, Qi Tian 0001
IEEE Trans. Multim.2
2015 Illumination Robust Color Naming via Label Propagation
abstract
Color composition is an important property for many computer vision tasks like image retrieval and object classification. In this paper we address the problem of inferring the color composition of the intrinsic reflectance of objects, where the shadows and highlights may change the observed color dramatically. We achieve this through color label propagation without recovering the intrinsic reflectance beforehand. Specifically, the color labels are propagated between regions sharing the same reflectance, and the direction of propagation is promoted to be from regions under full illumination and normal view angles to abnormal regions. We detect shadowed and highlighted regions as well as pairs of regions that have similar reflectance. A joint inference process is adopted to trim the inconsistent identities and connections. For evaluation we collect three datasets of images under noticeable highlights and shadows. Experimental results show that our model can effectively describe the color composition of real-world images.
Yuanliu Liu, Zejian Yuan, Badong Chen, Jianru Xue, Nanning Zheng 0001
ICCV4
2015 State-statistical model based trajectory-band planning in urban environment
abstract
In the traditional trajectory planning methods, a feasible, collision-free trajectory is generated to guide the vehicle. But generally the vehicle cannot follow the trajectory without tracking deviation because of the vehicle kinematical constraints and the performance of control algorithm. In this paper, State-Statistical Model (SSM) based trajectory-band planning method is proposed to predict the vehicle motion during the vehicle tracks the trajectory. In this method, the statistics of historical states are used to build the SSM which is a normal distribution model of tracking deviation in different segments of curvature radius and velocity. According to the SSM, the inaccessible states of vehicle can be obtained to search the best trajectory and the tracking deviation boundary can be calculated on the trajectory. Then the best trajectory is used as the base line to generate the trajectory-band of which the halfband width is the deviation boundary value. As a result, the trajectory-band can represent the maximum range of vehicle motion accurately.
Chao Ma 0024, Jing Yang 0014, Jianru Xue, Yuehu Liu
Intelligent Vehicles Symposium3
2015 A constrained VFH algorithm for motion planning of autonomous vehicles
abstract
The Vector Field Histogram (VFH) is a classical motion planning algorithm which is widely used to handle the trajectory planning problem of mobile robots. However, the traditional VFH algorithm is rarely applied to autonomous vehicles due to the vehicle's well-known non-holonomic constraints, especially in urban environments. To address this problem, we propose a constrained VFH algorithm which takes both kinematic and dynamic constraints of the vehicle into consideration. The goal is achieved via two contributions that concern both kinematic and dynamic constraints of the vehicle. First, we develop a new active region for VFH to guarantee that all states within the region are reachable for the vehicle. Second, we improve the cost function to guide the search to favor feasible motion direction for the vehicle. The proposed algorithm is extensively tested in various simulated urban environments, and experimental results validate its efficiency.
Panrang Qu, Jianru Xue, Chao Ma 0024
Intelligent Vehicles Symposium2
2015 Fast Democratic Aggregation and Query Fusion for Image Search
abstract
In image search using local features, to avoid indexing each feature individually, encoding methods are popularly adopted to embed and aggregate local features of an image into a compact vector. Democratic aggregation with triangulation embedding (T-embedding) exhibits significant retrieval accuracy improvement over previous works. However, it suffers high computational complexity. To address this problem and consistently improve the retrieval performance, we propose a new democratic method to accelerate aggregating step without accuracy lost. We also embed weak spatial context in the kernel construction to depress co-occurrence caused by local feature detector. Furthermore, we enhance the retrieval performance with an efficient query fusion strategy. The evaluation on public datasets shows that our democratic aggregation is an order of magnitude faster than the original democratic aggregation with comparable retrieval accuracy, and the query fusion achieves a significant accuracy improvement over previous works.
Zhanning Gao, Jianru Xue, Wengang Zhou 0001, Shanmin Pang, Qi Tian 0001
ICMR2
2015 Optimized truncation model for adaptive compressive sensing acquisition of images
abstract
The sparsity of the input signal is important for compressive sensing (CS) reconstruction in CS system. In this paper, we establish an optimized truncation model to determine the number of the sparsified coefficients to be truncated in CS acquisition according to the sampling rate. The proposed truncation model suits for signals of any dimension. With the truncation model, the sparsity of the signal can be optimized by properly truncating the small elements of the sparsified coefficients. Furthermore we propose an adaptive CS acquisition solution based on the truncation model to reduce the noise folding effect. The proposed solution is verified for CS acquisition of natural images. Simulation results show that the proposed solution achieves significant improvement of the reconstructed image quality by 0.7~1.4 dB on average compared with existing solutions.
Xiangwei Li, Xuguang Lan, Meng Yang 0002, Jianru Xue, Nanning Zheng 0001
VCIP4
2015 Accurate non-rigid registration based on heuristic tree for registering point sets with large deformation
Shaoyi Du, Chunjia Zhang, Meifeng Xu, Jianru Xue
Neurocomputing5
2015 Image re-ranking with an alternating optimization
Shanmin Pang, Jianru Xue, Zhanning Gao, Qi Tian 0001
Neurocomputing2
2015 Authentication and copyright protection watermarking scheme for H.264 based on visual saliency and secret sharing
Lihua Tian, Nanning Zheng 0001, Jianru Xue, Ce Li 0001
Multim. Tools Appl.3
2015 A robust approach to detect digital forgeries by exploring correlation patterns
Lu Li 0009, Jianru Xue, Lihua Tian
Pattern Anal. Appl.2
2015 A Visual Model-Based Perceptual Image Hash for Content Authentication
abstract
Perceptual image hash has been widely investigated in an attempt to solve the problems of image content authentication and content-based image retrieval. In this paper, we combine statistical analysis methods and visual perception theory to develop a real perceptual image hash method for content authentication. To achieve real perceptual robustness and perceptual sensitivity, the proposed method uses Watson's visual model to extract visually sensitive features that play an important role in the process of humans perceiving image content. We then generate robust perceptual hash code by combining image-block-based features and key-point-based features. The proposed method achieves a tradeoff between perceptual robustness to tolerate content-preserving manipulations and a wide range of geometric distortions and perceptual sensitivity to detect malicious tampering. Furthermore, it has the functionality to detect compromised image regions. Compared with state-of-the-art schemes, the proposed method obtains a better comprehensive performance in content-based image tampering detection and localization.
Kemu Pang, Xiaorui Zhou, Lu Li 0009, Jianru Xue
IEEE Trans. Inf. Forensics Secur.6
2015 Efficient Sampling-Based Motion Planning for On-Road Autonomous Driving
abstract
This paper introduces an efficient motion planning method for on-road driving of the autonomous vehicles, which is based on the rapidly exploring random tree (RRT) algorithm. RRT is an incremental sampling-based algorithm and is widely used to solve the planning problem of mobile robots. However, due to the meandering path, the inaccurate terminal state, and the slow exploration, it is often inefficient in many applications such as autonomous vehicles. To address these issues and considering the realistic context of on-road autonomous driving, we propose a fast RRT algorithm that introduces a rule-template set based on the traffic scenes and an aggressive extension strategy of search tree. Both improvements lead to a faster and more accurate RRT toward the goal state compared with the basic RRT algorithm. Meanwhile, a model-based prediction postprocess approach is adopted, by which the generated trajectory can be further smoothed and a feasible control sequence for the vehicle would be obtained. Furthermore, in the environments with dynamic obstacles, an integrated approach of the fast RRT algorithm and the configuration-time space can be used to improve the quality of the planned trajectory and the replanning. A large number of experimental results illustrate that our method is fast and efficient in solving planning queries of on-road autonomous driving and demonstrate its superior performances over previous approaches.
Jianru Xue, Kuniaki Kawabata, Jihua Zhu, Chao Ma 0024, Nanning Zheng 0001
IEEE Trans. Intell. Transp. Syst.2
2014 Video Object Discovery and Co-segmentation with Extremely Weak Supervision
Le Wang 0003, Gang Hua 0001, Rahul Sukthankar, Jianru Xue, Nanning Zheng 0001
ECCV (4)4
2014 Facial age estimation from web photos using multiple-instance learning
abstract
One of the main bottle-necks in traditional facial age estimation is the lack of training sample problem. The rapid development of Internet provides us new chance to solve this problem. Unlimited number of facial images with their age labels can be collected through web mining technique. These images together with their surrounding text description make up the simplest cross-media data representation. In this paper, we model this problem within a Multiple Instance Learning (MIL) framework, and a novel algorithm named Witness based Multiple Instance Regression (WMIR) is proposed. The "witness" faces in the group photos are found together with their age label and confidence. A probabilistic weighted Support Vector Regression (pw-SVR) method is designed to utilize these cross-media data for learning a more robust age estimator. Experimental results upon both the synthetic data and real web data have verified the advantage of our algorithm compared with other related methods.
Jianru Xue
ICME4
2014 Real-time global localization of intelligent road vehicles in lane-level via lane marking detection and shape registration
abstract
In this paper, we propose an accurate and real-time positioning method for intelligent road vehicles in urban environments. The proposed method uses a robust lane marking detection algorithm, as well as an efficient shape registration algorithm between the detected lane markings and a GPS based road shape prior, to improve the robustness and accuracy of global localization of a road vehicle. We exploit both the state-of-the-art technologies of visual localization based on lane marking detection and the wide availability of Global Positioning System (GPS) based localization. We show that by formulating the positioning problem in a relative sense, we can estimate the vehicle localization in real-time and bound its absolute error in centimeter-level by a cross validation scheme. The validation scheme integrates the vision based lane marking detection with the shape registration, and improves the performance of the overall localization system. The GPS localization can be refined by using lane marking detection when the GPS suffers from frequent satellite signal masking or blockage, while lane marking detection is validated and completing by the GPS based road shape prior when it does not work well in adverse weather conditions or with poor lane signature. We extensively evaluate the proposed method with a single forward-looking camera mounted on an autonomous vehicle which travels at 60km/h through several urban street scenes.
Dixiao Cui, Jianru Xue, Shaoyi Du, Nanning Zheng 0001
IROS2
2014 Image Re-ranking with an Alternating Optimization
abstract
In this work, we propose an efficient image re-ranking method, without additional memory cost compared with the baseline method~\cite{philbin2007object}, to re-rank all retrieved images. The motivation of the proposed method is that, there are usually many visual words in the query image that only give votes to irrelevant images. With this observation, we propose to only use visual words which can help to find relevant images to re-rank the retrieved images. To achieve the goal, we first find some similar images to the query by maximizing a quadratic function when given an initial ranking of the retrieved images. Then we select query visual words with an alternating optimization strategy: (1) at each iteration, select words based on the similar images that we have found and (2) in turn, update the similar images with the selected words. These two steps are repeated until convergence. Experimental results on standard benchmark datasets show that the proposed method outperforms spatial based re-ranking methods.
Shanmin Pang, Jianru Xue, Zhanning Gao, Qi Tian 0001
ACM Multimedia2
2014 Exploiting local linear geometric structure for identifying correct matches
Shanmin Pang, Jianru Xue, Qi Tian 0001, Nanning Zheng 0001
Comput. Vis. Image Underst.2
2014 A statistical feature based approach to distinguish PRCG from photographs
Bingchao Xu, Lu Li 0009, Jianru Xue
Comput. Vis. Image Underst.5
2014 Video object segmentation with shape cue based on spatiotemporal superpixel neighbourhood
abstract
In this study, the authors present a method to extract moving objects in image sequences. The proposed approach is based on a graph cuts algorithm defined on a spatiotemporal superpixel neighbourhood. Presegmented superpixels are partitioned into foreground and background while preserving temporal and spatial coherence. It achieves this goal by three steps. First, instead of operating at pixel level, the superpixels are advocated as basic units of the authors segmentation scheme. Second, within the graph cuts framework, two superpixel‐based data terms and two superpixel‐based smoothness terms are proposed to solve segmentation problem. Finally, the proposed method yields the segmentation of all the superpixels within video volume by the graph cuts algorithm. To illustrate the advantages of this approach, the quantitative and qualitative results are compared with other state‐of‐the‐art methods. The experimental results show that the proposed method gives better performance of segmentation with respect to these methods.
Nanning Zheng 0001, Jianru Xue, Xuguang Lan, Ce Li 0001
IET Comput. Vis.3
2014 Object segmentation and key-pose based summarization for motion video
Jianru Xue, Xuguang Lan, Ce Li 0001, Nanning Zheng 0001
Multim. Tools Appl.2
2014 Joint Segmentation and Recognition of Categorized Objects From Noisy Web Image Collection
abstract
The segmentation of categorized objects addresses the problem of joint segmentation of a single category of object across a collection of images, where categorized objects are referred to objects in the same category. Most existing methods of segmentation of categorized objects made the assumption that all images in the given image collection contain the target object. In other words, the given image collection is noise free. Therefore, they may not work well when there are some noisy images which are not in the same category, such as those image collections gathered by a text query from modern image search engines. To overcome this limitation, we propose a method for automatic segmentation and recognition of categorized objects from noisy Web image collections. This is achieved by cotraining an automatic object segmentation algorithm that operates directly on a collection of images, and an object category recognition algorithm that identifies which images contain the target object. The object segmentation algorithm is trained on a subset of images from the given image collection which are recognized to contain the target object with high confidence, while training the object category recognition model is guided by the intermediate segmentation results obtained from the object segmentation algorithm. This way, our co-training algorithm automatically identifies the set of true positives in the noisy Web image collection, and simultaneously extracts the target objects from all the identified images. Extensive experiments validated the efficacy of our proposed approach on four datasets: 1) the Weizmann horse dataset, 2) the MSRC object category dataset, 3) the iCoseg dataset, and 4) a new 30-categories dataset including 15,634 Web images with both hand-annotated category labels and ground truth segmentation labels. It is shown that our method compares favorably with the state-of-the-art, and has the ability to deal with noisy image collections.
Le Wang 0003, Gang Hua 0001, Jianru Xue, Zhanning Gao, Nanning Zheng 0001
IEEE Trans. Image Process.3
2013 Moment feature based forensic detection of resampled digital images
abstract
Forensic detection of resampled digital images has become an important technology among many others to establish the integrity of digital visual content. This paper proposes a moment feature based method to detect resampled digital images. Rather than concentrating on the positions of characteristic resampling peaks, we utilize a moment feature to exploit the periodic interpolation characteristics in the frequency domain. Not only the positions of resampling peaks but also the amplitude distribution is taken into consideration. With the extracted moment feature, a trained SVM classifier is used to detect resampled digital images. Extensive experimental results show the validity and efficiency of the proposed method.
Lu Li 0009, Jianru Xue, Nanning Zheng 0001
ACM Multimedia2
2013 Locality preserving verification for image search
abstract
Establishing correct correspondences between two images has a wide range of applications, such as 2D and 3D registration, structure from motion, and image retrieval. In this paper, we propose a new matching method based on spatial constraints. The proposed method has linear time complexity, and is efficient when applying it to image retrieval. The main assumption behind our method is that, the local geometric structure among a feature point and its neighbors, is not easily affected by both geometric and photometric transformations, and thus should be preserved in their corresponding images. We model this local geometric structure by linear coefficients that reconstruct the point from its neighbors. The method is flexible, as it can not only estimate the number of correct matches between two images efficiently, but also determine the correctness of each match accurately. Furthermore, it is simple and easy to be implemented. When applying the proposed method on re-ranking images in an image search engine, it outperforms the-state-of-the-art techniques.
Shanmin Pang, Jianru Xue, Nanning Zheng 0001, Qi Tian 0001
ACM Multimedia2
2013 Universal and low-complexity quantizer design for compressive sensing image coding
abstract
Compressive sensing imaging (CSI) is a new framework for image coding, which enables acquiring and compressing a scene simultaneously. The CS encoder shifts the bulk of the system complexity to the decoder efficiently. Ideally, implementation of CSI provides lossless compression in image coding. In this paper, we consider the lossy compression of the CS measurements in CSI system. We design a universal quantizer for the CS measurements of any input image. The proposed method firstly establishes a universal probability model for the CS measurements in advance, without knowing any information of the input image. Then a fast quantizer is designed based on this established model. Simulation result demonstrates that the proposed method has nearly optimal rate-distortion (R~D) performance, meanwhile, maintains a very low computational complexity at the CS encoder.
Xiangwei Li, Xuguang Lan, Meng Yang 0002, Jianru Xue, Nanning Zheng 0001
VCIP4
2013 Automatic salient object extraction with contextual cue and its applications to recognition and alpha matting
Jianru Xue, Le Wang 0003, Nanning Zheng 0001, Gang Hua 0001
Pattern Recognit.1
2012 Large-Scale Bundle Adjustment by Parameter Vector Partition
Shanmin Pang, Jianru Xue, Le Wang 0003, Nanning Zheng 0001
ACCV (4)2
2012 Concurrent segmentation of categorized objects from an image collection
Le Wang 0003, Jianru Xue, Nanning Zheng 0001, Gang Hua 0001
ICPR2
2012 A Novel Image Signature Method for Content Authentication
abstract
We proposed an image signature method for content authentication, which applies a hierarchical approach to construct an image signature. In the first level, DWT and DCT are used to extract image features; then these features are encrypted by using sub-keys that are generated by a cryptographically hash function. In the second level, Karhunen–Loeve transformation is used to reduce the signature length. The main features of the proposed method are as follows: (i) It achieves a trade-off between robustness and tampering sensitivity. (ii) It provides a tool for image tampering detection and tampering localization. (iii) It can be used to detect the thumbnail of the large image to improve detection efficiency. (iv) It provides the compact signature, and the signature length is independent of the image size. Experimental results show that proposed method is robust for content-preserving manipulations such as JPEG compression, adding noise, filtering and Gamma correction, etc.
Nanning Zheng 0001, Jianru Xue, Zhenli Liu
Comput. J.3
2012 Image forensic signature for content authenticity analysis
Jianru Xue, Zhenqiang Zheng, Zhenli Liu
J. Vis. Commun. Image Represent.2
2012 PSF Estimation via Gradient Domain Correlation
abstract
This paper proposes an efficient method to estimate the point spread function (PSF) of a blurred image using image gradients spatial correlation. A patch-based image degradation model is proposed for estimating the sample covariance matrix of the gradient domain natural image. Based on the fact that the gradients of clean natural images are approximately uncorrelated to each other, we estimated the autocorrelation function of the PSF from the covariance matrix of gradient domain blurred image using the proposed patch-based image degradation model. The PSF is computed using a phase retrieval technique to remove the ambiguity introduced by the absence of the phase. Experimental results show that the proposed method significantly reduces the computational burden in PSF estimation, compared with existing methods, while giving comparable blurring kernel.
Jianru Xue, Nanning Zheng 0001
IEEE Trans. Image Process.2
2011 A Robust Approach to Detect Tampering by Exploring Correlation Patterns
Lu Li 0009, Jianru Xue, Lihua Tian
CAIP (2)2
2011 Automatic salient object extraction with contextual cue
abstract
We present a method for automatically extracting salient object from a single image, which is cast in an energy minimization framework. Unlike most previous methods that only leverage appearance cues, we employ an auto-context cue as a complementary data term. Benefitting from a generic saliency model for bootstrapping, the segmentation of the salient object and the learning of the auto-context model are iteratively performed without any user intervention. Upon convergence, we obtain not only a clear separation of the salient object, but also an auto-context classifier which can be used to recognize the same type of object in other images. Our experiments on four benchmarks demonstrated the efficacy of the added contextual cue. It is shown that our method compares favorably with the state-of-the-art, some of which even embraced user interactions.
Le Wang 0003, Jianru Xue, Nanning Zheng 0001, Gang Hua 0001
ICCV2
2011 Fast and robust isotropic scaling iterative closest point algorithm
abstract
The iterative closest point (ICP) algorithm is an accurate approach for the registration between two point sets on the same scale. However, it can not handle the case with different scales. This paper proposes a fast and robust ICP algorithm for isotropic scaling point sets registration (FRISICP). In order to accurately and directly estimate the scale factor without any constraints, we introduce a bidirection distance measurement method into the least square (LS) problem. Then to keep computational efficiency when the number of points in the set increasing, we further introduce a sparse-to-dense hierarchical model in ICP algorithm to speed up the isotropic scaling point set matching process. Experimental results demonstrate that the proposed FRISICP method outperforms other algorithms on both 2D and 3D point sets.
Ce Li 0001, Jianru Xue, Nanning Zheng 0001, Shaoyi Du, Jihua Zhu
ICIP2
2011 3D spatio-temporal graph cuts for video objects segmentation
abstract
In this paper, we present a method to extract moving objects in monocular image sequences. The proposed method is based on graph cuts defined on a spatio-temporal region adjacency graph (RAG). First, we initially over-segment each frame in the video, and take the over-segmented regions as the vertices in the 3D spatio-temporal graph. Second, multiple cues are fused together to extract objects accurately. Finally, accurate foreground/background segmentation are efficiently achieved by binary graph cut. The experimental results showed that the proposed method improved the performance of segmentation with respect to the popular methods.
Jianru Xue, Nanning Zheng 0001, Xuguang Lan, Ce Li 0001
ICIP2
2011 Auto-generated strokes for motion segmentation
abstract
We propose a new approach to motion segmentation that is based on auto-generated strokes. The novelty of the approach is twofold. First, inspired by recent work of other researchers we formulate the problem as that of interactive segmentation. Instead of inputting the strokes by the user, the strokes in our approach are auto-generated. The second novelty of the paper is formulation in which, unlike in many other motion segmentation algorithms, we do not use complex algorithm which fuses the output of multiple cues to segment foreground objects, a simple and effective maximum hybrid similarity method is presented. The maximum hybrid similarity does not need to set the threshold in advance. Experimental results have shown the superiority of the proposed method in extracting moving objects.
Jianru Xue, Ce Li 0001, Xuguang Lan, Nanning Zheng 0001
ISCAS2
2011 Nonparametric bottom-up saliency detection using hypercomplex spectral contrast
abstract
Saliency detection is an useful technique for image semantic analysis such as auto image segmentation, image retargeting, advertising design and image compression. Inspired by two existing saliency detection algorithms, named spectral residual (SR) and phase spectrum of quaternion Fourier transform (PQFT), we propose a new bottom-up saliency detection method which is featured with the introduction of hypercomplex spectral contrast (HSC) in saliency detection. The proposed HSC algorithm introduces the HSV color image vector space in hypercomplex number, and is better comprehensive to consider amplitude spectral contrast into saliency model as well as phase spectral contrast. Meanwhile, we also incorporate the human vision nonuniform sampling into our model, which is a common phenomenon that directs visual attention to the logarithmic center of image in natural scenes. Experimental results on two public saliency detection datasets show that our approach performs better than four state-of-the art approaches remarkably.
Ce Li 0001, Jianru Xue, Nanning Zheng 0001
ACM Multimedia2
2011 Key object-based static video summarization
abstract
In this paper, we present a system for object-based video summarization facilitated by an efficient video object segmentation system. We eliminate the redundancy not only from spatial and temporal domain, but also from content domain. First, we detect shot boundaries and extract video objects by a 3D graph-based algorithm. Once the objects are obtained, the shape of the objects need to be represented. The key objects are extracted in a global manner by K-means clustering of shapes. Experimental results on the proposed object-based scheme combined with efficient video object segmentation show desirable summarization.
Jianru Xue, Xuguang Lan, Ce Li 0001, Nanning Zheng 0001
ACM Multimedia2
2011 An integrated visual saliency-based watermarking approach for synchronous image authentication and copyright protection
Lihua Tian, Nanning Zheng 0001, Jianru Xue, Ce Li 0001
Signal Process. Image Commun.3
2011 Proto-Object Based Rate Control for JPEG2000: An Approach to Content-Based Scalability
abstract
The JPEG2000 system provides scalability with respect to quality, resolution and color component in the transfer of images. However, scalability with respect to semantic content is still lacking. We propose a biologically plausible salient region based bit allocation mechanism within the JPEG2000 codec for the purpose of augmenting scalability with respect to semantic content. First, an input image is segmented into several salient proto-objects (a region that possibly contains a semantically meaningful physical object) and background regions (a region that contains no object of interest) by modeling visual focus of attention on salient proto-objects. Then, a novel rate control scheme distributes a target bit rate to each individual region according to its saliency, and constructs quality layers of proto-objects for the purpose of more precise truncation comparable to original quality layers in the standard. Empirical results show that the suggested approach adds to the JPEG2000 system scalability with respect to content as well as the functionality of selectively encoding, decoding, and manipulation of each individual proto-object in the image, with only some slightly trivial modifications to the JPEG2000 standard. Furthermore, the proposed rate control approach efficiently reduces the computational complexity and memory usage, as well as maintains the high quality of the image to a level comparable to the conventional post-compression rate distortion (PCRD) optimum truncation algorithm for JPEG2000.
Jianru Xue, Ce Li 0001, Nanning Zheng 0001
IEEE Trans. Image Process.1
2010 Arbitrary ROI-Based Wavelet Video Coding
abstract
An arbitrary shape region of interest (ROI) coding is presented for scalable wavelet video codec in this paper. The padding of macroblock and polygon matching are employed to estimate the motion of the ROI. The motion vectors derived are set as the motion trajectory of the samples to generate one-dimensional temporal signal, filtered to reduce the temporal redundancy by using motion compensated temporal filtering for arbitrary shape ROI. The Reconstructed quality of the ROI coding can be significantly improved at low bit rate, compared to non-ROI coding. The efficiency of the MCTF based on arbitrary ROI is compared with that of the video object coding in MPEG-4. The ability of the MCTF to reduce the temporal redundancy is better than or comparable to that MPEG-4 to some extent.
Xuguang Lan, Nanning Zheng 0001, Miao Hui, Jianru Xue
DCC5
2010 PSF Estimation via Covariance Matching
Jianru Xue, Nanning Zheng 0001
ICASSP2
2010 Scaling iterative closest point algorithm for registration of m-D point sets
Shaoyi Du, Nanning Zheng 0001, Shihui Ying, Jianru Xue
J. Vis. Commun. Image Represent.5
2009 Hierarchical Model for Joint Detection and Tracking of Multi-target
Jianru Xue, Nanning Zheng 0001
ACCV (2)1
2009 A Peer-to-Peer Architecture for Live Streaming with DRM
abstract
DRM is becoming more and more important for P2P live streaming. In this paper, a manageable overlay network architecture with DRM, is proposed for live streaming. The system consists of register server, index servers, supernodes and peers. The register server authorizes the peers, assign the key and index server list; The index server acts as the centralized index server to store the peer list, program list, and buffer information of peers. Supernodes are special peers which store the bigger buffer of live streaming. Each peer periodically exchanges data availability information with the assigned index server. The peer can retrieve correspondingly unavailable data from partners given by index server using the proposed scheduling algorithm. It also supports DRM in register server. The proposed system has been demonstrated based on CERNET of China. Good streaming quality can be achieved due to its global optimization and the digital right of video contents can be protected to some extent.
Xuguang Lan, Jianru Xue, Lihua Tian, Nanning Zheng 0001
CCNC2
2009 Joint Network-Source Video Coding Based on Lagrangian Rate Allocation
abstract
Joint network-source video coding (JNSC) is targeted to achieve the optimum delivery of a video source to a number of destinations over network with capacity constraints. In this paper, a practical scalable multiple description coding is proposed for JNSC, based on Lagrangian rate allocation and scalable video coding. After the spatiotemporal wavelet transformation of input video sequence and the bit plane coding and context-based adaptive binary arithmetic coding, jointing network-source coding is performed on the coding passes of the code blocks (CB) using Lagrangian rate allocation.
Xuguang Lan, Nanning Zheng 0001, Jianru Xue, Ce Li 0001, Songlin Zhao
DCC3
2008 A CAVLC-Based Blind Watermarking Method for H.264/AVC Compressed Video
abstract
Streaming media services have been applied in many applications. At the same time, the security of it should be considered. To protect the video content, watermarking data (image) is usually embedded into the quantized DCT coefficients of video. At the same time, the watermarked video should keep the fidelity and the bit-rate. While for high efficient compression video such as H.264/AVC it is very difficult, because just one bit alteration may widely affect the video content and the bit-rate. A CAVLC-based blind watermarking method for H.264/AVC compressed video is proposed. The watermarking data is only embedded into the last non-zero and non-trailing AC coefficient in context adaptive variable length coding (CAVLC) of H.264/AVC. With this kind of embedding, the artifact due to the embedding could be reduced efficiently by CAVLC. Experimental results show that on average, the introduced distortion by watermark embedding is less than 0.5 dB, and the increased stream bit rate is only 0.1%.
Lihua Tian, Nanning Zheng 0001, Jianru Xue
APSCC3
2008 Scalable Multiple Description Coding Based on Bitplane and Lagrangian Rate Allocation
abstract
Scalable video coding is applied to meet the heterogeneity of networks, fluctuation of bandwidth, and diversity of end users. But if the base layer is damaged or missed, the scalable video coding will no longer be effective. To address this problem, a scalable multiple description of video is proposed based on the bit plane and rate allocation of scalable wavelet video coding in an error-prone transmission environment. After spatiotemporal transformation, three-dimensional bit planes of the wavelet frame are sampled in odd and even order except the most significant bit to generate multiple descriptions, which are encoded into scalable video streaming by using entropy coding and Lagragian rate allocation. Combined with multiple path transport, the performance of proposed scalable multiple description coding is demonstrated in Ns2.
Xuguang Lan, Nanning Zheng 0001, Songlin Zhao, Weike Chen, Jianru Xue
CCNC5
2008 A Peer-to-Peer Architecture Based on Scalable Video Coding
abstract
Combining the advantages of the centralized P2P structure and the data-driven structure with scalable video coding, we propose an adaptive P2P architecture for live video based on scalable wavelet video coding over Internet. The core operations are simple: every peer periodically exchanges data availability and bandwidth information with the central server, acting as the centralized index, which selects and sends a set of partners that have expected data to the demanding node. And the central server classifies one peer to a certain level according to the peer's downloading bandwidth which coordinates with the layer level of the video data encoded using scalable wavelet coding. The peer retrieves correspondingly unavailable data from partners according to availability information. There are three principal advantages of this architecture: 1) easy to manage, as the server authorizes, classifies, and clusters each new peer according to its bandwidth as it joins the overlay network, and thus the servers maintain a global structure; 2) efficiency in dynamically heterogeneous networks, as users with different processing ability under heterogeneous networks can retrieve the adaptive data from their partner peers according to their downloading bandwidths, and 3) robustness and resilience, as the partner peers can adapt to quick switching among multi-suppliers, and the scalable video transmission is adaptive. A scheduling algorithm is proposed to enable efficient and continuous scalable streaming of low to high bandwidth content with different service levels over heterogeneous networks.
Xuguang Lan, Nanning Zheng 0001, Jianru Xue, Weike Chen, Songlin Zhao
DCC3
2008 Linear regression models for DCT domain approximate filtering and deblurring
abstract
This paper presents two linear regression models by exploring the relationships between DCT coefficients of the original image and filtered image. The first model is to scale the DCT coefficients of the original image in order to approximate the operation of 2-D spatial domain filtering. The second model is to predict the original image from the filtered image in a similar manner. We show that the first model is used for DCT domain filtering, while the second model can be used for fast DCT domain image deblurring. Both of them are easy to implement on compressed formats of DCT-based compression methods (JPEG, MPEG, H.26X) by using decoding quantization tables that are different from the encoding quantization tables.
Nanning Zheng 0001, Jianru Xue, Xuguang Lan
ICME3
2008 Manageable peer-to-peer architecture for video-on-demand
abstract
An efficiently manageable overlay network architecture, called AIRVoD, is proposed for video-on- demand, based on a distributed-centralized P2P network. The system consists of distributed servers, peers, SuperNodes and sub-SuperNodes. The distributed central servers act as the centralized index server to store the peer list, program list, and buffer information of peers. Each newly joined peer in AIRVoD periodically exchanges data availability information with the central server. Some powerful peers are selected to be sub-SuperNodes which store a larger part of the demanded program. The demanding peers can retrieve correspondingly unavailable data from partners selected from the central server, sub- SuperNodes and SuperNodes that have the original programs to supply the available data, by using the proposed parallel scheduling algorithm. There are four characteristics of this architecture: 1) easy to globally balance and manage: central server can identify, and cluster each newly joined peer, and allocate the load in the whole peer network; 2) efficient for dynamic networks: data transmission is dynamically determined according to data availability which can be derived from central server; 3) resilient, as the partnerships can adapt to quick switching among multi-suppliers under global balance; and 4) highly cost-effective: the powerful peers are taken full advantage to be larger suppliers. AIRVoD has been demonstrated based on CERNET OF CHINA.
Xuguang Lan, Nanning Zheng 0001, Jianru Xue, Weike Chen
IPDPS3
2008 Image smoothing and sharpening based on nonlinear diffusion equation
Fang Dai, Nanning Zheng 0001, Jianru Xue
Signal Process.3
2008 Tracking Multiple Visual Targets via Particle-Based Belief Propagation
abstract
Multiple-target tracking in video (MTTV) presents a technical challenge in video surveillance applications. In this paper, we formulate the MTTV problem using dynamic Markov network (DMN) techniques. Our model consists of three coupled Markov random fields: 1) a field for the joint state of the multitarget; 2) a binary random process for the existence of each individual target; and 3) a binary random process for the occlusion of each dual adjacent target. To make the inference tractable, we introduce two robust functions that eliminate the two binary processes. We then propose a novel belief propagation (BP) algorithm called particle-based BP and embed it into a Markov chain Monte Carlo approach to obtain the maximum a posteriori estimation in the DMN. With a stratified sampler, we incorporate the information obtained from a learned bottom-up detector (e.g., support-vector-machine-based classifier) and the motion model of the target into the message propagation. Other low-level visual cues such as motion and shape can be easily incorporated into our framework to obtain better tracking results. We have performed extensive experimental verification, and the results suggest that our method is comparable to the state-of-art multitarget tracking methods in all the cases we tested.
Jianru Xue, Nanning Zheng 0001, Jason Geng, Xiaopin Zhong
IEEE Trans. Syst. Man Cybern. Part B1
2007 A peer-to-peer architecture for efficient live scalable media streaming on internet
abstract
This paper presents a manageable overlay network architecture SVCP2P for live scalable media streaming. Every peer in SVCP2P periodically exchanges data availability information with one of distributed central servers which act as the centralized index for storing peer list, program list and buffer information of peers. An efficient scheduling algorithm is proposed, which achieves real-time and continuous transmission of the scalable streaming. There are three characteristics of this architecture: 1) easy management; 2) efficient to heterogeneous network because of the scalable media streaming adapting to the heterogeneous demand; and 3) robust and resilient. We have examined the SVCP2P which has been implemented based on the IP Internet over LAN, and the results demonstrate the efficiency of SVCP2P.
Xuguang Lan, Nanning Zheng 0001, Jianru Xue, Xiaoguang Wu
ACM Multimedia3
2006 Tracking Targets Via Particle Based Belief Propagation
Jianru Xue, Nanning Zheng 0001, Xiaopin Zhong
ACCV (1)1
2006 Pseudo Measurement Based Multiple Model Approach for Robust Player Tracking
Xiaopin Zhong, Nanning Zheng 0001, Jianru Xue
ACCV (2)3
2006 Graphical Model based Cue Integration Strategy for Head Tracking
abstract
To achieve robust system, more and more vision researchers take into account fusing multiple visual cues. In this paper, we propose a novel strategy to integrate multiple naive cues for head tracking. Firstly, a cue dependency model is constructed via graphical model. Secondly, a new inference procedure based on non-parametric belief propagation is built for cue integration. The work presented is thus a general framework easy to extend for other computer vision research problems. Experimental results imply that the strategy we propose is effective, and it is robust without estimation of cue reliability. 1
Xiaopin Zhong, Jianru Xue, Nanning Zheng 0001
BMVC2
2006 Sequential stratified sampling belief propagation for multiple targets tracking
Jianru Xue, Nanning Zheng 0001, Xiaopin Zhong
Sci. China Ser. F Inf. Sci.1
2005 Sequential Stratified Sampling Belief Propagation for Multiple Targets Tracking
Jianru Xue, Nanning Zheng 0001, Xiaopin Zhong
ICIC (1)1