VLDB 2026 Research / reviewers in the wild / expert
Zhixiong Nan
dblp:187/2031
· DBLP profile ↗
33ranked-venue papers
13as first author
24since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 26 · 10 first-author · 17 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 5 first-author · 10 since 2021Systems, architecture and hardware · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | AbductiveMLLM: Boosting Visual Abductive Reasoning Within MLLMsabstractVisual abductive reasoning (VAR) is a challenging task that requires AI systems to infer the most likely explanation for incomplete visual observations. While recent MLLMs develop strong general-purpose multimodal reasoning capabilities, they remain fall short in abductive inference, as compared to human beings. To bridge this gap, we draw inspiration from the interplay between verbal and pictorial abduction in human cognition, and propose to strengthen abduction of MLLMs by mimicking such dual-mode behavior. Concretely, we introduce AbductiveMLLM comprising of two synergistic components: REASONER and IMAGINER. The REASONER operates in the verbal domain. It first explores a broad space of possible explanations using a blind LLM and then prunes visually incongruent hypotheses based on cross-modal causal alignment. The remaining hypotheses are introduced into the MLLM as targeted priors, steering its reasoning toward causally coherent explanations. The IMAGINER, on the other hand, further guides MLLMs by emulating human-like pictorial thinking. It conditions a text-to-image diffusion model on both the input video and the REASONER’s output embeddings to “imagine” plausible visual scenes that correspond to verbal explanation, thereby enriching MLLMs' contextual grounding. The two components are trained jointly in an end-to-end manner. Experiments on standard VAR benchmarks show that AbductiveMLLM achieves state-of-the-art performance, consistently outperforming traditional solutions and advanced MLLMs. Boyu Chang, Qi Wang 0009, Zhixiong Nan, Yazhou Yao, Tianfei Zhou |
AAAI | 4 |
| 2026 | PGMamba: A Physical Model-Guided Global Mamba for Underwater Image EnhancementabstractUnderwater image enhancement (UIE) aims to address image degradation caused by water absorption and scattering effects. Despite significant progress in deep learning-based UIE methods, existing approaches still face key challenges due to the neglect of physical imaging principle. Moreover, while current Mamba models achieve global modeling via multi-directional scanning, their local sequential strategy lacks sufficient global context. To this end, we propose a novel Physical Model-Guided Global Mamba (PGMamba) that combines the efficient sequential modeling capability of Mamba with underwater imaging physical model. Specifically, we first design a Spatial-Aware Global Mamba (SAGMamba) that achieves efficient long-range dependency modeling through a spatial-aware ranking strategy with global context information. Second, we develop a Physical Model-Guided Feed-Forward Network (PMGFFN) that explicitly incorporates underwater optical imaging principles into the network architecture. Extensive experimental results and comprehensive ablation studies demonstrate the outstanding performance and importance of our proposed method. Zijun Tan, Chuan Fu, Tan Guo, Zhixiong Nan, Pengzhan Zhou, Xinggan Peng, Fulin Luo |
AAAI | 4 |
| 2026 | Lane detection with vanishing box based dynamic anchor generation mechanism
Zhixiong Nan, Wanying Xu, Tao Xiang 0001 |
Neurocomputing | 1 |
| 2026 | A lane detection model with knowledge guided anchor feature enhancement mechanism
Zhixiong Nan, Wanying Xu, Fulin Luo, Tao Xiang 0001 |
Pattern Recognit. | 1 |
| 2025 | MI-DETR: An Object Detection Model with Multi-time Inquiries MechanismabstractBased on analyzing the character of cascaded decoder architecture commonly adopted in existing DETR-like models, this paper proposes a new decoder architecture. The cascaded decoder architecture constrains object queries to update in the cascaded direction, only enabling object queries to learn relatively-limited information from image features. However, the challenges for object detection in natural scenes (e.g., extremely-small, heavily-occluded, and confusingly mixed with the background) require an object detection model to fully utilize image features, which motivates us to propose a new decoder architecture with the parallel Multi-time Inquiries (MI) mechanism. MI mechanism is very simple, enabling object queries to parallelly perform multi-time inquiries to learn more comprehensive information from image features. Our MI based model, MI-DETR, outperforms all existing DETR-like models on COCO benchmark under different backbones and training epochs, achieving +2.3 AP and +0.6 AP improvements compared to the most representative model DINO and SOTA model Relation-DETR under ResNet-50 backbone. Zhixiong Nan, Xianghong Li, Jifeng Dai, Tao Xiang 0001 |
CVPR | 1 |
| 2025 | Dynamic facial expression recognition in the wild via Multi-Snippet Spatiotemporal Learning
Yang Lü, Fuchun Zhang, Zongnan Ma, Zhixiong Nan |
Neurocomputing | 5 |
| 2025 | HEOI: Human Attention Prediction in Natural Daily Life With Fine-Grained Human-Environment-Object Interaction ModelabstractThis paper handles the problem of human attention prediction in natural daily life from the third-person view. Due to the significance of this topic in various applications, researchers in the computer vision community have proposed many excellent models in the past few decades, and many models have begun to focus on natural daily life scenarios in recent years. However, existing mainstream models usually ignore a basic fact that human attention is a typical interdisciplinary concept. Specifically, the mainstream definition is direction-level or pixel-level, while many interdisciplinary studies argue the object-level definition. Additionally, the mainstream model structure converges to the dual-pathway architecture or its variants, while the majority of interdisciplinary studies claim attention is involved in the human-environment interaction procedure. Grounded on solid theories and studies in interdisciplinary fields including computer vision, cognition, neuroscience, psychology, and philosophy, this paper proposes a fine-grained Human-Environment-Object Interaction (HEOI) model, which for the first time integrates multi-granularity human cues to predict human attention. Our model is explainable and lightweight, and validated to be effective by a wide range of comparison, ablation, and visualization experiments on two public datasets. Zhixiong Nan, Leiyu Jia, Bin Xiao 0002 |
IEEE Trans. Image Process. | 1 |
| 2024 | EMVCC: Enhanced Multi-View Contrastive Clustering for Hyperspectral ImagesabstractCross-view consensus representation plays a critical role in hyperspectral images (HSIs) clustering. Recent multi-view contrastive cluster methods utilize contrastive loss to extract contextual consensus representation. However, these methods have a fatal flaw: contrastive learning may treat similar heterogeneous views as positive sample pairs and dissimilar homogeneous views as negative sample pairs. At the same time, the data representation via self-supervised contrastive loss is not specifically designed for clustering. Thus, to tackle this challenge, we propose a novel multi-view clustering method, i.e., Enhanced Multi-View Contrastive Clustering (EMVCC). First, the spatial multi-view is designed to learn the diverse features for contrastive clustering, and the globally relevant information of spectrum-view is extracted by Transformer, enhancing the spatial multi-view differences between neighboring samples. Then, a joint self-supervised loss is designed to constrain the consensus representation from different perspectives to efficiently avoid false negative pairs. Specifically, to preserve the diversity of multi-view information, the features are enhanced by using probabilistic contrastive loss, and the data is projected into a semantic representation space, ensuring that the similar samples in this space are closer in distance. Finally, we design a novel clustering loss that aligns the view feature representation with high confidence pseudo-labels for promoting the network to learn cluster-friendly features. In the training process, the joint self-supervised loss is used to optimize the cross-view features.Abundant experiment studies on numerous benchmarks verify the superiority of EMVCC in comparison to some state-of-the-art clustering methods. The codes are available at https://github.com/YiLiu1999/EMVCC. Fulin Luo, Yi Liu 0038, Xiuwen Gong, Zhixiong Nan, Tan Guo |
ACM Multimedia | 4 |
| 2024 | On-Road Object Importance Estimation: A New Dataset and A Model with Multi-Fold Top-Down GuidanceabstractThis paper addresses the problem of on-road object importance estimation, which utilizes video sequences captured from the driver's perspective as the input. Although this problem is significant for safer and smarter driving systems, the exploration of this problem remains limited. On one hand, publicly-available large-scale datasets are scarce in the community. To address this dilemma, this paper contributes a new large-scale dataset named Traffic Object Importance (TOI). On the other hand, existing methods often only consider either bottom-up feature or single-fold guidance, leading to limitations in handling highly dynamic and diverse traffic scenarios. Different from existing methods, this paper proposes a model that integrates multi-fold top-down guidance with the bottom-up feature. Specifically, three kinds of top-down guidance factors (i.e., driver intention, semantic context, and traffic rule) are integrated into our model. These factors are important for object importance estimation, but none of the existing methods simultaneously consider them. To our knowledge, this paper proposes the first on-road object importance estimation model that fuses multi-fold top-down guidance factors with bottom-up feature. Extensive experiments demonstrate that our model outperforms state-of-the-art methods by large margins, achieving 23.1% Average Precision (AP) improvement compared with the recently proposed model (i.e., Goal). Zhixiong Nan, Yilong Chen 0004, Tianfei Zhou, Tao Xiang 0001 |
NeurIPS | 1 |
| 2024 | DI-MaskDINO: A Joint Object Detection and Instance Segmentation ModelabstractThis paper is motivated by an interesting phenomenon: the performance of object detection lags behind that of instance segmentation (i.e., performance imbalance) when investigating the intermediate results from the beginning transformer decoder layer of MaskDINO (i.e., the SOTA model for joint detection and segmentation). This phenomenon inspires us to think about a question: will the performance imbalance at the beginning layer of transformer decoder constrain the upper bound of the final performance? With this question in mind, we further conduct qualitative and quantitative pre-experiments, which validate the negative impact of detection-segmentation imbalance issue on the model performance. To address this issue, this paper proposes DI-MaskDINO model, the core idea of which is to improve the final performance by alleviating the detection-segmentation imbalance. DI-MaskDINO is implemented by configuring our proposed De-Imbalance (DI) module and Balance-Aware Tokens Optimization (BATO) module to MaskDINO. DI is responsible for generating balance-aware query, and BATO uses the balance-aware query to guide the optimization of the initial feature tokens. The balance-aware query and optimized feature tokens are respectively taken as the Query and Key&Value of transformer decoder to perform joint object detection and instance segmentation. DI-MaskDINO outperforms existing joint object detection and instance segmentation models on COCO and BDD100K benchmarks, achieving +1.2 $AP^{box}$ and +0.9 $AP^{mask}$ improvements compared to SOTA joint detection and segmentation model MaskDINO. In addition, DI-MaskDINO also obtains +1.0 $AP^{box}$ improvement compared to SOTA object detection model DINO and +3.0 $AP^{mask}$ improvement compared to SOTA segmentation model Mask2Former. Zhixiong Nan, Xianghong Li, Tao Xiang 0001, Jifeng Dai |
NeurIPS | 1 |
| 2024 | Intention action anticipation model with guide-feedback loop mechanism
Zongnan Ma, Fuchun Zhang, Zhixiong Nan |
Knowl. Based Syst. | 3 |
| 2024 | Third-Person View Attention Prediction in Natural Scenarios With Weak Information Dependency and Human-Scene Interaction MechanismabstractFirst-person view attention has been widely studied in computer science domain since 1990s while third-person view attention in natural scenarios begins to gain the intensive interest until 2015. This paper focuses on the problem of third-person view attention prediction in natural scenarios where a human freely performs daily activities without constraints. To handle the two insuffiencies of existing methods: (i) assuming some extra information (except for input images) are given in advance and (ii) ignoring the importance of human-scene interaction, this paper proposes a model with weak information dependency, which helps to alleviate annotation costs. In addition, a transformer-based human-scene interaction mechanism is proposed to explore the global and long-dependency contexts between the human and scene. The pipeline of the proposed model is firstly extracting human and scene features, then inferring human attention probability map by fusing human and scene features via a transformer-based network, and finally predicting human attention object based on human attention probability map and object detection. The experiments on two public datasets validate the effectiveness of our model. Zhixiong Nan, Tao Xiang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2024 | Video Violence Rating: A Large-Scale Public Database and A Multimodal Rating ModelabstractRecognizing violence in videos is significant for the automatic identification and assessment of violence content to restrict the access to violence for specific audiences such as children. Existing methods focus on violence detection, which is only able to recognize whether there exists violence or not. Differently, this paper handles the problem of video violence rating, which provides a more granular classification of violence levels. However, there is no publicly available database for video violence rating since it asks for fine-grained violence level annotations. Therefore, this paper introduces a large-scale violence rating database, which will be publicly released. Furthermore, we propose a multimodal violence rating model. Different from existing models, our model makes use of the token-based interaction and contrastive learning techniques. The token-based interaction is able to strengthen the feature representations and make full use of multimodal features. The contrastive learning can improve the performance of the model. To evaluate our model, a wide range of experiments are conducted, and experiment results show that our model outperforms existing methods. We have made the VioShot database publicly available for downloading athttps://sites.google.com/site/xiangtaooo/. Tao Xiang 0001, Hongyan Pan, Zhixiong Nan |
IEEE Trans. Multim. | 3 |
| 2023 | FBLNet: FeedBack Loop Network for Driver Attention PredictionabstractThe problem of predicting driver attention from the driving perspective is gaining increasing research focus due to its remarkable significance for autonomous driving and assisted driving systems. The driving experience is extremely important for safe driving, a skilled driver is able to effortlessly predict oncoming danger (before it becomes salient) based on the driving experience and quickly pay attention to the corresponding zones. However, the nonobjective driving experience is difficult to model, so a mechanism simulating the driver experience accumulation procedure is absent in existing methods, and the current methods usually follow the technique line of saliency prediction methods to predict driver attention. In this paper, we propose a FeedBack Loop Network (FBLNet), which attempts to model the driving experience accumulation procedure. By over-and-over iterations, FBLNet generates the incremental knowledge that carries rich historically-accumulative and long-term temporal information. The incremental knowledge in our model is like the driving experience of humans. Under the guidance of the incremental knowledge, our model fuses the CNN feature and Transformer feature that are extracted from the input image to predict driver attention. Our model exhibits a solid advantage over existing methods, achieving an outstanding performance improvement on two driver attention benchmark datasets. Yilong Chen 0004, Zhixiong Nan, Tao Xiang 0001 |
ICCV | 2 |
| 2023 | A Fast and Map-Free Model for Trajectory Prediction in TrafficsabstractTo handle the two shortcomings of existing methods, (i) nearly all models rely on high-definition (HD) maps, yet the map information is not always available in real traffic scenes and HD map-building is expensive and time-consuming and (ii) existing models usually focus on improving prediction accuracy at the expense of reducing computing efficiency, yet the efficiency is crucial for various real applications, this paper proposes an efficient trajectory prediction model that is not dependent on traffic maps. The core idea of our model is encoding single-agent's spatial-temporal information in the first stage and exploring multi-agents' spatial-temporal interactions in the second stage. By comprehensively utilizing attention mechanism, LSTM, graph convolution network and temporal transformer in the two stages, our model is able to learn rich dynamic and interaction information of all agents. Our model achieves the highest performance when comparing with existing map-free methods and also exceeds most map-based state-of-the-art methods on the Argoverse dataset. In addition, our model also exhibits a faster inference speed than the baseline methods. Junhong Xiang, Jingmin Zhang, Zhixiong Nan |
IROS | 3 |
| 2023 | ECCA: Efficient Correntropy-Based Clustering Algorithm With Orthogonal Concept FactorizationabstractOne of the hottest topics in unsupervised learning is how to efficiently and effectively cluster large amounts of unlabeled data. To address this issue, we propose an orthogonal conceptual factorization (OCF) model to increase clustering effectiveness by restricting the degree of freedom of matrix factorization. In addition, for the OCF model, a fast optimization algorithm containing only a few low-dimensional matrix operations is given to improve clustering efficiency, as opposed to the traditional CF optimization algorithm, which involves dense matrix multiplications. To further improve the clustering efficiency while suppressing the influence of the noises and outliers distributed in real-world data, an efficient correntropy-based clustering algorithm (ECCA) is proposed in this article. Compared with OCF, an anchor graph is constructed and then OCF is performed on the anchor graph instead of directly performing OCF on the original data, which can not only further improve the clustering efficiency but also inherit the advantages of the high performance of spectral clustering. In particular, the introduction of the anchor graph makes ECCA less sensitive to changes in data dimensions and still maintains high efficiency at higher data dimensions. Meanwhile, for various complex noises and outliers in real-world data, correntropy is introduced into ECCA to measure the similarity between the matrix before and after decomposition, which can greatly improve the clustering effectiveness and robustness. Subsequently, a novel and efficient half-quadratic optimization algorithm was proposed to quickly optimize the ECCA model. Finally, extensive experiments on different real-world datasets and noisy datasets show that ECCA can archive promising effectiveness and robustness while achieving tens to thousands of times the efficiency compared with other state-of-the-art baselines. Ben Yang, Xuetao Zhang 0001, Feiping Nie 0001, Badong Chen, Fei Wang 0008, Zhixiong Nan, Nanning Zheng 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 6 |
| 2022 | Efficient correntropy-based multi-view clustering with anchor graph embedding
Ben Yang, Xuetao Zhang 0001, Badong Chen, Feiping Nie 0001, Zhiping Lin 0001, Zhixiong Nan |
Neural Networks | 6 |
| 2022 | Hierarchical Motion Planning for Autonomous Driving in Large-Scale Complex ScenariosabstractMotion planning algorithms, an essential part of the autonomous driving system, have been extensively studied. However, in large-scale complex scenarios, how to develop an optimal path to comply with the requirements of smoothness and safety remains a vital issue. In this study, a hierarchical search spacial scales-based hybrid A* (termed as HHA*) motion planning method is proposed, capable of efficiently generating smooth and safe paths. The proposed HHA* method covers two stages. First, the search space is divided on a coarse scale to generate local goals. Subsequently, the novel heuristic function and exploration strategies are adopted in the fine-scale search space to generate paths like that with a human driver guided by the local goals. Moreover, with the usage of the clothoid, the smoothness of the generated path is improved to be G2–continuous (i.e., curvature continuous), which fits the vehicle’s kinematic constraints without the need for later smoothing. Numerous experimental results from the simulation and on-road tests indicate that the proposed method can effectively perform motion planning that meets smoothness and safety in large-scale complex scenarios. Songyi Zhang, Zhiqiang Jian, Xiaodong Deng, Shi-tao Chen, Zhixiong Nan, Nanning Zheng 0001 |
IEEE Trans. Intell. Transp. Syst. | 5 |
| 2021 | DSP-Net: Dense-to-Sparse Proposal Generation Approach for 3D Object Detection on Point CloudabstractObject proposals generated based on sparse points from the raw point cloud have been widely used in 3D object detection. However, following the above scheme, most existing proposal generators have two problems, one is that the features for proposal generation constrain the detection performance by containing insufficient information; the other is that the sparse points obtained from the raw point cloud are misaligned with their corresponding objects in location and feature aspects. In this paper, we propose a dense-to-sparse proposal generation approach for 3D object detection, which can deal with the two problems simultaneously. Our approach utilizes the 3D CNN backbone to output dense features as a supplement to the original sparse point features for proposal generation. Besides, an object-aware feature pooling module is designed to address the misalignment between sparse points and corresponding objects. Experiments on the KITTI dataset show that our method outperforms the existing sparse-style methods and other published state-of-the-art methods. Xinrui Yan, Shi-tao Chen, Zhixiong Nan, Jingmin Xin, Nanning Zheng 0001 |
IJCNN | 4 |
| 2021 | MPR-Net: Multi-Scale Key Points Regression for Lane DetectionabstractLane detection is usually regarded as a semantic segmentation task, however, segmentation-based methods require expensive computational costs, and it is difficult to segment the lanes with heavy noises. Therefore, this paper proposes a lane detection method based on multi-scale key point regression, which directly uses the position of the key points of the lane instead of predicting the pixel-wise outputs. First, we design a lightweight backbone to extract a series of feature maps with a forward view image as input, and then apply a multi-scale fusion network on these feature maps to obtain the location and confidence information of the key points of the lane. Finally, a clustering and curve fitting mechanism with quadratic inverse proportion are used to obtain the final lane detection. Our proposed model can recognize dashed lane markings and handle many extreme scenarios where lanes are completely occluded or heavily noised. In addition, our model uses a relatively explicit framework, which contributes to ensuring real-time performance at 30Hz. In order to prove our method's performance, we conduct experiments on the TuSimple benchmark and RVD dataset, and results demonstrate that our method achieves competitive results compared with other methods. Dantong Zhu, Shengqi Wang, Shi-tao Chen, Zhixiong Nan, Nanning Zheng 0001 |
IV | 5 |
| 2021 | X-GGM: Graph Generative Modeling for Out-of-distribution Generalization in Visual Question AnsweringabstractEncouraging progress has been made towards Visual Question Answering (VQA) in recent years, but it is still challenging to enable VQA models to adaptively generalize to out-of-distribution (OOD) samples. Intuitively, recompositions of existing visual concepts (i.e., attributes and objects) can generate unseen compositions in the training set, which will promote VQA models to generalize to OOD samples. In this paper, we formulate OOD generalization in VQA as a compositional generalization problem and propose a graph generative modeling-based training scheme (X-GGM) to handle the problem implicitly. X-GGM leverages graph generative modeling to iteratively generate a relation matrix and node representations for the predefined graph that utilizes attribute-object pairs as nodes. Furthermore, to alleviate the unstable training issue in graph generative modeling, we propose a gradient distribution consistency loss to constrain the data distribution with adversarial perturbations and the generated distribution. The baseline VQA model (LXMERT) trained with the X-GGM scheme achieves state-of-the-art OOD performance on two standard VQA OOD benchmarks, i.e., VQA-CP v2 and GQA-OOD. Extensive ablation studies demonstrate the effectiveness of X-GGM components. Jingjing Jiang, Ziyi Liu 0001, Yifan Liu 0014, Zhixiong Nan, Nanning Zheng 0001 |
ACM Multimedia | 4 |
| 2021 | Predicting short-term next-active-object through visual attention and hand position
Jingjing Jiang, Zhixiong Nan, Hui Chen 0036, Shi-tao Chen, Nanning Zheng 0001 |
Neurocomputing | 2 |
| 2021 | A joint object detection and semantic segmentation model with cross-attention and inner-attention mechanisms
Zhixiong Nan, Jizhi Peng, Jingjing Jiang, Hui Chen 0036, Ben Yang, Jingmin Xin, Nanning Zheng 0001 |
Neurocomputing | 1 |
| 2021 | Predicting Task-Driven Attention via Integrating Bottom-Up Stimulus and Top-Down GuidanceabstractTask-free attention has gained intensive interest in the computer vision community while relatively few works focus on task-driven attention (TDAttention). Thus this paper handles the problem of TDAttention prediction in daily scenarios where a human is doing a task. Motivated by the cognition mechanism that human attention allocation is jointly controlled by the top-down guidance and bottom-up stimulus, this paper proposes a cognitively-explanatory deep neural network model to predict TDAttention. Given an image sequence, bottom-up features, such as human pose and motion, are firstly extracted. At the same time, the coarse-grained task information and fine-grained task information are embedded as a top-down feature. The bottom-up features are then fused with the top-down feature to guide the model to predict TDAttention. Two public datasets are re-annotated to make them qualified for TDAttention prediction, and our model is widely compared with other models on the two datasets. In addition, some ablation studies are conducted to evaluate the individual modules in our model. Experiment results demonstrate the effectiveness of our model. Zhixiong Nan, Jingjing Jiang, Xiaofeng Gao 0002, Sanping Zhou, Weiliang Zuo, Ping Wei 0001, Nanning Zheng 0001 |
IEEE Trans. Image Process. | 1 |
| 2020 | A Deep Model for Joint Object Detection and Semantic Segmentation in Traffic ScenesabstractObject detection and semantic segmentation are two fundamental techniques of various applications in the fields of Intelligent Vehicles (IV) and Advanced Driving Assistance System (ADAS). Early studies separately handle these two problems. In this paper, inspired by some recent works, we propose a deep neural network model for joint object detection and semantic segmentation. Given an image, an encoder-decoder convolution network extracts a set of feature maps, these feature maps are shared by the detection branch and the segmentation branch to jointly carry out the object detection and semantic segmentation. In the detection branch, we design a PriorBox initialization mechanism to propose more object candidates. In the segmentation branch, we use the multi-scale atrous convolution to explore the global and local semantic information in traffic scenes. Benefiting from the PriorBox Initialization Mechanism (PBIM) and Multi-Scale Atrous Convolution (MSAC), our model presents the competitive performance. In the experiments, we widely compare with several recently-proposed methods on the public Cityscapes dataset, achieving the highest accuracy. In addition, to verify the robustness and generalization of our model, the extension experiments are also conducted on the well-known VOC2012 dataset. Jizhi Peng, Zhixiong Nan, Linhai Xu, Jingmin Xin, Nanning Zheng 0001 |
IJCNN | 2 |
| 2020 | MuRF-Net: Multi-Receptive Field Pillars for 3D Object Detection from Point CloudabstractIn this paper, we propose a point cloud based 3D object detection framework that accounts for both contextual and local information by leveraging multi-receptive field pillars, named as MuRF-Net. Recently, common pipelines can be divided into a voxel-based feature encoder and an object detector. During the feature encoding steps, contextual information is neglected, which is critical for the 3D object detection task. Thus, the encoded features are not suitable to input to the subsequent object detector. To address this challenge, we propose the MuRF-Net with a multi-receptive field voxelization mechanism to capture both contextual and local information. After the voxelization, the voxelized points (pillars) are processed by a feature encoder, and a channel-wise feature reconfiguration module is proposed to combine the features with different receptive fields using a lateral enhanced fusion network. In addition, to handle the increase of memory and computational cost brought by multi-receptive field voxelization, a dynamic voxel encoder is applied taking advantage of the sparseness of the point cloud. Experiments on the KITTI benchmark for both 3D object and Bird's Eye View (BEV) detection tasks on car class are conducted and MuRF-Net achieved the state-of-the-art results compared with other voxel-based methods. Besides, the MuRF-Net can achieve nearly real-time speed with 20Hz. Xinrui Yan, Shi-tao Chen, Jinpeng Dong, Ziyi Liu 0001, Zhixiong Nan, Nanning Zheng 0001 |
IV | 8 |
| 2020 | Traffic Agent Trajectory Prediction Using Social Convolution and Attention MechanismabstractThe trajectory prediction is significant for the decision-making of autonomous driving vehicles. In this paper, we propose a model to predict the trajectories of target agents around an autonomous vehicle. The main idea of our method is considering the history trajectories of the target agent and the influence of surrounding agents on the target agent. To this end, we encode the target agent history trajectories as an attention mask and construct a social map to encode the interactive relationship between the target agent and its surrounding agents. Given a trajectory sequence, the LSTM networks are firstly utilized to extract the features for all agents, based on which the attention mask and social map are formed. Then, the attention mask and social map are fused to get the fusion feature map, which is processed by the social convolution to obtain a fusion feature representation. Finally, this fusion feature is taken as the input of a variable-length LSTM to predict the trajectory of the target agent. We note that the variable-length LSTM enables our model to handle the case that the number of agents in the sensing scope is highly dynamic in traffic scenes. To verify the effectiveness of our method, we widely compare with several methods on a public dataset, achieving a 20% error decrease. In addition, the model satisfies the real-time requirement with the 32 fps. Tao Yang 0032, Zhixiong Nan, Shi-tao Chen, Nanning Zheng 0001 |
IV | 2 |
| 2020 | A Driving Behavior Recognition Model with Bi-LSTM and Multi-Scale CNNabstractIn autonomous driving, perceiving the driving behaviors of surrounding agents is important for the ego-vehicle to make a reasonable decision. In this paper, we propose a neural network model based on trajectories information for driving behavior recognition. Unlike existing trajectory-based methods that recognize the driving behavior using the handcrafted features or directly encoding the trajectory, our model involves a Multi-Scale Convolutional Neural Network (MSCNN) module to automatically extract the high-level features which are supposed to encode the rich spatial and temporal information. Given a trajectory sequence of an agent as the input, firstly, the Bi-directional Long Short Term Memory (Bi-LSTM) module and the MSCNN module respectively process the input, generating two features, and then the two features are fused to classify the behavior of the agent. We evaluate the proposed model on the public BLVD dataset, achieving a satisfying performance. Zhixiong Nan, Tao Yang 0032, Yifan Liu 0014, Nanning Zheng 0001 |
IV | 2 |
| 2020 | ATV Navigation in Complex and Unstructured Environment Containing StairsabstractSelf-driving and robotic technologies have been widely used in recent years. However, before L5 autonomy has been mature, it is more important and meaningful to apply these technologies to some wheeled robots on specific scenes. All-terrain vehicles (ATV) are widely used in urban search and rescue. An unstructured complex scene on which an ATV works usually contains many unusual obstacles, such as stairs. As a result, it is important for autonomous ATVs to detect, localize, and traverse stairs. In this paper, a real-time outdoor stair detection and localization method is proposed. A VLP-16 LIDAR is used to collect environment data, and the stairs are detected and localized using a single frame of LIDAR data by their geometric features, such as slope and parallel edges. A stair navigation strategy is also proposed in this paper. 3-axis attitudes measured by IMU are used during stair climbing on the basis of the coupling between roll and yaw on a slope. Experiments are carried out and the results prove the robustness and accuracy of detection and localization algorithm. The navigation strategy is proven to be safe and feasible. Kongtao Zhu, Junxiang Zhan, Shi-tao Chen, Zhixiong Nan, Tangyike Zhang, Dantong Zhu, Nanning Zheng 0001 |
IV | 4 |
| 2020 | Learning to infer human attention in daily activities
Zhixiong Nan, Tianmin Shu, Shu Wang 0002, Ping Wei 0001, Song-Chun Zhu, Nanning Zheng 0001 |
Pattern Recognit. | 1 |
| 2019 | Recognizing Unseen Attribute-Object Pair with Generative ModelabstractIn this paper, we are studying the problem of recognizing attribute-object pairs that do not appear in the training dataset, which is called unseen attribute-object pair recognition. Existing methods mainly learn a discriminative classifier or compose multiple classifiers to tackle this problem, which exhibit poor performance for unseen pairs. The key reasons for this failure are 1) they have not learned an intrinsic attributeobject representation, and 2) the attribute and object are processed either separately or equally so that the inner relation between the attribute and object has not been explored. To explore the inner relation of attribute and object as well as the intrinsic attribute-object representation, we propose a generative model with the encoder-decoder mechanism that bridges visual and linguistic information in a unified end-to-end network. The encoder-decoder mechanism presents the impressive potential to find an intrinsic attribute-object feature representation. In addition, combining visual and linguistic features in a unified model allows to mine the relation of attribute and object. We conducted extensive experiments to compare our method with several state-of-the-art methods on two challenging datasets. The results show that our method outperforms all other methods. Zhixiong Nan, Yang Liu 0266, Nanning Zheng 0001, Song-Chun Zhu |
AAAI | 1 |
| 2019 | Scene-Guided Region Proposal Re-ranking Method for On-road Vehicle Candidate GenerationabstractVehicle candidate generation is important for vehicle detection. Existing vehicle detection studies usually employ general-purpose region proposal methods to generate vehicle candidates, which do not consider the specificity of on-road vehicles in traffic scenes. In this paper, we propose a model to re-rank the candidates that are generated by general-purpose region proposal methods. Our model considers the specificity of on-road vehicle candidate generation in traffic scenes by encoding global-local semantic context and location-size geometric compatibility. In the experiments, we test our model on three art-of-the-state region proposal methods using two public datasets. The results show the significant performance improvement is gained after applying our model. Zhixiong Nan, Jiawei He 0002, Ping Wei 0001, Linhai Xu, Hongbin Sun 0001, Nanning Zheng 0001 |
IV | 1 |
| 2018 | Exploring the Potential of Using Semantic Context and Common Sense in On-Road Vehicle DetectionabstractVehicle detection is an important research topic for autonomous driving community. Since the great success of deep learning on object detection, almost all vehicle detection methods go along with this line. However, deep learning methods heavily rely on the training data, and the whole mechanism is like a “black box” Therefore, in this paper, we explore a vehicle detection method using traffic semantic context and human common sense instead of relying on the training data. To verify our idea, we compare our method with two classic machine learning methods as well as three state- of-the-art deep learning methods on a dataset collected in real traffics. The results show that our method outperforms others on this dataset. The deep learning methods may exceed ours after enlarging the training data or testing on more complicated datasets. However, the main contribution of this paper is providing inspiration for learning methods, and we believe their performance can be greatly improved after considering the idea of this paper. Zhixiong Nan, Menghan Pan, Xiao Wang 0002, Ping Wei 0001, Linhai Xu, Hongbin Sun 0001, Jingmin Xin, Nanning Zheng 0001 |
Intelligent Vehicles Symposium | 1 |