Nanning Zheng 0001

dblp:07/256-1 · also Nan-Ning Zheng 0001 · DBLP profile ↗
← Back
632ranked-venue papers
7as first author
256since 2021 · last 2026
0000-0003-1608-8257ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 360 · 4 first-author · 168 since 2021Graphics, computer vision, multimedia, augmented reality and games · 254 · 1 first-author · 84 since 2021Systems, architecture and hardware · 78 · 1 first-author · 28 since 2021Applied, interdisciplinary, general and emerging computing · 51 · 1 first-author · 23 since 2021Databases, data management, data science and information retrieval · 15 · 1 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 10 · 5 since 2021Computer networks · 3 · 1 since 2021Security and privacy · 1 · 1 since 2021Software engineering, systems software and programming languages · 1
YearPublicationVenuePosition
2026 EVOKE: Efficient and High-Fidelity EEG-to-Video Reconstruction via Decoupling Implicit Neural Representation
abstract
Visual neural decoding is an important research topic at the intersection of cognitive neuroscience and machine learning. While recent progress has been made in EEG-based neural decoding, reconstructing dynamic visual content remains challenging. In the field of EEG decoding, current models either utilize pre-trained encoders for feature extraction or employ graph neural networks to represent the spatio-temporal information embedding, resulting in poor model representation and high complexity. We propose EVOKE -- an innovative framework for zero-shot decoding of high-fidelity videos from EEG signals. EVOKE employs Implicit Neural Representations to perform complete spatial modeling of EEG and continuously decouples information in the EEG-INR perceptual space. Additionally, we construct a Hierarchical-aware Attention Module (HAM) to decode EEG from three feature anchors: visual, semantic, motion, and progressively control task inference. The Motion Attention Flow (MAF) we developed overcomes the limitations of capturing motion features in dynamic stimuli, creating a more robust representation that enhances reconstruction consistency. Comprehensive experiments prove that SOTA performance of EVOKE (0.353 SSIM, 0.715 CLIP-pcc). We provide an effective method for converting brain activity into rich visual experiences and set a new benchmark for brain multimodal generation.
Haodong Jing, Panqi Yang, Dongyao Jiang, Nanning Zheng 0001
AAAI5
2026 UniHOI: Unified Human-Object Interaction Understanding via Unified Token Space
abstract
In the field of human-object interaction (HOI), detection and generation are two dual tasks that have traditionally been addressed separately, hindering the development of comprehensive interaction understanding. To address this, we propose UniHOI, which jointly models HOI detection and generation via a unified token space, thereby effectively promoting knowledge sharing and enhancing generalization. Specifically, we introduce a symmetric interaction-aware attention module and a unified semi-supervised learning paradigm, enabling effective bidirectional mapping between images and interaction semantics even under limited annotations. Extensive experiments demonstrate that UniHOI achieves state-of-the-art performance in both HOI detection and generation. Specifically, UniHOI improves accuracy by 4.9% on long-tailed HOI detection and boosts interaction metrics by 42.0% on open-vocabulary generation tasks.
Panqi Yang, Haodong Jing, Nanning Zheng 0001
AAAI3
2026 Think before Go: Hierarchical Reasoning for Image-goal Navigation
abstract
Pengna Li, Kangyi Wu, Shaoqing Xu, Fang Li, Lin Zhao, Long Chen, Zhi-Xin Yang, Nanning Zheng. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Pengna Li, Kangyi Wu, Shaoqing Xu, Long Chen 0005, Nanning Zheng 0001
ACL (1)8
2026 Interactive design of developable surfaces by patch-based learning
Chaoyun Wang, Jianlei Wang, Chengcheng Tang, Nanning Zheng 0001, Caigui Jiang
Comput. Aided Des.4
2026 InstrucRobo: Object-centric multi-instruction decoupling model for explainable robotic manipulation
Panqi Yang, Haodong Jing, Nanning Zheng 0001
Eng. Appl. Artif. Intell.3
2026 UniBVR: Balancing visual and reasoning abilities in unified 3D scene understanding
Panqi Yang, Haodong Jing, Nanning Zheng 0001
Neurocomputing3
2026 Enhancing Value Decomposition With Target Transformation in Cooperative Multi-Agent Reinforcement Learning
abstract
The increasing need for cooperation among intelligent machines has heightened the importance of cooperative multi-agent reinforcement learning (MARL). However, a dominant class of cooperative MARL approaches relies on monotonic value decomposition, which enables scalable decentralized execution but restricts the representable class of joint action-values. However, existing remedies bias learning targets toward high-value samples, which can be fragile under stochastic returns because optimistic emphasis may amplify lucky but suboptimal trajectories. To solve this challenge, we propose Target Transformation, which maps non-monotonic and stochastic learning targets into a monotonic-representable surrogate while preserving the optimal joint action. Building on this idea, we develop Uncertainty-aware Target Transformation (UT2) with value-based and policy-based instantiations that combine an uncertainty estimator with a best-individual coordination envelope. Experiments on diverse cooperative MARL benchmarks show that UT2 improves both performance and stability over strong baselines, with larger gains as non-monotonicity and stochasticity increase.
Zeyang Liu 0001, Lipeng Wan 0003, Shiguang Sun, Xue Sui, Xingyu Chen 0001, Xuguang Lan, Nanning Zheng 0001
IEEE Trans. Pattern Anal. Mach. Intell.7
2026 MonoA2: Adaptive depth with augmented head for monocular 3D object detection
Jinpeng Dong, Sanping Zhou, Jingjing Jiang, Weiliang Zuo, Shi-tao Chen, Nanning Zheng 0001
Pattern Recognit.8
2026 Query-enhanced motion transformer with dilated static query and bridged dynamic query
Miao Kang, Liushuai Shi, Ke Ye, Sanping Zhou, Nanning Zheng 0001
Pattern Recognit.5
2026 Efficient DOA Estimation Based on Coprime Array Interpolation With Deep Unfolding Network
abstract
Coprime arrays increase the degrees of freedom for direction of arrival (DOA) estimation, but virtual-array gaps require costly filling procedures. In this letter, a new DOA estimation method is proposed for coprime arrays based on interpolation and a deep unfolding network. The reconstruction of the interpolated virtual array covariance matrix is formulated as a rank minimization problem and solved using an ADMM-based deep unfolding network with stage-wise learnable parameters and an unsupervised loss inspired by ADMM convergence criteria. Finally, root-MUSIC is employed for DOA estimation. Simulations demonstrate the effectiveness of the proposed method in terms of both computational efficiency and estimation performance.
Zhuoqian Jiang, Jingmin Xin, Weiliang Zuo, Nanning Zheng 0001, Akira Sano
IEEE Signal Process. Lett.4
2026 CylinderPlane: A General Cylindrical Representation for 360° 3D Content Generation
abstract
While Tri-plane representation has greatly advanced the development of 3D generative models, problems rooted in its inherent structure, such as multi-face artifacts caused by sharing the same features in symmetric regions, limit its ability to generate complete 360° views. In this paper, we propose CylinderPlane, a novel representation based on the cylindrical coordinate system, to achieve high-quality, artifact-free panoramic image synthesis. Unlike the inevitable feature entanglement in the Cartesian coordinate-based representation, the cylindrical coordinate system explicitly disentangles features at different angles. Consequently, our representation effectively eliminates feature ambiguity and ensures multi-view consistency across full 360°. We further develop a nested cylinder representation that combines cylinder planes of varying radii to achieve multi-scale feature fusion. This design not only addresses the limitations of Tri-plane in modeling complex geometries and varying resolutions, but also mitigates the polar discontinuity inherent in a single cylinder plane. Moreover, our versatile representation can be seamlessly integrated into various generative frameworks and rendering pipelines. Extensive experiments on both synthetic datasets and unstructured in-the-wild images demonstrate that our representation outperforms the existing methods.
Ru Jia, Xiaozhuang Ma, Jianji Wang 0001, Nanning Zheng 0001
IEEE Trans. Circuits Syst. Video Technol.4
2026 Mind-VAD: Brain-Inspired fMRI-to-Video Precise Reconstruction via Cross-Modal Autoregressive Diffusion
abstract
Decoding visual information from brain activity is important and challenging. Existing studies have successfully reconstructed static images from fMRI signals, but fMRI-based dynamic visual reconstruction still has limitations: the low temporal resolution of fMRI limits the coherence of the video, the lack of fine-grained alignment of brain with visual features at different scales, and the deviation of diffusion modeling from the logic of visual encoding. Research has shown that the visual cortex cognitive system streams perceptual and semantic information to different brain regions for processing. Inspired by this, we propose Mind-VAD, a novel dynamic visual reconstruction paradigm that redefines the brain decoding process as cross-modal generation guided by visual stream. Specifically, we design a time-sensitive encoder to accurately achieve temporal localization, and learn different dynamic visual representations of fMRI through an autoencoder with different scales. Then proceed with full-sequence, faster video autoregressive diffusion generation through a coarse-to-fine process. We include the latest AIGC Benchmark in evaluation. Extensive experiments show that Mind-VAD realized accurate and smooth reconstruction, achieves SOTA performance (87.3% classification, 27.5% semantic improvement) on several downstream tasks and more demanding metrics. Overall, Mind-VAD provides a neural interpretable and powerful framework for visual decoding and its potential applications.
Haodong Jing, Wenjie Gao 0001, Dongyao Jiang, Shuai Huang 0002, Nanning Zheng 0001
IEEE Trans. Circuits Syst. Video Technol.6
2026 DAMind: Zero-Shot Visual Cross-Domain Alignment and Representation for EEG Decoding
abstract
To efficiently assist humans in various tasks, it is crucial to accurately decode and understand the rich information embedded in brain's visual cognition. Existing brain-driven research often fails to overcome the challenge of small target data domains, and the lack of explicit semantic, spatial, and other information constraints on feature extractors prevents brain decoding models from learning uniform cross-domain representations, leading to degradation of their performance in unseen domains. To overcome these limitations, we propose DAMind, a multimodal EEG-based model for robust visual cross-domain alignment and decoding. Our approach integrates VLM with brain-inspired cognitive mechanisms, leveraging the strong image-text representation abilities to learn both fine-grained primary visual features and high-level semantic concepts from neural signals, provide effective visual fine-tuning using the visual guidance mechanism. DAMind introduces a stepwise EEG encoding process aligned with visual processing, and employs an instruction-based learning strategy for effective cross-domain zero-shot transfer. Its robust architecture efficiently achieves good generalization performance, enabling the mapping of EEG signals from multiple domains to a unified learning domain. We construct a comprehensive EEG decoding benchmark EBench, DAMind achieves state-of-the-art results on several visual tasks, and outperforms the baseline in zero-shot setting.
Haodong Jing, Panqi Yang, Shuai Huang 0002, Badong Chen, Nanning Zheng 0001
IEEE Trans. Image Process.7
2026 CNDM: Customized Noise Diffusion Model for Trajectory Prediction
abstract
Precise and reliable multi-agent trajectory prediction is fundamental to enabling safe autonomous navigation. While denoising diffusion models have demonstrated remarkable capability in capturing the multimodal uncertainty inherent in this task, their reliance on a generic, isotropic Gaussian noise prior poses a critical limitation. This uniform prior is fundamentally misaligned with the highly structured, heterogeneous uncertainty of agent motion—shaped by dynamics, type, and environmental constraints—forcing the denoising network to implicitly relearn complex motion priors from scratch. To bridge this gap, we propose the Customized Noise Diffusion Model (CNDM), a novel framework that introduces a learned, agent-specific noise prior. At the core of CNDM is a Prior-Guidance Network (PGN) that distills an agent’s history, type, and scene context into a parametric, anisotropic Gaussian distribution. This customized prior provides a physically-grounded starting point for the diffusion process. To enable efficient training on such anisotropic noise, we leverage a Mahalanobis whitening transformation to standardize the denoising task. Extensive experiments on the Waymo Open Motion and Argoverse 2 datasets show that CNDM achieves competitive performance against state-of-the-art methods, excelling particularly in capturing diverse motion patterns and improving probabilistic calibration. Ablation studies confirm that the performance gains are directly attributable to our customized noise design, underscoring the importance of integrating structured domain knowledge into the generative foundation of diffusion models.
Entao Chang, Jiawei Fu 0001, Wenjie Gao 0001, Shi-tao Chen, Jingmin Xin, Nanning Zheng 0001
IEEE Trans. Intell. Transp. Syst.6
2026 FlowCalib: Targetless Infrastructure LiDAR-Camera Extrinsic Calibration Based on Optical Flow and Scene Flow
abstract
Recently, multi-sensor fusion-based vehicle infrastructure cooperative perception has aroused extensive attention due to the demands for the safety of autonomous driving and traffic monitoring. An accurate calibration between different sensors is a critical foundation for most sensor fusion systems. For LiDAR-camera calibration, high accuracy can be achieved with the help of artificial calibration targets, such as a checkerboard. However, unlike autonomous vehicles, roadside sensors monitor traffic scenes with continuous traffic flow from a fixed viewpoint, posing challenges for conventional calibration methods. There, a calibration method suitable for roadside scenes is required for infrastructure sensors. In this paper, we propose FlowCalib, a novel targetless infrastructure LiDAR-camera spatial calibration method through alignment of scene flow and optical flow. The main idea is to leverage the inherent consistency of moving objects in traffic flow across two types of sensor data. Firstly, the moving objects are extracted by optical flow and scene flow. Then, the extrinsic parameters are obtained in two steps: rough calibration and calibration refinement. In rough calibration, the center and motion flow of each moving instance are calculated by clustering methods separately in the point cloud and image. Based on this, the possible initial value set of extrinsic parameters is estimated by two-step parameter sampling. The initial parameters are obtained by distance of center and motion flow in point cloud and image based scoring. Subsequently, the extrinsic parameters are refined by optimization of instance alignment loss and flow alignment loss of moving objects. In the end, quantitative and qualitative experiments are conducted to validate the effectiveness of the algorithm across both simulated datasets and real-world datasets.
Renwei Hai, Yanqing Shen, Shi-tao Chen, Jingmin Xin, Nanning Zheng 0001
IEEE Trans. Intell. Transp. Syst.6
2025 See Through Their Minds: Learning Transferable Brain Decoding Models from Cross-Subject fMRI
abstract
Deciphering visual content from fMRI sheds light on the human vision system, but data scarcity and noise limit brain decoding model performance. Traditional approaches rely on subject-specific models, which are sensitive to training sample size. In this paper, we address data scarcity by proposing shallow subject-specific adapters to map cross-subject fMRI data into unified representations. A shared deep decoding model then decodes these features into the target feature space. We use both visual and textual supervision for multi-modal brain decoding and integrate high-level perception decoding with pixel-wise reconstruction guided by high-level perceptions. Our extensive experiments reveal several interesting insights: 1) Training with cross-subject fMRI benefits both high-level and low-level decoding models; 2) Merging high-level and low-level information improves reconstruction performance at both levels; 3) Transfer learning is effective for new subjects with limited training data by training new adapters; 4) Decoders trained on visually-elicited brain activity can generalize to decode imagery-induced activity, though with reduced performance.
Guibo Zhu, Haodong Jing, Nanning Zheng 0001
AAAI5
2025 Unveiling Multi-View Anomaly Detection: Intra-view Decoupling and Inter-view Fusion
abstract
Anomaly detection has garnered significant attention for its extensive industrial application value. Most existing methods focus on single-view scenarios and fail to detect anomalies hidden in blind spots, leaving a gap in addressing the demands of multi-view detection in practical applications. Ensemble of multiple single-view models is a typical way to tackle the multi-view situation, but it overlooks the correlations between different views. In this paper, we propose a novel multi-view anomaly detection framework, Intra-view Decoupling and Inter-view Fusion (IDIF), to explore correlations among views. Our method contains three key components: 1) a proposed Consistency Bottleneck module extracting the common features of different views through information compression and mutual information maximization; 2) an Implicit Voxel Construction module fusing features of different views with prior knowledge represented in the form of voxels; and 3) a View-wise Dropout training strategy enabling the model to learn how to cope with missing views during test. The proposed IDIF achieves state-of-the-art performance on three datasets. Extensive ablation studies also demonstrate the superiority of our methods.
Yiyang Lian, Meiqin Liu 0001, Nanning Zheng 0001, Ping Wei 0001
AAAI5
2025 Beyond Image Classification: A Video Benchmark and Dual-Branch Hybrid Discrimination Framework for Compositional Zero-Shot Learning
abstract
Human reasoning naturally combines concepts to identify unseen compositions, a capability that Compositional Zero-Shot Learning (CZSL) aims to replicate in machine learning models. However, we observe that focusing solely on typical image classification tasks in CZSL may limit models’ compositional generalization potential. To address this, we introduce C-EgoExo, a video-based benchmark, along with a compositional action recognition task to enable more comprehensive evaluations. Inspired by human reasoning processes, we propose a Dual-branch Hybrid Discrimination (DHD) framework, featuring two branches that decode visual inputs in distinct observation sequences. Through a cross-attention mechanism and a contextual dependency encoder, DHD effectively mitigates challenges posed by conditional variance. We further design a Copula-based orthogonal decoding loss to counteract contextual interference in primitive decoding. Our approach demonstrates outstanding performance across diverse CZSL tasks, excelling in both image-based and video-based modalities and in attribute-object and action-object compositions, setting a new benchmark for CZSL evaluation.
Dongyao Jiang, Haodong Jing, Nanning Zheng 0001
CVPR4
2025 Beyond Single-Modal Boundary: Cross-Modal Anomaly Detection through Visual Prototype and Harmonization
abstract
Anomaly detection is a significant task for its application and research value. While existing methods have made impressive progress within the same modality, cross-modal anomaly detection remains an open and challenging problem. In this paper, we propose a cross-modal anomaly detection model that is trained using data from a variety of existing modalities and can be generalized well to unseen modalities. The model consists of three major components: 1) the Transferable Visual Prototype directly learns normal/abnormal semantics in visual space; 2) the Prototype Harmonization strategy adaptively utilizes the Transferable Visual Prototypes from various modalities for inference on the unknown modality; 3) the Visual Discrepancy Inference under the few-shot setting enhances performance. In the zero-shot setting, the proposed method achieves AUROC improvements of 4.1%, 6.1%, 7.6%, and 6.8% over the best competing methods in the RGB, 3D, MRI/CT, and Thermal modalities, respectively. In the few-shot setting, our model also achieves the highest AUROC/AP on ten datasets in four modalities, substantially outperforming existing methods. Codes are available at https://github.com/Kerio99/CMAD.
Ping Wei 0001, Yiyang Lian, Nanning Zheng 0001
CVPR5
2025 ForestLPR: LiDAR Place Recognition in Forests Attentioning Multiple BEV Density Images
abstract
Place recognition is essential to maintain global consistency in large-scale localization systems. While research in urban environments has progressed significantly using LiDARs or cameras, applications in natural forest-like environments remain largely under-explored. Furthermore, forests present particular challenges due to high self-similarity and substantial variations in vegetation growth over time. In this work, we propose a robust LiDAR-based place recognition method for natural forests, ForestLPR. We hypothesize that a set of cross-sectional images of the forest’s geometry at different heights contains the information needed to recognize revisiting a place. The cross-sectional images are represented by bird’s-eye view (BEV) density images of horizontal slices of the point cloud at different heights. Our approach utilizes a visual transformer as the shared backbone to produce sets of local descriptors and introduces a multi-BEV interaction module to attend to information at different heights adaptively. It is followed by an aggregation layer that produces a rotation-invariant place descriptor. We evaluated the efficacy of our method extensively on real-world data from public benchmarks as well as robotic datasets and compared it against the state-of-the-art (SOTA) methods. The results indicate that ForestLPR has consistently good performance on all evaluations and achieves an average increase of 7.38% and 9.11% on Recall@1 over the closest competitor on intra-sequence loop closure detection and inter-sequence re-localization, respectively, validating our hypothesis1.
Yanqing Shen, Turcan Tuna, Marco Hutter 0001, Cesar Dario Cadena Lerma, Nanning Zheng 0001
CVPR5
2025 Refiner: Fine-grained Cross-modal Concepts Refinement for Compositional Zero-Shot Learning
abstract
Recent Compositional Zero-Shot Learning (CZSL) methods increasingly adopt the pre-trained vision-language models to capture the contextual relations between image and text spaces. However, the single-class-token design from Transformer-based encoder inevitably captures contextual information from unrelated objects and background, thus hindering the modeling of fine-grained class-specific visual features. Suffering from cross-modal gap, prior methods also struggle to improve compositional recognition performance. To address these issues, we propose a fine-grained cross-modal concepts refinement framework, termed as Refiner, which comprises two pivotal components: (i) the fine-grained concepts refinement of image embeddings to capture state-object context within visual scenes, and (ii) the cross-modal information fusion to mitigate the modality gap. By leveraging learnable query vectors to capture region-specific semantic information pertinent to composition labels, our approach refines visual representations with fine-grained state-object context information. As for cross-modal information fusion, we construct a robust image-to-text mapping by aligning visual embeddings with states, objects, and compositions, respectively. Extensive experiments demonstrate that our Refiner achieves new state-of-the-art performance across all popular benchmarks in both closed- and open-world settings.
Haodong Jing, Hui Chen 0036, Nanning Zheng 0001
ICASSP5
2025 DAMap: Distance-Aware MapNet for High Quality HD Map Construction
Jinpeng Dong, Yutong Lin, Jingwen Fu, Sanping Zhou, Nanning Zheng 0001
ICCV6
2025 Beyond Brain Decoding: Visual-Semantic Reconstructions to Mental Creation Extension Based on fMRI
Haodong Jing, Dongyao Jiang, Haibo Hua, Nanning Zheng 0001
ICCV6
2025 FreqPDE: Rethinking Positional Depth Embedding for Multi-View 3D Object Detection Transformers
abstract
Detecting 3D objects accurately from multi-view 2D images is a challenging yet essential task in the field of autonomous driving. Current methods resort to integrating depth prediction to recover the spatial information for object query decoding, which necessitates explicit supervision from LiDAR points during the training phase. However, the predicted depth quality is still unsatisfactory such as depth discontinuity of object boundaries and indistinction of small objects, which are mainly caused by the sparse supervision of projected points and the use of high-level image features for depth prediction. Besides, cross-view consistency and scale invariance are also overlooked in previous methods. In this paper, we introduce Frequency-aware Positional Depth Embedding (FreqPDE) to equip 2D image features with spatial information for 3D detection transformer decoder, which can be obtained through three main modules. Specifically, the Frequency-aware Spatial Pyramid Encoder (FSPE) constructs a feature pyramid by combining high-frequency edge clues and low-frequency semantics from different levels respectively. Then the Cross-view Scale-invariant Depth Predictor (CSDP) estimates the pixel-level depth distribution with cross-view and efficient channel attention mechanism. Finally, the Positional Depth Encoder (PDE) combines the 2D image features and 3D position embeddings to generate the 3D depth-aware features for query decoding. Additionally, hybrid depth supervision is adopted for complementary depth learning from both metric and distribution aspects. Extensive experiments conducted on the nuScenes dataset demonstrate the effectiveness and superiority of our proposed method.
Haisheng Su, Feixiang Song, Sanping Zhou, Wei Wu 0021, Junchi Yan, Nanning Zheng 0001
ICCV7
2025 On the Statistical Mechanisms of Distributional Compositional Generalization
abstract
Distributional Compositional Generalization (DCG) refers to the ability to tackle tasks from new distributions by leveraging the knowledge of concepts learned from supporting distributions. In this work, we aim to explore the statistical mechanisms of DCG, which have been largely overlooked in previous studies. By statistically formulating the problem, this paper seeks to address two key research questions: 1) Can a method to one DCG problem be applicable to another? 2) What statistical properties can indicate a learning algorithm's capacity for knowledge composition in DCG tasks? \textbf{To address the first question}, an invariant measure is proposed to provide a dimension where all different methods converge. This measure underscores the critical role of data in enabling improvements without trade-offs. \textbf{As for the second question}, we reveal that by decoupling the impacts of insufficient data and knowledge composition, the ability of the learning algorithm to compose knowledge relies on the compatibility and sensitivity between the learning algorithm and the composition rule. In summary, the statistical analysis of the generalization mechanisms provided in this paper deepens our understanding of compositional generalization, offering a complementary evidence on the importance of data in DCG task.
Jingwen Fu, Nanning Zheng 0001
ICML2
2025 PlaneHEC: Efficient Hand-Eye Calibration for Multi-View Robotic Arm via Any Point Cloud Plane Detection
abstract
Hand-eye calibration is an important task in vision-guided robotic systems and is crucial for determining the transformation matrix between the camera coordinate system and the robot end-effector. Existing methods, for multi-view robotic systems, usually rely on accurate geometric models or manual assistance, generalize poorly, and can be very complicated and inefficient. Therefore, in this study, we propose PlaneHEC, a generalized hand-eye calibration method that does not require complex models and can be accomplished using only depth cameras, which achieves the optimal and fastest calibration results using arbitrary planar surfaces like walls and tables. PlaneHEC introduces hand-eye calibration equations based on planar constraints, which makes it strongly interpretable and generalizable. PlaneHEC also uses a comprehensive solution that starts with a closed-form solution and improves it with iterative optimization, which greatly improves accuracy. We comprehensively evaluated the performance of PlaneHEC in both simulated and real-world environments and compared the results with other point-cloud-based calibration methods, proving its superiority. Our approach achieves universal and fast calibration with an innovative design of computational models, providing a strong contribution to the development of multi-agent systems and embodied intelligence.
Haodong Jing, Yang Liao, Nanning Zheng 0001
ICRA5
2025 Novel AI Frameworks for ESG Performance Prediction: ESG-RAG and SSA-iLSTM
abstract
Environmental, Social, and Governance (ESG) issues are becoming a focal point in global investment decision-making. However, accurately analyzing and predicting a company’s ESG performance remains a complex challenge due to the fragmented and heterogeneous nature of ESG data. Traditional approaches often rely on a single predictive model and conventional time series techniques, which face limitations in handling low-frequency, discontinuous, and unstructured data while lacking standardized processing methods. To address these challenges, this study proposes the ESG-RAG (Retrieval-Augmented Generation) framework for in-depth information mining, enabling the extraction of implicit insights from ESG reports to alleviate data sparsity issues. Additionally, a novel temporal prediction framework, SSA-iLSTM (Singular Spectrum Analysis integrated with iTransformer and Long Short-Term Memory), is introduced. The framework processes the target variable’s temporal signals through the iTransformer module while decomposing covariates using SSA and processing them with the LSTM model. A cross-attention mechanism fuses the outputs from both modules, enabling dynamic feature integration across multiple variables. Experimental results using datasets from Wind (Chinese companies) and Refinitiv (US companies) demonstrate that ESG-RAG and SSA-iLSTM significantly outperform baseline models in predicting ESG performance. This approach not only enhances prediction accuracy but also provides investors with valuable ESG insights, supporting more informed investment decisions.
Lingyan Zhao, Nanning Zheng 0001
IJCNN5
2025 SAMap: Semantic Alignment for HD Map Detection Domain Generalization Under Varying Weather and Lighting
abstract
High-definition (HD) maps are crucial for autonomous driving systems. Despite recent advances in learning-based HD map prediction methods, these approaches experience significant performance degradation when encountering unseen weather or lighting conditions due to feature distribution discrepancies (domain gaps) of input images. To address this issue, we propose SAMap, a novel map learning framework that enhances domain generalization capabilities of existing models by reducing domain discrepancies in input images. SAMap innovatively introduces a Semantic Aligner, an image-to-image transformation module that aligns images from different domains into a unified domain space while preserving semantic consistency. To train this aligner, we leverage Vision-Language Models (VLMs) that have acquired image-text alignment capabilities. Specifically, we first train a Prompt Learner that combines handcrafted and learnable prompts to capture domain-invariant semantic information. We then train Semantic Aligner through dual supervision mechanisms: a content preservation loss that maintains feature consistency across transformations and a semantic alignment loss that leverages VLM’s encoders to align transformed images with domain-invariant textual representations. Adequate experiments on the NuScenes dataset demonstrate that when integrated with three existing HD map prediction methods, SAMap achieves a performance improvement of up to 11.6% on unseen domains (rain or night conditions), effectively validating its generalization capabilities across domains.
Wenjie Gao 0001, Haodong Jing, Jiawei Fu 0001, Shi-tao Chen, Nanning Zheng 0001
IROS6
2025 Towards Extrinsic Dexterity Grasping in Unrestricted Environments
abstract
Grasping large and flat objects (e.g., a book or a pan) is often regarded as an ungraspable task, which poses significant challenges due to the unreachable grasping poses. Prior research has exploited environmental interactions through Extrinsic Dexterity, utilizing external structures such as walls or table edges to facilitate object grasping. However, they are confined to task-specific policies while neglecting semantic perception and planning to identify optimal pre-grasp configurations. This limits their operational versatility, impeding effective adaptation to varied extrinsic dexterity constraints. In this work, we present ExDiff, a robot manipulation approach for extrinsic dexterity grasping in unrestricted environments. It utilizes Vision-Language Models (VLMs) to perceive the environmental state and generate instructions, followed by a Goal-Conditioned Action Diffusion (GCAD) model to predict the sequence of low-level actions. This diffusion model learns the low-level policy, conditioned on high-level instructions and cumulative rewards, which improves the generation of robot actions. Simulation experiments and real-world deployment results demonstrate that ExDiff effectively performs ungraspable tasks and generalizes to previously unseen target objects and scenes. Videos at - https://exdiff.github.io/index.html
Chengzhong Ma, Houxue Yang, Hanbo Zhang, Zeyang Liu 0001, Xuguang Lan, Nanning Zheng 0001
IROS8
2025 Modeling Human-like Driving Behavior Based on Maximum Entropy Deep Inverse Reinforcement Learning
abstract
Modeling expert driving behavior is crucial for the successful implementation of human-like autonomous driving. In this paper, we propose a new sampling-based Maximum Entropy Deep Inverse Reinforcement Learning (MEDIRL) framework. It leverages naturalistic human driving data to train the reward model and thus evaluates driving behaviors from the reward of sampled candidate trajectories. The proposed framework utilizes deep neural networks to learn the feature-reward mapping, which offers superior fitting capabilities compared to traditional linear reward functions. A polynomial trajectory sampler for long-term decision making and a dynamic window trajectory sampler for short-term planning are adopted to simplify the calculation of partition function in the MEDIRL algorithm. In addition, the proposed framework offers a solution to the probability estimation of driving behaviors by calculating the likelihood of sampled candidate trajectories based on their reward values. Comparative experiments are conducted on the NGSIM US-101 Highway dataset, and the experimental results demonstrate the superiority of the proposed model in personalizing reward functions, as well as the applicability of the proposed method in modeling driving behaviors across various time horizons.
Jiamin Shi, Tangyike Zhang, Shi-tao Chen, Nanning Zheng 0001, Jingmin Xin
IROS4
2025 HybridPlane: A General 4D Representation for Dynamic Scene Reconstruction
abstract
Despite recent advances in dynamic scene reconstruction, challenges from imbalanced camera distribution and inaccurate pose estimation in real-world datasets still persist, undermining the spatiotemporal consistency of reconstruction. In this paper, we propose HybridPlane, a novel representation that leverages the complementary advantages of cylindrical and Cartesian coordinate systems to achieve high-quality dynamic scene synthesis. Unlike Cartesian projection, which shares identical features in symmetric regions, cylindrical projection explicitly disentangles features from different viewpoints, thereby improving robustness against imbalanced camera distributions. Moreover, the synergy between these two coordinate systems in both projection and representational capacity enhances the model's ability to capture complex motions and fine-grained details. We further adopt the dynamic positional encoding strategy to enhance the smoothness of temporal interpolation under inaccurate camera poses by progressively regulating high-frequency signals without incurring additional computational overhead. Extensive experiments demonstrate that our versatile representation can be seamlessly integrated into various rendering pipelines, outperforming the previous methods in reconstruction quality while reducing computational and memory costs by approximately one-third.
Ru Jia, Xiaoqian Liang, Xubin Duan, Jianji Wang 0001, Nanning Zheng 0001
ACM Multimedia5
2025 Riemannian Consistency Model
abstract
Consistency models are a class of generative models that enable few-step generation for diffusion and flow matching models. While consistency models have achieved promising results on Euclidean domains like images, their applications to Riemannian manifolds remain challenging due to the curved geometry. In this work, we propose the Riemannian Consistency Model (RCM), which, for the first time, enables few-step consistency modeling while respecting the intrinsic manifold constraint imposed by the Riemannian geometry. Leveraging the covariant derivative and exponential-map-based parameterization, we derive the closed-form solutions for both discrete- and continuous-time training objectives for RCM. We then demonstrate theoretical equivalence between the two variants of RCM: Riemannian consistency distillation (RCD) that relies on a teacher model to approximate the marginal vector field, and Riemannian consistency training (RCT) that utilizes the conditional vector field for training. We further propose a simplified training objective that eliminates the need for the complicated differential calculation. Finally, we provide a unique kinematics perspective for interpreting the RCM objective, offering new theoretical angles. Through extensive experiments, we manifest the superior generative quality of RCM in few-step generation on various non-Euclidean manifolds, including flat-tori, spheres, and the 3D rotation group SO(3), spanning a variety of crucial real-world applications such as RNA and protein generation.
Chaoran Cheng, Xiangxin Zhou, Nanning Zheng 0001
NeurIPS5
2025 VideoVLA: Video Generators Can Be Generalizable Robot Manipulators
abstract
Generalization in robot manipulation is essential for deploying robots in open-world environments and advancing toward artificial general intelligence. While recent Vision-Language-Action (VLA) models leverage large pre-trained understanding models for perception and instruction following, their ability to generalize to novel tasks, objects, and settings remains limited. In this work, we present VideoVLA, a simple approach that explores the potential of transforming large video generation models into robotic VLA manipulators. Given a language instruction and an image, VideoVLA predicts an action sequence as well as the future visual outcomes. Built on a multi-modal Diffusion Transformer, VideoVLA jointly models video, language, and action modalities, using pre-trained video generative models for joint visual and action forecasting. Our experiments show that high-quality imagined futures correlate with reliable action predictions and task success, highlighting the importance of visual imagination in manipulation. VideoVLA demonstrates strong generalization, including imitating other embodiments' skills and handling novel objects. This dual-prediction strategy—forecasting both actions and their visual consequences—explores a paradigm shift in robot learning and unlocks generalization capabilities in manipulation systems.
Yichao Shen 0001, Fangyun Wei, Zhiying Du, Yaobo Liang, Yan Lu 0001, Jiaolong Yang, Nanning Zheng 0001, Baining Guo
NeurIPS7
2025 Relationship detection for manipulation in object stacking scene with fully connected CRF
Mengyuan Ding, Chenjie Yang, Xuguang Lan, Nanning Zheng 0001
Neurocomputing5
2025 SparseSCIGaussian: Sparse snapshot compressed images 3D Gaussian splatting
Xuan Wang 0009, Nanning Zheng 0001, Caigui Jiang
Neurocomputing3
2025 Exploring inter- and intra-modal relations in compositional zero-shot learning
Hui Chen 0036, Haodong Jing, Nanning Zheng 0001
Neurocomputing5
2025 COSDA: Covariance regularized semantic data augmentation for self-supervised visual representation learning
Hui Chen 0036, Jingjing Jiang, Nanning Zheng 0001
Knowl. Based Syst.4
2025 Pinpointing visual content: Disentangled features in multimodal model for EEG representation learning and decoding
Haodong Jing, Panqi Yang, Haibo Hua, Nanning Zheng 0001
Knowl. Based Syst.5
2025 Style mixup enhanced disentanglement learning for unsupervised domain adaptation in medical image segmentation
Zhuotong Cai, Jingmin Xin, Chenyu You, Peiwen Shi, Siyuan Dong, Nicha C. Dvornek, Nanning Zheng 0001, James S. Duncan
Medical Image Anal.7
2025 StructVPR++: Distill Structural and Semantic Knowledge With Weighting Samples for Visual Place Recognition
abstract
Visual place recognition is a challenging task for autonomous driving and robotics, which is usually considered as an image retrieval problem. A commonly used two-stage strategy involves global retrieval followed by re-ranking using patch-level descriptors. Most deep learning-based methods in an end-to-end manner cannot extract global features with sufficient semantic information from RGB images. In contrast, re-ranking can utilize more explicit structural and semantic information in one-to-one matching process, but it is time-consuming. To bridge the gap between global retrieval and re-ranking and achieve a good trade-off between accuracy and efficiency, we propose StructVPR++, a framework that embeds structural and semantic knowledge into RGB global representations via segmentation-guided distillation. Our key innovation lies in decoupling label-specific features from global descriptors, enabling explicit semantic alignment between image pairs without requiring segmentation during deployment. Furthermore, we introduce a sample-wise weighted distillation strategy that prioritizes reliable training pairs while suppressing noisy ones. Experiments on four benchmarks demonstrate that StructVPR++ surpasses state-of-the-art global methods by 5-23% in Recall@1 and even outperforms many two-stage approaches, achieving real-time efficiency with a single RGB input.
Yanqing Shen, Sanping Zhou, Jingwen Fu, Ruotong Wang 0005, Shi-tao Chen, Nanning Zheng 0001
IEEE Trans. Pattern Anal. Mach. Intell.6
2025 Relation-Specific Feature Augmentation for unbiased scene graph generation
Jianji Wang 0001, Hui Chen 0036, Nanning Zheng 0001
Pattern Recognit.5
2025 PRGS: Patch-to-Region Graph Search for Visual Place Recognition
Weiliang Zuo, Liguo Liu, Yanqing Shen, Fuhua Xiang, Jingmin Xin, Nanning Zheng 0001
Pattern Recognit.7
2025 Dual Graph Attention Networks for Multi-View Visual Manipulation Relationship Detection and Robotic Grasping
abstract
Visual manipulation relationship detection facilitates robots to achieve safe, orderly, and efficient grasping tasks. However, most existing algorithms only model object-level or relational-level dependency individually, lacking sufficient global information, which is difficult to handle different types of reasoning errors, especially in complex environments with multi-object stacking and occlusion. To solve the above problems, we propose Dual Graph Attention Networks (Dual-GAT) for visual manipulation relationship detection, with an object-level graph network for capturing object-level dependencies and a relational-level graph network for capturing relational triplets-level interactions. The attention mechanism assigns different weights to different dependencies, obtains more accurate global context information for reasoning, and gets a manipulation relationship graph. In addition, we use multi-view feature fusion to improve the occluded object features, then enhance the relationship detection performance in multi-object scenes. Finally, our method is deployed on the robot to construct a multi-object grasping system, which can be well applied to stacking environments. Experimental results on the datasets VMRD and REGRAD show that our method significantly outperforms others. Note to Practitioners—The motivation of this research paper is to design efficient and accurate visual manipulation relationship reasoning methods to accomplish relevant grasping tasks in complex stacked scenes with multiple objects. The problem requires judging the positional relationship between objects in a stacked scene and determining the appropriate grasping order. In order to accurately grasp the target objects, this paper proposes a Dual Graph Attention Network for visual manipulation relationship detection, which utilizes object-level and relationship-level dependencies to obtain accurate global information for better performance. Meanwhile, multi-view feature fusion effectively improves the object occlusion problem. Our method can be combined with grasping detection to realize the robot grasping task in a real environment with multiple objects stacked in cluttered scenes. Practitioners can apply our method to real-time robot operating systems.
Mengyuan Ding, Yaorui Shi, Xuguang Lan, Nanning Zheng 0001
IEEE Trans Autom. Sci. Eng.5
2025 A Novel Dense Object Detector With Scale Balanced Sample Assignment and Refinement
abstract
Scale variation of objects remains one of the crucial challenges in object detection. Currently, conventional dense detectors with fixed receptive fields and label weights are not conducive to the detection of multi-scale objects. However, the design limitations of unbalanced label weights and fixed refinement for multi-scale objects and multi-tasks in these studies make it difficult to achieve better detection performance. In this paper, we propose a novel dense detector named Balanced FCOS which consists of two components: Balanced Label Assignment (BLA) and Flexible Shape-based Refinement (FSR). The BLA implements scale-balanced sample assignment by introducing reweighting factors consisting of localization and classification scores into the label assignment. Low-quality but high-weight samples can be weakened by the BLA. Furthermore, we design a cross-reweighting mechanism in the BLA to ensure score consistency between classification and localization. The FSR implements scale-balanced sample refinement by learning flexible sample points’ offsets for multi-scale objects and multi-tasks based on objects’ coarse features to get more discriminative features with appropriate receptive field. In addition, better features obtained by FSR are beneficial to get better classification and localization scores, which can be used by BLA to produce accurate label weights. Only equipped with the BLA, we can achieve 41.7/46.6 AP under R50/R101-FCOS without any additional parameters. When combining the BLA with the FSR, our Balanced FCOS achieves SOTA results among dense detectors on the COCO test-dev set. Experiments conducted on other heads (T-Head, DyHead), detectors (DINO), and datasets (AI-TOD) further demonstrate the effectiveness of our method.
Jinpeng Dong, Dingyi Yao, Sanping Zhou, Nanning Zheng 0001
IEEE Trans. Circuits Syst. Video Technol.5
2025 AFSIFormer: Adaptive Frequency-Spatial Interaction Attention Mechanism for Aerial Image Semantic Segmentation
Jie Hui, Wenyu Mi, Jianji Wang 0001, Yuanyang Cao, Nanning Zheng 0001
IEEE Trans. Geosci. Remote. Sens.6
2025 GRRSIS: Generalized Referring Remote Sensing Image Segmentation
abstract
Referring Remote Sensing Image Segmentation (RRSIS) is a challenging task that involves segmenting target instances within a top-view image guided by a natural language expression. Existing classic RRSIS methods commonly support target expressions only, i.e., the target described by the expression is present in the image. No-target expressions are excluded. Under this constraint, the model may face significant challenges. For instance, a small error, such as a typographical mistake, could cause a complete failure of the model. To overcome this issue, in this paper, we introduce a new benchmark called Generalized Referring Remote Sensing Image Segmentation (GRRSIS), which extends classic RRSIS by allowing expressions to refer to no-target objects. Towards this, we construct the first large-scale dataset for GRRSIS, called GRRSIS-D, which includes multi-target, single-target, and no-target expressions. Core challenges in GRRSIS stem from the fact that objects in aerial images often occupy only a small number of pixels, exhibit significant orientation variations, and present varying levels of recognition difficulty. To tackle these challenges, we propose an Oriented-aware Multi-Scale Network with an Adaptive Angle Sensing module that integrates Adaptive Rotated Convolution and a gating mechanism to capture diverse object orientations while suppressing irrelevant features for more accurate representations. Additionally, we introduce a novel Online Hard Case Mining Loss, which allocates varying levels of attention to foreground and background regions and reshapes the standard loss by down-weighting well-segmented examples, effectively addressing the issues caused by low pixel occupancy and uneven sample difficulty. The proposed approach achieves state-of-the-art performance on both the newly introduced GRRSIS and classic RRSIS tasks.
Wenyu Mi, Jianji Wang 0001, Fuzhen Zhuang, Nanning Zheng 0001
IEEE Trans. Geosci. Remote. Sens.4
2025 B-AVIBench: Toward Evaluating the Robustness of Large Vision-Language Model on Black-Box Adversarial Visual-Instructions
abstract
Large Vision-Language Models (LVLMs) have shown significant progress in responding well to visual-instructions from users. However, these instructions, encompassing images and text, are susceptible to both intentional and inadvertent attacks. Despite the critical importance of LVLMs’ robustness against such threats, current research in this area remains limited. To bridge this gap, we introduce B-AVIBench, a framework designed to analyze the robustness of LVLMs when facing various Black-box Adversarial Visual-Instructions (B-AVIs), including four types of image-based B-AVIs, ten types of text-based B-AVIs, and nine types of content bias B-AVIs (such as gender, violence, cultural, and racial biases, among others). We generate 316K B-AVIs encompassing five categories of multimodal capabilities (ten tasks) and content bias. We then conduct a comprehensive evaluation involving 14 open-source LVLMs to assess their performance. B-AVIBench also serves as a convenient tool for practitioners to evaluate the robustness of LVLMs against B-AVIs. Our findings and extensive experimental results shed light on the vulnerabilities of LVLMs, and highlight that inherent biases exist even in advanced closed-source LVLMs like GeminiProVision and GPT-4V. This underscores the importance of enhancing the robustness, security, and fairness of LVLMs. The source code and benchmark are available athttps://github.com/zhanghao5201/B-AVIBench.
Hao Zhang 0117, Wenqi Shao, Ping Luo 0002, Yu Qiao 0001, Nanning Zheng 0001, Kaipeng Zhang
IEEE Trans. Inf. Forensics Secur.7
2025 Leveraging Anchor-Based LiDAR 3D Object Detection via Point Assisted Sample Selection
abstract
3D object detection based on LiDAR point cloud and prior anchor boxes is a critical technology for autonomous driving environment perception and understanding. Nevertheless, an overlooked practical issue in existing methods is the ambiguity in training sample allocation based on box Intersection over Union (IoUbox). This problem impedes further enhancements in the performance of anchor-based LiDAR 3D object detectors. To tackle this challenge, this paper introduces a new training sample selection method that utilizes point cloud distribution for anchor sample quality measurement, named Point Assisted Sample Selection (PASS). This method has undergone rigorous evaluation on four widely utilized datasets. Experimental results demonstrate that the application of PASS elevates the average precision of anchor-based LiDAR 3D object detectors to a novel state-of-the-art, thereby proving the effectiveness of the proposed approach. The codes will be made available at https://github.com/ XJTU-Haolin/Point_Assisted_Sample_Selection.
Shi-tao Chen, Nanning Zheng 0001
IEEE Trans. Intell. Transp. Syst.3
2025 RankTuning: Cross-Image Partial Tuning Strategies for Rank Optimization in Visual Place Recognition
abstract
Aiming to estimate the location, a common strategy of Visual Place Recognition (VPR) involves utilizing global retrieval to get top-k candidates first and performing local feature matching in candidates for reranking. Although local reranking methods bring performance gains, they need a lot of computational overhead. To narrow the performance gap between global retrieval and local reranking methods with little cost, one method is to rerank candidates with global features. However, previous works only utilized the information from positive samples in candidates, ignoring the fact that negative samples can also provide useful information. To this end, we propose RankTuning, a method that aggregates all the information from candidates using global features for reranking. Specifically, we design a cross-image interaction module that allows all candidates to interact with others to enhance the discriminative power of features. Furthermore, to drive the training of this module, we propose Generalized Recall loss to handle hard samples with a better gradient strategy. Experimental results demonstrate that our method can be easily inserted into existing architectures and achieve state-of-the-art performance. Meanwhile, our method does not require additional storage overhead, and the matching latency is only 6.3% of that of the current fastest local reranking method. The code is released athttps://github.com/LKELN/RankTuning.git
Liguo Liu, Weiliang Zuo, Jingwen Fu, Yanqing Shen, Jingmin Xin, Nanning Zheng 0001
IEEE Trans. Intell. Transp. Syst.6
2025 BrainCLIP: Brain Representation via CLIP for Generic Natural Visual Stimulus Decoding
abstract
Functional Magnetic Resonance Imaging (fMRI) presents challenges due to limited paired samples and low signal-to-noise ratios, particularly in tasks involving reconstructing natural images or decoding their semantic content. To address these challenges, we introduce BrainCLIP, an innovative fMRI-based brain decoding model. BrainCLIP leverages Contrastive Language-Image Pre-training's (CLIP) cross-modal generalization abilities to bridge brain activity, images, and text for the first time. Our experiments demonstrate CLIP's effectiveness in diverse brain decoding tasks, including zero-shot visual category decoding, fMRI-image/text alignment, and fMRI-to-image generation. The core objective of BrainCLIP is to train a mapping network that translates fMRI patterns into a unified CLIP embedding space, achieved through visual and textual supervision integration. Our experiments highlight that this approach significantly enhances performance in tasks such as fMRI-text alignment and fMRI-based image generation. Notably, BrainCLIP surpasses BraVL, a recent multi-modal method, in zero-shot visual category decoding. Moreover, BrainCLIP demonstrates strong capability in reconstructing visual stimuli with high semantic fidelity, competing favorably with state-of-the-art methods in capturing high-level semantic features during fMRI-based natural image reconstruction.
Liangjun Chen, Guibo Zhu, Badong Chen, Nanning Zheng 0001
IEEE Trans. Medical Imaging6
2025 Robust Noisy Label Learning via Two-Stream Sample Distillation
abstract
Noisy label learning aims to learn robust networks under the supervision of noisy labels, which plays a critical role in deep learning. Existing work either conducts sample selection or label correction to deal with noisy labels during the model training process. In this paper, we design a simple yet effective sample selection framework, termed Two-Stream Sample Distillation (TSSD), for noisy label learning, which can extract more high-quality samples with clean labels to improve the robustness of network training. Firstly, a novel Parallel Sample Division (PSD) module is designed to generate acertaintraining set with sufficient reliable positive and negative samples by jointly considering the sample structure in feature space and the human prior in loss space. Secondly, a novel Meta Sample Purification (MSP) module is further designed to mine adequate semi-hard samples from the remaininguncertaintraining set by learning a strong meta classifier with extra golden data. As a result, more and more high-quality samples will be distilled from the noisy training set to train networks robustly in every iteration. Extensive experiments on four benchmark datasets, including CIFAR-10, CIFAR-100, Tiny-ImageNet and Clothing-1M, show that our method has achieved state-of-the-art results over its competitors.
Sihan Bai, Sanping Zhou, Le Wang 0003, Nanning Zheng 0001
IEEE Trans. Multim.5
2025 Meta Pairwise Relationship Distillation for Unsupervised Person Re-Identification
abstract
Unsupervised person re-identification (Re-ID) is challenging due to the lack of ground-truth labels. Most existing methods rely on pseudo labels estimated via iterative clustering and thus are highly susceptible to performance penalties incurred by the inaccurate estimated number of clusters. Alternatively, we utilize the sample pairs with pairwise pseudo labels to guide the feature learning to avoid the dilemma of determining cluster numbers. In this article, we propose a meta pairwise relationship distillation (MPRD) method that incorporates a graph convolutional network (GCN) to provide high-fidelity pairwise relationships to supervise the model training. A small amount of metadata with very-confidence pairwise relationships and the unlabeled pairs with the provided pseudo pairwise relationships participate in the GCN training. Besides, we introduce a hard sample deduction (HSD) module to timely mine the sample pairs with error-prone pairwise pseudo labels to mitigate the misled optimization by noisy labels. Furthermore, since the features of each positive pair represent the same person, we design a positive pair alignment (PPA) module to reduce the redundant information in each feature, which is achieved by minimizing the difference between each positive pair's feature distributions. Extensive experiments on the Market-1501, DukeMTMC-reID, and MSMT17 datasets show that our method outperforms the state-of-the-art unsupervised methods.
Haoxuanye Ji, Le Wang 0003, Sanping Zhou, Wei Tang 0016, Nanning Zheng 0001, Gang Hua 0001
IEEE Trans. Neural Networks Learn. Syst.5
2025 Improving Offline Reinforcement Learning With in-Sample Advantage Regularization for Robot Manipulation
abstract
Offline reinforcement learning (RL) aims to learn the possible policy from a fixed dataset without real-time interactions with the environment. By avoiding the risky exploration of the robot, this approach is expected to significantly improve the robot's learning efficiency and safety. However, due to errors in value estimation from out-of-distribution actions, most offline RL algorithms constrain or regularize the policy to the actions contained within the dataset. The cost of such methods is the introduction of new hyperparameters and additional complexity. In this article, we aim to adapt offline RL to robotic manipulation with minimal changes and to avoid evaluating out-of-distribution actions as much as possible. Therefore, we improve offline RL with in-sample advantage regularization (ISAR). To mitigate the impact of unseen actions, the ISAR learns the state-value function only with the dataset sample to regress the optimal action-value function. Our method calculates the advantage function of action-state pairs based on in-sample value estimation and adds a behavior cloning (BC) regularization term in the policy update. This improves sample efficiency with minimal changes, resulting in a simple and easy-to-implement method. The experiments of the D4RL robot benchmark and multigoal sparse rewards robotic tasks show that the ISAR achieves excellent performance comparable to current state-of-the-art algorithms without the need for complex parameter tuning and too much training time. In addition, we demonstrate the effectiveness of our method on a real-world robot platform.
Chengzhong Ma, Deyu Yang, Zeyang Liu 0001, Houxue Yang, Xingyu Chen 0001, Xuguang Lan, Nanning Zheng 0001
IEEE Trans. Neural Networks Learn. Syst.8
2025 Semantic Consistency Reasoning for 3-D Object Detection in Point Clouds
abstract
Point cloud-based 3-D object detection is a significant and critical issue in numerous applications. While most existing methods attempt to capitalize on the geometric characteristics of point clouds, they neglect the internal semantic properties of point and the consistency between the semantic and geometric clues. We introduce a semantic consistency (SC) mechanism for 3-D object detection in this article, by reasoning about the semantic relations between 3-D object boxes and its internal points. This mechanism is based on a natural principle: the semantic category of a 3-D bounding box should be consistent with the categories of all points within the box. Driven by the SC mechanism, we propose a novel SC network (SCNet) to detect 3-D objects from point clouds. Specifically, the SCNet is composed of a feature extraction module, a detection decision module, and a semantic segmentation module. In inference, the feature extraction and the detection decision modules are used to detect 3-D objects. In training, the semantic segmentation module is jointly trained with the other two modules to produce more robust and applicable model parameters. The performance is greatly boosted through reasoning about the relations between the output 3-D object boxes and segmented points. The proposed SC mechanism is model-agnostic and can be integrated into other base 3-D object detection models. We test the proposed model on three challenging indoor and outdoor benchmark datasets: ScanNetV2, SUN RGB-D, and KITTI. Furthermore, to validate the universality of the SC mechanism, we implement it in three different 3-D object detectors. The experiments show that the performance is impressively improved and the extensive ablation studies also demonstrate the effectiveness of the proposed model.
Wenwen Wei, Ping Wei 0001, Zhimin Liao, Jialu Qin, Xiang Cheng 0001, Meiqin Liu 0001, Nanning Zheng 0001
IEEE Trans. Neural Networks Learn. Syst.7
2024 IS-DARTS: Stabilizing DARTS through Precise Measurement on Candidate Importance
abstract
Among existing Neural Architecture Search methods, DARTS is known for its efficiency and simplicity. This approach applies continuous relaxation of network representation to construct a weight-sharing supernet and enables the identification of excellent subnets in just a few GPU days. However, performance collapse in DARTS results in deteriorating architectures filled with parameter-free operations and remains a great challenge to the robustness. To resolve this problem, we reveal that the fundamental reason is the biased estimation of the candidate importance in the search space through theoretical and experimental analysis, and more precisely select operations via information-based measurements. Furthermore, we demonstrate that the excessive concern over the supernet and inefficient utilization of data in bi-level optimization also account for suboptimal results. We adopt a more realistic objective focusing on the performance of subnets and simplify it with the help of the informationbased measurements. Finally, we explain theoretically why progressively shrinking the width of the supernet is necessary and reduce the approximation error of optimal weights in DARTS. Our proposed method, named IS-DARTS, comprehensively improves DARTS and resolves the aforementioned problems. Extensive experiments on NAS-Bench-201 and DARTS-based search space demonstrate the effectiveness of IS-DARTS.
Hongyi He, Longjun Liu, Haonan Zhang 0002, Nanning Zheng 0001
AAAI4
2024 Voxel or Pillar: Exploring Efficient Point Cloud Representation for 3D Object Detection
abstract
Efficient representation of point clouds is fundamental for LiDAR-based 3D object detection. While recent grid-based detectors often encode point clouds into either voxels or pillars, the distinctions between these approaches remain underexplored. In this paper, we quantify the differences between the current encoding paradigms and highlight the limited vertical learning within. To tackle these limitations, we propose a hybrid detection framework named Voxel-Pillar Fusion (VPF), which synergistically combines the unique strengths of both voxels and pillars. To be concrete, we first develop a sparse voxel-pillar encoder that encodes point clouds into voxel and pillar features through 3D and 2D sparse convolutions respectively, and then introduce the Sparse Fusion Layer (SFL), facilitating bidirectional interaction between sparse voxel and pillar features. Our computationally efficient, fully sparse method can be seamlessly integrated into both dense and sparse detectors. Leveraging this powerful yet straightforward representation, VPF delivers competitive performance, achieving real-time inference speeds on the nuScenes and Waymo Open Dataset.
Sanping Zhou, Jinpeng Dong, Nanning Zheng 0001
AAAI5
2024 GSO-Net: Grid Surface Optimization via Learning Geometric Constraints
abstract
In the context of surface representations, we find a natural structural similarity between grid surface and image data. Motivated by this inspiration, we propose a novel approach: encoding grid surfaces as geometric images and using image processing methods to address surface optimization-related problems. As a result, we have created the first dataset for grid surface optimization and devised a learning-based grid surface optimization network specifically tailored to geometric images, addressing the surface optimization problem through a data-driven learning of geometric constraints paradigm. We conduct extensive experiments on developable surface optimization, surface flattening, and surface denoising tasks using the designed network and datasets. The results demonstrate that our proposed method not only addresses the surface optimization problem better than traditional numerical optimization methods, especially for complex surfaces, but also boosts the optimization speed by multiple orders of magnitude. This pioneering study successfully applies deep learning methods to the field of surface optimization and provides a new solution paradigm for similar tasks, which will provide inspiration and guidance for future developments in the field of discrete surface optimization. The code and dataset are available at https://github.com/chaoyunwang/GSO-Net.
Chaoyun Wang, Jingmin Xin, Nanning Zheng 0001, Caigui Jiang
AAAI3
2024 Text Grouping Adapter: Adapting Pre-Trained Text Detector for Layout Analysis
abstract
Significant progress has been made in scene text detection models since the rise of deep learning, but scene text layout analysis, which aims to group detected text instances as paragraphs, has not kept pace. Previous works either treated text detection and grouping using separate models, or train a model from scratch while using a unified one. All of them have not yet made full use of the already well-trained text detectors and easily obtainable detection datasets. In this paper, we present Text Grouping Adapter (TGA), a module that can enable the utilization of various pretrained text detectors to learn layout analysis, allowing us to adopt a well-trained text detector right off the shelf or just fine-tune it efficiently. Designed to be compatible with various text detector architectures, TGA takes detected text regions and image features as universal inputs to as-semble text instance features. To capture broader contextual information for layout analysis, we propose to predict text group masks from text instance features by one-to-many assignment. Our comprehensive experiments demonstrate that, even with frozen pretrained models, incorporating our TGA into various pretrained text detectors and text spotters can achieve superior layout analysis performance, simultaneously inheriting generalized text detection ability from pretraining. In the case of full parameter fine-tuning, we can further improve layout analysis performance.
Tianci Bi, Zhizheng Zhang 0004, Wenxuan Xie, Cuiling Lan, Yan Lu 0001, Nanning Zheng 0001
CVPR7
2024 Hidden States in LLMs Improve EEG Representation Learning and Visual Decoding
abstract
Analyzing brain signals and reconstructing visual stimuli from the brain can facilitate further exploration on cognitive functions of the human brain, which have attracted strong interest in neuroscience and artificial intelligence. However, due to defects such as complex noises and the lack of alignment accuracy, efficient methods for extracting information from Electroencephalogram (EEG) signals are still very limited, making it difficult to perform EEG visual decoding tasks. Our study shows a way to handle the issues by proposing a new method for EEG representation learning and visual decoding, thus completing end-to-end image reconstruction tasks from EEG signals. We utilize the ability of semantic extraction and prediction of large language models (LLMs) to enhance the performance of EEG feature extraction. For semantic representation learning, we align EEG signals with target semantic embeddings, which are obtained from hidden states of Large Language Model Meta AI 2 (LLaMa-2) by inputting descriptions of images into the model. We also extract visual features from EEG signals to improve the quantity of the reconstructed images at low levels. Then we fuse semantic features and visual features by applying a pre-trained diffusion model and finally generate the corresponding images. We are the first to incorporate the LLM into EEG visual decoding tasks. Our method achieves the state-of-the-art result of EEG classification accuracy and the quality of reconstructed images on ImageNet-EEG datasets. In one word, our work is an important step forward in the field of exploiting the relationship between language models and human visual cognition. Our codes are available at https://github.com/lay-atsa/llm4eeg.
Aoyang Liu, Haodong Jing, Nanning Zheng 0001
ECAI5
2024 AugDETR: Improving Multi-scale Learning for Detection Transformer
Jinpeng Dong, Yutong Lin, Sanping Zhou, Nanning Zheng 0001
ECCV (24)5
2024 PMT: Progressive Mean Teacher via Exploring Temporal Consistency for Semi-Supervised Medical Image Segmentation
Sanping Zhou, Le Wang 0003, Nanning Zheng 0001
ECCV (68)4
2024 MRSP: Learn Multi-representations of Single Primitive for Compositional Zero-Shot Learning
Dongyao Jiang, Hui Chen 0036, Haodong Jing, Nanning Zheng 0001
ECCV (65)5
2024 AMA: An Analytical Approach to Maximizing the Efficiency of Deep Learning on Versal AI Engine
abstract
The traditional cache-based multi-core architecture represented by CUDA has been plagued by the “memory wall” problem, and a large number of applications represented by large language model inference are unable to meet the high computational intensity requirements. The Versal AI Engine architecture provides a variety of rich inter-core connections in the processor array, increasing many data reuse opportunities and potentially alleviating the “memory wall” issue. However, traditional parallel programming models cannot be directly applied to this architecture, and how to map computations to achieve high computational utilization becomes a new challenge. To address this, we propose AMA, a hierarchical performance analysis model built on the Versal AI Engine architecture, designed to maximize the efficiency of typical deep learning applications. Experiments show that AMA modeling is accurate and efficient. On the VCK190 platform, we achieved a matrix multiplication throughput of 5867.29 GFLOPS in fp32 and 88.55 TOPS in int8, and a convolution throughput of 99.6770 TOPS. In terms of energy efficiency, AMA achieved 142.68 GFLOPS/W in fp32 precision and 1.416 TOPS/W in int8 matrix multiplication. Compared to the current state-of-the-art methods, we achieved a $\mathbf{1 4. 9 9 \%}$ increase in throughput and a 22.92% increase in energy efficiency, providing new analytical performance model and practical guidance for efficient deep learning deployment on AI Engine.
Xiaodong Deng, Longjun Liu, Nanning Zheng 0001
FPL6
2024 CuES: Conditional Uncorrelation-based Characteristic Enhancement and Fusion of Electrical Signals
abstract
The lifespan of a generator greatly depends on the quality and aging of its stator bar insulation material. Aging of insulation materials can lead to premature equipment failure and significant material loss, resulting in substantial economic losses. However, existing methods for predicting the lifespan of electronic wire bars have several drawbacks, such as slow training speed, the need for a large amount of training data, and a tendency to overfit. To address this issue, we propose a characteristic enhancement algorithm based on conditional uncorrelation. This algorithm leverages characteristic enhancement to generate an extensive dataset and utilizes subset selection to identify relevant electrical parameters for predicting the remaining life span of the stator bar’s main insulation configurations. Experimental results demonstrate the advantages of our research compared to deep learning models. Our approach offers a promising solution for accurately predicting the remaining life of stator bar insulation, thereby facilitating effective maintenance planning and minimizing economic losses.
Haohao Cai, Xichun Liu, Jianji Wang 0001, Nanning Zheng 0001
FUSION7
2024 Symmetric Consistency with Cross-Domain Mixup for Cross-Modality Cardiac Segmentation
abstract
Accurate cardiac segmentation in cross-modality images plays an important role in the quantitative analysis of the heart to diagnose cardiovascular diseases. However, achieving high performance in cross-modality segmentation is hindered by the time-consuming annotation and modality gap. While some approaches employ Unsupervised Domain Adaptation (UDA) through adversarial learning to address the issue, it still remains challenging due to the instability of the adversarial generative models. In this work, we propose Symmetric Consistency with Cross-Domain Mixup (SCCDM), integrated with the teacher-student model for cross-modality cardiac segmentation. Specifically, we introduce symmetric consistency across the domains for two mixed data to diversify the data distribution from both the source domain and target domain. Extensive experiments on a public cardiac dataset demonstrate that SCCDM achieves superior domain adaptation performance for cardiac segmentation compared to state-of-the-art methods.
Zhuotong Cai, Jingmin Xin, Siyuan Dong, John A. Onofrey, Nanning Zheng 0001, James S. Duncan
ICASSP5
2024 V-DETR: DETR with Vertex Relative Position Encoding for 3D Object Detection
abstract
We introduce a highly performant 3D object detector for point clouds using the DETR framework. The prior attempts all end up with suboptimal results because they fail to learn accurate inductive biases from the limited scale of training data. In particular, the queries often attend to points that are far away from the target objects, violating the locality principle in object detection. To address the limitation, we introduce a novel 3D Vertex Relative Position Encoding (3DV-RPE) method which computes position encoding for each point based on its relative position to the 3D boxes predicted by the queries in each decoder layer, thus providing clear information to guide the model to focus on points near the objects, in accordance with the principle of locality. Furthermore, we have systematically refined our pipeline, including data normalization, to better align with the task requirements. Our approach demonstrates remarkable performance on the demanding ScanNetV2 benchmark, showcasing substantial enhancements over the prior state-of-the-art CAGroup3D. Specifically, we achieve an increase in $AP_{25}$ from $75.1\%$ to $77.8\%$ and in ${AP}_{50}$ from $61.3\%$ to $66.0\%$.
Yichao Shen 0001, Zigang Geng, Yuhui Yuan, Yutong Lin, Chunyu Wang 0001, Han Hu 0001, Nanning Zheng 0001, Baining Guo
ICLR8
2024 Breaking through the learning plateaus of in-context learning in Transformer
abstract
In-context learning, i.e., learning from context examples, is an impressive ability of Transformer. Training Transformers to possess this in-context learning skill is computationally intensive due to the occurrence of *learning plateaus*, which are periods within the training process where there is minimal or no enhancement in the model's in-context learning capability. To study the mechanism behind the learning plateaus, we conceptually separate a component within the model's internal representation that is exclusively affected by the model's weights. We call this the “weights component”, and the remainder is identified as the “context component”. By conducting meticulous and controlled experiments on synthetic tasks, we note that the persistence of learning plateaus correlates with compromised functionality of the weights component. Recognizing the impaired performance of the weights component as a fundamental behavior that drives learning plateaus, we have developed three strategies to expedite the learning of Transformers. The effectiveness of these strategies is further confirmed in natural language processing tasks. In conclusion, our research demonstrates the feasibility of cultivating a powerful in-context learning ability within AI systems in an eco-friendly manner.
Jingwen Fu, Tao Yang 0032, Yuwang Wang, Yan Lu 0001, Nanning Zheng 0001
ICML5
2024 Self-Consistency Training for Density-Functional-Theory Hamiltonian Prediction
abstract
Predicting the mean-field Hamiltonian matrix in density functional theory is a fundamental formulation to leverage machine learning for solving molecular science problems. Yet, its applicability is limited by insufficient labeled data for training. In this work, we highlight that Hamiltonian prediction possesses a self-consistency principle, based on which we propose self-consistency training, an exact training method that does not require labeled data. It distinguishes the task from predicting other molecular properties by the following benefits: (1) it enables the model to be trained on a large amount of unlabeled data, hence addresses the data scarcity challenge and enhances generalization; (2) it is more efficient than running DFT to generate labels for supervised training, since it amortizes DFT calculation over a set of queries. We empirically demonstrate the better generalization in data-scarce and out-of-distribution scenarios, and the better efficiency over DFT labeling. These benefits push forward the applicability of Hamiltonian prediction to an ever-larger scale.
Chang Liu 0030, Zun Wang 0006, Xinran Wei, Siyuan Liu 0005, Nanning Zheng 0001, Bin Shao 0002, Tie-Yan Liu
ICML6
2024 Learning Segmented 3D Gaussians via Efficient Feature Unprojection for Zero-Shot Neural Scene Segmentation
Bin Dou, Yongjia Ma, Zejian Yuan, Nanning Zheng 0001
ICONIP (8)6
2024 Complementing Onboard Sensors with Satellite Maps: A New Perspective for HD Map Construction
abstract
High-definition (HD) maps play a crucial role in autonomous driving systems. Recent methods have attempted to construct HD maps in real-time using vehicle onboard sensors. Due to the inherent limitations of onboard sensors, which include sensitivity to detection range and susceptibility to occlusion by nearby vehicles, the performance of these methods significantly declines in complex scenarios and long-range detection tasks. In this paper, we explore a new perspective that boosts HD map construction through the use of satellite maps to complement onboard sensors. We initially generate the satellite map tiles for each sample in nuScenes and release a complementary dataset for further research. To enable better integration of satellite maps with existing methods, we propose a hierarchical fusion module, which includes feature-level fusion and BEV-level fusion. The feature-level fusion, composed of a mask generator and a masked cross-attention mechanism, is used to refine the features from onboard sensors. The BEV-level fusion mitigates the coordinate differences between features obtained from onboard sensors and satellite maps through an alignment module. The experimental results on the augmented nuScenes showcase the seamless integration of our module into three existing HD map construction methods. The satellite maps and our proposed module notably enhance their performance in both HD map semantic segmentation and instance detection tasks. Our code will be available at https://github.com/xjtu-csgao/SatforHDMap.
Wenjie Gao 0001, Jiawei Fu 0001, Yanqing Shen, Haodong Jing, Shi-tao Chen, Nanning Zheng 0001
ICRA6
2024 POAQL: A Partially Observable Altruistic Q-Learning Method for Cooperative Multi-Agent Reinforcement Learning
abstract
Multi-Agent Path Finding (MAPF) is an important issue in multi-agent cooperation. Many studies apply MultiAgent Reinforcement Learning (MARL) to solve MAPF in partially observable settings. The objective of cooperative MARL is to maximize the cumulative team reward. Nevertheless, in partially observable settings, the team reward is misleading due to unpredictable factors from the behavior and state of unobserved agents. To address this issue, we propose a Partially Observable Altruistic Q-learning (POAQL) method. POAQL considers the cumulative reward of the observed subteam instead of the whole team, where Altruistic Q-learning plays an important role in learning the subteam action value. In addition, we design a new conflict resolution without additional guidance to emphasize the cooperative nature of MARL frameworks. Experimental results show that POAQL outperforms existing reinforcement learning methods in terms of efficiency and performance.
Lesong Tao, Miao Kang, Jinpeng Dong, Songyi Zhang, Ke Ye, Shi-tao Chen, Nanning Zheng 0001
ICRA7
2024 Vehicle Trajectory Prediction with Soft Behavior Constraints
abstract
Trajectory prediction plays a crucial role in autonomous driving, but it is challenging due to the multi-modal nature of future trajectories. Behavior information is frequently employed to capture more diverse modalities of future trajectories. Traditional behavior information is typically hard-encoded, which is often inaccurate and inadequate for reflecting future multimodality. Therefore, we introduce the concept of soft vehicle behavior, which is represented as a probability distribution over a predefined comprehensive set of behaviors. This approach allows for a more rational depiction of vehicle behavior and captures potential future driving modalities. Based on it, we propose a new soft-behavior-constrained vehicle trajectory prediction framework. The framework consists of a backbone and a lightweight and plug-and-play behavior prediction module, which is used to imbue soft behavior constraints to assist in representation learning. We integrated the behavior prediction module into five representative trajectory predictors and achieved improvements of at least 4.2% in minFDE(K=5) on the nuScenes dataset and 0.5% in minFDE(K=6) on the Argoverse 1 motion forecasting dataset. These universal increments prove the effectiveness and generalizability of soft behavior constraints in vehicle trajectory prediction.
Ke Ye, Sanping Zhou, Miao Kang, Jingwen Fu, Nanning Zheng 0001
IROS5
2024 Fast Multi-Class Vehicle Cooperative Path Optimization in Complex Urban V2X Transportation: A Novel Parallel Multi-Agent Reinforcement Learning Approach
abstract
Urban road traffic systems are advancing into sophisticated networks, underscoring the importance of real-time collaborative decision-making. This study tackles the intricate challenge of cooperative path planning under complex urban conditions, taking into account a variety of vehicle types and their respective priorities. While conventional path planning techniques struggle with such intricate coordination, reinforcement learning, though theoretically capable, is hindered by its limited model reusability and protracted training times. To address these issues, we present a novel parallel multi-agent reinforcement learning strategy for path planning that is adaptable to various vehicle types. The problem is initially cast as a multi-agent Markov Decision Process (MDP), followed by the introduction of a parallel training approach within the Q-learning framework. This approach leverages tensor computation to transform the Q-table, state, and reward, thereby markedly accelerating the training process. Empirical simulations demonstrate the approach’s efficacy, achieving a 0.84% reduction in training time (from approximately 771.611 seconds to 0.654 seconds), achieving a 93.94% lower probability of path overlap though the total distance increased by 7.69%.
Shi-tao Chen, Shuyang Cai, Ziheng Tang, Donghe Li, Nanning Zheng 0001
IV5
2024 The Optimal Horizon Model Predictive Control Planning for Autonomous Vehicles in Dynamic Environments
abstract
A primary challenge in autonomous driving is achieving safe and efficient trajectory planning in complex dynamic environments. This task requires adherence to traffic laws and vehicle dynamics models as well as an understanding of the spatial distributions and behavior of various traffic participants in densely populated areas. Model Predictive Control (MPC) and its variants typically employ a fixed prediction horizon, which results in limited adaptability in dynamic environments. A long prediction horizon escalates computational costs, while a short prediction horizon may impact real-time performance adversely. To tackle this challenge, our study introduces an optimal horizon MPC planning approach. This method incorporates a sliding horizon window founded on reinforcement learning and interactive MPC planning, making it versatile for a variety of driving scenarios. Additionally, our approach implicitly models the spatio-temporal interactions among traffic participants, thereby enriching the information pool for effective planning. Rigorous tests and validations conducted using the real-world dataset nuPlan affirm that our proposed method delivers robust planning performance, facilitating safe and efficient trajectory planning for autonomous vehicles.
Shi-tao Chen, Jiamin Shi, Nanning Zheng 0001
IV4
2024 Human-Like Reverse Parking using Deep Reinforcement Learning with Attention Mechanism
abstract
This study explores efficient and safe Automated Valet Parking (AVP) strategies in unstructured and dynamic environments. Existing approaches utilizing reinforcement learning neglected the interaction between dynamic agents and ego vehicle, and disregarded human driving patterns, leading to their ineffectiveness in unstructured dynamic environments. We propose a novel hybrid attention mechanism that comprehends the mixed interactions between static and dynamic elements, aiding autonomous vehicles in advanced planning. We implemented a guidance system based on human preferences, eliminating the need for expert data and expediting the training process via intermediate planning stages, thereby facilitating parking maneuvers akin to human drivers. The model was trained and validated in a range of parking situations. The experimental outcomes indicate that our method possesses robust adaptability and navigation skills in static and dynamic environments.
Zhuo Qiu, Shi-tao Chen, Jiamin Shi, Nanning Zheng 0001
IV5
2024 Class-Aware Mutual Mixup with Triple Alignments for Semi-supervised Cross-Domain Segmentation
Zhuotong Cai, Jingmin Xin, Tianyi Zeng, Siyuan Dong, Nanning Zheng 0001, James S. Duncan
MICCAI (8)5
2024 Refracting Once is Enough: Neural Radiance Fields for Novel-View Synthesis of Real Refractive Objects
abstract
Neural Radiance Fields (NeRF) have shown promise in novel view synthesis, but it still face challenges when applied to refractive objects. The presence of refraction disrupts multiview consistency, often resulting in renderings that are either blurred or distorted. Recent methods alleviate this challenge by introducing external supervision, such as mask images and Index of Refraction. However,acquiring such information is often impractical,limiting the application of NeRF-like models to complex scenes with refracting elementsand yielding unsatisfactory results. To address these limitations, we introduce RoseNeRF (Refracting once is enough for NeRF), a novel method that simplifies the complex interaction of rays within objects to a single refraction event. We design the refraction network that efficiently maps a ray in the 4D light field to its refracted counterpart, better modeling curved ray paths. Furthermore, we introduce a regularization strategy to ensure the reversibility of optical paths, which is anchored in physical world theorems. To help it easier for the network to learn the highly view-dependent appearance of refractive objects, we also propose novel density decoding strategies. Our method is designed for seamless integration into most NeRF-like frameworks and has demonstrated state-of-the-art performance without any additional information on both the Eikonal Fields' dataset and Shiny dataset.
Xiaoqian Liang, Jianji Wang 0001, Yuanliang Lu, Xubin Duan, Xichun Liu, Nanning Zheng 0001
ICMR6
2024 Make Your LLM Fully Utilize the Context
abstract
While many contemporary large language models (LLMs) can process lengthy input, they still struggle to fully utilize information within the long context, known as the *lost-in-the-middle* challenge. We hypothesize that it stems from insufficient explicit supervision during the long-context training, which fails to emphasize that any position in a long context can hold crucial information. Based on this intuition, our study presents **information-intensive (IN2) training**, a purely data-driven solution to overcome lost-in-the-middle. Specifically, IN2 training leverages a synthesized long-context question-answer dataset, where the answer requires (1) **fine-grained information awareness** on a short segment (~128 tokens) within a synthesized long context (4K-32K tokens), and (2) the **integration and reasoning** of information from two or more short segments. Through applying this information-intensive training on Mistral-7B, we present **FILM-7B** (FIll-in-the-Middle). To thoroughly assess the ability of FILM-7B for utilizing long contexts, we design three probing tasks that encompass various context styles (document, code, and structured-data context) and information retrieval patterns (forward, backward, and bi-directional retrieval). The probing results demonstrate that FILM-7B can robustly retrieve information from different positions in its 32K context window. Beyond these probing tasks, FILM-7B significantly improves the performance on real-world long-context tasks (e.g., 23.5->26.9 F1 score on NarrativeQA), while maintaining a comparable performance on short-context tasks (e.g., 59.3->59.2 accuracy on MMLU).
Shengnan An, Zexiong Ma, Zeqi Lin, Nanning Zheng 0001, Jian-Guang Lou, Weizhu Chen
NeurIPS4
2024 TPR: Topology-Preserving Reservoirs for Generalized Zero-Shot Learning
abstract
Pre-trained vision-language models (VLMs) such as CLIP have shown excellent performance for zero-shot classification. Based on CLIP, recent methods design various learnable prompts to evaluate the zero-shot generalization capability on a base-to-novel setting. This setting assumes test samples are already divided into either base or novel classes, limiting its application to realistic scenarios. In this paper, we focus on a more challenging and practical setting: generalized zero-shot learning (GZSL), i.e., testing with no information about the base/novel division. To address this challenging zero-shot problem, we introduce two unique designs that enable us to classify an image without the need of knowing whether it comes from seen or unseen classes. Firstly, most existing methods only adopt a single latent space to align visual and linguistic features, which has a limited ability to represent complex visual-linguistic patterns, especially for fine-grained tasks. Instead, we propose a dual-space feature alignment module that effectively augments the latent space with a novel attribute space induced by a well-devised attribute reservoir. In particular, the attribute reservoir consists of a static vocabulary and learnable tokens complementing each other for flexible control over feature granularity. Secondly, finetuning CLIP models (e.g., prompt learning) on seen base classes usually sacrifices the model's original generalization capability on unseen novel classes. To mitigate this issue, we present a new topology-preserving objective that can enforce feature topology structures of the combined base and novel classes to resemble the topology of CLIP. In this manner, our model will inherit the generalization ability of CLIP through maintaining the pairwise class angles in the attribute space. Extensive experiments on twelve object recognition datasets demonstrate that our model, termed Topology-Preserving Reservoir (TPR), outperforms strong baselines including both prompt learning and conventional generative-based zero-shot methods.
Hui Chen 0036, Yanbin Liu 0003, Nanning Zheng 0001, Xin Yu 0002
NeurIPS4
2024 Molecule Design by Latent Prompt Transformer
abstract
This work explores the challenging problem of molecule design by framing it as a conditional generative modeling task, where target biological properties or desired chemical constraints serve as conditioning variables. We propose the Latent Prompt Transformer (LPT), a novel generative model comprising three components: (1) a latent vector with a learnable prior distribution modeled by a neural transformation of Gaussian white noise; (2) a molecule generation model based on a causal Transformer, which uses the latent vector as a prompt; and (3) a property prediction model that predicts a molecule's target properties and/or constraint values using the latent prompt. LPT can be learned by maximum likelihood estimation on molecule-property pairs. During property optimization, the latent prompt is inferred from target properties and constraints through posterior sampling and then used to guide the autoregressive molecule generation. After initial training on existing molecules and their properties, we adopt an online learning algorithm to progressively shift the model distribution towards regions that support desired target properties. Experiments demonstrate that LPT not only effectively discovers useful molecules across single-objective, multi-objective, and structure-constrained optimization tasks, but also exhibits strong sample efficiency.
Deqian Kong, Jianwen Xie, Edouardo Honig, Shuanghong Xue, Pei Lin, Sanping Zhou, Nanning Zheng 0001, Ying Nian Wu
NeurIPS10
2024 Neural P3M: A Long-Range Interaction Modeling Enhancer for Geometric GNNs
abstract
Geometric graph neural networks (GNNs) have emerged as powerful tools for modeling molecular geometry. However, they encounter limitations in effectively capturing long-range interactions in large molecular systems. To address this challenge, we introduce **Neural P$^3$M**, a versatile enhancer of geometric GNNs to expand the scope of their capabilities by incorporating mesh points alongside atoms and reimaging traditional mathematical operations in a trainable manner. Neural P$^3$M exhibits flexibility across a wide range of molecular systems and demonstrates remarkable accuracy in predicting energies and forces, outperforming on benchmarks such as the MD22 dataset. It also achieves an average improvement of 22% on the OE62 dataset while integrating with various architectures. Codes are available at https://github.com/OnlyLoveKFC/Neural_P3M.
Chaoran Cheng, Shaoning Li, Yuxuan Ren, Bin Shao 0002, Pheng-Ann Heng, Nanning Zheng 0001
NeurIPS8
2024 Diffusion Model with Cross Attention as an Inductive Bias for Disentanglement
abstract
Disentangled representation learning strives to extract the intrinsic factors within the observed data. Factoring these representations in an unsupervised manner is notably challenging and usually requires tailored loss functions or specific structural designs. In this paper, we introduce a new perspective and framework, demonstrating that diffusion models with cross-attention itself can serve as a powerful inductive bias to facilitate the learning of disentangled representations. We propose to encode an image into a set of concept tokens and treat them as the condition of the latent diffusion model for image reconstruction, where cross attention over the concept tokens is used to bridge the encoder and the U-Net of the diffusion model. We analyze that the diffusion process inherently possesses the time-varying information bottlenecks. Such information bottlenecks and cross attention act as strong inductive biases for promoting disentanglement. Without any regularization term in the loss function, this framework achieves superior disentanglement performance on the benchmark datasets, surpassing all previous methods with intricate designs. We have conducted comprehensive ablation studies and visualization analyses, shedding a light on the functioning of this model. We anticipate that our findings will inspire more investigation on exploring diffusion model for disentangled representation learning towards more sophisticated data analysis and understanding.
Tao Yang 0032, Cuiling Lan, Yan Lu 0001, Nanning Zheng 0001
NeurIPS4
2024 Correlation Information Bottleneck: Towards Adapting Pretrained Multimodal Models for Robust Visual Question Answering
Jingjing Jiang, Ziyi Liu 0001, Nanning Zheng 0001
Int. J. Comput. Vis.3
2024 Open-Vocabulary Animal Keypoint Detection with Semantic-Feature Matching
Hao Zhang 0117, Lumin Xu, Shenqi Lai, Wenqi Shao, Nanning Zheng 0001, Ping Luo 0002, Yu Qiao 0001, Kaipeng Zhang
Int. J. Comput. Vis.5
2024 Hierarchical Bayesian Causality Network to Extract High-Level Semantic Information in Visual Cortex
abstract
Functional MRI (fMRI) is a brain signal with high spatial resolution, and visual cognitive processes and semantic information in the brain can be represented and obtained through fMRI. In this paper, we design single-graphic and matched/unmatched double-graphic visual stimulus experiments and collect 12 subjects' fMRI data to explore the brain's visual perception processes. In the double-graphic stimulus experiment, we focus on the high-level semantic information as "matching", and remove tail-to-tail conjunction by designing a model to screen the matching-related voxels. Then, we perform Bayesian causal learning between fMRI voxels based on the transfer entropy, establish a hierarchical Bayesian causal network (HBcausalNet) of the visual cortex, and use the model for visual stimulus image reconstruction. HBcausalNet achieves an average accuracy of 70.57% and 53.70% in single- and double-graphic stimulus image reconstruction tasks, respectively, higher than HcorrNet and HcasaulNet. The results show that the matching-related voxel screening and causality analysis method in this paper can extract the "matching" information in fMRI, obtain a direct causal relationship between matching information and fMRI, and explore the causal inference process in the brain. It suggests that our model can effectively extract high-level semantic information in brain signals and model effective connections and visual perception processes in the visual cortex of the brain.
Ming Du 0001, Haodong Jing, Nanning Zheng 0001
Int. J. Neural Syst.5
2024 Understanding mobile GUI: From pixel-words to screen-sentences
Jingwen Fu, Yuwang Wang, Wenjun Zeng 0001, Nanning Zheng 0001
Neurocomputing5
2024 KPDet: Keypoint-based 3D object detection with Parametric Radius Learning
Sanping Zhou, Xinrui Yan, Nanning Zheng 0001
Neurocomputing4
2024 Residual feature learning with hierarchical calibration for gaze estimation
Zhengdan Yin, Sanping Zhou, Le Wang 0003, Gang Hua 0001, Nanning Zheng 0001
Mach. Vis. Appl.6
2024 G2-MonoDepth: A General Framework of Generalized Depth Inference From Monocular RGB+X Data
abstract
Monocular depth inference is a fundamental problem for scene perception of robots. Specific robots may be equipped with a camera plus an optional depth sensor of any type and located in various scenes of different scales, whereas recent advances derived multiple individual sub-tasks. It leads to additional burdens to fine-tune models for specific robots and thereby high-cost customization in large-scale industrialization. This article investigates a unified task of monocular depth inference, which infers high-quality depth maps from all kinds of input raw data from various robots in unseen scenes. A basic benchmark G2-MonoDepth is developed for this task, which comprises four components: (a) a unified data representation RGB+X to accommodate RGB plus raw depth with diverse scene scale/semantics, depth sparsity ([0%, 100%]) and errors (holes/noises/blurs), (b) a novel unified loss to adapt to diverse depth sparsity/errors of input raw data and diverse scales of output scenes, (c) an improved network to well propagate diverse scene scales from input to output, and (d) a data augmentation pipeline to simulate all types of real artifacts in raw depth maps for training. G2-MonoDepth is applied in three sub-tasks including depth estimation, depth completion with different sparsity, and depth enhancement in unseen scenes, and it always outperforms SOTA baselines on both real-world data and synthetic data.
Haotian Wang 0009, Meng Yang 0002, Nanning Zheng 0001
IEEE Trans. Pattern Anal. Mach. Intell.3
2024 Adversarial Attack and Defense in Deep Ranking
abstract
Deep Neural Network classifiers are vulnerable to adversarial attacks, where an imperceptible perturbation could result in misclassification. However, the vulnerability of DNN-based image ranking systems remains under-explored. In this paper, we propose two attacks against deep ranking systems, i.e., Candidate Attack and Query Attack, that can raise or lower the rank of chosen candidates by adversarial perturbations. Specifically, the expected ranking order is first represented as a set of inequalities. Then a triplet-like objective function is designed to obtain the optimal perturbation. Conversely, an anti-collapse triplet defense is proposed to improve the ranking model robustness against all proposed attacks, where the model learns to prevent the adversarial attack from pulling the positive and negative samples close to each other. To comprehensively measure the empirical adversarial robustness of a ranking model with our defense, we propose an empirical robustness score, which involves a set of representative attacks against ranking models. Our adversarial ranking attacks and defenses are evaluated on MNIST, Fashion-MNIST, CUB200-2011, CARS196, and Stanford Online Products datasets. Experimental results demonstrate that our attacks can effectively compromise a typical deep ranking system. Nevertheless, our defense can significantly improve the ranking system's robustness and simultaneously mitigate a wide range of attacks.
Le Wang 0003, Zhenxing Niu, Qilin Zhang 0004, Nanning Zheng 0001, Gang Hua 0001
IEEE Trans. Pattern Anal. Mach. Intell.5
2024 Transfer easy to hard: Adversarial contrastive feature learning for unsupervised person re-identification
Haoxuanye Ji, Le Wang 0003, Sanping Zhou, Wei Tang 0016, Nanning Zheng 0001, Gang Hua 0001
Pattern Recognit.5
2024 TransOSV: Offline Signature Verification with Transformers
Ping Wei 0001, Zeyu Ma 0005, Changkai Li, Nanning Zheng 0001
Pattern Recognit.5
2024 Bidirectional feature learning network for RGB-D salient object detection
abstract
RGB-D salient object detection aims to perform the pixel-wise localization of salient objects from both RGB and depth images, whose challenge mainly comes from how to learn complementary features from each modality. Existing works often use increasingly large models for performance enhancement, which need large memory and time consumption in practice. In this paper, we propose a simple yet effective B idirectional F eature L earning Net work (BFLNet) for RGB-D salient object detection under limited memory and time conditions. To achieve accurate performance with lightweight backbone networks , an effective B idirectional F eature F usion (BFF) module is designed to merge features from both RGB and depth streams, in which the cross-modal fusions and cross-scale fusions are jointly conducted to fuse the immediate features in multiple scales and multiple modals. What is more, a simple D ual C onsistency L oss (DCL) function is designed to prompt cross-modal fusion by keeping the consistency between cross-modal target predictions. Extensive experiments on four benchmark datasets demonstrate that our method has achieved the state-of-the-art performance with high efficiency in RGB-D salient object detection. Code will be available at https://github.com/nightsky-nostar/BFLNet .
Ye Niu, Sanping Zhou, Yonghao Dong, Le Wang 0003, Jinjun Wang, Nanning Zheng 0001
Pattern Recognit.6
2024 FMGNet: An efficient feature-multiplex group network for real-time vision task
Hao Zhang 0117, Kaipeng Zhang, Nanning Zheng 0001, Shenqi Lai
Pattern Recognit.4
2024 DQ-STP: An Efficient Sparse On-Device Training Processor Based on Low-Rank Decomposition and Quantization for DNN
abstract
Due to the bottleneck problems such as scenario-varying application, significant data communication overhead and privacy protection between off-line training and on-line inference, intelligent edge devices capable of adaptively fine-tuning the deep neural network (DNN) models for specific tasks have become the most urgent need. However, the computational cost is intolerable for ordinary on-device training (ODT), which inspires us to explore an efficient ODT processor, named DQ-STP. In this paper, we leverage a series of optimization techniques using software-hardware co-design. On the one hand, the proposed design incorporates SVD-based low-rank decomposition,$2^{n}$quantization and ACBN algorithm on the software side. This unifies the sparse computing mode of convolutional layers and enhancing weight sparsity. On the other hand, the proposed design effectively leverages data sparsity on the hardware side through four techniques: 1) The flag compressed sparse row is proposed to compress input feature maps and gradient maps. 2) A unified processing element (PE) array comprising shifters and adders is proposed to expedite forward and error propagation steps. 3) The PE arrays for error propagation and weight gradients generation are separated to enhance throughput. 4) A sparse alignment strategy is proposed to further enhance PE utilization. Through these software and hardware co-optimization, the proposed DQ-STP achieves an area efficiency and peak energy efficiency of 41.2 GOPS/mm2 and 90.63 TOPS/W. In comparison to state-of-the-art reference designs, the proposed DQ-STP demonstrates a$2.19\times $improvement in normalized area efficiency and a$1.85\times $enhancement in energy efficiency.
Baoting Li, Danqing Zhang, Xuchong Zhang, Hongbin Sun 0001, Nanning Zheng 0001
IEEE Trans. Circuits Syst. I Regul. Pap.7
2024 Knowledge Graph Enhancement for Fine-Grained Zero-Shot Learning on ImageNet21K
abstract
Fine-grained Zero-shot Learning on the large-scale dataset ImageNet21K is an important task that has promising perspectives in many real-world scenarios. One typical solution is to explicitly model the knowledge passing using a Knowledge Graph (KG) to transfer knowledge from seen to unseen instances. By analyzing the hierarchical structure and the word descriptions on ImageNet21K, we find that the noisy semantic information, the sparseness of seen classes, and the lack of supervision of unseen classes make the knowledge passing insufficient, which limits the KG-based fine-grained ZSL. To resolve this problem, in this paper, we enhance the knowledge passing from three aspects. First, we use more powerful models such as the Large Language Model and Vision-Language Model to get more reliable semantic embeddings. Then we propose a strategy that globally enhances the knowledge graph based on the convex combination relationship of the semantic embeddings. It effectively connects the edges between the non-kinship seen and unseen classes that have strong correlations while assigning an importance score to each edge. Based on the enhanced knowledge graph, we further present a novel regularizer that locally enhances the knowledge passing during training. We extensively conducted comparative evaluations to demonstrate the advantages of our method over state-of-the-art approaches.
Xingyu Chen 0001, Zeyang Liu 0001, Lipeng Wan 0003, Xuguang Lan, Nanning Zheng 0001
IEEE Trans. Circuits Syst. Video Technol.6
2024 Selective Transfer Learning of Cross-Modality Distillation for Monocular 3D Object Detection
abstract
Monocular 3D object detection is a promising yet ill-posed task for autonomous vehicles due to the lack of accurate depth information. Cross-modality knowledge distillation could effectively transfer depth information from LiDAR to image-based network. However, modality gap between image and LiDAR seriously limits its accuracy. In this paper, we systematically investigate the negative transfer problem induced by modality gap in cross-modality distillation for the first time, including not only the architecture inconsistency issue but more importantly the feature overfitting issue. We propose a selective learning approach named MonoSTL to overcome these issues, which encourages positive transfer of depth information from LiDAR while alleviates the negative transfer on image-based network. On the one hand, we utilize similar architectures to ensure spatial alignment of features between image-based and LiDAR-based networks. On the other hand, we develop two novel distillation modules, namely Depth-Aware Selective Feature Distillation (DASFD) and Depth-Aware Selective Relation Distillation (DASRD), which selectively learn positive features and relationships of objects by integrating depth uncertainty into feature and relation distillations, respectively. Our approach can be seamlessly integrated into various CNN-based and DETR-based models, where we take three recent models on KITTI and a recent model on NuScenes for validation. Extensive experiments show that our approach considerably improves the accuracy of the base models and thereby achieves the best accuracy compared with all recently released SOTA models. The code is released on https://github.com/DingCodeLab/MonoSTL.
Meng Yang 0002, Nanning Zheng 0001
IEEE Trans. Circuits Syst. Video Technol.3
2024 FS-Depth: Focal-and-Scale Depth Estimation From a Single Image in Unseen Indoor Scene
abstract
It has long been an ill-posed problem to predict absolute depth maps from single images in unseen scenes. We observe that it is essentially due to not only the scale-ambiguous problem, but more importantly, the focal-ambiguous problem that decreases the generalization ability of monocular depth estimation. That is, images may be captured by cameras of different focal lengths in scenes of different scales. In this paper, we develop a focal-and-scale depth estimation model to well learn absolute depth maps from single images in unseen indoor scenes. First, a relative depth estimation network is adopted to learn relative depths from single images with diverse scales. Second, multi-scale features are generated by mapping a single focal length value to focal length features and concatenating them with intermediate features of different scales in relative depth estimation. Finally, relative depths and multi-scale features are jointly fed into an absolute depth estimation network. Our model is enabled to be well trained on either a single dataset or a mixed dataset with diverse focal lengths and scene scales by a dual-directional alignment strategy. In addition, a new pipeline is developed to augment the diversity of focal lengths of public datasets, which are often captured with cameras of the same or similar focal lengths. The experiments verify that our model trained on NYUDv2 significantly improves the generalization ability of monocular depth estimation by 32%/14% (RMSE) on three unseen datasets with/without data augmentation compared with state-of-the-art (SOTA) baselines, and well alleviates the deformation problem of depth maps in 3D view. The generalization ability is further improved by 16% when the model is trained on a mixture of NYUDv2 and SUNRGBD. In addition, our model maintains a SOTA accuracy, when it is trained and tested on NYUDv2 similar to existing models. The code is released onhttps://github.com/wcrwcrwcr/FS-Depth-v1.
Chengrui Wei, Meng Yang 0002, Nanning Zheng 0001
IEEE Trans. Circuits Syst. Video Technol.4
2024 Cross Time-Frequency Transformer for Temporal Action Localization
abstract
Most modern approaches in temporal action localization (TAL) mainly focus on time domain information, while neglecting the advantages of information from other domains. How to effectively utilize information from different domains and their interactions in a reasonable manner has been an attractive yet challenging issue in TAL. In this paper, we propose a novel cross time-frequency Transformer model (TFFormer) for TAL. A dual-branch network architecture is designed to capture the time and frequency features at multiple scales, using the multi-scale transformer in the time branch and the DB1 Discrete Wavelet Transform (DWT) in the frequency branch. To fuse these features from different domains, we propose a cross time-frequency attention mechanism that includes a time pathway and a frequency pathway, enhancing the interaction between the temporal and frequency features. Furthermore, a gated control mechanism is designed to aggregate features from different scales, characterizing the respective contributions of features at different scales. We also design a new regression loss function for locating the time boundaries. Extensive experiments were carried out on four challenging benchmark datasets, including two third-person datasets and two first-person datasets. The proposed method achieves impressive results on these datasets. Specifically, TFFormer achieves an average mAP of 23.2% on Ego4D and 25.6% on EPIC-Kitchens 100, which outperform previous state-of-the-arts by a large margin. It also obtains competitive results on ActivityNet v1.3 and THUMOS14, with an average mAP of 36.2% and 67.8%. We also conducted extensive ablation studies to validate the effectiveness of each component in the proposed method.
Ping Wei 0001, Nanning Zheng 0001
IEEE Trans. Circuits Syst. Video Technol.3
2024 FFINet: Future Feedback Interaction Network for Motion Forecasting
abstract
Motion forecasting plays a crucial role in autonomous driving, with the aim of predicting the future reasonable motions of traffic agents. Most existing methods mainly model the historical interactions between agents and the environment, and predict multi-modal trajectories in a feedforward process, ignoring potential trajectory changes caused by future interactions between agents. In this paper, we propose a novel Future Feedback Interaction Network (FFINet) to aggregate the current, observations and potential future interaction features for trajectory prediction. Firstly, we employ different spatial-temporal encoders to embed the decomposed position vectors and the current position of each scene, providing rich features for the subsequent cross-temporal aggregation. Secondly, the relative interaction and cross-temporal aggregation strategies are sequentially adopted to integrate features in the current fusion module, observation interaction module, future feedback module and global fusion module, in which the future feedback module can enable the understanding of pre-action by feeding the influence of preview information to feedforward prediction. Thirdly, the comprehensive interaction features are further fed into final predictor to generate the joint predicted trajectories of multiple agents. Extensive experimental results show that our FFINet achieves the state-of-the-art performance on Argoverse 1 and Argoverse 2 motion forecasting benchmarks.
Miao Kang, Shengqi Wang, Sanping Zhou, Ke Ye, Jingjing Jiang, Nanning Zheng 0001
IEEE Trans. Intell. Transp. Syst.6
2024 Clothoid-Based Reference Path Reconstruction for HD Map Generation
abstract
High-definition (HD) map is one of the key assets for autonomous driving, which supports various modules such as behavior prediction and motion planning of autonomous vehicles by providing accurate and rich geometric and semantic information. However, at present, the scalability and computational efficiency of HD map generation cannot meet the needs of highly automated driving. Specifically, efficiently obtaining the optimal parameters of the road’s reference path is still an open problem. In this paper, we propose a fast and robust path reconstruction method, which compresses the dense points of a reference line into sparse parameters with minimal loss of information. The reconstructed path consists of segmented linear curvature contours, which are straight lines, circular arcs, and clothoids. The optimum result is obtained through linear programming for short-path reconstruction, and for the long paths, a fast progressive reconstruction approach is used to find a feasible solution. Experimental results on both randomly generated data and the GPS-collected trajectories show that compared with existing methods, the proposed method can generate more accurate path reconstruction, and the computational time is greatly reduced.
Songyi Zhang, Runsheng Wang, Zhiqiang Jian, Nanning Zheng 0001, Masayoshi Tomizuka
IEEE Trans. Intell. Transp. Syst.5
2024 Single-Shot and Multi-Shot Feature Learning for Multi-Object Tracking
abstract
Multi-Object Tracking (MOT) remains a vital component of intelligent video analysis, which aims to locate targets and maintain a consistent identity for each target throughout a video sequence. Existing works usually learn a discriminative feature representation, such as motion and appearance, to associate the detections across frames, which are easily affected by mutual occlusion and background clutter in practice. In this paper, we propose a simple yet effective two-stage feature learning paradigm to jointly learn single-shot and multi-shot features for different targets, so as to achieve robust data association in the tracking process. For the detections without being associated, we design a novel single-shot feature learning module to extract discriminative features of each detection, which can efficiently associate targets between adjacent frames. For the tracklets being lost several frames, we design a novel multi-shot feature learning module to extract discriminative features of each tracklet, which can accurately refind these lost targets after a long period. Once equipped with a simple data association logic, the resulting VisualTracker can perform robust MOT based on the single-shot and multi-shot feature representations. Extensive experimental results demonstrate that our method has achieved significant improvements on MOT17 and MOT20 datasets while reaching state-of-the-art performance on DanceTrack dataset.
Sanping Zhou, Le Wang 0003, Jinjun Wang, Nanning Zheng 0001
IEEE Trans. Multim.6
2024 3D Scene Graph Generation From Point Clouds
abstract
Scene graph generation is a significant and challenging task for scene understanding. Most existing methods are confined to the 2D space (i.e. images) or additional use of segmentation information, while neglecting the richer spatial and geometric information of 3D space. In this paper, we propose a novel method to generate scene graphs from 3D point clouds. Specifically, our model consists of three parts: a point feature extraction backbone, a box head, and a relation head. The feature extraction backbone extracts base features directly from raw point clouds, and the box head produces detected 3D bounding boxes. Final 3D scene graphs are obtained from the relation head which takes the extracted features and 3D boxes as inputs. We also design a point RoI module which sequentially processes points inside 3D boxes with a bidirectional LSTM. To further leverage the geometric characteristics of point clouds, we propose a location attention module which learns the influence of relative locations between objects. We introduce the RelationScanNet dataset with densely annotated semantic and geometric relationships, which extends one of the most widely used dataset ScanNetV2 in 3D indoor scene understanding. We test the proposed method on the RelationScanNet dataset and 3DSSG dataset. The results prove the strength of our method.
Wenwen Wei, Ping Wei 0001, Jialu Qin, Zhimin Liao, Shuaijie Wang, Xiang Cheng 0001, Meiqin Liu 0001, Nanning Zheng 0001
IEEE Trans. Multim.8
2024 Gated Multi-Scale Transformer for Temporal Action Localization
abstract
Temporal action localization (TAL) is a critical task in video understanding. Effectively utilizing multi-scale information and handling interactions across various scales have consistently posed challenging issues within the realm of TAL. In this paper, we propose a novel gated multi-scale Transformer model (TransGMC) for temporal action localization. A gated control mechanism is designed to filter and aggregate the information at different scales, by which the contributions of contexts at different temporal scales are well characterized. To enhance the feature representation at each temporal scale, the rich global-local contexts are extracted at each temporal scale. A cascade attention module that contains two seamlessly integrated channel attention and moment attention is proposed for capturing global temporal contexts. We utilize a new regression loss function for locating the time boundaries. We conducted experiments on four challenging benchmark datasets, including two third-person view datasets and two first-person view datasets. Our method achieves an average mAP of 67.5% on THUMOS14, 36.1% on ActivityNet v1.3, 24.9% on EPIC-Kitchens 100, and 23.2% on Ego4D, which all outperform the previous state-of-the-arts methods. Extensive ablation studies also validate the effectiveness of the proposed method. Code is available at https://github.com/EdenGabriel/TransGMC.
Ping Wei 0001, Ziyang Ren, Nanning Zheng 0001
IEEE Trans. Multim.4
2024 Inverse Adversarial Diversity Learning for Network Ensemble
abstract
Network ensemble aims to obtain better results by aggregating the predictions of multiple weak networks, in which how to keep the diversity of different networks plays a critical role in the training process. Many existing approaches keep this kind of diversity either by simply using different network initializations or data partitions, which often requires repeated attempts to pursue a relatively high performance. In this article, we propose a novel inverse adversarial diversity learning (IADL) method to learn a simple yet effective ensemble regime, which can be easily implemented in the following two steps. First, we take each weak network as a generator and design a discriminator to judge the difference between the features extracted by different weak networks. Second, we present an inverse adversarial diversity constraint to push the discriminator to cheat generators that all the resulting features of the same image are too similar to distinguish each other. As a result, diverse features will be extracted by these weak networks through a min-max optimization. What is more, our method can be applied to a variety of tasks, such as image classification and image retrieval, by applying a multitask learning objective function to train all these weak networks in an end-to-end manner. We conduct extensive experiments on the CIFAR-10, CIFAR-100, CUB200-2011, and CARS196 datasets, in which the results show that our method significantly outperforms most of the state-of-the-art approaches.
Sanping Zhou, Jinjun Wang, Le Wang 0003, Xingyu Wan, Siqi Hui, Nanning Zheng 0001
IEEE Trans. Neural Networks Learn. Syst.6
2023 How Do In-Context Examples Affect Compositional Generalization?
abstract
Shengnan An, Zeqi Lin, Qiang Fu, Bei Chen, Nanning Zheng, Jian-Guang Lou, Dongmei Zhang. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Shengnan An, Zeqi Lin, Qiang Fu 0015, Bei Chen 0008, Nanning Zheng 0001, Jian-Guang Lou, Dongmei Zhang 0001
ACL (1)5
2023 Unsupervised Domain Adaptation by Cross-Prototype Contrastive Learning for Medical Image Segmentation
abstract
Unsupervised Domain Adaptation (UDA), which aligns the labeled source distribution to the unlabeled target distribution, has shown remarkable achievement in the medical image segmentation task. Previous UDA methods unilaterally consider the global distribution alignment through explicit category-based loss while good separation and discrimination of class are insufficiently explored, resulting in the sub-aligned distribution across domains. In this paper, we propose cross-prototype contrastive learning method (CPCL) for UDA segmentation through class centroid alignment. Specifically, to reduce the intra-class distance and increase the inter-class distance, we first introduce prototype-feature contrastive learning to align the pixel-level features and the same-class global prototype across domains. Secondly, we further present prototype-prototype contrastive learning to align the same class prototypes between the source domain and target domain for compact category centroid and better global domain distribution alignment. Extensive experiments on two public cardiac datasets demonstrate that the proposed CPCL achieves superior domain adaptation performance as compared with the state-of-the-art.
Zhuotong Cai, Jingmin Xin, Siyuan Dong, Chenyu You, Peiwen Shi, Tianyi Zeng, John A. Onofrey, Nanning Zheng 0001, James S. Duncan
BIBM9
2023 Cervical Cytology Classification with Coarse Labels Based on Two-Stage Weakly Supervised Contrastive Learning Framework
abstract
Deep learning methods have achieved remarkable success in various tasks from cervical cytology images. However, for the gigapixel whole slide images (WSIs), the acquisition of annotations is a time-consuming and labor-intensive task requiring a high level of expertise. While the sparse distribution of malignant cells and the factors above pose great difficulties to label thousands of patches divided from the WSI, it is much easier to obtain the coarse labels at the WSI level. In this paper, we propose a novel weakly supervised contrastive learning framework, which utilizes only coarse labels from the WSIs for cervical cytology patch classification. The proposed framework consists of two stages, including the representation learning stage and the classifier finetuning stage. In the first stage, to effectively exploit useful information of coarse labels, we devise a re-weight cross-entropy loss, which can fast warm up the training and reduce the inexact supervision from the coarse labels simultaneously. To further excavate features bypassing the coarse labels, we propose a self-supervised contrastive loss, where the random augmentation and the mean teacher architecture enrich the external variations, and help better extract representations through patch similarities. In the second stage, based on ensemble predictions and uncertainty selections, reliable pseudo labels are generated for the inaccurate labels to finetune the classifier, with better performance achieved. Extensive experiments on the in-house dataset demonstrate that the proposed method is more efficient than other state-of-the-art methods. Our code is available on https://github.com/chaisiyii/WSCL.
Siyi Chai, Jingmin Xin, Jiayi Wu 0002, Hongxuan Yu, Zhaohai Liang, Nanning Zheng 0001
BIBM7
2023 MixPHM: Redundancy-Aware Parameter-Efficient Tuning for Low-Resource Visual Question Answering
abstract
Recently, finetuning pretrained vision-language models (VLMs) has been a prevailing paradigm for achieving state-of-the-art performance in VQA. However, as VLMs scale, it becomes computationally expensive, storage inefficient, and prone to overfitting when tuning full model parameters for a specific task in low-resource settings. Although current parameter-efficient tuning methods dramatically reduce the number of tunable parameters, there still exists a significant performance gap with full finetuning. In this paper, we propose MixPHM, a redundancy-aware parameter-efficient tuning method that outperforms full finetuning in low-resource VQA. Specifically, MixPHM is a lightweight module implemented by multiple PHM-experts in a mixture-of-experts manner. To reduce parameter redundancy, we reparameterize expert weights in a low-rank subspace and share part of the weights inside and across MixPHM. Moreover, based on our quantitative analysis of representation redundancy, we propose Redundancy Regularization, which facilitates MixPHM to reduce task-irrelevant redundancy while promoting task-relevant correlation. Experiments conducted on VQA v2, GQA, and OK-VQA with different low-resource settings show that our MixPHM outperforms state-of-the-art parameter-efficient methods and is the only one consistently surpassing full finetuning.
Jingjing Jiang, Nanning Zheng 0001
CVPR2
2023 StructVPR: Distill Structural Knowledge with Weighting Samples for Visual Place Recognition
abstract
Visual place recognition (VPR) is usually considered as a specific image retrieval problem. Limited by existing training frameworks, most deep learning-based works cannot extract sufficiently stable global features from RGB images and rely on a time-consuming re-ranking step to exploit spatial structural information for better performance. In this paper, we propose StructVPR, a novel training architecture for VPR, to enhance structural knowledge in RGB global features and thus improve feature stability in a constantly changing environment. Specifically, StructVPR uses segmentation images as a more definitive source of structural knowledge input into a CNN network and applies knowledge distillation to avoid online segmentation and inference of seg-branch in testing. Considering that not all samples contain high-quality and helpful knowledge, and some even hurt the performance of distillation, we partition samples and weigh each sample's distillation loss to enhance the expected knowledge precisely. Finally, StructVPR achieves impressive performance on several benchmarks using only global retrieval and even outperforms many two-stage approaches by a large margin. After adding additional re-ranking, ours achieves state-of-the-art performance while maintaining a low computational cost.
Yanqing Shen, Sanping Zhou, Jingwen Fu, Ruotong Wang 0005, Shi-tao Chen, Nanning Zheng 0001
CVPR6
2023 Skill-Based Few-Shot Selection for In-Context Learning
abstract
In-context learning is the paradigm that adapts large language models to downstream tasks by providing a few examples.Few-shot selectionselecting appropriate examples for each test instance separately-is important for in-context learning.In this paper, we propose SKILL-KNN, a skill-based few-shot selection method for in-context learning.The key advantages of SKILL-KNN include: (1) it addresses the problem that existing methods based on pre-trained embeddings can be easily biased by surface natural language features that are not important for the target task; (2) it does not require training or fine-tuning of any models, making it suitable for frequently expanding or changing example banks.The key insight is to optimize the inputs fed into the embedding model, rather than tuning the model itself.Technically, SKILL-KNN generates the skill-based descriptions for each test case and candidate example by utilizing a pre-processing few-shot prompting, thus eliminating unimportant surface features.Experimental results across five cross-domain semantic parsing datasets and six backbone models show that SKILL-KNN significantly outperforms existing methods.* Work done during the internship at Microsoft.DB Schema: employee (name, age, city, …) … Question: Which cities do more than one employee under age 30 come from?Prompting-Based Rewriting Off-the-Shelf Embedding Model Skill-Based Descriptions Input Query Candidate 1:This task requires the greater-than and less-than constraints.Candidate 2:This task requires to apply two constraints on one selected column.Candidate … … Example Bank Skill-Based Selection DB Schema: cinema (name, capacity, location, …) … Question: Find the locations that have more than one movie theater with capacity above 300.SQL Query: SELECT … WHERE capacity > 300 GROUP BY location HAVING count(*) > 1 DB Schema: endowment (school, donator, amount, …) … Question : Find the number of schools that have more than one donator whose donation amount is less than 8
Shengnan An, Zeqi Lin, Qiang Fu 0015, Bei Chen 0008, Nanning Zheng 0001, Weizhu Chen, Jian-Guang Lou
EMNLP6
2023 MLF-DET: Multi-Level Fusion for Cross-Modal 3D Object Detection
Zewei Lin, Yanqing Shen, Sanping Zhou, Shi-tao Chen, Nanning Zheng 0001
ICANN (7)5
2023 Cross-Graph Transformer Network for Temporal Sentence Grounding
Jiahui Shang, Ping Wei 0001, Nanning Zheng 0001
ICANN (6)3
2023 Temporal Deformable Transformer for Action Localization
Haoying Wang, Ping Wei 0001, Meiqin Liu 0001, Nanning Zheng 0001
ICANN (6)4
2023 Inverse Compositional Learning for Weakly-supervised Relation Grounding
abstract
Video relation grounding (VRG) is a significant and challenging problem in the domains of cross-modal learning and video understanding. In this study, we introduce a novel approach called inverse compositional learning (ICL) for weakly-supervised video relation grounding. Our approach represents relations at both the holistic and partial levels, formulating VRG as a joint optimization problem that encompasses reasoning at both levels. For holistic-level reasoning, we propose an inverse attention mechanism and a compositional encoder to generate compositional relevance features. Additionally, we introduce an inverse loss to evaluate and learn the relevance between visual features and relation features. At the partial-level reasoning, we introduce a grounding by classification scheme. By leveraging the learned holistic-level features and partial-level features, we train the entire model in an end-to-end manner. We conduct evaluations on two challenging datasets and demonstrate the substantial superiority of our proposed method over state-of-the-art methods. Extensive ablation studies confirm the effectiveness of our approach.
Ping Wei 0001, Zeyu Ma 0005, Nanning Zheng 0001
ICCV4
2023 DETR Does Not Need Multi-Scale or Locality Design
abstract
This paper presents an improved DETR detector that maintains a "plain" nature: using a single-scale feature map and global cross-attention calculations without specific locality constraints, in contrast to previous leading DETR-based detectors that reintroduce architectural inductive biases of multi-scale and locality into the decoder. We show that two simple technologies are surprisingly effective within a plain design to compensate for the lack of multi-scale feature maps and locality constraints. The first is a box-to-pixel relative position bias (BoxRPB) term added to the cross-attention formulation, which well guides each query to attend to the corresponding object region while also providing encoding flexibility. The second is masked image modeling (MIM)-based backbone pre-training which helps learn representation with fine-grained localization ability and proves crucial for remedying dependencies on the multi-scale feature maps. By incorporating these technologies and recent advancements in training and problem formation, the improved "plain" DETR showed exceptional improvements over the original DETR detector. By leveraging the Object365 dataset for pre-training, it achieved 63.9 mAP accuracy using a Swin-L backbone, which is highly competitive with state-of-the-art detectors which all heavily rely on multi-scale feature maps and region-based feature extraction. Code will be available at https://github.com/impiga/Plain-DETR.
Yutong Lin, Yuhui Yuan, Zheng Zhang 0022, Nanning Zheng 0001, Han Hu 0001
ICCV5
2023 Does Deep Learning Learn to Abstract? A Systematic Probing Framework
Shengnan An, Zeqi Lin, Bei Chen 0008, Qiang Fu 0015, Nanning Zheng 0001, Jian-Guang Lou
ICLR5
2023 DBQ-SSD: Dynamic Ball Query for Efficient 3D Object Detection
Lin Song 0002, Weixin Mao, Xiaoping Li 0005, Hongbin Sun 0001, Jian Sun 0001, Nanning Zheng 0001
ICLR9
2023 MMRDN: Consistent Representation for Multi-View Manipulation Relationship Detection in Object-Stacked Scenes
abstract
Manipulation relationship detection (MRD) aims to guide the robot to grasp objects in the right order, which is important to ensure the safety and reliability of grasping in object stacked scenes. Previous works infer manipulation relationship by deep neural network trained with data collected from a predefined view, which has limitation in visual dislocation in unstructured environments. Multi-view data provide more comprehensive information in space, while a challenge of multi-view MRD is domain shift. In this paper, we propose a novel multi-view fusion framework, namely multi-view MRD network (MMRDN), which is trained by 2D and 3D multi-view data. We project the 2D data from different views into a common hidden space and fit the embeddings with a set of Von-Mises-Fisher distributions to learn the consistent representations. Besides, taking advantage of position information within the 3D data, we select a set of$K$Maximum Vertical Neighbors (KMVN) points from the point cloud of each object pair, which encodes the relative position of these two objects. Finally, the features of multi-view 2D and 3D data are concatenated to predict the pairwise relationship of objects. Experimental results on the challenging REGRAD dataset show that MMRDN outperforms the state-of-the-art methods in multi-view MRD tasks. The results also demonstrate that our model trained by synthetic data is capable to transfer to real-world scenarios.
Lipeng Wan 0003, Xingyu Chen 0001, Xuguang Lan, Nanning Zheng 0001
ICRA6
2023 TEMI-MOT: Towards Efficient Multi-Modality Instance-Aware Feature Learning for 3D Multi-Object Tracking
abstract
3D multi-object tracking is one of the key technologies of autonomous driving, which aims to ensure that autonomous driving vehicles accurately perceive the movements and intentions of surrounding traffic participants. In recent years, some 3D multi-object tracking methods based on multimodality have been proposed. Although these methods improve the accuracy of object association in the tracking process, these methods are still difficult to effectively deal with the problems of feature ambiguity due to occlusion, incorrect feature alignment between different modalities, and confusion of adjacent target features caused by coarse-grained feature maps. To address these problems, we propose a new multi-modality feature learning method for 3D multi-object tracking, named TEMI-MOT, which is composed of three modules in series: the point-guided image feature sampler, the instance-aware feature encoder, and the tracking pipeline. The point-guided feature sampler realizes the alignment between the point cloud and image features, the instance-aware feature encoder fuses the aligned image features with each object's points to generate the discriminative instance-aware features, and the tracking pipeline finally outputs the results based on instance-aware features and G-IoU geometric similarities. Our approach achieves state-of-the-art results on the nuScenes dataset among the methods using CenterPoint detections. The experimental results show that the proposed method has better robustness and effectiveness for 3D multi-object tracking.
Sanping Zhou, Jinpeng Dong, Nanning Zheng 0001
IJCNN4
2023 InteractionNet: Joint Planning and Prediction for Autonomous Driving with Transformers
abstract
Planning and prediction are two important modules of autonomous driving and have experienced tremendous advancement recently. Nevertheless, most existing methods regard planning and prediction as independent and ignore the correlation between them, leading to the lack of consideration for interaction and dynamic changes of traffic scenarios. To address this challenge, we propose InteractionNet, which leverages transformer to share global contextual reasoning among all traffic participants to capture interaction and interconnect planning and prediction to achieve joint. Besides, InteractionNet deploys another transformer to help the model pay extra attention to the perceived region containing critical or unseen vehicles. InteractionNet outperforms other baselines in several benchmarks, especially in terms of safety, which benefits from the joint consideration of planning and forecasting. The code will be available at https://github.com/fujiawei0724/InteractionNet.
Jiawei Fu 0001, Yanqing Shen, Zhiqiang Jian, Shi-tao Chen, Jingmin Xin, Nanning Zheng 0001
IROS6
2023 Prioritized Planning for Target-Oriented Manipulation via Hierarchical Stacking Relationship Prediction
abstract
In scenarios involving grasping multiple targets, the learning of stacking relationships between objects is fundamental for robots to execute safely and efficiently. However, current methods lack subdivision for the hierarchy of stacking relationship types. In scenes where objects are mostly stacked in an orderly manner, they are incapable of performing human-like and high-efficient grasping decisions. This paper proposes a perception-planning method to distinguish different stacking forms between objects and generate prioritized manipulation sequences based on given target designations. We utilize a Hierarchical Stacking Relationship Network (HSRN) to discriminate the hierarchy of stacking and generate a refined Stacking Relationship Tree (SRT) for relationship description. Considering objects with high stacking stability can be processed together if necessary, we introduce an elaborate decision-making planner based on Partially Observable Markov Decision Process (POMDP), which leverages observations and generates the least grasp-consuming decision chain with robustness and is suitable for simultaneously specifying multiple targets. To verify our work, we set the scene to the dining table and augment REGRAD dataset for network training. Experiments show that our method effectively generates grasping decisions that conform to human requirements, and improves the implementation efficiency compared with existing methods on the basis of guaranteeing success rate.
Zewen Wu, Xingyu Chen 0001, Chengzhong Ma, Xuguang Lan, Nanning Zheng 0001
IROS6
2023 Efficient Lane-changing Behavior Planning via Reinforcement Learning with Imitation Learning Initialization
abstract
Robust lane-changing behavior planning is critical to ensuring the safety and comfort of autonomous vehicles. In this paper, we proposed an efficient and robust vehicle lane-changing behavior decision-making method based on reinforcement learning (RL) and imitation learning (IL) initialization which learns the potential lane-changing driving mechanisms from driving mechanism from the interactions between vehicle and environment, so as to simplify the manual driving modeling and have good adaptability to the dynamic changes of lane-changing scene. Our method further makes the following improvements on the basis of the Proximal Policy Optimization (PPO) algorithm: (1) A dynamic hybrid reward mechanism for lane-changing tasks is adopted; (2) A state space construction method based on fuzzy logic and deformation pose is presented to enable behavior planning to learn more refined tactical decision-making; (3) An RL initialization method based on imitation learning which only requires a small amount of scene data is introduced to solve the low efficiency of RL learning under sparse reward. Experiments on the SUMO show the effectiveness of the proposed method, and the test on the CARLA simulator also verifies the generalization ability of the method.
Jiamin Shi, Tangyike Zhang, Junxiang Zhan, Shi-tao Chen, Jingmin Xin, Nanning Zheng 0001
IV6
2023 Design Hybrid Computing Architecture for Accelerating Point Cloud Registration
abstract
High-precision simultaneous localization and mapping (SLAM) is one of the core technologies of unmanned driving. LiDAR-based SLAM algorithms are often complex and computationally intensive, and usually are deployed on high performance CPU or GPU computing architecture with high power consumption and low energy efficiency ratio, which is not conducive to vehicle-level applications. In this paper, we design and implement a low power CPU and FPGA hybrid computing architecture for accelerating the key algorithm of LiDAR-based localization scheme. More specifically, we propose a software and hardware co-design strategy: (1) we first propose chain representation as a new type of map representation, which uses the depth discontinuity region as the segmentation location to segment the point cloud data. Our method not only reduces noise issues for down-sampling operation in point cloud representation, but also has the same computational and storage overhead as point cloud representation. (2) We further exploit the inherent parallelism in the algorithms to design a pipeline hardware architecture, which can effectively improve the speed of the algorithm in the embedded platform. Deployed on the Xilinx ZCU102 platform, our system achieves 24.4x and 3.2x speedups compared to the ARM Cortex A53 processor and the Intel i7-10700 processor, respectively, at 4.204W power consumption without severely degrading the final output quality.
Xiao Wang 0002, Xiaodong Deng, Yingxiang Li, Shi-tao Chen, Longjun Liu, Nanning Zheng 0001
IV6
2023 Interpretable Driver Fatigue Estimation Based on Hierarchical Symptom Representations
Jiaqin Lin, Shaoyi Du, Yuying Liu 0007, Nanning Zheng 0001
MMM (2)6
2023 Learning Trajectories are Generalization Indicators
abstract
This paper explores the connection between learning trajectories of Deep Neural Networks (DNNs) and their generalization capabilities when optimized using (stochastic) gradient descent algorithms. Instead of concentrating solely on the generalization error of the DNN post-training, we present a novel perspective for analyzing generalization error by investigating the contribution of each update step to the change in generalization error. This perspective enable a more direct comprehension of how the learning trajectory influences generalization error. Building upon this analysis, we propose a new generalization bound that incorporates more extensive trajectory information. Our proposed generalization bound depends on the complexity of learning trajectory and the ratio between the bias and diversity of training set. Experimental observations reveal that our method effectively captures the generalization error throughout the training process. Furthermore, our approach can also track changes in generalization error when adjustments are made to learning rates and label noise levels. These results demonstrate that learning trajectory information is a valuable indicator of a model's generalization capabilities.
Jingwen Fu, Zhizheng Zhang 0004, Dacheng Yin, Yan Lu 0001, Nanning Zheng 0001
NeurIPS5
2023 Closing the gap between the upper bound and lower bound of Adam's iteration complexity
abstract
Recently, Arjevani et al. [1] establish a lower bound of iteration complexity for the first-order optimization under an $L$-smooth condition and a bounded noise variance assumption. However, a thorough review of existing literature on Adam's convergence reveals a noticeable gap: none of them meet the above lower bound. In this paper, we close the gap by deriving a new convergence guarantee of Adam, with only an $L$-smooth condition and a bounded noise variance assumption. Our results remain valid across a broad spectrum of hyperparameters. Especially with properly chosen hyperparameters, we derive an upper bound of the iteration complexity of Adam and show that it meets the lower bound for first-order optimizers. To the best of our knowledge, this is the first to establish such a tight upper bound for Adam's convergence. Our proof utilizes novel techniques to handle the entanglement between momentum and adaptive learning rate and to convert the first-order term in the Descent Lemma to the gradient norm, which may be of independent interest.
Jingwen Fu, Huishuai Zhang, Nanning Zheng 0001, Wei Chen 0034
NeurIPS4
2023 Geometric Transformer with Interatomic Positional Encoding
abstract
The widespread adoption of Transformer architectures in various data modalities has opened new avenues for the applications in molecular modeling. Nevertheless, it remains elusive that whether the Transformer-based architecture can do molecular modeling as good as equivariant GNNs. In this paper, by designing Interatomic Positional Encoding (IPE) that parameterizes atomic environments as Transformer's positional encodings, we propose Geoformer, a novel geometric Transformer to effectively model molecular structures for various molecular property prediction. We evaluate Geoformer on several benchmarks, including the QM9 dataset and the recently proposed Molecule3D dataset. Compared with both Transformers and equivariant GNN models, Geoformer outperforms the state-of-the-art (SoTA) algorithms on QM9, and achieves the best performance on Molecule3D for both random and scaffold splits. By introducing IPE, Geoformer paves the way for molecular geometric modeling based on Transformer architecture. Codes are available at https://github.com/microsoft/AI2BMD/tree/Geoformer.
Shaoning Li, Tong Wang 0014, Bin Shao 0002, Nanning Zheng 0001, Tie-Yan Liu
NeurIPS5
2023 DisDiff: Unsupervised Disentanglement of Diffusion Probabilistic Models
abstract
Targeting to understand the underlying explainable factors behind observations and modeling the conditional generation process on these factors, we connect disentangled representation learning to diffusion probabilistic models (DPMs) to take advantage of the remarkable modeling ability of DPMs. We propose a new task, disentanglement of (DPMs): given a pre-trained DPM, without any annotations of the factors, the task is to automatically discover the inherent factors behind the observations and disentangle the gradient fields of DPM into sub-gradient fields, each conditioned on the representation of each discovered factor. With disentangled DPMs, those inherent factors can be automatically discovered, explicitly represented and clearly injected into the diffusion process via the sub-gradient fields. To tackle this task, we devise an unsupervised approach, named DisDiff, and for the first time achieving disentangled representation learning in the framework of DPMs. Extensive experiments on synthetic and real-world datasets demonstrate the effectiveness of DisDiff.
Tao Yang 0032, Yuwang Wang, Yan Lu 0001, Nanning Zheng 0001
NeurIPS4
2023 Consistency Inspection for Assembly of Bolt on Engine Using Multi-view Stereo
abstract
Assembly inspection is a crucial aspect of smart manufacturing to guarantee the quality of products. To address the challenges posed by complex workshop scenes, an assembly inspection algorithm based on multi-view stereo is proposed. Specifically, to enhance the accuracy of reconstruction, the multi-view stereo method utilizes segmentation attention-assisted depth estimation and point cloud fusion to mitigate the noise in the reconstructed point cloud. And the proposed method introduces normalized depth loss to improve the reconstruction ability of foreground pixels. Futhermore, this paper presents a novel 3D point cloud-based size estimation algorithm capable of accurately estimating the size of small objects in complex working scenes. Experimental results demonstrate the efficacy of the improved multi-view stereo algorithm in assembly scenes and verify the superiority of the proposed size estimation algorithm.
Minglv Jiang, Yuanliang Lu, Jianji Wang 0001, Nanning Zheng 0001
SMC5
2023 Curriculum classification network based on margin balancing multi-loss and ensemble learning
Shaoyi Du, Yuying Liu 0007, Xijing Wang, Yuting Chi, Nanning Zheng 0001, Yucheng Guo
Future Gener. Comput. Syst.5
2023 Transformer-Based Approach Via Contrastive Learning for Zero-Shot Detection
abstract
Zero-shot detection (ZSD) aims to locate and classify unseen objects in pictures or videos by semantic auxiliary information without additional training examples. Most of the existing ZSD methods are based on two-stage models, which achieve the detection of unseen classes by aligning object region proposals with semantic embeddings. However, these methods have several limitations, including poor region proposals for unseen classes, lack of consideration of semantic representations of unseen classes or their inter-class correlations, and domain bias towards seen classes, which can degrade overall performance. To address these issues, the Trans-ZSD framework is proposed, which is a transformer-based multi-scale contextual detection framework that explicitly exploits inter-class correlations between seen and unseen classes and optimizes feature distribution to learn discriminative features. Trans-ZSD is a single-stage approach that skips proposal generation and performs detection directly, allowing the encoding of long-term dependencies at multiple scales to learn contextual features while requiring fewer inductive biases. Trans-ZSD also introduces a foreground-background separation branch to alleviate the confusion of unseen classes and backgrounds, contrastive learning to learn inter-class uniqueness and reduce misclassification between similar classes, and explicit inter-class commonality learning to facilitate generalization between related classes. Trans-ZSD addresses the domain bias problem in end-to-end generalized zero-shot detection (GZSD) models by using balance loss to maximize response consistency between seen and unseen predictions, ensuring that the model does not bias towards seen classes. The Trans-ZSD framework is evaluated on the PASCAL VOC and MS COCO datasets, demonstrating significant improvements over existing ZSD models.
Wei Liu 0220, Hui Chen 0036, Jianji Wang 0001, Nanning Zheng 0001
Int. J. Neural Syst.5
2023 Multi-view Contour-constrained Transformer Network for Thin-cap Fibroatheroma Identification
abstract
Identification and detection of thin-cap fibroatheroma (TCFA) from intravascular optical coherence tomography (IVOCT) images is critical for treatment of coronary heart diseases. Recently, deep learning methods have shown promising successes in TCFA identification. However, most methods usually do not effectively utilize multi-view information or incorporate prior domain knowledge. In this paper, we propose a multi-view contour-constrained transformer network (MVCTN) for TCFA identification in IVOCT images. Inspired by the diagnosis process of cardiologists, we use contour constrained self-attention modules (CCSM) to emphasize features corresponding to salient regions (i.e., vessel walls) in an unsupervised manner and enhance the visual interpretability based on class activation mapping (CAM). Moreover, we exploit transformer modules (TM) to build global-range relations between two views (i.e., polar and Cartesian views) to effectively fuse features at multiple feature scales. Experimental results on a semi-public dataset and an in-house dataset demonstrate that the proposed MVCTN outperforms other single-view and multi-view methods. Lastly, the proposed MVCTN can also provide meaningful visualization for cardiologists via CAM.
Jingmin Xin, Jiayi Wu 0002, Yangyang Deng, Ruisheng Su, Wiro J. Niessen, Nanning Zheng 0001, Theo van Walsum
Neurocomputing7
2023 Batch normalization-free weight-binarized SNN based on hardware-saving IF neuron
Guanchao Qiao, Nanning Zheng 0001, Yue Zuo, Pujun Zhou, M. L. Sun, Shaogang Hu, Qi Yu 0002
Neurocomputing2
2023 Multi-weight susceptible-infected model for predicting COVID-19 in China
Jun Zhang 0003, Nanning Zheng 0001, Dingyi Yao, Jianji Wang 0001, Jingmin Xin
Neurocomputing2
2023 Multi-scale interaction transformer for temporal action proposal generation
Jiahui Shang, Ping Wei 0001, Nanning Zheng 0001
Image Vis. Comput.4
2023 RepCo: Replenish sample views with better consistency for contrastive learning
Longjun Liu, Yi Zhang 0140, Puhang Jia, Haonan Zhang 0002, Nanning Zheng 0001
Neural Networks6
2023 Learning to Infer Unseen Single-/ Multi-Attribute-Object Compositions With Graph Networks
abstract
Inferring the unseen attribute-object composition is critical to make machines learn to decompose and compose complex concepts like people. Most existing methods are limited to the composition recognition of single-attribute-object, and can hardly learn relations between the attributes and objects. In this paper, we propose an attribute-object semantic association graph model to learn the complex relations and enable knowledge transfer between primitives. With nodes representing attributes and objects, the graph can be constructed flexibly, which realizes both single- and multi-attribute-object composition recognition. In order to reduce mis-classifications of similar compositions (e.g., scratched screen and broken screen), driven by the contrastive loss, the anchor image feature is pulled closer to the corresponding label feature and pushed away from other negative label features. Specifically, a novel balance loss is proposed to alleviate the domain bias, where a model prefers to predict seen compositions. In addition, we build a large-scale Multi-Attribute Dataset (MAD) with 116,099 images and 8,030 label categories for inferring unseen multi-attribute-object compositions. Along with MAD, we propose two novel metrics Hard and Soft to give a comprehensive evaluation in the multi-attribute setting. Experiments on MAD and two other single-attribute-object benchmarks (MIT-States and UT-Zappos50K) demonstrate the effectiveness of our approach.
Hui Chen 0036, Jingjing Jiang, Nanning Zheng 0001
IEEE Trans. Pattern Anal. Mach. Intell.3
2023 Representing Multimodal Behaviors With Mean Location for Pedestrian Trajectory Prediction
abstract
Representing multimodal behaviors is a critical challenge for pedestrian trajectory prediction. Previous methods commonly represent this multimodality with multiple latent variables repeatedly sampled from a latent space, encountering difficulties in interpretable trajectory prediction. Moreover, the latent space is usually built by encoding global interaction into future trajectory, which inevitably introduces superfluous interactions and thus leads to performance reduction. To tackle these issues, we propose a novel Interpretable Multimodality Predictor (IMP) for pedestrian trajectory prediction, whose core is to represent a specific mode by its mean location. We model the distribution of mean location as a Gaussian Mixture Model (GMM) conditioned on sparse spatio-temporal features, and sample multiple mean locations from the decoupled components of GMM to encourage multimodality. Our IMP brings four-fold benefits: 1) Interpretable prediction to provide semantics about the motion behavior of a specific mode; 2) Friendly visualization to present multimodal behaviors; 3) Well theoretical feasibility to estimate the distribution of mean locations supported by the central-limit theorem; 4) Effective sparse spatio-temporal features to reduce superfluous interactions and model temporal continuity of interaction. Extensive experiments validate that our IMP not only outperforms state-of-the-art methods but also can achieve a controllable prediction by customizing the corresponding mean location.
Liushuai Shi, Le Wang 0003, Chengjiang Long, Sanping Zhou, Wei Tang 0016, Nanning Zheng 0001, Gang Hua 0001
IEEE Trans. Pattern Anal. Mach. Intell.6
2023 Adaptive Two-Stream Consensus Network for Weakly-Supervised Temporal Action Localization
abstract
Weakly-supervised temporal action localization (W-TAL) aims to classify and localize all action instances in untrimmed videos under only video-level supervision. Without frame-level annotations, it is challenging for W-TAL methods to clearly distinguish actions and background, which severely degrades the action boundary localization and action proposal scoring. In this paper, we present an adaptive two-stream consensus network (A-TSCN) to address this problem. Our A-TSCN features an iterative refinement training scheme: a frame-level pseudo ground truth is generated and iteratively updated from a late-fusion activation sequence, and used to provide frame-level supervision for improved model training. Besides, we introduce an adaptive attention normalization loss, which adaptively selects action and background snippets according to video attention distribution. By differentiating the attention values of the selected action snippets and background snippets, it forces the predicted attention to act as a binary selection and promotes the precise localization of action boundaries. Furthermore, we propose a video-level and a snippet-level uncertainty estimator, and they can mitigate the adverse effect caused by learning from noisy pseudo ground truth. Experiments conducted on the THUMOS14, ActivityNet v1.2, ActivityNet v1.3, and HACS datasets show that our A-TSCN outperforms current state-of-the-art methods, and even achieves comparable performance with several fully-supervised methods.
Yuanhao Zhai 0001, Le Wang 0003, Wei Tang 0016, Qilin Zhang 0004, Nanning Zheng 0001, David S. Doermann, Junsong Yuan 0001, Gang Hua 0001
IEEE Trans. Pattern Anal. Mach. Intell.5
2023 Towards Trajectory Forecasting From Detection
abstract
Trajectory forecasting for traffic participants (e.g., vehicles) is critical for autonomous platforms to make safe plans. Currently, most trajectory forecasting methods assume that object trajectories have been extracted and directly develop trajectory predictors based on the ground truth trajectories. However, this assumption does not hold in practical situations. Trajectories obtained from object detection and tracking are inevitably noisy, which could cause serious forecasting errors to predictors built on ground truth trajectories. In this paper, we propose to predict trajectories directly based on detection results without relying on explicitly formed trajectories. Different from traditional methods which encode the motion cues of an agent based on its clearly defined trajectory, we extract the motion information only based on the affinity cues among detection results, in which an affinity-aware state update mechanism is designed to manage the state information. In addition, considering that there could be multiple plausible matching candidates, we aggregate the states of them. These designs take the uncertainty of association into account which relax the undesirable effect of noisy trajectory obtained from data association and improve the robustness of the predictor. Extensive experiments validate the effectiveness of our method and its generalization ability to different detectors or forecasting schemes.
Pu Zhang 0001, Lei Bai 0001, Jianwu Fang, Jianru Xue, Nanning Zheng 0001, Wanli Ouyang
IEEE Trans. Pattern Anal. Mach. Intell.6
2023 ContextLoc++: A Unified Context Model for Temporal Action Localization
abstract
Effectively tackling the problem of temporal action localization (TAL) necessitates a visual representation that jointly pursues two confounding goals, i.e., fine-grained discrimination for temporal localization and sufficient visual invariance for action classification. We address this challenge by enriching the local, global and multi-scale contexts in the popular two-stage temporal localization framework. Our proposed model, dubbed ContextLoc++, can be divided into three sub-networks: L-Net, G-Net, and M-Net. L-Net enriches the local context via fine-grained modeling of snippet-level features, which is formulated as a query-and-retrieval process. Furthermore, the spatial and temporal snippet-level features, functioning as keys and values, are fused by temporal gating. G-Net enriches the global context via higher-level modeling of the video-level representation. In addition, we introduce a novel context adaptation module to adapt the global context to different proposals. M-Net further fuses the local and global contexts with multi-scale proposal features. Specially, proposal-level features from multi-scale video snippets can focus on different action characteristics. Short-term snippets with fewer frames pay attention to the action details while long-term snippets with more frames focus on the action variations. Experiments on the THUMOS14 and ActivityNet v1.3 datasets validate the efficacy of our method against existing state-of-the-art TAL algorithms.
Zixin Zhu, Le Wang 0003, Wei Tang 0016, Nanning Zheng 0001, Gang Hua 0001
IEEE Trans. Pattern Anal. Mach. Intell.4
2023 Surface normal and Gaussian weight constraints for indoor depth structure completion
Dongran Ren, Meng Yang 0002, Jiangfan Wu, Nanning Zheng 0001
Pattern Recognit.4
2023 Coarse-to-fine feature representation based on deformable partition attention for melanoma identification
Dong Zhang 0009, Jing Yang 0014, Shaoyi Du, Hongcheng Han, Yuyan Ge, Longfei Zhu, Ce Li 0001, Meifeng Xu, Nanning Zheng 0001
Pattern Recognit.9
2023 An Energy-and-Area-Efficient CNN Accelerator for Universal Powers-of-Two Quantization
abstract
CNN model computation on edge devices is tightly restricted to the limited resource and power budgets, which motivates the low-bit quantization technology to compress CNN models into 4-bit or lower format to reduce the model size and increase hardware efficiency. Most current low-bit quantization methods use uniform quantization that maps weight and activation values onto evenly-distributed levels, which usually results in accuracy loss due to distribution mismatch. Meanwhile, some non-uniform quantization methods propose specialized representation that can better match various distribution shapes but are usually difficult to be efficiently accelerated on hardware. In order to achieve low-bit quantization with high accuracy and hardware efficiency, this paper proposes Universal Power-of-Two (UPoT), a novel low-bit quantization method that represents values as the addition of multiple power-of-two values selected from a series of subsets. By updating the subset contents, UPoT can provide adaptive quantization levels for various distributions. For each CNN model layer, UPoT automatically searches for the optimized distribution that minimizes the quantization error. Moreover, we design an efficient accelerator system with specifically optimized power-of-two multipliers and requantization units. Evaluations show that the proposed architecture can provide high-performance CNN inference with reduced circuit area and energy, and outperforms several mainstream CNN accelerators with higher ($8\times $–$65\times $) area efficiency and ($2\times $–$19\times $) energy efficiency. Further experiments of 4/3/2-bit quantization on ResNet18/50, MobileNet_V2 and EfficientNet models show that our UPoT can achieve high model accuracy which greatly outperform other state-of-the-art low-bit quantization methods by 0.3%–6%. The results indicate that our approach provides a highly-efficient accelerator for low-bit CNN model quantization with low hardware overheads and good model accuracy.
Tian Xia 0008, Boran Zhao, Gelin Fu, Wenzhe Zhao 0001, Nanning Zheng 0001, Pengju Ren
IEEE Trans. Circuits Syst. I Regul. Pap.6
2023 Optimizing FPGA-Based DNN Accelerator With Shared Exponential Floating-Point Format
abstract
In recent years, low-precision fixed-point computation has become a widely used technique for neural network inference on FPGAs. However, this approach has some limitations, as certain neural networks are difficult to quantify using fixed-point arithmetic, such as those involved in super-resolution scaling, image denoising, and other scenarios that lack sufficient conditions for fine-tuning. Furthermore, deploying a floating-point precision neural network directly on an FPGA would lead to significant hardware overhead and low computational efficiency. To address this issue, this paper proposes an FPGA-friendly floating-point data format that achieves the same storage density as int8 without sacrificing inference accuracy or requiring fine-tuning. Additionally, this paper presents an FPGA-based neural network accelerator that is compatible with the proposed format, utilizing DSP resources to increase the number of DSP cascading from 7 to 16, and solving the back-to-back accumulation issue of floating-point numbers. This design achieves comparable resource consumption and execution efficiency to those of 8-bit fixed-point accelerators. Experimental results demonstrate that the accelerator proposed in this study achieves the same accuracy as the native floating point on multiple neural networks without fine-tuning, and remains high computing performance. When deployed on the Xilinx ZU9P, the performance achieves 4.072 TFlops at 250 MHz, which outperforms the previous works, including the Xilinx official DPU.
Wenzhe Zhao 0001, Qiwei Dang, Tian Xia 0008, Nanning Zheng 0001, Pengju Ren
IEEE Trans. Circuits Syst. I Regul. Pap.5
2023 Multimodal Pedestrian Trajectory Prediction Using Probabilistic Proposal Network
abstract
Forecasting multiple pedestrian trajectories is a challenging task for real-world applications, as the motion patterns of pedestrian are essentially stochastic and uncertain. Previous works have demonstrated that predicting diverse goals in advance can effectively improve the performance of pedestrian trajectory prediction. However, these methods are either unable to perform probabilistic and high-efficiency trajectory prediction, or mainly rely on the predefined template trajectories which are not high-performance and insufficient to represent the possible pedestrian behaviors. In this paper, we propose a new Probabilistic Proposal Network (PPNet) to concentrate on the generation of goals and the utilization of goal guidance. PPNet firstly generates multiple weighted goals based on the diverse latent intentions automatically obtained by unsupervised learning, and then designs the goal-conditioned Transformer networks to predict probabilistic proposals as the final trajectories. Extensive experimental results on ETH/UCY datasets and Stanford Drone Dataset indicate that PPNet achieves both state-of-the-art performance and high efficiency on pedestrian trajectory prediction.
Weihuang Chen, Lingyang Xue, Jinghai Duan, Hongbin Sun 0001, Nanning Zheng 0001
IEEE Trans. Circuits Syst. Video Technol.6
2023 Onboard Sensors-Based Self-Localization for Autonomous Vehicle With Hierarchical Map
abstract
Localization is a fundamental and crucial module for autonomous vehicles. Most of the existing localization methodologies, such as signal-dependent methods (RTK-GPS and Bluetooth), simultaneous localization and mapping (SLAM), and map-based methods, have been utilized in outdoor autonomous driving vehicles and indoor robot positioning. However, they suffer from severe limitations, such as signal-blocked scenes of GPS, computing resource occupation explosion in large-scale scenarios, intolerable time delay, and registration divergence of SLAM/map-based methods. In this article, a self-localization framework, without relying on GPS or any other wireless signals, is proposed. We demonstrate that the proposed homogeneous normal distribution transform algorithm and two-way information interaction mechanism could achieve centimeter-level localization accuracy, which reaches the requirement of autonomous vehicle localization for instantaneity and robustness. In addition, benefitting from hardware and software co-design, the proposed localization approach is extremely light-weighted enough to be operated on an embedded computing system, which is different from other LiDAR localization methods relying on high-performance CPU/GPU. Experiments on a public dataset (Baidu Apollo SouthBay dataset) and real-world verified the effectiveness and advantages of our approach compared with other similar algorithms.
Yanqing Shen, Yuedong Yang, Xiaodong Deng, Shi-tao Chen, Jingmin Xin, Nanning Zheng 0001
IEEE Trans. Cybern.7
2023 RGB-Guided Depth Map Recovery by Two-Stage Coarse-to-Fine Dense CRF Models
abstract
Depth maps generally suffer from large erroneous areas even in public RGB-Depth datasets. Existing learning-based depth recovery methods are limited by insufficient high-quality datasets and optimization-based methods generally depend on local contexts not to effectively correct large erroneous areas. This paper develops an RGB-guided depth map recovery method based on the fully connected conditional random field (dense CRF) model to jointly utilize local and global contexts of depth maps and RGB images. A high-quality depth map is inferred by maximizing its probability conditioned upon a low-quality depth map and a reference RGB image based on the dense CRF model. The optimization function is composed of redesigned unary and pairwise components, which constraint local structure and global structure of depth map, respectively, with the guidance of RGB image. In addition, the texture-copy artifacts problem is handled by two-stage dense CRF models in a coarse-to-fine way. A coarse depth map is first recovered by embedding RGB image in a dense CRF model in unit of $3\times 3$ blocks. It is refined afterward by embedding RGB image in another model in unit of individual pixels and restricting the model mainly work in discontinued regions. Extensive experiments on six datasets verify that the proposed method considerably outperforms a dozen of baseline methods in correcting erroneous areas and diminishing texture-copy artifacts of depth maps.
Haotian Wang 0009, Meng Yang 0002, Ce Zhu, Nanning Zheng 0001
IEEE Trans. Image Process.4
2023 LiVLR: A Lightweight Visual-Linguistic Reasoning Framework for Video Question Answering
abstract
Video Question Answering (VideoQA), aiming to correctly answer a given question based on understanding multimodal video content, is challenging due to the richness of the video content. From the perspective of video understanding, a complete VideoQA framework needs to understand the video content at different semantic levels and flexibly integrate diverse video content to distill question-related content. To this end, we propose a Lightweight Visual-Linguistic Reasoning framework named$\text{LiVLR}$. Specifically,$\text{LiVLR}$first utilizes graph-based visual and linguistic encoders to obtain multi-grained visual and linguistic representations, respectively. Subsequently, the obtained representations are integrated with the devised Diversity-aware Visual-Linguistic Reasoning module ($\text{DaVL}$).$\text{DaVL}$distinguishes different types of representations with the learnable index embedding in graph embedding. Therefore,$\text{DaVL}$can flexibly adjust the importance of different representations when generating the question-related joint representation. The proposed$\text{LiVLR}$is lightweight and shows its performance advantage on three VideoQA benchmarks, MRSVTT-QA, KnowIT VQA, and TVQA. Extensive ablation studies demonstrate the effectiveness of the key components of$\text{LiVLR}$.
Jingjing Jiang, Ziyi Liu 0001, Nanning Zheng 0001
IEEE Trans. Multim.3
2023 Multi-Panda Tracking
abstract
Multi-Panda Tracking (MPT) is a video-based tracking task for panda individuals, which is conducive to the observation and measurement of distribution and status of pandas. Different from tracking general objects such as pedestrians and vehicles, MPT is extremely challenging due to the indistinguishable appearances and diversified postures of pandas. In this case, existing tracking methods cannot appropriately tackle with the excessive occlusion between different panda individuals, hence suffering from identity switch, missing and inaccurate detections. To address these problems, we propose a simple yet effective MPT framework in the tracking-by-detection paradigm, which is benefited both from a short-term prediction filtering module and a discriminative feature learning network. In particular, the short-term prediction filtering module introduces similarity learning to enhance the temporal consistency among detections, which is capable of supplementing the missing detections and discarding false positive detections. Besides, the discriminative feature learning network leverages a two-branch network to learn both local and global discriminative features, so as to distinguish different panda individuals with a very similar appearance with a subtle difference. To evaluate the proposed method, we annotate a large-scale MPT dataset, named PANDA2021, which is particularly challenging due to the similar appearance and dramatic occlusion between panda individuals. Experiments on PANDA2021 demonstrate that the proposed MPT method significantly outperforms the competing methods. Moreover, experimental results on pedestrian tracking dataset MOT16 further demonstrate that the proposed MPT method achieves comparative performance with competing methods.
Le Wang 0003, Sanping Zhou, Nanning Zheng 0001
IEEE Trans. Multim.4
2023 Adaptive Ladder Loss for Learning Coherent Visual-Semantic Embedding
abstract
For visual-semantic embedding, the existing methods normally treat the relevance between queries and candidates in a bipolar way – relevant or irrelevant, and all “irrelevant” candidates are uniformly pushed away from the query by an equal margin in the embedding space, regardless of their various proximity to the query. This practice disregards relatively discriminative information and could lead to suboptimal ranking in the retrieval results and poorer user experience, especially in the long-tail query scenario where a matching candidate may not necessarily exist. In this paper, we introduce a continuous variable to model the relevance degree between queries and multiple candidates, and propose to learn a coherent embedding space, where candidates with higher relevance degrees are mapped closer to the query than those with lower relevance degrees. In particular, the new ladder loss is proposed by extending the triplet loss inequality to a more general inequality chain, which implements variable push-away margins according to respective relevance degrees. To adapt to the varying mini-batch statistics and improve the efficiency of the ladder loss, we also propose a Silhouette score-based method to adaptively decide the ladder level and hence the underlying inequality chain. In addition, a proper Coherent Score metric is proposed to better measure the ranking results including those “irrelevant” candidates. Extensive experiments on multiple datasets validate the efficacy of our proposed method, which achieves significant improvement over existing state-of-the-art methods.
Le Wang 0003, Zhenxing Niu, Qilin Zhang 0004, Nanning Zheng 0001
IEEE Trans. Multim.5
2023 Hierarchical Model Compression via Shape-Edge Representation of Feature Maps - an Enlightenment From the Primate Visual System
abstract
The cumbersome computation of deep neural networks (DNNs) limits their practical deployment on resource-constrained mobile multimedia devices. To deploy DNNs on devices with limited computing resources, model compression techniques are leveraged to accelerate the networks, where network pruning can improve the inference efficiency of DNNs by removing redundant weights and structures. As one of the important components of DNNs, the feature maps (FMs) can be leveraged to evaluate the importance of network structures for DNN pruning. However, previous methods neglect to fully explore the characteristics of FMs in network pruning. In this paper, we investigate the high capacity and resource efficient analogy-ventral dual-pathway primates visual system (PVS) to propose a hierarchical pruning framework (dubbed as HPSE). In an efficient PVS, the analog pathway analyzes low-frequency information to facilitate the high-frequency information inference in ventral stream. In HPSE, we extract the low-frequency shape information and high-frequency edge information from FMs to present a novel pruning pipeline that resembles the analysis mechanism of PVS. In particular, we first imitate the analogy pathway to group different FMs in each layer by calculating the shape-feature overlap. Secondly, we leverage the edge information modulated by the grouping results of the first step to prune the network. The effectiveness of HPSE is verified by pruning various DNNs on different benchmarks. For example, for ResNet-56 on CIFAR-10, HPSE reduces 52.9% of FLOPs with a slight accuracy improvement; for ResNet-50 on ImageNet, we achieve 54.3%-FLOPs drop with only 0.49% Top-1 accuracy loss.
Haonan Zhang 0002, Longjun Liu, Bingyao Kang, Nanning Zheng 0001
IEEE Trans. Multim.4
2023 P$^{2}$-GAN: Efficient Stroke Style Transfer Using Single Style Image
abstract
Style transfer is a useful image synthesis technique that can re-render given image into another artistic style while preserving its content information. Generative Adversarial Network (GAN) is a widely adopted framework toward this task for its better representation ability on local style patterns than the traditional Gram-matrix based methods. However, most previous methods rely on sufficient amount of pre-collected style images to train the model. In this paper, a novel Patch Permutation GAN (P$^{2}$-GAN) network that can efficiently learn the stroke style from a single style image is proposed. We use patch permutation to generate multiple training samples from the given style image. A patch discriminator that can simultaneously process patch-wise images and natural images seamlessly is designed. We also propose a local texture descriptor based criterion to quantitatively evaluate the style transfer quality. Experimental results showed that our method can produce finer quality re-renderings from single style image with improved computational efficiency compared with many state-of-the-arts methods.
Zhentan Zheng, Nanning Zheng 0001
IEEE Trans. Multim.3
2023 ECCA: Efficient Correntropy-Based Clustering Algorithm With Orthogonal Concept Factorization
abstract
One of the hottest topics in unsupervised learning is how to efficiently and effectively cluster large amounts of unlabeled data. To address this issue, we propose an orthogonal conceptual factorization (OCF) model to increase clustering effectiveness by restricting the degree of freedom of matrix factorization. In addition, for the OCF model, a fast optimization algorithm containing only a few low-dimensional matrix operations is given to improve clustering efficiency, as opposed to the traditional CF optimization algorithm, which involves dense matrix multiplications. To further improve the clustering efficiency while suppressing the influence of the noises and outliers distributed in real-world data, an efficient correntropy-based clustering algorithm (ECCA) is proposed in this article. Compared with OCF, an anchor graph is constructed and then OCF is performed on the anchor graph instead of directly performing OCF on the original data, which can not only further improve the clustering efficiency but also inherit the advantages of the high performance of spectral clustering. In particular, the introduction of the anchor graph makes ECCA less sensitive to changes in data dimensions and still maintains high efficiency at higher data dimensions. Meanwhile, for various complex noises and outliers in real-world data, correntropy is introduced into ECCA to measure the similarity between the matrix before and after decomposition, which can greatly improve the clustering effectiveness and robustness. Subsequently, a novel and efficient half-quadratic optimization algorithm was proposed to quickly optimize the ECCA model. Finally, extensive experiments on different real-world datasets and noisy datasets show that ECCA can archive promising effectiveness and robustness while achieving tens to thousands of times the efficiency compared with other state-of-the-art baselines.
Ben Yang, Xuetao Zhang 0001, Feiping Nie 0001, Badong Chen, Fei Wang 0008, Zhixiong Nan, Nanning Zheng 0001
IEEE Trans. Neural Networks Learn. Syst.7
2023 A Comprehensive Performance Model of Sparse Matrix-Vector Multiplication to Guide Kernel Optimization
abstract
Sparse Matrix-Vector Multiplication (SpMV) is important in scientific and industrial applications and remains a well-known challenge for modern CPUs due to high sparsity and irregularity. Many researchers try to improve SpMV performance by designing dedicated data formats and computation patterns. However, out-of-order superscalar CPUs have complex micro-architectures where exist complicated interactions and restrictions among software and hardware factors. It is hard to systematically study the effectiveness of optimization methods on the overall performance, as its benefits may be undermined by other factors. In this paper, we thoroughly study the execution of SpMV on modern CPUs and propose a comprehensive performance model to reveal the critical factors and their relationships. Specifically, we first study the coding characteristics of SpMV kernels to identify key factors worthy of attention. Then we model the execution of SpMV as two overlapped parts: CPU pipeline and memory latency. Both are carefully modeled with related hardware and software factors. We also model SIMD performance with the usage of specific SIMD instructions and vector registers. Experiments show that our model matches the actual execution of real-world processors. Guided by the model, we propose SpV8, a novel SpMV kernel that optimizes critical factors to improve computation efficiency and memory bandwidth. Experiments on Intel/AMD x86 and ARM AArch64 platforms show that SpV8 outperforms several state-of-the-art approaches with large margins, achieving average$3.4\times$over Intel Math Kernel Library and$1.4\times$over the best existing approach. Such results indicate that the proposed model is capable of valuable guidance for efficient SpMV optimizations.
Tian Xia 0008, Gelin Fu, Zhongpei Luo, Lucheng Zhang, Wenzhe Zhao 0001, Nanning Zheng 0001, Pengju Ren
IEEE Trans. Parallel Distributed Syst.8
2023 Milestones in Autonomous Driving and Intelligent Vehicles - Part I: Control, Computing System Design, Communication, HD Map, Testing, and Human Behaviors
abstract
Interest in autonomous driving (AD) and intelligent vehicles (IVs) is growing at a rapid pace due to the convenience, safety, and economic benefits. Although a number of surveys have reviewed research achievements in this field, they are still limited in specific tasks and lack systematic summaries and research directions in the future. Our work is divided into three independent articles and the first part is a survey of surveys (SoS) for total technologies of AD and IVs that involves the history, summarizes the milestones, and provides the perspectives, ethics, and future research directions. This is the second part (Part I for this technical survey) to review the development of control, computing system design, communication, high-definition map (HD map), testing, and human behaviors in IVs. In addition, the third part (Part II for this technical survey) is to review the perception and planning sections. The objective of this article is to involve all the sections of AD, summarize the latest technical milestones, and guide abecedarians to quickly understand the development of AD and IVs. Combining the SoS and Part II, we anticipate that this work will bring novel and diverse insights to researchers and abecedarians, and serve as a bridge between past and future.
Long Chen 0005, Yuchen Li 0004, Chao Huang 0006, Yang Xing 0002, Daxin Tian, Li Li 0013, Zhongxu Hu, Siyu Teng, Chen Lv 0001, Jinjun Wang, Dongpu Cao, Nanning Zheng 0001, Fei-Yue Wang 0001
IEEE Trans. Syst. Man Cybern. Syst.12
2023 Milestones in Autonomous Driving and Intelligent Vehicles - Part II: Perception and Planning
abstract
A growing interest in autonomous driving (AD) and intelligent vehicles (IVs) is fueled by their promise for enhanced safety, efficiency, and economic benefits. While previous surveys have captured progress in this field, a comprehensive and forward-looking summary is needed. Our work fills this gap through three distinct articles. The first part, a “survey of surveys” (SoS), outlines the history, surveys, ethics, and future directions of AD and IV technologies. The second part, “Milestones in AD and IVs Part I: Control, Computing System Design, Communication, high-definition map (HD map), Testing, and Human Behaviors” delves into the development of control, computing system, communication, HD map, testing, and human behaviors in IVs. This part, the third part, reviews perception and planning in the context of IVs. Aiming to provide a comprehensive overview of the latest advancements in AD and IVs, this work caters to both newcomers and seasoned researchers. By integrating the SoS and Part I, we offer unique insights and strive to serve as a bridge between past achievements and future possibilities in this dynamic field.
Long Chen 0005, Siyu Teng, Bai Li 0002, Xiaoxiang Na, Yuchen Li 0004, Jinjun Wang, Dongpu Cao, Nanning Zheng 0001, Fei-Yue Wang 0001
IEEE Trans. Syst. Man Cybern. Syst.9
2023 ACBN: Approximate Calculated Batch Normalization for Efficient DNN On-Device Training Processor
abstract
Batch normalization (BN) has been established as a very effective component in deep learning, largely helping accelerate the convergence of deep neural network (DNN) training. Nevertheless, its hardware architecture has not received much attention in the field of DNN on-device training processors. Several previous designs incur either high off-chip memory traffic or high circuit complexity, and hence have deficiencies in terms of hardware efficiency and performance. This article proposes approximately calculated BN (ACBN) to achieve a much better tradeoff between hardware efficiency and performance for DNN on-device training processors. The accuracy and convergence rate of the proposed ACBN have been extensively evaluated using four typical DNN models. Compared with the state-of-the-art reference design, the hardware simulation results show the proposed ACBN can at least reduce floating point operations by 22.2% and save external memory access by 33.3% on average. Moreover, the proposed ACBN introduces 63.6% data sparsity for the backward propagation of BN layers of VGG16 on average. To the best of our knowledge, we are the first to introduce data sparsity for the backward propagation of BN layers. The ACBN module is implemented on Zynq UltraScale+ ZCU102 system-on-chip (SoC) field-programmable gate array (FPGA), and the results show that the implementation of ACBN hardware module saves 33.9% look-up table (LUT), 49.4% flip-flop (FF), 75% digital signal processor (DSP), and reduces the power by 12.4% compared with the reference design while achieving better performance.
Baoting Li, Fujie Luo, Xuchong Zhang, Hongbin Sun 0001, Nanning Zheng 0001
IEEE Trans. Very Large Scale Integr. Syst.6
2023 HIPU: A Hybrid Intelligent Processing Unit With Fine-Grained ISA for Real-Time Deep Neural Network Inference Applications
abstract
Neural network algorithms have shown superior performance over conventional algorithms, leading to the designation and deployment of dedicated accelerators in practical scenarios. Coarse-grained accelerators achieve high performance but can support only a limited number of predesigned operators, which cannot cover the flexible operators emerging in modern neural network algorithms. Therefore, fine-grained accelerators, such as instruction set architecture (ISA)-based accelerators, have become a hot research topic due to their sufficient flexibility to cover the unpredefined operators. The main challenges for fine-grained accelerators include the undesired long delays of single-image inference when performing multibatch inference, as well as the difficulty of meeting real-time constraints when processing multiple tasks simultaneously. This article proposes a hybrid intelligent processing unit (HIPU) to address the aforementioned problems. Specifically, we design a novel conversion-free data format, expanding the single-instruction multiple-data (SIMD) instruction set and optimizing the microarchitecture design to improve the performance. We also arrange the inference schedule to guarantee scalability on multicores. The experimental results show that the proposed accelerator maintains high multiply–accumulation (MAC) utilization for all common operators and achieves high performance with 4–$7\times $speedup against NVIDIA RTX2080Ti GPU. Finally, the proposed accelerator is manufactured using TSMC 28-nm technology, achieving 1 GHz for each core, with a peak performance of 13 TOPS.
Wenzhe Zhao 0001, Guoming Yang, Tian Xia 0008, Nanning Zheng 0001, Pengju Ren
IEEE Trans. Very Large Scale Integr. Syst.5
2022 Construct Effective Geometry Aware Feature Pyramid Network for Multi-Scale Object Detection
abstract
Feature Pyramid Network (FPN) has been widely adopted to exploit multi-scale features for scale variation in object detection. However, intrinsic defects in most of the current methods with FPN make it difficult to adapt to the feature of different geometric objects. To address this issue, we introduce geometric prior into FPN to obtain more discriminative features. In this paper, we propose Geometry-aware Feature Pyramid Network (GaFPN), which mainly consists of the novel Geometry-aware Mapping Module and Geometry-aware Predictor Head.The Geometry-aware Mapping Module is proposed to make full use of all pyramid features to obtain better proposal features by the weight-generation subnetwork. The weights generation subnetwork generates fusion weight for each layer proposal features by using the geometric information of the proposal. The Geometry-aware Predictor Head introduces geometric prior into predictor head by the embedding generation network to strengthen feature representation for classification and regression. Our GaFPN can be easily extended to other two-stage object detectors with feature pyramid and applied to instance segmentation task. The proposed GaFPN significantly improves detection performance compared to baseline detectors with ResNet-50-FPN: +1.9, +2.0, +1.7, +1.3, +0.8 points Average Precision (AP) on Faster-RCNN, Cascade R-CNN, Dynamic R-CNN, SABL, and AugFPN respectively on MS COCO dataset.
Jinpeng Dong, Songyi Zhang, Shi-tao Chen, Nanning Zheng 0001
AAAI5
2022 Social Interpretable Tree for Pedestrian Trajectory Prediction
abstract
Understanding the multiple socially-acceptable future behaviors is an essential task for many vision applications. In this paper, we propose a tree-based method, termed as Social Interpretable Tree (SIT), to address this multi-modal prediction task, where a hand-crafted tree is built depending on the prior information of observed trajectory to model multiple future trajectories. Specifically, a path in the tree from the root to leaf represents an individual possible future trajectory. SIT employs a coarse-to-fine optimization strategy, in which the tree is first built by high-order velocity to balance the complexity and coverage of the tree and then optimized greedily to encourage multimodality. Finally, a teacher-forcing refining operation is used to predict the final fine trajectory. Compared with prior methods which leverage implicit latent variables to represent possible future trajectories, the path in the tree can explicitly explain the rough moving behaviors (e.g., go straight and then turn right), and thus provides better interpretability. Despite the hand-crafted tree, the experimental results on ETH-UCY and Stanford Drone datasets demonstrate that our method is capable of matching or exceeding the performance of state-of-the-art methods. Interestingly, the experiments show that the raw built tree without training outperforms many prior deep neural network based approaches. Meanwhile, our method presents sufficient flexibility in long-term prediction and different best-of-K predictions.
Liushuai Shi, Le Wang 0003, Chengjiang Long, Sanping Zhou, Fang Zheng 0009, Nanning Zheng 0001, Gang Hua 0001
AAAI6
2022 LGD: Label-Guided Self-Distillation for Object Detection
abstract
In this paper, we propose the first self-distillation framework for general object detection, termed LGD (Label-Guided self-Distillation). Previous studies rely on a strong pretrained teacher to provide instructive knowledge that could be unavailable in real-world scenarios. Instead, we generate an instructive knowledge by inter-and-intra relation modeling among objects, requiring only student representations and regular labels. Concretely, our framework involves sparse label-appearance encoding, inter-object relation adaptation and intra-object knowledge mapping to obtain the instructive knowledge. They jointly form an implicit teacher at training phase, dynamically dependent on labels and evolving student representations. Modules in LGD are trained end-to-end with student detector and are discarded in inference. Experimentally, LGD obtains decent results on various detectors, datasets, and extensive tasks like instance segmentation. For example in MS-COCO dataset, LGD improves RetinaNet with ResNet-50 under 2x single-scale training from 36.2% to 39.0% mAP (+ 2.8%). It boosts much stronger detectors like FCOS with ResNeXt-101 DCN v2 under 2x multi-scale training from 46.1% to 47.9% (+ 1.8%). Compared with a classical teacher-based method FGFI, LGD not only performs better without requiring pretrained teacher but also reduces 51% training cost beyond inherent student learning.
Peizhen Zhang, Zijian Kang, Tong Yang 0005, Xiangyu Zhang 0005, Nanning Zheng 0001, Jian Sun 0001
AAAI5
2022 Learning Disentangled Classification and Localization Representations for Temporal Action Localization
abstract
A common approach to Temporal Action Localization (TAL) is to generate action proposals and then perform action classification and localization on them. For each proposal, existing methods universally use a shared proposal-level representation for both tasks. However, our analysis indicates that this shared representation focuses on the most discriminative frames for classification, e.g., ``take-offs" rather than ``run-ups" in distinguishing ``high jump" and ``long jump", while frames most relevant to localization, such as the start and end frames of an action, are largely ignored. In other words, such a shared representation can not simultaneously handle both classification and localization tasks well, and it makes precise TAL difficult. To address this challenge, this paper disentangles the shared representation into classification and localization representations. The disentangled classification representation focuses on the most discriminative frames, and the disentangled localization representation focuses on the action phase as well as the action start and end. Our model could be divided into two sub-networks, i.e., the disentanglement network and the context-based aggregation network. The disentanglement network is an autoencoder to learn orthogonal hidden variables of classification and localization. The context-based aggregation network aggregates the classification and localization representations by modeling local and global contexts. We evaluate our proposed method on two popular benchmarks for TAL, which outperforms all state-of-the-art methods.
Zixin Zhu, Le Wang 0003, Wei Tang 0016, Ziyi Liu 0001, Nanning Zheng 0001, Gang Hua 0001
AAAI5
2022 TransVPR: Transformer-Based Place Recognition with Multi-Level Attention Aggregation
abstract
Visual place recognition is a challenging task for applications such as autonomous driving navigation and mobile robot localization. Distracting elements presenting in complex scenes often lead to deviations in the perception of visual place. To address this problem, it is crucial to integrate information from only task-relevant regions into image representations. In this paper, we introduce a novel holistic place recognition model, TransVPR, based on vision Transformers. It benefits from the desirable property of the self-attention operation in Transformers which can naturally aggregate task-relevant features. Attentions from multiple levels of the Transformer, which focus on different regions of interest, are further combined to generate a global image representation. In addition, the output tokens from Transformer layers filtered by the fused attention mask are considered as key-patch descriptors, which are used to perform spatial matching to re-rank the candidates retrieved by the global image features. The whole model allows end-to-end training with a single objective and image-level supervision. TransVPR achieves state-of-the-art performance on several real-world benchmarks while maintaining low computational time and storage requirements.
Ruotong Wang 0005, Yanqing Shen, Weiliang Zuo, Sanping Zhou, Nanning Zheng 0001
CVPR5
2022 Learning to Refactor Action and Co-occurrence Features for Temporal Action Localization
abstract
The main challenge of Temporal Action Localization is to retrieve subtle human actions from various co-occurring ingredients, e.g., context and background, in an untrimmed video. While prior approaches have achieved substantial progress through devising advanced action detectors, they still suffer from these co-occurring ingredients which often dominate the actual action content in videos. In this paper, we explore two orthogonal but complementary aspects of a video snippet, i.e., the action features and the co-occurrence features. Especially, we develop a novel auxiliary task by decoupling these two types of features within a video snippet and recombining them to generate a new feature representation with more salient action information for accurate action localization. We term our method RefactorNet, which first explicitly factorizes the action content and regularizes its co-occurrence features, and then synthesizes a new action-dominated video representation. Extensive experimental results and ablation studies on THUMOS14 and ActivityNet v 1.3 demonstrate that our new representation, combined with a simple action detector, can significantly improve the action localization performance.
Le Wang 0003, Sanping Zhou, Nanning Zheng 0001, Wei Tang 0016
CVPR4
2022 Asymmetric Relation Consistency Reasoning for Video Relation Grounding
Ping Wei 0001, Jiapeng Li 0003, Zeyu Ma 0005, Jiahui Shang, Nanning Zheng 0001
ECCV (35)6
2022 Towards Building A Group-based Unsupervised Representation Disentanglement Framework
Tao Yang 0032, Xuanchi Ren, Yuwang Wang, Wenjun Zeng 0001, Nanning Zheng 0001
ICLR5
2022 Offline Signature Verification with Transformers
abstract
Signature verification is a frequently-used forensics technology. Although the previous convolution neural network (CNN) based methods have made a great progress, the limitation of local neighborhood operation of CNN impedes reasoning about the relation of global signature strokes. To overcome this weakness, in this paper, we propose a novel holistic-part unified model named TransOSV based on the transformer framework. Signature images are encoded into patch sequences by the proposed holistic encoder to learn global representation. Considering the subtle local difference between the genuine signature and forged signature, we design a contrast based part decoder that is utilized to learn discriminative part features. To reduce the influence of sample imbalance, we formulate a new focal contrast loss function. Extensive experimental results and ablation studies prove the potential of the proposed model.
Ping Wei 0001, Zeyu Ma 0005, Changkai Li, Nanning Zheng 0001
ICME5
2022 HOIG: End-to-End Human-Object Interactions Grounding with Transformers
abstract
Visual grounding is a crucial and challenging problem in many applications. While it has been extensively investigated over the past years, human-centric grounding with multiple instances is still an open problem. In this paper, we introduce a new task of Human-Object Interactions (HOI) Grounding to localize all the referring human-object pair instances in an image with a given ⟨human, interaction, object⟩ phrase. We design an encoder-decoder architecture to model the task as a set prediction problem based on transformers. A vision-language alignment module and a grounding decoder are designed to learn accurate cross-modal contexts and interactions. Our model accomplishes alignment and prediction in an end-to-end manner without pre-trained detectors or post-processing. Experiments on two challenging datasets prove the strength of our model.
Zeyu Ma 0005, Ping Wei 0001, Nanning Zheng 0001
ICME4
2022 Relation Reasoning for Video Pedestrian Trajectory Prediction
abstract
Pedestrian trajectory prediction is a challenging and important task in many applications, which aims to predict future pedestrians' trajectory coordinates from the input historical data. The existing methods usually use ready-made trajectory coordinates as inputs, which is, however, unavailable in video-based scenarios. In this paper, we propose a relation reasoning hypergraph (RRH) model to directly predict multiple pedestrian trajectories from raw videos. It is a challenging issue for the input and output are in different modalities and a video may contain multiple pedestrians. Our model integrates historical trajectory tracking, pedestrian relation reasoning, and future trajectory prediction into one framework. For capturing the subtle social relationships among pedestrians, we design a relation reasoning hypergraph network. We tested the proposed method on two public pedestrians datasets and the performance demonstrates the power of the model.
Haowen Tang, Ping Wei 0001, Jiapeng Li 0003, Nanning Zheng 0001
ICME5
2022 SE(3) Equivariant Graph Neural Networks with Complete Local Frames
abstract
Group equivariance (e.g. SE(3) equivariance) is a critical physical symmetry in science, from classical and quantum physics to computational biology. It enables robust and accurate prediction under arbitrary reference transformations. In light of this, great efforts have been put on encoding this symmetry into deep neural networks, which has been shown to improve the generalization performance and data efficiency for downstream tasks. Constructing an equivariant neural network generally brings high computational costs to ensure expressiveness. Therefore, how to better trade-off the expressiveness and computational efficiency plays a core role in the design of the equivariant deep learning models. In this paper, we propose a framework to construct SE(3) equivariant graph neural networks that can approximate the geometric quantities efficiently. Inspired by differential geometry and physics, we introduce equivariant local complete frames to graph neural networks, such that tensor information at given orders can be projected onto the frames. The local frame is constructed to form an orthonormal basis that avoids direction degeneration and ensure completeness. Since the frames are built only by cross product operations, our method is computationally efficient. We evaluate our method on two tasks: Newton mechanics modeling and equilibrium molecule conformation generation. Extensive experimental results demonstrate that our model achieves the best or competitive performance in two types of datasets.
Weitao Du, Yuanqi Du, Wei Chen 0034, Nanning Zheng 0001, Bin Shao 0002, Tie-Yan Liu
ICML6
2022 Greedy based Value Representation for Optimal Coordination in Multi-agent Reinforcement Learning
abstract
Due to the representation limitation of the joint Q value function, multi-agent reinforcement learning methods with linear value decomposition (LVD) or monotonic value decomposition (MVD) suffer from relative overgeneralization. As a result, they can not ensure optimal consistency (i.e., the correspondence between individual greedy actions and the best team performance). In this paper, we derive the expression of the joint Q value function of LVD and MVD. According to the expression, we draw a transition diagram, where each self-transition node (STN) is a possible convergence. To ensure the optimal consistency, the optimal node is required to be the unique STN. Therefore, we propose the greedy-based value representation (GVR), which turns the optimal node into an STN via inferior target shaping and eliminates the non-optimal STNs via superior experience replay. Theoretical proofs and empirical results demonstrate that given the true Q values, GVR ensures the optimal consistency under sufficient exploration. Besides, in tasks where the true Q values are unavailable, GVR achieves an adaptive trade-off between optimality and stability. Our method outperforms state-of-the-art baselines in experiments on various benchmarks.
Lipeng Wan 0003, Zeyang Liu 0001, Xingyu Chen 0001, Xuguang Lan, Nanning Zheng 0001
ICML5
2022 Parametric Path Optimization for Wheeled Robots Navigation
abstract
Collision risk and smoothness are the most important factors in global path planning. Currently, planning methods that reduce global path collision risk and improve its smoothness through numerical optimization have achieved good results. However, these methods cannot always optimize the path. The reason is all points on the path are considered as decision variables, which leads to the high dimensionality of the defined optimization problem. Therefore, we propose a novel global path optimization method. The method characterizes the path as a parametric curve and then optimizes the curve's parameters with a defined objective function, which successfully reduces the dimension of optimization problem. The proposed method is compared with baseline and state-of-the-art methods. Experimental results show the path optimized by our method is not only optimal in collision risk, but also in efficiency and smoothness. Furthermore, the proposed method is also implemented and tested in both simulation and real robots.
Zhiqiang Jian, Songyi Zhang, Shi-tao Chen, Nanning Zheng 0001
ICRA5
2022 A Continuous Learning Approach for Probabilistic Human Motion Prediction
abstract
Human Motion Prediction (HMP) plays a crucial role in safe Human-Robot-Interaction (HRI). Currently, the majority of HMP algorithms are trained by massive pre-collected data. As the training data only contains a few pre-defined motion patterns, these methods cannot handle the unfamiliar motion patterns. Moreover, the pre-collected data are usually non-interactive, which does not consider the real-time responses of collaborators. As a result, these methods usually perform unsatisfactorily in real HRI scenarios. To solve this problem, in this paper, we propose a novel Continual Learning (CL) approach for probabilistic HMP which makes the robot continually learns during its interaction with collaborators. The proposed approach consists of two steps. First, we leverage a Bayesian Neural Network to model diverse uncertainties of observed human motions for collecting online interactive data safely. Then we take Experience Replay and Knowledge Distillation to elevate the model with new experiences while maintaining the knowledge learned before. We first evaluate our approach on a large-scale benchmark dataset Human3.6m. The experimental results show that our approach achieves a lower prediction error compared with the baselines methods. Moreover, our approach could continually learn new motion patterns without forgetting the learned knowledge. We further conduct real-scene experiments using Kinect DK. The results show that our approach can learn the human kinematic model from scratch, which effectively secures the interaction.
Shihong Wang, Xingyu Chen 0001, Xuguang Lan, Nanning Zheng 0001
ICRA6
2022 Foreground-attention in neural decoding: Guiding Loop-Enc-Dec to reconstruct visual stimulus images from fMRI
abstract
The reconstruction of visual stimulus images from functional Magnetic Resonance Imaging (fMRI) has received extensive attention in recent years, which provides a possibility to interpret the human brain. Due to the high-dimensional and high-noise characteristics of fMRI data, how to extract stable, reliable and useful information from fMRI data for image reconstruction has become a challenging problem. Inspired by the mechanism of human visual attention, in this paper, we propose a novel method of reconstructing visual stimulus images, which first decodes human visual salient region from fMRI, we define human visual salient region as foreground attention (F-attention), and then reconstructs the visual images guided by F-attention. Because the human brain is strongly wound into sulci and gyri, some spatially adjacent voxels are not connected in practice. Therefore, it is necessary to consider the global information when decoding fMRI, so we introduce the self-attention module for capturing global information into the process of decoding F-attention. In addition, in order to obtain more loss constraints in the training process of encoder-decoder, we also propose a new training strategy called Loop-Enc-Dec. The experimental results show that the F-attention decoder decodes the visual attention from fMRI successfully, and the Loop-Enc-Dec guided by F-attention can also well reconstruct the visual stimulus images.
Mingyang Sheng, Nanning Zheng 0001
IJCNN4
2022 CMB: A Novel Structural Re-parameterization Block without Extra Training Parameters
abstract
Structural re-parameterization is a raising field, which aims at improving the performance of convolutional neural networks (CNNs) through training an over-parameterization model and transferring it into a compact inference model. However, the performance improvements of prior structural re-parameterization works often come at the cost of heavy extra training resources, which increases carbon emissions and limits the potential applications on large-scale industrial tasks. To this end, first, we conduct experiments with a series of blocks composed of multiple identical branches to investigate the mechanism behind the structural re-parameterization, and then provide an interpretation. Moreover, motivated by the studies of effective receptive fields in the biological visual systems and neural networks, we propose a novel compact block named circular mask block (CMB). Given a neural network, we replace the regular convolutional layer with CMB to construct a training architecture, which can be trained to gain an accuracy boost with No extra training parameters and limited extra training FLOPs. After training, the training architecture can be transformed into the original architecture for inference. Extensive experiments are performed on CIFAR-10 and ImageNet to evaluate the effectiveness of our method. For example, we improve 0.85% top-1 accuracy of ResNet-50 on ImageNet without extra training parameters and only 11.32M extra training FLOPs, which saves 434x training FLOPs compared with prior works.
Hengyi Zhou, Longjun Liu, Haonan Zhang 0002, Hongyi He, Nanning Zheng 0001
IJCNN5
2022 Pedestrian Intention Prediction Based on Traffic-Aware Scene Graph Model
abstract
Anticipating the future behavior of pedestrians is a crucial part of deploying Automated Driving Systems (ADS) in urban traffic scenarios. Most recent works utilize a convolutional neural network (CNN) to extract visual information, which is then input to a recurrent neural network (RNN) along with pedestrian-specific features like location and speed to obtain temporal features. However, the majority of these approaches lack the ability to parse the relationships of the related objects in the specific traffic scene, which leads to omitting the interactions between the pedestrians and the interactions between the pedestrians and the traffic. For this purpose, we propose a graph-structured model which can dig out pedestrians' dynamic constraints by constructing a traffic-aware scene graph within each frame. In addition, to capture pedestrian movement more effectively, we also introduce a temporal feature representation model, which first uses inter-frame and intra-frame GRU (II-GRU) to mine inter-frame information and intra-frame information together, and then employs a novel attention mechanism to adaptively generate attention weights. Extensive experiments on the JAAD and PIE datasets prove that our proposed model is effective in reaching and enhancing the state-of-the-art performance.
Xingchen Song, Miao Kang, Sanping Zhou, Jianji Wang 0001, Yishu Mao 0003, Nanning Zheng 0001
IROS6
2022 Rethinking the Mechanism of the Pattern Pruning and the Circle Importance Hypothesis
abstract
Network pruning is an effective and widely-used model compression technique. Pattern pruning is a new sparsity dimension pruning approach whose compression ability has been proven in some prior works. However, a detailed study on "pattern" and pattern pruning is still lacking. In this paper, we analyze the mechanism behind pattern pruning. Our analysis reveals that the effectiveness of pattern pruning should be attributed to finding the less important weights even before training. Then, motivated by the fact that the retinal ganglion cells in the biological visual system have approximately concentric receptive fields, we further investigate and propose the Circle Importance Hypothesis to guide the design of efficient patterns. We also design two series of special efficient patterns - circle patterns and semicircle patterns. Moreover, inspired by the neural architecture search technique, we propose a novel one-shot gradient-based pattern pruning algorithm. Besides, we also expand depthwise convolutions with our circle patterns, which improves the accuracy of networks with little extra memory cost. Extensive experiments are performed to validate our hypotheses and the effectiveness of the proposed methods. For example, we reduce the 44.0% FLOPS of ResNet-56 while improving its accuracy to 94.38% on CIFAR-10. And we reduce the 41.0% FLOPS of ResNet-18 with only a 1.11% accuracy drop on ImageNet.
Hengyi Zhou, Longjun Liu, Haonan Zhang 0002, Nanning Zheng 0001
ACM Multimedia4
2022 Could Giant Pre-trained Image Models Extract Universal Representations?
abstract
Frozen pretrained models have become a viable alternative to the pretraining-then-finetuning paradigm for transfer learning. However, with frozen models there are relatively few parameters available for adapting to downstream tasks, which is problematic in computer vision where tasks vary significantly in input/output format and the type of information that is of value. In this paper, we present a study of frozen pretrained models when applied to diverse and representative computer vision tasks, including object detection, semantic segmentation and video action recognition. From this empirical analysis, our work answers the questions of what pretraining task fits best with this frozen setting, how to make the frozen setting more flexible to various downstream tasks, and the effect of larger model sizes. We additionally examine the upper bound of performance using a giant frozen pretrained model with 3 billion parameters (SwinV2-G) and find that it reaches competitive performance on a varied set of major benchmarks with only one shared frozen base network: 60.0 box mAP and 52.2 mask mAP on COCO object detection test-dev, 57.6 val mIoU on ADE20K semantic segmentation, and 81.7 top-1 accuracy on Kinetics-400 action recognition. With this work, we hope to bring greater attention to this promising path of freezing pretrained image models.
Yutong Lin, Zheng Zhang 0022, Han Hu 0001, Nanning Zheng 0001, Stephen Lin 0001, Yue Cao 0001
NeurIPS5
2022 Visual Concepts Tokenization
abstract
Obtaining the human-like perception ability of abstracting visual concepts from concrete pixels has always been a fundamental and important target in machine learning research fields such as disentangled representation learning and scene decomposition. Towards this goal, we propose an unsupervised transformer-based Visual Concepts Tokenization framework, dubbed VCT, to perceive an image into a set of disentangled visual concept tokens, with each concept token responding to one type of independent visual concept. Particularly, to obtain these concept tokens, we only use cross-attention to extract visual information from the image tokens layer by layer without self-attention between concept tokens, preventing information leakage across concept tokens. We further propose a Concept Disentangling Loss to facilitate that different concept tokens represent independent visual concepts. The cross-attention and disentangling loss play the role of induction and mutual exclusion for the concept tokens, respectively. Extensive experiments on several popular datasets verify the effectiveness of VCT on the tasks of disentangled representation learning and scene decomposition. VCT achieves the state of the art results by a large margin.
Tao Yang 0032, Yuwang Wang, Yan Lu 0001, Nanning Zheng 0001
NeurIPS4
2022 Improved drug-target interaction prediction with intermolecular graph transformer
abstract
The identification of active binding drugs for target proteins (referred to as drug-target interaction prediction) is the key challenge in virtual screening, which plays an essential role in drug discovery. Although recent deep learning-based approaches achieve better performance than molecular docking, existing models often neglect topological or spatial of intermolecular information, hindering prediction performance. We recognize this problem and propose a novel approach called the Intermolecular Graph Transformer (IGT) that employs a dedicated attention mechanism to model intermolecular information with a three-way Transformer-based architecture. IGT outperforms state-of-the-art (SoTA) approaches by 9.1% and 20.5% over the second best option for binding activity and binding pose prediction, respectively, and exhibits superior generalization ability to unseen receptor proteins than SoTA approaches. Furthermore, IGT exhibits promising drug screening ability against severe acute respiratory syndrome coronavirus 2 by identifying 83.1% active drugs that have been validated by wet-lab experiments with near-native predicted binding poses. Source code and datasets are available at https://github.com/microsoft/IGT-Intermolecular-Graph-Transformer.
Siyuan Liu 0005, Liang He 0010, Bin Shao 0002, Jian Yin 0001, Nanning Zheng 0001, Tie-Yan Liu, Tong Wang 0014
Briefings Bioinform.7
2022 DFSNet: Dividing-fuse deep neural networks with searching strategy for distributed DNN architecture
Wenxuan Hou 0001, Longjun Liu, Haonan Zhang 0002, Hongbin Sun 0001, Nanning Zheng 0001
Neurocomputing5
2022 Learning to predict diverse trajectory from human motion patterns
Miao Kang, Jingwen Fu, Sanping Zhou, Songyi Zhang, Nanning Zheng 0001
Neurocomputing5
2022 EvoSTGAT: Evolving spatiotemporal graph attention networks for pedestrian trajectory prediction
Haowen Tang, Ping Wei 0001, Jiapeng Li 0003, Nanning Zheng 0001
Neurocomputing4
2022 End-to-end learning of self-rectification and self-supervised disparity prediction for stereo vision
Xuchong Zhang, Han Zhai, Hongbin Sun 0001, Nanning Zheng 0001
Neurocomputing6
2022 CMD: controllable matrix decomposition with global optimization for deep neural network compression
Haonan Zhang 0002, Longjun Liu, Hengyi Zhou, Hongbin Sun 0001, Nanning Zheng 0001
Mach. Learn.5
2022 Subjective low-light image enhancement based on a foreground saliency map model
Pengcheng Hao, Meng Yang 0002, Nanning Zheng 0001
Multim. Tools Appl.3
2022 Weakly Supervised Temporal Action Localization Through Contrast Based Evaluation Networks
abstract
Given only video-level action categorical labels during training, weakly-supervised temporal action localization (WS-TAL) learns to detect action instances and locates their temporal boundaries in untrimmed videos. Compared to its fully supervised counterpart, WS-TAL is more cost-effective in data labeling and thus favorable in practical applications. However, the coarse video-level supervision inevitably incurs ambiguities in action localization, especially in untrimmed videos containing multiple action instances. To overcome this challenge, we observe that significant temporal contrasts among video snippets, e.g., caused by temporal discontinuities and sudden changes, often occur around true action boundaries. This motivates us to introduce a Contrast-based Localization EvaluAtioN Network (CleanNet), whose core is a new temporal action proposal evaluator, which provides fine-grained pseudo supervision by leveraging the temporal contrasts among snippet-level classification predictions. As a result, the uncertainty in locating action instances can be resolved via evaluating their temporal contrast scores. Moreover, the new action localization module is an integral part of CleanNet which enables end-to-end training. This is in contrast to many existing WS-TAL methods where action localization is merely a post-processing step. Besides, we also explore the usage of temporal contrast on temporal action proposal (TAP) generation task, which we believe is the first attempt with the weak supervision setting. Experiments on the THUMOS14, ActivityNet v1.2 and v1.3 datasets validate the efficacy of our method against existing state-of-the-art WS-TAL algorithms.
Ziyi Liu 0001, Le Wang 0003, Qilin Zhang 0004, Wei Tang 0016, Nanning Zheng 0001, Gang Hua 0001
IEEE Trans. Pattern Anal. Mach. Intell.5
2022 Group Sampling for Scale Invariant Face Detection
abstract
Detectors based on deep learning tend to detect multi-scale objects on a single input image for efficiency. Recent works, such as FPN and SSD, generally use feature maps from multiple layers with different spatial resolutions to detect objects at different scales, e.g., high-resolution feature maps for small objects. However, we find that objects at all scales can also be well detected with features from a single layer of the network. In this paper, we carefully examine the factors affecting detection performance across a large range of scales, and conclude that the balance of training samples, including both positive and negative ones, at different scales is the key. We propose a group sampling method which divides the anchors into several groups according to the scale, and ensure that the number of samples for each group is the same during training. Our approach using only one single layer of FPN as features is able to advance the state-of-the-arts. Comprehensive analysis and extensive experiments have been conducted to show the effectiveness of the proposed method. Moreover, we show that our approach is favorably applicable to other tasks, such as object detection on COCO dataset, and to other detection pipelines, such as YOLOv3, SSD and R-FCN. Our approach, evaluated on face detection benchmarks including FDDB and WIDER FACE datasets, achieves state-of-the-art results without bells and whistles.
Xiang Ming, Fangyun Wei, Ting Zhang 0002, Dong Chen 0003, Nanning Zheng 0001, Fang Wen 0001
IEEE Trans. Pattern Anal. Mach. Intell.5
2022 Social-Aware Pedestrian Trajectory Prediction via States Refinement LSTM
abstract
In the task of pedestrian trajectory prediction, social interaction could be one of the most complicated factors since it is difficult to be interpreted through simple rules. Recent studies have shown a great ability of LSTM networks in learning social behaviors from datasets, e.g., introducing LSTM hidden states of the neighbors at the last time step into LSTM recursion. However, those methods depend on previous neighboring features which lead to a delayed observation. In this paper, we propose a data-driven states refinement LSTM network (SR-LSTM) to enable the utilization of the current intention of neighbors through a message passing framework. Moreover, the model performs in the form of self-updating by jointly refining the current states of all participants, rather than an input-output mechanism served by feature concatenation. In the process of states refinement, a social-aware information selection module consisting of an element-wise motion gate and a pedestrian-wise attention is designed to serve as the guidance of the message passing process. Considering the pedestrian walking space as a graph where each pedestrian is a node and each pedestrian pair with an edge, spatial-edge LSTMs are further exploited to enhance the model capacity, where two kinds of LSTMs interact with each other so that states of them are interactively refined. Experimental results on four widely used pedestrian trajectory datasets, ETH, UCY, PWPD, and NYGC demonstrate the effectiveness of the proposed model.
Pu Zhang 0001, Jianru Xue, Pengfei Zhang 0005, Nanning Zheng 0001, Wanli Ouyang
IEEE Trans. Pattern Anal. Mach. Intell.4
2022 Loss functions for pose guided person image generation
Haoyue Shi 0002, Le Wang 0003, Nanning Zheng 0001, Gang Hua 0001, Wei Tang 0016
Pattern Recognit.3
2022 Adaptive Disparity Candidates Prediction Network for Efficient Real-Time Stereo Matching
abstract
Efficient real-time disparity estimation is critical for the application of stereo vision systems in various areas. Recently, stereo network based on coarse-to-fine method has largely relieved the memory constraints and speed limitations of large-scale network models. Nevertheless, all of the previous coarse-to-fine designs employ constant offsets and three or more stages to progressively refine the coarse disparity map, still resulting in unsatisfactory computation accuracy and inference time when deployed on mobile devices. This paper claims that the coarse matching errors can be corrected efficiently with fewer stages as long as more accurate disparity candidates can be provided. Therefore, we propose a dynamic offset prediction module to meet different correction requirements of diverse objects and design an efficient two-stage framework. In addition, a disparity-independent convolution is proposed to regularize the compact cost volume efficiently and further improve the overall performance. The disparity quality and efficiency of various stereo networks are evaluated on multiple datasets and platforms. Evaluation results demonstrate that, the disparity error rate of the proposed network achieves 2.66% and 2.71% on KITTI 2012 and 2015 test sets respectively, where the computation speed is$2\times $faster than the state-of-the-art lightweight models on high-end and source-constrained GPUs.
He Dai, Xuchong Zhang, Hongbin Sun 0001, Nanning Zheng 0001
IEEE Trans. Circuits Syst. Video Technol.5
2022 FCHP: Exploring the Discriminative Feature and Feature Correlation of Feature Maps for Hierarchical DNN Pruning and Compression
abstract
Pruning can remove the redundant parameters and structures of Deep Neural Networks (DNNs) to reduce inference time and memory overhead. As one of the important components of DNN, feature maps (FMs) have been widely used in network pruning. However, previous approaches do not fully investigate the discriminative features in FMs, and also do not explicitly utilize all the features associated with each layer in the pruning procedure. In this paper, we explore the discriminative feature of FMs and explicitly investigate the two-adjacent-layer features of each layer to propose a three-phase hierarchical pruning framework, dubbed as FCHP. Firstly, we decompose each FM into several components to extract the discriminative feature. After that, since pruning each layer is related to the FMs of adjacent layers, we explicitly calculate the feature correlation of discriminative features of two adjacent layers, and then use the feature correlation to cluster FMs into several hierarchies to guide subsequent pruning. Finally, we compute the content of discriminative features, and remove channels corresponding to FMs with fewer discriminative features in each hierarchy, respectively. In the experiment, we prune DNNs with the multiple types of architecture on different benchmarks, and the results have achieved the state-of-the-arts in terms of compressed parameters and FLOPs drop. For example, as for ResNet-56 on CIFAR-10, FCHP respectively obtains 50% of parameters and FLOPs reduction with negligible accuracy loss. Besides, as for ResNet-50 on ImageNet, FCHP reduces 40.5% of parameters and 44.1% of FLOPs with 0.43% of Top-1 accuracy drop.
Haonan Zhang 0002, Longjun Liu, Hengyi Zhou, Liang Si, Hongbin Sun 0001, Nanning Zheng 0001
IEEE Trans. Circuits Syst. Video Technol.6
2022 Conditional Uncorrelation and Efficient Subset Selection in Sparse Regression
abstract
Given$m~d$-dimensional responsors and$n~d$-dimensional predictors, sparse regression finds at most$k$predictors for each responsor for linear approximation,$1\leq k \leq d-1$. The key problem in sparse regression is subset selection, which usually suffers from high computational cost. In recent years, many improved approximate methods of subset selection have been published. However, less attention has been paid to the nonapproximate method of subset selection, which is very necessary for many questions in data analysis. Here, we consider sparse regression from the view of correlation and propose the formula of conditional uncorrelation. Then, an efficient nonapproximate method of subset selection is proposed in which we do not need to calculate any coefficients in the regression equation for candidate predictors. By the proposed method, the computational complexity is reduced from$O([{1}/{6}]{k^{3}}\!+(m+1)k^{2}\!+\!mkd)$to$O([{1}/{6}]{k^{3}}\!+[{1}/{2}](m+1)k^{2})$for each candidate subset in sparse regression. Because the dimension$d$is generally the number of observations or experiments and large enough, the proposed method can greatly improve the efficiency of nonapproximate subset selection. We also apply the proposed method in real scenarios of dental age assessment and sparse coding to validate the efficiency of the proposed method.
Jianji Wang 0001, Qi Liu 0010, Shaoyi Du, Yu-Cheng Guo, Nanning Zheng 0001, Fei-Yue Wang 0001
IEEE Trans. Cybern.6
2022 Depth Map Recovery Based on a Unified Depth Boundary Distortion Model
abstract
Depth maps acquired by either physical sensors or learning methods are often seriously distorted due to boundary distortion problems, including missing, fake, and misaligned boundaries (compared with RGB images). An RGB-guided depth map recovery method is proposed in this paper to recover true boundaries in seriously distorted depth maps. Therefore, a unified model is first developed to observe all these kinds of distorted boundaries in depth maps. Observing distorted boundaries is equivalent to identifying erroneous regions in distorted depth maps, because depth boundaries are essentially formed by contiguous regions with different intensities. Then, erroneous regions are identified by separately extracting local structures of RGB image and depth map with Gaussian kernels and comparing their similarity on the basis of the SSIM index. A depth map recovery method is then proposed on the basis of the unified model. This method recovers true depth boundaries by iteratively identifying and correcting erroneous regions in recovered depth map based on the unified model and a weighted median filter. Because RGB image generally includes additional textural contents compared with depth maps, texture-copy artifacts problem is further addressed in the proposed method by restricting the model works around depth boundaries in each iteration. Extensive experiments are conducted on five RGB-depth datasets including depth map recovery, depth super-resolution, depth estimation enhancement, and depth completion enhancement. The results demonstrate that the proposed method considerably improves both the quantitative and visual qualities of recovered depth maps in comparison with fifteen competitive methods. Most object boundaries in recovered depth maps are corrected accurately, and kept sharply and well aligned with the ones in RGB images.
Haotian Wang 0009, Meng Yang 0002, Xuguang Lan, Ce Zhu, Nanning Zheng 0001
IEEE Trans. Image Process.5
2022 Three Principles to Determine the Right-of-Way for AVs: Safe Interaction With Humans
abstract
Autonomous vehicles (AVs) are widely believed to be good for improving transportation safety and efficiency. However, recent fatal accidents by some of their prototypes remind us that there are no operationalizable and quantitative safe driving strategies available for an AV in a wide range of situations to avoid collisions. In contrast with many recent studies that focused on ethical considerations when AVs are facing unavoidable harms, we study how to proactively prevent collisions by setting up a set of decision rules for AVs to determine the right-of-way efficiently. Notably, we summarize three essential principles for AVs designing to increase driving safety, and establish a rule-based nine-step communication-decision model to implement them. Our method is constructed by analyzing how human drivers solve potential conflicts. The decision rules are designed to be ambiguity-free and readily computable with the least communication so that human drivers and AVs could easily understand each other in terms of their behaviors and intentions of. We have demonstrated the effectiveness of our method by comparing it with some alternative approaches.
Li Li 0013, Can Zhao 0004, Xiao Wang 0002, Zhiheng Li 0001, Long Chen 0005, Nanning Zheng 0001, Fei-Yue Wang 0001
IEEE Trans. Intell. Transp. Syst.7
2022 A Flexible and Explainable Vehicle Motion Prediction and Inference Framework Combining Semi-Supervised AOG and ST-LSTM
abstract
Accurate trajectory prediction of surrounding vehicles is important for automated vehicles. To solve several existing problems of maneuver-based trajectory prediction, we propose four targeted solutions and establish a trajectory prediction model that integrates semi-supervised And-or Graph (AOG) and Spatio-temporal LSTM (ST-LSTM). To reduce the dependence on the well-labeled dataset, we introduce the concept of sub-maneuvers to improve the classifications of vehicle movements based on the given rough maneuver labels. AOG is used as the backbone of the probabilistic motion inference considering sub-maneuvers. We only define the basic units and inference logics of AOG and design a semi-supervised approach to directly learn the sub-maneuvers and the inference model structure from the training data, without manually specifying the structure (layers and nodes) of the inference model. This approach helps to avoid excessive artificial design or biases. The learned hierarchical motion inference model improves the interpretability of the overall trajectory prediction process. To utilize vehicle interaction information and further yield more accurate prediction, we adopt two different methods to consider vehicle interaction in the two sub-models (maneuver recognition and trajectory prediction). The experiment on NGSIM I-80 dataset shows that the maneuver-based model proposed in this paper (AOG-ST and refined AOG-ST-TB) performs more accurate trajectory prediction results. Although the AOG-ST seems clumsy and slow, we show that it is a flexible and quick model for trajectory prediction for various driving scenarios through the discussion and experiment.
Shengzhe Dai, Zhiheng Li 0001, Li Li 0013, Nanning Zheng 0001, Shuofeng Wang
IEEE Trans. Intell. Transp. Syst.4
2022 Multi-Model-Based Local Path Planning Methodology for Autonomous Driving: An Integrated Framework
abstract
Autonomous driving systems (ADSs) need to be able to respond quickly to changes in the dynamic traffic scenario. However, regardless of the changes occurring in traffic scenes, the current local path planning frameworks of ADSs are based on the fixed frequency re-planning path (i.e., running their planning algorithms repeatedly). This planning method makes it difficult to provide a reasonable traveling path, agility, and comfort for driverless vehicles in changing traffic scenarios. Therefore, this article performs an in-depth analysis of the problems of traditional planning frameworks which use a fixed frequency to replan the path and proposes a novel path planning framework that is universal based on multiple-models. The proposed framework divides the planning process into several layers, each of which has different functions. With this framework, the ADS can adaptively adjust the planning process according to the changes in traffic scenes and then provide different path planning algorithms to ensure its safety and flexibility in the process of driving. Moreover, the problems caused by the traditional planning framework can be solved. This framework has been applied to the autonomous vehicle “Pioneer”, which won first place in the 2019 China Intelligent Vehicle Future Challenge (IVFC). The effectiveness and rationality of the integrated framework of local path planning proposed in this article were verified by a large number of tests in real-world traffic scenarios.
Zhiqiang Jian, Shi-tao Chen, Songyi Zhang, Yu Chen 0040, Nanning Zheng 0001
IEEE Trans. Intell. Transp. Syst.5
2022 SceGene: Bio-Inspired Traffic Scenario Generation for Autonomous Driving Testing
abstract
The core value of simulation-based autonomy tests is to create densely extreme traffic scenarios to test the performance and robustness of the algorithms and systems. Test scenarios are usually designed or extracted manually from the real-world data, which is inefficient with a remarkable domain gap compared with testing in real scenarios. Therefore, it is crucial to automatically generate realistic and diverse dynamic traffic scenarios making autonomy tests efficient. Moreover, scenario generation is expected to be interpretable, controllable, and diversified, which can be hard to achieve simultaneously by methods based on rules or deep networks. In this paper, we propose a dynamic traffic scenario generation method called SceGene, inspired by genetic inheritance and mutation processes in biological intelligence. SceGene applies biological processes, such as crossover and mutation, to exchange and mutate the content of scenarios, and involves the natural selection process to control generation direction. SceGene has three main parts: 1) a new representation method for describing the traffic scenarios’ feature; 2) a new scenario generation algorithm based on crossover, mutation, and selection; and 3) an abnormal scenario information repair method based on the microscopic driving model. Evaluation on the public traffic scenario dataset shows that SceGene can ensure highly realistic and diversified scenario generation in an interpretable and controllable way, significantly improving the efficiency of the simulation-based autonomy tests.
Ao Li 0006, Shi-tao Chen, Nanning Zheng 0001, Masayoshi Tomizuka
IEEE Trans. Intell. Transp. Syst.4
2022 From Human Driving to Automated Driving: What Do We Know About Drivers?
abstract
Humanlike automated driving (AD) strategies which are inspired by drivers’ cognition ways may show advantages in dealing with complicated scenarios. However, many humanlike AD strategies just mimic drivers’ behaviors or some specific characteristic. Learning algorithms are powerful technics to realize these strategies, but the architectures in learning-based strategies are too simple or with no detailed foundations. Therefore, we mean to summarize drivers’ cognition characteristics and design a comprehensive and well-founded architecture for humanlike AD solutions. We review the massive studies about drivers with human driving or AD and summarize the characteristics from three perspectives, cognition foundation, cognition process, and cognition strategies. As for cognition foundation, we propose a simple analogy to show the working mechanisms of biological neural networks; as for cognition foundation, the important role of previous experience is highlighted; as for cognition strategies, we discuss drivers’ cognition compensation strategies under the influences of environment, vehicle automation, and personal states systematically. After the above review of drivers’ characteristics, we classify the methods to model drivers. We find that models based on cognition processes can maintain more cognition details, and thus we design a driving-dedicated cognitive architecture. This architecture works by the cooperation of several modules including long-term memory, management module, and so on. It has solid theoretical and factual foundations and can reflect drivers’ cognition characteristics comprehensively. Finally, we discuss what needs to be done in the near future for us to improve humanlike AD solutions gradually.
Shi-tao Chen, Jingyue Zheng, Masayoshi Tomizuka, Nanning Zheng 0001, Jianqiang Wang 0003
IEEE Trans. Intell. Transp. Syst.5
2022 Hierarchical Motion Planning for Autonomous Driving in Large-Scale Complex Scenarios
abstract
Motion planning algorithms, an essential part of the autonomous driving system, have been extensively studied. However, in large-scale complex scenarios, how to develop an optimal path to comply with the requirements of smoothness and safety remains a vital issue. In this study, a hierarchical search spacial scales-based hybrid A* (termed as HHA*) motion planning method is proposed, capable of efficiently generating smooth and safe paths. The proposed HHA* method covers two stages. First, the search space is divided on a coarse scale to generate local goals. Subsequently, the novel heuristic function and exploration strategies are adopted in the fine-scale search space to generate paths like that with a human driver guided by the local goals. Moreover, with the usage of the clothoid, the smoothness of the generated path is improved to be G2–continuous (i.e., curvature continuous), which fits the vehicle’s kinematic constraints without the need for later smoothing. Numerous experimental results from the simulation and on-road tests indicate that the proposed method can effectively perform motion planning that meets smoothness and safety in large-scale complex scenarios.
Songyi Zhang, Zhiqiang Jian, Xiaodong Deng, Shi-tao Chen, Zhixiong Nan, Nanning Zheng 0001
IEEE Trans. Intell. Transp. Syst.6
2022 Generalized Zero-Shot Learning Via Multi-Modal Aggregated Posterior Aligning Neural Network
abstract
The visual-semantic gap between the visual space (visual features) and semantic space (semantic attributes) is one of the main problems in the Generalized Zero-Shot Learning (GZSL) task. The essence of this problem is that the structure of manifolds in these two spaces is inconsistent, which makes it difficult to learn embeddings that unify visual features and semantic attributes for similarity measurement. In this work, we tackle this problem by proposing a multi-modal aggregated posterior aligning neural network based on Wasserstein Auto-encoders (WAE) which learns a shared latent space for visual features and semantic attributes. The key to our approach is that the aggregated posterior distribution of the latent representations encoded from visual features of each class is encouraged to be aligned with a Gaussian distribution predicted by the corresponding semantic attribute in the latent space. On one hand, requiring the latent manifolds of visual features and semantic attributes to be consistent preserves the inter-class association between seen and unseen classes. On the other hand, the aggregated posterior of each class is directly defined as a Gaussian in the latent space, which provides a reliable way to synthesize latent features for training classification models. Using the AWA1, AWA2, CUB, aPY, FLO, and SUN benchmark datasets, we extensively conducted comparative evaluations to demonstrate the advantages of our method over state-of-the-art approaches.
Xingyu Chen 0001, Jin Li 0011, Xuguang Lan, Nanning Zheng 0001
IEEE Trans. Multim.4
2022 Action Coherence Network for Weakly-Supervised Temporal Action Localization
abstract
Weakly-supervised Temporal Action Localization (W-TAL) aims at simultaneously classifying and locating all action instances with only video-level supervision. However, current W-TAL methods have two limitations. First, they ignore the difference in video representations between an action instance and its surrounding background when generating and scoring action proposals. Second, the unique characteristics of the RGB frames and optical flow are largely ignored when fusing these two modalities. To address these problems, an Action Coherence Network (ACN) is proposed in this paper. Its core is a new coherence loss which exploits both classification predictions and video content representations to supervise action boundary regression and thus leads to more accurate action localization results. Besides, the proposed ACN explicitly takes into account the specific characteristics of RGB frames and optical flow by training two separate sub-networks, each of which is able to generate modality-specific action proposals independently. Finally, to take advantage of the complementary action proposals generated by two streams, a novel fusion module is introduced to reconcile them and obtain the final action localization results. Experiments on the THUMOS14 and ActivityNet datasets show that our ACN outperforms the state-of-the-art W-TAL methods, and is even comparable to some recent fully-supervised methods. Particularly, ACN achieves a mean average precision of 26.4% on the THUMOS14 dataset under the IoU threshold 0.5.
Yuanhao Zhai 0001, Le Wang 0003, Wei Tang 0016, Qilin Zhang 0004, Nanning Zheng 0001, Gang Hua 0001
IEEE Trans. Multim.5
2022 Neighborhood Geometric Structure-Preserving Variational Autoencoder for Smooth and Bounded Data Sources
abstract
Many data sources, such as human poses, lie on low-dimensional manifolds that are smooth and bounded. Learning low-dimensional representations for such data is an important problem. One typical solution is to utilize encoder-decoder networks. However, due to the lack of effective regularization in latent space, the learned representations usually do not preserve the essential data relations. For example, adjacent video frames in a sequence may be encoded into very different zones across the latent space with holes in between. This is problematic for many tasks such as denoising because slightly perturbed data have the risk of being encoded into very different latent variables, leaving output unpredictable. To resolve this problem, we first propose a neighborhood geometric structure-preserving variational autoencoder (SP-VAE), which not only maximizes the evidence lower bound but also encourages latent variables to preserve their structures as in ambient space. Then, we learn a set of small surfaces to approximately bound the learned manifold to deal with holes in latent space. We extensively validate the properties of our approach by reconstruction, denoising, and random image generation experiments on a number of data sources, including synthetic Swiss roll, human pose sequences, and facial expression images. The experimental results show that our approach learns more smooth manifolds than the baselines. We also apply our approach to the tasks of human pose refinement and facial expression image interpolation where it gets better results than the baselines.
Xingyu Chen 0001, Chunyu Wang 0001, Xuguang Lan, Nanning Zheng 0001, Wenjun Zeng 0001
IEEE Trans. Neural Networks Learn. Syst.4
2022 RGB-D Point Cloud Registration Based on Salient Object Detection
abstract
We propose a robust algorithm for aligning rigid, noisy, and partially overlapping red green blue-depth (RGB-D) point clouds. To address the problems of data degradation and uneven distribution, we offer three strategies to increase the robustness of the iterative closest point (ICP) algorithm. First, we introduce a salient object detection (SOD) method to extract a set of points with significant structural variation in the foreground, which can avoid the unbalanced proportion of foreground and background point sets leading to the local registration. Second, registration algorithms that rely only on structural information for alignment cannot establish the correct correspondences when faced with the point set with no significant change in structure. Therefore, a bidirectional color distance (BCD) is designed to build precise correspondence with bidirectional search and color guidance. Third, the maximum correntropy criterion (MCC) and trimmed strategy are introduced into our algorithm to handle with noise and outliers. We experimentally validate that our algorithm is more robust than previous algorithms on simulated and real-world scene data in most scenarios and achieve a satisfying 3-D reconstruction of indoor scenes.
Teng Wan, Shaoyi Du, Wenting Cui, Runzhao Yao, Yuyan Ge, Ce Li 0001, Yue Gao 0002, Nanning Zheng 0001
IEEE Trans. Neural Networks Learn. Syst.8
2022 Perturbation of Spike Timing Benefits Neural Network Performance on Similarity Search
abstract
Perturbation has a positive effect, as it contributes to the stability of neural systems through adaptation and robustness. For example, deep reinforcement learning generally engages in exploratory behavior by injecting noise into the action space and network parameters. It can consistently increase the agent's exploration ability and lead to richer sets of behaviors. Evolutionary strategies also apply parameter perturbations, which makes network architecture robust and diverse. Our main concern is whether the notion of synaptic perturbation introduced in a spiking neural network (SNN) is biologically relevant or if novel frameworks and components are desired to account for the perturbation properties of artificial neural systems. In this work, we first review part of the locality-sensitive hashing (LSH) of similarity search, the FLY algorithm, as recently published in Science, and propose an improved architecture, time-shifted spiking LSH (TS-SLSH), with the consideration of temporal perturbations of the firing moments of spike pulses. Experiment results show promising performance of the proposed method and demonstrate its generality to various spiking neuron models. Therefore, we expect temporal perturbation to play an active role in SNN performance.
Ziru Wang, Badong Chen, Nanning Zheng 0001, Pengju Ren
IEEE Trans. Neural Networks Learn. Syst.5
2022 Multinetwork Collaborative Feature Learning for Semisupervised Person Reidentification
abstract
Person reidentification (Re-ID) aims at matching images of the same identity captured from the disjoint camera views, which remains a very challenging problem due to the large cross-view appearance variations. In practice, the mainstream methods usually learn a discriminative feature representation using a deep neural network, which needs a large number of labeled samples in the training process. In this article, we design a simple yet effective multinetwork collaborative feature learning (MCFL) framework to alleviate the data annotation requirement for person Re-ID, which can confidently estimate the pseudolabels of unlabeled sample pairs and consistently learn the discriminative features of input images. To keep the precision of pseudolabels, we further build a novel self-paced collaborative regularizer to extensively exchange the weight information of unlabeled sample pairs between different networks. Once the pseudolabels are correctly estimated, we take the corresponding sample pairs into the training process, which is beneficial to learn more discriminative features for person Re-ID. Extensive experimental results on the Market1501, DukeMTMC, and CUHK03 data sets have shown that our method outperforms most of the state-of-the-art approaches.
Sanping Zhou, Jinjun Wang, Deyu Meng, Le Wang 0003, Nanning Zheng 0001
IEEE Trans. Neural Networks Learn. Syst.6
2021 Weakly Supervised Temporal Action Localization Through Learning Explicit Subspaces for Action and Context
abstract
Weakly-supervised Temporal Action Localization (WS-TAL) methods learn to localize temporal starts and ends of action instances in a video under only video-level supervision. Existing WS-TAL methods rely on deep features learned for action recognition. However, due to the mismatch between classification and localization, these features cannot distinguish the frequently co-occurring contextual background, i.e., the context, and the actual action instances. We term this challenge action-context confusion, and it will adversely affect the action localization accuracy. To address this challenge, we introduce a framework that learns two feature subspaces respectively for actions and their context. By explicitly accounting for action visual elements, the action instances can be localized more precisely without the distraction from the context. To facilitate the learning of these two feature subspaces with only video-level categorical labels, we leverage the predictions from both spatial and temporal streams for snippets grouping. In addition, an unsupervised learning task is introduced to make the proposed module focus on mining temporal information. The proposed approach outperforms state-of-the-art WS-TAL methods on three benchmarks, i.e., THUMOS14, ActivityNet v1.2 and v1.3 datasets.
Ziyi Liu 0001, Le Wang 0003, Wei Tang 0016, Junsong Yuan 0001, Nanning Zheng 0001, Gang Hua 0001
AAAI5
2021 ACSNet: Action-Context Separation Network for Weakly Supervised Temporal Action Localization
abstract
The object of Weakly-supervised Temporal Action Localization (WS-TAL) is to localize all action instances in an untrimmed video with only video-level supervision. Due to the lack of frame-level annotations during training, current WS-TAL methods rely on attention mechanisms to localize the foreground snippets or frames that contribute to the video-level classification task. This strategy frequently confuse context with the actual action, in the localization result. Separating action and context is a core problem for precise WS-TAL, but it is very challenging and has been largely ignored in the literature. In this paper, we introduce an Action-Context Separation Network (ACSNet) that explicitly takes into account context for accurate action localization. It consists of two branches (i.e., the Foreground-Background branch and the Action-Context branch). The Foreground-Background branch first distinguishes foreground from background within the entire video while the Action-Context branch further separates the foreground as action and context. We associate video snippets with two latent components (i.e., a positive component and a negative component), and their different combinations can effectively characterize foreground, action and context. Furthermore, we introduce extended labels with auxiliary context categories to facilitate the learning of action-context separation. Experiments on THUMOS14 and ActivityNet v1.2/v1.3 datasets demonstrate the ACSNet outperforms existing state-of-the-art WS-TAL methods by a large margin.
Ziyi Liu 0001, Le Wang 0003, Qilin Zhang 0004, Wei Tang 0016, Junsong Yuan 0001, Nanning Zheng 0001, Gang Hua 0001
AAAI6
2021 Semantic Consistency Networks for 3D Object Detection
abstract
Detecting 3D objects from point clouds is a significant yet challenging issue in many applications. While most existing approaches seek to leverage geometric information of point clouds, few studies accommodate the inherent semantic characteristics of each point and the consistency between the geometric and semantic cues. In this work, we propose a novel semantic consistency network (SCNet) driven by a natural principle: the class of a predicted 3D bounding box should be consistent with the classes of all the points inside this box. Specifically, our SCNet consists of a feature extraction structure, a detection decision structure, and a semantic segmentation structure. In inference, the feature extraction and the detection decision structures are used to detect 3D objects. In training, the semantic segmentation structure is jointly trained with the other two structures to produce more robust and applicative model parameters. A novel semantic consistency loss is proposed to regulate the output 3D object boxes and the segmented points to boost the performance. Our model is evaluated on two challenging datasets and achieves comparable results to the state-of-the-art methods.
Wenwen Wei, Ping Wei 0001, Nanning Zheng 0001
AAAI3
2021 End-to-End Object Detection With Fully Convolutional Network
abstract
Mainstream object detectors based on the fully convolutional network has achieved impressive performance. While most of them still need a hand-designed non-maximum suppression (NMS) post-processing, which impedes fully end-to-end training. In this paper, we give the analysis of discarding NMS, where the results reveal that a proper label assignment plays a crucial role. To this end, for fully convolutional detectors, we introduce a Prediction-aware One-To-One (POTO) label assignment for classification to enable end-to-end detection, which obtains comparable performance with NMS. Besides, a simple 3D Max Filtering (3DMF) is proposed to utilize the multi-scale features and improve the discriminability of convolutions in the local region. With these techniques, our end-to-end framework achieves competitive performance against many state-of-the-art detectors with NMS on COCO and CrowdHuman datasets. The code is available at https://github.com/Megvii-BaseDetection/DeFCN.
Lin Song 0002, Hongbin Sun 0001, Jian Sun 0001, Nanning Zheng 0001
CVPR6
2021 SpV8: Pursuing Optimal Vectorization and Regular Computation Pattern in SpMV
abstract
Sparse Matrix-Vector Multiplication (SpMV) plays an important role in many scientific and industry applications, and remains a well-known challenge due to the high sparsity and irregularity. Most existing researches on SpMV try to pursue high vectorization efficiency. However, such approaches may suffer from non-negligible speculation penalty due to their irregular computation patterns. In this paper, we propose SpV8, a novel approach that optimizes both speculation and vectorization in SpMV. Specifically, SpV8 analyzes data distribution in different matrices and row panels, and accordingly applies optimization method that achieves the maximal vectorization with regular computation patterns. We evaluate SpV8 on Intel Xeon CPU and compare with multiple state-of-art SpMV algorithms using 71 sparse matrices. The results show that SpV8 achieves up to 10× speedup (average 2.8×) against the standard MKL SpMV routine, and up to 2.4× speedup (average 1.4×) against the best existing approach. Moreover, SpMV features very low preprocessing overhead in all compared approaches, which indicates SpV8 is highly-applicable in real-world applications.
Tian Xia 0008, Wenzhe Zhao 0001, Nanning Zheng 0001, Pengju Ren
DAC4
2021 Correlation-Based Robust Linear Regression with Iterative Outlier Removal
abstract
Here we consider linear regression from the view of correlation and propose a robust regression algorithm. The main idea of this work is from the fact that the inliers lying in a low dimensional subspace are mostly correlated, and the presence of outliers leads to the decrease of correlation. We design an iterative outlier removal algorithm based on correlation, by which the outliers can be effectively removed in a normal-distributed or uniform-distributed data set. Finally, the linear equation is calculated based on the remaining points. The experiment results show that the proposed method outperforms the state-of-the-art approaches. In some cases in which outliers are more than inliers, the proposed method can still obtain the real formulas.
Jianji Wang 0001, Yuanjie Li, Nanning Zheng 0001
ICASSP5
2021 Meta Pairwise Relationship Distillation for Unsupervised Person Re-identification
abstract
Unsupervised person re-identification (Re-ID) remains challenging due to the lack of ground-truth labels. Existing methods often rely on estimated pseudo labels via iterative clustering and classification, and they are unfortunately highly susceptible to performance penalties incurred by the inaccurate estimated number of clusters. Alternatively, we propose the Meta Pairwise Relationship Distillation (MPRD) method to estimate the pseudo labels of sample pairs for unsupervised person Re-ID. Specifically, it consists of a Convolutional Neural Network (CNN) and Graph Convolutional Network (GCN), in which the GCN estimates the pseudo labels of sample pairs based on the current features extracted by CNN, and the CNN learns better features by involving high-fidelity positive and negative sample pairs imposed by GCN. To achieve this goal, a small amount of labeled samples are used to guide GCN training, which can distill meta knowledge to judge the difference in the neighborhood structure between positive and negative sample pairs. Extensive experiments on Market-1501, DukeMTMC-reID and MSMT17 datasets show that our method outperforms the state-of-the-art approaches.
Haoxuanye Ji, Le Wang 0003, Sanping Zhou, Wei Tang 0016, Nanning Zheng 0001, Gang Hua 0001
ICCV5
2021 Unlimited Neighborhood Interaction for Heterogeneous Trajectory Prediction
abstract
Understanding complex social interactions among agents is a key challenge for trajectory prediction. Most existing methods consider the interactions between pairwise traffic agents or in a local area, while the nature of interactions is unlimited, involving an uncertain number of agents and non-local areas simultaneously. Besides, they treat heterogeneous traffic agents the same, namely those among agents of different categories, while neglecting people’s diverse reaction patterns toward traffic agents in different categories. To address these problems, we propose a simple yet effective Unlimited Neighborhood Interaction Network (UNIN), which predicts trajectories of heterogeneous agents in multiple categories. Specifically, the proposed unlimited neighborhood interaction module generates the fused-features of all agents involved in an interaction simultaneously, which is adaptive to any number of agents and any range of interaction area. Meanwhile, a hierarchical graph attention module is proposed to obtain category-to-category interaction and agent-to-agent interaction. Finally, parameters of a Gaussian Mixture Model are estimated for generating the future trajectories. Extensive experimental results on benchmark datasets demonstrate a significant performance improvement of our method over the state-of-the-art methods.
Fang Zheng 0009, Le Wang 0003, Sanping Zhou, Wei Tang 0016, Zhenxing Niu, Nanning Zheng 0001, Gang Hua 0001
ICCV6
2021 Practical Relative Order Attack in Deep Ranking
abstract
Recent studies unveil the vulnerabilities of deep ranking models, where an imperceptible perturbation can trigger dramatic changes in the ranking result. While previous attempts focus on manipulating absolute ranks of certain candidates, the possibility of adjusting their relative order remains under-explored. In this paper, we formulate a new adversarial attack against deep ranking systems, i.e., the Order Attack, which covertly alters the relative order among a selected set of candidates according to an attacker-specified permutation, with limited interference to other unrelated candidates. Specifically, it is formulated as a triplet-style loss imposing an inequality chain reflecting the specified permutation. However, direct optimization of such white-box objective is infeasible in a real-world attack scenario due to various black-box limitations. To cope with them, we propose a Short-range Ranking Correlation metric as a surrogate objective for black-box Order Attack to approximate the white-box method. The Order Attack is evaluated on the Fashion-MNIST and Stanford-Online-Products datasets under both white-box and black-box threat models. The black-box attack is also successfully implemented on a major e-commerce platform. Comprehensive experimental evaluations demonstrate the effectiveness of the proposed methods, revealing a new type of ranking model vulnerability.
Le Wang 0003, Zhenxing Niu, Qilin Zhang 0004, Nanning Zheng 0001, Gang Hua 0001
ICCV6
2021 Enriching Local and Global Contexts for Temporal Action Localization
abstract
Effectively tackling the problem of temporal action localization (TAL) necessitates a visual representation that jointly pursues two confounding goals, i.e., fine-grained discrimination for temporal localization and sufficient visual invariance for action classification. We address this challenge by enriching both the local and global contexts in the popular two-stage temporal localization framework, where action proposals are first generated followed by action classification and temporal boundary regression. Our proposed model, dubbed ContextLoc, can be divided into three sub-networks: L-Net, G-Net and P-Net. L-Net enriches the local context via fine-grained modeling of snippet-level features, which is formulated as a query-and-retrieval process. G-Net enriches the global context via higher-level modeling of the video-level representation. In addition, we introduce a novel context adaptation module to adapt the global context to different proposals. P-Net further models the context-aware inter-proposal relations. We explore two existing models to be the P-Net in our experiments. The efficacy of our proposed method is validated by experimental results on the THUMOS14 (54.3% at [email protected]) and ActivityNet v1.3 (56.01% at [email protected]) datasets, which outperforms recent states of the art. Code is available at https://github.com/buxiangzhiren/ContextLoc.
Zixin Zhu, Wei Tang 0016, Le Wang 0003, Nanning Zheng 0001, Gang Hua 0001
ICCV4
2021 TAG-Reg: Iterative Accurate Global Registration Algorithm
abstract
In this paper, we propose an accurate global registration (TAG-Reg) algorithm for poor initialization and partially overlapping point clouds registration problem. Firstly, methods based on geometric structure information of points can get the accurate results, which is vulnerable to poor initialization. Meanwhile, existing features based global methods can solve poor initialization problem at a certain extent, but it cannot obtain accurate results. So, we combine the geometric structure information with feature as hybrid feature to solve poor initialization problem completely and obtain accurate results. Secondly, we introduce dynamic trimmed strategy combining with hybrid feature to deal with partially overlapping problem. Then, to improve the accuracy of our method, we utilize the probabilistic method to suppress noise. At last, we establish the TAG-Reg model and propose an iterative algorithm to solve this problem. Experimental results show that our TAG-Reg achieves state-of-the-art performance compared to existing non-deep learning and recent deep learning methods. Our source code will open at https://github.com/BiaoBiaoLi/TAG-Reg.
Qixing Xie, Shaoyi Du, Wenting Cui, Runzhao Yao, Yue Gao 0002, Nanning Zheng 0001
ICME7
2021 Probabilistic Human Motion Prediction via A Bayesian Neural Network
abstract
Human motion prediction is an important and challenging topic that has promising prospects in efficient and safe human-robot-interaction systems. Currently, the majority of the human motion prediction algorithms are based on deterministic models, which may lead to risky decisions for robots. To solve this problem, we propose a probabilistic model for human motion prediction in this paper. The key idea of our approach is to extend the conventional deterministic motion prediction neural network to a Bayesian one. On one hand, our model could generate several future motions when given an observed motion sequence. On the other hand, by calculating the Epistemic Uncertainty and the Heteroscedastic Aleatoric Uncertainty, our model could tell the robot if the observation has been seen before and also give the optimal result among all possible predictions. We extensively validate our approach on a large scale benchmark dataset Human3.6m. The experiments show that our approach performs better than deterministic methods. We further evaluate our approach in a Human-Robot-Interaction (HRI) scenario. The experimental results show that our approach makes the interaction more efficient and safer.
Xingyu Chen 0001, Xuguang Lan, Nanning Zheng 0001
ICRA4
2021 REGNet: REgion-based Grasp Network for End-to-end Grasp Detection in Point Clouds
abstract
Reliable robotic grasping in unstructured environments is a crucial but challenging task. The main problem is to generate the optimal grasp of novel objects from partial noisy observations. This paper presents an end-to-end grasp detection network taking one single-view point cloud as input to tackle the problem. Our network includes three stages: Score Network (SN), Grasp Region Network (GRN), and Refine Network (RN). Specifically, SN regresses point grasp confidence and selects positive points with high confidence. Then GRN conducts grasp proposal prediction on the selected positive points. RN generates more accurate grasps by refining proposals predicted by GRN. To further improve the performance, we propose a grasp anchor mechanism, in which grasp anchors with assigned gripper orientations are introduced to generate grasp proposals. Experiments demonstrate that REGNet achieves a success rate of 79.34% and a completion rate of 96% in real-world clutter, which significantly outperforms several state-of-the-art point-cloud based methods, including GPD, PointNetGPD, and S4G. The code is available at https://github.com/zhaobinglei/REGNet for 3D Grasping.
Binglei Zhao 0002, Hanbo Zhang, Xuguang Lan, Nanning Zheng 0001
ICRA6
2021 Hindsight Trust Region Policy Optimization
abstract
Reinforcement Learning (RL) with sparse rewards is a major challenge. We pro- pose Hindsight Trust Region Policy Optimization (HTRPO), a new RL algorithm that extends the highly successful TRPO algorithm with hindsight to tackle the challenge of sparse rewards. Hindsight refers to the algorithm’s ability to learn from information across goals, including past goals not intended for the current task. We derive the hindsight form of TRPO, together with QKL, a quadratic approximation to the KL divergence constraint on the trust region. QKL reduces variance in KL divergence estimation and improves stability in policy updates. We show that HTRPO has similar convergence property as TRPO. We also present Hindsight Goal Filtering (HGF), which further improves the learning performance for suitable tasks. HTRPO has been evaluated on various sparse-reward tasks, including Atari games and simulated robot control. Experimental results show that HTRPO consistently outperforms TRPO, as well as HPG, a state-of-the-art policy 14 gradient algorithm for RL with sparse rewards.
Hanbo Zhang, Cedar Site Bai, Xuguang Lan, David Hsu, Nanning Zheng 0001
IJCAI5
2021 Exploring Effective DNN Models for Forensic Age Estimation based on Panoramic Radiograph Images
abstract
Dental age estimation is widely used in forensic identification, but the accuracy of traditional methods cannot satisfy the demand for accuracy, especially for age estimation of adults. We introduce a deep learning-based methodology to estimate the age based on collected X-ray images of the teeth. We present a new dental dataset, which contains labeled orthopan-tomograms (OPGs) of 27,957 people, including 16,383 OPGs for females as well as 11,574 OPGs for males. All ages range from 0 to 93-year-old with a median of 27. The accuracy of the age labels is guaranteed by the ID card information. Aiming at the characteristics of the dental data itself, we explore various neural network elements that are effective for age estimation, including proper network depth, convolution kernel size, multi-branch structure, and the feature reusing of early layers. Based on the characteristic exploration, we further search models for dental age estimation by using the popular Neural Architecture Search (NAS) method. Experiment results show that our model achieves a mean absolute error (MAE) of 1.64 years, surpass all existing CNN models. Compared with Inception-v4 with an MAE of 1.70 and 20.46B FLOPs (inputs size 384×384), the FLOPs of our model can be reduced by 2.7 times (7.49B FLOPs). To our best knowledge, this is the first study for age estimation by exploring and searching the DNN model. Our results have surpassed legal medical expert-level performance (with an MAE of more than 2) for age estimation. Our methodology and results in this paper are very meaningful to forensic medicine for aging estimation with panoramic radiograph images.
Wenxuan Hou 0001, Longjun Liu, Jinxia Gao, Anguo Zhu, Keyang Pan, Hongbin Sun 0001, Nanning Zheng 0001
IJCNN7
2021 CAQ: Context-Aware Quantization via Reinforcement Learning
abstract
Model quantization is a crucial step for porting Deep Neural Networks (DNNs) on embedded devices to meet the limited computation and storage resources requirement. Traditional methods usually obtain the scaling factor and quantize the weights based on the information of single layer. However, our analysis indicate that these selection methods of scaling factor overlook the differences and dependencies among layers, leading to large truncation errors or zeroing errors, which is the main reason for the performance degradation. To this end, we propose a Context-Aware Quantization (CAQ) scheme, which formalizes the model quantization as a global optimization problem and leverages reinforcement learning to search for the optimal scaling factors based on the entire model. Further, we adopt shift-based scaling factors to narrow the search space to improve the search efficiency, additionally, it reduces the computational complexity during the inference phase, and also provides a simpler and more robust activation calibration solution. We extensively test our scheme on a wide range of Neural Networks, including ResNet 50/101/152, InceptionV3 and MobileNetV2 on ImageNet, the entire search process only takes about 1 hour on a single GeForce RTX 2080 Ti. Compared with the existed methods, Our scheme can get a better performance, which could maintain the post-quantization accuracy loss less than 0.25%, while reducing memory footprint by 5%-8% and multiply accumulate (MAC) operations by 2%-4%. Besides, we further show that the CAQ can be applied on other tasks, such as object detection and segmentation.
Zhijun Tu, Tian Xia 0008, Wenzhe Zhao 0001, Pengju Ren, Nanning Zheng 0001
IJCNN6
2021 Joint Critics Mechanism: A Universal Framework for Multi-targets Visual Navigation
abstract
Regarding to target-driven visual navigation problem, training a universal value function or policy function approximator is considered to be a fairly difficult task, because there may exits potential conflicts among different targets. When modeling navigation as a goal-conditional reinforcement learning problem, the algorithm can only support a relatively small number of goals, which limits the universality of the reinforcement learning based methods. In this work, we proposed a framework for multi-targets visual navigation, termed Joint Critics Mechanism, to better train the universal policy function approximator. Recognizing that target-specific network has better convergence, we use the target-specific value network to estimate the advantage of the target-universal policy network for better convergence. In this way, we avoid complexity of learning competitive targets and achieve a better convergence with a larger number of targets. For evaluation, we conduct experiments in realistic simulation environments and the results prove the rationality and effectiveness of our proposed framework.
Youzhuo Wang, Wenzhe Zhao 0001, Tian Xia 0008, Nanning Zheng 0001, Pengju Ren
IJCNN4
2021 A Lightweight sequence-based Unsupervised Loop Closure Detection
abstract
Stable, effective and lightweight loop closure detection is an always pursued goal in real-time SLAM systems, that can be ported on embedded processors and deployed on autonomous robotics. Deep learning methods have extended the expressive ability and adaptability of the descriptor, and sequence-based methods can greatly improve the matching accuracy. However, the increased computation complexity and storage bandwidth requirements of matching calculations for high-dimensional descriptor make it infeasible for real-time deployment, especially for robots that navigate in relatively big maps. To address this challenge, we propose a lightweight sequence-based unsupervised loop closure detection scheme. To be specific, Principal Component Analysis (PCA) is applied to squeeze the descriptor dimensions while maintaining sufficient expressive ability. Additionally, with the consideration of the image sequence and combining linear query with fast approximate nearest neighbor search to further reduce the execution time and improve the efficiency of sequence matching. We implement our method on CALC, a state-of-the-art unsupervised solution, and conduct experiments on NVIDIA TX2, results demonstrate that the accuracy has been improved by 5%, while the execution speed is 2× faster. Source code is available at https://github.com/Mingrui-Yu/Seq-CALC.
Fan Xiong, Wenzhe Zhao 0001, Nanning Zheng 0001, Pengju Ren
IJCNN5
2021 DSP-Net: Dense-to-Sparse Proposal Generation Approach for 3D Object Detection on Point Cloud
abstract
Object proposals generated based on sparse points from the raw point cloud have been widely used in 3D object detection. However, following the above scheme, most existing proposal generators have two problems, one is that the features for proposal generation constrain the detection performance by containing insufficient information; the other is that the sparse points obtained from the raw point cloud are misaligned with their corresponding objects in location and feature aspects. In this paper, we propose a dense-to-sparse proposal generation approach for 3D object detection, which can deal with the two problems simultaneously. Our approach utilizes the 3D CNN backbone to output dense features as a supplement to the original sparse point features for proposal generation. Besides, an object-aware feature pooling module is designed to address the misalignment between sparse points and corresponding objects. Experiments on the KITTI dataset show that our method outperforms the existing sparse-style methods and other published state-of-the-art methods.
Xinrui Yan, Shi-tao Chen, Zhixiong Nan, Jingmin Xin, Nanning Zheng 0001
IJCNN6
2021 MPR-Net: Multi-Scale Key Points Regression for Lane Detection
abstract
Lane detection is usually regarded as a semantic segmentation task, however, segmentation-based methods require expensive computational costs, and it is difficult to segment the lanes with heavy noises. Therefore, this paper proposes a lane detection method based on multi-scale key point regression, which directly uses the position of the key points of the lane instead of predicting the pixel-wise outputs. First, we design a lightweight backbone to extract a series of feature maps with a forward view image as input, and then apply a multi-scale fusion network on these feature maps to obtain the location and confidence information of the key points of the lane. Finally, a clustering and curve fitting mechanism with quadratic inverse proportion are used to obtain the final lane detection. Our proposed model can recognize dashed lane markings and handle many extreme scenarios where lanes are completely occluded or heavily noised. In addition, our model uses a relatively explicit framework, which contributes to ensuring real-time performance at 30Hz. In order to prove our method's performance, we conduct experiments on the TuSimple benchmark and RVD dataset, and results demonstrate that our method achieves competitive results compared with other methods.
Dantong Zhu, Shengqi Wang, Shi-tao Chen, Zhixiong Nan, Nanning Zheng 0001
IV6
2021 X-GGM: Graph Generative Modeling for Out-of-distribution Generalization in Visual Question Answering
abstract
Encouraging progress has been made towards Visual Question Answering (VQA) in recent years, but it is still challenging to enable VQA models to adaptively generalize to out-of-distribution (OOD) samples. Intuitively, recompositions of existing visual concepts (i.e., attributes and objects) can generate unseen compositions in the training set, which will promote VQA models to generalize to OOD samples. In this paper, we formulate OOD generalization in VQA as a compositional generalization problem and propose a graph generative modeling-based training scheme (X-GGM) to handle the problem implicitly. X-GGM leverages graph generative modeling to iteratively generate a relation matrix and node representations for the predefined graph that utilizes attribute-object pairs as nodes. Furthermore, to alleviate the unstable training issue in graph generative modeling, we propose a gradient distribution consistency loss to constrain the data distribution with adversarial perturbations and the generated distribution. The baseline VQA model (LXMERT) trained with the X-GGM scheme achieves state-of-the-art OOD performance on two standard VQA OOD benchmarks, i.e., VQA-CP v2 and GQA-OOD. Extensive ablation studies demonstrate the effectiveness of X-GGM components.
Jingjing Jiang, Ziyi Liu 0001, Yifan Liu 0014, Zhixiong Nan, Nanning Zheng 0001
ACM Multimedia5
2021 AKECP: Adaptive Knowledge Extraction from Feature Maps for Fast and Efficient Channel Pruning
abstract
Pruning can remove redundant parameters and structures of Deep Neural Networks (DNNs) to reduce inference time and memory overhead. As an important component of neural networks, the feature map (FM) has stated to be adopted for network pruning. However, the majority of FM-based pruning methods do not fully investigate effective knowledge in the FM for pruning. In addition, it is challenging to design a robust pruning criterion with a small number of images and achieve parallel pruning due to the variability of FMs. In this paper, we propose Adaptive Knowledge Extraction for Channel Pruning (AKECP), which can compress the network fast and efficiently. In AKECP, we first investigate the characteristics of FMs and extract effective knowledge with an adaptive scheme. Secondly, we formulate the effective knowledge of FMs to measure the importance of corresponding network channels. Thirdly, thanks to the effective knowledge extraction, AKECP can efficiently and simultaneously prune all the layers with extremely few or even one image. Experimental results show that our method can compress various networks on different datasets without introducing additional constraints, and it has advanced the state-of-the-arts. Notably, for ResNet-110 on CIFAR-10, AKECP achieves 59.9% of parameters and 59.8% of FLOPs reduction with negligible accuracy loss. For ResNet-50 on ImageNet, AKECP saves 40.5% of memory footprint and reduces 44.1% of FLOPs with only 0.32% of Top-1 accuracy drop.
Haonan Zhang 0002, Longjun Liu, Hengyi Zhou, Wenxuan Hou 0001, Hongbin Sun 0001, Nanning Zheng 0001
ACM Multimedia6
2021 Instance-Conditional Knowledge Distillation for Object Detection
abstract
Knowledge distillation has shown great success in classification, however, it is still challenging for detection. In a typical image for detection, representations from different locations may have different contributions to detection targets, making the distillation hard to balance. In this paper, we propose a conditional distillation framework to distill the desired knowledge, namely knowledge that is beneficial in terms of both classification and localization for every instance. The framework introduces a learnable conditional decoding module, which retrieves information given each target instance as query. Specifically, we encode the condition information as query and use the teacher's representations as key. The attention between query and key is used to measure the contribution of different features, guided by a localization-recognition-sensitive auxiliary task. Extensive experiments demonstrate the efficacy of our method: we observe impressive improvements under various settings. Notably, we boost RetinaNet with ResNet-50 backbone from $37.4$ to $40.7$ mAP ($+3.3$) under $1\times$ schedule, that even surpasses the teacher ($40.4$ mAP) with ResNet-101 backbone under $3\times$ schedule. Code has been released on https://github.com/megvii-research/ICD.
Zijian Kang, Peizhen Zhang, Xiangyu Zhang 0005, Jian Sun 0001, Nanning Zheng 0001
NeurIPS5
2021 Dynamic Grained Encoder for Vision Transformers
abstract
Transformers, the de-facto standard for language modeling, have been recently applied for vision tasks. This paper introduces sparse queries for vision transformers to exploit the intrinsic spatial redundancy of natural images and save computational costs. Specifically, we propose a Dynamic Grained Encoder for vision transformers, which can adaptively assign a suitable number of queries to each spatial region. Thus it achieves a fine-grained representation in discriminative regions while keeping high efficiency. Besides, the dynamic grained encoder is compatible with most vision transformer frameworks. Without bells and whistles, our encoder allows the state-of-the-art vision transformers to reduce computational complexity by 40%-60% while maintaining comparable performance on image classification. Extensive experiments on object detection and segmentation further demonstrate the generalizability of our approach. Code is available at https://github.com/StevenGrove/vtpack.
Lin Song 0002, Songyang Zhang 0001, Xuming He 0001, Hongbin Sun 0001, Jian Sun 0001, Nanning Zheng 0001
NeurIPS8
2021 Co-evolution Transformer for Protein Contact Prediction
abstract
Proteins are the main machinery of life and protein functions are largely determined by their 3D structures. The measurement of the pairwise proximity between amino acids of a protein, known as inter-residue contact map, well characterizes the structural information of a protein. Protein contact prediction (PCP) is an essential building block of many protein structure related applications. The prevalent approach to contact prediction is based on estimating the inter-residue contacts using hand-crafted coevolutionary features derived from multiple sequence alignments (MSAs). To mitigate the information loss caused by hand-crafted features, some recently proposed methods try to learn residue co-evolutions directly from MSAs. These methods generally derive coevolutionary features by aggregating the learned residue representations from individual sequences with equal weights, which is inconsistent with the premise that residue co-evolutions are a reflection of collective covariation patterns of numerous homologous proteins. Moreover, non-homologous residues and gaps commonly exist in MSAs. By aggregating features from all homologs equally, the non-homologous information may cause misestimation of the residue co-evolutions. To overcome these issues, we propose an attention-based architecture, Co-evolution Transformer (CoT), for PCP. CoT jointly considers the information from all homologous sequences in the MSA to better capture global coevolutionary patterns. To mitigate the influence of the non-homologous information, CoT selectively aggregates the features from different homologs by assigning smaller weights to non-homologous sequences or residue pairs. Extensive experiments on two rigorous benchmark datasets demonstrate the effectiveness of CoT. In particular, CoT achieves a $51.6\%$ top-L long-range precision score for the Free Modeling (FM) domains on the CASP14 benchmark, which outperforms the winner group of CASP14 contact prediction challenge by $9.8\%$.
Fusong Ju, Jianwei Zhu, Liang He 0010, Bin Shao 0002, Nanning Zheng 0001, Tie-Yan Liu
NeurIPS6
2021 Predicting short-term next-active-object through visual attention and hand position
Jingjing Jiang, Zhixiong Nan, Hui Chen 0036, Shi-tao Chen, Nanning Zheng 0001
Neurocomputing5
2021 Nesting spatiotemporal attention networks for action recognition
Jiapeng Li 0003, Ping Wei 0001, Nanning Zheng 0001
Neurocomputing3
2021 A joint object detection and semantic segmentation model with cross-attention and inner-attention mechanisms
Zhixiong Nan, Jizhi Peng, Jingjing Jiang, Hui Chen 0036, Ben Yang, Jingmin Xin, Nanning Zheng 0001
Neurocomputing7
2021 Associations between MSE and SSIM as cost functions in linear decomposition with application to bit allocation for sparse coding
Jianji Wang 0001, Nanning Zheng 0001, Badong Chen, José C. Príncipe, Fei-Yue Wang 0001
Neurocomputing3
2021 Graph-based temporal action co-localization from an untrimmed video
Le Wang 0003, Changbo Zhai, Qilin Zhang 0004, Wei Tang 0016, Nanning Zheng 0001, Gang Hua 0001
Neurocomputing5
2021 A multilevel fusion network for 3D object detection
Chunlong Xia, Ping Wei 0001, Wenwen Wei, Nanning Zheng 0001
Neurocomputing4
2021 A multi-cue guidance network for depth completion
Yongchi Zhang, Ping Wei 0001, Nanning Zheng 0001
Neurocomputing3
2021 HSC: Leveraging horizontal shortcut connections for improving accuracy and computational efficiency of lightweight CNN
Anguo Zhu, Longjun Liu, Wenxuan Hou 0001, Hongbin Sun 0001, Nanning Zheng 0001
Neurocomputing5
2021 Societal Intelligence for Safer and Smarter Transportation
abstract
Recent years have witnessed exciting developments in our transportation system with increasingly intelligent vehicles and infrastructure. The transportation system is envisioned to be highly heterogeneous, consisting of diverse participants with mixed intelligence and connectivity. Among them, autonomous vehicles have the highest intelligence and connectivity level and could contribute greatly to the operation of the transportation system in an efficient and reliable manner. However, the current design of autonomous driving techniques is mostly concerned with the autonomous vehicle at the individual level, and the overall transportation system does not provide proactive support to autonomous driving. In fact, the increasing intelligence and connectivity in transportation could be leveraged to significantly enhance the safety and efficiency of individual vehicles and the entire system. To facilitate this, vehicles need to interact and cooperate both among themselves and with the transportation infrastructure and management. In this article, we propose the societal intelligence (SI) framework. Different from the existing multientity intelligence frameworks, SI allows for much diverse interactions among the multiple entities at different levels and is thus suitable for transportation. In addition, we also render the driving process into four functional layers and demonstrate how the social intelligence framework can adapt to these layers, respectively.
Xiang Cheng 0001, Dongliang Duan, Liuqing Yang 0001, Nanning Zheng 0001
IEEE Internet Things J.4
2021 Single-Image super-resolution - When model adaptation matters
Yudong Liang, Radu Timofte, Jinjun Wang, Sanping Zhou, Yihong Gong, Nanning Zheng 0001
Pattern Recognit.6
2021 Efficient Repair Analysis Algorithm Exploration for Memory With Redundancy and In-Memory ECC
abstract
In-memory error correction code (ECC) is a promising technique to improve the yield and reliability of high density memory design. However, the use of in-memory ECC poses a new problem to memory repair analysis algorithm, which has not been explored before. This article first makes a quantitative evaluation and demonstrates that the straightforward algorithms for memory with redundancy and in-memory ECC have serious deficiency on either repair rate or repair analysis speed. Accordingly, an optimal repair analysis algorithm that leverages preprocessing/filter algorithms, hybrid search tree, and depth-first search strategy is proposed to achieve low computational complexity and optimal repair rate in the meantime. In addition, a heuristic repair analysis algorithm that uses a greedy strategy is proposed to efficiently find repair solutions. Experimental results demonstrate that the proposed optimal repair analysis algorithm can achieve optimal repair rate and increase the repair analysis speed by up to 105×105× compared with the straightforward exhaustive search algorithm. The proposed heuristic repair analysis algorithm is approximately 28 percent faster than the proposed optimal algorithm, at the expense of 5.8 percent repair rate loss.
Minjie Lv, Hongbin Sun 0001, Jingmin Xin, Nanning Zheng 0001
IEEE Trans. Computers4
2021 PIT: Processing-In-Transmission With Fine-Grained Data Manipulation Networks
abstract
In the domain of data parallel computation, most works focus on data flow optimization inside the PE array and favorable memory hierarchy to pursue the maximum parallelism and efficiency, while the importance of data contents has been overlooked for a long time. As we observe, for structured data, insights on the contents (i.e., their values and locations within a structured form) can greatly benefit the computation performance, as fine-grained data manipulation can be performed. In this paper, we claim that by providing a flexible and adaptive data path, an efficient architecture with capability of fine-grained data manipulation can be built. Specifically, we design SOM, a portable and highly-adaptive data transmission network, with the capability of operand sorting, non-blocking self-route ordering and multicasting. Based on SOM, we propose the processing-in-transmission architecture (PITA), which extends the traditional SIMD architecture to perform some fundamental data processing during its transmission, by embedding multiple levels of SOM networks on the data path. We evaluate the performance of PITA in two irregular computation problems. We first map the matrix inversion task onto PITA and show considerable performance gain can be achieved, resulting in 3x-20x speedup against Intel MKL, and 20x-40x against cuBLAS. Then we evaluate our PITA on sparse CNNs. The results indicate that PITA can greatly improve computation efficiency and reduce memory bandwidth pressure. We achieved 2x-9x speedup against several state-of-art accelerators on sparse CNN, where nearly 100 percent PE efficiency is maintained under high sparsity. We believe the concept of PIT is a promising computing paradigm that can enlarge the capability of traditional parallel architecture.
Pengchen Zong, Tian Xia 0008, Jianming Tong, Wenzhe Zhao 0001, Nanning Zheng 0001, Pengju Ren
IEEE Trans. Computers7
2021 Dynamic Dataflow Scheduling and Computation Mapping Techniques for Efficient Depthwise Separable Convolution Acceleration
abstract
Depthwise separable convolution (DSC) has become one of the essential structures for lightweight convolutional neural networks. Nevertheless, its hardware architecture has not received much attention. Several previous hardware designs incur either high off-chip memory traffic or large on-chip memory usage, and hence have deficiency in terms of hardware efficiency as well as performance. This paper proposes two efficient dynamic design techniques, i.e. adaptive row-based dataflow scheduling and adaptive computation mapping, to achieve a much better trade-off between hardware efficiency and performance for DSC-based lightweight CNN accelerator. The effectiveness and efficiency of the proposed dynamic design techniques have been extensively evaluated using six DSC-based lightweight CNNs. Compared with the reference architectures, the simulation results show the proposed architectural techniques can at least reduce on-chip buffer size by 50.4% and improve the performance of convolution calculation by 1.18× while maintaining the minimum off-chip memory traffic. MobileNetV2 is implemented on Zynq UltraScale+ ZCU102 SoC FPGA, and the results show the proposed accelerator can achieve 381.7 frames per second (fps), which is 1.43× of the reference design, and it can save about 36.3% on-chip buffer size compared with the reference design, while maintaining the same off-chip memory traffic.
Baoting Li, Xuchong Zhang, Longjun Liu, Hongbin Sun 0001, Nanning Zheng 0001
IEEE Trans. Circuits Syst. I Regul. Pap.7
2021 Correntropy-Based Multiview Subspace Clustering
abstract
Multiview subspace clustering, which aims to cluster the given data points with information from multiple sources or features into their underlying subspaces, has a wide range of applications in the communities of data mining and pattern recognition. Compared with the single-view subspace clustering, it is challenging to efficiently learn the structure of the representation matrix from each view and make use of the extra information embedded in multiple views. To address the two problems, a novel correntropy-based multiview subspace clustering (CMVSC) method is proposed in this article. The objective function of our model mainly includes two parts. The first part utilizes the Frobenius norm to efficiently estimate the dense connections between the points lying in the same subspace instead of following the standard compressive sensing approach. In the second part, the correntropy-induced metric (CIM) is introduced to characterize the noise in each view and utilize the information embedded in different views from an information-theoretic perspective. Furthermore, an efficient iterative algorithm based on the half-quadratic technique (HQ) and the alternating direction method of multipliers (ADMM) is developed to optimize the proposed joint learning problem, and extensive experimental results on six real-world multiview benchmarks demonstrate that the proposed methods can outperform several state-of-the-art multiview subspace clustering methods.
Lei Xing 0003, Badong Chen, Shaoyi Du, Yuantao Gu, Nanning Zheng 0001
IEEE Trans. Cybern.5
2021 Predicting Task-Driven Attention via Integrating Bottom-Up Stimulus and Top-Down Guidance
abstract
Task-free attention has gained intensive interest in the computer vision community while relatively few works focus on task-driven attention (TDAttention). Thus this paper handles the problem of TDAttention prediction in daily scenarios where a human is doing a task. Motivated by the cognition mechanism that human attention allocation is jointly controlled by the top-down guidance and bottom-up stimulus, this paper proposes a cognitively-explanatory deep neural network model to predict TDAttention. Given an image sequence, bottom-up features, such as human pose and motion, are firstly extracted. At the same time, the coarse-grained task information and fine-grained task information are embedded as a top-down feature. The bottom-up features are then fused with the top-down feature to guide the model to predict TDAttention. Two public datasets are re-annotated to make them qualified for TDAttention prediction, and our model is widely compared with other models on the two datasets. In addition, some ablation studies are conducted to evaluate the individual modules in our model. Experiment results demonstrate the effectiveness of our model.
Zhixiong Nan, Jingjing Jiang, Xiaofeng Gao 0002, Sanping Zhou, Weiliang Zuo, Ping Wei 0001, Nanning Zheng 0001
IEEE Trans. Image Process.7
2021 Tracking Beyond Detection: Learning a Global Response Map for End-to-End Multi-Object Tracking
abstract
Most of the existing Multi-Object Tracking (MOT) approaches follow the Tracking-by-Detection and Data Association paradigm, in which objects are firstly detected and then associated in the tracking process. In recent years, deep neural network has been utilized to obtain more discriminative appearance features for cross-frame association, and noticeable performance improvement has been reported. On the other hand, the Tracking-by-Detection framework is yet not completely end-to-end, which leads to huge computation and limited performance especially in the inference (tracking) process. To address this problem, we present an effective end-to-end deep learning framework which can directly take image-sequence/video as input and output the located and tracked objects of learned types. Specifically, a novel global response network is learned to project multiple objects in the image-sequence/video into a continuous response map, and the trajectory of each tracked object can then be easily picked out. The overall process is similar to how a detector inputs an image and outputs the bounding boxes of each detected object. Experimental results based on the MOT16 and MOT17 benchmarks show that our proposed on-line tracker achieves state-of-the-art performance on several tracking metrics.
Xingyu Wan, Jiakai Cao, Sanping Zhou, Jinjun Wang, Nanning Zheng 0001
IEEE Trans. Image Process.5
2021 Giant Panda Identification
abstract
The lack of automatic tools to identify giant panda makes it hard to keep track of and manage giant pandas in wildlife conservation missions. In this paper, we introduce a new Giant Panda Identification (GPID) task, which aims to identify each individual panda based on an image. Though related to the human re-identification and animal classification problem, GPID is extraordinarily challenging due to subtle visual differences between pandas and cluttered global information. In this paper, we propose a new benchmark dataset iPanda-50 for GPID. The iPanda-50 consists of 6, 874 images from 50 giant panda individuals, and is collected from panda streaming videos. We also introduce a new Feature-Fusion Network with Patch Detector (FFN-PD) for GPID. The proposed FFN-PD exploits the patch detector to detect discriminative local patches without using any part annotations or extra location sub-networks, and builds a hierarchical representation by fusing both global and local features to enhance the inter-layer patch feature interactions. Specifically, an attentional cross-channel pooling is embedded in the proposed FFN-PD to improve the identify-specific patch detectors. Experiments performed on the iPanda-50 datasets demonstrate the proposed FFN-PD significantly outperforms competing methods. Besides, experiments on other fine-grained recognition datasets (i.e., CUB-200-2011, Stanford Cars, and FGVC-Aircraft) demonstrate that the proposed FFN-PD outperforms existing state-of-the-art methods.
Le Wang 0003, Rizhi Ding, Yuanhao Zhai 0001, Qilin Zhang 0004, Wei Tang 0016, Nanning Zheng 0001, Gang Hua 0001
IEEE Trans. Image Process.6
2021 Hierarchical and Interactive Refinement Network for Edge-Preserving Salient Object Detection
abstract
Salient object detection has undergone a very rapid development with the blooming of Deep Neural Network (DNN), which is usually taken as an important preprocessing procedure in various computer vision tasks. However, the down-sampling operations, such as pooling and striding, always make the final predictions blurred at edges, which has seriously degenerated the performance of salient object detection. In this paper, we propose a simple yet effective approach, i.e., Hierarchical and Interactive Refinement Network (HIRN), to preserve the edge structures in detecting salient objects. In particular, a novel multi-stage and dual-path network structure is designed to estimate the salient edges and regions from the low-level and high-level feature maps, respectively. As a result, the predicted regions will become more accurate by enhancing the weak responses at edges, while the predicted edges will become more semantic by suppressing the false positives in background. Once the salient maps of edges and regions are obtained at the output layers, a novel edge-guided inference algorithm is introduced to further filter the resulting regions along the predicted edges. Extensive experiments on several benchmark datasets have been conducted, in which the results show that our method significantly outperforms a variety of state-of-the-art approaches.
Sanping Zhou, Jinjun Wang, Le Wang 0003, Jimuyang Zhang, Fei Wang 0037, Nanning Zheng 0001
IEEE Trans. Image Process.7
2021 A Theoretical Foundation of Intelligence Testing and Its Application for Intelligent Vehicles
abstract
Intelligent vehicle testing received quickly increasing attention due to the intermittent accidents of intelligent vehicle prototypes that occurred recently. In this paper, we investigate the theoretical underpinnings of such testing and establish a rigid analyzing framework for general intelligence testing problems by borrowing the ideas of Probably Approximately Correct (PAC) learning. Our focus is on the relationship between the number of sampled scenarios and the testing efficiency. We explain various existing algorithms within this new framework and clarify some misconceptions about the reasoning underpinning these methods. We show that intelligent vehicles are testable if the testing scenarios are well defined and appropriately sampled. Moreover, we propose a sampling strategy to generate new challenging scenarios to boost testing efficiency.
Li Li 0013, Nanning Zheng 0001, Fei-Yue Wang 0001
IEEE Trans. Intell. Transp. Syst.2
2021 Object Cosegmentation in Noisy Videos With Multilevel Hypergraph
abstract
With the target of simultaneously segmenting semantically related videos to identify the common objects, video object cosegmentation has attracted the attention of researchers in recent years. Existing methods are primarily based on pair-wise relations between adjacent pixels and regions, which are susceptible to performance degradation from object entries/exists or occlusions. Specifically, we refer these video frames without the common objects present as the “empty” frames. In this paper, we propose a multilevel hypergraph-based full Video object CoSegmentation (VCS) method, which incorporates high-level semantics and low-level appearance/motion/saliency to construct the hyperedge among multiple spatially and temporally adjacent regions. Specifically, the high-level semantic model fuses multiple object proposals from each frame instead of relying on a single object proposal per frame. A hypergraph cut is subsequently utilized to calculate the object cosegmentation. Experiments on four video object segmentation/cosegmentation datasets against state-of-the-art methods with both objective and subjective results manifest the effectiveness of the proposed VCS method, including the SegTrack and VCoSeg datasets without “empty” frames, the XJTU-Stevens dataset with 3.7% “empty” frames, and the Noisy-ViCoSeg dataset proposed together with our method with 30.3% “empty” frames.
Le Wang 0003, Qilin Zhang 0004, Zhenxing Niu, Nanning Zheng 0001, Gang Hua 0001
IEEE Trans. Multim.5
2021 Minimum Error Entropy Kalman Filter
abstract
To date, most linear and nonlinear Kalman filters (KFs) have been developed under the Gaussian assumption and the well-known minimum mean square error (MMSE) criterion. In order to improve the robustness with respect to impulsive (or heavy-tailed) non-Gaussian noises, the maximum correntropy criterion (MCC) has recently been used to replace the MMSE criterion in developing several robust Kalman-type filters. To deal with more complicated non-Gaussian noises such as noises from multimodal distributions, in this article, we develop a new Kalman-type filter, called minimum error entropy KF (MEE-KF), by using the minimum error entropy (MEE) criterion instead of the MMSE or MCC. Similar to the MCC-based KFs, the proposed filter is also an online algorithm with the recursive process, in which the propagation equations are used to give prior estimates of the state and covariance matrix, and a fixed-point algorithm is used to update the posterior estimates. In addition, the MEE extended KF (MEE-EKF) is also developed for performance improvement in the nonlinear situations. The high accuracy and strong robustness of MEE-KF and MEE-EKF are confirmed by experimental results.
Badong Chen, Lujuan Dang, Yuantao Gu, Nanning Zheng 0001, José C. Príncipe
IEEE Trans. Syst. Man Cybern. Syst.4
2021 A Real-Time Robotic Grasping Approach With Oriented Anchor Box
abstract
Grasping is an essential skill for robots to interact with humans and the environment. In this paper, we build a vision-based, robust, and real-time robotic grasping approach with fully convolutional neural network. The main component of our approach is a grasp detection network with oriented anchor boxes as detection priors. Because the orientation of detected grasps is significant, which determines the rotation angle configuration of the gripper, we propose the orientation anchor box mechanism to regress grasp angle based on predefined assumption instead of classification or regression without any priors. With oriented anchor boxes, the grasps can be predicted more accurately and efficiently. Besides, to accelerate the network training and further improve the performance of angle regression, angle matching is proposed during training instead of Jaccard index matching. Fivefold cross-validation results demonstrate that our proposed algorithm achieves an accuracy of 98.8% and 97.8% in image-wise split and object-wise split, respectively, and the speed of our detection algorithm is 67 frames per second (FPS) with GTX 1080Ti, outperforming all the current state-of-the-art grasp detection algorithms on Cornell Dataset both in speed and accuracy. Robotic experiments demonstrate the robustness and generalization ability in unseen objects and real-world environment, with the average success rate of 90.0% and 84.2% of familiar things and unseen things, respectively, on Baxter robot platform.
Hanbo Zhang, Xinwen Zhou, Xuguang Lan, Jin Li 0011, Nanning Zheng 0001
IEEE Trans. Syst. Man Cybern. Syst.6
2021 Exploring Highly Dependable and Efficient Datacenter Power System Using Hybrid and Hierarchical Energy Buffers
abstract
The massive and irregular load surges challenge datacenter power infrastructures. As a result, power mismatching between supply and demand has emerged as a crucial availability issue in modern datacenters which are either under-provisioned or powered by intermittent power sources. Recent proposals have employed energy storage devices such as the uninterruptible power supply (UPS) to address this issue. However, current approaches lack the capacity of efficiently handling the irregular and unpredictable power mismatches. In this paper, we propose Hybrid and Hierarchical Energy Buffering (HHEB), a novel heterogeneous and adaptive scheme that could enable various energy storage devices (ESDs) to be efficiently integrated into existing datacenters for dynamically dealing with power mismatches. Our techniques exploit the diverse characteristics of different ESDs and intelligent load assignment algorithms to improve the dependability and efficiency of datacenter power systems. We evaluate the HHEB design with a prototype. Compared with a homogenous battery energy buffering system, HHEB could improve energy efficiency by 39.7 percent, extend UPS lifetime by 4.7X, promote energy availability by 3.2X, reduce system downtime by 41 percent, and effectively improve the energy availability of various energy buffers in different hierarchies. It allows datacenters to adapt to various power supply anomalies, thereby improving operational efficiency, dependability and availability.
Longjun Liu, Hongbin Sun 0001, Chao Li 0009, Tao Li 0006, Jingmin Xin, Nanning Zheng 0001
IEEE Trans. Sustain. Comput.6
2020 Designing Efficient Shortcut Architecture for Improving the Accuracy of Fully Quantized Neural Networks Accelerator
abstract
Network quantization is an effective solution to compress Deep Neural Networks (DNN) that can be accelerated with custom circuit. However, existing quantization methods suffer from significant loss in accuracy. In this paper, we propose an efficient shortcut architecture to enhance the representational capability of DNN between different convolution layers. We further implement the shortcut hardware architecture to effectively improve the accuracy of fully quantized neural networks accelerator. The experimental results show that our shortcut architecture can obviously improve network accuracy while increasing very few hardware resources ( 0.11 × and 0.17 × for LUT and FF respectively) compared with the whole accelerator.
Baoting Li, Longjun Liu, Yanming Jin, Hongbin Sun 0001, Nanning Zheng 0001
ASP-DAC6
2020 Lymph Node Metastasis Classification Based on Semi-Supervised Multi-View Network
abstract
Lymphatic metastasis is one of the most common proliferation pathways of thyroid carcinoma. Accurate diagnosis of lymph nodes is of great significance to surgical planning and prognosis. Due to the continuous development of deep learning recently, computer-aided diagnosis (CAD) systems for thyroid cancer have made considerable progress, but the research on the effective diagnosis of lymphatic metastasis remains insufficient. Focusing on this issue, we propose a semi-supervised multi-view network to diagnose lymph node metastasis, which combines coarse-view and fine-view to obtain a more comprehensive description. This method consists of three parts as follows: 1) joint probabilistic labels of the nodule partition information are generated by fuzzy clustering and perform semi-supervised learning on coarse-view with real labels; 2) an attention mechanism based network for fine-view is designed to capture various differentiated local features in a pyramid manner; 3) the two parts are then combined to extract global and local features more effectively to derive more accurate diagnostic reasoning. Especially, the introduction of fuzzy logic greatly reduces the impact of the uncertainty of the generated labels, thereby ensuring the effectiveness of the pseudo-labels. Extensive experiments on our collected dataset demonstrate that the proposed method is more efficient than other state-of-the-art methods.
Yiwen Luo, Jingmin Xin, Junqin Feng, Litao Ruan, Nanning Zheng 0001
BIBM7
2020 Loss Functions for Person Image Generation
Haoyue Shi 0002, Le Wang 0003, Wei Tang 0016, Nanning Zheng 0001, Gang Hua 0001
BMVC4
2020 Semantics-Guided Neural Networks for Efficient Skeleton-Based Human Action Recognition
abstract
Skeleton-based human action recognition has attracted great interest thanks to the easy accessibility of the human skeleton data. Recently, there is a trend of using very deep feedforward neural networks to model the 3D coordinates of joints without considering the computational efficiency. In this paper, we propose a simple yet effective semantics-guided neural network (SGN) for skeleton-based action recognition. We explicitly introduce the high level semantics of joints (joint type and frame index) into the network to enhance the feature representation capability. In addition, we exploit the relationship of joints hierarchically through two modules, i.e., a joint-level module for modeling the correlations of joints in the same frame and a framelevel module for modeling the dependencies of frames by taking the joints in the same frame as a whole. A strong baseline is proposed to facilitate the study of this field. With an order of magnitude smaller model size than most previous works, SGN achieves the state-of-the-art performance on the NTU60, NTU120, and SYSU datasets.
Pengfei Zhang 0005, Cuiling Lan, Wenjun Zeng 0001, Junliang Xing, Jianru Xue, Nanning Zheng 0001
CVPR6
2020 A Boundary Based Out-of-Distribution Classifier for Generalized Zero-Shot Learning
Xingyu Chen 0001, Xuguang Lan, Fuchun Sun 0001, Nanning Zheng 0001
ECCV (24)4
2020 COCOA: Content-Oriented Configurable Architecture Based on Highly-Adaptive Data Transmission Networks
abstract
In domain of parallel computation, most works focus on optimizing PE organization or memory hierarchy to pursue the maximum efficiency, while the importance of data contents has been overlooked for a long time. Actually for structured data, insights on data contents (i.e. values and locations within a structured form) can greatly benefit the computation performance, as fine-grained data manipulation can be performed. In this paper, we claim that by providing a flexible and adaptive data path, an efficient architecture with capability of fine-grained data manipulation can be built. Specifically, we propose COCOA, a novel content-oriented configurable architecture, which integrates multi-functional data reorganization networks in traditional computing scheme to handle the contents of data during the transmission path, so that they can be processed more efficiently. We evaluate COCOA on various problems: complex matrix algorithm (matrix inversion) and sparse DNN. The results indicates that COCOA is versatile enough to achieve high computation efficiency in both cases.
Tian Xia 0008, Pengchen Zong, Jianming Tong, Wenzhe Zhao 0001, Nanning Zheng 0001, Pengju Ren
ACM Great Lakes Symposium on VLSI6
2020 Fine-Grained Giant Panda Identification
abstract
The image-based fine-grained identification of individual giant pandas (Ailuropoda melanoleuca) is an emerging technology, and it is extraordinarily challenging due to the extremely subtle visual differences between individual giant pandas and limited annotated training data. To address these challenges, we propose the Feature-Fusion Convolutional Neural Network with Patch Detector (FFCNN-PD) algorithm, which exploits the discriminative local patches and builds a hierarchical representation generated by fusing both global and local features. Specifically, an attentional cross-channel pooling is embedded in the FFCNN-PD to improve the class- specific patch detectors. In addition, we propose a new giant panda identification dataset (iPanda-30) to establish a benchmark. Experiments on the proposed iPanda-30 dataset and other fine-grained recognition datasets demonstrate the effectiveness of the FFCNN-PD algorithm against the existing state-of-the-arts.
Rizhi Ding, Le Wang 0003, Qilin Zhang 0004, Zhenxing Niu, Nanning Zheng 0001, Gang Hua 0001
ICASSP5
2020 Exploring Better Speculation and Data Locality in Sparse Matrix-Vector Multiplication on Intel Xeon
abstract
Sparse Matrix-Vector Multiplication (SpMV) is a fundamental workload of numerous applications. However, for today's high-end superscalar CPUs, such as Intel Xeon series, it is usually difficult to efficiently perform SpMV due to the irregular, matrix-dependent data access and computation pattern. While many researches focus on optimizing the memory bandwidth bound by improving data locality, this work dives into the execution of SpMV computation on Intel Xeon CPU and reveals that the bad-speculation penalty is significant in many sparse matrices and too expensive to be ignored. We study and characterize sparsity structure types that are more vulnerable to the cache miss penalty or the bad speculation penalty, respectively. Based on this insight, we proposed a fast preprocessing method, which divides the matrix into sub-matrices and determines the critical performance bound of sub-matrices according to the data distribution characteristics. On each submatrix, a combination of dedicated row reordering strategies is performed to efficiently alleviate its key performance bounds: bad speculation, cache miss, or both. Our matrix representation is based on standard Compressed Sparse Row (CSR) format, and can be easily adapted to existing SpMV libraries. Our approach is evaluated on Intel Xeon Gold 6146 Processor with a wide-range of matrices from the SuiteSparse benchmarks. The results demonstrate that the proposed approach achieves an average 1.8× speedup (up to 2.5×) on multi-threaded MKL Sparse Routines, with a quite low pre-processing cost. Additionally, when used in conjunction with MKL's original optimization method, our approach can further prompt the speedup, to average 3.6 × (up to 8.3 ×), This result indicates that our method can serve as a fast and wide-spectrum optimization method which is compatible with existing routines.
Tian Xia 0008, Wenzhe Zhao 0001, Nanning Zheng 0001, Pengju Ren
ICCD5
2020 Local-Global Interactive Network For Face Age Transformation
abstract
Face age transformation aims to generate a face image in the past or future and has been receiving increasing attention for its significance recently. This paper proposes a local-global interactive framework for long-span face age transformation. We divide a face image into five local parts and design a local generative network for each of them to learn the local changes of a face image. Meanwhile a global generative network is utilized to learn the global changes. We introduce an interactive network and an age classification network, which are respectively used to integrate the local and global features and maintain the corresponding age features in different age groups. Given a face image at a certain age, our network can produce a realistic image of face aging or rejuvenation. We test and evaluate the model on complex datasets. Extensive qualitative comparison experiments prove the effectiveness and potential of our proposed method.
Ping Wei 0001, Yongchi Zhang, Nanning Zheng 0001
ICPR5
2020 Inferring Tasks and Fluents in Videos by Learning Causal Relations
abstract
Recognizing time-varying object states in complex tasks is an important and challenging issue. In this paper, we propose a novel model to jointly infer object fluents and complex tasks in videos. A task is a complex human activity with specific goals and a fluent is defined as a time-varying object state. A hierarchical graph represents a task as a human action stream and multiple concurrent object fluents which vary as the human performs the actions. In this process, the human actions serve as the causes of object state changes which conversely reflect the effects of human actions. For a given input video, a causal sampling search algorithm is proposed to jointly infer the task category and the states of objects in each video frame. For model learning, a structural SVM framework is adopted to jointly train the task, fluent, cause, and effect parameters. We test the proposed method on a task and fluent dataset. Experimental results demonstrate the effectiveness of the proposed method.
Haowen Tang, Ping Wei 0001, Nanning Zheng 0001
ICPR4
2020 A Duplex Spatiotemporal Filtering Network for Video-based Person Re-identification
abstract
Video-based person re-identification plays important roles in surveillance video analysis. This paper proposes a duplex spatiotemporal filtering network (DSFN) to re-identify persons in videos. A video sequence is represented as a duplex spatiotemporal matrix. DSFN model containing a group of filters performs filtering at feature levels in both temporal and spatial dimensions, by which the model focuses on feature-level semantic information rather than image-level information as in the traditional filters. We propose sparse-orthogonal constraints to enforce the model to extract more discriminative features. DSFN seeks to characterize not only the appearance features but also dynamic information such as gaits embedded in video sequences and obtains performance improvement as a result. Experimental results show that the proposed method outperforms other comparison approaches.
Chong Zheng, Ping Wei 0001, Nanning Zheng 0001
ICPR3
2020 Autonomous Tool Construction with Gated Graph Neural Network
abstract
Autonomous tool construction is a significant but challenging task in robotics. This task can be interpreted as when given a reference tool, selecting some available candidate parts to reconstruct it. Most of the existing works perform tool construction in the form of action part and grasp part, which is only a specific construction pattern and limits its application to some extent. In general scenarios, a tool can be constructed in various patterns with different part pairs. Therefore, whether a part pair is most suitable for constructing the tool depends not only on itself, but on other parts in the same scene. To solve this problem, we construct a Gated Graph Neural Network (GGNN) to model the relations between all part pairs, so that we can select the candidate parts in consideration of the global information. Afterwards, we embed the constructed GGNN into a RCNN-like structure to finally accomplish tool construction. The whole model will be named Tool Construction Graph RCNN (TC-GRCNN). In addition, we develop a mechanism that can generate large-scale training and testing data in simulation environments, by which we can save the time of data collection and annotation. Finally, the proposed model is deployed on the physical robot. The experiment results show that TC-GRCNN can perform well in the general scenarios of tool construction.
Chenjie Yang, Xuguang Lan, Hanbo Zhang, Nanning Zheng 0001
ICRA4
2020 HeLPS: Heterogeneous LiDAR-based Positioning System for Autonomous Vehicle
abstract
LiDAR-based positioning systems are widely used in unmanned systems. However, affected by the high computational complexity of high-precision positioning algorithms, the current positioning system is supported by hardware with low power efficiency and thus hard to integrate into many platforms. In this paper, we analyze features of the positioning system in autonomous driving application, design, and apply Heterogeneous LiDAR-based Positioning System, HeLPS, with software-hardware co-design methodology to achieve better efficiency. Our contributions can be concluded in three aspects. Firstly, we design the CPU-FPGA heterogeneous positioning system accelerating Iterative Closest Point (ICP) algorithm and achieves improvements on both speed and power-efficiency. Secondly, we exploit the spatial locality in the point cloud and design a new compressed data structure for fast neighbor accessing. The experiment reports a significant speedup comparing with other data structures. Lastly, we explore the data access pattern in positioning application and develop a specific cache system and out-of-order execution system reducing memory burden. Our system is deployed on a small and cheap Xilinx Zynq7000 ARM+FPGA platform, which achieves 983.3x speedup compared with Cortex-A9 CPU, and 31.8x speedup compared with i7-7820 CPU, with only 2.37W power consumption.
Yuedong Yang, Xiaodong Deng, Yanqing Shen, Shi-tao Chen, Nanning Zheng 0001
IECON6
2020 Robust Sparse Channel Estimation Based on Maximum Mixture Correntropy Criterion
abstract
The sparse channel estimation problem is drawing increasing attention in broadband wireless communication. Sparsity nature in structure and noise is one of the most important issues in such problems. Researchers have devised several sparsity penalty terms such as zero-attracting (ZA) and correntropy induced metric (CIM) to exploit potential sparse structure information and improve sparse channel estimate accuracy. To combat impulsive (sparse) noises, several adaptive filtering algorithms based on the maximum correntropy criterion (MCC) have been developed, which can achieve excellent performance, especially in heavy-tailed noises. Recently, the concept of mixture correntropy and the maximum mixture correntropy criterion (MMCC) were proposed to further promote the robustness of the MCC against impulsive noises. In this paper, a new robust and sparse adaptive filtering algorithm is developed to estimate the sparse channels with impulsive noises by combining the MMCC and CIM penalty. Thanks to the desirable property of the mixture correntropy, the proposed method behaves quite well with excellent convergence performance. Simulation results show that the new method can outperform several existing methods in sparse channel estimation including the MCC based methods.
Mingfei Lu, Lei Xing 0003, Nanning Zheng 0001, Badong Chen
IJCNN3
2020 A Deep Model for Joint Object Detection and Semantic Segmentation in Traffic Scenes
abstract
Object detection and semantic segmentation are two fundamental techniques of various applications in the fields of Intelligent Vehicles (IV) and Advanced Driving Assistance System (ADAS). Early studies separately handle these two problems. In this paper, inspired by some recent works, we propose a deep neural network model for joint object detection and semantic segmentation. Given an image, an encoder-decoder convolution network extracts a set of feature maps, these feature maps are shared by the detection branch and the segmentation branch to jointly carry out the object detection and semantic segmentation. In the detection branch, we design a PriorBox initialization mechanism to propose more object candidates. In the segmentation branch, we use the multi-scale atrous convolution to explore the global and local semantic information in traffic scenes. Benefiting from the PriorBox Initialization Mechanism (PBIM) and Multi-Scale Atrous Convolution (MSAC), our model presents the competitive performance. In the experiments, we widely compare with several recently-proposed methods on the public Cityscapes dataset, achieving the highest accuracy. In addition, to verify the robustness and generalization of our model, the extension experiments are also conducted on the well-known VOC2012 dataset.
Jizhi Peng, Zhixiong Nan, Linhai Xu, Jingmin Xin, Nanning Zheng 0001
IJCNN5
2020 Multiscale Adaptation Fusion Networks for Depth Completion
abstract
Depth completion is becoming a particularly important yet challenging problem with the growingly rapid progress of depth sensing technologies. Depth completion aims to complete sparse and noisy depth images to generate dense depth images. In this paper, we propose a multiscale adaptation fusion network (MAFN) for depth completion. The depth features are fused with RGB features at multiple scales with adaptation modules, where a neighbour attention mechanism is designed to adapt the local structures of the RGB image and the depth image. The fusion and completion process are unified under the encoder-decoder framework which is learned in an end-to-end way. By exploiting the detailed structural relationships of RGB images and depth images, our MAFN model can accurately complete and restore the invalid depth values on the sparse depth images. We test the proposed method on the challenging KITTI depth completion benchmark. The experimental results prove the effectiveness and strength of the proposed method.
Yongchi Zhang, Ping Wei 0001, Nanning Zheng 0001
IJCNN4
2020 CoBigICP: Robust and Precise Point Set Registration using Correntropy Metrics and Bidirectional Correspondence
abstract
In this paper, we propose a novel probabilistic variant of iterative closest point (ICP) dubbed as CoBigICP. The method leverages both local geometrical information and global noise characteristics. Locally, the 3D structure of both target and source clouds are incorporated into the objective function through bidirectional correspondence. Globally, error metric of correntropy is introduced as noise model to resist outliers. Importantly, the close resemblance between normal-distributions transform (NDT) and correntropy is revealed. To ease the minimization step, an on-manifold parameterization of the special Euclidean group is proposed. Extensive experiments validate that CoBigICP outperforms several well-known and state-of-the-art methods.
Pengyu Yin, Di Wang 0028, Shaoyi Du, Shihui Ying, Yue Gao 0002, Nanning Zheng 0001
IROS6
2020 Mixed Test Environment-based Vehicle-in-the-loop Validation - A New Testing Approach for Autonomous Vehicles
abstract
The current test of autonomous driving technology requires extensive experimental verification, whether in simulation or on real roads. Netherthless, how to test autonomous vehicles thoroughly in a safe and comprehensive manner remains a major challenge. To achieve safer and more effective autonomous driving testing, this paper proposes a novel mixed test environment-based validation method with vehicle in the loop(ViL) : (1) our method supports more realistic drive safety tests in mixed scenarios which integrate the synthetic and the real-world scenarios. Synthetic scenarios offer complex traffic simulation with diverse road conditions. The real-world scenarios introduce the real autonomous driving vehicle, the real sensor suite as well as the test field to the test loop, having further bridged the gap between Hardware-in-the-Loop(HiL) testing and real road tests than ViL; (2) virtual perceptional results are simulated directly and delivered to the real vehicle in Unified Fusion Data Format(UFDF), without rendering virtual detection data for reduced resource consumption; (3) diverse test scenarios are configurable and reproducible with OSM-based High Definition(HD) map, enabling the simulation to be decoupled from a specific test filed or traffic facilities. A series of experiments on the application of our method have been demonstrated, and our approach is proved to be a promising drive safety testing technique before actual road testing.
Yu Chen 0040, Shi-tao Chen, Tong Xiao 0006, Songyi Zhang, Qian Hou, Nanning Zheng 0001
IV6
2020 Using Detection, Tracking and Prediction in Visual SLAM to Achieve Real-time Semantic Mapping of Dynamic Scenarios
abstract
In this paper, we propose a lightweight system, RDS-SLAM, based on ORB-SLAM2, which can accurately estimate poses and build semantic maps at object level for dynamic scenarios in real time using only one commonly used Intel Core i7 CPU. In RDS-SLAM, three major improvements, as well as major architectural modifications, are proposed to overcome the limitations of ORB-SLAM2. Firstly, it adopts a lightweight object detection neural network in key frames. Secondly, an efficient tracking and prediction mechanism is embedded into the system to remove the feature points belonging to movable objects in all incoming frames. Thirdly, a semantic octree map is built by probabilistic fusion of detection and tracking results, which enables a robot to maintain a semantic description at object level for potential interactions in dynamic scenarios. We evaluate RDS-SLAM in TUM RGB-D dataset, and experimental results show that RDS-SLAM can run with 30.3 ms per frame in dynamic scenarios using only an Intel Core i7 CPU, and achieves comparable accuracy compared with the state-of-the-art SLAM systems which heavily rely on both Intel Core i7 CPUs and powerful GPUs.
Xingyu Chen 0001, Jianru Xue, Jianwu Fang, Yuxin Pan, Nanning Zheng 0001
IV5
2020 An Efficient Sampling-Based Hybrid A* Algorithm for Intelligent Vehicles
abstract
In this paper, we propose an improved sampling-based hybrid A* (SBA*) algorithm for path planning of intelligent vehicles, which works efficiently in complex urban environments. Two main modifications are introduced into the traditional hybrid A* algorithm to improve its adaptivity in both structured and unstructured traffic scenes. Firstly, a hybrid potential field (HPF) model considering both traffic regulation and obstacle configuration is proposed to represent the vehicle's workspace, which is utilized as a heuristic function. Secondly, a set of directional motion primitives is generated by taking the prior topological structure of the workspace into account. The path planner using SBA* not only obeys traffic regulations in structured scenes but also is capable of exploring complex unstructured scenes rapidly. Finally, a post-optimization step is adopted to increase the feasibility of the path. The efficacy of the proposed algorithm is extensively validated and tested with an autonomous vehicle in real traffic scenes. The experimental results show that SBA* works well in complex urban environments.
Gengxin Li, Jianru Xue, Di Wang 0028, Zhongxing Tao, Nanning Zheng 0001
IV7
2020 Modeling Methodology of Driver-Vehicle-Environment System Dynamics in Mixed Driving Situation
abstract
The interactions between driver, vehicle and environment generate vehicle's behaviors, and the interactions of vehicle groups shape the traffic modes. Therefore, the conditionally or fully automated driving technologies should be developed and tested in the driver-vehicle-environment (DVE) system rather than being developed and tested individually. To build DVE system dynamics, firstly, we propose an architecture to cope with the complicated interactions in automated vehicles (AVs) and in mixed traffic situations. Then we summarize the driver behavior models and compare the differences of intelligence between human driver and automated vehicle. Finally, we summarize the feasible modeling approaches of DVE system into five categories. The primary distinctions are the modeling methods of human drivers' roles, which are realized by human driver per se, human driver's cognitive architecture, psychological motivation model, mechanism imitation, and specific mechanism transfer respectively. Taking the applications of human machine interface and AD strategy developments as examples, we analyze the benefits and drawbacks of these approaches.
Shi-tao Chen, Nanning Zheng 0001, Jianqiang Wang 0003
IV3
2020 MuRF-Net: Multi-Receptive Field Pillars for 3D Object Detection from Point Cloud
abstract
In this paper, we propose a point cloud based 3D object detection framework that accounts for both contextual and local information by leveraging multi-receptive field pillars, named as MuRF-Net. Recently, common pipelines can be divided into a voxel-based feature encoder and an object detector. During the feature encoding steps, contextual information is neglected, which is critical for the 3D object detection task. Thus, the encoded features are not suitable to input to the subsequent object detector. To address this challenge, we propose the MuRF-Net with a multi-receptive field voxelization mechanism to capture both contextual and local information. After the voxelization, the voxelized points (pillars) are processed by a feature encoder, and a channel-wise feature reconfiguration module is proposed to combine the features with different receptive fields using a lateral enhanced fusion network. In addition, to handle the increase of memory and computational cost brought by multi-receptive field voxelization, a dynamic voxel encoder is applied taking advantage of the sparseness of the point cloud. Experiments on the KITTI benchmark for both 3D object and Bird's Eye View (BEV) detection tasks on car class are conducted and MuRF-Net achieved the state-of-the-art results compared with other voxel-based methods. Besides, the MuRF-Net can achieve nearly real-time speed with 20Hz.
Xinrui Yan, Shi-tao Chen, Jinpeng Dong, Ziyi Liu 0001, Zhixiong Nan, Nanning Zheng 0001
IV9
2020 Traffic Agent Trajectory Prediction Using Social Convolution and Attention Mechanism
abstract
The trajectory prediction is significant for the decision-making of autonomous driving vehicles. In this paper, we propose a model to predict the trajectories of target agents around an autonomous vehicle. The main idea of our method is considering the history trajectories of the target agent and the influence of surrounding agents on the target agent. To this end, we encode the target agent history trajectories as an attention mask and construct a social map to encode the interactive relationship between the target agent and its surrounding agents. Given a trajectory sequence, the LSTM networks are firstly utilized to extract the features for all agents, based on which the attention mask and social map are formed. Then, the attention mask and social map are fused to get the fusion feature map, which is processed by the social convolution to obtain a fusion feature representation. Finally, this fusion feature is taken as the input of a variable-length LSTM to predict the trajectory of the target agent. We note that the variable-length LSTM enables our model to handle the case that the number of agents in the sensing scope is highly dynamic in traffic scenes. To verify the effectiveness of our method, we widely compare with several methods on a public dataset, achieving a 20% error decrease. In addition, the model satisfies the real-time requirement with the 32 fps.
Tao Yang 0032, Zhixiong Nan, Shi-tao Chen, Nanning Zheng 0001
IV5
2020 Improving 3D Object Detection via Joint Attribute-oriented 3D Loss
abstract
3D object detection has become a hot topic in intelligent vehicle applications in recent years. Generally, deep learning has been the primary framework used in 3D object detection, and regression of the object location and classification of the objectness are the two indispensable components. In the process of training, the ℓn(n=1,2) and the focal loss are considered as the frequent solutions to minimize the regression and classification loss, respectively. However, there are two problems to be solved in the existing methods. For regression component, there is a gap between evaluation metrics, e.g., 3D Intersection over Union (IoU), and the traditional regression loss. As for the classification component, confidence score exists ambiguous due to the binary label assignment of target. To solve these problems, we propose a loss by jointing 3D IoU and other geometric attributes (named as jointed attribute-oriented 3D loss), which can be directly used in optimizing the regression component. In addition, the jointed attribute-oriented 3D loss can assign a soft label for supervising the training of the classification. By incorporating the proposed loss function into several state-of-the-art 3D object detection methods, the significant performance improvement has been achieved on the KITTI benchmark.
Jianru Xue, Jian Dou, Yuxin Pan, Jianwu Fang, Di Wang 0028, Nanning Zheng 0001
IV7
2020 A Driving Behavior Recognition Model with Bi-LSTM and Multi-Scale CNN
abstract
In autonomous driving, perceiving the driving behaviors of surrounding agents is important for the ego-vehicle to make a reasonable decision. In this paper, we propose a neural network model based on trajectories information for driving behavior recognition. Unlike existing trajectory-based methods that recognize the driving behavior using the handcrafted features or directly encoding the trajectory, our model involves a Multi-Scale Convolutional Neural Network (MSCNN) module to automatically extract the high-level features which are supposed to encode the rich spatial and temporal information. Given a trajectory sequence of an agent as the input, firstly, the Bi-directional Long Short Term Memory (Bi-LSTM) module and the MSCNN module respectively process the input, generating two features, and then the two features are fused to classify the behavior of the agent. We evaluate the proposed model on the public BLVD dataset, achieving a satisfying performance.
Zhixiong Nan, Tao Yang 0032, Yifan Liu 0014, Nanning Zheng 0001
IV5
2020 ATV Navigation in Complex and Unstructured Environment Containing Stairs
abstract
Self-driving and robotic technologies have been widely used in recent years. However, before L5 autonomy has been mature, it is more important and meaningful to apply these technologies to some wheeled robots on specific scenes. All-terrain vehicles (ATV) are widely used in urban search and rescue. An unstructured complex scene on which an ATV works usually contains many unusual obstacles, such as stairs. As a result, it is important for autonomous ATVs to detect, localize, and traverse stairs. In this paper, a real-time outdoor stair detection and localization method is proposed. A VLP-16 LIDAR is used to collect environment data, and the stairs are detected and localized using a single frame of LIDAR data by their geometric features, such as slope and parallel edges. A stair navigation strategy is also proposed in this paper. 3-axis attitudes measured by IMU are used during stair climbing on the basis of the coupling between roll and yaw on a slope. Experiments are carried out and the results prove the robustness and accuracy of detection and localization algorithm. The navigation strategy is proven to be safe and feasible.
Kongtao Zhu, Junxiang Zhan, Shi-tao Chen, Zhixiong Nan, Tangyike Zhang, Dantong Zhu, Nanning Zheng 0001
IV7
2020 A Slow-I-Fast-P Architecture for Compressed Video Action Recognition
abstract
Compressed video action recognition has drawn growing attention for the storage and processing advantages of compressed videos over original raw videos. While the past few years have witnessed remarkable progress in this problem, most existing approaches rely on RGB frames from raw videos and require multi-step training. In this paper, we propose a novel Slow-I-Fast-P (SIFP) neural network model for compressed video action recognition. It consists of the slow I pathway receiving a sparse sampling I-frame clip and the fast P pathway receiving a dense sampling pseudo optical flow clip. An unsupervised estimation method and a new loss function are designed to generate pseudo optical flows in compressed videos. Our model eliminates the dependence on the traditional optical flows calculated from raw videos. The model is trained in an end-to-end way. The proposed method is evaluated on the challenging HMDB51 and UCF101 datasets. The extensive comparison results and ablation studies demonstrate the effectiveness and strength of the proposed method.
Jiapeng Li 0003, Ping Wei 0001, Yongchi Zhang, Nanning Zheng 0001
ACM Multimedia4
2020 Action Co-localization in an Untrimmed Video by Graph Neural Networks
Changbo Zhai, Le Wang 0003, Qilin Zhang 0004, Zhanning Gao, Zhenxing Niu, Nanning Zheng 0001, Gang Hua 0001
MMM (1)6
2020 Compositional Generalization by Learning Analytical Expressions
abstract
Compositional generalization is a basic and essential intellective capability of human beings, which allows us to recombine known parts readily. However, existing neural network based models have been proven to be extremely deficient in such a capability. Inspired by work in cognition which argues compositionality can be captured by variable slots with symbolic functions, we present a refreshing view that connects a memory-augmented neural model with analytical expressions, to achieve compositional generalization. Our model consists of two cooperative neural modules, Composer and Solver, fitting well with the cognitive argument while being able to be trained in an end-to-end manner via a hierarchical reinforcement learning algorithm. Experiments on the well-known benchmark SCAN demonstrate that our model seizes a great ability of compositional generalization, solving all challenges addressed by previous works with 100% accuracies.
Qian Liu 0033, Shengnan An, Jian-Guang Lou, Bei Chen 0008, Zeqi Lin, Yan Gao 0002, Nanning Zheng 0001, Dongmei Zhang 0001
NeurIPS8
2020 Rethinking Learnable Tree Filter for Generic Feature Transform
abstract
The Learnable Tree Filter presents a remarkable approach to model structure-preserving relations for semantic segmentation. Nevertheless, the intrinsic geometric constraint forces it to focus on the regions with close spatial distance, hindering the effective long-range interactions. To relax the geometric constraint, we give the analysis by reformulating it as a Markov Random Field and introduce a learnable unary term. Besides, we propose a learnable spanning tree algorithm to replace the original non-differentiable one, which further improves the flexibility and robustness. With the above improvements, our method can better capture long range dependencies and preserve structural details with linear complexity, which is extended to several vision tasks for more generic feature transform. Extensive experiments on object detection/instance segmentation demonstrate the consistent improvements over the original version. For semantic segmentation, we achieve leading performance (82.1% mIoU) on the Cityscapes benchmark without bells-and whistles. Code is available at https://github.com/StevenGrove/LearnableTreeFilterV2.
Lin Song 0002, Zhengkai Jiang 0001, Xiangyu Zhang 0005, Hongbin Sun 0001, Jian Sun 0001, Nanning Zheng 0001
NeurIPS8
2020 Fine-Grained Dynamic Head for Object Detection
abstract
The Feature Pyramid Network (FPN) presents a remarkable approach to alleviate the scale variance in object representation by performing instance-level assignments. Nevertheless, this strategy ignores the distinct characteristics of different sub-regions in an instance. To this end, we propose a fine-grained dynamic head to conditionally select a pixel-level combination of FPN features from different scales for each instance, which further releases the ability of multi-scale feature representation. Moreover, we design a spatial gate with the new activation function to reduce computational complexity dramatically through spatially sparse convolutions. Extensive experiments demonstrate the effectiveness and efficiency of the proposed method on several state-of-the-art detection benchmarks. Code is available at https://github.com/StevenGrove/DynamicHead.
Lin Song 0002, Zhengkai Jiang 0001, Hongbin Sun 0001, Jian Sun 0001, Nanning Zheng 0001
NeurIPS7
2020 Spatiotemporal neural networks for action recognition based on joint loss
Chao Jing, Ping Wei 0001, Hongbin Sun 0001, Nanning Zheng 0001
Neural Comput. Appl.4
2020 Learning to infer human attention in daily activities
Zhixiong Nan, Tianmin Shu, Shu Wang 0002, Ping Wei 0001, Song-Chun Zhu, Nanning Zheng 0001
Pattern Recognit.7
2020 Robust and precise isotropic scaling registration algorithm using bi-directional distance and correntropy
Wenting Cui, Shaoyi Du, Teng Wan, Runzhao Yao, Yuying Liu 0007, Mengqi Han, Qingnan Mou, Yu-Cheng Guo, Nanning Zheng 0001
Pattern Recognit. Lett.9
2020 Visual manipulation relationship recognition in object-stacking scenes
Hanbo Zhang, Xuguang Lan, Xinwen Zhou, Nanning Zheng 0001
Pattern Recognit. Lett.6
2020 Algorithm and VLSI Architecture Co-Design on Efficient Semi-Global Stereo Matching
abstract
Semi-global matching (SGM) is favored for high accuracy real-time stereo matching design as it achieves a good trade-off between disparity image quality and computational complexity. Nevertheless, most of previous SGM designs so far are restricted to the real-time processing of small image resolution and disparity range, or achieve high throughput by simplifying the original algorithm at the penalty of significant disparity image quality degradation. We analyze that the major challenge to efficient SGM design is its memory architecture, including both on-chip memory cost and off-chip memory bandwidth. We address the memory architecture challenge by algorithm and architecture co-design. Based on two observed features of SGM algorithm, i.e. incompleteness and inaccuracy, this paper proposes several efficient techniques to reduce on-chip memory cost and compress off-chip memory bandwidth respectively. Moreover, we also design high throughput and pipelined architecture to implement the proposed techniques. The disparity image quality and hardware efficiency of the proposed SGM design are evaluated on both KITTI2015 and Middlebury V3 stereo datasets. Evaluation results demonstrate that, the throughput of the proposed circuit designs can easily achieve 1080P@30fps at the disparity range of 128, and can reduce the on-chip memory cost and off-chip memory bandwidth by up to 4× and 2× respectively while achieving better or the same disparity image quality, compared with the best reference design techniques.
Xuchong Zhang, He Dai, Hongbin Sun 0001, Nanning Zheng 0001
IEEE Trans. Circuits Syst. Video Technol.4
2020 Predicting COVID-19 in China Using Hybrid AI Model
abstract
The coronavirus disease 2019 (COVID-19) breaking out in late December 2019 is gradually being controlled in China, but it is still spreading rapidly in many other countries and regions worldwide. It is urgent to conduct prediction research on the development and spread of the epidemic. In this article, a hybrid artificial-intelligence (AI) model is proposed for COVID-19 prediction. First, as traditional epidemic models treat all individuals with coronavirus as having the same infection rate, an improved susceptible-infected (ISI) model is proposed to estimate the variety of the infection rates for analyzing the transmission laws and development trend. Second, considering the effects of prevention and control measures and the increase of the public's prevention awareness, the natural language processing (NLP) module and the long short-term memory (LSTM) network are embedded into the ISI model to build the hybrid AI model for COVID-19 prediction. The experimental results on the epidemic data of several typical provinces and cities in China show that individuals with coronavirus have a higher infection rate within the third to eighth days after they were infected, which is more in line with the actual transmission laws of the epidemic. Moreover, compared with the traditional epidemic models, the proposed hybrid AI model can significantly reduce the errors of the prediction results and obtain the mean absolute percentage errors (MAPEs) with 0.52%, 0.38%, 0.05%, and 0.86% for the next six days in Wuhan, Beijing, Shanghai, and countrywide, respectively.
Nanning Zheng 0001, Shaoyi Du, Jianji Wang 0001, Wenting Cui, Zijian Kang, Tao Yang 0032, Bin Lou, Yuting Chi, Hong Long, Mei Ma, Dong Zhang 0009, Jingmin Xin
IEEE Trans. Cybern.1
2020 A Joint Label Space for Generalized Zero-Shot Classification
abstract
The fundamental problem of Zero-Shot Learning (ZSL) is that the one-hot label space is discrete, which leads to a complete loss of the relationships between seen and unseen classes. Conventional approaches rely on using semantic auxiliary information, e.g. attributes, to re-encode each class so as to preserve the inter-class associations. However, existing learning algorithms only focus on unifying visual and semantic spaces without jointly considering the label space. More importantly, because the final classification is conducted in the label space through a compatibility function, the gap between attribute and label spaces leads to significant performance degradation. Therefore, this paper proposes a novel pathway that uses the label space to jointly reconcile visual and semantic spaces directly, which is named Attributing Label Space (ALS). In the training phase, one-hot labels of seen classes are directly used as prototypes in a common space, where both images and attributes are mapped. Since mappings can be optimized independently, the computational complexity is extremely low. In addition, the correlation between semantic attributes has less influence on visual embedding training because features are mapped into labels instead of attributes. In the testing phase, the discrete condition of label space is removed, and priori one-hot labels are used to denote seen classes and further compose labels of unseen classes. Therefore, the label space is very discriminative for the Generalized ZSL (GZSL), which is more reasonable and challenging for real-world applications. Extensive experiments on five benchmarks manifest improved performance over all of compared state-of-the-art methods.
Jin Li 0011, Xuguang Lan, Yang Long 0001, Yang Liu 0069, Xingyu Chen 0001, Ling Shao 0001, Nanning Zheng 0001
IEEE Trans. Image Process.7
2020 EleAtt-RNN: Adding Attentiveness to Neurons in Recurrent Neural Networks
abstract
Recurrent neural networks (RNNs) are capable of modeling temporal dependencies of complex sequential data. In general, current available structures of RNNs tend to concentrate on controlling the contributions of current and previous information. However, the exploration of different importance levels of different elements within an input vector is always ignored. We propose a simple yet effective Element-wise-Attention Gate (EleAttG), which can be easily added to an RNN block (e.g. all RNN neurons in an RNN layer), to empower the RNN neurons to have attentiveness capability. For an RNN block, an EleAttG is used for adaptively modulating the input by assigning different levels of importance, i.e., attention, to each element/dimension of the input. We refer to an RNN block equipped with an EleAttG as an EleAtt-RNN block. Instead of modulating the input as a whole, the EleAttG modulates the input at fine granularity, i.e., element-wise, and the modulation is content adaptive. The proposed EleAttG, as an additional fundamental unit, is general and can be applied to any RNN structures, e.g., standard RNN, Long Short-Term Memory (LSTM), or Gated Recurrent Unit (GRU). We demonstrate the effectiveness of the proposed EleAtt-RNN by applying it to different tasks including the action recognition, from both skeleton-based data and RGB videos, gesture recognition, and sequential MNIST classification. Experiments show that adding attentiveness through EleAttGs to RNN blocks significantly improves the power of RNNs.
Pengfei Zhang 0005, Jianru Xue, Cuiling Lan, Wenjun Zeng 0001, Zhanning Gao, Nanning Zheng 0001
IEEE Trans. Image Process.6
2020 Hierarchical U-Shape Attention Network for Salient Object Detection
abstract
Salient object detection aims at locating the most conspicuous objects in natural images, which usually acts as a very important pre-processing procedure in many computer vision tasks. In this paper, we propose a simple yet effective Hierarchical U-shape Attention Network (HUAN) to learn a robust mapping function for salient object detection. Firstly, a novel attention mechanism is formulated to improve the well-known U-shape network [1], in which the memory consumption can be extensively reduced and the mask quality can be significantly improved by the resulting U-shape Attention Network (UAN). Secondly, a novel hierarchical structure is constructed to well bridge the low-level and high-level feature representations between different UANs, in which both the intra-network and inter-network connections are considered to explore the salient patterns from a local to global view. Thirdly, a novel Mask Fusion Network (MFN) is designed to fuse the intermediate prediction results, so as to generate a salient mask which is in higher-quality than any of those inputs. Our HUAN can be trained together with any backbone network in an end-to-end manner, and high-quality masks can be finally learned to represent the salient objects. Extensive experimental results on several benchmark datasets show that our method significantly outperforms most of the state-of-the-art approaches.
Sanping Zhou, Jinjun Wang, Jimuyang Zhang, Le Wang 0003, Shaoyi Du, Nanning Zheng 0001
IEEE Trans. Image Process.7
2020 Probability Density Rank-Based Quantization for Convex Universal Learning Machines
abstract
The distributions of input data are very important for learning machines, such as the convex universal learning machines (CULMs). The CULMs are a family of universal learning machines with convex optimization. However, the computational complexity is a crucial problem in CULMs, because the dimension of the nonlinear mapping layer (the hidden layer) of the CULMs is usually rather large in complex system modeling. In this article, we propose an efficient quantization method called Probability density Rank-based Quantization (PRQ) to decrease the computational complexity of CULMs. The PRQ ranks the data according to the estimated probability densities and then selects a subset whose elements are equally spaced in the ranked data sequence. We apply the PRQ to kernel ridge regression (KRR) and random Fourier feature recursive least squares (RFF-RLS), which are two typical algorithms of CULMs. The proposed method not only keeps the similarity of data distribution between the code book and data set but also reduces the computational cost by using the kd-tree. Meanwhile, for a given data set, the method yields deterministic quantization results, and it can also exclude the outliers and avoid too many borders in the code book. This brings great convenience to practical applications of the CULMs. The proposed PRQ is evaluated on several real-world benchmark data sets. Experimental results show satisfactory performance of PRQ compared with some state-of-the-art methods.
Zhengda Qin, Badong Chen, Yuantao Gu, Nanning Zheng 0001, José C. Príncipe
IEEE Trans. Neural Networks Learn. Syst.4
2019 Video Imprint Segmentation for Temporal Action Detection in Untrimmed Videos
abstract
We propose a temporal action detection by spatial segmentation framework, which simultaneously categorize actions and temporally localize action instances in untrimmed videos. The core idea is the conversion of temporal detection task into a spatial semantic segmentation task. Firstly, the video imprint representation is employed to capture the spatial/temporal interdependences within/among frames and represent them as spatial proximity in a feature space. Subsequently, the obtained imprint representation is spatially segmented by a fully convolutional network. With such segmentation labels projected back to the video space, both temporal action boundary localization and per-frame spatial annotation can be obtained simultaneously. The proposed framework is robust to variable lengths of untrimmed videos, due to the underlying fixed-size imprint representations. The efficacy of the framework is validated in two public action detection datasets.
Zhanning Gao, Le Wang 0003, Qilin Zhang 0004, Zhenxing Niu, Nanning Zheng 0001, Gang Hua 0001
AAAI5
2019 Recognizing Unseen Attribute-Object Pair with Generative Model
abstract
In this paper, we are studying the problem of recognizing attribute-object pairs that do not appear in the training dataset, which is called unseen attribute-object pair recognition. Existing methods mainly learn a discriminative classifier or compose multiple classifiers to tackle this problem, which exhibit poor performance for unseen pairs. The key reasons for this failure are 1) they have not learned an intrinsic attributeobject representation, and 2) the attribute and object are processed either separately or equally so that the inner relation between the attribute and object has not been explored. To explore the inner relation of attribute and object as well as the intrinsic attribute-object representation, we propose a generative model with the encoder-decoder mechanism that bridges visual and linguistic information in a unified end-to-end network. The encoder-decoder mechanism presents the impressive potential to find an intrinsic attribute-object feature representation. In addition, combining visual and linguistic features in a unified model allows to mine the relation of attribute and object. We conducted extensive experiments to compare our method with several state-of-the-art methods on two challenging datasets. The results show that our method outperforms all other methods.
Zhixiong Nan, Yang Liu 0266, Nanning Zheng 0001, Song-Chun Zhu
AAAI3
2019 Object Affordances Graph Network for Action Recognition
Haoliang Tan, Le Wang 0003, Qilin Zhang 0004, Zhanning Gao, Nanning Zheng 0001, Gang Hua 0001
BMVC5
2019 Compressing Unknown Images With Product Quantizer for Efficient Zero-Shot Classification
abstract
For Zero-Shot Learning (ZSL), the Nearest Neighbor (NN) search is generally conducted for classification, which may cause unacceptable computational complexity for large-scale datasets. To compress zero-shot classes by the trained quantizer for efficient search, it tends to induce large quantization error because distributions between seen and unseen classes are different. However, as semantic attributes of classes are available in ZSL, both seen and unseen classes have the same distribution for one specific property, e.g., animals have or not have spots. Based on this intuition, a Product Quantization Zero-Shot Learning (PQZSL) method is proposed to learn embeddings as well as quantizers to compress visual features into compact codes for Approximate NN (ANN) search. Particularly, visual features are projected into an orthogonal semantic space, and then the Product Quantization (PQ) is utilized to quantize individual properties. Experimental results on five benchmark datasets demonstrate that unseen classes are represented by the Cartesian product of quantized properties with little quantization error. As classes in orthogonal common space are more discriminative, the classification based on PQZSL achieves state-of-the-art performance in Generalized Zero-Shot Learning (GZSL) task, meanwhile, the speed of ANN search is 10-100 times higher than traditional NN search.
Jin Li 0011, Xuguang Lan, Yang Liu 0069, Le Wang 0003, Nanning Zheng 0001
CVPR5
2019 SR-LSTM: State Refinement for LSTM Towards Pedestrian Trajectory Prediction
abstract
In crowd scenarios, reliable trajectory prediction of pedestrians requires insightful understanding of their social behaviors. These behaviors have been well investigated by plenty of studies, while it is hard to be fully expressed by hand-craft rules. Recent studies based on LSTM networks have shown great ability to learn social behaviors. However, many of these methods rely on previous neighboring hidden states but ignore the important current intention of the neighbors. In order to address this issue, we propose a data-driven state refinement module for LSTM network (SR-LSTM), which activates the utilization of the current intention of neighbors, and jointly and iteratively refines the current states of all participants in the crowd through a message passing mechanism. To effectively extract the social effect of neighbors, we further introduce a social-aware information selection mechanism consisting of an element-wise motion gate and a pedestrian-wise attention to select useful message from neighboring pedestrians. Experimental results on two public datasets, i.e. ETH and UCY, demonstrate the effectiveness of our proposed SR-LSTM and we achieve state-of-the-art results.
Pu Zhang 0001, Wanli Ouyang, Pengfei Zhang 0005, Jianru Xue, Nanning Zheng 0001
CVPR5
2019 Weakly Supervised Temporal Action Localization Through Contrast Based Evaluation Networks
abstract
Weakly-supervised temporal action localization (WS-TAL) is a promising but challenging task with only video-level action categorical labels available during training. Without requiring temporal action boundary annotations in training data, WS-TAL could possibly exploit automatically retrieved video tags as video-level labels. However, such coarse video-level supervision inevitably incurs confusions, especially in untrimmed videos containing multiple action instances. To address this challenge, we propose the Contrast-based Localization EvaluAtioN Network (CleanNet) with our new action proposal evaluator, which provides pseudo-supervision by leveraging the temporal contrast in snippet-level action classification predictions. Essentially, the new action proposal evaluator enforces an additional temporal contrast constraint so that high-evaluation-score action proposals are more likely to coincide with true action instances. Moreover, the new action localization module is an integral part of CleanNet which enables end-to-end training. This is in contrast to many existing WS-TAL methods where action localization is merely a post-processing step. Experiments on THUMOS14 and ActivityNet datasets validate the efficacy of CleanNet against existing state-ofthe- art WS-TAL algorithms.
Ziyi Liu 0001, Le Wang 0003, Qilin Zhang 0004, Zhanning Gao, Zhenxing Niu, Nanning Zheng 0001, Gang Hua 0001
ICCV6
2019 Preference Relationship-Based CrossCMN Scheme for Answer Ranking in Community QA
abstract
Community question answering (CQA) systems aim to provide users with high-quality answers. Nevertheless, unreliable answers are often returned to users in CQA systems, and the phenomenon causes that users have to browse multiple answers to find the best one. To improve such problem, we design a novel scheme, named PW-CrossCMN. The scheme ranks the candidate answers by pair-wise approach based on numerous historical documents. In the scheme, we apply the preference relationship into deep learning framework. Specifically, the scheme consists of two phases. In phase 1, the scheme extracts the features via automated feature engineering to construct the preference vectors and then divides the vectors into balanced positive and negative training samples based on the preference relationship. In phase 2, we build the CrossCMN model, which implements the multi-network parallel convolution and the cross forward propagation of full-connected layers, to achieve training and prediction tasks. Moreover, the multi-layer perception (MLP) is introduced to extract combination features in the prediction module. We perform extensive experiments on two typical datasets, and the results show that our scheme has more excellent performance in answer ranking task compared with several state-of-the-art baselines. In addition, we have released the relevant codes.
Jianji Wang 0001, Xuguang Lan, Nanning Zheng 0001
ICDM4
2019 Exploring Hardware Friendly Bottleneck Architecture in CNN for Embedded Computing Systems
abstract
In this paper, we explore how to design lightweight CNN architecture for embedded computing systems. We propose L-Mobilenet model for ZYNQ based hardware platform. L-Mobilenet can adapt well to hardware computing and accelerating, and its network structure is inspired by the state-of-the-art work of Inception-Resnet and Mobilenet-V2, which can effectively reduce parameters and delay while maintaining the accuracy of inference. We deploy our L-Mobilenet model to GPU and ZYNQ embedded platform for fully evaluating the performance of our design. By measuring with cifar10 and cifar100 datasets, L-Mobilenet model is able to gain 3× speed up and 3.7× fewer parameters than MobileNet-V2 while maintaining a similar accuracy. It also can obtain 2× speed up and 1.5× fewer parameters than Shufflenet-V2 while maintaining the same accuracy. Experiments show that our network model can obtain better performance because of the special considerations for hardware accelerating and software-hardware co-design strategies in our L-Mobilenet bottleneck architecture.
Xing Lei, Longjun Liu, Hongbin Sun 0001, Nanning Zheng 0001
ICIP5
2019 Har Enhanced Weakly-Supervised Semantic Segmentation Coupled with Adversarial Learning
abstract
Semantic segmentation is a challenging computer visual task which needs enormous pixel-level annotation data. But collecting a large amount of pixel-level annotation data is labor intensive. To address this issue, our work focuses on weakly-supervised learning approach which combines the adversarial learning and localization ability of classification model together, in this way, data with different annotations can be fully utilized. Specifically, the adversarial learning encourages the high order spatial consistences thus offers a relatively reliable initial confidence map. And we find that the hybrid atrous rate (HAR) can improve the localization ability of the classification model, thus indicate more precise object-related regions, which serves as strong supervision information. We conduct experiments with different settings to demonstrate the effectiveness of this weakly-supervised learning approach. The results show that our approach can improve the performance of baseline adversarial learning from 73.2 to 75.1 (mIOU), which is pretty effective.
Leiyuan Ma, Ziyi Liu 0001, Nanning Zheng 0001, Jianji Wang 0001
ICIP3
2019 Action Coherence Network for Weakly Supervised Temporal Action Localization
abstract
Most prominent temporal action localization methods are of the fully-supervised type, which rely heavily on frame-level labels, which could be prohibitively expensive to annotate. Thanks to recent developments on the Weakly-supervised Temporal Action Localization (W-TAL), this alternative paradigm requires only video-level labels in training, alleviating such annotation efforts. Specifically, we present Action Coherence Network (ACN) for W-TAL, which features a new coherence loss that better supervises action boundary learning and facilitate proposal regression. In addition, a purpose-built fusion module is proposed for localization inference based on features extracted by two streams of convolutional neural network. Overall, the proposed ACN achieves state-of-the-art W-TAL performance on two challenging datasets (THU-MOS14 and ActivityNet1.2, particularly ACN attains mAP of 24.2% on THUMOS14 under IoU threshold 0.5), which is approaching some recent fully-supervised TAL methods.
Yuanhao Zhai 0001, Le Wang 0003, Ziyi Liu 0001, Qilin Zhang 0004, Gang Hua 0001, Nanning Zheng 0001
ICIP6
2019 Face Age Transformation with Progressive Residual Adversarial Autoencoder
abstract
Face age transformation is an important issue in many applications. While unidirectional and short-span face ageing has achieved remarkable progress, it remains a challenging problem to generate both younger-look and older-look face images over long age span. In this paper, we present a progressive residual adversarial autoencoder (PRAA) model for bidirectional and long-span face age transformation. Given an input face image, our model aims to synthesize face images of its younger looks (face rejuvenation) and older looks (face ageing). The PRAA contains adversarial generators and discriminators where age information is encoded as latent features. It adopts a residual-encoding strategy by which the original face images and the residual face images are jointly encoded. We adopt a progressive multi-scale method to train the network, by which our model can capture both the global structure and the local detail changes in face age transformation. We test our model on challenging data and the experimental results prove the strength of our method.
Xuexiang Zhang, Ping Wei 0001, Nanning Zheng 0001
IJCNN3
2019 Precise Correntropy-based 3D Object Modelling With Geometrical Traffic Prior
abstract
Robust 3D perception using LiDAR is of prime importance for robotics, and its fundamental core lies in precise object modelling resisting to noise and outliers. In this paper, a precise 3D object modelling algorithm is designed especially for the intelligent vehicles. The proposed algorithm is advantageous by leveraging the crucial traffic geometrical prior of road surface profile, and both the noise and outliers are elegantly handled by robust correntropy-based metric. More specifically, the road surface correction (RSC) method transforms each individual LiDAR measurement from its locally planar road surface to a globally ideal plane. This procedure essentially guarantees the reduction of vehicle's motion from arbitrary 3D motion to physically feasible 2D motion. To deal with the noise and outliers, a correntropy-based multi-frame matching (CorrMM) algorithm is proposed which has a robust objective function with respect to point-to-plane residual error. An efficient solver inspired by M-estimator and retraction technique on Lie group is developed, which elegantly converts the optimization of highly non-linear objective function into a simple quadratic programming (QP) problem. Extensive experimental results validate that the proposed algorithm attains more crisper 3D object models than several state-of-the-art algorithms on a challenging real traffic dataset.
Di Wang 0028, Jianru Xue, Yinghan Jin, Nanning Zheng 0001, Masayoshi Tomizuka
IROS5
2019 Task-oriented Grasping in Object Stacking Scenes with CRF-based Semantic Model
abstract
In task-oriented grasping, the robot is supposed to manipulate the objects in a task-compatible manner, which is more important but more challenging than just stably grasping. However, most of existing works perform task-oriented grasping only in single object scenes. This greatly limits their practical application in real world scenes, in which there are usually multiple stacked objects with serious overlaps and occlusions. To perform task-oriented grasping in object stacking scenes, in this paper, we firstly build a synthetic dataset named Object Stacking Grasping Dataset (OSGD) for task-oriented grasping in object stacking scenes. Secondly, a Conditional Random Field (CRF) is constructed to model the semantic contents in object regions. The modelled semantic contents can be illustrated as incompatibility of task labels and continuity of task regions. This proposed approach can greatly reduce the interference of overlaps and occlusions in object stacking scenes. To embed the CRF-based semantic model into our grasp detection network, we implement the inference process of CRFs as a RNN so that the whole model, Task-oriented Grasping CRFs (TOG-CRFs) can be trained end to end. Finally, in object stacking scenes, the constructed model can help robot achieve 69.4% success rate for task-oriented grasping.
Chenjie Yang, Xuguang Lan, Hanbo Zhang, Nanning Zheng 0001
IROS4
2019 A Multi-task Convolutional Neural Network for Autonomous Robotic Grasping in Object Stacking Scenes
abstract
Autonomous robotic grasping plays an important role in intelligent robotics. However, how to help the robot grasp specific objects in object stacking scenes is still an open problem, because there are two main challenges for autonomous robots: (1) it is a comprehensive task to know what and how to grasp; (2) it is hard to deal with the situations in which the target is hidden or covered by other objects. In this paper, we propose a multi-task convolutional neural network for autonomous robotic grasping, which can help the robot find the target, make the plan for grasping and finally grasp the target step by step in object stacking scenes. We integrate vision-based robotic grasping detection and visual manipulation relationship reasoning in one single deep network and build the autonomous robotic grasping system. Experimental results demonstrate that with our model, Baxter robot can autonomously grasp the target with a success rate of 90.6%, 71.9% and 59.4% in object cluttered scenes, familiar stacking scenes and complex stacking scenes respectively.
Hanbo Zhang, Xuguang Lan, Cedar Site Bai, Lipeng Wan 0003, Chenjie Yang, Nanning Zheng 0001
IROS6
2019 ROI-based Robotic Grasp Detection for Object Overlapping Scenes
abstract
Grasp detection considering the affiliations between grasps and their owner in object overlapping scenes is a necessary and challenging task for the practical use of the robotic grasping approach. In this paper, a robotic grasp detection algorithm named ROI-GD is proposed to provide a feasible solution to this problem based on Region of Interest (ROI), which is the region proposal for objects. ROI-GD uses features from ROIs to detect grasps instead of the whole scene. It has two stages: the first stage is to provide ROIs in the input image and the second-stage is the grasp detector based on ROI features. We also contribute a multi-object grasp dataset, (a) which is much larger than Cornell Grasp Dataset, by labeling Visual Manipulation Relationship Dataset. Experimental results demonstrate that ROI-GD performs much better in object overlapping scenes and at the meantime, remains comparable with state-of-the-art grasp detection algorithms on Cornell Grasp Dataset and Jacquard Dataset. Robotic experiments demonstrate that ROI-GD can help robots grasp the target in single-object and multi-object scenes with the overall success rates of 92.5% and 83.8% respectively.
Hanbo Zhang, Xuguang Lan, Cedar Site Bai, Xinwen Zhou, Nanning Zheng 0001
IROS6
2019 REcache: Efficient Sustainable Energy Management Circuits and Policies for Computing Systems
abstract
The rapidly growing computing systems, such as AI server cluster, IoT devices etc. are facing increasing energy expenditure pressure and the warning of carbon footprint. Designing eco-friendly computing systems which integrated renewable energy sources have attracted considerable attentions recently. Existing schemes either incur green energy efficiency degradation or sacrifice workload performance. This paper proposes REcache (Renewable Energy cache), a sustainable energy management scheme to efficiently utilize green energy for computing systems. Compared to previous proposals, we present a dedicated circuit and energy-aware management policies to coordinate energy harvesting, power management and workload scheduling. We evaluate our scheme through both prototyping and simulation. The experimental results show that the REcache could effectively improve the energy availability 10%, workload performance 5% for different workloads on average.
Longjun Liu, Hongbin Sun 0001, Nanning Zheng 0001, Tao Li 0006
ISCAS4
2019 A Hardware-Efficient Post-Processing Algorithm for Motion Compensated Frame Rate Up-Conversion
abstract
Post-processing is an important module in motion compensated frame rate up-conversion design, and has a direct impact on the image quality of interpolated frame. However, how to balance between image quality and computational efficiency is still very challenging for post-processing, especially in hardware design. This paper proposes a hardware-efficient post-processing algorithm which leverages the temporal and spatial constraints to locally refine interpolated pixels. Moreover, we employ the quantization and approximation techniques to further reduce the computational intensity of the proposed post-processing algorithm. The quality of interpolated frame has been greatly improved both in objective and subjective aspects. The proposed post-processing algorithm is extensively evaluated by a set of video test sequences. Evaluation results demonstrate that, compared with the reference designs, the proposed algorithm can improve PSNR by at least 1.69 dB with comparable computational complexity.
Yunqi Mi, Hongbin Sun 0001, Nanning Zheng 0001
ISCAS5
2019 High-Definition Map Combined Local Motion Planning and Obstacle Avoidance for Autonomous Driving
abstract
Local motion planning plays an important role in an autonomous driving system. And applying mature local motion planning methods to real traffic scenarios with regular constraints is one of the keys to the applications of autonomous vehicles. In this paper, we present a local motion planning method combined with High-Definition (HD) maps. Through the HD map defined by OpenStreetMap, the local motion planner can obtain the prior knowledge of traffic scenarios and achieve path planning and optimization accordingly. In order to improve the safety and comfort of the obstacle avoiding process, we also propose an inertia-like path selection algorithm based on this planning method. We evaluated the proposed method on our designed autonomous driving experimental platform `Pioneer' and participated in the 2018 Intelligent Vehicles Future Challenge. In the competition, the `Pioneer' successfully completed all the races and won the championship without any manual intervention.
Zhiqiang Jian, Songyi Zhang, Shi-tao Chen, Nanning Zheng 0001
IV5
2019 Scene-Guided Region Proposal Re-ranking Method for On-road Vehicle Candidate Generation
abstract
Vehicle candidate generation is important for vehicle detection. Existing vehicle detection studies usually employ general-purpose region proposal methods to generate vehicle candidates, which do not consider the specificity of on-road vehicles in traffic scenes. In this paper, we propose a model to re-rank the candidates that are generated by general-purpose region proposal methods. Our model considers the specificity of on-road vehicle candidate generation in traffic scenes by encoding global-local semantic context and location-size geometric compatibility. In the experiments, we test our model on three art-of-the-state region proposal methods using two public datasets. The results show the significant performance improvement is gained after applying our model.
Zhixiong Nan, Jiawei He 0002, Ping Wei 0001, Linhai Xu, Hongbin Sun 0001, Nanning Zheng 0001
IV7
2019 Robust Extrinsic Parameter Calibration of 3D LIDAR Using Lie Algebras
abstract
In the field of autonomous driving, multi-beam light detection and ranging (3D LIDAR) system and global navigation satellite system/integrated inertial navigation system (GNSS/INS) are widely used in high-definition map construction, localization and obstacle detection. As 3D LIDAR system and INS have their own coordinate systems, the calibration of the two mentioned systems is required. In this paper, a novel algorithm for calibrating the coordinate system of 3D LIDAR and INS is proposed, which consists of three parts. The first procedure is to project two point clouds to the world coordinate system based on the initial transform matrix between 3D LIDAR and INS with the real-time data from INS. Then optimal point-to-point correspondences can be found between two frames of point cloud data through registration method. Finally, the loss function is constructed with the sum of the Euclidean distances of the corresponding points and optimized by using perturbation model of Lie algebras, so as to obtain the optimal transform matrix. With different given initial calibration parameters, test results of both simulation and real experiments validate the proposed algorithm and quantify its accuracy and robustness.
Yanqing Shen, Tangyike Zhang, Songyi Zhang, Yongbo Huo, Shi-tao Chen, Nanning Zheng 0001
IV7
2019 Learnable Tree Filter for Structure-preserving Feature Transform
abstract
Learning discriminative global features plays a vital role in semantic segmentation. And most of the existing methods adopt stacks of local convolutions or non-local blocks to capture long-range context. However, due to the absence of spatial structure preservation, these operators ignore the object details when enlarging receptive fields. In this paper, we propose the learnable tree filter to form a generic tree filtering module that leverages the structural property of minimal spanning tree to model long-range dependencies while preserving the details. Furthermore, we propose a highly efficient linear-time algorithm to reduce resource consumption. Thus, the designed modules can be plugged into existing deep neural networks conveniently. To this end, tree filtering modules are embedded to formulate a unified framework for semantic segmentation. We conduct extensive ablation studies to elaborate on the effectiveness and efficiency of the proposed method. Specifically, it attains better performance with much less overhead compared with the classic PSP block and Non-local operation under the same backbone. Our approach is proved to achieve consistent improvements on several benchmarks without bells-and-whistles. Code and models are available at https://github.com/StevenGrove/TreeFilter-Torch.
Lin Song 0002, Gang Yu 0002, Hongbin Sun 0001, Jian Sun 0001, Nanning Zheng 0001
NeurIPS7
2019 Autonomous driving: cognitive construction and situation understanding
Shi-tao Chen, Zhiqiang Jian, Yu Chen 0040, Zhuoli Zhou, Nanning Zheng 0001
Sci. China Inf. Sci.6
2019 Guest Editorial Special Issue on IoT on the Move: Enabling Technologies and Driving Applications for Internet of Intelligent Vehicles (IoIV)
abstract
The new era of the Internet of Things (IoT) is prompting the evolution of conventional vehicle ad-hoc networks (VANETs) into the Internet of Intelligent Vehicles (IoIV). Different from VANETs, where a vehicle is essentially considered as a node disseminating messages, the emerging IoIV paradigm is expected to regard each vehicle as a smart object equipped with a powerful multisensor platform, unprecedented communication capability, computing units, and Internet protocol (IP)-based connectivity. As such, the vehicles in IoIV are highly efficient in a broad array of vehicular and transportation applications. As a unique subset of general purpose IoT, IoIV can benefit from the existing research on VANET, which lays the foundation toward a more pervasive and ubiquitous communications and networking core that is essential for IoIV. Nevertheless, research in many aspects of IoIV, especially those that are application-driven and data-oriented ones, is still at its infancy.
Liuqing Yang 0001, Xiang Cheng 0001, Mounir Ghogho, Ender Ayanoglu, Tiejun Huang 0001, Nanning Zheng 0001
IEEE Internet Things J.6
2019 Video Imprint
abstract
A new unified video analytics framework (ER3) is proposed for complex event retrieval, recognition and recounting, based on the proposed video imprint representation, which exploits temporal correlations among image features across video frames. With the video imprint representation, it is convenient to reverse map back to both temporal and spatial locations in video frames, allowing for both key frame identification and key areas localization within each frame. In the proposed framework, a dedicated feature alignment module is incorporated for redundancy removal across frames to produce the tensor representation, i.e., the video imprint. Subsequently, the video imprint is individually fed into both a reasoning network and a feature aggregation module, for event recognition/recounting and event retrieval tasks, respectively. Thanks to its attention mechanism inspired by the memory networks used in language modeling, the proposed reasoning network is capable of simultaneous event category recognition and localization of the key pieces of evidence for event recounting. In addition, the latent structure in our reasoning network highlights the areas of the video imprint, which can be directly used for event recounting. With the event retrieval task, the compact video representation aggregated from the video imprint contributes to better retrieval results than existing state-of-the-art methods.
Zhanning Gao, Le Wang 0003, Nebojsa Jojic, Zhenxing Niu, Nanning Zheng 0001, Gang Hua 0001
IEEE Trans. Pattern Anal. Mach. Intell.5
2019 View Adaptive Neural Networks for High Performance Skeleton-Based Human Action Recognition
abstract
Skeleton-based human action recognition has recently attracted increasing attention thanks to the accessibility and the popularity of 3D skeleton data. One of the key challenges in action recognition lies in the large variations of action representations when they are captured from different viewpoints. In order to alleviate the effects of view variations, this paper introduces a novel view adaptation scheme, which automatically determines the virtual observation viewpoints over the course of an action in a learning based data driven manner. Instead of re-positioning the skeletons using a fixed human-defined prior criterion, we design two view adaptive neural networks, i.e., VA-RNN and VA-CNN, which are respectively built based on the recurrent neural network (RNN) with the Long Short-term Memory (LSTM) and the convolutional neural network (CNN). For each network, a novel view adaptation module learns and determines the most suitable observation viewpoints, and transforms the skeletons to those viewpoints for the end-to-end recognition with a main classification network. Ablation studies find that the proposed view adaptive models are capable of transforming the skeletons of various views to much more consistent virtual viewpoints. Therefore, the models largely eliminate the influence of the viewpoints, enabling the networks to focus on the learning of action-specific features and thus resulting in superior performance. In addition, we design a two-stream scheme (referred to as VA-fusion) that fuses the scores of the two networks to provide the final prediction, obtaining enhanced performance. Moreover, random rotation of skeleton sequences is employed to improve the robustness of view adaptation models and alleviate overfitting during training. Extensive experimental evaluations on five challenging benchmarks demonstrate the effectiveness of the proposed view-adaptive networks and superior performance over state-of-the-art approaches.
Pengfei Zhang 0005, Cuiling Lan, Junliang Xing, Wenjun Zeng 0001, Jianru Xue, Nanning Zheng 0001
IEEE Trans. Pattern Anal. Mach. Intell.6
2019 Correntropy based scale ICP algorithm for robust point set registration
Zongze Wu 0001, Hongchen Chen, Shaoyi Du, Minyue Fu 0001, Nanning Zheng 0001
Pattern Recognit.6
2019 Semi-supervised person re-identification using multi-view clustering
Xiaomeng Xin, Jinjun Wang, Ruji Xie, Sanping Zhou, Wenli Huang 0004, Nanning Zheng 0001
Pattern Recognit.6
2019 Design Space Exploration of Neural Network Activation Function Circuits
abstract
The widespread application of artificial neural networks has prompted researchers to experiment with field-programmable gate array and customized ASIC designs to speed up their computation. These implementation efforts have generally focused on weight multiplication and signal summation operations, and less on activation functions used in these applications. Yet, efficient hardware implementations of nonlinear activation functions like exponential linear units (ELU), scaled ELU (SELU), and hyperbolic tangent (tanh), are central to designing effective neural network accelerators, since these functions require lots of resources. In this paper, we explore efficient hardware implementations of activation functions using purely combinational circuits, with a focus on two widely used nonlinear activation functions, i.e., SELU and tanh. Our experiments demonstrate that neural networks are generally insensitive to the precision of the activation function. The results also prove that the proposed combinational circuit-based approach is very efficient in terms of speed and area, with negligible accuracy loss on the MNIST, CIFAR-10, and IMAGE NET benchmarks. Synopsys design compiler synthesis results show that circuit designs for tanh and SELU can save between ${\times 3.13\sim \times 7.69}$ and ${ {\times 4.45\sim \times 8.45}}$ area compared to the look-up table/memory-based implementations, and can operate at 5.14 GHz and 4.52 GHz using the 28-nm SVT library, respectively. The implementation is available at: https://github.com/ThomasMrY/ActivationFunctionDemo.
Tao Yang 0032, Yadong Wei, Zhijun Tu, Haolun Zeng, Michel A. Kinsy, Nanning Zheng 0001, Pengju Ren
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.6
2019 Joint Task Difficulties Estimation and Testees Ranking for Intelligence Evaluation
abstract
In this paper, we study the testing tasks evaluation and testees ranking problem, in which tasks have different difficulty levels, and testees have different capabilities.We assume that a testee may have a probability to pass a certain task so as to allow certain uncertainty. The goal of this problem is to simultaneously determine the relative difficulty level of each testing task and the relative capability of every testee, purely based on the test outcome. We design two models to solve this problem. The first one assumes that the test outcome follows a certain Bernoulli distribution; while the second one assumes that the test outcome follows a certain Bernoulli distribution with the beta distribution-type a priori knowledge. Then, we form the original problem into likelihood estimation problems and solve them by using coordinate descent algorithms. We show that the beta distribution-type a priori knowledge is needed, when we only carry out a limited number of tests due to time and financial budgets. All these findings are useful to intelligence tests. Finally, we discuss how to extend this statistical learning model for more general cases as well as in a specific case in the field of Computational Social Systems like artificial social cognition evaluation.
Chi Zhang 0020, Yuehu Liu, Li Li 0013, Nanning Zheng 0001, Fei-Yue Wang 0001
IEEE Trans. Comput. Soc. Syst.4
2019 NIPM-sWMF: Toward Efficient FPGA Design for High-Definition Large-Disparity Stereo Matching
abstract
Large disparity stereo matching is critical to the application of a stereo vision system especially for outdoor scenes. Nevertheless, how to efficiently design high accuracy large-disparity stereo matching on a field-programmable gate array (FPGA) is still a grand challenge. The computational complexity of previously proposed stereo matching is inevitably proportional to disparity range; hence their hardware designs become very inefficient when the disparity range is large. Motivated by the original PatchMatch and weighted median filtering (WMF) algorithms, this paper proposes a non-iterative PatchMatch and separable WMF (NIPM-sWMF) algorithm to significantly reduce the computational complexity of stereo matching and make it independent of disparity range. Moreover, we also propose a fully pipelined architecture design on FPGA that employs several hardware techniques to efficiently implement the proposed NIPM-sWMF. The disparity quality of the proposed NIPM-sWMF algorithm is evaluated on both KITTI2015 and Middlebury V3 stereo data sets, and the proposed architecture design is implemented and synthesized on Xilinx FPGA. Evaluation results demonstrate that the proposed NIPM-sWMF design on FPGA reaches the real-time performance of 1920 × 1080@60 Hz at the disparity range of 128, and can achieve almost the same disparity estimation accuracy, 4.5× processing throughput, while reducing the hardware cost of LUT, Register, DSP, and BRAM by 40%, 47%, 100%, and 68%, respectively, compared with the reference stereo matching design. Therefore, the proposed NIPM-sWMF design is an efficient way to address the challenge of large-disparity stereo matching.
Xuchong Zhang, Hongbin Sun 0001, Shiqiang Chen, Lin Song 0002, Nanning Zheng 0001
IEEE Trans. Circuits Syst. Video Technol.5
2019 On the Crossroad of Artificial Intelligence: A Revisit to Alan Turing and Norbert Wiener
abstract
To give a high-level summary to current approaches for implementing artificial intelligence (AI), we explain the key commonalities and major differences between Turing's approach and Wiener's approach in this perspective. Especially, the problems, successful achievements, limitations, and future research directions of existing approaches that follow Weiner's ideas are addressed, respectively, aiming to provide readers with a good start point and a roadmap. Some other related topics, for example, the role of human experts in developing AI, are also discussed to seek potential solutions for some existing difficulties.
Li Li 0013, Nanning Zheng 0001, Fei-Yue Wang 0001
IEEE Trans. Cybern.2
2019 Cross-View Person Identification Based on Confidence-Weighted Human Pose Matching
abstract
Cross-view person identification (CVPI) from multiple temporally synchronized videos taken by multiple wearable cameras from different, varying views is a very challenging but important problem, which has attracted more interest recently. Current state-of-the-art performance of CVPI is achieved by matching appearance and motion features across videos, while the matching of pose features does not work effectively given the high inaccuracy of the 3D pose estimation on videos/images collected in the wild. To address this problem, we first introduce a new metric of confidence to the estimated location of each human-body joint in 3D human pose estimation. Then, a mapping function, which can be hand-crafted or learned directly from the datasets, is proposed to combine the inaccurately estimated human pose and the inferred confidence metric to accomplish CVPI. Specifically, the joints with higher confidence are weighted more in the pose matching for CVPI. Finally, the estimated pose information is integrated into the appearance and motion features to boost the CVPI performance. In the experiments, we evaluate the proposed method on three wearable-camera video datasets and compare the performance against several other existing CVPI methods. The experimental results show the effectiveness of the proposed confidence metric, and the integration of pose, appearance, and motion produces a new state-of-the-art CVPI performance.
Guoqiang Liang 0001, Xuguang Lan, Xingyu Chen 0001, Song Wang 0002, Nanning Zheng 0001
IEEE Trans. Image Process.6
2019 Discriminative Feature Learning With Foreground Attention for Person Re-Identification
abstract
The performance of person re-identification (Re-ID) has been seriously affected by the large cross-view appearance variations caused by mutual occlusions and background clutter. Hence, learning a feature representation that can adaptively emphasize the foreground persons becomes very critical to solve the person Re-ID problem. In this paper, we propose a simple yet effective foreground attentive neural network (FANN) to learn a discriminative feature representation for person Re-ID, which can adaptively enhance the positive side of foreground and weaken the negative side of background. Specifically, a novel foreground attentive subnetwork is designed to drive the network’s attention, in which a decoder network is used to reconstruct the binary mask by using a novel local regression loss function, and an encoder network is regularized by the decoder network to focus its attention on the foreground persons. The resulting feature maps of encoder network are further fed into the body part subnetwork and feature fusion subnetwork to learn discriminative features. Besides, a novel symmetric triplet loss function is introduced to supervise feature learning, in which the intra-class distance is minimized and the inter-class distance is maximized in each triplet unit, simultaneously. Training our FANN in a multi-task learning framework, a discriminative feature representation can be learned to find out the matched reference to each probe among various candidates in the gallery. Extensive experimental results on several public benchmark datasets are evaluated, which have shown clear improvements of our method over the state-of-the-art approaches.
Sanping Zhou, Jinjun Wang, Deyu Meng, Yudong Liang, Yihong Gong, Nanning Zheng 0001
IEEE Trans. Image Process.6
2019 Learning Composite Latent Structures for 3D Human Action Representation and Recognition
abstract
3D human action representation and recognition are important issues in many multimedia applications. While latent state approaches have been widely used for action modeling, previous works assume the latent states of actions are single attribute. This assumption is inaccurate for representing structures of complex actions. In this paper, we propose that latent states have composite attributes and introduce a novel composite latent structure (CLS) model to represent and recognize 3D human actions with skeleton sequences. A human action is modeled with a hierarchical graph, which represents the action sequence as sequential atomic actions. An atomic action is represented as a composite latent state, which is composed of a latent semantic attribute and a latent geometric attribute. A discriminative EM-like algorithm is proposed to learn the model parameters and the composite latent structures of human actions. Given a 3D skeleton sequence, a composite attribute iterative programming algorithm is proposed to recognize the action and infer the action's latent temporal structure. We evaluate the proposed method on three challenging 3D action datasets-MSR 3D Action Dataset, Multiview 3D Event Dataset, and UTKinect-Action 3D Dataset. Extensive experimental results on these datasets demonstrate the effectiveness and advantage of the proposed method.
Ping Wei 0001, Hongbin Sun 0001, Nanning Zheng 0001
IEEE Trans. Multim.3
2019 SynBF: A New Bilateral Filter for Postremoval of Noise From Synthesis Views in 3-D Video
abstract
In 3-D video systems, noise in the texture and depth videos of reference views may not be removed (Scenario 1) or not fully removed (Scenario 2) by prefiltering methods before the view synthesis procedure. In these scenarios, the noise is transferred to the generated synthesis view. After investigating the noise model of the synthesis view, we conclude that the noise in the synthesis view not only causes fluctuation in the photometric values of pixels in the range domain but also additionally shifts the positions of neighbor pixels in the spatial domain compared to that in natural images. It consequently damages the textural content near edges in the synthesis view, for which the popular local filters of natural images, that is, the bilateral filter (BF) and the guided filter, do not work well. In this paper, we develop a new local filter for the synthesis view (named SynBF) after it has been generated, which has a similar expression as that of the BF but not the exact same weight terms. On one hand, the spatial term in the classical BF is directly reused due to its robustness to noise, which gives high weights to spatially closed pixels of to-be-filtered pixels. On the other hand, a reliability term is designed that gives high weights to pixels that are unlikely to be affected by noise. It is inspired by the finding that not all pixels are significantly affected by noise in the synthesis view. In this way, true edge profiles are protected in the filtering process. Experiments are conducted on a set of synthesis views for both scenarios above and compared to the two local filters, which verifies its effectiveness in removing noise and protecting edge profiles. The proposed method can be considered as a supplement to prefiltering methods of texture/depth videos in 3-D video systems.
Meng Yang 0002, Nanning Zheng 0001
IEEE Trans. Multim.2
2019 Efficient Estimation of View Synthesis Distortion for Depth Coding Optimization
abstract
Depth coding in depth-based three-dimensional (3-D) video is unique in that its quality is measured by view synthesis distortion (VSD) rather than the depth distortion itself, which further complicates the coding optimization as the VSD is related to quality of both the associated depth and texture videos. In this paper, an efficient VSD estimation scheme is developed to measure the effect of depth errors on the VSD for a block given its depth distortion in mean-squared error. Unlike other relevant VSD models which involve computationally intensive parameter training or Fourier transform, the proposed scheme is free of parameter training, while taking the advantage of integer 4 × 4 discrete Cosine transform to replace Fourier transform, thus well-saving computational cost and diminishing sensitivity to training dataset of video. The proposed scheme is then incorporated on the coding unit basis into the rate-distortion optimization for depth coding optimization, coupled with adapting quantization parameter accordingly to accommodate local effect of the depth errors on the VSD. Experimental results show that our solution obtains better results in depth coding than three testing solutions, on the platform of H.264/AVC reference software JM16.0. Benefiting from the efficiency of the VSD estimation, low coding complexity is obtained as well. The proposed solution is further evaluated on the reference software HTM13.0 of the latest 3-D high-efficiency video coding standard, exhibiting better and comparable results compared against the HTM codec with the view synthesis optimization disabled and enabled, respectively.
Meng Yang 0002, Ce Zhu, Xuguang Lan, Nanning Zheng 0001
IEEE Trans. Multim.4
2019 Quantized Minimum Error Entropy Criterion
abstract
Comparing with traditional learning criteria, such as mean square error, the minimum error entropy (MEE) criterion is superior in nonlinear and non-Gaussian signal processing and machine learning. The argument of the logarithm in Renyi's entropy estimator, called information potential (IP), is a popular MEE cost in information theoretic learning. The computational complexity of IP is, however, quadratic in terms of sample number due to double summation. This creates the computational bottlenecks, especially for large-scale data sets. To address this problem, in this paper, we propose an efficient quantization approach to reduce the computational burden of IP, which decreases the complexity from O(N2) to O(MN) with M ≪ N. The new learning criterion is called the quantized MEE (QMEE). Some basic properties of QMEE are presented. Illustrative examples with linear-in-parameter models are provided to verify the excellent performance of QMEE.
Badong Chen, Lei Xing 0003, Nanning Zheng 0001, José C. Príncipe
IEEE Trans. Neural Networks Learn. Syst.3
2019 Fine-Grained Image Classification Using Modified DCNNs Trained by Cascaded Softmax and Generalized Large-Margin Losses
abstract
We develop a fine-grained image classifier using a general deep convolutional neural network (DCNN). We improve the fine-grained image classification accuracy of a DCNN model from the following two aspects. First, to better model the h -level hierarchical label structure of the fine-grained image classes contained in the given training data set, we introduce h fully connected (fc) layers to replace the top fc layer of a given DCNN model and train them with the cascaded softmax loss. Second, we propose a novel loss function, namely, generalized large-margin (GLM) loss, to make the given DCNN model explicitly explore the hierarchical label structure and the similarity regularities of the fine-grained image classes. The GLM loss explicitly not only reduces between-class similarity and within-class variance of the learned features by DCNN models but also makes the subclasses belonging to the same coarse class be more similar to each other than those belonging to different coarse classes in the feature space. Moreover, the proposed fine-grained image classification framework is independent and can be applied to any DCNN structures. Comprehensive experimental evaluations of several general DCNN models (AlexNet, GoogLeNet, and VGG) using three benchmark data sets (Stanford car, fine-grained visual classification-aircraft, and CUB-200-2011) for the fine-grained image classification task demonstrate the effectiveness of our method.
Weiwei Shi 0003, Yihong Gong, De Cheng, Nanning Zheng 0001
IEEE Trans. Neural Networks Learn. Syst.5
2019 Efficient Compression-Based Line Buffer Design for Image/Video Processing Circuits
abstract
Line buffer is a typical and major on-chip memory design architecture for image/video processing circuits. As it usually occupies very large on-chip circuit area, it is of great importance to reduce its hardware cost through efficient architecture design. Data compression is a promising technique to improve the hardware efficiency of line buffer architecture. Nevertheless, the previously proposed data compression technique for line buffer architecture only exploits fixed length code (FLC), which actually has the deficiency on compression performance. Instead, this paper explores to efficiently use variable length code in line buffer architecture. By restricting variable length coding within small compression granularity (CG), the proposed compression algorithm not only significantly improves compression performance but also meets the specific requirements in line buffer architecture design. The simple compression algorithm further enables the efficient and fully pipelined VLSI architecture and circuits. Experimental results demonstrate that the proposed compression algorithm achieves 6.67-dB peak signal-to-noise ratio improvement at the compression ratio of 50% and the CG of 16 pixels, compared with FLC design. The VLSI circuits of the proposed compression can achieve the throughput of 4K × 2K at 60 fps with reasonable hardware cost. The use of the proposed compression technique in line buffer architecture can significantly reduce on-chip memory cost while maintaining satisfactory visual quality.
Longjun Liu, Hongbin Sun 0001, Nanning Zheng 0001
IEEE Trans. Very Large Scale Integr. Syst.5
2018 Cross-View Person Identification by Matching Human Poses Estimated With Confidence on Each Body Joint
abstract
Cross-view person identification (CVPI) from multiple temporally synchronized videos taken by multiple wearable cameras from different, varying views is a very challenging but important problem, which has attracted more interests recently. Current state-of-the-art performance of CVPI is achieved by matching appearance and motion features across videos, while the matching of pose features does not work effectively given the high inaccuracy of the 3D human pose estimation on videos/images collected in the wild. In this paper, we introduce a new metric of confidence to the 3D human pose estimation and show that the combination of the inaccurately estimated human pose and the inferred confidence metric can be used to boost the CVPI performance---the estimated pose information can be integrated to the appearance and motion features to achieve the new state-of-the-art CVPI performance. More specifically, the estimated confidence metric is measured at each human-body joint and the joints with higher confidence are weighted more in the pose matching for CVPI. In the experiments, we validate the proposed method on three wearable-camera video datasets and compare the performance against several other existing CVPI methods.
Guoqiang Liang 0001, Xuguang Lan, Song Wang 0002, Nanning Zheng 0001
AAAI5
2018 Where and Why Are They Looking? Jointly Inferring Human Attention and Intentions in Complex Tasks
abstract
This paper addresses a new problem - jointly inferring human attention, intentions, and tasks from videos. Given an RGB-D video where a human performs a task, we answer three questions simultaneously: 1) where the human is looking - attention prediction; 2) why the human is looking there - intention prediction; and 3) what task the human is performing - task recognition. We propose a hierarchical model of human-attention-object (HAO) which represents tasks, intentions, and attention under a unified framework. A task is represented as sequential intentions which transition to each other. An intention is composed of the human pose, attention, and objects. A beam search algorithm is adopted for inference on the HAO graph to output the attention, intention, and task results. We built a new video dataset of tasks, intentions, and attention. It contains 14 task classes, 70 intention categories, 28 object classes, 809 videos, and approximately 330,000 frames. Experiments show that our approach outperforms existing approaches.
Ping Wei 0001, Yang Liu 0266, Tianmin Shu, Nanning Zheng 0001, Song-Chun Zhu
CVPR4
2018 Kernelized Subspace Pooling for Deep Local Descriptors
abstract
Representing local image patches in an invariant and discriminative manner is an active research topic in computer vision. It has recently been demonstrated that local feature learning based on deep Convolutional Neural Network (CNN) can significantly improve the matching performance. Previous works on learning such descriptors have focused on developing various loss functions, regularizations and data mining strategies to learn discriminative CNN representations. Such methods, however, have little analysis on how to increase geometric invariance of their generated descriptors. In this paper, we propose a descriptor that has both highly invariant and discriminative power. The abilities come from a novel pooling method, dubbed Subspace Pooling (SP) which is invariant to a range of geometric deformations. To further increase the discriminative power of our descriptor, we propose a simple distance kernel integrated to the marginal triplet loss that helps to focus on hard examples in CNN training. Finally, we show that by combining SP with the projection distance metric [13], the generated feature descriptor is equivalent to that of the Bilinear CNN model [22], but outperforms the latter with much lower memory and computation consumptions. The proposed method is simple, easy to understand and achieves good performance. Experimental results on several patch matching benchmarks show that our method outperforms the state-of-the-arts significantly.
Xing Wei 0001, Yihong Gong, Nanning Zheng 0001
CVPR4
2018 Transductive Semi-Supervised Deep Learning Using Min-Max Features
Weiwei Shi 0003, Yihong Gong, Chris Ding, Zhiheng Ma, Nanning Zheng 0001
ECCV (5)6
2018 Grassmann Pooling as Compact Homogeneous Bilinear Pooling for Fine-Grained Visual Classification
Xing Wei 0001, Yihong Gong, Jiawei Zhang 0002, Nanning Zheng 0001
ECCV (3)5
2018 Adding Attentiveness to the Neurons in Recurrent Neural Networks
Pengfei Zhang 0005, Jianru Xue, Cuiling Lan, Wenjun Zeng 0001, Zhanning Gao, Nanning Zheng 0001
ECCV (9)6
2018 Accelerate GPU Concurrent Kernel Execution by Mitigating Memory Pipeline Stalls
abstract
Following the advances in technology scaling, graphics processing units (GPUs) incorporate an increasing amount of computing resources and it becomes difficult for a single GPU kernel to fully utilize the vast GPU resources. One solution to improve resource utilization is concurrent kernel execution (CKE). Early CKE mainly targets the leftover resources. However, it fails to optimize the resource utilization and does not provide fairness among concurrent kernels. Spatial multitasking assigns a subset of streaming multiprocessors (SMs) to each kernel. Although achieving better fairness, the resource underutilization within an SM is not addressed. Thus, intra-SM sharing has been proposed to issue thread blocks from different kernels to each SM. However, as shown in this study, the overall performance may be undermined in the intra-SM sharing schemes due to the severe interference among kernels. Specifically, as concurrent kernels share the memory subsystem, one kernel, even as computing-intensive, may starve from not being able to issue memory instructions in time. Besides, severe L1 D-cache thrashing and memory pipeline stalls caused by one kernel, especially a memory-intensive one, will impact other kernels, further hurting the overall performance. In this study, we investigate various approaches to overcome the aforementioned problems exposed in intra-SM sharing. We first highlight that cache partitioning techniques proposed for CPUs are not effective for GPUs. Then we propose two approaches to reduce memory pipeline stalls. The first is to balance memory accesses of concurrent kernels. The second is to limit the number of inflight memory instructions issued from individual kernels. Our evaluation shows that the proposed schemes significantly improve the weighted speedup of two state-of-the-art intra-SM sharing schemes, Warped-Slicer and SMK, by 24.6% and 27.2% on average, respectively, with lightweight hardware overhead.
Hongwen Dai, Chao Li 0004, Chen Zhao 0009, Fei Wang 0008, Nanning Zheng 0001, Huiyang Zhou
HPCA6
2018 Joint Spatio-Temporal Action Localization in Untrimmed Videos with Per-Frame Segmentation
abstract
Inspired by the recent spatio-temporal action localization efforts with tubelets (sequences of bounding boxes), we present a new spatio-temporal action detector Segment-tube, which consists of sequences of per-frame segmentation masks. The proposed Segment-tube detector can temporally pinpoint the starting/ending frame of each action class in the presence of preceding/subsequent interference actions in untrimmed videos. Simultaneously, the Segment-tube detector produces per-frame segmentation masks instead of bounding boxes, offering superior spatial accuracy to tubelets. This is achieved by alternating iterative optimization between temporal action localization and spatial action segmentation. Experimental results on multiple datasets validate the efficacy of the proposed detector.
Xuhuan Duan, Le Wang 0003, Changbo Zhai, Nanning Zheng 0001, Qilin Zhang 0004, Zhenxing Niu, Gang Hua 0001
ICIP4
2018 Video Object Co-Segmentation from Noisy Videos by a Multi-Level Hypergraph Model
abstract
Defined as simultaneously segmenting a set of related videos to identify the common objects, video co-segmentation has attracted the attention of researchers in recent years. Existing methods are primarily based on pair-wise relations between adjacent pixels/regions, which are susceptible to performance degradation from “empty” video frames (e.g., due to transient/intermittent common objects). In this paper, a new multilevel hypergraph based method, termed the full Video object Co-Segmentation method (VCS), is proposed, which incorporates both a high-level semantics object model and a low-level appearance/motion/saliency object model to construct the hyperedge among multiple spatially and temporally adjacent regions. Specifically, the high-level semantic model fuses multiple object proposals from each frame instead of relying on a single object proposal per frame. A hypergraph cut is subsequently utilized to calculate the object co-segmentation. Experiments on three datasets demonstrate the efficacy of the proposed VCS method.
Le Wang 0003, Qilin Zhang 0004, Nanning Zheng 0001, Gang Hua 0001
ICIP4
2018 Augmented Space Linear Model
abstract
The linear model uses the space defined by the input to project the target or desired signal and find the optimal set of model parameters. When the problem is nonlinear, the adaption requires nonlinear models for good performance, but it becomes slower and more cumbersome. In this paper, we propose a linear model called Augmented Space Linear Model (ASLM), which uses the full joint space of input and desired signal as the projection space and approaches the performance of nonlinear models. This new algorithm takes advantage of the linear solution, and corrects the estimate for the current testing phase input with the error assigned to the input space neighborhood in the training phase. This algorithm can solve the nonlinear problem with the computational efficiency of linear methods, which can be regarded as a trade off between accuracy and computational complexity. Making full use of the training data, the proposed augmented space model may provide a new way to improve many modeling tasks.
Zhengda Qin, Badong Chen, Nanning Zheng 0001, José C. Príncipe
IJCNN3
2018 Spiking Locality-Sensitive Hash: Spiking Computation with Phase Encoding Method
abstract
A novel similarity search method, named spiking locality sensitive hash (SLSH), a forward spiking neuron network(SNN) is proposed in this paper. The SLSH architecture is composed of successively connected encoding and fully connected layer. We optimize phase encoding to maximize the difference between corresponding pixels of any two different images. Then we test the performance of the encoding method and the SLSH model on graphic datasets. Experimental results prove that improved phase encoding method based on the difference exhibits the accuracy of 100%, 100% and 92%, which has superiority over previous phase encoding whose accuracies are 93%, 78% and 55% when the noise level is 5%, 20% and 40% respectively. Furthermore, experiments demonstrate that SLSH method is more capable than the traditional Locality-Sensitive Hash(LSH) and the FLY algorithm published in SCIENCE in similarity search. The mean average precision of SLSH is twice of FLY algorithm when the hash length is 5. In addition, the SLSH achieves a good recognition performance even under the influence of noise for MNIST, SVHN and SIFT datasets.
Ziru Wang, Zhiwei Dong, Nanning Zheng 0001, Pengju Ren
IJCNN4
2018 Accurate Mix-Norm-Based Scan Matching
abstract
Highly accurate mapping and localization is of prime importance for mobile robotics, and its core lies in efficient scan matching. Previous research are focusing on designing a robust objective function and the residual error distribution is often ignored or simply assumed as unitary or mixture of simple distributions. In this paper, a mixture of exponential power (MoEP) distributions is proposed to approximate the residual error distribution. The objective function induced by MoEP-based residual error modelling ensembles a mix-norm-based scan matching (MiNoM), which enhances the matching accuracy and convergence characteristic. Both the parameters of transformation (rotation and translation) and residual error distribution are estimated efficiently via an EM-like algorithm. The optimization of MiNoM is iteratively achieved via two phases: An on-line parameter learning (OPL) phase to learn residual error distribution for better representation according to the likelihood field model (LFM), and an iteratively reweighted least squares (IRLS) phase to attain transformation for accuracy and efficiency. Extensive experimental results validate that the proposed MiNoM out-performs several state-of-the-art scan matching algorithms in both convergence characteristic and matching accuracy.
Di Wang 0028, Jianru Xue, Zhongxing Tao, Dixiao Cui, Shaoyi Du, Nanning Zheng 0001
IROS7
2018 Fully Convolutional Grasp Detection Network with Oriented Anchor Box
abstract
In this paper, we present a real-time approach to predict multiple grasping poses for a parallel-plate robotic gripper using RGB images. A model with oriented anchor box mechanism is proposed and a new matching strategy is used during the training process. An end-to-end fully convolutional neural network is employed in our work. The network consists of two parts: the feature extractor and multi-grasp predictor. The feature extractor is a deep convolutional neural network. The multi-grasp predictor regresses grasp rectangles from predefined oriented rectangles, called oriented anchor boxes, and classifies the rectangles into graspable and ungraspable. On the standard Cornell Grasp Dataset, our model achieves an accuracy of 97.74% and 96.61% on image-wise split and object-wise split respectively, and outperforms the latest state-of-the-art approach by 1.74% on image-wise split and 0.51% on object-wise split.
Xinwen Zhou, Xuguang Lan, Hanbo Zhang, Nanning Zheng 0001
IROS6
2018 Autonomous Vehicle Testing and Validation Platform: Integrated Simulation System with Hardware in the Loop
abstract
With the development of autonomous driving, offline testing remains an important process allowing low-cost and efficient validation of vehicle performance and vehicle control algorithms in multiple virtual scenarios. This paper aims to propose a novel simulation platform with hardware in the loop (HIL). This platform comprises of four layers: the vehicle simulation layer, the virtual sensors layer, the virtual environment layer and the Electronic Control Unit (ECU) layer for hardware control. Our platform has attained multiple capabilities: (1) it enables the construction and simulation of kinematic car models, various sensors and virtual testing fields; (2) it performs a closed-loop evaluation of scene perception, path planning, decision-making and vehicle control algorithms, whilst also having multi-agent interaction system; (3) it further enables rapid migrations of control and decision-making algorithms from the virtual environment to real self-driving cars. In order to verify the effectiveness of our simulation platform, several experiments have been performed with self-defined car models in virtual scenarios of a public road and an open parking lot and the results are substantial.
Yu Chen 0040, Shi-tao Chen, Tangyike Zhang, Songyi Zhang, Nanning Zheng 0001
Intelligent Vehicles Symposium5
2018 Model-Based Decision Making With Imagination for Autonomous Parking
abstract
Autonomous parking technology is a key concept within autonomous driving research. This paper will propose an imaginative autonomous parking algorithm to solve issues concerned with parking. The proposed algorithm consists of three parts: an imaginative model for anticipating results before parking, an improved rapid-exploring random tree (RRT) for planning a feasible trajectory from a given start point to a parking lot, and a path smoothing module for optimizing the efficiency of parking tasks. Our algorithm is based on a real kinematic vehicle model; which makes it more suitable for algorithm application on real autonomous cars. Furthermore, due to the introduction of the imagination mechanism, the processing speed of our algorithm is ten times faster than that of traditional methods, permitting the realization of real-time planning simultaneously. In order to evaluate the algorithm's effectiveness, we have compared our algorithm with traditional RRT, within three different parking scenarios. Ultimately, results show that our algorithm is more stable than traditional RRT and performs better in terms of efficiency and quality.
Ziyue Feng, Shi-tao Chen, Yu Chen 0040, Nanning Zheng 0001
Intelligent Vehicles Symposium4
2018 A Novel Approach for Detecting Road Based on Two-Stream Fusion Fully Convolutional Network
abstract
Road detection is one of the most basic tasks of autonomous driving systems. At present, researches on this issue mainly take two kinds of data as input,i.e., LIDAR point clouds and RGB images from cameras. To make best use of the advantages and bypass the disadvantages of these two kinds of data, we propose a novel network, namely two- stream fusion fully convolutional network (TSF-FCN), which can take advantage of both the accurate location information from LIDAR point clouds and rich appearance information from RGB images. One stream of this network is LIDAR stream which aggregates multi-scale contextual information from LIDAR point clouds. The other stream is RGB stream which is used for extracting features from RGB images. To fuse the two streams, the feature maps of RGB stream are converted to a bird-view representation to concatenate with that of LIDAR stream. In this way, the two kinds of data can complement each other for detecting road. To verify the efficacy of our TSF-FCN, experiments are carried on KITTI- ROAD benchmark and competitive performance is achieved compared with state-of-the-art methods.
Ziyi Liu 0001, Jingmin Xin, Nanning Zheng 0001
Intelligent Vehicles Symposium4
2018 Exploring the Potential of Using Semantic Context and Common Sense in On-Road Vehicle Detection
abstract
Vehicle detection is an important research topic for autonomous driving community. Since the great success of deep learning on object detection, almost all vehicle detection methods go along with this line. However, deep learning methods heavily rely on the training data, and the whole mechanism is like a “black box” Therefore, in this paper, we explore a vehicle detection method using traffic semantic context and human common sense instead of relying on the training data. To verify our idea, we compare our method with two classic machine learning methods as well as three state- of-the-art deep learning methods on a dataset collected in real traffics. The results show that our method outperforms others on this dataset. The deep learning methods may exceed ours after enlarging the training data or testing on more complicated datasets. However, the main contribution of this paper is providing inspiration for learning methods, and we believe their performance can be greatly improved after considering the idea of this paper.
Zhixiong Nan, Menghan Pan, Xiao Wang 0002, Ping Wei 0001, Linhai Xu, Hongbin Sun 0001, Jingmin Xin, Nanning Zheng 0001
Intelligent Vehicles Symposium8
2018 Leveraging Spatio-Temporal Evidence and Independent Vision Channel to Improve Multi-Sensor Fusion for Vehicle Environmental Perception
abstract
For intelligent vehicles, multi-sensor fusion is of great importance to perceive traffic environment with high accuracy and robustness. In this paper, we propose two effective methods, i.e. spatio-temporal evidence generating and independent vision channel, to improve multi-sensor track-level fusion for vehicle environmental perception. The spatio-temporal evidence includes instantaneous evidence, tracking evidence and tracks matching evidence to improve existence fusion. Independent vision channel leverages the specific advantage of vision processing on object recognition to improve classification fusion. The proposed methods are evaluated by using the multi-sensor dataset collected from real traffic environment. Experimental results demonstrate that the proposed methods can significantly improve the multi-sensor track-level fusion in terms of both detection accuracy and classification accuracy.
Juwang Shi, Wenxiu Wang, Xiao Wang 0002, Hongbin Sun 0001, Xuguang Lan, Jingmin Xin, Nanning Zheng 0001
Intelligent Vehicles Symposium7
2018 Modeling and Predicting Vehicle Motion Activities by Using And-Or Graph
abstract
The ability of modeling and predicting vehicle motion activities is important for automated vehicles. In this paper, we propose an And-Or Graph based model to give a simple and clear description of motion activities. Compared to other models, this new model relaxes the Markov property requirement in transition between activities and is thus more flexible. The parameters of this model can be easily learned from data. Using the trained new model, we can predict the on-going motion activity label and its corresponding probability. Experiments show that a high prediction accuracy (97%) can be achieved by this new model.
Shuofeng Wang, Li Li 0013, Nanning Zheng 0001, Dongpu Cao
Intelligent Vehicles Symposium3
2018 Efficient Rectangle Fitting of Sparse Laser Data for Robust On-Road Obiect Detection
abstract
On-road object detection is one of the most important tasks for the autonomous driving of intelligent vehicle. Nevertheless, the previous methods based on 2D LIDAR sensor only focus on the detection of vehicles, and show severe limitations on the detection of other objects. Accordingly, this paper proposes an on-road object detection method, which employs rectangle fitting and concavity determination to improve the robustness of ob- ject detection. The proposed approaches are extensively evaluated by using the sparse laser data collected by 2D LIDAR from real traffic environment. Experimental results demonstrate that the proposed rectangle fitting outperforms the previous approaches in terms of both detection accuracy and computational efficiency.
Zhaohong Xiang, Xiao Wang 0002, Hongbin Sun 0001, Jinming Xin, Nanning Zheng 0001
Intelligent Vehicles Symposium7
2018 Automatic salient object sequence rebuilding for video segment analysis
Haibin Duan, Zejian Yuan, Nanning Zheng 0001
Sci. China Inf. Sci.5
2018 A novel spiking neural network of receptive field encoding with groups of neurons decision
abstract
Human information processing depends mainly on billions of neurons which constitute a complex neural network, and the information is transmitted in the form of neural spikes. In this paper, we propose a spiking neural network (SNN), named MD-SNN, with three key features: (1) using receptive field to encode spike trains from images; (2) randomly selecting partial spikes as inputs for each neuron to approach the absolute refractory period of the neuron; (3) using groups of neurons to make decisions. We test MD-SNN on the MNIST data set of handwritten digits, and results demonstrate that: (1) Different sizes of receptive fields influence classification results significantly. (2) Considering the neuronal refractory period in the SNN model, increasing the number of neurons in the learning layer could greatly reduce the training time, effectively reduce the probability of over-fitting, and improve the accuracy by 8.77%. (3) Compared with other SNN methods, MD-SNN achieves a better classification; compared with the convolution neural network, MD-SNN maintains flip and rotation invariance (the accuracy can remain at 90.44% on the test set), and it is more suitable for small sample learning (the accuracy can reach 80.15% for 1000 training samples, which is 7.8 times that of CNN).
Ziru Wang, Si-yu Yu, Badong Chen, Nanning Zheng 0001, Pengju Ren
Frontiers Inf. Technol. Electron. Eng.5
2018 Deep feature learning via structured graph Laplacian embedding for person re-identification
De Cheng, Yihong Gong, Xiaojun Chang, Weiwei Shi 0003, Alex Hauptmann 0001, Nanning Zheng 0001
Pattern Recognit.6
2018 Entropy and orthogonality based deep discriminative feature learning for object recognition
Weiwei Shi 0003, Yihong Gong, De Cheng, Nanning Zheng 0001
Pattern Recognit.5
2018 Deep self-paced learning for person re-identification
Sanping Zhou, Jinjun Wang, Deyu Meng, Xiaomeng Xin, Yihong Gong, Nanning Zheng 0001
Pattern Recognit.7
2018 VLSI Architecture Exploration of Guided Image Filtering for 1080P@60Hz Video Processing
abstract
Guided image filtering (GIF) is a promising edge-preserving filtering technique that has been applied in a variety of applications. Nevertheless, an efficient very-large-scale integration (VLSI) architecture design of GIF is still very challenging for the real-time processing of full-high definition videos. Previously proposed architectures are somewhat inefficient in terms of either on-chip memory usage or off-chip memory bandwidth. This paper aims to improve the balance between on-chip memory usage and off-chip memory bandwidth through architecture exploration. Three critical architectural tradeoffs in the VLSI design of GIF are explored, and two efficient VLSI architectures, namely sequential line-based and parallel line-based architectures, are proposed. Experimental results demonstrate that the proposed VLSI design only consumes 34.1-K logic gates, 25.4-KB on-chip memories, and 373-MB/s off-chip memory bandwidth while achieving a real-time video processing of 1080P@60Hz at the maximum clock frequency of 297-MHz. Moreover, the proposed VLSI circuits are fully pipelined and synchronized to the pixel clock of output video, so can be seamlessly integrated into diverse real-time video processing systems.
Xuchong Zhang, Hongbin Sun 0001, Shiqiang Chen, Nanning Zheng 0001
IEEE Trans. Circuits Syst. Video Technol.4
2018 Hierarchical and Parallel Pipelined Heterogeneous SoC for Embedded Vision Processing
abstract
Object recognition is widely used in vision computing for various applications. Traditional CPU and application specific integrated circuit for vision computing cannot provide high performance and enough flexibility, which limit the use of vision systems. In this paper, a hierarchical and parallel pipelined heterogeneous chip for object recognition is proposed to achieve high flexibility, high performance, and area efficiency. In addition, a reformulation of 3D position estimation is proposed. The method uses single precision to achieve the short computing time and accuracy requirement. The hardware resource is small. Application-specific components, such as connected component information extractor and information extraction accelerator, are designed for high performance. Reconfiguration processors and application-specific instruction set processor are introduced to improve flexibility. These components are connected to hierarchical parallel buses. The chip is fabricated in 180-nm CMOS technology and occupies 72.25 mm2with 1.09M bits on-chip memory. It delivers 204 GOPS + 665M FLOPS operations. The results show that this hierarchical and parallel pipelined heterogeneous chip is suitable for embedded vision systems.
Bin Zhang 0022, Chen Zhao 0009, Jizhong Zhao, Nanning Zheng 0001
IEEE Trans. Circuits Syst. Video Technol.5
2018 Robust Learning With Kernel Mean p-Power Error Loss
abstract
Correntropy is a second order statistical measure in kernel space, which has been successfully applied in robust learning and signal processing. In this paper, we define a nonsecond order statistical measure in kernel space, called the kernel mean- power error (KMPE), including the correntropic loss (C-Loss) as a special case. Some basic properties of KMPE are presented. In particular, we apply the KMPE to extreme learning machine (ELM) and principal component analysis (PCA), and develop two robust learning algorithms, namely ELM-KMPE and PCA-KMPE. Experimental results on synthetic and benchmark data show that the developed algorithms can achieve better performance when compared with some existing methods.
Badong Chen, Lei Xing 0003, Harry Qin, Nanning Zheng 0001
IEEE Trans. Cybern.5
2018 Joint Video Object Discovery and Segmentation by Coupled Dynamic Markov Networks
abstract
It is a challenging task to extract segmentation mask of a target from a single noisy video, which involves object discovery coupled with segmentation. To solve this challenge, we present a method to jointly discover and segment an object from a noisy video, where the target disappears intermittently throughout the video. Previous methods either only fulfill video object discovery, or video object segmentation presuming the existence of the object in each frame. We argue that jointly conducting the two tasks in a unified way will be beneficial. In other words, video object discovery and video object segmentation tasks can facilitate each other. To validate this hypothesis, we propose a principled probabilistic model, where two dynamic Markov networks are coupled-one for discovery and the other for segmentation. When conducting the Bayesian inference on this model using belief propagation, the bi-directional message passing reveals a clear collaboration between these two inference tasks. We validated our proposed method in five data sets. The first three video data sets, i.e., the SegTrack data set, the YouTube-objects data set, and the Davis data set, are not noisy, where all video frames contain the objects. The two noisy data sets, i.e., the XJTU-Stevens data set, and the Noisy-ViDiSeg data set, newly introduced in this paper, both have many frames that do not contain the objects. When compared with state of the art, it is shown that although our method produces inferior results on video data sets without noisy frames, we are able to obtain better results on video data sets with noisy frames.
Ziyi Liu 0001, Le Wang 0003, Gang Hua 0001, Qilin Zhang 0004, Zhenxing Niu, Ying Wu 0001, Nanning Zheng 0001
IEEE Trans. Image Process.7
2018 Data-Driven State-Increment Statistical Model and Its Application in Autonomous Driving
abstract
The aim of trajectory planning is to generate a feasible, collision-free trajectory to guide an autonomous vehicle from the initial state to the goal state safely. However, it is difficult to guarantee that the trajectory is feasible for the vehicle and the real path of the vehicle is collision-free when the vehicle follows the trajectory. In this paper, a state-increment statistical model (SISM) is proposed to describe the kinodynamic constraints of a vehicle by modeling the controller, the actuator, and the vehicle model jointly. The SISM consists of Gaussian distributions of lateral error increments in all state subspaces which are composed of the curvature radius, the velocity, and the lateral error. It is a data-driven modeling approach that can improve the SISM via increasing the number of samples of the increment-state, which is composed of the state and its corresponding increment of the lateral error. According to the SISM, the experience cost functions are designed to evaluate the trajectories for searching the best one with the lowest cost, and the real path can be predicted directly according to the planned trajectory and the vehicle state. The predicted path can be utilized effectually to evaluate the safety of the vehicle motion.
Chao Ma 0024, Jianru Xue, Yuehu Liu, Jing Yang 0014, Nanning Zheng 0001
IEEE Trans. Intell. Transp. Syst.6
2018 Worst Case Driven Display Frame Compression for Energy-Efficient Ultra-HD Display Processing
abstract
Display frame compression is an effective technique to address the challenge of external memory access in ultrahigh definition video display system. Nevertheless, previously proposed display frame compression designs are inadequate in terms of either energy efficiency or throughput. This paper aims to exploit the algorithm and very large scale integration (VLSI) architecture of a worst case driven display frame compression. By using a prediction-and-compression framework and a semi-fixed length coding scheme, the proposed design can achieve the much better balance between compression efficiency and throughput, and substantially reduce the bandwidth requirement and energy consumption of external memory system in the meanwhile. Extensive experiments demonstrate that the proposed display frame compression achieves 5.7-dB peak signal-to-noise ratio improvement, 3.1% compression ratio reduction, 3 × throughput, and 66.4% hardware cost saving, compared with the best previous work. In addition, the proposed VLSI design can support the throughput of 4 K × 2 K@60 Hz and reduce at least 17.6% energy consumption of external memory system by exploiting dynamic voltage and frequency scaling, compared with conventional display frame compression works.
Qiubo Chen, Hongbin Sun 0001, Nanning Zheng 0001
IEEE Trans. Multim.3
2018 Large Margin Learning in Set-to-Set Similarity Comparison for Person Reidentification
abstract
Person reidentification aims at matching images of the same person across disjoint camera views, which is a challenging problem in multimedia analysis, multimedia editing, and content-based media retrieval communities. The major challenge lies in how to preserve similarity of the same person across video footages with large appearance variations, while discriminating different individuals. To address this problem, conventional methods usually consider the pairwise similarity between persons by only measuring the point-to-point distance. In this paper, we propose using a deep learning technique to model a novel set-to-set (S2S) distance, in which the underline objective focuses on preserving the compactness of intraclass samples for each camera view, while maximizing the margin between the intraclass set and interclass set. The S2S distance metric consists of three terms, namely, the class-identity term, the relative distance term, and the regularization term. The class-identity term keeps the intraclass samples within each camera view gathering together, the relative distance term maximizes the distance between the intraclass class set and interclass set across different camera views, and the regularization term smoothes the parameters of the deep convolutional neural network. As a result, the final learned deep model can effectively find out the matched target to the probe object among various candidates in the video gallery by learning discriminative and stable feature representations. Using the CUHK01, CUHK03, PRID2011, and Market1501 benchmark datasets, we extensively conducted comparative evaluations to demonstrate the advantages of our method over the state-of-the-art approaches.
Sanping Zhou, Jinjun Wang, Qiqi Hou, Yihong Gong, Nanning Zheng 0001
IEEE Trans. Multim.6
2018 Improving CNN Performance Accuracies With Min-Max Objective
abstract
We propose a novel method for improving performance accuracies of convolutional neural network (CNN) without the need to increase the network complexity. We accomplish the goal by applying the proposed Min-Max objective to a layer below the output layer of a CNN model in the course of training. The Min-Max objective explicitly ensures that the feature maps learned by a CNN model have the minimum within-manifold distance for each object manifold and the maximum between-manifold distances among different object manifolds. The Min-Max objective is general and able to be applied to different CNNs with insignificant increases in computation cost. Moreover, an incremental minibatch training procedure is also proposed in conjunction with the Min-Max objective to enable the handling of large-scale training data. Comprehensive experimental evaluations on several benchmark data sets with both the image classification and face verification tasks reveal that employing the proposed Min-Max objective in the training process can remarkably improve performance accuracies of a CNN model in comparison with the same model trained without using this objective.
Weiwei Shi 0003, Yihong Gong, Jinjun Wang, Nanning Zheng 0001
IEEE Trans. Neural Networks Learn. Syst.5
2018 Training DCNN by Combining Max-Margin, Max-Correlation Objectives, and Correntropy Loss for Multilabel Image Classification
abstract
In this paper, we build a multilabel image classifier using a general deep convolutional neural network (DCNN). We propose a novel objective function that consists of three parts, i.e., max-margin objective, max-correlation objective, and correntropy loss. The max-margin objective explicitly enforces that the minimum score of positive labels must be larger than the maximum score of negative labels by a predefined margin, which not only improves accuracies of the multilabel classifier, but also eases the threshold determination. The max-correlation objective can make the DCNN model learn a latent semantic space, which maximizes the correlations between the feature vectors of the training samples and their corresponding ground-truth label vectors projected into this space. Instead of using the traditional softmax loss, we adopt the correntropy loss from the information theory field to minimize the training errors of the DCNN model. The proposed framework can be end-to-end trained. Comprehensive experimental evaluations on Pascal VOC 2007 and MIR Flickr 25K multilabel benchmark data sets with four DCNN models, i.e., AlexNet, VGG-16, GoogLeNet, and ResNet demonstrate that the proposed objective function can remarkably improve the performance accuracies of a DCNN model for the task of multilabel image classification.
Weiwei Shi 0003, Yihong Gong, Nanning Zheng 0001
IEEE Trans. Neural Networks Learn. Syst.4
2018 Exploring Customizable Heterogeneous Power Distribution and Management for Datacenter
abstract
Large-scale datacenters are facing increasing pressure of capping their carbon emission and power cost. Many leading-edge studies have started to explore server clusters running on multiple power sources. Existing approaches do not sufficiently consider the fine-grained power delivery to satisfy diverse requirements in datacenter, especially in the multi-tenant/colocation datacenter, which may yield low energy utilization. To address the emerging trend and new requirements, this article proposes a novel Datacenter inner Power Switch Network (DiPSN) to improve datacenter power efficiency and user satisfaction. DiPSN is a reconfigurable and easy-to-scale-out power architecture, which enables datacenter to distribute various power sources in a fine-grained manner. Moreover, a tailored machine learning based power source management framework is proposed for DiPSN to dynamically optimize user customized performance metrics and maximize datacenter revenue. Compared with conventional single-switch power distribution system, our DiPSN can be configured to improve solar energy utilization by 39.6 percent, reduce utility power cost by 11.1 percent and improve workload performance by 33.8 percent. Meanwhile, our design can extend battery lifetime by 9.3 percent. This work could provide valuable guidelines for designing heterogeneous power distribution architecture and management methodology in datacenters for improving user-customizable efficiency, sustainability and economy.
Longjun Liu, Hongbin Sun 0001, Chao Li 0009, Yang Hu 0001, Tao Li 0006, Nanning Zheng 0001
IEEE Trans. Parallel Distributed Syst.6
2018 A Limb-Based Graphical Model for Human Pose Estimation
abstract
Modeling the relationship among human joints is one of the most important components in human pose estimation. Most of previous methods define this relationship as a geometric constraint on the relative locations of two neighboring joints. In this constraint, the local appearance of the region connecting two neighboring joints is ignored. However, discarding this image appearance leads to some severe problems, such as double-counting and localization failure when the human pose is rare in the training dataset. Moreover, this image appearance, called human limb, plays an important role in human pose estimation in human visual system. Due to these reasons, we propose to solve a new task: human limb detection, which aims at detecting and representing this local image appearance. We combine this task with human joint localization as a unified framework. After getting the initial detections, we design a two-steps graphical model to capture the spatial relationship among human joints and limbs in a coarse to fine way. We evaluate the proposed method on two widely used datasets for human pose estimation: 1) frame labeled in cinema and 2) leeds sports pose datasets. The experiments results show the effectiveness of our method.
Guoqiang Liang 0001, Xuguang Lan, Jiang Wang 0001, Jianji Wang 0001, Nanning Zheng 0001
IEEE Trans. Syst. Man Cybern. Syst.5
2017 POSTER: Accelerate GPU Concurrent Kernel Execution by Mitigating Memory Pipeline Stalls
abstract
In this study, we demonstrate that the performance may be undermined in the state-of-the-art intra-SM sharing schemes for concurrent kernel execution (CKE) on GPUs, due to the interference among concurrent kernels. We highlight that cache partitioning techniques proposed for CPUs are not effective for GPUs. Then we propose to balance memory accesses and limit the number of inflight memory instructions issued from concurrent kernels to reduce memory pipeline stalls. Our proposed schemes significantly improve the performance of two state-of-the-art intra-SM sharing schemes, Warped-Slicer and SMK.
Hongwen Dai, Chao Li 0004, Chen Zhao 0009, Fei Wang 0008, Nanning Zheng 0001, Huiyang Zhou
PACT6
2017 ER3: A Unified Framework for Event Retrieval, Recognition and Recounting
abstract
We develop a unified framework for complex event retrieval, recognition and recounting. The framework is based on a compact video representation that exploits the temporal correlations in image features. Our feature alignment procedure identifies and removes the feature redundancies across frames and outputs an intermediate tensor representation we call video imprint. The video imprint is then fed into a reasoning network, whose attention mechanism parallels that of memory networks used in language modeling. The reasoning network simultaneously recognizes the event category and locates the key pieces of evidence for event recounting. In event retrieval tasks, we show that the compact video representation aggregated from the video imprint achieves significantly better retrieval accuracy compared with existing methods. We also set new state of the art results in event recognition tasks with an additional benefit: The latent structure in our reasoning network highlights the areas of the video imprint and can be directly used for event recounting. As video imprint maps back to locations in the video frames, the network allows not only the identification of key frames but also specific areas inside each frame which are most influential to the decision process.
Zhanning Gao, Gang Hua 0001, Dongqing Zhang, Nebojsa Jojic, Le Wang 0003, Jianru Xue, Nanning Zheng 0001
CVPR7
2017 Point to Set Similarity Based Deep Feature Learning for Person Re-Identification
abstract
Person re-identification (Re-ID) remains a challenging problem due to significant appearance changes caused by variations in view angle, background clutter, illumination condition and mutual occlusion. To address these issues, conventional methods usually focus on proposing robust feature representation or learning metric transformation based on pairwise similarity, using Fisher-type criterion. The recent development in deep learning based approaches address the two processes in a joint fashion and have achieved promising progress. One of the key issues for deep learning based person Re-ID is the selection of proper similarity comparison criteria, and the performance of learned features using existing criterion based on pairwise similarity is still limited, because only P2P distances are mostly considered. In this paper, we present a novel person Re-ID method based on P2S similarity comparison. The P2S metric can jointly minimize the intra-class distance and maximize the inter-class distance, while back-propagating the gradient to optimize parameters of the deep model. By utilizing our proposed P2S metric, the learned deep model can effectively distinguish different persons by learning discriminative and stable feature representations. Comprehensive experimental evaluations on 3DPeS, CUHK01, PRID2011 and Market1501 datasets demonstrate the advantages of our method over the state-of-the-art approaches.
Sanping Zhou, Jinjun Wang, Yihong Gong, Nanning Zheng 0001
CVPR5
2017 View Adaptive Recurrent Neural Networks for High Performance Human Action Recognition from Skeleton Data
abstract
Skeleton-based human action recognition has recently attracted increasing attention due to the popularity of 3D skeleton data. One main challenge lies in the large view variations in captured human actions. We propose a novel view adaptation scheme to automatically regulate observation viewpoints during the occurrence of an action. Rather than re-positioning the skeletons based on a human defined prior criterion, we design a view adaptive recurrent neural network (RNN) with LSTM architecture, which enables the network itself to adapt to the most suitable observation viewpoints from end to end. Extensive experiment analyses show that the proposed view adaptive RNN model strives to (1) transform the skeletons of various views to much more consistent viewpoints and (2) maintain the continuity of the action rather than transforming every frame to the same position with the same body orientation. Our model achieves significant improvement over the state-of-the-art approaches on three benchmark datasets.
Pengfei Zhang 0005, Cuiling Lan, Junliang Xing, Wenjun Zeng 0001, Jianru Xue, Nanning Zheng 0001
ICCV6
2017 Discriminative Dictionary Learning With Ranking Metric Embedded for Person Re-Identification
abstract
The goal of person re-identification (Re-Id) is to match pedestrians captured from multiple non-overlapping cameras. In this paper, we propose a novel dictionary learning based method with the ranking metric embedded, for person Re-Id. A new and essential ranking graph Laplacian term is introduced, which minimizes the intra-personal compactness and maximizes the inter-personal dispersion in the objective. Different from the traditional dictionary learning based approaches and their extensions, which just use the same or not information, our proposed method can explore the ranking relationship among the person images, which is essential for such retrieval related tasks. Simultaneously, one distance measurement has been explicitly learned in the model to further improve the performance. Since we have reformulated these ranking constraints into the graph Laplacian form, the proposed method is easy-to-implement but effective. We conduct extensive experiments on three widely used person Re-Id benchmark datasets, and achieve state-of-the-art performances.
De Cheng, Xiaojun Chang, Li Liu 0031, Alex Hauptmann 0001, Yihong Gong, Nanning Zheng 0001
IJCAI6
2017 Inferring Human Attention by Learning Latent Intentions
abstract
This paper addresses the problem of inferring 3D human attention in RGB-D videos at scene scale. 3D human attention describes where a human is looking in 3D scenes. We propose a probabilistic method to jointly model attention, intentions, and their interactions. Latent intentions guide human attention which conversely reveals the intention features. This mutual interaction makes attention inference a joint optimization with latent intentions. An EM-based approach is adopted to learn the latent intentions and model parameters. Given an RGB-D video with 3D human skeletons, a joint-state dynamic programming algorithm is utilized to jointly infer the latent intentions, the 3D attention directions, and the attention voxels in scene point clouds. Experiments on a new 3D human attention dataset prove the strength of our method.
Ping Wei 0001, Dan Xie 0005, Nanning Zheng 0001, Song-Chun Zhu
IJCAI3
2017 Random fourier feature kernel recursive least squares
abstract
In this paper, we investigate the nonlinear, finite dimensional and data independent random Fourier feature expansions that can approximate the popular Gaussian kernel. With recursive least squares algorithm, we develop the Random Fourier Feature Recursive Least Squares algorithm (RFF-RLS), which shows significant performance improvements in simulations when compared with several other online kernel learning algorithms such as Kernel Least Mean Square (KLMS) and Kerne Recursive Least Squares (KRLS). Our results confirm that the RFF-RLS can achieve desirable performance with low computational cost. As for the random Fourier features, the randomization generally results in redundancy. We use an algorithm, namely, Vector Quantization with Information Theoretic Learning (VQIT) to decrease the dictionary size. The resulting sparse dictionary can match the original data distribution well. The RFF-RLS with VQIT can outperform the RFF-RLS without VQIT.
Zhengda Qin, Badong Chen, Nanning Zheng 0001
IJCNN3
2017 sWMF: Separable weighted median filter for efficient large-disparity stereo matching
abstract
Although large disparity stereo matching is critical to the practical application of stereo vision system especially for outdoor scenes, its efficient hardware design is still a grand challenge. Motivated by the discovery that well-designed weighted median filter (WMF) can achieve satisfactory accuracy with simple box-filter aggregation, this paper proposes a separable weighted median filter (sWMF) that only has the computational complexity of O(r) and is independent of disparity range. Moreover, the proposed sWMF can be efficiently implemented as a fully pipelined architecture. Evaluation results demonstrate that, at the penalty of only 0.06% disparity error rate, the proposed sWMF design can save 12.9% Slice LUTs, 76.7% DSPs and 64.0% Block RAMs at the disparity range of 128, compared with previous WMF implementation on FPGA.
Shiqiang Chen, Xuchong Zhang, Hongbin Sun 0001, Nanning Zheng 0001
ISCAS4
2017 Master general parking skill via deep learning
abstract
Parking is one basic function of autonomous vehicles. However, parking still remains difficult to be implemented, since it requires to generate a relatively long-term series of actions to reach a certain objective under complicated constraints. One recently proposed method used deep neural networks(DNN) to learn the relationship between the actual parking trajectories and the corresponding steering actions, so as to find the best parking trajectory via direct recalling. However, this method can only handle a special vehicle whose dynamic parameters are well known. In this paper, we use transfer learning technique to further extend this direct trajectory planning method and master general parking skills. We aim to mimic how human drivers make parking by using a specially designed deep neural network. The first few layers of this DNN contain the general parking trajectory planning knowledge for all kinds of vehicles; while the last few layers of this DNN can be quickly tuned to adapt various kinds of vehicles. Numerical tests show that, combining transfer learning and direct trajectory planning solution, our new approach enables automated vehicles to convey the knowledge of trajectory planning from one vehicle to another with a few try-and-tests.
Yilun Lin 0002, Li Li 0013, Xingyuan Dai, Nanning Zheng 0001, Fei-Yue Wang 0001
Intelligent Vehicles Symposium4
2017 Video Search via Ranking Network with Very Few Query Exemplars
De Cheng, Lu Jiang 0004, Yihong Gong, Nanning Zheng 0001, Alex Hauptmann 0001
MMM (2)4
2017 Exemplar-Guided Similarity Learning on Polynomial Kernel Feature Map for Person Re-identification
Dapeng Chen, Zejian Yuan, Jingdong Wang 0001, Badong Chen, Gang Hua 0001, Nanning Zheng 0001
Int. J. Comput. Vis.6
2017 Active Rectification of Curved Document Images Using Structured Beams
Gaofeng Meng, Shiming Xiang, Chunhong Pan, Nanning Zheng 0001
Int. J. Comput. Vis.4
2017 Salient Object Detection: A Discriminative Regional Feature Integration Approach
Jingdong Wang 0001, Huaizu Jiang, Zejian Yuan, Ming-Ming Cheng, Xiaowei Hu 0003, Nanning Zheng 0001
Int. J. Comput. Vis.6
2017 Part-aware trajectories association across non-overlapping uncalibrated cameras
De Cheng, Yihong Gong, Jinjun Wang, Qiqi Hou, Nanning Zheng 0001
Neurocomputing5
2017 A vision-centered multi-sensor fusing approach to self-localization and obstacle perception for robotic cars
abstract
Most state-of-the-art robotic cars’ perception systems are quite different from the way a human driver understands traffic environments. First, humans assimilate information from the traffic scene mainly through visual perception, while the machine perception of traffic environments needs to fuse information from several different kinds of sensors to meet safety-critical requirements. Second, a robotic car requires nearly 100% correct perception results for its autonomous driving, while an experienced human driver works well with dynamic traffic environments, in which machine perception could easily produce noisy perception results. In this paper, we propose a vision-centered multi-sensor fusing framework for a traffic environment perception approach to autonomous driving, which fuses camera, LIDAR, and GIS information consistently via both geometrical and semantic constraints for efficient self-localization and obstacle perception. We also discuss robust machine vision algorithms that have been successfully integrated with the framework and address multiple levels of machine vision techniques, from collecting training data, efficiently processing sensor data, and extracting low-level features, to higher-level object and environment mapping. The proposed framework has been tested extensively in actual urban scenes with our self-developed robotic cars for eight years. The empirical results validate its robustness and efficiency.
Jianru Xue, Di Wang 0028, Shaoyi Du, Dixiao Cui, Nanning Zheng 0001
Frontiers Inf. Technol. Electron. Eng.6
2017 Hybrid-augmented intelligence: collaboration and cognition
abstract
The long-term goal of artificial intelligence (AI) is to make machines learn and think like human beings. Due to the high levels of uncertainty and vulnerability in human life and the open-ended nature of problems that humans are facing, no matter how intelligent machines are, they are unable to completely replace humans. Therefore, it is necessary to introduce human cognitive capabilities or human-like cognitive models into AI systems to develop a new form of AI, that is, hybrid-augmented intelligence. This form of AI or machine intelligence is a feasible and important developing model. Hybrid-augmented intelligence can be divided into two basic models: one is human-in-the-loop augmented intelligence with human-computer collaboration, and the other is cognitive computing based augmented intelligence, in which a cognitive model is embedded in the machine learning system. This survey describes a basic framework for human-computer collaborative hybrid-augmented intelligence, and the basic elements of hybrid-augmented intelligence based on cognitive computing. These elements include intuitive reasoning, causal models, evolution of memory and knowledge, especially the role and basic principles of intuitive reasoning for complex problem solving, and the cognitive learning framework for visual scene understanding based on memory and reasoning. Several typical applications of hybrid-augmented intelligence in related fields are given.
Nanning Zheng 0001, Ziyi Liu 0001, Pengju Ren, Shi-tao Chen, Si-yu Yu, Jianru Xue, Badong Chen, Fei-Yue Wang 0001
Frontiers Inf. Technol. Electron. Eng.1
2017 Pose-and-illumination-invariant face representation via a triplet-loss trained deep reconstruction model
Xingyu Chen 0001, Xuguang Lan, Guoqiang Liang 0001, Nanning Zheng 0001
Multim. Tools Appl.5
2017 Fast additive quantization for vector compression in nearest neighbor search
Jin Li 0011, Xuguang Lan, Jiang Wang 0001, Meng Yang 0002, Nanning Zheng 0001
Multim. Tools Appl.5
2017 Multi-Timescale Collaborative Tracking
abstract
We present the multi-timescale collaborative tracker for single object tracking. The tracker simultaneously utilizes different types of "forces", namely attraction, repulsion and support, to take advantage of their complementary strengths. We model the three forces via three components that are learned from the sample sets with different timescales. The long-term descriptive component attracts the target sample, while the medium-term discriminative component repulses the target from the background. They are collaborated in the appearance model to benefit each other. The short-term regressive component combines the votes of the auxiliary samples to predict the target's position, forming the context-aware motion model. The appearance model and the motion model collaboratively determine the target state, and the optimal state is estimated by a novel coarse-to-fine search strategy. We have conducted an extensive set of experiments on the standard 50 video benchmark. The results confirm the effectiveness of each component and their collaboration, outperforming current state-of-the-art methods.
Dapeng Chen, Zejian Yuan, Gang Hua 0001, Jingdong Wang 0001, Nanning Zheng 0001
IEEE Trans. Pattern Anal. Mach. Intell.5
2017 Video Object Discovery and Co-Segmentation with Extremely Weak Supervision
abstract
We present a spatio-temporal energy minimization formulation for simultaneous video object discovery and co-segmentation across multiple videos containing irrelevant frames. Our approach overcomes a limitation that most existing video co-segmentation methods possess, i.e., they perform poorly when dealing with practical videos in which the target objects are not present in many frames. Our formulation incorporates a spatio-temporal auto-context model, which is combined with appearance modeling for superpixel labeling. The superpixel-level labels are propagated to the frame level through a multiple instance boosting algorithm with spatial reasoning, based on which frames containing the target object are identified. Our method only needs to be bootstrapped with the frame-level labels for a few video frames (e.g., usually 1 to 3) to indicate if they contain the target objects or not. Extensive experiments on four datasets validate the efficacy of our proposed method: 1) object segmentation from a single video on the SegTrack dataset, 2) object co-segmentation from multiple videos on a video co-segmentation dataset, and 3) joint object discovery and co-segmentation from multiple videos containing irrelevant frames on the MOViCS dataset and XJTU-Stevens, a new dataset that we introduce in this paper. The proposed method compares favorably with the state-of-the-art in all of these experiments.
Le Wang 0003, Gang Hua 0001, Rahul Sukthankar, Jianru Xue, Zhenxing Niu, Nanning Zheng 0001
IEEE Trans. Pattern Anal. Mach. Intell.6
2017 Modeling 4D Human-Object Interactions for Joint Event Segmentation, Recognition, and Object Localization
abstract
In this paper, we present a 4D human-object interaction (4DHOI) model for solving three vision tasks jointly: i) event segmentation from a video sequence, ii) event recognition and parsing, and iii) contextual object localization. The 4DHOI model represents the geometric, temporal, and semantic relations in daily events involving human-object interactions. In 3D space, the interactions of human poses and contextual objects are modeled by semantic co-occurrence and geometric compatibility. On the time axis, the interactions are represented as a sequence of atomic event transitions with coherent objects. The 4DHOI model is a hierarchical spatial-temporal graph representation which can be used for inferring scene functionality and object affordance. The graph structures and parameters are learned using an ordered expectation maximization algorithm which mines the spatial-temporal structures of events from RGB-D video samples. Given an input RGB-D video, the inference is performed by a dynamic programming beam search algorithm which simultaneously carries out event segmentation, recognition, and object localization. We collected a large multiview RGB-D event dataset which contains 3,815 video sequences and 383,036 RGB-D frames captured by three RGB-D cameras. The experimental results on three challenging datasets demonstrate the strength of the proposed method.
Ping Wei 0001, Yibiao Zhao, Nanning Zheng 0001, Song-Chun Zhu
IEEE Trans. Pattern Anal. Mach. Intell.3
2017 Constructing Deep Sparse Coding Network for image classification
Shizhou Zhang, Jinjun Wang, Yihong Gong, Nanning Zheng 0001
Pattern Recognit.5
2017 A new compressive sensing video coding framework based on Gaussian mixture model
Xiangwei Li, Xuguang Lan, Meng Yang 0002, Jianru Xue, Nanning Zheng 0001
Signal Process. Image Commun.5
2017 Balanced Mixture of Deformable Part Models With Automatic Part Configurations
abstract
This paper presents a method to improve the traditional mixture of deformable part models (MDPM) method from the learning perspective. First, an object part configuration learning algorithm based on group sparsity constraint is introduced to automatically discover the object part number, size, and location. The algorithm imposes two additional regularization terms in addition to the standard hinge loss function. The first term focuses on automatic part selection and the second term focuses on automatic part placement. Second, this paper introduces an improved MDPM training framework. The framework applies a learned transformation to normalize the prediction score from each individual deformable part model (DPM) into a pseudo probability such that the partition of the entire object appearance feature space becomes less sensitive to the prior distributions of different DPMs. Finally, the two proposed improvements are combined and formulated under the expectation-maximization framework. We evaluate our method mainly using the PASCAL VOC2007 and VOC2010 detection benchmarks and show that the proposed learning algorithms could increase the detection mean AP score by 2.4% and 0.9%, respectively, on these two data sets when using the proposed part selection method and the training algorithm. We also present further in-depth analysis of the proposed algorithm in the experiments.
De Cheng, Yihong Gong, Jingjun Wang, Nanning Zheng 0001
IEEE Trans. Circuits Syst. Video Technol.4
2017 Correntropy Maximization via ADMM: Application to Robust Hyperspectral Unmixing
abstract
In hyperspectral images, some spectral bands suffer from low signal-to-noise ratio due to noisy acquisition and atmospheric effects, thus requiring robust techniques for the unmixing problem. This paper presents a robust supervised spectral unmixing approach for hyperspectral images. The robustness is achieved by writing the unmixing problem as the maximization of the correntropy criterion subject to the most commonly used constraints. Two unmixing problems are derived: the first problem considers the fully constrained unmixing, with both the nonnegativity and sum-to-one constraints, while the second one deals with the nonnegativity and the sparsity promoting of the abundances. The corresponding optimization problems are solved using an alternating direction method of multipliers (ADMM) approach. Experiments on synthetic and real hyperspectral images validate the performance of the proposed algorithms for different scenarios, demonstrating that the correntropy-based unmixing with ADMM is particularly robust against highly noisy outlier bands.
Fei Zhu 0001, Abderrahim Halimi, Paul Honeine, Badong Chen, Nanning Zheng 0001
IEEE Trans. Geosci. Remote. Sens.5
2017 Haze Removal Using the Difference- Structure-Preservation Prior
abstract
Fog cover is generally present in outdoor scenes, which limits the potential for efficient information extraction from images. In this paper, the goal of the developed algorithm is to obtain an optimal transmission map as well as to remove hazes from a single input image. To solve the problem, we meticulously analyze the optical model and recast the initial transmission map under an additional boundary prior. For better preservation of the results, the difference-structure-preservation dictionary could be learned, such that the local consistency features of the transmission map could be well preserved after coefficient shrinkage. Experimental results show that the method preserves the natural appearance of the image.
Lin-Yuan He, Jizhong Zhao, Nanning Zheng 0001, Duyan Bi
IEEE Trans. Image Process.3
2017 A Novel Method of Minimizing View Synthesis Distortion Based on Its Non-Monotonicity in 3D Video
abstract
In depth-based 3D video, the view synthesis distortion (VSD), is generally measured by modeling the effect of texture and depth errors separately. With such a development, it has been referred that the VSD changes monotonically with respect to to both the texture and depth distortions. In this paper, we find that the VSD does not always change monotonically with them by both theoretical analysis and experimental test, when the effect of the texture and depth errors is considered together. Specifically, first, we prove that the VSD is non-monotonic with the texture distortion. That is, the VSD increases with the increasing texture distortion at higher distortion range but conversely decreases with it at lower range. It is different from the general scenario that only considering the effect of the texture errors. We also analytically depict their relationship with low computational cost and identify the turning point at which the change of the VSD is converted. Second, we confirm that the VSD is always monotonic with the depth distortion, which is consistent with the general scenario that only considering the effect of the depth errors. The non-monotonicity property of the VSD can be utilized to improve the viewing performance of 3D video in relevant applications, since a minimal value of the VSD exists at the turning point. We conduct two applications for this purpose. First, it is used to generate the synthesis view of minimal distortion, which achieves 0.51-dB gain of PSNR on average for the tested scenarios. Second, it is used for lossy compression of texture videos in 3D video, which reduces the coding rate by 24% on average for the tested scenarios, meanwhile, keeps the VSD not increased simultaneously.
Meng Yang 0002, Nanning Zheng 0001, Ce Zhu, Fei Wang 0008
IEEE Trans. Image Process.2
2017 Online Variable Coding Length Product Quantization for Fast Nearest Neighbor Search in Mobile Retrieval
abstract
Quantization methods are crucial for efficient nearest neighbor search in many applications such as image, music, or product search. As mobile devices are becoming increasingly more popular, the quantization methods on mobile devices are more important, because a large portion of the search queries are becoming performed on mobile devices. One important characteristic of the communication on mobile devices is the inherent unreliability of their communication channels. In order to adapt the quality changes of the communication channels, we need to change the coding length of the quantization accordingly. The existing quantization methods use fixed-length codebooks, and it is expensive to retrain another codebook with different coding length. In this paper, we propose a novel variable length product quantization framework that consists of a set of fast universal scalar quantizers. The framework is capable of producing variable length quantization without retraining the codebook. Each data vector is transformed into a new space to reduce the correlation across dimensions. A proper number of bits is allocated to represent the scalar component in each dimension according to the given coding length. For each component, we estimate its probability density function (PDF) and design an efficient universal scalar quantizer based on the PDF and the allocated bits. To reduce distortion, we learn a Gaussian mixture model for the data. The experimental results show that, compared to state-of-the-art product quantization methods, our approach can construct the codebooks online for variable coding lengths and achieve the comparable performance.
Jin Li 0011, Xuguang Lan, Xiangwei Li, Jiang Wang 0001, Nanning Zheng 0001, Ying Wu 0001
IEEE Trans. Multim.5