Jingmin Xin

dblp:60/1120 · DBLP profile ↗
← Back
52ranked-venue papers
3as first author
28since 2021 · last 2026
0000-0003-1906-2327ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 21 · 13 since 2021Graphics, computer vision, multimedia, augmented reality and games · 15 · 3 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 13 · 9 since 2021Systems, architecture and hardware · 8 · 4 since 2021Software engineering, systems software and programming languages · 1Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 VAV-R1: Difficulty-Aware Multimodal Reasoning for Video Anomaly Validation
abstract
Video anomaly validation (VAV) serves as the final alarm validation step for false positive filtering in video anomaly detection (VAD), requiring both higher accuracy in anomaly identification and stronger justifications. Despite rapid VAD advancements, existing methods still lack sufficient interpretability and struggle with challenging cases. To address this, we propose VAV-R1, a multimodal reasoning model tailored for VAV. We first construct ThinkVAV, a dedicated benchmark with fine-grained reasoning annotations across diverse anomaly types. Furthermore, we introduce DA-GRPO, a difficulty-aware reinforcement learning strategy that prioritizes learning from more challenging cases. Extensive experiments demonstrate that VAV-R1 achieves SOTA performance across multiple tasks.
Qianhao Ren, Yutong Wang 0001, Zhongjiang He, Jingmin Xin, Hao Sun 0038
ICMR7
2026 Efficient DOA Estimation Based on Coprime Array Interpolation With Deep Unfolding Network
abstract
Coprime arrays increase the degrees of freedom for direction of arrival (DOA) estimation, but virtual-array gaps require costly filling procedures. In this letter, a new DOA estimation method is proposed for coprime arrays based on interpolation and a deep unfolding network. The reconstruction of the interpolated virtual array covariance matrix is formulated as a rank minimization problem and solved using an ADMM-based deep unfolding network with stage-wise learnable parameters and an unsupervised loss inspired by ADMM convergence criteria. Finally, root-MUSIC is employed for DOA estimation. Simulations demonstrate the effectiveness of the proposed method in terms of both computational efficiency and estimation performance.
Zhuoqian Jiang, Jingmin Xin, Weiliang Zuo, Nanning Zheng 0001, Akira Sano
IEEE Signal Process. Lett.2
2026 CNDM: Customized Noise Diffusion Model for Trajectory Prediction
abstract
Precise and reliable multi-agent trajectory prediction is fundamental to enabling safe autonomous navigation. While denoising diffusion models have demonstrated remarkable capability in capturing the multimodal uncertainty inherent in this task, their reliance on a generic, isotropic Gaussian noise prior poses a critical limitation. This uniform prior is fundamentally misaligned with the highly structured, heterogeneous uncertainty of agent motion—shaped by dynamics, type, and environmental constraints—forcing the denoising network to implicitly relearn complex motion priors from scratch. To bridge this gap, we propose the Customized Noise Diffusion Model (CNDM), a novel framework that introduces a learned, agent-specific noise prior. At the core of CNDM is a Prior-Guidance Network (PGN) that distills an agent’s history, type, and scene context into a parametric, anisotropic Gaussian distribution. This customized prior provides a physically-grounded starting point for the diffusion process. To enable efficient training on such anisotropic noise, we leverage a Mahalanobis whitening transformation to standardize the denoising task. Extensive experiments on the Waymo Open Motion and Argoverse 2 datasets show that CNDM achieves competitive performance against state-of-the-art methods, excelling particularly in capturing diverse motion patterns and improving probabilistic calibration. Ablation studies confirm that the performance gains are directly attributable to our customized noise design, underscoring the importance of integrating structured domain knowledge into the generative foundation of diffusion models.
Entao Chang, Jiawei Fu 0001, Wenjie Gao 0001, Shi-tao Chen, Jingmin Xin, Nanning Zheng 0001
IEEE Trans. Intell. Transp. Syst.5
2026 FlowCalib: Targetless Infrastructure LiDAR-Camera Extrinsic Calibration Based on Optical Flow and Scene Flow
abstract
Recently, multi-sensor fusion-based vehicle infrastructure cooperative perception has aroused extensive attention due to the demands for the safety of autonomous driving and traffic monitoring. An accurate calibration between different sensors is a critical foundation for most sensor fusion systems. For LiDAR-camera calibration, high accuracy can be achieved with the help of artificial calibration targets, such as a checkerboard. However, unlike autonomous vehicles, roadside sensors monitor traffic scenes with continuous traffic flow from a fixed viewpoint, posing challenges for conventional calibration methods. There, a calibration method suitable for roadside scenes is required for infrastructure sensors. In this paper, we propose FlowCalib, a novel targetless infrastructure LiDAR-camera spatial calibration method through alignment of scene flow and optical flow. The main idea is to leverage the inherent consistency of moving objects in traffic flow across two types of sensor data. Firstly, the moving objects are extracted by optical flow and scene flow. Then, the extrinsic parameters are obtained in two steps: rough calibration and calibration refinement. In rough calibration, the center and motion flow of each moving instance are calculated by clustering methods separately in the point cloud and image. Based on this, the possible initial value set of extrinsic parameters is estimated by two-step parameter sampling. The initial parameters are obtained by distance of center and motion flow in point cloud and image based scoring. Subsequently, the extrinsic parameters are refined by optimization of instance alignment loss and flow alignment loss of moving objects. In the end, quantitative and qualitative experiments are conducted to validate the effectiveness of the algorithm across both simulated datasets and real-world datasets.
Renwei Hai, Yanqing Shen, Shi-tao Chen, Jingmin Xin, Nanning Zheng 0001
IEEE Trans. Intell. Transp. Syst.5
2026 PET Head Motion Estimation Using Supervised Deep Learning With Attention
abstract
Head movement poses a significant challenge in brain positron emission tomography (PET) imaging, resulting in image artifacts and tracer uptake quantification inaccuracies. Effective head motion estimation and correction are crucial for precise quantitative image analysis and accurate diagnosis of neurological disorders. Hardware-based motion tracking (HMT) has limited applicability in real-world clinical practice. To overcome this limitation, we propose a deep-learning head motion correction approach with cross-attention (DL-HMC++) to predict rigid head motion from one-second 3D PET raw data. DL-HMC++ is trained in a supervised manner by leveraging existing dynamic PET scans with gold-standard motion measurements from external HMT. We evaluate DL-HMC++ on two PET scanners (HRRT and mCT) and four radiotracers (18F-FDG,18F-FPEB,11C-UCB-J, and11C-LSN3172176) to demonstrate the effectiveness and generalization of the approach in large cohort PET studies. Quantitative and qualitative results demonstrate that DL-HMC++ consistently outperforms state-of-the-art data-driven motion estimation methods, producing motion-free images with clear delineation of brain structures and reduced motion artifacts that are indistinguishable from gold-standard HMT. Brain region of interest standard uptake value analysis exhibits average difference ratios between DL-HMC++ and gold-standard HMT to be 1.2±0.5% for HRRT and 0.5±0.2% for mCT. DL-HMC++ demonstrates the potential for data-driven PET head motion correction to remove the burden of HMT, making motion correction accessible to clinical populations beyond research settings. The code is available at https://github.com/maxxxxxxcai/DL-HMC-TMI.
Zhuotong Cai, Tianyi Zeng, Eléonore V. Lieffrig, Kathryn Fontaine, Chenyu You, Enette Mae Revilla, James S. Duncan, Jingmin Xin, Yihuan Lu, John A. Onofrey
IEEE Trans. Medical Imaging9
2025 Mind the Gap: Aligning Vision Foundation Models to Image Feature Matching
abstract
Leveraging the vision foundation models has emerged as a mainstream paradigm that improves the performance of image feature matching. However, previous works have ignored the misalignment when introducing the foundation models into feature matching. The misalignment arises from the discrepancy between the foundation models focusing on single-image understanding and the cross-image understanding requirement of feature matching. Specifically, 1) the embeddings derived from commonly used foundation models exhibit discrepancies with the optimal embeddings required for feature matching; 2) lacking an effective mechanism to leverage the single-image understanding ability into cross-image understanding. A significant consequence of the misalignment is they struggle when addressing multi-instance feature matching problems. To address this, we introduce a simple but effective framework, called IMD (Image feature Matching with a pre-trained Diffusion model) with two parts: 1) Unlike the dominant solutions employing contrastive-learning based foundation models that emphasize global semantics, we integrate the generative-based diffusion models to effectively capture instance-level details. 2) We leverage the prompt mechanism in generative model as a natural tunnel, propose a novel cross-image interaction prompting module to facilitate bidirectional information interaction between image pairs. To more accurately measure the misalignment, we propose a new benchmark called IMIM, which focuses on multi-instance scenarios. Our proposed IMD establishes a new state-of-the-art in commonly evaluated benchmarks, and the superior improvement 12% in IMIM indicates our method efficiently mitigates the misalignment.
Yuhan Liu 0006, Jingwen Fu, Yang Wu 0001, Kangyi Wu, Pengna Li, Jiayi Wu 0002, Sanping Zhou, Jingmin Xin
ICCV8
2025 Modeling Human-like Driving Behavior Based on Maximum Entropy Deep Inverse Reinforcement Learning
abstract
Modeling expert driving behavior is crucial for the successful implementation of human-like autonomous driving. In this paper, we propose a new sampling-based Maximum Entropy Deep Inverse Reinforcement Learning (MEDIRL) framework. It leverages naturalistic human driving data to train the reward model and thus evaluates driving behaviors from the reward of sampled candidate trajectories. The proposed framework utilizes deep neural networks to learn the feature-reward mapping, which offers superior fitting capabilities compared to traditional linear reward functions. A polynomial trajectory sampler for long-term decision making and a dynamic window trajectory sampler for short-term planning are adopted to simplify the calculation of partition function in the MEDIRL algorithm. In addition, the proposed framework offers a solution to the probability estimation of driving behaviors by calculating the likelihood of sampled candidate trajectories based on their reward values. Comparative experiments are conducted on the NGSIM US-101 Highway dataset, and the experimental results demonstrate the superiority of the proposed model in personalizing reward functions, as well as the applicability of the proposed method in modeling driving behaviors across various time horizons.
Jiamin Shi, Tangyike Zhang, Shi-tao Chen, Nanning Zheng 0001, Jingmin Xin
IROS5
2025 Style mixup enhanced disentanglement learning for unsupervised domain adaptation in medical image segmentation
Zhuotong Cai, Jingmin Xin, Chenyu You, Peiwen Shi, Siyuan Dong, Nicha C. Dvornek, Nanning Zheng 0001, James S. Duncan
Medical Image Anal.2
2025 Learning from open-set noisy labels based on multi-prototype modeling
Yue Zhang 0025, Chaowei Fang, Jiayi Wu 0002, Jingmin Xin
Pattern Recognit.6
2025 An information bottleneck approach for feature selection
Qi Zhang 0089, Mingfei Lu, Shujian Yu, Jingmin Xin, Badong Chen
Pattern Recognit.4
2025 PRGS: Patch-to-Region Graph Search for Visual Place Recognition
Weiliang Zuo, Liguo Liu, Yanqing Shen, Fuhua Xiang, Jingmin Xin, Nanning Zheng 0001
Pattern Recognit.6
2025 RankTuning: Cross-Image Partial Tuning Strategies for Rank Optimization in Visual Place Recognition
abstract
Aiming to estimate the location, a common strategy of Visual Place Recognition (VPR) involves utilizing global retrieval to get top-k candidates first and performing local feature matching in candidates for reranking. Although local reranking methods bring performance gains, they need a lot of computational overhead. To narrow the performance gap between global retrieval and local reranking methods with little cost, one method is to rerank candidates with global features. However, previous works only utilized the information from positive samples in candidates, ignoring the fact that negative samples can also provide useful information. To this end, we propose RankTuning, a method that aggregates all the information from candidates using global features for reranking. Specifically, we design a cross-image interaction module that allows all candidates to interact with others to enhance the discriminative power of features. Furthermore, to drive the training of this module, we propose Generalized Recall loss to handle hard samples with a better gradient strategy. Experimental results demonstrate that our method can be easily inserted into existing architectures and achieve state-of-the-art performance. Meanwhile, our method does not require additional storage overhead, and the matching latency is only 6.3% of that of the current fastest local reranking method. The code is released athttps://github.com/LKELN/RankTuning.git
Liguo Liu, Weiliang Zuo, Jingwen Fu, Yanqing Shen, Jingmin Xin, Nanning Zheng 0001
IEEE Trans. Intell. Transp. Syst.5
2024 GSO-Net: Grid Surface Optimization via Learning Geometric Constraints
abstract
In the context of surface representations, we find a natural structural similarity between grid surface and image data. Motivated by this inspiration, we propose a novel approach: encoding grid surfaces as geometric images and using image processing methods to address surface optimization-related problems. As a result, we have created the first dataset for grid surface optimization and devised a learning-based grid surface optimization network specifically tailored to geometric images, addressing the surface optimization problem through a data-driven learning of geometric constraints paradigm. We conduct extensive experiments on developable surface optimization, surface flattening, and surface denoising tasks using the designed network and datasets. The results demonstrate that our proposed method not only addresses the surface optimization problem better than traditional numerical optimization methods, especially for complex surfaces, but also boosts the optimization speed by multiple orders of magnitude. This pioneering study successfully applies deep learning methods to the field of surface optimization and provides a new solution paradigm for similar tasks, which will provide inspiration and guidance for future developments in the field of discrete surface optimization. The code and dataset are available at https://github.com/chaoyunwang/GSO-Net.
Chaoyun Wang, Jingmin Xin, Nanning Zheng 0001, Caigui Jiang
AAAI2
2024 RCNet: A Redundant Compression Network Using Information Bottleneck for Pathology Whole Slide Image Classification
abstract
The analysis of Whole Slide Images (WSIs) is vital for tumor diagnosis, and numerous deep learning methods have been extensively researched. However, the large size and multi-scale features of WSIs introduce substantial redundant information, leading to significant performance challenges. Existing methods have struggled to address this redundancy effectively. To reduce redundant information, we propose a novel Redundant Compression Network (RCNet) for WSI classification, incorporating an effective Multi-Scale Fusion Module (MSFM) and an Information Bottleneck Compression Module (IBCM). Specifically, the MSFM filters the redundancy in multi-scale features by emphasizing critical scales, while the IBCM eliminates redundant instances through compression using the information bottleneck technique. We evaluate our method on two datasets, the public DigestPath2019 dataset and a private lung pathology dataset. On the public dataset, our approach achieves improvements of at least 2.4% in F1-score, 2.3% in accuracy, 3.8% in Recall and 1.0% in AUC compared to state-of-the-art methods. Additionally, on the private dataset, our approach achieves improvements of at least 1.6% in F1-score, 1.2% in accuracy and 0.8% in AUC. These results demonstrate the effectiveness of our method in improving WSI classification by reducing redundancy and focusing on the most essential information.
Hongxuan Yu, Jiayi Wu 0002, Jichen Xu, Shuhao Wang, Siyi Chai, Jingmin Xin
BIBM7
2024 Symmetric Consistency with Cross-Domain Mixup for Cross-Modality Cardiac Segmentation
abstract
Accurate cardiac segmentation in cross-modality images plays an important role in the quantitative analysis of the heart to diagnose cardiovascular diseases. However, achieving high performance in cross-modality segmentation is hindered by the time-consuming annotation and modality gap. While some approaches employ Unsupervised Domain Adaptation (UDA) through adversarial learning to address the issue, it still remains challenging due to the instability of the adversarial generative models. In this work, we propose Symmetric Consistency with Cross-Domain Mixup (SCCDM), integrated with the teacher-student model for cross-modality cardiac segmentation. Specifically, we introduce symmetric consistency across the domains for two mixed data to diversify the data distribution from both the source domain and target domain. Extensive experiments on a public cardiac dataset demonstrate that SCCDM achieves superior domain adaptation performance for cardiac segmentation compared to state-of-the-art methods.
Zhuotong Cai, Jingmin Xin, Siyuan Dong, John A. Onofrey, Nanning Zheng 0001, James S. Duncan
ICASSP2
2024 Class-Aware Mutual Mixup with Triple Alignments for Semi-supervised Cross-Domain Segmentation
Zhuotong Cai, Jingmin Xin, Tianyi Zeng, Siyuan Dong, Nanning Zheng 0001, James S. Duncan
MICCAI (8)2
2023 Unsupervised Domain Adaptation by Cross-Prototype Contrastive Learning for Medical Image Segmentation
abstract
Unsupervised Domain Adaptation (UDA), which aligns the labeled source distribution to the unlabeled target distribution, has shown remarkable achievement in the medical image segmentation task. Previous UDA methods unilaterally consider the global distribution alignment through explicit category-based loss while good separation and discrimination of class are insufficiently explored, resulting in the sub-aligned distribution across domains. In this paper, we propose cross-prototype contrastive learning method (CPCL) for UDA segmentation through class centroid alignment. Specifically, to reduce the intra-class distance and increase the inter-class distance, we first introduce prototype-feature contrastive learning to align the pixel-level features and the same-class global prototype across domains. Secondly, we further present prototype-prototype contrastive learning to align the same class prototypes between the source domain and target domain for compact category centroid and better global domain distribution alignment. Extensive experiments on two public cardiac datasets demonstrate that the proposed CPCL achieves superior domain adaptation performance as compared with the state-of-the-art.
Zhuotong Cai, Jingmin Xin, Siyuan Dong, Chenyu You, Peiwen Shi, Tianyi Zeng, John A. Onofrey, Nanning Zheng 0001, James S. Duncan
BIBM2
2023 Cervical Cytology Classification with Coarse Labels Based on Two-Stage Weakly Supervised Contrastive Learning Framework
abstract
Deep learning methods have achieved remarkable success in various tasks from cervical cytology images. However, for the gigapixel whole slide images (WSIs), the acquisition of annotations is a time-consuming and labor-intensive task requiring a high level of expertise. While the sparse distribution of malignant cells and the factors above pose great difficulties to label thousands of patches divided from the WSI, it is much easier to obtain the coarse labels at the WSI level. In this paper, we propose a novel weakly supervised contrastive learning framework, which utilizes only coarse labels from the WSIs for cervical cytology patch classification. The proposed framework consists of two stages, including the representation learning stage and the classifier finetuning stage. In the first stage, to effectively exploit useful information of coarse labels, we devise a re-weight cross-entropy loss, which can fast warm up the training and reduce the inexact supervision from the coarse labels simultaneously. To further excavate features bypassing the coarse labels, we propose a self-supervised contrastive loss, where the random augmentation and the mean teacher architecture enrich the external variations, and help better extract representations through patch similarities. In the second stage, based on ensemble predictions and uncertainty selections, reliable pseudo labels are generated for the inaccurate labels to finetune the classifier, with better performance achieved. Extensive experiments on the in-house dataset demonstrate that the proposed method is more efficient than other state-of-the-art methods. Our code is available on https://github.com/chaisiyii/WSCL.
Siyi Chai, Jingmin Xin, Jiayi Wu 0002, Hongxuan Yu, Zhaohai Liang, Nanning Zheng 0001
BIBM2
2023 InteractionNet: Joint Planning and Prediction for Autonomous Driving with Transformers
abstract
Planning and prediction are two important modules of autonomous driving and have experienced tremendous advancement recently. Nevertheless, most existing methods regard planning and prediction as independent and ignore the correlation between them, leading to the lack of consideration for interaction and dynamic changes of traffic scenarios. To address this challenge, we propose InteractionNet, which leverages transformer to share global contextual reasoning among all traffic participants to capture interaction and interconnect planning and prediction to achieve joint. Besides, InteractionNet deploys another transformer to help the model pay extra attention to the perceived region containing critical or unseen vehicles. InteractionNet outperforms other baselines in several benchmarks, especially in terms of safety, which benefits from the joint consideration of planning and forecasting. The code will be available at https://github.com/fujiawei0724/InteractionNet.
Jiawei Fu 0001, Yanqing Shen, Zhiqiang Jian, Shi-tao Chen, Jingmin Xin, Nanning Zheng 0001
IROS5
2023 Efficient Lane-changing Behavior Planning via Reinforcement Learning with Imitation Learning Initialization
abstract
Robust lane-changing behavior planning is critical to ensuring the safety and comfort of autonomous vehicles. In this paper, we proposed an efficient and robust vehicle lane-changing behavior decision-making method based on reinforcement learning (RL) and imitation learning (IL) initialization which learns the potential lane-changing driving mechanisms from driving mechanism from the interactions between vehicle and environment, so as to simplify the manual driving modeling and have good adaptability to the dynamic changes of lane-changing scene. Our method further makes the following improvements on the basis of the Proximal Policy Optimization (PPO) algorithm: (1) A dynamic hybrid reward mechanism for lane-changing tasks is adopted; (2) A state space construction method based on fuzzy logic and deformation pose is presented to enable behavior planning to learn more refined tactical decision-making; (3) An RL initialization method based on imitation learning which only requires a small amount of scene data is introduced to solve the low efficiency of RL learning under sparse reward. Experiments on the SUMO show the effectiveness of the proposed method, and the test on the CARLA simulator also verifies the generalization ability of the method.
Jiamin Shi, Tangyike Zhang, Junxiang Zhan, Shi-tao Chen, Jingmin Xin, Nanning Zheng 0001
IV5
2023 Multi-view Contour-constrained Transformer Network for Thin-cap Fibroatheroma Identification
abstract
Identification and detection of thin-cap fibroatheroma (TCFA) from intravascular optical coherence tomography (IVOCT) images is critical for treatment of coronary heart diseases. Recently, deep learning methods have shown promising successes in TCFA identification. However, most methods usually do not effectively utilize multi-view information or incorporate prior domain knowledge. In this paper, we propose a multi-view contour-constrained transformer network (MVCTN) for TCFA identification in IVOCT images. Inspired by the diagnosis process of cardiologists, we use contour constrained self-attention modules (CCSM) to emphasize features corresponding to salient regions (i.e., vessel walls) in an unsupervised manner and enhance the visual interpretability based on class activation mapping (CAM). Moreover, we exploit transformer modules (TM) to build global-range relations between two views (i.e., polar and Cartesian views) to effectively fuse features at multiple feature scales. Experimental results on a semi-public dataset and an in-house dataset demonstrate that the proposed MVCTN outperforms other single-view and multi-view methods. Lastly, the proposed MVCTN can also provide meaningful visualization for cardiologists via CAM.
Jingmin Xin, Jiayi Wu 0002, Yangyang Deng, Ruisheng Su, Wiro J. Niessen, Nanning Zheng 0001, Theo van Walsum
Neurocomputing2
2023 Multi-weight susceptible-infected model for predicting COVID-19 in China
Jun Zhang 0003, Nanning Zheng 0001, Dingyi Yao, Jianji Wang 0001, Jingmin Xin
Neurocomputing7
2023 Onboard Sensors-Based Self-Localization for Autonomous Vehicle With Hierarchical Map
abstract
Localization is a fundamental and crucial module for autonomous vehicles. Most of the existing localization methodologies, such as signal-dependent methods (RTK-GPS and Bluetooth), simultaneous localization and mapping (SLAM), and map-based methods, have been utilized in outdoor autonomous driving vehicles and indoor robot positioning. However, they suffer from severe limitations, such as signal-blocked scenes of GPS, computing resource occupation explosion in large-scale scenarios, intolerable time delay, and registration divergence of SLAM/map-based methods. In this article, a self-localization framework, without relying on GPS or any other wireless signals, is proposed. We demonstrate that the proposed homogeneous normal distribution transform algorithm and two-way information interaction mechanism could achieve centimeter-level localization accuracy, which reaches the requirement of autonomous vehicle localization for instantaneity and robustness. In addition, benefitting from hardware and software co-design, the proposed localization approach is extremely light-weighted enough to be operated on an embedded computing system, which is different from other LiDAR localization methods relying on high-performance CPU/GPU. Experiments on a public dataset (Baidu Apollo SouthBay dataset) and real-world verified the effectiveness and advantages of our approach compared with other similar algorithms.
Yanqing Shen, Yuedong Yang, Xiaodong Deng, Shi-tao Chen, Jingmin Xin, Nanning Zheng 0001
IEEE Trans. Cybern.6
2022 Multi-View Information Bottleneck Without Variational Approximation
abstract
By "intelligently" fuse the complementary information across different views, multi-view learning is able to improve the performance of classification task. In this work, we extend the information bottleneck principle to supervised multi-view learning scenario and use the recently proposed matrix-based Rényi’s α-order entropy functional to optimize the resulting objective directly, without the necessity of variational approximation or adversarial training. Empirical results in both synthetic and real-world datasets suggest that our method enjoys improved robustness to noise and redundant information in each view, especially given limited training samples. Code is available at https://github.com/archy666/MEIB.
Qi Zhang 0089, Shujian Yu, Jingmin Xin, Badong Chen
ICASSP3
2021 DSP-Net: Dense-to-Sparse Proposal Generation Approach for 3D Object Detection on Point Cloud
abstract
Object proposals generated based on sparse points from the raw point cloud have been widely used in 3D object detection. However, following the above scheme, most existing proposal generators have two problems, one is that the features for proposal generation constrain the detection performance by containing insufficient information; the other is that the sparse points obtained from the raw point cloud are misaligned with their corresponding objects in location and feature aspects. In this paper, we propose a dense-to-sparse proposal generation approach for 3D object detection, which can deal with the two problems simultaneously. Our approach utilizes the 3D CNN backbone to output dense features as a supplement to the original sparse point features for proposal generation. Besides, an object-aware feature pooling module is designed to address the misalignment between sparse points and corresponding objects. Experiments on the KITTI dataset show that our method outperforms the existing sparse-style methods and other published state-of-the-art methods.
Xinrui Yan, Shi-tao Chen, Zhixiong Nan, Jingmin Xin, Nanning Zheng 0001
IJCNN5
2021 A joint object detection and semantic segmentation model with cross-attention and inner-attention mechanisms
Zhixiong Nan, Jizhi Peng, Jingjing Jiang, Hui Chen 0036, Ben Yang, Jingmin Xin, Nanning Zheng 0001
Neurocomputing6
2021 Efficient Repair Analysis Algorithm Exploration for Memory With Redundancy and In-Memory ECC
abstract
In-memory error correction code (ECC) is a promising technique to improve the yield and reliability of high density memory design. However, the use of in-memory ECC poses a new problem to memory repair analysis algorithm, which has not been explored before. This article first makes a quantitative evaluation and demonstrates that the straightforward algorithms for memory with redundancy and in-memory ECC have serious deficiency on either repair rate or repair analysis speed. Accordingly, an optimal repair analysis algorithm that leverages preprocessing/filter algorithms, hybrid search tree, and depth-first search strategy is proposed to achieve low computational complexity and optimal repair rate in the meantime. In addition, a heuristic repair analysis algorithm that uses a greedy strategy is proposed to efficiently find repair solutions. Experimental results demonstrate that the proposed optimal repair analysis algorithm can achieve optimal repair rate and increase the repair analysis speed by up to 105×105× compared with the straightforward exhaustive search algorithm. The proposed heuristic repair analysis algorithm is approximately 28 percent faster than the proposed optimal algorithm, at the expense of 5.8 percent repair rate loss.
Minjie Lv, Hongbin Sun 0001, Jingmin Xin, Nanning Zheng 0001
IEEE Trans. Computers3
2021 Exploring Highly Dependable and Efficient Datacenter Power System Using Hybrid and Hierarchical Energy Buffers
abstract
The massive and irregular load surges challenge datacenter power infrastructures. As a result, power mismatching between supply and demand has emerged as a crucial availability issue in modern datacenters which are either under-provisioned or powered by intermittent power sources. Recent proposals have employed energy storage devices such as the uninterruptible power supply (UPS) to address this issue. However, current approaches lack the capacity of efficiently handling the irregular and unpredictable power mismatches. In this paper, we propose Hybrid and Hierarchical Energy Buffering (HHEB), a novel heterogeneous and adaptive scheme that could enable various energy storage devices (ESDs) to be efficiently integrated into existing datacenters for dynamically dealing with power mismatches. Our techniques exploit the diverse characteristics of different ESDs and intelligent load assignment algorithms to improve the dependability and efficiency of datacenter power systems. We evaluate the HHEB design with a prototype. Compared with a homogenous battery energy buffering system, HHEB could improve energy efficiency by 39.7 percent, extend UPS lifetime by 4.7X, promote energy availability by 3.2X, reduce system downtime by 41 percent, and effectively improve the energy availability of various energy buffers in different hierarchies. It allows datacenters to adapt to various power supply anomalies, thereby improving operational efficiency, dependability and availability.
Longjun Liu, Hongbin Sun 0001, Chao Li 0009, Tao Li 0006, Jingmin Xin, Nanning Zheng 0001
IEEE Trans. Sustain. Comput.5
2020 Lymph Node Metastasis Classification Based on Semi-Supervised Multi-View Network
abstract
Lymphatic metastasis is one of the most common proliferation pathways of thyroid carcinoma. Accurate diagnosis of lymph nodes is of great significance to surgical planning and prognosis. Due to the continuous development of deep learning recently, computer-aided diagnosis (CAD) systems for thyroid cancer have made considerable progress, but the research on the effective diagnosis of lymphatic metastasis remains insufficient. Focusing on this issue, we propose a semi-supervised multi-view network to diagnose lymph node metastasis, which combines coarse-view and fine-view to obtain a more comprehensive description. This method consists of three parts as follows: 1) joint probabilistic labels of the nodule partition information are generated by fuzzy clustering and perform semi-supervised learning on coarse-view with real labels; 2) an attention mechanism based network for fine-view is designed to capture various differentiated local features in a pyramid manner; 3) the two parts are then combined to extract global and local features more effectively to derive more accurate diagnostic reasoning. Especially, the introduction of fuzzy logic greatly reduces the impact of the uncertainty of the generated labels, thereby ensuring the effectiveness of the pseudo-labels. Extensive experiments on our collected dataset demonstrate that the proposed method is more efficient than other state-of-the-art methods.
Yiwen Luo, Jingmin Xin, Junqin Feng, Litao Ruan, Nanning Zheng 0001
BIBM2
2020 A Deep Model for Joint Object Detection and Semantic Segmentation in Traffic Scenes
abstract
Object detection and semantic segmentation are two fundamental techniques of various applications in the fields of Intelligent Vehicles (IV) and Advanced Driving Assistance System (ADAS). Early studies separately handle these two problems. In this paper, inspired by some recent works, we propose a deep neural network model for joint object detection and semantic segmentation. Given an image, an encoder-decoder convolution network extracts a set of feature maps, these feature maps are shared by the detection branch and the segmentation branch to jointly carry out the object detection and semantic segmentation. In the detection branch, we design a PriorBox initialization mechanism to propose more object candidates. In the segmentation branch, we use the multi-scale atrous convolution to explore the global and local semantic information in traffic scenes. Benefiting from the PriorBox Initialization Mechanism (PBIM) and Multi-Scale Atrous Convolution (MSAC), our model presents the competitive performance. In the experiments, we widely compare with several recently-proposed methods on the public Cityscapes dataset, achieving the highest accuracy. In addition, to verify the robustness and generalization of our model, the extension experiments are also conducted on the well-known VOC2012 dataset.
Jizhi Peng, Zhixiong Nan, Linhai Xu, Jingmin Xin, Nanning Zheng 0001
IJCNN4
2020 Predicting COVID-19 in China Using Hybrid AI Model
abstract
The coronavirus disease 2019 (COVID-19) breaking out in late December 2019 is gradually being controlled in China, but it is still spreading rapidly in many other countries and regions worldwide. It is urgent to conduct prediction research on the development and spread of the epidemic. In this article, a hybrid artificial-intelligence (AI) model is proposed for COVID-19 prediction. First, as traditional epidemic models treat all individuals with coronavirus as having the same infection rate, an improved susceptible-infected (ISI) model is proposed to estimate the variety of the infection rates for analyzing the transmission laws and development trend. Second, considering the effects of prevention and control measures and the increase of the public's prevention awareness, the natural language processing (NLP) module and the long short-term memory (LSTM) network are embedded into the ISI model to build the hybrid AI model for COVID-19 prediction. The experimental results on the epidemic data of several typical provinces and cities in China show that individuals with coronavirus have a higher infection rate within the third to eighth days after they were infected, which is more in line with the actual transmission laws of the epidemic. Moreover, compared with the traditional epidemic models, the proposed hybrid AI model can significantly reduce the errors of the prediction results and obtain the mean absolute percentage errors (MAPEs) with 0.52%, 0.38%, 0.05%, and 0.86% for the next six days in Wuhan, Beijing, Shanghai, and countrywide, respectively.
Nanning Zheng 0001, Shaoyi Du, Jianji Wang 0001, Wenting Cui, Zijian Kang, Tao Yang 0032, Bin Lou, Yuting Chi, Hong Long, Mei Ma, Dong Zhang 0009, Jingmin Xin
IEEE Trans. Cybern.16
2018 A Novel Approach for Detecting Road Based on Two-Stream Fusion Fully Convolutional Network
abstract
Road detection is one of the most basic tasks of autonomous driving systems. At present, researches on this issue mainly take two kinds of data as input,i.e., LIDAR point clouds and RGB images from cameras. To make best use of the advantages and bypass the disadvantages of these two kinds of data, we propose a novel network, namely two- stream fusion fully convolutional network (TSF-FCN), which can take advantage of both the accurate location information from LIDAR point clouds and rich appearance information from RGB images. One stream of this network is LIDAR stream which aggregates multi-scale contextual information from LIDAR point clouds. The other stream is RGB stream which is used for extracting features from RGB images. To fuse the two streams, the feature maps of RGB stream are converted to a bird-view representation to concatenate with that of LIDAR stream. In this way, the two kinds of data can complement each other for detecting road. To verify the efficacy of our TSF-FCN, experiments are carried on KITTI- ROAD benchmark and competitive performance is achieved compared with state-of-the-art methods.
Ziyi Liu 0001, Jingmin Xin, Nanning Zheng 0001
Intelligent Vehicles Symposium3
2018 Exploring the Potential of Using Semantic Context and Common Sense in On-Road Vehicle Detection
abstract
Vehicle detection is an important research topic for autonomous driving community. Since the great success of deep learning on object detection, almost all vehicle detection methods go along with this line. However, deep learning methods heavily rely on the training data, and the whole mechanism is like a “black box” Therefore, in this paper, we explore a vehicle detection method using traffic semantic context and human common sense instead of relying on the training data. To verify our idea, we compare our method with two classic machine learning methods as well as three state- of-the-art deep learning methods on a dataset collected in real traffics. The results show that our method outperforms others on this dataset. The deep learning methods may exceed ours after enlarging the training data or testing on more complicated datasets. However, the main contribution of this paper is providing inspiration for learning methods, and we believe their performance can be greatly improved after considering the idea of this paper.
Zhixiong Nan, Menghan Pan, Xiao Wang 0002, Ping Wei 0001, Linhai Xu, Hongbin Sun 0001, Jingmin Xin, Nanning Zheng 0001
Intelligent Vehicles Symposium7
2018 Leveraging Spatio-Temporal Evidence and Independent Vision Channel to Improve Multi-Sensor Fusion for Vehicle Environmental Perception
abstract
For intelligent vehicles, multi-sensor fusion is of great importance to perceive traffic environment with high accuracy and robustness. In this paper, we propose two effective methods, i.e. spatio-temporal evidence generating and independent vision channel, to improve multi-sensor track-level fusion for vehicle environmental perception. The spatio-temporal evidence includes instantaneous evidence, tracking evidence and tracks matching evidence to improve existence fusion. Independent vision channel leverages the specific advantage of vision processing on object recognition to improve classification fusion. The proposed methods are evaluated by using the multi-sensor dataset collected from real traffic environment. Experimental results demonstrate that the proposed methods can significantly improve the multi-sensor track-level fusion in terms of both detection accuracy and classification accuracy.
Juwang Shi, Wenxiu Wang, Xiao Wang 0002, Hongbin Sun 0001, Xuguang Lan, Jingmin Xin, Nanning Zheng 0001
Intelligent Vehicles Symposium6
2018 Point-wise saliency detection on 3D point clouds via covariance descriptors
Yu Guo 0006, Fei Wang 0008, Jingmin Xin
Vis. Comput.3
2017 Robust echo state networks based on correntropy induced loss function
Yu Guo 0006, Fei Wang 0008, Badong Chen, Jingmin Xin
Neurocomputing4
2017 A hierarchical recursive method for text detection in natural scene images
Yonghong Song, Yuanlin Zhang 0001, Jingmin Xin
Multim. Tools Appl.4
2017 Managing Battery Aging for High Energy Availability in Green Datacenters
abstract
Energy storage devices (ESD), such as UPS batteries, have been repurposed in datacenter as a promising tuning knob for peak power shaving and power cost reducing. However, batteries progressively aging due to irregular usage patterns, which result in less effective capacity and even pose serious threat to server availability. Nevertheless, prior proposals largely ignore the aging issues of battery which may lead to low energy availability for datacenter servers. To fill this critical void, we thoroughly investigate battery aging on a heavily instrumented prototype system over an observation period of ten months. We propose Battery Anti-Aging Treatment Plus (BAAT-P), a novel power delivery architecture included aging management algorithms from the perspective of computing system to hide, reduce, mitigate and plan the battery aging effects for high energy availability in datacenter. Our techniques exploit diverse battery aging mechanisms and dynamic aging management algorithms to provide system-level availability guarantee for datacenter. We evaluate the BAAT-P design with a real prototype. Compared with a battery powered datacenter without aging management policies, the results show that BAAT-P can extend battery lifetime by 72 percent, reduce battery cost by 33 percent and effectively improve energy availability for datacenter servers while maintaining workload performance for the performance critical workloads.
Longjun Liu, Hongbin Sun 0001, Chao Li 0009, Tao Li 0006, Jingmin Xin, Nanning Zheng 0001
IEEE Trans. Parallel Distributed Syst.5
2016 SVM classification of microaneurysms with imbalanced dataset based on borderline-SMOTE and data cleaning techniques
abstract
Microaneurysms are the earliest clinic signs of diabetic retinopathy, and many algorithms were developed for the automatic classification of these specific pathology. However, the imbalanced class distribution of dataset usually causes the classification accuracy of true microaneurysms be low. Therefore, by combining the borderline synthetic minority over-sampling technique (BSMOTE) with the data cleaning techniques such as Tomek links and Wilson’s edited nearest neighbor rule (ENN) to resample the imbalanced dataset, we propose two new support vector machine (SVM) classification algorithms for the microaneurysms. The proposed BSMOTE-Tomek and BSMOTE-ENN algorithms consist of: 1) the adaptive synthesis of the minority samples in the neighborhood of the borderline, and 2) the remove of redundant training samples for improving the efficiency of data utilization. Moreover, the modified SVM classifier with probabilistic outputs is used to divide the microaneurysm candidates into two groups: true microaneurysms and false microaneurysms. The experiments with a public microaneurysms database shows that the proposed algorithms have better classification performance including the receiver operating characteristic (ROC) curve and the free-response receiver operating characteristic (FROC) curve.
Qingjie Wang, Jingmin Xin, Jiayi Wu 0002, Nanning Zheng 0001
ICMV2
2016 Hierarchical learning of large-margin metrics for large-scale image classification
Jingmin Xin, Peixiang Dong, Jianping Fan 0001
Neurocomputing3
2016 On-Road Vehicle Detection and Tracking Using MMW Radar and Monovision Fusion
abstract
With the potential to increase road safety and provide economic benefits, intelligent vehicles have elicited a significant amount of interest from both academics and industry. A robust and reliable vehicle detection and tracking system is one of the key modules for intelligent vehicles to perceive the surrounding environment. The millimeter-wave radar and the monocular camera are two vehicular sensors commonly used for vehicle detection and tracking. Despite their advantages, the drawbacks of these two sensors make them insufficient when used separately. Thus, the fusion of these two sensors is considered as an efficient way to address the challenge. This paper presents a collaborative fusion approach to achieve the optimal balance between vehicle detection accuracy and computational efficiency. The proposed vehicle detection and tracking design is extensively evaluated with a real-world data set collected by the developed intelligent vehicle. Experimental results show that the proposed system can detect on-road vehicles with 92.36% detection rate and 0% false alarm rate, and it only takes ten frames (0.16 s) for the detection and tracking of each vehicle. This system is installed on Kuafu-II intelligent vehicle for the fourth and fifth autonomous vehicle competitions, which is called “Intelligent Vehicle Future Challenge” in China.
Xiao Wang 0002, Linhai Xu, Hongbin Sun 0001, Jingmin Xin, Nanning Zheng 0001
IEEE Trans. Intell. Transp. Syst.4
2016 Integrated Longitudinal and Lateral Control for Kuafu-II Autonomous Vehicle
abstract
Over the past decades, there has been significant research effort dedicated to the development of autonomous vehicles and advanced driver assistance systems. The driving control system, which is responsible for trajectory tracking and driving safety, is one of the most important technologies for autonomous vehicles. This paper describes the design of driving control system, including both longitudinal and lateral controllers, for the Kuafu-II autonomous vehicle. Compared with most of the previous researches that inevitably require a large amount of parameters, the presented control system design in this paper integrates several typical and efficient controllers to significantly reduce the system sensitivity to these parameters, and it is able to achieve the system robustness under diversified circumstances. The effectiveness of the presented control system design has been extensively evaluated under simulation and on road tests.
Linhai Xu, Yingzhou Wang, Hongbin Sun 0001, Jingmin Xin, Nanning Zheng 0001
IEEE Trans. Intell. Transp. Syst.4
2016 RE-UPS: an adaptive distributed energy storage system for dynamically managing solar energy in green datacenters
Longjun Liu, Hongbin Sun 0001, Chao Li 0009, Yang Hu 0001, Jingmin Xin, Nanning Zheng 0001, Tao Li 0006
J. Supercomput.5
2015 Logic-DRAM co-design to efficiently repair stacked DRAM with unused spares
abstract
Three dimensional (3D) integration is promising to provide dramatic performance and energy efficiency improvement to 3D logic-DRAM integrated computing system, but also poses significant challenge to the yield and reliability. By leveraging logic-DRAM co-design, this paper exploits the cost efficient approach to repair 3D integration induced defective cells in stacked DRAM with unused spares. In particular, we propose to make the DRAM array open its redundancy to off-chip access by small architecture modification, and further design the defective address comparison and redundant address remapping with very efficient architecture on logic die to achieve the equivalent memory repair. Simulation results have demonstrated that the proposed repair technique for DRAM after die stacking is able to significantly alleviate the yield loss, with very low area and power consumption overhead and negligible timing penalty.
Minjie Lv, Hongbin Sun 0001, Jingmin Xin, Nanning Zheng 0001
ASP-DAC3
2015 HEB: deploying and managing hybrid energy buffers for improving datacenter efficiency and economy
abstract
Today, an increasing number of applications and services are being hosted by large-scale data centers. The massive and irregular load surges challenge data center power infrastructures. As a result, power mismatching between supply and demand has emerged as a crucial issue in modern data centers which are either under-provisioned or powered by intermittent power sources. Recent proposals have employed energy storage devices such as the uninterruptible power supply (UPS) systems to address this issue. However, current approaches lack the capacity of efficiently handling the irregular and unpredictable power mismatches.
Longjun Liu, Chao Li 0009, Hongbin Sun 0001, Yang Hu 0001, Juncheng Gu, Tao Li 0006, Jingmin Xin, Nanning Zheng 0001
ISCA7
2015 Natural scene text detection with multi-layer segmentation and higher order conditional random field based analysis
Yonghong Song, Yuanlin Zhang 0001, Jingmin Xin
Pattern Recognit. Lett.4
2014 A computationally efficient source localization method for a mixture of near-field and far-field narrowband signals
abstract
In this paper, we consider the source localization for a mixed near-field (NF) and far-field (FF) narrowband signals impinging on a uniform linear array (ULA) with the symmetrical geometric configuration. A computationally efficient direction-of-arrivals (DOAs) and range estimation method for the mixed NF and FF signals is proposed, where the DOAs of the NF and FF signals are estimated separately, and the computationally burdensome eigendecomposition is avoided. Comparing to some existent methods, the proposed method can separate the NF signals from the FF signals more efficiently, and consequently the estimation performance is improved. The effectiveness of the proposed method is verified though numerical examples.
Weiliang Zuo, Jingmin Xin, Jiasong Wang, Nanning Zheng 0001, Akira Sano
ICASSP2
2011 Linear Pose Estimation Algorithm Based on Quaternion
Yongjian He, Caigui Jiang, Chengwei Hu, Jingmin Xin, Fei Wang 0008
ICIC (1)4
2005 Subspace-based adaptive direction estimation and tracking in multipath environment
abstract
A new computationally efficient subspace-based algorithm is proposed for estimating and tracking the directions of coherent narrowband signals impinging on a uniform linear array (ULA). Specifically the null space is estimated using the least-mean-square (LMS) or normalized LMS (NLMS) algorithm, and the directions are updated using the approximate Newton method. By studying the convergence analyses of the LMS and NLMS algorithms, where the "weight" is in the form of a matrix and there is a correlation between the "additive noise" and "input data" in the updating equation, the step-size stability conditions are derived explicitly. Further the tracking of crossing directions of moving signals is considered. The theoretical analyses and effectiveness of the proposed algorithm are verified.
Jingmin Xin, Naoyuki Hirosaki, Hiroyuki Tsuji, Yoji Ohashi, Akira Sano
ICASSP (4)1
2002 Direction estimation of coherent signals using spatial signature
abstract
A computationally efficient spatial signature-based (SS) method is proposed for estimating the directions of arrival of coherent narrowband signals impinging on a uniform linear array. The normalized SS of the coherent signals is blindly estimated from the principal eigenvector of array covariance matrix and then is used to estimate the directions with a modified Kumaresan-Prony method, where a linear prediction model is combined with "subarray" averaging. The proposed method not only has the maximum permissible array aperture and computational simplicity, it also better resolves closely spaced coherent signals with a small length of data and at a lower signal-to-noise ratio.
Jingmin Xin, Akira Sano
IEEE Signal Process. Lett.1
2000 Regularization approach to direction estimation of coherent narrow-band signals
abstract
This paper addresses the problem of estimating the directions of coherent narrow-band signals impinging on a uniform linear array when the number of signals is unknown. By incorporating the linear prediction (LP) model with a subarray scheme, the directions of coherent signals can be estimated from the zeros of the corresponding prediction polynomial. To obtain a reliable estimation of the LP coefficients, we introduce multiple regularization parameters into the corrected least squares (CLS) estimation. The analytical expressions of the mean-squares-error (MSE) of the regularized CLS estimate and the optimal regularization parameters are derived. It is clarified that the number of signals can be determined by comparing the optimal parameters with the eigenvalues. An iterative regularization algorithm is developed for direction estimation without any a priori knowledge, where the number of coherent signals and the noise variance are estimated from the noisy data simultaneously.
Jingmin Xin, Akira Sano
ICASSP1
1995 Stability analysis of robust adaptive filter used in feedforward and feedback compensation
abstract
This paper proposes robust adaptive algorithms for adjusting coefficients of an adaptive filter which is used for feedforward control together with a feedback compensator. The filtered-x algorithm which is widely employed in adaptive signal processing cannot always assure the stability. Stability-guaranteed adaptive algorithms are given in time domain and frequency domain on a basis of the strictly positive real property of adaptive systems in the presence of unknown disturbances. Time domain algorithms can be applied to an IIR or FIR adaptive filter, while frequency domain approaches use a structure of adaptive frequency sampling filter. Numerical simulations and experimentally obtained results exhibited significant improvement on convergency and stability of the proposed adaptive algorithms in application to adaptive active noise control.
Makoto Kajiki, Jingmin Xin, Hiromitsu Ohmori, Akira Sano
ICASSP2