EDBT 2026 Demo / reviewers in the wild / expert
Jinsheng Xiao
dblp:28/6940
· DBLP profile ↗
60ranked-venue papers
22as first author
39since 2021 · last 2027
0000-0002-5403-1895ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 33 · 11 first-author · 24 since 2021Graphics, computer vision, multimedia, augmented reality and games · 19 · 8 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 9 · 3 first-author · 8 since 2021Databases, data management, data science and information retrieval · 3 · 2 since 2021Computer networks · 1 · 1 first-author · 1 since 2021Security and privacy · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2027 | PCGDNet: Reliability-aware glass detection with physics-inspired transparent-surface consistency from monocular RGB images
Jinsheng Xiao, Wenfei Wu |
Expert Syst. Appl. | 3 |
| 2026 | LSTT: Long short-term feature enhancement transformer for video small object detection
Jinsheng Xiao, Ruidi Chen, Wei Yang 0043 |
Expert Syst. Appl. | 1 |
| 2026 | SQKformer: Spiking sparse QKformer with adaptive batch normalization for membrane potential
Yunhua Chen, Zequan Xie, Jinyu Zhong, Pinghua Chen, Jinsheng Xiao |
Neurocomputing | 5 |
| 2026 | Enhanced depth completion for transparent objects via failure-aware RGB-D pretraining
Wenfei Wu, Jinsheng Xiao, Danqi Meng |
J. Vis. Commun. Image Represent. | 4 |
| 2026 | C3Net: A cross-modal collaborative calibration of features for object detection using frames and events
Yunhua Chen, Jinyu Zhong, Zequan Xie, Jinsheng Xiao, Pinghua Chen |
Neural Networks | 5 |
| 2026 | STKPS-Net: Spatio-Temporal Key Patch Selection Network for Few Shot Anomalous Action RecognitionabstractFor providing timely warnings and preventing potential damages, it is crucial to detect anomalous actions that threaten public safety through surveillance cameras. Compared to normal actions, anomalous actions often occupy only a small portion of surveillance videos and exhibit more complex manifestations in terms of time and space. Considering that normal action recognition methods fail to highlight crucial information from small-sized patches, we propose the Spatio-temporal Key Patch Selection Network (STKPS-Net). It includes a spatially adaptive key patch selection module to select small but informative patches, and a long-short feature map spatio-temporal relation module to capture dynamic changes in anomalous actions. Additionally, a spatio-temporal refined loss is introduced to enhance fine-grained feature learning. Experimental results on the HMDB51, Kinetics, and UCF-Crime v2 datasets show that our STKPS-Net achieves state-of-the-art performance in few-shot anomalous action recognition, outperforming the most competitive methods by 1.2% on the anomalous action dataset UCF-Crime v2. Jinsheng Xiao, Ruidi Chen, Xingyu Gao 0001, Hailong Shi, Zhongyuan Wang 0001 |
IEEE Trans. Inf. Forensics Secur. | 1 |
| 2026 | Efficient Topology-Aware Motion Planning for AVP in Large-Scale Occupancy MapabstractWith the development of autonomous driving technology, autonomous valet parking (AVP) has become a key technology to solve the problem of urban parking. Current commercial AVP systems generally adopt solutions based on semantic maps, which achieve high-precision parking in small-scale scenarios. However, when the parking environment is expanded to large underground parking lots, semantic and occupancy grid maps face bottleneck problems such as a sharp drop in path generation efficiency and delayed parking space retrieval response. In addition, traditional High-definition maps (HD maps) rely on manual annotation and complex post-processing. In response to the above challenges, this article proposes an efficient adaptive topology plan for AVP in large-scale occupancy map: first, a scale-adaptive index model based on the R-tree structure is constructed to achieve hierarchical storage and dynamic resolution selection of grid map data; secondly, a multi - scale feature fusion topology aware method is designed to generate the environment topology; finally, a multi-path parallel hybrid A${}^{\ast }$algorithm is proposed for efficient planning. A comparison of our framework with both traditional and state-of-the-art methods shows that the framework is capable of enhancing planning efficiency and reducing average path generation time in large parking lots. Through simulations and real-world tests, the method has been shown to reduce path search time whilst generating paths that are easier to track with less tracking error. Jian Zhou 0011, Fuyu Nie, Haoran Li 0022, Jinsheng Xiao |
IEEE Trans. Intell. Transp. Syst. | 6 |
| 2025 | SCNet: Spatio-temporal Feature Aggregation and Cross-modal Interactive Encoding Network for DAVIS Object DetectionabstractDAVIS cameras, which output both event streams and frames simultaneously, are increasingly being used to address the primary object detection challenges posed by complex lighting and motion blur. Nevertheless, fully leveraging the abundant temporal information and effectively fusing data from these two modalities remains a formidable challenge. In this paper, we first design a multi-scale spatio-temporal aggregation (MSTA) module to distill richer semantic information from event frames. Secondly, we assimilate and harness the strengths of YOLOv8 and RT-DETR to develop an innovative encoder with Multi-scale Cross-modal dynamic Interactive fusion and multi-level feature interactive Fusion (MCIF). In MCIF, we propose a dynamic channel switching and spatial attention with learnable fusing factors (DCF-CSSA) to improve the complementary interaction of cross-modal features. Extensive experiments demonstrate that our approach (which we call SCNet) significantly outperforms existing state-of-the-art (SOTA) object detection methods that fuse events and frames, achieving an mAP50 improvement of 6.2% on PKU-DAVIS-SOD and 12% on DESC-MOD, both contain a large number of samples with challenging lighting conditions and motion blur. Yunhua Chen, Jinyu Zhong, Pinghua Chen, Jinsheng Xiao |
ICMR | 5 |
| 2025 | HGCM: A hybrid radar feature and GNSS based continuous mapping methodabstractOnline mapping and localization are critical technologies in autonomous driving and robotics systems, currently being a prominent research topic. However, the majority of proposed methods have primarily focused on camera and LiDAR (Light Detection and Ranging) sensors, which are inadequate for addressing more challenging weather conditions. Radar-based simultaneous localization and mapping (SLAM) is essential for autonomous navigation in complex and dynamic environments where other sensors may fail. However, there is a gap in accuracy and stability compared with LiDAR-based SLAM, thus limiting its application for autonomous driving, drones and robotics. High noise levels, low feature stability and the lack of global constraints would result in the deterioration of pose estimation. This paper presents a Simultaneous Localization and Mapping (SLAM) system that combines hybrid features from a spinning radar and Global Navigation Satellite System (GNSS) measurements to achieve accurate mapping results in complex urban environments. To ensure intra-frame and inter-frame stability of pose estimation, we applied structural noise removal, feature ranking and motion buffer techniques. Focusing on improving trajectory consistency, we introduce a GNSS factor and a motion factor to provide stable global constraints and mitigate the impact of GNSS noise on the system, respectively. Experimental results on the Mulran dataset, which includes complex urban scenarios, demonstrate that our system achieves state-of-the-art performance, with an average ATE(Absolute Trajectory Error) of 2.71m and a standard deviation of 1.10m. Youchen Tang, Jian Zhou 0011, Maosheng Yan, Yuanxian Huang, Jinsheng Xiao |
Expert Syst. Appl. | 6 |
| 2025 | Human key point detection method based on enhanced receptive field and transformer
Hongyu Liang, Wenjuan Xie, Jinsheng Xiao |
Neurocomputing | 4 |
| 2025 | Optimizing Multi-AAV Formation Cooperative Control Strategies With MCDDPG ApproachabstractIn recent years, the application of autonomous aerial vehicle (AAV) devices in military, industrial, and civilian sectors has become increasingly widespread. Consequently, research on multi-AAV formation cooperative control strategies has garnered significant attention. However, current multiagent reinforcement learning algorithms often struggle with unguided exploration, making it challenging for agents to develop efficient action strategies for complex collaborative tasks. To address this issue, this article introduces a multicritic deep deterministic policy gradient (MCDDPG) algorithm. This algorithm designs a multicritic (MC) structure based on the DDPG algorithm. This structure guides AAVs using physical models for tracking and obstacle avoidance, while deep learning models are employed to facilitate cooperative coordination among AAVs. Furthermore, to address the weight allocation issue among different Critic modules in the MC structure, a dynamic difficulty priority weight optimization algorithm is implemented. This enhances the algorithm’s collaborative capabilities. To validate the collaborative planning capability of the proposed algorithm, a simulation scenario involving multicoupled tasks is designed in the multiagent particle environment (MPE). In this scenario, the MCDDPG algorithm demonstrates the fastest convergence speed and the optimal collaborative strategy, outperforming other state-of-the-art multiagent deep reinforcement learning (MADRL) algorithms currently in use. Jinsheng Xiao, Bolun Yan, Honggang Xie, Qiuze Yu, Linkun Li, Yuan-Fang Wang |
IEEE Internet Things J. | 1 |
| 2025 | Probabilistic memory auto-encoding network for abnormal behavior detection in surveillance videoabstractAbnormal behavior detection in surveillance video, as one of the essential functions in the intelligent surveillance system, plays a vital role in anti-terrorism, maintaining stability, and ensuring social security. Aiming at the problem of extremely imbalance between normal behavior data and abnormal behavior data, the probabilistic memory model-based network is designed to learn from the distribution of normal behaviors and guide the detection of abnormal behavior. An auto-encoding model is employed as the backbone network, and the gap between the predicted future frame and the real frame is used to measure the degree of abnormality. An autoregressive conditional probability estimation model and a normal distribution memory model are employed as auxiliary modules, to achieve the prediction of normal frames. When extracting temporal and spatial features in the backbone network, the causal three-dimensional convolution and time-dimension shared fully connected layers are used to avoid future information leakage and ensure the timing of information. In addition, from the perspective of probability entropy and behavioral modality diversity, autoregressive probability model is proposed to fit the distribution of input normal frame, so the network converges to the low entropy state of the normal behavior distribution. The memory module stores the feature of normal behavior in historical data, and injects the current input data. The memory vector and the encoding vector are concatenated along the time dimension and input to the decoder, realizing normal frame prediction. Using public datasets, ablation and comparison experiments show that the proposed algorithm has significant advantages in anomaly detection. Jinsheng Xiao, Qiuze Yu, Honggang Xie, Yuan-Fang Wang |
Neural Networks | 1 |
| 2025 | Revisiting the Learning Stage in Range View Representation for Autonomous DrivingabstractLiDAR segmentation is crucial for autonomous driving perception. Range view methods have been widely adopted for these applications due to their intuitiveness and ease of implementation. However, the inherent shortcomings of the range view approach (e.g., assuming that point clouds within the same pixel of a range image have the same semantic class) make it difficult to perform accurate fine-grained segmentation tasks, thus limiting its potential in practical applications. To address these issues, we propose RangeFusion, an end-to-end framework that greatly improves the ability to learn and process LiDAR point clouds from range views by employing a multispatial learning model. A novel range-scan space (RSS) is proposed to address the inability of existing range view methods to accurately aggregate features of neighboring points. This space achieves accurate and efficient neighboring point feature aggregation with linear time complexity. In addition, a supervised label smoothing method called multilevel feature selection heads (MFSHs) is designed, which achieves more fine-grained semantic prediction by subdividing the full point cloud into multisemantic hierarchical subclouds and adaptively fusing the features with confidence filtering. The performance of the proposed method was evaluated on several benchmarks, including SemanticKITTI and nuScenes. On these two datasets, mean intersection over union (mIoU) scores of 67.9% and 80.2% were achieved, respectively. This demonstrates that the proposed approach outperforms existing range view- and multiview-based approaches while maintaining efficient performance at 26.5 FPS. In addition, real road data were collected for testing. The code is available athttps://github.com/Wansit99/RangeFusion. Jinsheng Xiao, Siyuan Wang 0011, Jian Zhou 0011, Ziyin Zeng, Ruijia Chen |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2025 | VPDNet: Virtual Point Density-Aware Network for Multimodal 3-D Object DetectionabstractLidar has become a prevalent sensor for 3D object detection in autonomous driving. However, the sparse and irregular nature of point cloud data obtained from Lidar necessitates the adoption of cross-modal object detection methods to enhance detection accuracy and stability. Nonetheless, the disparate representations of point clouds and images hinder comprehensive data fusion, leading to suboptimal performance. This paper proposes a novel multimodal density-aware 3D object detection method, VPDNet, and leverages depth completion-generated virtual point clouds to address the challenges of data fusion. To mitigate the interference caused by inaccurate depth completion, we introduce a virtual point cloud enhancement module. This module utilizes weighted features of virtual point clouds based on image segmentation information to suppress invalid information while retaining valid information. Furthermore, to further capture information from images and point clouds, we design an interactive attention fusion module that integrates features from virtual point clouds and real point clouds by adjusting attention weights. Additionally, the characteristics of Lidar dictate that collected point cloud data varies unevenly with distance, and point density, an essential feature, is often overlooked. Therefore, we propose a density-aware module that combines fused voxel features with kernel density estimation and point density information to extract spatial local features. Experiments conducted on the KITTI dataset demonstrate that our proposed 3D object detection algorithm improves pedestrian detection results by 6.93% compared to the baseline network. The code is available at https://github.com/Alexandriahui/VPDNet. Binghui Yang, Jinsheng Xiao |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2025 | RoCalib: Large-Scale Autonomous Geo-Calibration for Roadside Lidar With High-Definition MapabstractTo address the challenges of low overlap and viewing direction differences in large-scale roadside LiDAR calibration, this paper proposes RoCalib, a novel automatic roadside LiDAR calibration method based on the high-definition (HD) map. This method enables geographic registration of the LiDAR without any specific target and improves the efficiency and safety of the large-scale roadside facility calibration and maintenance. First, a novel virtual reprojection model is designed to construct a virtual mapping from the HD map to LiDAR, reducing representation differences. Based on this, a universal spatial context descriptor is introduced, applicable to various LiDAR systems, facilitating rapid retrieval of LiDAR positions within the HD map. Finally, based on the multi-feature optimization method considering the road structure, the fine registration and parameter calibration of the roadside LiDAR and the HD map are completed. The proposed framework is validated on simulated, public, and self-collected datasets, demonstrating that this method can automatically and accurately achieve multi-LiDAR geographic calibration, yielding superior performance. Cong Duan, Jian Zhou 0011, Zhen Dong 0005, Youchen Tang, Jinsheng Xiao |
IEEE Trans. Intell. Transp. Syst. | 6 |
| 2025 | ReID-FSAI: Person Re-Identification Network Fused With Semantic and Attribute InformationabstractPersonal belongings information (e.g., backpacks and reticules) and attribute descriptions (e.g., gender and age) provide critical discriminative cues for person re-identification (Re-ID) tasks. However, existing Re-ID algorithms leveraging additional semantic models often fail to accurately recognize personal belongings and suffer from noisy attribute predictions derived from global or local features, as they inadequately exploit attribute correlations. To address these challenges, we propose a novel person re-identification network, ReID-FSAI, which fuses personal belongings information and attribute descriptions from isolated semantic regions. ReID-FSAI integrates personal belongings areas identified through feature clustering with semantic parsing results from an auxiliary semantic model. By treating the generated semantic regions as body labels, our network refines global features into precise semantic features and accurately predicts attribute information from these regions. Furthermore, ReID-FSAI employs a reweighting model to enhance the confidence in specific attributes, improving attribute prediction accuracy. By combining predictions of attributes and personal belongings with global features, our approach significantly improves the representation ability of pedestrians. Experimental evaluations on the Market-1501 and DukeMTMC-reID datasets demonstrate that ReID-FSAI achieves superior performance in both person re-ID and attribute prediction, surpassing state-of-the-art methods. Jinsheng Xiao, Qiuze Yu, Zhongyuan Wang 0001, Yuan-Fang Wang |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2025 | MIM: High-Definition Maps Incorporated Multi-View 3D Object Detectionabstract3D object detection has aroused increasing interest as a crucial component of autonomous driving systems. While recent works have explored various multi-modal fusion methods to enhance accuracy and robustness, fusing multi-view images and high-definition (HD) maps remains uncharted. Inspired by our previous work, we endeavor to introduce HD maps to camera-based detection, prompting the design of a new framework. To address this, we first analyze the function of HD maps in object detection to understand their benefits and the rationale for their fusion. From this analysis, we identify key disparities in view, semantics, and scale, leading to the development of MIM, a framework for HD Maps Incorporated Multi-view 3D object detection. HD maps are enriched in semantics by sampling unlabeled areas and encoding them into map features. Simultaneously, multi-view images are transformed into features in bird’s-eye view (BEV) using the adopted baseline. These features are then fused using attention mechanisms to align scales. Experiments conducted on the nuScenes dataset demonstrate that MIM outperforms camera-based methods. Moreover, an in-depth analysis investigates how HD maps impact object detection regarding each semantic layer. The results underscore the operational intricacies of HD maps in perception, setting the stage for future research. Code is available athttps://github.com/WHU-xjs/MIM-3D-Det. Jinsheng Xiao, Jian Zhou 0011, Ziyue Tian, Hongping Zhang, Yuan-Fang Wang |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2024 | REDIR: Refocus-Free Event-Based De-occlusion Image Reconstruction
Hailong Shi, Jinsheng Xiao, Xingyu Gao 0001 |
ECCV (80) | 4 |
| 2024 | Group IF Units with Membrane Potential Sharing for High-Accuracy Low-Latency Spiking Neural NetworksabstractSpiking neural networks (SNNs) have attracted much attention due to their low energy consumption and fast inference on neuromorphic hardware. Currently, the most effective way to implement deep SNNs is through ANN-SNN conversion, which combines the maturity of ANNs with the fast inference of SNNs, achieving accuracy comparable to ANNs on large-scale datasets. However, to achieve this accuracy, a large number of time steps are usually required, which undermines the low energy consumption and fast inference advantages of SNNs. When the number of time steps is very limited, the converted SNNs face a severe drop in accuracy, especially in the case of a single time step, which severely restricts the practical application of SNNs. In this paper, we analyze the main reasons why ANN- SNN cannot achieve lossless conversion with very limited number of time steps, and then proposes a group IF units with shared membrane potential to achieve lossless conversion of ANN-SNN at extremely low latency. We validate the effectiveness of our method on CIFAR-10, CIFAR-100, and ImageNet. Experiments show that our proposed method outperforms the existing state- of-the-art methods at the same number of time steps. Moreover, we have achieved an nearly lossless conversion of ANN-SNN in a single time step to get a high-accuracy SNN for the first time. For example, for VGG-16, our method only loses 0.8% and 3.15% of the accuracy compared to ANNs on CIFAR-10/100 in a single time step. Zhenxiong Ye, Weiming Zeng, Yunhua Chen, Jinsheng Xiao, Irwin King |
IJCNN | 5 |
| 2024 | Efficient Spatio-temporal Event Representation Based on Kalman Filtering and Linear Weighted TimestampsabstractEvent cameras, serving as innovative vision sensors, provide novel insights and approaches for image classification tasks. These cameras generate a sparse and discrete event stream, with each individual event carrying minimal information. Consequently, it becomes crucial to convert this event stream into a suitable event representation that aids in feature extraction and object recognition. In this research, we present a computationally efficient spatio-temporal event representation that not only preserves the spatio-temporal information of events in its entirety but also simplifies the computation of temporal information through linearly weighted timestamps. Furthermore, we propose an adaptive segmentation method for event streams. This method generates time bins that exhibit high robustness to motion speed by integrating both global and local distribution information of event counts. To verify the efficacy of our proposed method, we conducted experiments on three publicly available datasets. The results demonstrate that our method surpasses other methods on both N-Caltech101 and CIFAR10-DVS, with enhancements of 1.1% and 4.1% respectively, and produces competitive results on N-CARS. Jinyu Zhong, Weiming Zeng, Yunhua Chen, Jinsheng Xiao, Irwin King |
IJCNN | 4 |
| 2024 | A lane-level localization method via the lateral displacement estimation model on expressway
Jian Zhou 0011, Quanhua Dong, Yaoan Bian, Zhijiang Li, Jinsheng Xiao |
Expert Syst. Appl. | 6 |
| 2024 | Offline writer identification approach using moment features and high-order correlation functions
Ayixiamu Litifu, Jinsheng Xiao |
J. Vis. Commun. Image Represent. | 2 |
| 2024 | High-performance deep spiking neural networks via at-most-two-spike exponential coding
Yunhua Chen, Ren Feng, Zhimin Xiong, Jinsheng Xiao, Jian K. Liu |
Neural Networks | 4 |
| 2024 | Re-decoupling the classification branch in object detectors for few-class scenes
Jie Hua 0005, Zhongyuan Wang 0001, Qin Zou 0001, Jinsheng Xiao, Xin Tian 0006 |
Pattern Recognit. | 4 |
| 2024 | PointNAT: Large-Scale Point Cloud Semantic Segmentation via Neighbor Aggregation With TransformerabstractGiven the prominence of 3D sensors in recent years, 3D point clouds are worthy to be further investigated for environment perception and scene understanding. Learning accurate local and global contexts in point clouds is pivotal for semantic segmentation, and neighbor aggregation and Transformers have achieved notable success in local and global perception in point cloud analysis, respectively. Nevertheless, studying each independently is far from the optimal solution for comprehensive feature learning. To address this, we take a novel step towards investigating and integrating the structures of neighbor aggregation and Transformers. In this paper, we introduce Point Neighbor Aggregation with Transformer (PointNAT), a conceptually straightforward and effective approach aiming to enhance the performance of 3D point cloud semantic segmentation. PointNAT consists of a Neighbor Aggregation Block (NAB) for local perception, a Point Transformer Block (PTB) for global modeling, and a Hybrid Block to connect NABs and PTBs. NABs effectively learn complex local features at varying scales through an improved neighbor aggregation operation and a multi-head mechanism. PTBs efficiently perform global attention using a small set of learnable key points. Hybrid Blocks serve as high-and-low frequency signal hybridizers, merging the strengths of these two blocks by adaptively assigning hybrid weights to local and global contexts. We have evaluated the performance of PointNAT with state-of-the-art networks on several benchmarks, including S3DIS, Toronto3D, and SensatUrban. PointNAT achieves mIoU scores of 77.8%, 84.7%, and 65.2% in these three dataset, respectively. Furthermore, it outperforms the baseline approach PointNeXt by 3.0%, 1.3%, and 4.2%, respectively, while utilizing only 59.9% of the parameters and 15.2% of the FLOPs. The results demonstrate PointNAT’s superior ability in accurately segmenting large-scale 3D point cloud scenes, emphasizing its potential to advance environment perception and scene understanding. Our code is available at https://github.com/zeng-ziyin/PointNAT. Ziyin Zeng, Huan Qiu, Jian Zhou 0011, Zhen Dong 0005, Jinsheng Xiao |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2023 | LSTFE-Net: Long Short-Term Feature Enhancement Network for Video Small Object DetectionabstractVideo small object detection is a difficult task due to the lack of object information. Recent methods focus on adding more temporal information to obtain more potent high-level features, which often fail to specify the most vital information for small objects, resulting in insufficient or inappropriate features. Since information from frames at different positions contributes differently to small objects, it is not ideal to assume that using one universal method will extract proper features. We find that context information from the long-term frame and temporal information from the short-term frame are two useful cues for video small object detection. To fully utilize these two cues, we propose a long short-term feature enhancement network (LSTFE-Net) for video small object detection. First, we develop a plug-and-play spatiotemporal feature alignment module to create temporal correspondences between the short-term and current frames. Then, we propose a frame selection module to select the long-term frame that can provide the most additional context information. Finally, we propose a long short-term feature aggregation module to fuse long short-term features. Compared to other state-of-the-art methods, our LSTFE-Net achieves 4.4% absolute boosts in AP on the FL-Drones dataset. More details can be found at https://github.com/xiaojs18/LSTFE-Net. Jinsheng Xiao, Yuanxu Wu, Yunhua Chen, Zhongyuan Wang 0001, Jiayi Ma 0001 |
CVPR | 1 |
| 2023 | LSA3D: Lightweight Separate Asynchronous 3D Convolutional Neural Network for Gait Recognition
Jianyu Chen 0008, Zhongyuan Wang 0001, Kangli Zeng, Jinsheng Xiao, Zhen Han 0002 |
ICANN (10) | 4 |
| 2023 | RMPE:Reducing Residual Membrane Potential Error for Enabling High-Accuracy and Ultra-low-latency Spiking Neural Networks
Yunhua Chen, Zhimin Xiong, Ren Feng, Pinghua Chen, Jinsheng Xiao |
ICONIP (3) | 5 |
| 2023 | Tiny object detection with context enhancement and feature purification
Jinsheng Xiao, Haowen Guo, Jian Zhou 0011, Qiuze Yu, Yunhua Chen, Zhongyuan Wang 0001 |
Expert Syst. Appl. | 1 |
| 2023 | FDLR-Net: A feature decoupling and localization refinement network for object detection in remote sensing images
Jinsheng Xiao, Yuntao Yao, Jian Zhou 0011, Haowen Guo, Qiuze Yu, Yuan-Fang Wang |
Expert Syst. Appl. | 1 |
| 2023 | HeadPose-Softmax: Head pose adaptive curriculum learning loss for deep face recognition
Jifan Yang, Zhongyuan Wang 0001, Baojin Huang, Jinsheng Xiao, Chao Liang 0001, Zhen Han 0002, Hua Zou 0002 |
Pattern Recognit. | 4 |
| 2022 | ECL: Exclusive Curriculum Learning for Video Super-ResolutionabstractVideo super-resolution (VSR) problem has gained a soaring development along with deep learning methods. However, the further progress requires the blessing of more complex architectures. Unlike them, this paper promotes VSR performance from a new perspective of sample difficulty. We propose an exclusive curriculum learning strategy for VSR, which can improve the representation power without noticeable computation increment. Specifically, this paper memorizes the performance track of every sample and calculate a customized weight for each sample according to it. In this way, the model can automatically concentrate on the easy samples first and gradually focus on the hard ones. Experimental analysis on training process and benchmark datasets demonstrate that our method can substantially boost the performance with a superior convergence speed and a limited number of parameters. Sicheng Hu, Zhongyuan Wang 0001, Peng Yi 0002, Zheng He 0001, Jinsheng Xiao, Jing Xiao 0004 |
ICME | 5 |
| 2022 | DANet: Image Deraining via Dynamic Association LearningabstractRain streaks and background components in a rainy input are highly correlated, making the deraining task a composition of the rain streak removal and background restoration. However, the correlation of these two components is barely considered, leading to unsatisfied deraining results. To this end, we propose a dynamic associated network (DANet) to achieve the association learning between rain streak removal and background recovery. There are two key aspects to fulfill the association learning: 1) DANet unveils the latent association knowledge between rain streak prediction and background texture recovery, and leverages it as an extra prior via an associated learning module (ALM) to promote the texture recovery. 2) DANet introduces the parametric association constraint for enhancing the compatibility of deraining model with background reconstruction, enabling it to be automatically learned from the training data. Moreover, we observe that the sampled rainy image enjoys the similar distribution to the original one. We thus propose to learn the rain distribution at the sampling space, and exploit super-resolution to reconstruct high-frequency background details for computation and memory reduction. Our proposed DANet achieves the approximate deraining performance to the state-of-the-art MPRNet but only requires 52.6\% and 23\% inference time and computational cost, respectively. Kui Jiang, Zhongyuan Wang 0001, Zheng Wang 0007, Peng Yi 0002, Junjun Jiang, Jinsheng Xiao, Chia-Wen Lin |
IJCAI | 6 |
| 2022 | An adaptive threshold mechanism for accurate and efficient deep spiking convolutional neural networks
Yunhua Chen, Yingchao Mai, Ren Feng, Jinsheng Xiao |
Neurocomputing | 4 |
| 2022 | Two-stage unsupervised facial image quality measurement
Guangcheng Wang, Zhongyuan Wang 0001, Baojin Huang, Kui Jiang, Zheng He 0001, Hancheng Zhu, Jinsheng Xiao, Xin Tian 0006 |
Inf. Sci. | 7 |
| 2022 | A Robust Descriptor Based on Modality-Independent Neighborhood Information for Optical-SAR Image MatchingabstractDue to the intensity differences and speckle noise, automatic optical-synthetic aperture radar (SAR) image matching is still a challenging task. This letter addresses this problem by proposing a novel descriptor (MaskMIND) with three different modes using modality independent neighborhood information. This descriptor aims to sample and active relative structural information to improve accuracy and precision. In addition, the gradient maps are calculated respectively in pretreatment to eliminate noise. Then the corresponding metric, which takes into account the increasing positional uncertainty with distance, is defined using the sum of squared differences (SSD) accelerated by fast Fourier transform (FFT). Our methods are effective because of its relativeness and abstractness. The experimental results in five optical-SAR image pairs show that our methods have great performance and potentialities. Compared with CFOG, which is the state-of-the-art method, the accuracy of our sMaskMIND-grids is improved by 12% on average. Qiuze Yu, Wensen Zhao, Yuxuan Jiang 0003, Ruikai Wang, Jinsheng Xiao |
IEEE Geosci. Remote. Sens. Lett. | 5 |
| 2022 | Generative adversarial network with hybrid attention and compromised normalization for multi-scene image conversion
Jinsheng Xiao, Shuhao Zhang 0008, Yuntao Yao, Zhongyuan Wang 0001, Yongqin Zhang, Yuan-Fang Wang |
Neural Comput. Appl. | 1 |
| 2022 | Cross-domain heterogeneous residual network for single image super-resolution
Qinghui Zhu, Yongqin Zhang, Juanjuan Yin, Ruyi Wei, Jinsheng Xiao, Deqiang Xiao, Guoying Zhao 0001 |
Neural Networks | 6 |
| 2021 | Writer identification using redundant writing patterns and dual-factor analysis of variance
Ayixiamu Litifu, Jinsheng Xiao |
Appl. Intell. | 3 |
| 2020 | Single image dehazing based on learning of haze layers
Jinsheng Xiao, Mengyao Shen, Junfeng Lei, Jinglong Zhou, Reinhard Klette, Haigang Sui |
Neurocomputing | 1 |
| 2020 | Novel shrinking residual convolutional neural network for efficient accurate stereo matching
Junfeng Lei, Jinsheng Xiao, Yunhua Chen |
J. Vis. Commun. Image Represent. | 4 |
| 2020 | A comprehensive approach for road marking detection and recognition
Huyin Zhang, Jinsheng Xiao, Shejie Lu, Reinhard Klette, Mohammad Norouzifard |
Multim. Tools Appl. | 3 |
| 2019 | An improved image mixed noise removal algorithm based on super-resolution algorithm and CNN
Huyin Zhang, Jinsheng Xiao, Shejie Lu, Mohammad Norouzifard |
Neural Comput. Appl. | 3 |
| 2019 | Improving the Antinoise Ability of DNNs via a Bio-Inspired Noise Adaptive Activation Function Rand SoftplusabstractAlthough deep neural networks (DNNs) have led to many remarkable results in cognitive tasks, they are still far from catching up with human-level cognition in antinoise capability. New research indicates how brittle and susceptible current models are to small variations in data distribution. In this letter, we study the stochasticity-resistance character of biological neurons by simulating the input-output response process of a leaky integrate-and-fire (LIF) neuron model and proposed a novel activation function, rand softplus (RSP), to model the response process. In RSP, a scale factor [Formula: see text] is employed to mimic the stochasticity-adaptability of biological neurons, thereby enabling the antinoise capability of a DNN to be improved by the novel activation function. We validated the performance of RSP with a 19-layer residual network (ResNet) and a 19-layer visual geometry group (VGG) on facial expression recognition data sets and compared it with other popular activation functions, such as rectified linear units (ReLU), softplus, leaky ReLU (LReLU), exponential linear unit (ELU), and noisy softplus (NSP). The experimental results show that RSP is applied to VGG-19 or ResNet-19, and the average recognition accuracy under five different noise levels exceeds the other functions on both of the two facial expression data sets; in other words, RSP outperforms the other activation functions in noise resistance. Compared with the application in ResNet-19, the application of RSP in VGG-19 can improve a network's antinoise performance to a greater extent. In addition, RSP is easier to train compared to NSP because it has only one parameter to be calculated automatically according to the input data. Therefore, this work provides the deep learning community with a novel activation function that can better deal with overfitting problems. Yunhua Chen, Yingchao Mai, Jinsheng Xiao |
Neural Comput. | 3 |
| 2018 | An Image Rain Removal algorithm based on the depth of field and sparse codingabstractRainfall weather can always seriously deteriorate the quality of the outdoor monitoring system image. Since the decomposition based methods do not need to impose any restrictions on the types of rain, they have a wider application in removing the rain streaks. However, they still have the problems of rain residues in the low frequency component, and mis-matching the background and the rain streaks with the same gradient in the high frequency. In this condition, we propose an image rain removal algorithm based on the depth of field and sparse coding. The algorithm includes four steps: image decomposition, dictionary learning, atomic clustering based on Principal Component Analysis and Support Vector Machine, image revising based on the depth of field saliency map. Firstly, the image is decomposed by using the combination of bilateral filtering and short-time Fourier transform, so that the contour in the low-frequency part of the image can be better preserved. The depth of field saliency map of the image is utilized to eliminate the rain residues in the low frequency components, and also to solve the problem of mis-matching the background and the rain streaks with the same gradient in the high frequency components. The experimental results demonstrate that the proposed algorithm performs better both in rain removal and preserving the detailed information of the image than current methods. Junfeng Lei, Shangyue Zhang, Wentao Zou, Jinsheng Xiao, Yunhua Chen, Haigang Sui |
ICPR | 4 |
| 2018 | Single-image Dehazing Algorithm Based on Convolutional Neural NetworksabstractThe paper proposes a new single image dehazing method based on a convolutional neural network. Our method directly learns an end to end mapping between the haze images and their corresponding haze layers (i.e. residual images between haze images and non-haze images). A convolutional neural network takes the haze image as an input and the residual image as an output. Then, a recovered dehazed image can be obtained by removing the residual from the haze image. Residual learning allows the network to directly estimate the initial haze layer with relatively high learning rates, which reduce computational complexity and speed-up the convergence process. Since the initial haze layer is only approximate, we use a guided filter to refine this layer to avoid halos and block artefacts, which makes the recovered image more similar to a real scene. The algorithms are tested on fog images with different fog densities. Comparisons are provided with other dehazing algorithms. Experiments demonstrate that the proposed algorithm outperforms state-of-the-art methods on both synthetic and real-world images, qualitatively and quantitatively. Jinsheng Xiao, Enyu Liu, Junfeng Lei, Reinhard Klette |
ICPR | 1 |
| 2018 | Lane Detection and Road Surface Reconstruction Based on Multiple Vanishing Point & SymposiaabstractLane detection algorithm based on monocular camera is one of the most popular methods in recent years, which can meet the requirement of real-time and robust for autonomous vehicle. In this way, the position of lane markers can be transferred from perspective space to road space base on the planar road assumption. However, large numbers of road scenes, especially the up and down slope road environment, cannot meet this requirement.In this paper, we propose a multiple vanishing point detection method to reconstruct the road space in slope scenes. In order to improve the accuracy of vanishing point estimation, the road images are decomposed into near and far regions. We extract candidate lane markers in near region by using multiscale convolution kernel and Hough Transform at first. Then, the lane markers in far region can be detected based on the result of near region. At last, different vanishing points are extracted in near and far regions. With the help of a vanishing point based on camera model, we can project both of near and far regions into road space. The experiment is conducted on our self-driving car `TuLian' in campus environment. Jian Zhou 0011, Jinsheng Xiao, Weicheng Zeng |
Intelligent Vehicles Symposium | 5 |
| 2018 | Frame Interpolation Algorithm Using Improved 3-D Recursive Search
Honggang Xie, Jinsheng Xiao, Qian Jia |
PRCV (1) | 3 |
| 2018 | Blind video denoising via texture-aware noise estimation
Jinsheng Xiao, Hong Tian, Yongqin Zhang, Yongqiang Zhou, Junfeng Lei |
Comput. Vis. Image Underst. | 1 |
| 2018 | Video denoising algorithm based on improved dual-domain filtering and 3D block matchingabstractThis study introduces an algorithm for video denoising based on improved dual‐domain filtering and 3D block matching. The wavelet thresholding based on 3D block matching is introduced to make full use of the correlation of video sequence in order to apply dual‐domain filtering to the video. A layered approach is used that attempts denoising in both a base layer and a detail layer. The result of wavelet thresholding based on 3D block matching is used as a guide image to make the base layer smoother. Shrinkage of short‐time Fourier transform coefficients further decreases the noise in the detail layer. Experimental results show that the authors’ algorithm generates a better base layer and detail layer than the traditional dual‐domain filtering algorithms. The subjective and objective comparisons of different algorithms also prove that the proposed algorithm performs better for video denoising. Jinsheng Xiao, Wentao Zou, Shangyue Zhang, Junfeng Lei, Yuan-Fang Wang |
IET Image Process. | 1 |
| 2018 | Kernel Wiener filtering model with low-rank approximation for image denoising
Yongqin Zhang, Jinsheng Xiao, Jiaying Liu 0001, Zongming Guo, Xiaopeng Zong |
Inf. Sci. | 2 |
| 2018 | Single image rain removal based on depth of field and sparse coding
Jinsheng Xiao, Wentao Zou, Yunhua Chen, Junfeng Lei |
Pattern Recognit. Lett. | 1 |
| 2017 | Image Noise Estimation Based on Principal Component Analysis and Variance-Stabilizing Transformation
Huying Zhang, Jinsheng Xiao, Jian Zhou 0011 |
ICIG (3) | 4 |
| 2017 | Lane Detection Based on Road Module and Extended Kalman Filter
Jinsheng Xiao, Wentao Zou, Reinhard Klette |
PSIVT | 1 |
| 2017 | Scene-aware image dehazing based on sky-segmented dark channel priorabstractAn improved dehazing algorithm based on dark channel theory is proposed, in order to solve the problems of colour distortion and halo effect which still exists in dark channel prior algorithm. The dark channel prior theory may lead to colour distortion in sky region. Firstly, the guided filter is used to refine the segmentation of the sky region, and the atmospheric light is estimated accurately. Then, the median filter is used to obtain the detailed edge information. So a more clear transmission can be gotten which effectively suppress the halo problem. Finally, the gamma correction is applied to enhance image lightness with an empirically selected gamma parameter. The experimental results show that the proposed algorithm can effectively remove the haze. It can correct the colour distortion of the sky area and eliminate the halo effect at the edge of the scene. Jinsheng Xiao, Yongqin Zhang, Enyu Liu, Junfeng Lei |
IET Image Process. | 1 |
| 2017 | Detail enhancement of image super-resolution based on detail synthesis
Jinsheng Xiao, Enyu Liu, Yuan-Fang Wang, Wenbin Jiang 0001 |
Signal Process. Image Commun. | 1 |
| 2016 | Adaptive shock filter for image super-resolution and enhancement
Jinsheng Xiao, Guanlin Pang, Yongqin Zhang, Yuli Kuang, Yixiang Wang |
J. Vis. Commun. Image Represent. | 1 |
| 2016 | Multi-focus image fusion based on depth extraction with inhomogeneous diffusion equation
Jinsheng Xiao, Yongqin Zhang, Baiyu Zou, Junfeng Lei, Qingquan Li 0001 |
Signal Process. | 1 |
| 2015 | Learning block-structured incoherent dictionaries for sparse representation
Yongqin Zhang, Jinsheng Xiao, Shuhong Li, Caiyun Shi, Guoxi Xie |
Sci. China Inf. Sci. | 2 |
| 2014 | Hierarchical tone mapping based on image colour appearance modelabstractTo solve the problem of low efficiency and poor effect of the current tone mapping methods for the high dynamic range images, the authors propose a hierarchical tone mapping algorithm based on colour appearance model. The discrete Gaussian kernel is used to speed up the bilateral filter. The operation of tone compression in RGB colour space is adopted to correct the colour casts. The extreme values of the pixels are also adjusted in the detail layer. Moreover, after the tone mapping, the colour saturation is enhanced in the image regions of rich details and sharp edges. Experimental results show that the proposed algorithm with less computational cost reduces the halo effect significantly, and achieves the natural colour and the rich details. It outperforms the state‐of‐the‐art methods in terms of visual quality and objective indicators. Jinsheng Xiao, Guoxiong Liu, Shih-Lung Shaw, Yongqin Zhang |
IET Comput. Vis. | 1 |