Xichuan Zhou

dblp:116/3393 · DBLP profile ↗
← Back
57ranked-venue papers
22as first author
37since 2021 · last 2027
0000-0002-3304-3045ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 26 · 10 first-author · 15 since 2021Applied, interdisciplinary, general and emerging computing · 19 · 9 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 4 first-author · 9 since 2021Systems, architecture and hardware · 3 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-authorComputer networks · 1 · 1 since 2021
YearPublicationVenuePosition
2027 ST-JSCC: Synergizing structural and textural dependencies for robust and efficient image transmission
Rulong He, Mingyang Wan, Haoming Luo, Xichuan Zhou, Haijun Liu 0001
Signal Process.5
2026 Remote sensing optical image matching through neighborhood-aware global propagation in graph neural networks
Yanchun Liu, Gemine Vivone, Jing Nie 0001, Haijun Liu 0001, Xichuan Zhou, Lihui Chen 0002
Eng. Appl. Artif. Intell.5
2026 Multiscale wavelet-based spatial-spectral compression network for hyperspectral image
Mingyang Wan, Aibin Peng, Xiangfei Shen, Rulong He, Lihui Chen 0002, Haijun Liu 0001, Xichuan Zhou
Eng. Appl. Artif. Intell.9
2026 PM-adapter: MoE based dynamic denoising fine-tuning for thermal infrared object detection
Haijun Liu 0001, Boya Wei, Jing Nie 0001, Suju Li, Xichuan Zhou
Neurocomputing7
2026 Sparse gain adaptation with dual-domain fusion network for multimodal object detection
Xichuan Zhou, Boya Wei, Cong Mao, Lihui Chen 0002, Haijun Liu 0001, Jin Xie 0005, Jing Nie 0001
Neurocomputing1
2026 RA-PTQ: Reparameterization-Aware Post-Training Quantization for accurate vision transformers in low-bit scenarios
Rui Ding 0009, Sihuan Zhao, Jing Nie 0001, Lihui Chen 0002, Haijun Liu 0001, Xichuan Zhou
Knowl. Based Syst.7
2026 Multimodality Image Registration With Modality Distillation
abstract
Multimodal image registration aims to spatially align images from different modalities at the pixel level. However, due to the nonlinear relationship of radiation intensities caused by different imaging modalities, achieving high accuracy in multimodal image registration presents a significant challenge. Additionally, the presence of both global transformations (i.e., large-scale rigid affine transformations) and local distortions (i.e., small-scale nonrigid deformations) between paired images further complicates the registration process. This article addressed the challenge resulting from modality differences through modality distillation. Specifically, a teacher (i.e., a homomodal image registration model) is trained to guide the student (i.e., a multimodal image registration model). Besides, this article simultaneously aligned large-scale rigid and small-scale nonrigid deformations by predicting deformation flow from both global and local features, thereby achieving high-precision registration. Furthermore, this proposed method incorporated a deformation mask during training to mitigate the negative impact of black edges in the obtained registration results on model performance. Experimental results demonstrate that the proposed method delivers state-of-the-art registration accuracy across various multimodal datasets, with ablation studies confirming the effectiveness of each component. The codes will be available at https://github.com/2351056918/Multimodality-Image-Registration-with-Modailty-Distillation.
Xichuan Zhou, Jicheng Zhao, Lihui Chen 0002, Gemine Vivone, Yanchun Liu, Jing Nie 0001, Haijun Liu 0001
IEEE Trans. Neural Networks Learn. Syst.1
2026 CAPilot: A High-Performance and High-Reliability Communication Middleware for Autonomous Driving
abstract
With the swift advancement of artificial intelligence technology, autonomous driving has increasingly emerged as a pivotal technology in the future of transportation. Real-time data exchange and processing across modules in autonomous driving systems necessitate efficient and reliable communication middleware. However, existing communication methods suffer from delay, congestion and packet loss when dealing with high-frequency and large-data-volume transmission tasks, significantly impairing system performance and security. To reduce communication latency and CPU overhead, a multi-mode adaptive high-performance and high-reliability communication middleware CAPilot is proposed. Firstly, a novel shared-memory communication architecture is proposed, comprising a Data Pool, an Event Notification Index Pool, and a Cycle Index Pool. The Data Pool employs a lock-free mechanism to avert deadlock and starvation issues, while addressing frame-skipping using a real-time maintenance and discriminative approach. Event-triggered and period-triggered data acquisition strategies proficiently circumvent data security concerns and performance limitations inherent in conventional shared memory connectivity. Then, to mitigate the overhead associated with dynamic broadcasts within the constrained embedded resources of the network, an adaptive communication scheme is proposed. This scheme incorporates a profile-based static communication encoding that automatically determines the optimal communication method based on the environments of the communicating entities. Finally, the intra-process pointer passing method is optimised by introducing a dual adaptive buffered ring queue, which facilitates bulk data retrieval without using locks. Experimental results show that CAPilot outperforms existing communication middlewares such as ROS2, CyberRT, and DDS in terms of communication latency, message throughput, message frame loss rate, and resource utilisation. These advancements suggest that CAPilot is well-suited for extensive deployment in diverse autonomous driving applications.
Pinzhong Qin, Changquan Xue, Jing Nie 0001, Haijun Liu 0001, Xichuan Zhou
ACM Trans. Internet Techn.6
2025 EigenSR: Eigenimage-Bridged Pre-Trained RGB Learners for Single Hyperspectral Image Super-Resolution
abstract
Single hyperspectral image super-resolution (single-HSI-SR) aims to improve the resolution of a single input low-resolution HSI. Due to the bottleneck of data scarcity, the development of single-HSI-SR lags far behind that of RGB natural images. In recent years, research on RGB SR has shown that models pre-trained on large-scale benchmark datasets can greatly improve performance on unseen data, which may stand as a remedy for HSI. But how can we transfer the pre-trained RGB model to HSI, to overcome the data-scarcity bottleneck? Because of the significant difference in the channels between the pre-trained RGB model and the HSI, the model cannot focus on the correlation along the spectral dimension, thus limiting its ability to utilize on HSI. Inspired by the HSI spatial-spectral decoupling, we propose a new framework that first fine-tunes the pre-trained model with the spatial components (known as eigenimages), and then infers on unseen HSI using an iterative spectral regularization (ISR) to maintain the spectral correlation. The advantages of our method lie in: 1) we effectively inject the spatial texture processing capabilities of the pre-trained RGB model into HSI while keeping spectral fidelity, 2) learning in the spectral-decorrelated domain can improve the generalizability to spectral-agnostic data, and 3) our inference in the eigenimage domain naturally exploits the spectral low-rank property of HSI, thereby reducing the complexity. This work bridges the gap between pre-trained RGB models and HSI via eigenimages, addressing the issue of limited HSI training data, hence the name EigenSR. Extensive experiments show that EigenSR outperforms the state-of-the-art (SOTA) methods in both spatial and spectral metrics.
Xi Su, Xiangfei Shen, Mingyang Wan, Jing Nie 0001, Lihui Chen 0002, Haijun Liu 0001, Xichuan Zhou
AAAI7
2025 Hybrid cross-modality fusion network for medical image segmentation with contrastive learning
Xichuan Zhou, Jing Nie 0001, Haijun Liu 0001, Fu Liang, Lihui Chen 0002, Jin Xie 0005
Eng. Appl. Artif. Intell.1
2025 E-TransConvNet: An enhanced transformer and convolutional network for medical image segmentation from ultrasound and CT images
Chukwuemeka Clinton Atabansi, Jing Nie 0001, Jiachen Huang, Haijun Liu 0001, Jin Xie 0005, Xichuan Zhou
Expert Syst. Appl.7
2025 Distribution-modulated binary neural network for image classification
Yingcheng Lin, Rui Ding 0009, Haijun Liu 0001, Xichuan Zhou
Image Vis. Comput.5
2025 Progressive fine-to-coarse reconstruction for accurate low-bit post-training quantization in vision transformers
Rui Ding 0009, Liang Yong, Sihuan Zhao, Jing Nie 0001, Lihui Chen 0002, Haijun Liu 0001, Xichuan Zhou
Neural Networks7
2025 Binary Neural Networks With Feature Information Retention for Efficient Image Classification
abstract
Although binary neural networks (BNNs) enjoy extreme compression ratios, there are significant accuracy gap compared with full-precision models. Previous works propose various strategies to reduce the information loss induced by the binarization process, improving the performance of binary neural networks to some extent. However, in this letter, we argue that few studies try to alleviate this problem from the structure perspective, resulting in inferior performance. To this end, we propose a novel Feature Information Retention Network named FIRNet, which incorporates an extra path to propagate the untouched informative feature maps. Specifically, the FIRNet splits the input feature maps into two groups, one of which is fed into the normal layers and another kept untouched for information retention. Then we utilize the concatenation, shuffle and pooling operations to process these features with 64× memory saving. Finally, with only a 1.7% complexity increase, a FIR fusion layer is proposed to aggregate the features from two branches. Experimental results demonstrate that our proposed method achieves 1.0% Top-1 accuracy improvement over the baseline model and outperforms other state-of-the-art BNNs on the ImageNet dataset.
Rui Ding 0009, Haijun Liu 0001, Xichuan Zhou
IEEE Signal Process. Lett.4
2025 High-Fidelity Pansharpening via Trigeminal Pyramid Decoding of CNN-Transformer Encoded Features
abstract
Spectral and spatial fidelity remains a longstanding challenge in the field of pansharpening, which aims to generate high-resolution multispectral (HRMS) images by integrating high-resolution panchromatic (PAN) images with low-resolution multispectral (LRMS) images. This study proposes a high-fidelity pansharpening network that utilizes bidirectional trigeminal pyramid decoding of features encoded by a CNN-Transformer architecture. Specifically, local and global features at multiple scales are initially extracted using a CNN-Transformer encoder to facilitate multi-scale feature fusion. Subsequently, we design a decoder based on bidirectional trigeminal pyramids to achieve a high-fidelity fusion output. One reverse decoding pyramid decodes the fused features of LRMS and PAN images from the encoder. One spectral feature pyramid is employed to enhance the spectral information of the reverse decoding pyramid, while the last spatial feature pyramid is utilized to enrich the spatial information, thereby improving the overall spectral and spatial fidelity of the fused output. Furthermore, content-guided attention (CGA) is incorporated to adaptively integrate the spectral and spatial feature pyramids into the reverse decoding pyramid. Extensive experiments demonstrate that our network surpasses the comparative state-of-the-art (SOTA) methods in both qualitative and quantitative evaluations. The code is available at https://github.com/songvvvv/pansharpening.
Lihui Chen 0002, Tianxin Song, Lihua Jian, Di Zhang 0002, Gemine Vivone, Xichuan Zhou
IEEE Trans. Geosci. Remote. Sens.6
2025 Bi-SSFormer: An Ultralightweight Binary Spectral-Spatial Transformer for Hyperspectral Image Classification
Rui Ding 0009, Yanchun Liu, Baoliang Wang, Lihui Chen 0002, Haijun Liu 0001, Gemine Vivone, Xichuan Zhou
IEEE Trans. Geosci. Remote. Sens.9
2025 MIT-SAM: Medical Image-Text SAM With Mutually Enhanced Heterogeneous Features Fusion for Medical Image Segmentation
abstract
In recent times, leveraging lesion text as supplementary data to enhance the performance of medical image segmentation models has garnered attention. Previous approaches only used attention mechanisms to integrate image and text features, while not effectively utilizing the highly condensed textual semantic information in improving the fused features, resulting in inaccurate lesion segmentation. This paper introduces a novel approach, the Medical Image-Text Segment Anything Model (MIT-SAM), for text-assisted medical image segmentation. Specifically, we introduce the SAM-enhanced image encoder and a Bert-based text encoder to extract heterogeneous features. To better leverage the highly condensed textual semantic information for heterogeneous feature fusion, such as crucial details like position and quantity, we propose the image-text interactive fusion (ITIF) block and self-supervised text reconstruction (SSTR) method. The ITIF block facilitates the mutual enhancement of homogeneous information among heterogeneous features and the SSTR method empowers the model to capture crucial details concerning lesion text, including location, quantity, and other key aspects. Experimental results demonstrate that our proposed model achieves state-of-the-art performance on the QaTa-COV19 and MosMedData+ datasets.
Xichuan Zhou, Lingfeng Yan, Rui Ding 0009, Chukwuemeka Clinton Atabansi, Jing Nie 0001, Lihui Chen 0002, Haijun Liu 0001
IEEE J. Biomed. Health Informatics1
2024 GOENet: Group Operations Enhanced Binary Neural Network for Efficient Image Classification
abstract
There exists an innegligible performance gap between the binary neural networks and their full-precision counterparts, which prevents their deployment on real-world applications. Recently, plenty of researchers strive to solve this problem by incorporating more binary subnets, improving the representational power with acceptable complexity increase. However, current methods make structure design and parallel acceleration sophisticated and non-trivial and may degenerate the representational power due to the isomorphic subnets. Besides, the final feature fusion method is sub-optimal. In this letter, we propose a simple yet effective binary neural network named GOENet enhanced by group operations. The multiple parallel binary subnets use group-wise binary thresholds to improve feature diversity and are merged into the group convolutional layers. The inter-subnet connections are implemented with the parameter-free group shuffle operations, improving the model parallelism and performance. We also propose the group feature summation module as a better fusion method for its efficiency and effectiveness. Experimental results on ImageNet show that our proposed method outperforms the SOTA GroupNet by 1.7% Top-1 accuracy with 0.84× and 0.89× saving in computation and memory complexity.
Rui Ding 0009, Haijun Liu 0001, Xichuan Zhou
IEEE Signal Process. Lett.4
2024 Theory and Low-Power Design of Moving Accumulative Sign Filter
abstract
A novel down-sampling filter named moving accumulative sign filter (MASF) is proposed for low-power down-sampling of large-scale binary and ternary data. Besides, the MASF has greatly circuit realization advantages than state-of-the-art cascaded-integrator-comb (CIC) filter, especially in the area of low-power design. The theory of MASF is proposed and introduced comprehensively, including the algorithm model, transfer function, and frequency response characteristics. The pipeline voting architecture is applied to the implementation of the MASF to improve the speed of data processing, which simplifies the circuit structure and reduce the power consumption. The MASF circuits of general application based on pipeline voting are designed for binary and ternary signals only using D flip-flop and logic gates. The area and power consumption of MASF are reduced by 86% and 88% compared with CIC filter under the same conditions on FPGA. What’s more, a hardware-friendly pooling algorithm named polar-pooling is proposed based on MASF for binary and ternary feature maps, which greatly reduces the time and space complexity of pooling. Compared with max-pooling and average-pooling, the processing time of polar-pooling is reduced by more than 75% for a$200\times 200$binary image. The two-stage MASF circuit for ternary signal processing is implemented at 40-nm CMOS process, compared with state-of-the-arts cascade-of-integrators filter which cascading two integrators, the normalized power consumption of proposed two-stage MASF circuit has 67% reduction and the area has 75% reduction.
Yingjun Xia, Jianjiang Luo, Peng Yin 0004, Dengwei Yan, Xichuan Zhou, Amine Bermak, Fang Tang
IEEE Trans. Circuits Syst. I Regul. Pap.5
2024 MSNet: Self-Supervised Multiscale Network With Enhanced Separation Training for Hyperspectral Anomaly Detection
abstract
Hyperspectral anomaly detection (HAD) has attracted increasing attention due to its economical and efficient applications. The main challenge lies in the data-starved problem of hyperspectral images (HSIs) and the costliness of manual annotation, making it heavily reliant on the model’s adaptability and robustness to unseen scenes under limited samples. Self-supervised learning offers a solution to this urgency via mining meaningful representations from the data itself. One promising paradigm is leveraging untrained neural networks to reconstruct the background component for revealing anomalous information. Its capability stems from the network architecture and the training process rather than learning from expensive and strongly domain-dependent data, which is naturally applicable to HAD. In this article, to handle the urgent requirement for self-supervised learning in HAD, we propose a multiscale network (termed MSNet) that detects anomalies with enhanced separation training. The network architecture consists of several multiscale convolutional encoder-decoder (CED) layers, considering the spatial characteristics of the anomalies. To suppress the anomalies during background reconstruction, we adopt a new separation training strategy by introducing a soft separator for better practicality on larger datasets. Extensive experiments conducted on five commonly used datasets and the HAD100 dataset, demonstrate the superiority of our method over its counterparts. Our code is available athttps://github.com/enter-i-username/MSNet.
Haijun Liu 0001, Xi Su, Xiangfei Shen, Xichuan Zhou
IEEE Trans. Geosci. Remote. Sens.4
2024 MeSAM: Multiscale Enhanced Segment Anything Model for Optical Remote Sensing Images
abstract
Segment anything model (SAM) has been widely applied to various downstream tasks for its excellent performance and generalization capability. However, SAM exhibits three limitations related to remote sensing semantic segmentation task: 1) the image encoders excessively lose high-frequency information, such as object boundaries and textures, resulting in rough segmentation masks; 2) due to being trained on natural images, SAM faces difficulty in accurately recognizing objects with large-scale variations and uneven distribution in remote sensing images; 3) the output tokens used for mask prediction are trained on natural images and not applicable to remote sensing image segmentation. In this paper, we explore an efficient paradigm for applying SAM to the semantic segmentation of remote sensing images. Furthermore, we propose MeSAM, a new SAM fine-tuning method more suitable for remote sensing images to adapt it to semantic segmentation tasks. Our method first introduces an inception mixer into the image encoder to effectively preserve high-frequency features. Secondly, by designing a mask decoder with remote-sensing correction and incorporating multiscale connections, we make up the difference in SAM from natural images to remote sensing images. Experimental results demonstrated that our method significantly improves the segmentation accuracy of SAM for remote sensing images, outperforming some state-of-the-art methods. The code will be available at https://github.com/Magic-lem/MeSAM.
Xichuan Zhou, Fu Liang, Lihui Chen 0002, Haijun Liu 0001, Gemine Vivone, Jocelyn Chanussot
IEEE Trans. Geosci. Remote. Sens.1
2024 BTC-Net: Efficient Bit-Level Tensor Data Compression Network for Hyperspectral Image
abstract
Now it is still a challenge to compress high-throughput hyperspectral tensor image data on lightweight air-carried/spaceborne remote sensing systems, primarily due to insufficient computational resources and limited transmission bandwidth. To address this challenge, we propose a bit-level tensor data compression network (BTC-Net) that provides higher compression performance by leveraging a data-driven lightweight quantized neural encoder with two-stage bit compression. The BTC-Net achieves semantic near-lossless high reconstruction quality at low compression bit rates thanks to its optimized decoder, which uses a channel-wise attention-based enhancement module to recover hyperspectral tensor data. Experimental results on different hyperspectral datasets show that the BTC-Net could achieve an extremely low compression bit rate of fewer than 0.04 bits per pixel per band (bpppb) with state-of-the-art reconstruction performances. The demo of BTC-Net will be publicly available online at: https://github.com/zx20173646/BTCNet.
Xichuan Zhou, Xuan Zou, Xiangfei Shen, Wenjia Wei, Haijun Liu 0001
IEEE Trans. Geosci. Remote. Sens.1
2023 Low-cost real-time VLSI system for high-accuracy optical flow estimation using biological motion features and random forests
Cong Shi 0003, Junxian He, Shrinivas J. Pundlik, Xichuan Zhou, Nanjian Wu, Gang Luo 0003
Sci. China Inf. Sci.4
2023 InfraNet: Accurate forehead temperature measurement framework for people in the wild with monocular thermal infrared camera
Xichuan Zhou, Dongshan Lei, Chunqiao Long, Jing Nie 0001, Haijun Liu 0001
Neural Networks1
2023 Efficient Hyperspectral Sparse Regression Unmixing With Multilayers
abstract
The sparse regression method is known for its ability to unmix hyperspectral data, but it can be computationally expensive and accurately insufficient due to the large scale and high coherence of the spectral library. To address this issue, a new approach called layered sparse regression unmixing (termed LSU) has been proposed in this paper. This method involves breaking down the sparse unmixing process into multilayers, each of which interactively learns a row-sparsity-promoting abundance matrix and fine-tunes active library atoms based on measured activeness. By doing so, LSU outputs both a learned abundance matrix and an optimal library that can best model each mixed pixel in the scene. The proposed LSU can be efficiently solved by the alternating direction method of the multipliers framework. Experimental results obtained from simulated and real hyperspectral images demonstrate the effectiveness of LSU. The demo of the proposed LSU will be publicly available at https://github.com/XiangfeiShen/Layered_Sparse_Regression_Unmixing.
Xiangfei Shen, Lihui Chen 0002, Haijun Liu 0001, Xi Su, Wenjia Wei, Xichuan Zhou
IEEE Trans. Geosci. Remote. Sens.7
2023 Matrix Factorization With Framelet and Saliency Priors for Hyperspectral Anomaly Detection
abstract
Hyperspectral anomaly detection aims to separate sparse anomalies from low-rank background components. A variety of detectors have been proposed to identify anomalies, but most of them tend to emphasize characterizing backgrounds with multiple types of prior knowledge and limited information on anomaly components. To tackle these issues, this article simultaneously focuses on two components and proposes a matrix factorization method with framelet and saliency priors to handle the anomaly detection problem. We first employ a framelet to characterize nonnegative background representation coefficients, as they can jointly maintain sparsity and piecewise smoothness after framelet decomposition. We then exploit saliency prior knowledge to measure each pixel’s potential to be an anomaly. Finally, we incorporate the pure pixel index (PPI) with Reed-Xiaoli’s (RX) method to possess representative dictionary atoms. We solve the optimization problem using a block successive upper-bound minimization (BSUM) framework with guaranteed convergence. Experiments conducted on benchmark hyperspectral datasets demonstrate that the proposed method outperforms some state-of-the-art anomaly detection methods.
Xiangfei Shen, Haijun Liu 0001, Jing Nie 0001, Xichuan Zhou
IEEE Trans. Geosci. Remote. Sens.4
2023 Cellular Binary Neural Network for Accurate Image Classification and Semantic Segmentation
abstract
This paper presents the Cellular Binary Neural Network (CBNN), which is an efficient deep neural network with binary weights and activations. To address the challenge of performance drop caused by low-precision representation, the CBNN adopts multiple subnets which are connected via learnable global lateral paths. The introduced lateral connections are assumed to be sparse and grouped with respect to different source layers. The inter-network lateral connections and inner-network parameters are simultaneously optimized by the distributional loss, classification loss and the group sparse regularization term. Experiments on the CIFAR-10 and ImageNet datasets showed that, by incorporating optimized group-sparse lateral paths, the CBNN outperformed many state-of-the-art binary neural networks in terms of classification accuracy. Besides, to verify the generalization of the proposed binary model, we extended the CBNN on semantic segmentation task. CBNN takes advantage of the multiple subnets to derive the more informative feature maps which are computed by the parallel aggregation in the last convolution block. Experiments on PASCAL VOC segmentation dataset demonstrated that, under the same segmentation settings, the proposed method achieved the superior performance over other compared networks and even the full-precision counterpart.
Xichuan Zhou, Rui Ding 0009, Wenjia Wei, Haijun Liu 0001
IEEE Trans. Multim.1
2022 Superpixel-Guided Local Sparsity Prior for Hyperspectral Sparse Regression Unmixing
abstract
Sparse regression relaxes the difficulties of blind unmixing of hyperspectral data thanks to the spectral library. Many investigations, however, attach importance to global priors such as sparsity and low-rankness. This letter proposes a local-global-based sparse regression unmixing method, called LGSU, by introducing a local sparsity regularization to help boost the unmixing performance that only considers global sparsity. The proposed LGSU first uses a superpixel-based technique to yield a set of homogeneous superpixels for guiding local sparse regularization purposes. LGSU then considers a traditional ℓ1regularization to enhance global sparsity. Coupling with local and global sparsity constraints, the proposed LGSU can effectively estimate the abundance of a given image via the alternating direction method of multipliers. Experimental results obtained from synthetic and real hyperspectral images demonstrate the effectiveness of the proposed algorithm.
Xiangfei Shen, Haijun Liu 0001, Xinzheng Zhang 0002, Xichuan Zhou
IEEE Geosci. Remote. Sens. Lett.5
2022 LC-BiDet: Laterally Connected Binary Detector With Efficient Image Processing
abstract
Recently, binary neural networks have received increasing interest in object detection community, since compared to the full-precision detection networks, they can reduce the memory and computation requirements significantly due to the efficient XNOR and BITCOUNT operations introduced by binarization. However, the final detection performance of binary detectors always suffers from a severe degradation compared with the full-precision counterparts. In this letter, to achieve a better trade-off between the inference efficiency and the detection performance, we propose to learn a laterally connected binary detector based on the multiple parallel binary subnets with a group sparse regularization term. Firstly, we simply adopt the parallel structure to improve the representation capacity of binary detectors. Secondly, we introduce the dense lateral connections between binary subnets in each convolutional layer to increase the information flows. Finally, to reduce the redundancy of the dense connections, we take the L21 norm as the regularization term to optimize the lateral connections in an end-to-end manner. Experimental results on PASCAL VOC and COCO datasets show that our proposed laterally connected binary detector could outperform the other state-ofthe- art binary detectors.
Xichuan Zhou, Rui Ding 0009, Haijun Liu 0001
IEEE Signal Process. Lett.1
2022 Toward Weak Signal Analysis in Hyperspectral Data: An Efficient Unmixing Perspective
abstract
Many unmixing methods hold the assumption that endmembers correspond to major land-covers, but not true for some unmixing tasks where observed minor object signals corresponding to some special types of endmembers are relatively weak. When there exist weak signals that have low intensity potentially caused by subtle mixing abundance fractions regarding the endmembers of minor objects, the traditional unmixing techniques may fail. This paper pioneers weak signal scenarios in hyperspectral unmixing using an efficient method called HyperWeak. Specifically, HyperWeak involves a sparse nonnegative matrix factorization model that contains two main parts, where the unsupervised part estimates the endmember and abundance matrices, and the supervised part ensures the minimal degradation of prior knowledge. To enhance the robustness of the HyperWeak model, this paper considers a reweighted sparsity constraint to boost the sparseness of the abundance matrix. For effectively solving optimization problems, Nesterov’s optimal gradient method is used in this paper. Experiments conducted on synthetic and real hyperspectral images indicate that HyperWeak can improve the unmixing performances of hyperspectral data in weak signal situations.
Xiangfei Shen, Haijun Liu 0001, Fangyuan Ge, Xichuan Zhou
IEEE Trans. Geosci. Remote. Sens.5
2021 Optimizing Information Theory Based Bitwise Bottlenecks for Efficient Mixed-Precision Activation Quantization
abstract
Recent researches on information theory shed new light on the continuous attempts to open the black box of neural signal encoding. Inspired by the problem of lossy signal compression for wireless communication, this paper presents a Bitwise Bottleneck approach for quantizing and encoding neural network activations. Based on the rate-distortion theory, the Bitwise Bottleneck attempts to determine the most significant bits in activation representation by assigning and approximating the sparse coefficients associated with different bits. Given the constraint of a limited average code rate, the bottleneck minimizes the distortion for optimal activation quantization in a flexible layer-by-layer manner. Experiments over ImageNet and other datasets show that, by minimizing the quantization distortion of each layer, the neural network with bottlenecks achieves the state-of-the-art accuracy with low-precision activation. Meanwhile, by reducing the code rate, the proposed method can improve the memory and computational efficiency by over six times compared with the deep neural network with standard single-precision representation. The source code is available on GitHub: https://github.com/CQUlearningsystemgroup/BitwiseBottleneck.
Xichuan Zhou, Cong Shi 0003, Haijun Liu 0001
AAAI1
2021 Learning to Binarize Convolutional Neural Networks with Adaptive Neural Encoder
abstract
The high computational complexity and memory consumption of the deep Convolutional Neural Networks (CNNs) restrict their deployability in resource-limited embedded devices. To address this challenge, emerging solutions are proposed for neural network quantization and compression. Among them, Binary Neural Networks (BNNs) show their potential in reducing computational and memory complexity; however, they suffer from considerable performance degradation. One of the major causes is their non-differentiable discrete quantization implemented using a fixed sign function, which leads to output distribution distortion. In this paper, instead of using the fixed and naive sign function, we propose a novel adaptive Neural Encoder (NE), which learns to quantize the full-precision weights as binary values. Inspired by the research of neural network distillation, a distribution loss is introduced as a regularizer to minimize the Kullback-Leibler divergence between the outputs of the full-precision model and the encoded binary model. With an end-to-end backpropagation training process, the adaptive neural encoder, along with the binary convolutional neural network, could reach convergence iteratively. Comprehensive experiments with different network structures and datasets show that the proposed method can improve the performance of the baselines and also outperform many state-of-the-art approaches. The source code of the proposed method is publicly available at https://github.com/CQUlearningsystemgroup/LearningToBinarize.
Fangyuan Ge, Rui Ding 0009, Haijun Liu 0001, Xichuan Zhou
IJCNN5
2021 A Heterogeneous Spiking Neural Network for Computationally Efficient Face Recognition
abstract
Computational efficiency is critical to many mobile and always-on face recognition applications. To this end, a heterogeneous spiking neural network (SNN) is proposed for face recognition. To obtain high recognition accuracy at minimal computational overheads, the heterogeneous SNN consists of an encoding subnet for sparse image feature encoding and classification subnet for feature classification. The experimental results suggest that the proposed heterogeneous algorithm can achieve high recognition accuracy on small datasets of human face samples with labeled identities at a high computational efficiency with very low neuronal activities. The proposed SNN is promising for low-cost mobile or always-on systems with strictly constrained resource and energy budgets.
Xichuan Zhou, Zhenghua Zhou, Zhengqing Zhong, Jianyi Yu, Tengxiao Wang, Min Tian 0003, Cong Shi 0003
ISCAS1
2021 CompSNN: A lightweight spiking neural network based on spatiotemporally compressive spike features
Tengxiao Wang, Cong Shi 0003, Xichuan Zhou, Yingcheng Lin, Junxian He, Ping Gan, Ping Li 0042, Ying Wang 0001, Nanjian Wu, Gang Luo 0003
Neurocomputing3
2021 Hapke Data Augmentation for Deep Learning-Based Hyperspectral Data Analysis With Limited Samples
abstract
The emerging technology of deep neural networks has been proven to be successful for hyperspectral image analysis. However, it is still a great challenge to apply the deep learning method for quantitatively retrieving mineralogical composition, because typical deep neural networks generally require thousands of labeled samples for training, while only a few mineral samples can be acquired and examined for quantitative examination in practice. To address this challenge, this letter proposes a training data augmentation approach which incorporates the prior-knowledge of hyperspectral reflectance characteristics using the classic Hapke equations. Experiments over both laboratory and airborne hyperspectral remote sensing data show that the proposed method outperforms the widely used approaches for quantitative mineral analysis.
Fangyuan Ge, Yingjun Zhao, Ming Li 0082, Cong Shi 0003, Dong Li 0007, Xichuan Zhou
IEEE Geosci. Remote. Sens. Lett.8
2021 Strong but Simple Baseline With Dual-Granularity Triplet Loss for Visible-Thermal Person Re-Identification
abstract
This letter presents a conceptually simple and effective dual-granularity triplet loss for visible-thermal person re-identification (VT-ReID). Generally, ReID models are always trained with the sample-based triplet loss and identification loss from the fine granularity level. Further, center-based loss could be introduced to encourage the intra-class compactness and inter-class discrimination from the coarse granularity level. Our proposed dual-granularity triplet loss well organizes the sample-based triplet loss and center-based triplet loss in a hierarchical fine to coarse granularity manner, just with some simple configurations of typical operations, such as pooling and batch normalization. Experiments on RegDB and SYSU-MM01 datasets show that with only the global features our dual-granularity triplet loss can improve the VT-ReID performance by a significant margin. It can be a strong VT-ReID baseline to boost future research with high quality.
Haijun Liu 0001, Yanxia Chai, Xiaoheng Tan, Dong Li 0007, Xichuan Zhou
IEEE Signal Process. Lett.5
2021 Parameter Sharing Exploration and Hetero-Center Triplet Loss for Visible-Thermal Person Re-Identification
abstract
This paper focuses on the visible-thermal cross-modality person re-identification (VT Re-ID) task, whose goal is to match person images between the daytime visible modality and the nighttime thermal modality. The two-stream network is usually adopted to address the cross-modality discrepancy, the most challenging problem for VT Re-ID, by learning the multi-modality person features. In this paper, we explore how many parameters a two-stream network should share, which is still not well investigated in the existing literature. By splitting the ResNet50 model to construct the modality-specific feature extraction network and modality-sharing feature embedding network, we experimentally demonstrate the effect of parameter sharing of two-stream network for VT Re-ID. Moreover, in the framework of part-level person feature learning, we propose the hetero-center triplet loss to relax the strict constraint of traditional triplet loss by replacing the comparison of theanchor to all the other samplesby theanchor center to all the other centers. With extremely simple means, the proposed method can significantly improve the VT Re-ID performance. The experimental results on two datasets show that our proposed method distinctly outperforms the state-of-the-art methods by large margins, especially on the RegDB dataset achieving superior performance, rank1/mAP/mINP 91.05%/83.28%/68.84%. It can be a new baseline for VT Re-ID, with a simple but effective strategy.
Haijun Liu 0001, Xiaoheng Tan, Xichuan Zhou
IEEE Trans. Multim.3
2020 When Single Event Upset Meets Deep Neural Networks: Observations, Explorations, and Remedies
abstract
Deep Neural Network has proved its potential in various perception tasks and hence become an appealing option for interpretation and data processing in security sensitive systems. However, security-sensitive systems demand not only high perception performance, but also design robustness under various circumstances. Unlike prior works that study network robustness from software level, we investigate from hardware perspective about the impact of Single Event Upset (SEU) induced parameter perturbation (SIPP) on neural networks. We systematically define the fault models of SEU and then provide the definition of sensitivity to SIPP as the robustness measure for the network. We are then able to analytically explore the weakness of a network and summarize the key findings for the impact of SIPP on different types of bits in a floating point parameter, layer-wise robustness within the same network and impact of network depth. Based on those findings, we propose two remedy solutions to protect DNNs from SIPPs, which can mitigate accuracy degradation from 28% to 0.27% for ResNet with merely 0.24-bit SRAM area overhead per parameter.
Zheyu Yan, Yiyu Shi 0001, Wang Liao 0001, Masanori Hashimoto, Xichuan Zhou, Cheng Zhuo
ASP-DAC5
2020 Probability Weighted Compact Feature for Domain Adaptive Retrieval
abstract
Domain adaptive image retrieval includes single-domain retrieval and cross-domain retrieval. Most of the existing image retrieval methods only focus on single-domain retrieval, which assumes that the distributions of retrieval databases and queries are similar. However, in practical application, the discrepancies between retrieval databases often taken in ideal illumination/pose/background/camera conditions and queries usually obtained in uncontrolled conditions are very large. In this paper, considering the practical application, we focus on challenging cross-domain retrieval. To address the problem, we propose an effective method named Probability Weighted Compact Feature Learning (PWCF), which provides inter-domain correlation guidance to promote cross-domain retrieval accuracy and learns a series of compact binary codes to improve the retrieval speed. First, we derive our loss function through the Maximum A Posteriori Estimation (MAP): Bayesian Perspective (BP) induced focal-triplet loss, BP induced quantization loss and BP induced classification loss. Second, we propose a common manifold structure between domains to explore the potential correlation across domains. Considering the original feature representation is biased due to the inter-domain discrepancy, the manifold structure is difficult to be constructed. Therefore, we propose a new feature named Histogram Feature of Neighbors (HFON) from the sample statistics perspective. Extensive experiments on various benchmark databases validate that our method outperforms many state-of-the-art image retrieval methods for domain adaptive image retrieval. The source code is available at {https://github.com/fuxianghuang1/PWCF}.
Fuxiang Huang, Lei Zhang 0038, Yang Yang 0002, Xichuan Zhou
CVPR4
2020 MoNet3D: Towards Accurate Monocular 3D Object Localization in Real Time
abstract
Monocular multi-object detection and localization in 3D space has been proven to be a challenging task. The MoNet3D algorithm is a novel and effective framework that can predict the 3D position of each object in a monocular image, and draw a 3D bounding box on each object. The MoNet3D method incorporates the prior knowledge of spatial geometric correlation of neighboring objects into the deep neural network training process, in order to improve the accuracy of 3D object localization. Experiments over the KITTI data set show that the accuracy of predicting the depth and horizontal coordinate of the object in 3D space can reach 96.25% and 94.74%, respectively. Meanwhile, the method can realize the real-time image processing capability of 27.85 FPS. Our code is publicly available at https://github.com/CQUlearningsystemgroup/YicongPeng
Xichuan Zhou, Yicong Peng, Chunqiao Long, Fengbo Ren, Cong Shi 0003
ICML1
2020 Build a compact binary neural network through bit-level sensitivity and data pruning
Yixing Li, Xichuan Zhou, Fengbo Ren
Neurocomputing3
2019 An Efficient Compressive Convolutional Network for Unified Object Detection and Image Compression
abstract
This paper addresses the challenge of designing efficient framework for real-time object detection and image compression. The proposed Compressive Convolutional Network (CCN) is basically a compressive-sensing-enabled convolutional neural network. Instead of designing different components for compressive sensing and object detection, the CCN optimizes and reuses the convolution operation for recoverable data embedding and image compression. Technically, the incoherence condition, which is the sufficient condition for recoverable data embedding, is incorporated in the first convolutional layer of the CCN model as regularization; Therefore, the CCN convolution kernels learned by training over the VOC and COCO image set can be used for data embedding and image compression. By reusing the convolution operation, no extra computational overhead is required for image compression. As a result, the CCN is 3.1 to 5.0 fold more efficient than the conventional approaches. In our experiments, the CCN achieved 78.1 mAP for object detection and 3.0 dB to 5.2 dB higher PSNR for image compression than the examined compressive sensing approaches.
Xichuan Zhou, Shujun Liu, Yingcheng Lin, Lei Zhang 0038, Cheng Zhuo
AAAI1
2019 Sparse representation of classified patches for CS-MRI reconstruction
Jianxin Cao, Shujun Liu, Hongqing Liu 0002, Xiaoheng Tan, Xichuan Zhou
Neurocomputing5
2019 A deep manifold learning approach for spatial-spectral classification with limited labeled training samples
Xichuan Zhou, Fang Tang, Yingjun Zhao, Lei Zhang 0038, Dong Li 0007
Neurocomputing1
2019 A Deep Learning Approach for Targeted Contrast-Enhanced Ultrasound Based Prostate Cancer Detection
abstract
The important role of angiogenesis in cancer development has driven many researchers to investigate the prospects of noninvasive cancer diagnosis based on the technology of contrast-enhanced ultrasound (CEUS) imaging. This paper presents a deep learning framework to detect prostate cancer in the sequential CEUS images. The proposed method uniformly extracts features from both the spatial and the temporal dimensions by performing three-dimensional convolution operations, which captures the dynamic information of the perfusion process encoded in multiple adjacent frames for prostate cancer detection. The deep learning models were trained and validated against expert delineations over the CEUS images recorded using two types of contrast agents, i.e., the anti-PSMA based agent targeted to prostate cancer cells and the non-targeted blank agent. Experiments showed that the deep learning method achieved over 91 percent specificity and 90 percent average accuracy over the targeted CEUS images for prostate cancer detection, which was superior ( ) than previously reported approaches and implementations.
Fan Yang 0021, Xichuan Zhou, Yanli Guo, Fang Tang, Fengbo Ren, Jishun Guo, Shuiwang Ji
IEEE ACM Trans. Comput. Biol. Bioinform.3
2018 Joint spectral-spatial hyperspectral image classification based on hierarchical subspace switch ensemble learning algorithm
Yongming Li 0003, Tingjie Xie, Shujun Liu, Xichuan Zhou, Xinzheng Zhang 0002
Appl. Intell.6
2018 MRI reconstruction via enhanced group sparsity and nonconvex regularization
Shujun Liu, Jianxin Cao, Hongqing Liu 0002, Xichuan Zhou, Zhengzhou Li
Neurocomputing4
2018 CS-MRI reconstruction via group-based eigenvalue decomposition and estimation
Shujun Liu, Jianxin Cao, Hongqing Liu 0002, Xiaoheng Tan, Xichuan Zhou
Neurocomputing6
2018 Group sparsity with orthogonal dictionary and nonconvex regularization for exact MRI reconstruction
Shujun Liu, Jianxin Cao, Hongqing Liu 0002, Xiaoheng Tan, Xichuan Zhou
Inf. Sci.5
2018 A Spatial-Temporal Method to Detect Global Influenza Epidemics Using Heterogeneous Data Collected from the Internet
abstract
The 2009 influenza pandemic teaches us how fast the influenza virus could spread globally within a short period of time. To address the challenge of timely global influenza surveillance, this paper presents a spatial-temporal method that incorporates heterogeneous data collected from the Internet to detect influenza epidemics in real time. Specifically, the influenza morbidity data, the influenza-related Google query data and news data, and the international air transportation data are integrated in a multivariate hidden Markov model, which is designed to describe the intrinsic temporal-geographical correlation of influenza transmission for surveillance purpose. Respective models are built for 106 countries and regions in the world. Despite that the WHO morbidity data are not always available for most countries, the proposed method achieves 90.26 to 97.10 percent accuracy on average for real-time detection of global influenza epidemics during the period from January 2005 to December 2015. Moreover, experiment shows that, the proposed method could even predict an influenza epidemic before it occurs with 89.20 percent accuracy on average. Timely international surveillance results may help the authorities to prevent and control the influenza disease at the early stage of a global influenza pandemic.
Xichuan Zhou, Fan Yang 0021, Qin Li 0006, Fang Tang, Shengdong Hu, Zhi Lin 0002, Lei Zhang 0038
IEEE ACM Trans. Comput. Biol. Bioinform.1
2018 DANoC: An Efficient Algorithm and Hardware Codesign of Deep Neural Networks on Chip
abstract
Deep neural networks (NNs) are the state-of-the-art models for understanding the content of images and videos. However, implementing deep NNs in embedded systems is a challenging task, e.g., a typical deep belief network could exhaust gigabytes of memory and result in bandwidth and computational bottlenecks. To address this challenge, this paper presents an algorithm and hardware codesign for efficient deep neural computation. A hardware-oriented deep learning algorithm, named the deep adaptive network, is proposed to explore the sparsity of neural connections. By adaptively removing the majority of neural connections and robustly representing the reserved connections using binary integers, the proposed algorithm could save up to 99.9% memory utility and computational resources without undermining classification accuracy. An efficient sparse-mapping-memory-based hardware architecture is proposed to fully take advantage of the algorithmic optimization. Different from traditional Von Neumann architecture, the deep-adaptive network on chip (DANoC) brings communication and computation in close proximity to avoid power-hungry parameter transfers between on-board memory and on-chip computational units. Experiments over different image classification benchmarks show that the DANoC system achieves competitively high accuracy and efficiency comparing with the state-of-the-art approaches.
Xichuan Zhou, Shengli Li 0007, Fang Tang, Shengdong Hu, Zhi Lin 0002, Lei Zhang 0038
IEEE Trans. Neural Networks Learn. Syst.1
2017 Deep Learning With Grouped Features for Spatial Spectral Classification of Hyperspectral Images
abstract
This letter presents a novel deep learning algorithm for feature extraction from the hyperspectral images. The proposed method takes advantage of the knowledge that the features of the spatial-spectral data naturally fall into an array of groups with respect to different spectral bands. Aiming to reduce the influence of redundant spectral bands adaptively using unlabeled hyperspectral data, we incorporate the group information in the training algorithm of the deep neural network via a regularized weight-decay process. Experiments over different benchmarks of hyperspectral images show that the proposed method provides competitive solution with the state-of-the-art approaches.
Xichuan Zhou, Shengli Li 0007, Fang Tang, Shengdong Hu, Shujun Liu
IEEE Geosci. Remote. Sens. Lett.1
2016 Global influenza surveillance with Laplacian multidimensional scaling
abstract
The Global Influenza Surveillance Network is crucial for monitoring epidemic risk in participating countries. However, at present, the network has notable gaps in the developing world, principally in Africa and Asia where laboratory capabilities are limited. Moreover, for the last few years, various influenza viruses have been continuously emerging in the resource-limited countries, making these surveillance gaps a more imminent challenge. We present a spatial-transmission model to estimate epidemic risks in the countries where only partial or even no surveillance data are available. Motivated by the observation that countries in the same influenza transmission zone divided by the World Health Organization had similar transmission patterns, we propose to estimate the influenza epidemic risk of an unmonitored country by incorporating the surveillance data reported by countries of the same transmission zone. Experiments show that the risk estimates are highly correlated with the actual influenza morbidity trends for African and Asian countries. The proposed method may provide the much-needed capability to detect, assess, and notify potential influenza epidemics to the developing world.
Xichuan Zhou, Fang Tang, Qin Li 0006, Shengdong Hu, Yunjian Jia
Frontiers Inf. Technol. Electron. Eng.1
2012 Towards the Optimal Discriminant Subspace
abstract
Dimensionality reduction is a common practice in many learning and Intelligence applications. However, most existing methods use the dimension of the target subspace as a parameter, making it hard to decide which subspace is optimal for classification. In this paper, we address the challenge of learning the optimal subspace for the nearest neighbor classification. We focus on labeled data and assume that the data for each class lie on respective sub-manifolds. To separate each sub-manifold, the labels of the data are used to learn the subspace where neighboring points of the same class keep close and those of different classes are disassociated. The sub-manifold separating method is first proposed as linear projection. For more complicated nonlinear situation, we generalize the algorithm using the kernel method. A group of experiments on data representation and classification are performed to evaluate he effectiveness of the proposed approaches.
Xichuan Zhou, Ping Gan
Web Intelligence1
2012 Largemargin classification for combating disguise attacks on spam filters
abstract
This paper addresses the challenge of large margin classification for spam filtering in the presence of an adversary who disguises the spam mails to avoid being detected. In practice, the adversary may strategically add good words indicative of a legitimate message or remove bad words indicative of spam. We assume that the adversary could afford to modify a spam message only to a certain extent, without damaging its utility for the spammer. Under this assumption, we present a large margin approach for classification of spam messages that may be disguised. The proposed classifier is formulated as a second-order cone programming optimization. We performed a group of experiments using the TREC 2006 Spam Corpus. Results showed that the performance of the standard support vector machine (SVM) degrades rapidly when more words are injected or removed by the adversary, while the proposed approach is more stable under the disguise attack.
Xichuan Zhou, Haibin Shen
J. Zhejiang Univ. Sci. C1
2011 Integrating outlier filtering in large margin training
abstract
Large margin classifiers such as support vector machines (SVM) have been applied successfully in various classification tasks. However, their performance may be significantly degraded in the presence of outliers. In this paper, we propose a robust SVM formulation which is shown to be less sensitive to outliers. The key idea is to employ an adaptively weighted hinge loss that explicitly incorporates outlier filtering in the SVM training, thus performing outlier filtering and classification simultaneously. The resulting robust SVM formulation is non-convex. We first relax it into a semi-definite programming which admits a global solution. To improve the efficiency, an iterative approach is developed. We have performed experiments using both synthetic and real-world data. Results show that the performance of the standard SVM degrades rapidly when more outliers are included, while the proposed robust SVM training is more stable in the presence of outliers.
Xichuan Zhou, Haibin Shen, Jieping Ye
J. Zhejiang Univ. Sci. C1
2010 Notifiable infectious disease surveillance with data collected by search engine
abstract
Notifiable infectious diseases are a major public health concern in China, causing about five million illnesses and twelve thousand deaths every year. Early detection of disease activity, when followed by a rapid response, can reduce both social and medical impact of the disease. We aim to improve early detection by monitoring health-seeking behavior and disease-related news over the Internet. Specifically, we counted unique search queries submitted to the Baidu search engine in 2008 that contained disease-related search terms. Meanwhile we counted the news articles aggregated by Baidu’s robot programs that contained disease-related keywords. We found that the search frequency data and the news count data both have distinct temporal association with disease activity. We adopted a linear model and used searches and news with 1–200-day lead time as explanatory variables to predict the number of infections and deaths attributable to four notifiable infectious diseases, i.e., scarlet fever, dysentery, AIDS, and tuberculosis. With the search frequency data and news count data, our approach can quantitatively estimate up-to-date epidemic trends 10–40 days ahead of the release of Chinese Centers for Disease Control and Prevention (Chinese CDC) reports. This approach may provide an additional tool for notifiable infectious disease surveillance.
Xichuan Zhou, Haibin Shen
J. Zhejiang Univ. Sci. C1