EDBT 2026 Demo / reviewers in the wild / expert
Haijun Liu 0001
dblp:40/2619-1
· DBLP profile ↗
50ranked-venue papers
10as first author
38since 2021 · last 2027
0000-0001-5782-4543ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 23 · 4 first-author · 17 since 2021Graphics, computer vision, multimedia, augmented reality and games · 20 · 6 first-author · 12 since 2021Applied, interdisciplinary, general and emerging computing · 11 · 1 first-author · 11 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2027 | ST-JSCC: Synergizing structural and textural dependencies for robust and efficient image transmission
Rulong He, Mingyang Wan, Haoming Luo, Xichuan Zhou, Haijun Liu 0001 |
Signal Process. | 6 |
| 2026 | Remote sensing optical image matching through neighborhood-aware global propagation in graph neural networks
Yanchun Liu, Gemine Vivone, Jing Nie 0001, Haijun Liu 0001, Xichuan Zhou, Lihui Chen 0002 |
Eng. Appl. Artif. Intell. | 4 |
| 2026 | Multiscale wavelet-based spatial-spectral compression network for hyperspectral image
Mingyang Wan, Aibin Peng, Xiangfei Shen, Rulong He, Lihui Chen 0002, Haijun Liu 0001, Xichuan Zhou |
Eng. Appl. Artif. Intell. | 8 |
| 2026 | PM-adapter: MoE based dynamic denoising fine-tuning for thermal infrared object detection
Haijun Liu 0001, Boya Wei, Jing Nie 0001, Suju Li, Xichuan Zhou |
Neurocomputing | 1 |
| 2026 | Sparse gain adaptation with dual-domain fusion network for multimodal object detection
Xichuan Zhou, Boya Wei, Cong Mao, Lihui Chen 0002, Haijun Liu 0001, Jin Xie 0005, Jing Nie 0001 |
Neurocomputing | 7 |
| 2026 | RA-PTQ: Reparameterization-Aware Post-Training Quantization for accurate vision transformers in low-bit scenarios
Rui Ding 0009, Sihuan Zhao, Jing Nie 0001, Lihui Chen 0002, Haijun Liu 0001, Xichuan Zhou |
Knowl. Based Syst. | 6 |
| 2026 | Multimodality Image Registration With Modality DistillationabstractMultimodal image registration aims to spatially align images from different modalities at the pixel level. However, due to the nonlinear relationship of radiation intensities caused by different imaging modalities, achieving high accuracy in multimodal image registration presents a significant challenge. Additionally, the presence of both global transformations (i.e., large-scale rigid affine transformations) and local distortions (i.e., small-scale nonrigid deformations) between paired images further complicates the registration process. This article addressed the challenge resulting from modality differences through modality distillation. Specifically, a teacher (i.e., a homomodal image registration model) is trained to guide the student (i.e., a multimodal image registration model). Besides, this article simultaneously aligned large-scale rigid and small-scale nonrigid deformations by predicting deformation flow from both global and local features, thereby achieving high-precision registration. Furthermore, this proposed method incorporated a deformation mask during training to mitigate the negative impact of black edges in the obtained registration results on model performance. Experimental results demonstrate that the proposed method delivers state-of-the-art registration accuracy across various multimodal datasets, with ablation studies confirming the effectiveness of each component. The codes will be available at https://github.com/2351056918/Multimodality-Image-Registration-with-Modailty-Distillation. Xichuan Zhou, Jicheng Zhao, Lihui Chen 0002, Gemine Vivone, Yanchun Liu, Jing Nie 0001, Haijun Liu 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 7 |
| 2026 | CAPilot: A High-Performance and High-Reliability Communication Middleware for Autonomous DrivingabstractWith the swift advancement of artificial intelligence technology, autonomous driving has increasingly emerged as a pivotal technology in the future of transportation. Real-time data exchange and processing across modules in autonomous driving systems necessitate efficient and reliable communication middleware. However, existing communication methods suffer from delay, congestion and packet loss when dealing with high-frequency and large-data-volume transmission tasks, significantly impairing system performance and security. To reduce communication latency and CPU overhead, a multi-mode adaptive high-performance and high-reliability communication middleware CAPilot is proposed. Firstly, a novel shared-memory communication architecture is proposed, comprising a Data Pool, an Event Notification Index Pool, and a Cycle Index Pool. The Data Pool employs a lock-free mechanism to avert deadlock and starvation issues, while addressing frame-skipping using a real-time maintenance and discriminative approach. Event-triggered and period-triggered data acquisition strategies proficiently circumvent data security concerns and performance limitations inherent in conventional shared memory connectivity. Then, to mitigate the overhead associated with dynamic broadcasts within the constrained embedded resources of the network, an adaptive communication scheme is proposed. This scheme incorporates a profile-based static communication encoding that automatically determines the optimal communication method based on the environments of the communicating entities. Finally, the intra-process pointer passing method is optimised by introducing a dual adaptive buffered ring queue, which facilitates bulk data retrieval without using locks. Experimental results show that CAPilot outperforms existing communication middlewares such as ROS2, CyberRT, and DDS in terms of communication latency, message throughput, message frame loss rate, and resource utilisation. These advancements suggest that CAPilot is well-suited for extensive deployment in diverse autonomous driving applications. Pinzhong Qin, Changquan Xue, Jing Nie 0001, Haijun Liu 0001, Xichuan Zhou |
ACM Trans. Internet Techn. | 5 |
| 2025 | EigenSR: Eigenimage-Bridged Pre-Trained RGB Learners for Single Hyperspectral Image Super-ResolutionabstractSingle hyperspectral image super-resolution (single-HSI-SR) aims to improve the resolution of a single input low-resolution HSI. Due to the bottleneck of data scarcity, the development of single-HSI-SR lags far behind that of RGB natural images. In recent years, research on RGB SR has shown that models pre-trained on large-scale benchmark datasets can greatly improve performance on unseen data, which may stand as a remedy for HSI. But how can we transfer the pre-trained RGB model to HSI, to overcome the data-scarcity bottleneck? Because of the significant difference in the channels between the pre-trained RGB model and the HSI, the model cannot focus on the correlation along the spectral dimension, thus limiting its ability to utilize on HSI. Inspired by the HSI spatial-spectral decoupling, we propose a new framework that first fine-tunes the pre-trained model with the spatial components (known as eigenimages), and then infers on unseen HSI using an iterative spectral regularization (ISR) to maintain the spectral correlation. The advantages of our method lie in: 1) we effectively inject the spatial texture processing capabilities of the pre-trained RGB model into HSI while keeping spectral fidelity, 2) learning in the spectral-decorrelated domain can improve the generalizability to spectral-agnostic data, and 3) our inference in the eigenimage domain naturally exploits the spectral low-rank property of HSI, thereby reducing the complexity. This work bridges the gap between pre-trained RGB models and HSI via eigenimages, addressing the issue of limited HSI training data, hence the name EigenSR. Extensive experiments show that EigenSR outperforms the state-of-the-art (SOTA) methods in both spatial and spectral metrics. Xi Su, Xiangfei Shen, Mingyang Wan, Jing Nie 0001, Lihui Chen 0002, Haijun Liu 0001, Xichuan Zhou |
AAAI | 6 |
| 2025 | Hybrid cross-modality fusion network for medical image segmentation with contrastive learning
Xichuan Zhou, Jing Nie 0001, Haijun Liu 0001, Fu Liang, Lihui Chen 0002, Jin Xie 0005 |
Eng. Appl. Artif. Intell. | 5 |
| 2025 | E-TransConvNet: An enhanced transformer and convolutional network for medical image segmentation from ultrasound and CT images
Chukwuemeka Clinton Atabansi, Jing Nie 0001, Jiachen Huang, Haijun Liu 0001, Jin Xie 0005, Xichuan Zhou |
Expert Syst. Appl. | 5 |
| 2025 | MambaSOD: Dual Mamba-driven cross-modal fusion network for RGB-D Salient Object Detection
Yue Zhan, Zhihong Zeng, Haijun Liu 0001, Xiaoheng Tan, Yinli Tian |
Neurocomputing | 3 |
| 2025 | Distribution-modulated binary neural network for image classification
Yingcheng Lin, Rui Ding 0009, Haijun Liu 0001, Xichuan Zhou |
Image Vis. Comput. | 4 |
| 2025 | Progressive fine-to-coarse reconstruction for accurate low-bit post-training quantization in vision transformers
Rui Ding 0009, Liang Yong, Sihuan Zhao, Jing Nie 0001, Lihui Chen 0002, Haijun Liu 0001, Xichuan Zhou |
Neural Networks | 6 |
| 2025 | Binary Neural Networks With Feature Information Retention for Efficient Image ClassificationabstractAlthough binary neural networks (BNNs) enjoy extreme compression ratios, there are significant accuracy gap compared with full-precision models. Previous works propose various strategies to reduce the information loss induced by the binarization process, improving the performance of binary neural networks to some extent. However, in this letter, we argue that few studies try to alleviate this problem from the structure perspective, resulting in inferior performance. To this end, we propose a novel Feature Information Retention Network named FIRNet, which incorporates an extra path to propagate the untouched informative feature maps. Specifically, the FIRNet splits the input feature maps into two groups, one of which is fed into the normal layers and another kept untouched for information retention. Then we utilize the concatenation, shuffle and pooling operations to process these features with 64× memory saving. Finally, with only a 1.7% complexity increase, a FIR fusion layer is proposed to aggregate the features from two branches. Experimental results demonstrate that our proposed method achieves 1.0% Top-1 accuracy improvement over the baseline model and outperforms other state-of-the-art BNNs on the ImageNet dataset. Rui Ding 0009, Haijun Liu 0001, Xichuan Zhou |
IEEE Signal Process. Lett. | 3 |
| 2025 | Bi-SSFormer: An Ultralightweight Binary Spectral-Spatial Transformer for Hyperspectral Image Classification
Rui Ding 0009, Yanchun Liu, Baoliang Wang, Lihui Chen 0002, Haijun Liu 0001, Gemine Vivone, Xichuan Zhou |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2025 | MIT-SAM: Medical Image-Text SAM With Mutually Enhanced Heterogeneous Features Fusion for Medical Image SegmentationabstractIn recent times, leveraging lesion text as supplementary data to enhance the performance of medical image segmentation models has garnered attention. Previous approaches only used attention mechanisms to integrate image and text features, while not effectively utilizing the highly condensed textual semantic information in improving the fused features, resulting in inaccurate lesion segmentation. This paper introduces a novel approach, the Medical Image-Text Segment Anything Model (MIT-SAM), for text-assisted medical image segmentation. Specifically, we introduce the SAM-enhanced image encoder and a Bert-based text encoder to extract heterogeneous features. To better leverage the highly condensed textual semantic information for heterogeneous feature fusion, such as crucial details like position and quantity, we propose the image-text interactive fusion (ITIF) block and self-supervised text reconstruction (SSTR) method. The ITIF block facilitates the mutual enhancement of homogeneous information among heterogeneous features and the SSTR method empowers the model to capture crucial details concerning lesion text, including location, quantity, and other key aspects. Experimental results demonstrate that our proposed model achieves state-of-the-art performance on the QaTa-COV19 and MosMedData+ datasets. Xichuan Zhou, Lingfeng Yan, Rui Ding 0009, Chukwuemeka Clinton Atabansi, Jing Nie 0001, Lihui Chen 0002, Haijun Liu 0001 |
IEEE J. Biomed. Health Informatics | 8 |
| 2024 | PSE-Net: Channel pruning for Convolutional Neural Networks with parallel-subnets estimator
Shiguang Wang, Tao Xie 0010, Haijun Liu 0001, Xingcheng Zhang, Jian Cheng 0003 |
Neural Networks | 3 |
| 2024 | GOENet: Group Operations Enhanced Binary Neural Network for Efficient Image ClassificationabstractThere exists an innegligible performance gap between the binary neural networks and their full-precision counterparts, which prevents their deployment on real-world applications. Recently, plenty of researchers strive to solve this problem by incorporating more binary subnets, improving the representational power with acceptable complexity increase. However, current methods make structure design and parallel acceleration sophisticated and non-trivial and may degenerate the representational power due to the isomorphic subnets. Besides, the final feature fusion method is sub-optimal. In this letter, we propose a simple yet effective binary neural network named GOENet enhanced by group operations. The multiple parallel binary subnets use group-wise binary thresholds to improve feature diversity and are merged into the group convolutional layers. The inter-subnet connections are implemented with the parameter-free group shuffle operations, improving the model parallelism and performance. We also propose the group feature summation module as a better fusion method for its efficiency and effectiveness. Experimental results on ImageNet show that our proposed method outperforms the SOTA GroupNet by 1.7% Top-1 accuracy with 0.84× and 0.89× saving in computation and memory complexity. Rui Ding 0009, Haijun Liu 0001, Xichuan Zhou |
IEEE Signal Process. Lett. | 3 |
| 2024 | AirSOD: A Lightweight Network for RGB-D Salient Object DetectionabstractSalient object detection (SOD) aims to identify the most prominent regions in images. However, the large model sizes, high computational costs, and slow inference speeds of existing RGB-D SOD models have hindered their deployment on real-world embedded devices. To address this issue, we propose a novel method named AirSOD, which is committed to lightweight RGB-D SOD. Specifically, we first design a hybrid feature extraction network, which includes the first three stages of MobileNetV2 and our Parallel Attention-Shift convolution (PAS) module. Using the novel PAS module enables capturing both long-range dependencies and local information to enhance the representation learning while significantly reducing the number of parameters and computational complexity. Secondly, we propose a Multi-level and Multi-modal feature Fusion (MMF) module to facilitate feature fusion, and a Multi-path enhancement for Feature Refinement (MFR) decoder for feature integration. The proposed method significantly reduces the model size by 63%, decreases the computational complexity by 43%, and improves the inference speed by 43% compared with the cutting-edge model (MobileSal). We test our AirSOD on six widely-used RGB-D SOD datasets. Extensive experimental results demonstrate that our method obtains satisfactory performance. The source codes will be made available. Zhihong Zeng, Haijun Liu 0001, Fenglei Chen, Xiaoheng Tan |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2024 | MSNet: Self-Supervised Multiscale Network With Enhanced Separation Training for Hyperspectral Anomaly DetectionabstractHyperspectral anomaly detection (HAD) has attracted increasing attention due to its economical and efficient applications. The main challenge lies in the data-starved problem of hyperspectral images (HSIs) and the costliness of manual annotation, making it heavily reliant on the model’s adaptability and robustness to unseen scenes under limited samples. Self-supervised learning offers a solution to this urgency via mining meaningful representations from the data itself. One promising paradigm is leveraging untrained neural networks to reconstruct the background component for revealing anomalous information. Its capability stems from the network architecture and the training process rather than learning from expensive and strongly domain-dependent data, which is naturally applicable to HAD. In this article, to handle the urgent requirement for self-supervised learning in HAD, we propose a multiscale network (termed MSNet) that detects anomalies with enhanced separation training. The network architecture consists of several multiscale convolutional encoder-decoder (CED) layers, considering the spatial characteristics of the anomalies. To suppress the anomalies during background reconstruction, we adopt a new separation training strategy by introducing a soft separator for better practicality on larger datasets. Extensive experiments conducted on five commonly used datasets and the HAD100 dataset, demonstrate the superiority of our method over its counterparts. Our code is available athttps://github.com/enter-i-username/MSNet. Haijun Liu 0001, Xi Su, Xiangfei Shen, Xichuan Zhou |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2024 | MeSAM: Multiscale Enhanced Segment Anything Model for Optical Remote Sensing ImagesabstractSegment anything model (SAM) has been widely applied to various downstream tasks for its excellent performance and generalization capability. However, SAM exhibits three limitations related to remote sensing semantic segmentation task: 1) the image encoders excessively lose high-frequency information, such as object boundaries and textures, resulting in rough segmentation masks; 2) due to being trained on natural images, SAM faces difficulty in accurately recognizing objects with large-scale variations and uneven distribution in remote sensing images; 3) the output tokens used for mask prediction are trained on natural images and not applicable to remote sensing image segmentation. In this paper, we explore an efficient paradigm for applying SAM to the semantic segmentation of remote sensing images. Furthermore, we propose MeSAM, a new SAM fine-tuning method more suitable for remote sensing images to adapt it to semantic segmentation tasks. Our method first introduces an inception mixer into the image encoder to effectively preserve high-frequency features. Secondly, by designing a mask decoder with remote-sensing correction and incorporating multiscale connections, we make up the difference in SAM from natural images to remote sensing images. Experimental results demonstrated that our method significantly improves the segmentation accuracy of SAM for remote sensing images, outperforming some state-of-the-art methods. The code will be available at https://github.com/Magic-lem/MeSAM. Xichuan Zhou, Fu Liang, Lihui Chen 0002, Haijun Liu 0001, Gemine Vivone, Jocelyn Chanussot |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2024 | BTC-Net: Efficient Bit-Level Tensor Data Compression Network for Hyperspectral ImageabstractNow it is still a challenge to compress high-throughput hyperspectral tensor image data on lightweight air-carried/spaceborne remote sensing systems, primarily due to insufficient computational resources and limited transmission bandwidth. To address this challenge, we propose a bit-level tensor data compression network (BTC-Net) that provides higher compression performance by leveraging a data-driven lightweight quantized neural encoder with two-stage bit compression. The BTC-Net achieves semantic near-lossless high reconstruction quality at low compression bit rates thanks to its optimized decoder, which uses a channel-wise attention-based enhancement module to recover hyperspectral tensor data. Experimental results on different hyperspectral datasets show that the BTC-Net could achieve an extremely low compression bit rate of fewer than 0.04 bits per pixel per band (bpppb) with state-of-the-art reconstruction performances. The demo of BTC-Net will be publicly available online at: https://github.com/zx20173646/BTCNet. Xichuan Zhou, Xuan Zou, Xiangfei Shen, Wenjia Wei, Haijun Liu 0001 |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2023 | MDL-NAS: A Joint Multi-domain Learning Framework for Vision TransformerabstractIn this work, we introduce MDL-NAS, a unified frame-work that integrates multiple vision tasks into a manageable supernet and optimizes these tasks collectively under diverse dataset domains. MDL-NAS is storage-efficient since multiple models with a majority of shared parameters can be deposited into a single one. Technically, MDL-NAS constructs a coarse-to-fine search space, where the coarse search space offers various optimal architectures for different tasks while the fine search space provides fine-grained parameter sharing to tackle the inherent obstacles of multi-domain learning. In the fine search space, we suggest two parameter sharing policies, i.e., sequential sharing policy and mask sharing policy. Compared with previous works, such two sharing policies allow for the partial sharing and non-sharing of parameters at each layer of the network, hence attaining real fine-grained parameter sharing. Finally, we present a joint-subnet search algorithm that finds the optimal architecture and sharing parameters for each task within total resource constraints, challenging the traditional practice that downstream vision tasks are typically equipped with backbone networks designed for image classification. Experimentally, we demonstrate that MDL-NAS families fitted with non-hierarchical or hierarchical transformers deliver competitive performance for all tasks compared with state-of-the-art methods while maintaining efficient storage deployment and computation. We also demonstrate that MDL-NAS allows incremental learning and evades catastrophic forgetting when generalizing to a new task. Shiguang Wang, Tao Xie 0010, Jian Cheng 0003, Xingcheng Zhang, Haijun Liu 0001 |
CVPR | 5 |
| 2023 | InfraNet: Accurate forehead temperature measurement framework for people in the wild with monocular thermal infrared camera
Xichuan Zhou, Dongshan Lei, Chunqiao Long, Jing Nie 0001, Haijun Liu 0001 |
Neural Networks | 5 |
| 2023 | Efficient Hyperspectral Sparse Regression Unmixing With MultilayersabstractThe sparse regression method is known for its ability to unmix hyperspectral data, but it can be computationally expensive and accurately insufficient due to the large scale and high coherence of the spectral library. To address this issue, a new approach called layered sparse regression unmixing (termed LSU) has been proposed in this paper. This method involves breaking down the sparse unmixing process into multilayers, each of which interactively learns a row-sparsity-promoting abundance matrix and fine-tunes active library atoms based on measured activeness. By doing so, LSU outputs both a learned abundance matrix and an optimal library that can best model each mixed pixel in the scene. The proposed LSU can be efficiently solved by the alternating direction method of the multipliers framework. Experimental results obtained from simulated and real hyperspectral images demonstrate the effectiveness of LSU. The demo of the proposed LSU will be publicly available at https://github.com/XiangfeiShen/Layered_Sparse_Regression_Unmixing. Xiangfei Shen, Lihui Chen 0002, Haijun Liu 0001, Xi Su, Wenjia Wei, Xichuan Zhou |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2023 | Matrix Factorization With Framelet and Saliency Priors for Hyperspectral Anomaly DetectionabstractHyperspectral anomaly detection aims to separate sparse anomalies from low-rank background components. A variety of detectors have been proposed to identify anomalies, but most of them tend to emphasize characterizing backgrounds with multiple types of prior knowledge and limited information on anomaly components. To tackle these issues, this article simultaneously focuses on two components and proposes a matrix factorization method with framelet and saliency priors to handle the anomaly detection problem. We first employ a framelet to characterize nonnegative background representation coefficients, as they can jointly maintain sparsity and piecewise smoothness after framelet decomposition. We then exploit saliency prior knowledge to measure each pixel’s potential to be an anomaly. Finally, we incorporate the pure pixel index (PPI) with Reed-Xiaoli’s (RX) method to possess representative dictionary atoms. We solve the optimization problem using a block successive upper-bound minimization (BSUM) framework with guaranteed convergence. Experiments conducted on benchmark hyperspectral datasets demonstrate that the proposed method outperforms some state-of-the-art anomaly detection methods. Xiangfei Shen, Haijun Liu 0001, Jing Nie 0001, Xichuan Zhou |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2023 | Cellular Binary Neural Network for Accurate Image Classification and Semantic SegmentationabstractThis paper presents the Cellular Binary Neural Network (CBNN), which is an efficient deep neural network with binary weights and activations. To address the challenge of performance drop caused by low-precision representation, the CBNN adopts multiple subnets which are connected via learnable global lateral paths. The introduced lateral connections are assumed to be sparse and grouped with respect to different source layers. The inter-network lateral connections and inner-network parameters are simultaneously optimized by the distributional loss, classification loss and the group sparse regularization term. Experiments on the CIFAR-10 and ImageNet datasets showed that, by incorporating optimized group-sparse lateral paths, the CBNN outperformed many state-of-the-art binary neural networks in terms of classification accuracy. Besides, to verify the generalization of the proposed binary model, we extended the CBNN on semantic segmentation task. CBNN takes advantage of the multiple subnets to derive the more informative feature maps which are computed by the parallel aggregation in the last convolution block. Experiments on PASCAL VOC segmentation dataset demonstrated that, under the same segmentation settings, the proposed method achieved the superior performance over other compared networks and even the full-precision counterpart. Xichuan Zhou, Rui Ding 0009, Wenjia Wei, Haijun Liu 0001 |
IEEE Trans. Multim. | 5 |
| 2022 | Superpixel-Guided Local Sparsity Prior for Hyperspectral Sparse Regression UnmixingabstractSparse regression relaxes the difficulties of blind unmixing of hyperspectral data thanks to the spectral library. Many investigations, however, attach importance to global priors such as sparsity and low-rankness. This letter proposes a local-global-based sparse regression unmixing method, called LGSU, by introducing a local sparsity regularization to help boost the unmixing performance that only considers global sparsity. The proposed LGSU first uses a superpixel-based technique to yield a set of homogeneous superpixels for guiding local sparse regularization purposes. LGSU then considers a traditional ℓ1regularization to enhance global sparsity. Coupling with local and global sparsity constraints, the proposed LGSU can effectively estimate the abundance of a given image via the alternating direction method of multipliers. Experimental results obtained from synthetic and real hyperspectral images demonstrate the effectiveness of the proposed algorithm. Xiangfei Shen, Haijun Liu 0001, Xinzheng Zhang 0002, Xichuan Zhou |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2022 | Semantic Segmentation for High-Resolution Remote-Sensing Images via Dynamic Graph Context ReasoningabstractSemantic segmentation for high-resolution remote-sensing (HRRS) images is one of the most challenging tasks in remote-sensing images understanding. Capturing long-range dependencies in feature representations is crucial for semantic segmentation. Recent graph-based global reasoning networks (GloRe) focus on modeling the global contextual relationship between latent nodes based on fully connected graph in interaction space. However, such a dense operation is susceptible to redundant features. Most importantly, it treats each node equally, ignoring the contextual relationship between nodes in graphs. In this work, we propose to explore more effective contextual representations in semantic segmentation by introducing dynamic graph contextual reasoning module overGloRe, dubbed DGCR. It incorporates local semantic information that represents the relationships between nodes to perform long-range contextual reasoning. More specifically, to provide effectively and flexible reasoning in graph-based reasoning approaches, we construct$k$-nearest neighbor (KNN) graphs rather than fully connected graphs using only the$k$closest nodes depends on pairwise semantic distance. Extensive experiments on the International Society for Photogrammetry and Remote Sensing (ISPRS) Vaihingen and Potsdam datasets demonstrate the effectiveness and superiority of our proposed DGCR module over other state-of-the-art methods. Yanzhou Su, Jian Cheng 0003, Wen Wang 0012, Haiwei Bai, Haijun Liu 0001 |
IEEE Geosci. Remote. Sens. Lett. | 5 |
| 2022 | A Fourier-Based Semantic Augmentation for Visible-Thermal Person Re-IdentificationabstractThis letter introduces a novel Fourier-based data augmentation strategy for visible-thermal person re-identification (VT-ReID). Different from some existing methods which are proposed from the perspective of network structure and loss functions, our method aims to fully consider the semantic information from the perspective of data preprocessing. The main hypothesis is that the phase component in the Fourier domain contains high-level semantic information and the amplitude component contains low-level modality awareness information. In order to make the model pay more attention to semantic information learning, we design a simple but effective Fourier-based semantic augmentation (FSA) module, which can be inserted seamlessly into any existing models. Extensive experiments on RegDB and SYSU-MM01 datasets have shown that our proposed method can improve the VT-ReID performance significantly and achieve state-of-the-art performance. Xiaoheng Tan, Yanxia Chai, Fenglei Chen, Haijun Liu 0001 |
IEEE Signal Process. Lett. | 4 |
| 2022 | LC-BiDet: Laterally Connected Binary Detector With Efficient Image ProcessingabstractRecently, binary neural networks have received increasing interest in object detection community, since compared to the full-precision detection networks, they can reduce the memory and computation requirements significantly due to the efficient XNOR and BITCOUNT operations introduced by binarization. However, the final detection performance of binary detectors always suffers from a severe degradation compared with the full-precision counterparts. In this letter, to achieve a better trade-off between the inference efficiency and the detection performance, we propose to learn a laterally connected binary detector based on the multiple parallel binary subnets with a group sparse regularization term. Firstly, we simply adopt the parallel structure to improve the representation capacity of binary detectors. Secondly, we introduce the dense lateral connections between binary subnets in each convolutional layer to increase the information flows. Finally, to reduce the redundancy of the dense connections, we take the L21 norm as the regularization term to optimize the lateral connections in an end-to-end manner. Experimental results on PASCAL VOC and COCO datasets show that our proposed laterally connected binary detector could outperform the other state-ofthe- art binary detectors. Xichuan Zhou, Rui Ding 0009, Haijun Liu 0001 |
IEEE Signal Process. Lett. | 3 |
| 2022 | A Novel Multidimensional Domain Deep Learning Network for SAR Ship DetectionabstractSince only the spatial feature information of ship target is utilized, the current deep learning-based synthetic aperture radar (SAR) ship detection approaches cannot achieve a satisfactory performance, especially in the case of multiscale or rotations, and the complex background. To overcome these issues, a novel multidimensional domain deep learning network for SAR ship detection is developed in this work to exploit the spatial and frequency-domain complementary features. The proposed method consists of the following main three steps. First, to learn hierarchical spatial features, the feature pyramid network (FPN) is adopted to produce ship target spatial multiscale characteristics with a top-down structure. Second, with a polar Fourier transform, the rotation-invariant features of SAR ship targets are obtained in the frequency domain. After that, a novel spatial-frequency characteristics fusion network is then presented, which seeks to learn more compact feature representations across different domains by updating the parameters of sub-networks interactively. The detection results are obtained due to utilizing the multidimensional domain information, and we evaluate the effectiveness of the proposed method using the existing SAR ship detection data set (SSDD). The results of the proposed method outperform other convolutional neural network (CNN)-based algorithms, especially for multiscale and rotation ship targets under complex backgrounds. Dong Li 0007, Quanhuan Liang, Hongqing Liu 0001, Haijun Liu 0001, Guisheng Liao |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2022 | Toward Weak Signal Analysis in Hyperspectral Data: An Efficient Unmixing PerspectiveabstractMany unmixing methods hold the assumption that endmembers correspond to major land-covers, but not true for some unmixing tasks where observed minor object signals corresponding to some special types of endmembers are relatively weak. When there exist weak signals that have low intensity potentially caused by subtle mixing abundance fractions regarding the endmembers of minor objects, the traditional unmixing techniques may fail. This paper pioneers weak signal scenarios in hyperspectral unmixing using an efficient method called HyperWeak. Specifically, HyperWeak involves a sparse nonnegative matrix factorization model that contains two main parts, where the unsupervised part estimates the endmember and abundance matrices, and the supervised part ensures the minimal degradation of prior knowledge. To enhance the robustness of the HyperWeak model, this paper considers a reweighted sparsity constraint to boost the sparseness of the abundance matrix. For effectively solving optimization problems, Nesterov’s optimal gradient method is used in this paper. Experiments conducted on synthetic and real hyperspectral images indicate that HyperWeak can improve the unmixing performances of hyperspectral data in weak signal situations. Xiangfei Shen, Haijun Liu 0001, Fangyuan Ge, Xichuan Zhou |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2021 | Optimizing Information Theory Based Bitwise Bottlenecks for Efficient Mixed-Precision Activation QuantizationabstractRecent researches on information theory shed new light on the continuous attempts to open the black box of neural signal encoding. Inspired by the problem of lossy signal compression for wireless communication, this paper presents a Bitwise Bottleneck approach for quantizing and encoding neural network activations. Based on the rate-distortion theory, the Bitwise Bottleneck attempts to determine the most significant bits in activation representation by assigning and approximating the sparse coefficients associated with different bits. Given the constraint of a limited average code rate, the bottleneck minimizes the distortion for optimal activation quantization in a flexible layer-by-layer manner. Experiments over ImageNet and other datasets show that, by minimizing the quantization distortion of each layer, the neural network with bottlenecks achieves the state-of-the-art accuracy with low-precision activation. Meanwhile, by reducing the code rate, the proposed method can improve the memory and computational efficiency by over six times compared with the deep neural network with standard single-precision representation. The source code is available on GitHub: https://github.com/CQUlearningsystemgroup/BitwiseBottleneck. Xichuan Zhou, Cong Shi 0003, Haijun Liu 0001 |
AAAI | 4 |
| 2021 | Learning to Binarize Convolutional Neural Networks with Adaptive Neural EncoderabstractThe high computational complexity and memory consumption of the deep Convolutional Neural Networks (CNNs) restrict their deployability in resource-limited embedded devices. To address this challenge, emerging solutions are proposed for neural network quantization and compression. Among them, Binary Neural Networks (BNNs) show their potential in reducing computational and memory complexity; however, they suffer from considerable performance degradation. One of the major causes is their non-differentiable discrete quantization implemented using a fixed sign function, which leads to output distribution distortion. In this paper, instead of using the fixed and naive sign function, we propose a novel adaptive Neural Encoder (NE), which learns to quantize the full-precision weights as binary values. Inspired by the research of neural network distillation, a distribution loss is introduced as a regularizer to minimize the Kullback-Leibler divergence between the outputs of the full-precision model and the encoded binary model. With an end-to-end backpropagation training process, the adaptive neural encoder, along with the binary convolutional neural network, could reach convergence iteratively. Comprehensive experiments with different network structures and datasets show that the proposed method can improve the performance of the baselines and also outperform many state-of-the-art approaches. The source code of the proposed method is publicly available at https://github.com/CQUlearningsystemgroup/LearningToBinarize. Fangyuan Ge, Rui Ding 0009, Haijun Liu 0001, Xichuan Zhou |
IJCNN | 4 |
| 2021 | Strong but Simple Baseline With Dual-Granularity Triplet Loss for Visible-Thermal Person Re-IdentificationabstractThis letter presents a conceptually simple and effective dual-granularity triplet loss for visible-thermal person re-identification (VT-ReID). Generally, ReID models are always trained with the sample-based triplet loss and identification loss from the fine granularity level. Further, center-based loss could be introduced to encourage the intra-class compactness and inter-class discrimination from the coarse granularity level. Our proposed dual-granularity triplet loss well organizes the sample-based triplet loss and center-based triplet loss in a hierarchical fine to coarse granularity manner, just with some simple configurations of typical operations, such as pooling and batch normalization. Experiments on RegDB and SYSU-MM01 datasets show that with only the global features our dual-granularity triplet loss can improve the VT-ReID performance by a significant margin. It can be a strong VT-ReID baseline to boost future research with high quality. Haijun Liu 0001, Yanxia Chai, Xiaoheng Tan, Dong Li 0007, Xichuan Zhou |
IEEE Signal Process. Lett. | 1 |
| 2021 | Parameter Sharing Exploration and Hetero-Center Triplet Loss for Visible-Thermal Person Re-IdentificationabstractThis paper focuses on the visible-thermal cross-modality person re-identification (VT Re-ID) task, whose goal is to match person images between the daytime visible modality and the nighttime thermal modality. The two-stream network is usually adopted to address the cross-modality discrepancy, the most challenging problem for VT Re-ID, by learning the multi-modality person features. In this paper, we explore how many parameters a two-stream network should share, which is still not well investigated in the existing literature. By splitting the ResNet50 model to construct the modality-specific feature extraction network and modality-sharing feature embedding network, we experimentally demonstrate the effect of parameter sharing of two-stream network for VT Re-ID. Moreover, in the framework of part-level person feature learning, we propose the hetero-center triplet loss to relax the strict constraint of traditional triplet loss by replacing the comparison of theanchor to all the other samplesby theanchor center to all the other centers. With extremely simple means, the proposed method can significantly improve the VT Re-ID performance. The experimental results on two datasets show that our proposed method distinctly outperforms the state-of-the-art methods by large margins, especially on the RegDB dataset achieving superior performance, rank1/mAP/mINP 91.05%/83.28%/68.84%. It can be a new baseline for VT Re-ID, with a simple but effective strategy. Haijun Liu 0001, Xiaoheng Tan, Xichuan Zhou |
IEEE Trans. Multim. | 1 |
| 2020 | Enhancing the discriminative feature learning for visible-thermal cross-modality person re-identification
Haijun Liu 0001, Jian Cheng 0003, Wen Wang 0012, Yanzhou Su, Haiwei Bai |
Neurocomputing | 1 |
| 2020 | Multi-Scale Based Context-Aware Net for Action DetectionabstractWe address the problem of action detection in continuous untrimmed video streams, based on the two-stage framework: one stage for action proposals generation and the other for proposals classification and refinement. The context features inside and outside a candidate region (proposal) are critical for classification in action detection. Therefore, effective integration of these features with different scales has become a fundamental problem. We contend that different action instances and candidate proposals may need different context features. To address this issue, we present a novel multiple scales based context-aware net (MSCA-Net) to effectively classify the action proposals for action detection in this paper. For each candidate action proposal, MSCA-Net takes its multiple regions with different temporal scales as input and then generates suitable context features. Based on the “candidate-control” mechanism of LSTM, the proposed MSCA-Net specially adopts the two-branch structure: Branch1 generates multi-scale context features for each candidate proposal, whereas Branch2 utilizes the context-aware gate function to control the message passing. Extensive experiments on THUMOS’14, Charades daily and ActivityNet action detection datasets, demonstrate the effectiveness of the designed structure and show how these context features influence the detection results. Haijun Liu 0001, Shiguang Wang, Wen Wang 0012, Jian Cheng 0003 |
IEEE Trans. Multim. | 1 |
| 2019 | Gallery based k-reciprocal-like re-ranking for heavy cross-camera discrepancy in person re-identification
Haijun Liu 0001, Jian Cheng 0003 |
Neurocomputing | 1 |
| 2018 | Temporal Action Detection by Joint Identification-VerificationabstractTemporal action detection aims at not only recognizing action category but also detecting start time and end time for each action instance in an untrimmed video. The key challenge of this task is to accurately classify the actions and determine the temporal boundaries of each action instance. In temporal action detection benchmark: THUMOS 2014, large variations exist in the same action category while many similarities exist in different action categories, which always limit the performance of temporal action detection. To address this problem, we propose to use joint Identification-Verification network to reduce the intra-action variations and enlarge inter-action differences. The joint Identification-Verification network is a siamese network based on 3D ConvNets, which can simultaneously predict the action categories and the similarity scores for the input pairs of video proposal segments. Extensive experimental results on the challenging THUMOS 2014 dataset demonstrate the effectiveness of our proposed method compared to the existing state-of-art methods for temporal action detection in untrimmed videos. We further demonstrate that our model is a general framework by evaluating our approach on Charades dataset. Wen Wang 0012, Haijun Liu 0001, Shiguang Wang, Jian Cheng 0003 |
ICPR | 3 |
| 2018 | Visualizing deep neural network by alternately image blurring and deblurring
Feng Wang 0015, Haijun Liu 0001, Jian Cheng 0003 |
Neural Networks | 2 |
| 2018 | Additive Margin Softmax for Face VerificationabstractIn this letter, we propose a conceptually simple and intuitive learning objective function, i.e., additive margin softmax, for face verification. In general, face verification tasks can be viewed as metric learning problems, even though lots of face verification models are trained in classification schemes. It is possible when a large-margin strategy is introduced into the classification model to encourage intraclass variance minimization. As one alternative, angular softmax has been proposed to incorporate the margin. In this letter, we introduce another kind of margin to the softmax loss function, which is more intuitive and interpretable. Experiments on LFW and MegaFace show that our algorithm performs better when the evaluation criteria are designed for very low false alarm rate. Feng Wang 0015, Jian Cheng 0003, Weiyang Liu, Haijun Liu 0001 |
IEEE Signal Process. Lett. | 4 |
| 2018 | Sequential Subspace Clustering via Temporal Smoothness for Sequential Data SegmentationabstractThis paper develops a novel sequential subspace clustering method for sequential data. Inspired by the state-of-the-art methods, ordered subspace clustering, and temporal subspace clustering, we design a novel local temporal regularization term based on the concept of temporal predictability. Through minimizing the short-term variance on historical data, it can recover the temporal smoothness relationships in sequential data. Moreover, we claim that the local temporal regularization is more important than the global structural regularization for a specific task, such as sequential subspace clustering, which leads to a concise minimization objective function. To solve the bi-convex objective function, a simple and efficient optimization algorithm based on the alternate convex search method is devised to jointly learn the coding matrix and the dictionary. Furthermore, five baseline methods are also devised for comparison with our proposed method from different aspects. Extensive experimental results and comparisons with the state-of-the-art methods on three data sets demonstrate the effectiveness of the proposed temporal smoothness sequential subspace clustering method for sequential data. Haijun Liu 0001, Jian Cheng 0003, Feng Wang 0015 |
IEEE Trans. Image Process. | 1 |
| 2018 | Pedestrian Detection via Body Part Semantic and Contextual Information With DNNabstractPedestrian detection has achieved great improve-ments in recent years, while complex occlusion handling and high-accurate localization are still the most important problems. To take advantage of the body part semantic information and the contextual information for pedestrian detection, we propose the part and context network (PCN) in this paper. A PCN is composed of three branches: the basic branch; the part branch; and the context branch. It specially utilizes two branches to detect the pedestrians through the body part semantic information and the contextual information, respectively. In the part branch, the semantic information of body parts can communicate with each other via long short-term memory (LSTM). In the context branch, we adopt a local competition mechanism (maxout) for adaptive context scale selection. By combining the outputs of all branches, we develop a strong complementary pedestrian detector with a lower miss rate and higher localization accuracy, especially for the occlusion pedestrian. The combination of the body part semantic information and the contextual information in pedestrian detection is fully explored in this paper. Comprehensive evaluations on three challenging pedestrian detection datasets (i.e., Caltech, INRIA and KITTI) well demonstrate the effectiveness of our proposed PCN. Code for PCN is publicly available on GitHub https://github.com/sunnyxiaohu/pcn_pedestrian. Shiguang Wang, Jian Cheng 0003, Haijun Liu 0001, Feng Wang 0015, Hui Zhou 0005 |
IEEE Trans. Multim. | 3 |
| 2017 | Sequential Subspace Clustering via Temporal SmoothnessabstractThis paper develops a novel sequential subspace clustering method for sequential data. Inspired by state-of-the-art methods ordered subspace clustering (OSC) and temporal subspace clustering (TSC), we design a novel local temporal regularization term based on the concept of temporal predictability, which is measured by short-term variance against long-term variance, to recover the temporal smoothness relationships in sequential data. To solve the bi-convex objective function, a simple and efficient optimization algorithm based on the alternate convex search (ACS) method is devised to jointly learn the codings matrix and dictionary. Extensive experimental results and comparisons with state-of-the-art methods on gesture and face datasets demonstrate the effectiveness of the proposed temporal smoothness sequential subspace clustering method for sequential data. Haijun Liu 0001, Jian Cheng 0003, Feng Wang 0015 |
FG | 1 |
| 2017 | Kinship verification based on status-aware projection learningabstractKinship verification for parent-child is considered to be an asymmetric metric process, in which parents and children are associated with different status where the parents are priorly known to be significantly older than the children. To address the asymmetric metric learning, a status-aware projection learning (SaPL) method is proposed for facial image-based kinship verification, especially for the parent-child kinship. SaPL learns two status-specific projections to capture the significant appearance commonality between parents and children, respectively. Each status-specific projection consists of two components: a common component shared by the two status projections and a status-specific component. SaPL generally outperforms the one Mahalanobis distance metric. Extensive experimental results and comparisons with state-of-the-art approaches demonstrate the effectiveness of the proposed SaPL for kinship verification. Haijun Liu 0001, Jian Cheng 0003, Feng Wang 0015 |
ICIP | 1 |
| 2015 | Silhouette Analysis for Human Action Recognition Based on Supervised Temporal t-SNE and Incremental LearningabstractThis paper develops a human action recognition method for human silhouette sequences based on supervised temporal t-stochastic neighbor embedding (ST-tSNE) and incremental learning. Inspired by the SNE and its variants, ST-tSNE is proposed to learn the underlying relationship between action frames in a manifold, where the class label information and temporal information are introduced to well represent those frames from the same action class. As to the incremental learning, an important step for action recognition, we introduce three methods to perform the low-dimensional embedding of new data. Two of them are motivated by local methods, locally linear embedding and locality preserving projection. Those two techniques are proposed to learn explicit linear representations following the local neighbor relationship, and their effectiveness is investigated for preserving the intrinsic action structure. The rest one is based on manifold-oriented stochastic neighbor projection to find a linear projection from high-dimensional to low-dimensional space capturing the underlying pattern manifold. Extensive experimental results and comparisons with the state-of-the-art methods demonstrate the effectiveness and robustness of the proposed ST-tSNE and incremental learning methods in the human action silhouette analysis. Jian Cheng 0003, Haijun Liu 0001, Feng Wang 0015, Hongsheng Li 0001, Ce Zhu |
IEEE Trans. Image Process. | 2 |
| 2014 | Silhouette analysis for human action recognition based on maximum spatio-temporal dissimilarity embedding
Jian Cheng 0003, Haijun Liu 0001, Hongsheng Li 0001 |
Mach. Vis. Appl. | 2 |