EDBT 2026 Demo / reviewers in the wild / expert
Zhonghao Chen
dblp:288/8236
· DBLP profile ↗
26ranked-venue papers
8as first author
26since 2021 · last 2026
0000-0002-3524-2742ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 10 · 3 first-author · 10 since 2021Systems, architecture and hardware · 8 · 2 first-author · 8 since 2021Artificial intelligence and machine learning · 4 · 2 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 4 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Causally-Grounded Dual-Path Attention Intervention for Object Hallucination Mitigation in LVLMsabstractObject hallucination remains a critical challenge in Large Vision-Language Models (LVLMs), where models generate content inconsistent with visual inputs. Existing language-decoder based mitigation approaches often regulate visual or textual attention independently, overlooking their interaction as two key causal factors. To address this, we propose Owl (Bi-mOdal attention reWeighting for Layer-wise hallucination mitigation), a causally-grounded framework that models hallucination process via a structural causal graph, treating decomposed visual and textual attentions as mediators. We introduce VTACR (Visual-to-Textual Attention Contribution Ratio), a novel metric that quantifies the modality contribution imbalance during decoding. Our analysis reveals that hallucinations frequently occur in low-VTACR scenarios, where textual priors dominate and visual grounding is weakened. To mitigate this, we design a fine-grained attention intervention mechanism that dynamically adjusts token- and layer-wise attention guided by VTACR signals. Finally, we propose a dual-path contrastive decoding strategy: one path emphasizes visually grounded predictions, while the other amplifies hallucinated ones -- letting visual truth shine and hallucination collapse. Experimental results on the POPE and CHAIR benchmarks show that Owl achieves significant hallucination reduction, setting a new SOTA in faithfulness while preserving vision-language understanding capability. Our code is available at https://github.com/CikZ2023/OWL Liu Yu 0001, Zhonghao Chen, Ping Kuang, Zhikun Feng, Fan Zhou 0002, Gillian Dobbie |
AAAI | 2 |
| 2026 | When RDMA Goes Long-Haul: Characterization, Modeling, and Verbs-Level Emulation with Implications for Federated LearningabstractLong-haul Remote Direct Memory Access (RDMA) is rapidly emerging as a viable mechanism for extending high-performance communication beyond traditional datacenter environments into wide-area networks (WANs). However, increased round-trip time (RTT) fundamentally alters RDMA behavior, shifting execution from open-loop injection toward a closed-loop, window-limited regime governed by resource constraints and concurrency effects. Consequently, performance characteristics diverge from established datacenter assumptions, rendering existing models insufficient for predicting performance at WAN scale. Understanding these dynamics requires real-world characterization, yet geographically distributed long-haul testbeds remain expensive, scarce, and difficult to reproduce. Yuke Li 0003, Zhonghao Chen, Xiaoyi Lu 0001 |
HPDC | 2 |
| 2026 | Complementary information-guided interactive fusion network for HSI and LiDAR data joint classification
Shufang Xu, Qiyuan Xue, Zhonghao Chen, Shuyu Fei, Hongmin Gao 0001 |
Expert Syst. Appl. | 3 |
| 2026 | Seed-to-Semantics: Few-Shot Prototype-Guided Progressive Learning for Hyperspectral and LiDAR ClassificationabstractDeep learning-based fusion of hyperspectral images (HSI) and LiDAR has achieved strong performance in multimodal remote sensing classification, but its success is heavily constrained by the high cost of pixel-wise annotation. In extremely label-scarce regimes, such as 2-5 labeled samples per class, conventional deep models are prone to severe overfitting, while standard semi-supervised learning (SSL) methods often suffer from confirmation bias because pseudo-labels are generated from unstable early-stage representations. To address these challenges, we propose Prototype-Guided Progressive Learning (PGPL), a unified framework for few-shot HSI-LiDAR classification. Instead of relying solely on model confidence in latent space, PGPL first constructs a reliable initialization pool directly in the original data domain using spectral-angle and elevation-consistency cues, and then progressively expands the training set through class-balanced pseudo-label admission and temporal confidence stabilization. In this way, the framework improves pseudo-label reliability during both initialization and subsequent self-training. Extensive experiments on three benchmark datasets demonstrate that PGPL consistently outperforms state-of-the-art supervised and semi-supervised baselines under the corresponding 2-5-shot settings, achieving overall accuracy gains of 4.64% points on Houston, 1.16% on Trento, and 3.92% on MUUFL over the strongest competing methods, while also yielding higher pseudo-label purity. The source code will be publicly available at https://github.com/zhangyiyan001/PGPL. Hongmin Gao 0001, Weiping Ding 0001, Pedram Ghamisi, Zhonghao Chen, Bing Zhang 0001 |
IEEE Trans. Image Process. | 6 |
| 2025 | DSC-ROM: A Fully Digital Sparsity-Compressed Compute-in-ROM Architecture for on-Chip Deployment of Large-Scale DNNsabstractCompute-in-Memory (CiM) is a promising technique for energy-efficient deep neural network (DNN) inference to miti-gate the memory bottleneck. Unfortunately, conventional SRAM-based CiM has a low density and limited on-chip capacity, resulting in undesired weight reloading from off-chip DRAM. The emerging high-density ROM-based CiM architecture has recently revealed the opportunity of deploying large-scale DNNs on-chip, with optional assisting SRAM to ensure moderate flexibility. However, prior analog-domain ROM CiM still suffers from limited memory density improvement and low computing area efficiency due to stringent array structure and large A/D converter (ADC) overhead. This paper presents DSC-ROM, a fully digital sparsity-compressed compute-in-ROM architecture to address these challenges. DSC-ROM introduces a fully synthesizable macro-level design methodology that achieves a record-high memory density of 27.9 Mb/mm2in a 28nm CMOS technology. Experimental results show that the macro area efficiency of DSC-ROM improves by 5.6-6.6x compared with prior analog-based ROM CiM. Furthermore, a novel weight fine-tuning technique is proposed to ensure task transfer flexibility and reduce required assisting SRAM cells by 94.4%. Experimental results show that DSC-ROM designed for ResNet-18 pre-trained on ImageNet dataset achieves <0.5% accuracy loss in CIFAR-10 and FER2013, compared with the fully SRAM-based CiM. Zhonghao Chen, Yongpan Liu, Huazhong Yang, Xueqing Li 0002 |
DATE | 2 |
| 2025 | DANCE: Dual-Side Agile N:M Sparse Compressed Digital CiM Accelerator for Efficient Compound AIabstractCompound AI systems showcase impressive performance and versatility compared to single AI models by combining large language models (LLMs) with various smaller expert models. The major bottleneck of compound AI lies in frequent data movement due to the massive parameters and dynamic routing mechanisms. Compute-In-Memory (CiM) has demonstrated great potential to mitigate the memory wall. However, constrained by the rigid array structure, existing CiM accelerators struggle to meet more general and diverse model compression demands of compound AI, such as fine-grained pruning for expert models and outlier-aware quantization for LLM-based router models. The lack of support for agile model compression hinders the deployment of compound AI systems on CiM accelerators.To fully unlock the potential of CiM in accelerating compound AI, we present DANCE, a dual-side N:M sparse compressed digital CiM architecture with cross-layer co-optimizations: (i) At the circuit level, DANCE introduces a customized set-associative selection circuit to extract N:M sparse patterns for both weights and activations, maintaining high parallelism; (ii) At the architecture level, DANCE explores a novel design paradigm that integrates fine-grained pruning and outlier-aware quantization into a unified N:M sparsity compression framework. Experimental results show that DANCE achieves up to 4.36× energy efficiency improvement with <1% accuracy loss for ResNet-18 on CIFAR-100, and up to 2.59× energy efficiency improvement with <0.5 perplexity increase for Llama-7B on WikiText-2, compared to the conventional digital CiM baseline. Zhonghao Chen, Hongtao Zhong, Jianhe Deng, Mulin Shi, Yongpan Liu, Huazhong Yang, Xueqing Li 0002 |
ICCAD | 1 |
| 2025 | FedDES: Discrete Event Based Performance Simulation for Federated Learning SystemsabstractFederated Learning (FL) is a scalable and privacy-preserving paradigm well-suited for edge computing. Real-world FL deployments face substantial systems challenges such as compute variability and communication delays, motivating researchers to leverage simulation before real deployment. Most existing FL simulators, however, struggle to scale efficiently and incur long runtimes even for small workloads. To address this, we present FedDES, a high-fidelity, framework-agnostic discrete-event simulation platform that accurately models the runtime behavior of FL systems, including client training, communication overhead, network dynamics, and aggregation strategies. FedDES supports flexible configurations and diverse aggregation approaches, achieving simulation error within 2% of real deployments and delivering over 1000× speedup compared to prior tools. Large-scale experiments with up to 131,072 clients further show that the aggregation strategy critically affects performance, especially under heterogeneous and variable network conditions typical of edge environments. Zhonghao Chen, Weicong Chen 0002, Kibaek Kim, Guanpeng Li, Sheng Di, Xiaoyi Lu 0001 |
SEC | 1 |
| 2025 | Kung-Fu: An Energy-Efficient Compute-In-Memory Approach for Neural Network Inference Using Multi-Level Binary Computing FusionabstractCompute-In-Memory (CiM) is an emerging architecture designed to address the memory wall issue in deep neural network (DNN) inference. However, both the ADC in analog CiM (ACiM) and the adder trees in digital CiM (DCiM) contribute to significant energy and area overhead. In response to these challenges, binary neural networks (BNNs) have been proposed recently. Nevertheless, accuracy degradation poses a serious challenge to the application of BNNs in CiM due to errors in partial-sum accumulations. Furthermore, post-processing steps involving binary activation, such as ReLU, scaling, and bias addition, introduce redundant computing that cannot be effectively optimized by BNN-CiM.This work proposes a novel software-hardware co-optimization approach aimed at enabling an ADC-free analog CiM design while maintaining accuracy. Multi-Level binary computing fusion techniques comprising redundant load isolation based row fusion, in-array parallelism adaption based block fusion, and high-precision post-process elimination based layer fusion address the serious accuracy issues associated with conventional BNN algorithms. In contrast with past over 10% accuracy lost BNN-CiM on practical dataset CIFAR-10 and ImageNet, this work achieves more than 2.2x energy efficiency and 7.4x memory density than state-of-the-art with only 2% accuracy loss. Tianyu Liao, Zhonghao Chen, Yu Wang 0002, Huazhong Yang, Xueqing Li 0002 |
ISCAS | 4 |
| 2025 | HPC-R1: Characterizing R1-like Large Reasoning Models on HPCabstractLarge Reasoning Models (LRMs) are becoming increasingly popular as they offer advanced capabilities in logical inference, mathematical reasoning, and knowledge synthesis, even beyond those of standard language models. However, their complex training workflows present significant challenges in reproducibility, efficiency, and system-level optimization. This paper introduces HPC-R1, a comprehensive characterization of LRM training on the NERSC Perlmutter supercomputer, representing behavior on a Top500-ranked system. We analyze all major stages, including supervised fine-tuning (SFT), Group Relative Policy Optimization (GRPO)-based reinforcement learning (RL), autoregressive generation, and distillation using customized state-of-the-art frameworks. Our detailed performance analysis reveals key system inefficiencies and scaling behaviors. Through our in-depth analysis, we present 19 key observations across all stages, including 4 for SFT, 7 for GRPO-based RL, 6 for generation, and 2 for distillation. Based on these findings, we present several key recommendations to guide future HPC-AI system design. Adam Weingram, Zhonghao Chen, Hao Qi 0008, Xiaoyi Lu 0001 |
SC | 3 |
| 2025 | MDA-HTD: Mask-driven dual autoencoders meet hyperspectral target detection
Zhonghao Chen, Hongmin Gao 0001, Zhengtao Lu, Yao Ding 0010, Xin Li 0090, Bing Zhang 0001 |
Inf. Process. Manag. | 1 |
| 2025 | Multiscale Segmentation-Guided Fusion Network for Hyperspectral Image ClassificationabstractConvolution Neural Networks (CNNs) have demonstrated strong feature extraction capabilities in Euclidean spaces, achieving remarkable success in hyperspectral image (HSI) classification tasks. Meanwhile, Graph convolution networks (GCNs) effectively capture spatial-contextual characteristics by leveraging correlations in non-Euclidean spaces, uncovering hidden relationships to enhance the performance of HSI classification (HSIC). Methods combining GCNs with CNNs have achieved excellent results. However, existing GCN methods primarily rely on single-scale graph structures, limiting their ability to extract features across different spatial ranges. To address this issue, this paper proposes a multiscale segmentation-guided fusion network (MS2FN) for HSIC. This method constructs pixel-level graph structures based on multiscale segmentation data, enabling the GCN to extract features across various spatial ranges. Moreover, effectively utilizing features extracted from different spatial scales is crucial for improving classification performance. This paper adopts distinct processing strategies for different feature types to enhance feature representation. Comparative experiments demonstrate that the proposed method outperforms several state-of-the-art (SOTA) approaches in accuracy. The source code will be released at https://github.com/shengrunhua/MS2FN. Hongmin Gao 0001, Runhua Sheng, Yuanchao Su, Zhonghao Chen, Shufang Xu, Lianru Gao |
IEEE Trans. Image Process. | 4 |
| 2024 | TL2GH²T: Triple-Path Local-to-Global Network With Hybrid Head Transformer for Hyperspectral Change DetectionabstractWith the aid of transformers, significant progress has been achieved in hyperspectral image change detection (HSI-CD) in recent times. Nonetheless, most contemporary detection methods fail to incorporate diverse diagnostic features extracted from hyperspectral (HS) images. In addition, relying solely on algebraic-based techniques to extract information of difference is insufficient for achieving satisfactory detection performance. In this regard, we propose an innovative triple-path local-to-global network (TL2GN), complemented by a hybrid head transformer (HybridHT), called TL2GH2T, tailored for HSI-CD tasks. To be specific, TL2GH2T first investigates spatial, spectral, and spatial–spectral features from a local-to-global perspective. Then, a novel spatial and spectral token fusion (SSTF) module is developed to integrate the above three tokenized features, producing discriminative features from two HS images separately. Moreover, drawing inspiration from chromosomal crossover mechanisms, we propose a HybridHT. Its goal is to simultaneously learn cross correlation and self-correlation information of bitemporal features from a global perspective, producing highly discriminative distinctions. Our approach, validated through extensive experimentation on four varied HS benchmarks, exhibits exceptional performance in HSI-CD, outperforming contemporary methods in both visual and quantitative evaluations. Zhonghao Chen, Swalpa Kumar Roy, Hongmin Gao 0001, Yao Ding 0010, Xiongwu Xiao, Bing Zhang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2024 | Multiscale Random-Shape Convolution and Adaptive Graph Convolution Fusion Network for Hyperspectral Image ClassificationabstractConvolution neural networks (CNNs) are extensively utilized in hyperspectral image (HSI) classification due to their remarkable capability to extract features from patterns with fixed shapes. These networks have been shown to effectively capture features at the pixel level. However, the fixed shape of convolution kernels poses a challenge for CNNs to adapt to the diverse shapes found in HSIs. Graph neural networks (GNNs), particularly graph convolution networks (GCNs), possess robust feature extraction capabilities on graph structures and are extensively applied in HSI classification. However, one significant challenge in using GNNs is the selection of appropriate neighboring nodes for information aggregation. To address the existing challenges of GCN and CNN and leverage their respective advantages, this paper introduces a novel patch-based CNN-GCN fusion classification network, named multi-scale random-shape convolution and adaptive graph convolution fusion network (MRCAGCFN). It consists of a spectral transformation module and three main modules we proposed: a multi-scale random-shape convolution module for extracting convolution features, where the shape of the convolution kernel is randomized and a multi-scale approach is applied to enhance adaptability to data with diverse shapes; an adaptive feature-fusion graph convolution module for extracting graph convolution features, where the weights for neighborhood aggregation are learned adaptively to reduce feature fusion from dissimilar nodes and strengthen feature fusion from similar nodes; and an adaptive local feature processing module for processing features, where two different methods are employed to convert patch-level features to pixel-level features, thereby improving feature representation. MRCAGCFN combines the strengths of CNN and GCN while introducing enhancements to better accommodate diverse feature shapes. Experimental results on three HSI classification datasets demonstrate that our proposed MRCAGCFN outperforms some existing methods. The codes of our MRCAGCFN will be available at https://github.com/shengrunhua/MRCAGCFN. Hongmin Gao 0001, Runhua Sheng, Zhonghao Chen, Haiyun Liu, Shufang Xu, Bing Zhang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2024 | Airborne Small Target Detection Method Based on Multimodal and Adaptive Feature FusionabstractThe detection of airborne small targets amidst cluttered environments poses significant challenges. Factors such as the susceptibility of a single RGB image to interference from the environment in target detection and the difficulty of retaining small target information in detection necessitate the development of a new method to improve the accuracy and robustness of airborne small target detection. This article proposes a novel approach to achieve this goal by fusing RGB and infrared (IR) images, which is based on the existing fusion strategy with the addition of an attention mechanism. The proposed method employs the YOLO-SA network, which integrates a YOLO model optimized for the downsampling step with an enhanced image set. The fusion strategy employs an early fusion method to retain as much target information as possible for small target detection. To refine the feature extraction process, we introduce the self-adaptive characteristic aggregation fusion (SACAF) module, leveraging spatial and channel attention mechanisms synergistically to focus on crucial feature information. Adaptive weighting ensures effective enhancement of valid features while suppressing irrelevant ones. Experimental results indicate 1.8% and 3.5% improvements in mean average precision (mAP) over the LRAF-Net model and Infusion-Net detection network, respectively. Additionally, ablation studies validate the efficacy of the proposed algorithm’s network structure. Shufang Xu, Tianci Liu 0007, Zhonghao Chen, Hongmin Gao 0001 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2024 | Cognitive Fusion of Graph Neural Network and Convolutional Neural Network for Enhanced Hyperspectral Target DetectionabstractIn recent years, deep learning has emerged as a prominent technique in hyperspectral target detection (HTD). Extensive research has highlighted the potential of Graph Neural Network (GNN) as a promising framework for exploring non-Euclidean dependencies within hyperspectral imagery. However, GNN has not been introduced to HTD. Additionally, achieving a balanced training set while effectively suppressing background remains a challenge. Therefore, we propose the cognitive fusion of GNN and Convolutional Neural Network (CNN) for enhanced HTD (named as CFGC), which marks the first integration of GNN and CNN in HTD. Initially, using sparse subspace clustering and a similarity measurement strategy, we select the most representative background samples for HTD. Subsequently, linear interpolation combines the prior target with the Laplacian-weighted prior target, yielding abundant targets with meaningful transformations. Finally, a fused network of CNN and GNN is utilized for training both the prior target and the constructed training set. Significantly, the incorporation of attention mechanism in both the CNN and GNN branches stands out as a noteworthy advantage, augmenting the models’ ability to selectively prioritize crucial information. Four benchmark hyperspectral images have been used in extensive experiments, and the results demonstrate that CFGC exhibits superior performance in HTD. Shufang Xu, Sijie Geng, Zhonghao Chen, Hongmin Gao 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2024 | A Module-Level Configuration Methodology for Programmable Camouflaged LogicabstractLogic camouflage is a widely adopted technique that mitigates the threat of intellectual property (IP) piracy and overproduction in the integrated circuit (IC) supply chain. Camouflaged logic achieves functional obfuscation through physical-level ambiguity and post-manufacturing programmability. However, discussions on programmability are confined to the level of logic cells/gates, limiting the broader-scale application of logic camouflage. In this work, we propose a novel module-level configuration methodology for programmable camouflaged logic that can be implemented without additional hardware ports and with negligible resources. We prove theoretically that the configuration of the programmable camouflaged logic cells can be achieved through the inputs and netlist of the original module. Further, we propose a novel lightweight ferroelectric FET (FeFET)-based reconfigurable logic gate (rGate) family and apply it to the proposed methodology. With the flexible replacement and the proposed configuration-aware conversion algorithm, this work is characterized by the input-only programming scheme as well as the combination of high output error rate and point-function-like defense. Evaluations show an average of >95% of the alternative rGate location for camouflage, which is sufficient for the security-aware design. We illustrate the exponential complexity in function state traversal and the enhanced defense capability of locked blackbox against Boolean Satisfiability (SAT) attacks compared with key-based methods. We also preserve an evident output Hamming distance and introduce negligible hardware overheads in both gate-level and module-level evaluations under typical benchmarks. Zhonghao Chen, Yixin Xu 0001, Tongguang Yu, Ziheng Zheng, Enze Ye, Sumitha George, Huazhong Yang, Yongpan Liu, Kai Ni 0004, Narayanan Vijaykrishnan, Xueqing Li 0002 |
ACM Trans. Design Autom. Electr. Syst. | 2 |
| 2023 | ASMCap: An Approximate String Matching Accelerator for Genome Sequence Analysis Based on Capacitive Content Addressable MemoryabstractGenome sequence analysis is a powerful tool in medical and scientific research. Considering the inevitable sequencing errors and genetic variations, approximate string matching (ASM) has been adopted in practice for genome sequencing. However, with exponentially increasing bio-data, ASM hardware acceleration is facing severe challenges in improving the throughput and energy efficiency with the accuracy constraint.This paper presents ASMCap, an ASM acceleration approach for genome sequence analysis with hardware-algorithm co-optimization. At the circuit level, ASMCap adopts charge-domain computing based on the capacitive multi-level content addressable memories (ML-CAMs), and outperforms the state-of-the-art ML-CAM-based ASM accelerators EDAM with higher accuracy and energy efficiency. ASMCap also has misjudgment correction capability with two proposed hardware-friendly strategies, namely the Hamming-Distance Aid Correction (HDAC) for the substitution-dominant edits and the Threshold-Aware Sequence Rotation (TASR) for the consecutive indels. Evaluation results show that ASMCap can achieve an average of 1.2x (from 74.7% to 87.6%) and up to 1.8x (from 46.3% to 81.2%) higher F1score (the key metric of accuracy), 1.4x speedup, and 10.8x energy efficiency improvement compared with EDAM. Compared with the other ASM accelerators, including ResMA based on the comparison matrix, and SaVI based on the seeding strategy, ASMCap achieves an average improvement of 174x and 61x speedup, and 8.7e3x and 943x higher energy efficiency, respectively. Hongtao Zhong, Zhonghao Chen, Wenqin Huangfu, Yixin Xu 0001, Yongpan Liu, Narayanan Vijaykrishnan, Huazhong Yang, Xueqing Li 0002 |
DAC | 2 |
| 2023 | Local aggregation and global attention network for hyperspectral image classification with spectral-induced aligned superpixel segmentation
Zhonghao Chen, Guoyong Wu, Hongmin Gao 0001, Yao Ding 0010, Danfeng Hong, Bing Zhang 0001 |
Expert Syst. Appl. | 1 |
| 2023 | Grid Network: Feature Extraction in Anisotropic Perspective for Hyperspectral Image ClassificationabstractAbundant spectral signatures and spatial characteristics embedded in hyperspectral (HS) images enable the fine identification of land covers, attracting plenty of studies on feature extraction and feature utilization. Nevertheless, the high representative spectral and spatial features in the HS cube are unevenly distributed, which is failed to consider by many current methods. To conquer this shortcoming, we rethink the feature extraction of HS images from an anisotropic perspective and propose a novel model called grid network (GNet) for HS image classification. Beyond representing spectral-spatial features in three classic paradigms (simultaneously, hierarchically, and separately), GNet is capable of learning them in two new processes: multi-stage and multi-path. In this way, spectral and spatial features can be fully and balanced explored. More significantly, to make full use of low- and high-level features and avoid the existing semantic gap, we devise a spectral-spatial cross-level feature fusion module to model the relation between them. Extensive experiments, implemented on three HS datasets, demonstrate that the proposed GNet enables to acquire promising classification performance compared to state-of-the-art methods. The codes of this work will be available at https://github.com/zhonghaochen/GNet_Master for the sake of reproducibility. Zhonghao Chen, Danfeng Hong, Hongmin Gao 0001 |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2023 | Adaptively Dictionary Construction for Hyperspectral Target DetectionabstractThe task of hyperspectral images (HSIs) target detection is to identify whether the target spectral sequences present in the HSI. Recently, the topic of representation models has received much interest in hyperspectral target detection. The performance of representation models depends on whether the corresponding dictionary and sparse matrix can correctly recover the original spectrum. Therefore, the background dictionary of these models should contain the spectra of all classes except the target spectrum; i.e., the dictionary should be overcomplete. However, most representation models cannot satisfy this condition. Moreover, due to the potentially large spectral similarity between the target and the background, representation models perform poorly in background suppression. Aiming to solve these issues, a novel adaptively dictionary construction (ADC) strategy with background suppression sparse representation (BSSR) module is proposed in this letter, called adaptively dictionary construction for target detection (ADCTD). Specifically, the proposed ADC is adopted to segment the HSI into superpixels consisting of pixels with similar spectra. This process can be considered as an unsupervised coarse classification process, which can construct an overcomplete background dictionary. In addition, the BSSR is adopted to improve the separation of the target and background by a linear function. Experiments on three datasets demonstrate the superiority of the proposed ADCTD. Weibo Zhang, Zhonghao Chen, Hongmin Gao 0001 |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2023 | A Multidepth and Multibranch Network for Hyperspectral Target Detection Based on Band SelectionabstractDeep learning (DL) has recently risen to prominence in hyperspectral target detection (HTD). Nevertheless, how to tackle the extreme training sample imbalance together with achieving target highlighting and background suppression is challenging. Additionally, due to the spectral redundancy of hyperspectral imagery (HSI), it is a new course for HTD through band selection (BS) to retain crucial bands thereupon improving the subsequent detection performance. Accordingly, we propose a DL-based BS-HTD (DLBSTD) algorithm, incorporating DL-based BS with DL-based HTD for the first time. Most significantly, a multi-depth and multi-branch network (MDBN) for HTD based on a novel BS method is proposed. First of all, the BS method including an alternating local-global reconstruction network (ALGRN) and a correlation measurement strategy provides representative bands containing key target information for MDBN. For the training sample imbalance of MDBN, we develop a BS-based method to select multifarious representative background training samples and propose a target band random substitution (TBRS) strategy to augment an ample target training set. Lastly, the MDBN composed of a multi-depth feature extraction (MDFE) module, three fusion strategies, and the parallel local convolution and gated recurrent unit (Conv-GRU) fully taps the spectral feature relationships to highlight targets and suppress backgrounds. Compared with nine competitive HTD algorithms, we carry out plentiful experiments on four classical datasets exhibiting that the proposed DLBSTD has strong generalization and salient detection performance of target highlighting and background suppression. Hongmin Gao 0001, Zhonghao Chen, Shufang Xu, Danfeng Hong, Bing Zhang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2023 | Hyperspectral Target Detection via Spectral Aggregation and Separation Network With Target Band Random MaskabstractHyperspectral target detection (HTD) is a pixel-wise detection method based on limited prior targets and spectral differences, which has been widely studied and applied in many fields. Recently, deep learning (DL) plays an important role in hyperspectral imagery (HSI) processing. However, for HTD, the severe lack of class-balanced training sets is an enormous challenge. Meanwhile, it is difficult to suppress backgrounds while highlighting targets through the deep network. To address these issues, we propose a spectral aggregation and separation network (SASN) with a target band random mask (TBRM) for HTD in this paper. For the training sets of SASN, a multifarious representative background selection strategy (MRBS) is first proposed to obtain a multifarious and representative background training set. Next, aiming at the notorious class imbalance, a data augmentation (DA) method, TBRM, is proposed to generate adequate target training set by repeating randomly zero-masking the spectral bands of a prior target. Subsequently, in the training of SASN, residual connection and squeeze-and-excitation (SE) channel attention mechanism are applied to fully extract high discriminative features and nonlinear ones in the spectra. Besides, to better separate the targets and backgrounds, a triplet-soft loss function is presented, which makes the training in the direction of spectral separation of background samples from both the prior target and target samples. During testing, the trained SASN distinguishes the spectral similarities and differences simultaneously for highlighting targets and suppressing backgrounds. Moreover, extensive experimental results validate that the proposed method has superior detection performances, background suppression capacity, and separability compared with ten cutting-edge HTD algorithms on six benchmark HSI datasets. Hongmin Gao 0001, Zhonghao Chen, Feng Xu 0008, Danfeng Hong, Bing Zhang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2022 | Multiscale spectral-spatial cross-extraction network for hyperspectral image classificationabstractAbstract Convolutional neural networks (CNN) are becoming increasingly popular in modern remote sensing image classification tasks and have exhibited excellent results. For the existing CNN‐based hyperspectral image (HSI) classification methods, most of which extract spatial or spectral features separately by convolution. But nearly all of these methods ignore the fact that the weighted summation of convolution may lead to appear new features in another dimension. To address this issue, a novel multiscale spectral‐spatial cross‐extraction network (MSSCEN) is proposed for HSI classification. Specifically, the proposed MSSCEN introduces spectral‐spatial features cross extraction module (SSCEM), which fed extracted features from previous layer into spatial and spectral extraction branches separately again, so that the changes that occurred in the other domain after each convolution can be fully utilized. In addition, a new independent data augmentation module based on U‐Net is designed to mitigate the problem of limited labelled samples. The paper conducts experiments on three classic hyperspectral datasets and the results demonstrate that the proposed method achieves the best classification accuracy than other state‐of‐the‐art methods. Hongmin Gao 0001, Hongyi Wu, Zhonghao Chen |
IET Image Process. | 3 |
| 2022 | Shallow Network Based on Depthwise Overparameterized Convolution for Hyperspectral Image ClassificationabstractRecently, convolutional neural network (CNN) techniques have gained popularity as a tool for hyperspectral image classification (HSIC). To improve the feature extraction efficiency of HSIC under the condition of limited samples, the current methods generally use deep models with plenty of layers. However, deep network models are prone to overfitting and gradient vanishing problems when samples are limited. In addition, the spatial resolution decreases severely with deeper depth, which is very detrimental to spatial edge feature extraction. Therefore, this letter proposes a shallow model for HSIC, which is called a depthwise overparameterized convolutional neural network (DOCNN). To ensure the effective extraction of the shallow model, the depthwise overparameterized convolution (DO-Conv) kernel is introduced to extract the discriminative features. The DO-Conv kernel is composed of a standard convolution kernel and a depthwise convolution kernel, which can extract the spatial feature of the different channels individually and fuse the spatial features of the whole channels simultaneously. Moreover, to further reduce the loss of spatial edge features due to the convolution operation, a dense residual connection (DRC) structure is proposed to apply to the feature extraction part of the whole network. Experimental results obtained from three benchmark datasets show that the proposed method outperforms other state-of-the-art methods in terms of classification accuracy and computational efficiency. Hongmin Gao 0001, Zhonghao Chen |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2022 | Global to Local: A Hierarchical Detection Algorithm for Hyperspectral Image Target DetectionabstractHyperspectral image (HSI) has received considerable attention in the field of target detection due to its powerful ability to capture the spectral information of land covers, and plenty of detection algorithms have been explored. However, these methods generally leverage the difference between the spectrum of the target to be detected and the background spectrum to accomplish target detection, and so are susceptible to the problem of spectral variability. In this article, we propose a global-to-local hierarchical detection algorithm for HSI (G2LHTD). Firstly, extended morphological attribute profile (EMAP) is first used to model global spatial texture information from HSI. Subsequently, a diverse-direction constrained energy minimization (D2CEM) detector is developed to consider the spatial information within eight neighborhoods around each pixel in HSI, yielding comprehensive local spatial information. More substantially, to effectively discriminate the neighborhood information in diverse directions, we devise an adaptive neighborhood feature aggregation (ANFA) strategy, which will comprehensively evaluate the significance of neighborhood information in diverse directions. As a result, the spatial features of HSI can be comprehensively considered for hyperspectral target detection (HTD). Extensive experiments, conducted on four standard datasets, demonstrate the effectiveness of the proposed method. The codes of this work will be available at https://github.com/zhonghaocheng/G2LHTD_Master for the sake of reproducibility. Zhonghao Chen, Zhengtao Lu, Hongmin Gao 0001, Jia Zhao 0001, Danfeng Hong, Bing Zhang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2021 | An Attention Method to Introduce Prior Knowledge in Dialogue State Tracking
Zhonghao Chen |
ICONIP (3) | 1 |