EDBT 2026 Demo / reviewers in the wild / expert
Hongyu Chen 0003
dblp:28/3046-3
· DBLP profile ↗
13ranked-venue papers
4as first author
12since 2021 · last 2026
0000-0002-5329-8854ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 9 · 3 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Self-Supervised One-Step Diffusion Refinement for Snapshot Compressive ImagingabstractSnapshot compressive imaging (SCI) captures multispectral images (MSIs) using a single coded two-dimensional (2-D) measurement, but reconstructing high-fidelity MSIs from these compressed inputs remains a fundamentally ill-posed challenge. Recent diffusion-based methods improve quality but are limited by scarce MSI training data, domain shifts from RGB-pretrained models, and slow multi-step sampling. These drawbacks restrict their practicality in real-world applications. Unlike prior approaches that rely on expensive iterative refinement or subspace-based diffusion embeddings (e.g., DiffSCI, PSR-SCI)—we introduce a fundamentally different paradigm: a self-supervised One-Step Diffusion (OSD) framework designed specifically for SCI. The key novelty lies in using a single-step diffusion refiner to correct an initial reconstruction, eliminating iterative denoising entirely while preserving generative quality. Moreover, we adopt a self-supervised equivariant learning strategy to train both the predictor and refiner directly from raw 2-D measurements, enabling generalization to unseen domains without ground-truth MSI. To further address limited MSI data, we design a band-selection–driven distillation strategy that transfers core generative priors from large-scale RGB datasets, effectively bridging the domain gap. Extensive experiments confirm that our approach sets a new standard—yielding PSNR gains of 3.44dB, 1.61dB, and 0.28dB on the Harvard, NTIRE, and ICVL datasets respectively, while cutting reconstruction time from 8.9s to just 0.22s per image. These gains in efficiency and adaptability advance SCI reconstruction, enabling accurate and practical real-world deployment. Shaoguang Huang, Yunzhen Wang, Haijin Zeng, Hongyu Chen 0003, Hongyan Zhang 0001 |
AAAI | 4 |
| 2026 | Unsupervised High-Order Implicit Neural Representation With Line Attention for Metal Artifact ReductionabstractThe presence of metallic implants introduces bright and dark streaks that appear in computed tomography (CT) images, degrading image quality and interfering with medical diagnosis. To reduce these artifacts, deep learning approaches have been applied for metal-corrupted restoration, which usually requires a large amount of simulated degraded-clean pairs for training. To achieve metal artifact reduction (MAR) without reference images, implicit neural representation (INR) has emerged and shown capabilities for image restoration in an unsupervised manner. However, existing INR methods for MAR usually treat the spatial coordinates independently and ignore their correlation, resulting in detail loss and artifacts remaining. In this paper, we propose an INR-based unsupervised MAR framework and design a High-order Line Attention Network to capture local contextual and geometric representations from X-rays, which maps the spatial coordinates into discrete linear attenuation coefficients of imaged objects for artifact-free CT image reconstruction. The second-order feature interaction can effectively improve the spectral bias problems and fit low and high-frequency details of real signals well. The proposed line-attention module with linear complexity can establish global relationships among spatial point tokens from sampled rays. To provide more local contextual information, a multiple local adjacent ray sampling strategy is adopted to compose several sub-fan beams with more context as a training batch. With the help of these components, the unsupervised MAR framework can approximate the implicit continuous function to estimate measurements and generate artifact-free CT images. Simulated and real experiments indicated that the proposed approach achieved superior MAR performance compared with other state-of-the-art methods. Hongyu Chen 0003, Shaoguang Huang, Wei He 0003, Hongyan Zhang 0001 |
IEEE Trans. Medical Imaging | 1 |
| 2025 | Heterogeneous Data-based Cross-domain Few-shot Classification Method of Hyperspectral ImageabstractFew-shot learning (FSL) has been employed in hyperspectral image (HSI) classification, achieving excellent performance with limited training data. However, existing HSI few-shot classification methods often encounter the problem of insufficient domain-transferable knowledge learning that is either from natural images or HSI solely. In this paper, we propose a two-stage cross-domain few-shot classification method of HSI, which for the first time makes use of heterogeneous labeled natural images and HSIs in the source domain (SD) to support the classification of novel classes in the target HSI domain. We first use a large amount of labeled natural images at the first stage to pre-train a backbone, which will be used to extract the spatial feature of HSIs at the second stage with fine-tuning. In the second stage, we propose a cross-domain few-shot classification method, which allows for effective discriminative feature learning in the target HSI domain with the transferred knowledge of the old classes obtained from natural images and HSIs in the source domain. To obtain domain-transferable knowledge, FSL is employed on the HSI source and target domain. To deal with the domain shift problem, we propose a class-matching based cross-domain contrastive loss. In addition, we take into account the large spectral variations problem in the target HSI domain and introduce an instance-level self-supervised loss. Experimental results on real data sets demonstrate that our method outperforms the recent state-of-the-art. Shaoguang Huang, Hongyu Chen 0003, Hongyan Zhang 0001 |
ICASSP | 3 |
| 2025 | Hyperspectral Image Classification Based on a Locally Enhanced Transformer NetworkabstractRecently, transformer-based models have achieved remarkable performance in the hyperspectral image (HSI) classification. However, due to the limited training data, existing methods often show limited capability of capturing fine-grained local features. Although attempts have been made to solve this problem, the large amount of parameters imposes the risk of overfitting. In this paper, we propose a locally enhanced transformer network for HSI classification with fewer network parameters, which mainly consists of a multi-branch spatial-spectral tokenization (MSST) module and a dual-branch transformer encoder (DTE) module. The MSST generates effective spatialspectral tokens through diverse convolutions with a residual connection. The DTE consists of a global transformer branch and a locally enhanced transformer branch, which are used to capture the global and local spatial dependencies of HSI, respectively. Unlike the conventional self-attention module used in the global branch, we propose an improved multi-head selfattention (IMSA) module in the local branch by incorporating the local prior information of HSI with graph convolution, to enhance the local information extraction. To fuse the global and local features from the two branches, we introduce an adaptive strategy by using learnable weights for both branches. We devise our MSST and DTE with a shallow architecture, significantly reducing the number of parameters. Experimental results on benchmark datasets demonstrate that the proposed method outperforms the state-of-the-art. Shaoguang Huang, Hongyu Chen 0003, Siti Khairunniza-Bejo, Hongyan Zhang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2025 | Heterogeneous Data-Based Global-to-Local Cross-Domain Few-Shot Classification Method of Hyperspectral ImageabstractCross-domain few-shot learning (FSL) has shown promising performance in hyperspectral image classification (HSIC) under limited labeled data. However, existing approaches often suffer from insufficient meta-knowledge transfer due to reliance on a single source domain, and fail to fully bridge the domain gap owing to the use of single-level alignment strategies. In this paper, we propose a heterogeneous data-driven, global-to-local cross-domain FSL framework for HSIC, leveraging richly labeled natural RGB images and hyperspectral data as source domains to support classification in the target domain with only a few labeled samples. The proposed method consists of two stages. In the first stage, we employ natural RGB images to learn a powerful spatial feature extractor with FSL and self-supervised learning, which will be fine-tuned for HSI in the second stage. To alleviate the domain gap between the source and target HSI domains at the second cross-domain FSL stage, we propose a global-to-local domain adaptation strategy that performs alignment both at the domain level and class level, effectively reducing the learning bias toward the source HSI while promoting discriminative feature learning. Additionally, to address large intra-class variance in the target HSI domain, we introduce a self-supervised contrastive loss based on positive pairs only, enhancing the within-class representation compactness. Extensive experiments on three benchmark datasets demonstrate that our method outperforms the state-of-the-art. Shaoguang Huang, Hongyu Chen 0003, Hongyan Zhang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2024 | Spatial and Cluster Structural Prior-Guided Subspace Clustering for Hyperspectral ImageabstractSubspace clustering has achieved remarkable performance for hyperspectral image (HSI). However, existing methods are often computationally expensive and have limited ability to capture the intrinsic structural information of HSI. In this paper, we propose a structural prior-guided subspace clustering method, which simultaneously incorporates the local and non-local spatial information and the cluster prior information. Accordingly, three efficient regularizations are developed. Considering the local connectivity of pixels, we propose an ℓ2,1norm based constraint on the representation difference matrix to improve the homogeneity of clustering result. Next, to capture the non-local geometric structure of HSI, we propose a manifold-based regularization with an adaptively learned landmark graph. Furthermore, we explore the block-diagonal cluster structure of HSI and develop a landmark-based clustering constraint, which makes the representations more favorable for clustering. Our local constraint is imposed on all the data points due to its efficiency and the latter two are solely imposed on landmarks, leading to computationally efficient regularizations. Due to the local constraint, the manifold and cluster structure of the landmarks can be effectively propagated to all the data points. To make our model scalable to large-scale data, we learn a compact dictionary with an orthogonal constraint, significantly reducing the number of parameters. In addition, we propose a novel landmark selection method to support our landmark-based constraints using multi-scale super-pixel segmentation and clustering, which improves the uniformity and diversity of landmarks. We also develop an efficient algorithm to solve the proposed model. Experimental results demonstrate that our model outperforms the state-of-the-art. Shaoguang Huang, Haijin Zeng, Hongyu Chen 0003, Hongyan Zhang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2024 | SOSSF: Landsat-8 Image Synthesis on the Blending of Sentinel-1 and MODIS DataabstractLandsat optical sensor is crucial for the long-term observations of the Earth’s surface with a 30 m spatial resolution. However, the 16-day revisit cycle and severe atmospheric interference have impeded the monitoring of rapid surface changes. Spatiotemporal fusion (STF) is a classic method of predicting Landsat surface reflectance with multi-temporal and multi-source data, but it is limited by unpredictable temporal changes and cloudy Landsat-MODIS image pairs. Another emerging solution is synthetic aperture radar (SAR)-to-optical image translation (S2OIT), which always produces spectral distortions. To tackle these defects, we propose a new data-driven solution, SAR-optical data-based spatial–spectral fusion (SOSSF), which combines the high-spatial and cloud-free advantages of Sentinel-1 data and the high-spectral and high-temporal advantages of MODIS images to synthesize high-spatial and high-temporal Landsat-8 images. To achieve this solution, we first establish a worldwide benchmark dataset, namely SMILE, with various land cover types and all meteorological seasons, satisfying the big data requirements of deep learning. Second, we design an attention-based dual-path fusion network (ADFNet) to respectively extract and fully fuse spatial and spectral information from SAR-optical data. Extensive experiments suggest that the proposed SOSSF solution outperforms the state-of-the-art STF and S2OIT solutions, robustly performing in the continuously changing and frequently cloudy regions. The proposed ADFNet model achieves the best visual effect and the highest accuracy in different scenes, seasons, and bands. Furthermore, the proposed SOSSF solution is proven to be a practical way to simulate time-series and large-scale Landsat-8 surface reflectance, considerably enriching raw Landsat-8 products. Yu Xia 0032, Wei He 0003, Hongyu Chen 0003, Hongyan Zhang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2024 | Hider: A Hyperspectral Image Denoising Transformer With Spatial-Spectral Constraints for Hybrid Noise RemovalabstractHyperspectral image (HSI) qualities are limited by a mixture of Gaussian noise, impulse noise, stripes, and deadlines during the sensor imaging process, resulting in weak application performance. To enhance HSI qualities, methods based on convolutional neural networks have been successively applied to restore clean data from the observed data. However, the architecture of these methods lacks spectral and spatial constraints, and the convolution operators have limited receptive fields and inflexible model inferences. Thus, in this study, we propose an efficient end-to-end transformer, named HSI denoising transformer (Hider), for mixed HSI noise removal. First, a U-shaped 3-D transformer architecture is built for multiscale feature aggregation. Second, a multihead global spectral attention module within the spectral transformer block is designed to excavate information in different spectral patterns. Finally, an additional locally enhanced cross-spatial attention module within the spatial-spectral transformer block is constructed to build the long-range spatial relationship to avoid the high computational complexity of global spatial self-attention. Through the imposition of global correlations along spectrum and spatial self-similarity constraints on the transformer, our proposed Hider aims to capture long-range spatial contextual information and cluster objects with the same spectral pattern for HSI denoising. To verify the effectiveness and efficiency of Hider, we conducted extensive simulated and real experiments. The denoising results on both simulated and real-world datasets show that Hider achieves superior evaluation metrics and visual assessments compared with other state-of-the-art methods. Hongyu Chen 0003, Hongyan Zhang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2022 | Autoencoder in Autoencoder Network Based on Low-Rank Embedding for Anomaly Detection in Hyperspectral ImagesabstractThe main purpose of anomaly detection in hyperspectral images is to detect targets that are different from their surroundings. With the development of deep learning technology, anomaly detection in hyperspectral images using deep neural networks has drawn great attention in recent years. However, most of the existing deep learning-based anomaly detection algorithms fail to consider the low-rank properties of the background and underutilize the rich spectral information of the image. In this paper, we propose a novel autoencoder in autoencoder network based on low-rank module embedding for anomaly detection in hyperspectral images. Firstly, the background is purified by using the low-rank module (LRM), and then the image background is reconstructed by using autoencoder in autoencoder network (AiANet), which is a spatial-spectral dual encoding-decoding network. AiANet fully considers the differences between anomalies and backgrounds in spatial and spectral dimensions to better reconstruct the background. Finally, the anomaly appears in images as reconstruction errors. Our proposed method effectively exploits the low-rank property of the backgrounds and makes full use of the spectral information to extract pure backgrounds to separate anomalies. Experiments on two real hyperspectral images demonstrate that the proposed method outperforms the other competitors. Weinan Cao, Hongyan Zhang 0001, Wei He 0003, Hongyu Chen 0003, Ewe Hong Tat |
IGARSS | 4 |
| 2022 | MSBRNet: Multi-Scale Background Reconstruction Network with Low-Rank Embedding for Anomaly Detection in Hyperspectral ImagesabstractThe primary purpose of anomaly detection in hyperspectral images (HSI) is to detect different anomaly targets from their surrounding backgrounds. Recently, anomaly detection has been well developed by deep learning technology. However, previous works utilize deep neural networks as a feature extractor followed by a traditional detector to detect anomalies, which cannot separate background and anomaly effectively. In this paper, we try to extract low-rank background features using neural networks to make full use of the low-rank properties of the background and then reconstruct the background with these features to directly separate the anomaly from the background, so we propose an multi-scale background reconstruction network with low-rank embedding (MSBRNet) for anomaly detection in HSI. Firstly, we use a low-rank background features extraction module (LBM) to extract low-rank background features. Then the background is reconstructed using a multi-scale background reconstruction module (MBRM). Finally, we calculate the mean square error of the input image and the output background to measure the effect of background reconstruction, and anomalies appear as reconstruction errors. Experiments on two publicly available experimental datasets demonstrate the significant advantage of the proposed method over other competitors. Weinan Cao, Hongyan Zhang 0001, Wei He 0003, Hongyu Chen 0003, Ewe Hong Tat |
IGARSS | 4 |
| 2022 | A Mutual Information Domain Adaptation Network for Remotely Sensed Semantic SegmentationabstractAlthough deep learning has made semantic segmentation of very-high-resolution (VHR) remote sensing (RS) images practical and efficient, its large-scale application is still limited. Given the diversity of imaging sensors, acquisition conditions, and regional styles, a deep learning network well-trained on one source domain dataset often suffers from drastic performance drops when applied to other target domain datasets. Thus, we propose a novel end-to-end mutual information domain adaptation network (MIDANet) that can shift between semantic segmentation domains by integrating multitask learning in the convolutional neural networks within an entropy adversarial learning (EAL) framework. Through the joint learning of semantic segmentation and elevation estimation, the features extracted by MIDANet can concentrate more on the elevation clues while dropping the domain-variant information (i.e., texture, spectral information). First, one encoder is applied to excavate general semantic features. Two decoders that share the same architecture are used to perform pixel-level classification and digital surface model (DSM) regression. Second, feature interaction modules (FIMs) and a mutual information attention unit (MIAU) are designed to mine the latent relationships between the two tasks and enhance their feature representations. Finally, a final MIDANet is obtained for semantic segmentation that does not require any semantic segmentation labels in the target domain after the adversarial learning of the classification entropy at the output level. Extensive comparative experiments and ablation studies were conducted on the International Society for Photogrammetry and Remote Sensing (ISPRS) Potsdam and Vaihingen test datasets. The results show that MIDANet outperforms other state-of-the-art domain adaptation (DA) methods in both evaluation metrics and visual assessment. Hongyu Chen 0003, Hongyan Zhang 0001, Shengyang Li, Liangpei Zhang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2021 | LR-Net: Low-Rank Spatial-Spectral Network for Hyperspectral Image DenoisingabstractDue to the physical limitations of the imaging devices, hyperspectral images (HSIs) are commonly distorted by a mixture of Gaussian noise, impulse noise, stripes, and dead lines, leading to the decline in the performance of unmixing, classification, and other subsequent applications. In this paper, we propose a novel end-to-end low-rank spatial-spectral network (LR-Net) for the removal of the hybrid noise in HSIs. By integrating the low-rank physical property into a deep convolutional neural network (DCNN), the proposed LR-Net simultaneously enjoys the strong feature representation ability from DCNN and the implicit physical constraint of clean HSIs. Firstly, spatial-spectral atrous blocks (SSABs) are built to exploit spatial-spectral features of HSIs. Secondly, these spatial-spectral features are forwarded to a multi-atrous block (MAB) to aggregate the context in different receptive fields. Thirdly, the contextual features and spatial-spectral features from different levels are concatenated before being fed into a plug-and-play low-rank module (LRM) for feature reconstruction. With the help of the LRM, the workflow of low-rank matrix reconstruction can be streamlined in a differentiable manner. Finally, the low-rank features are utilized to capture the latent semantic relationships of the HSIs to recover clean HSIs. Extensive experiments on both simulated and real-world datasets were conducted. The experimental results show that the LR-Net outperforms other state-of-the-art denoising methods in terms of evaluation metrics and visual assessments. Particularly, through the collaborative integration of DCNNs and the low-rank property, the LR-Net shows strong stability and capacity for generalization. Hongyan Zhang 0001, Hongyu Chen 0003, Liangpei Zhang 0001 |
IEEE Trans. Image Process. | 2 |
| 2019 | Multi-Level Fusion of the Multi-Receptive Fields Contextual Networks and Disparity Network for Pairwise Semantic StereoabstractIn this paper, we propose a multi-level fusion framework to address the pairwise semantic stereo issue. For disparity estimation, we adopt the pyramid stereo matching network. For semantic segmentation, the single segmentation network is proposed with respect to the left image, along with the disparity fusion segmentation network for the combination of semantic features and disparity features. Specifically, the multi-receptive fusion block is designed and employed to fully extract and fuse the contextual information. Finally, the refined segmentation result is obtained via yet another fusion of the multi-model results. The proposed method achieved a mean intersection over union (mIoU) of 79.05%, an average endpoint error (EPE) of 1.3966, and an mIoU-3 of 77.75%, ranking first in the Pairwise Semantic Stereo Challenge of the 2019 IEEE GRSS Data Fusion Contest [1],[2]. Hongyu Chen 0003, Manhui Lin, Hongyan Zhang 0001, Gui-Song Xia, Xianwei Zheng, Liangpei Zhang 0001 |
IGARSS | 1 |