Junming Hou

dblp:01/1091 · DBLP profile ↗
← Back
17ranked-venue papers
5as first author
15since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 8 · 4 first-author · 8 since 2021Artificial intelligence and machine learning · 7 · 2 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 4 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2026 NODiff: Neural Operator Diffusion for Multispectral Image Fusion
abstract
Pansharpening is a powerful technique for generating high-resolution multispectral (HRMS) images by fusing currently available image pairs of low-resolution multispectral (LRMS) and texture-rich panchromatic (PAN) data, effectively addressing the physical constraints of satellite sensors. While recent generative diffusion models have demonstrated impressive performance gains in this domain, their prohibitive computational demands and training costs hinder practicality in resource-constrained remote sensing satellite systems. In this work, we propose NODiff, a novel diffusion framework that replaces the conventional attention-based denoising backbone with a neural operator, seamlessly integrating operator learning and generative modeling into an efficient yet effective solution for pansharpening. In practice, we implement our approach through a two-stage learning paradigm: First, we pretrain the proposed Neural Operator-based diffusion model to learn the high-resolution texture priors essential for pansharpening. Afterward, we freeze the pretrained parameters, and design a lightweight conditional detail guidance adapter to enable efficient fine-tuning for generating desired HRMS images. Meanwhile, a time-aware low-rank adaptation is introduced to dynamically refine high-frequency details potentially affected by spectral mode truncation. Extensive experiments on multiple benchmark datasets demonstrate that NODiff achieves competitive pansharpening performance while significantly reducing training and inference costs. Beyond pansharpening, our method provides new insights into building resource-efficient generative models.
Junming Hou, Ran Ran 0001, Sixing Chen, Xiaofeng Cong, Junling Li, Liang-Jian Deng
AAAI1
2026 Brightness-Aware Synthetic-to-Real Learning for Nighttime Hazy Image Enhancement
abstract
Nighttime hazy vision is severely limited by the presence of haze and multi-colored light sources. Different from the daytime image dehazing task which has been widely studied, less progress has been made in nighttime image dehazing. In this paper, through extensive analysis and experimentation, we find that game engine simulations offer strong real-world generalization but suffer from unrealistic brightness. To tackle this, we introduce a three-step, brightness-aware synthetic-to-real learning approach. First, we use supervised learning to train a spatial-frequency network (SFN) on synthetic data to produce pseudo-labels. With these pseudo-labels, we develop a semi-supervised dehazing model (SFN+) that minimizes domain discrepancy through a brightness consistency loss applied to local windows. Building on SFN+, we fine-tune the model for better vision using a relative brightness improvement strategy that accounts for color shifts from lighting and brightness shifts during enhancement (SFN++). Experiments on popular benchmark datasets confirm our method's superiority over state-of-the-art approaches.
Jie Gui, Xiaofeng Cong, Yu-Xin Zhang 0004, Junming Hou, Dacheng Tao
IEEE Trans. Pattern Anal. Mach. Intell.4
2025 Binarized Neural Network for Multi-spectral Image Fusion
abstract
Pan-sharpening technology refers to generating a high-resolution (HR) multi-spectral (MS) image with broad applications by fusing a low-resolution (LR) MS image and HR panchromatic (PAN) image. While deep learning approaches have shown impressive performance in pan-sharpening, they generally require extensive hardware with high memory and computational power, limiting their deployment on resource-constrained satellites. In this study, we investigate the use of binary neural networks (BNNs) for pan-sharpening and observe that binarization leads to distinct information degradation across different frequency components of an image. Building on this insight, we propose a novel binary pan-sharpening network, termed BNNPan, structured around the Prior-Integrated Binary Frequency (PIBF) module that features three key ingredients: Binary Wavelet Transform Convolution, Latent Diffusion Prior Compensation, and Channel-wise Distribution Calibration. Specifically, the first decomposes input features into distinct frequency components using Wavelet Transform, then applies a "divide-and-conquer" strategy to optimize binary feature learning for each component, informed by the corresponding full-precision residual statistics. The second integrates a latent diffusion prior to compensate for compromised information during binarization, while the third performs channel-wise calibration to further refine feature representation. Our BNNPan, developed with the proposed techniques, achieves promising pan-sharpening performance on multiple remote sensing datasets, surpassing state-of-the-art binarization algorithms.
Junming Hou, Ran Ran 0001, Xiaofeng Cong, Jian Wei You, Liang-Jian Deng
CVPR1
2025 Physics-informed Neural Operator for Pansharpening
abstract
Over the past decades, pansharpening has contributed greatly to numerous remote sensing applications, with methods evolving from theoretically grounded models to deep learning approaches and their hybrids. Though promising, existing methods rarely address pansharpening through the lens of underlying physical imaging processes. In this work, we revisit the spectral imaging mechanism and propose a novel physics‐informed neural operator framework for pansharpening, termed PINO, which faithfully models the end‐to‐end electro‐optical sensor process. Specifically, PINO operates as: (1) First, a spatial-spectral encoder pair is introduced to aggregate multi-granularity high-resolution panchromatic (PAN) and low-resolution multispectral (LRMS) features. (2) Subsequently, an iterative neural integral process utilizes these fused spatial-spectral characteristics to learn a continuous radiance field $L_i(x, y, \lambda)$ over spatial coordinates and wavelength, effectively emulating band-wise spectral integration. (3) Finally, the learned radiance field is modulated by the sensor’s spectral responsivity $R_b(\lambda)$ to produce physically consistent spatial–spectral fusion products. This physics-grounded fusion paradigm offers a principled solution for reconstructing high-resolution multispectral and hyperspectral images in accordance with sensor imaging physics, effectively harnessing the unique advantages of spectral data to better uncover real-world characteristics. Experiments on multiple benchmark datasets show that our method surpasses state-of-the-art fusion algorithms, achieving reduced spectral aberrations and finer spatial textures. Furthermore, extension to hyperspectral (HS) data demonstrates its generalizability and universality. The code will be available upon potential acceptance.
Junming Hou, Chenxu Wu, Xiaofeng Cong, Shangqi Deng, Junling Li, Liang-Jian Deng
NeurIPS2
2025 Bilateral Adaptive Evolution Transformer for Multispectral Image Fusion
abstract
Pansharpening is the critical technology for generating high-resolution (HR) multispectral (MS) images by learning the cross-modality complementary representations between the panchromatic (PAN) images and low-resolution (LR) MS images. Though methods based on convolutional neural networks (CNNs) have dominated the pansharpening community, they still suffer from the limited global modeling capability due to the inherent property of the convolutional operator. To remedy this common limitation, the transformer family has recently gained great popularity in this field. However, existing cascaded transformer designs inevitably introduce a heavy memory footprint and computational cost due to the dense dot-product self-attention (SA) computation. More importantly, these paradigms simply ignore the innate sparsity of remote sensing images, leading to information redundancy and a challenging optimization process. To alleviate these issues, we propose the bilateral adaptive evolution transformer (BAEFormer), which is built upon two core mechanisms: bilateral attention computation and adaptive attention evolution. Specifically, we first decompose the conventional quadratic complexity SA into linear-degree height and width computing at the first stage, respectively, which significantly reduces the computational complexity. Given the data-specific properties, furthermore, we devise a novel yet effective neighboring layer-dependent strategy to adaptively update the attention map of two spatial dimensions, thereby avoiding the repetitive SA computation while taking into account the dynamics toward the evolution of attention weights. Our model, called BAEFormer, outperforms other state-of-the-art pansharpening methods on various remote sensing datasets while showing fewer network parameters and computational requirements. The code is available athttps://github.com/coder-JMHou/BAEFormer.
Junming Hou, Chenxu Wu, Man Zhou 0003, Junling Li, Danfeng Hong
IEEE Trans. Geosci. Remote. Sens.1
2025 A General Cooperative Optimization Driven High-Frequency Enhancement Framework for Multispectral Image Fusion
abstract
Pan-sharpening essentially to boost the spatial resolution of a multispectral (MS) image guided by its paired panchromatic (PAN) image. In other words, this process intricately integrates the high-frequency components extracted from texture-rich PAN images into the low-resolution (LR) MS images, resulting in texture-rich MS images. Though existing deep learning (DL)-based techniques have made impressive performance compared with traditional algorithms, they still face challenges in accurately restoring high-frequency details in MS images, thus limiting overall pan-sharpening performance. In addition, reference high-resolution (HR) MS images are often underutilized, typically serving only as training labels. In this work, we present a general high-frequency enhancement framework for pan-sharpening, which is implemented through a cooperative optimization strategy using mutual information (MI) maximization and contrastive learning. Specifically, our model comprises two fundamental modules: the high-frequency feature alignment (HFFA) module and the high-frequency detail calibration (HFDC) module. The first employs MI maximization to align the high-frequency semantic statistical distribution between PAN images and reference HRMS images. The latter is designed to calibrate the high-frequency components of MS modality under the guidance of the PAN counterparts through the contrastive learning constraint, thereby producing more accurate high-frequency information on MS modality. By integrating the calibrated high-frequency features of MS modality and those of PAN modality, we can obtain a more comprehensive and precise high-frequency feature representation of these two modalities, facilitating the reconstruction of LRMS images. Our model, incorporating the aforementioned key elements, significantly surpasses other state-of-the-art (SOTA) techniques across multiple satellite datasets in both quantitative and qualitative experiments. Moreover, the real-world full-resolution and cross-sensor assessments testify to its exceptional generalization capabilities. The code is available athttps://github.com/Vcocoi/CONet.
Chentong Huang, Junming Hou, Chenxu Wu, Xiaofeng Cong, Man Zhou 0003, Junling Li, Danfeng Hong
IEEE Trans. Geosci. Remote. Sens.2
2025 Unrevealed Threats: Adversarial Robustness Analysis of Underwater Image Enhancement Models
abstract
Learning-based methods for underwater image enhancement (UWIE) have undergone extensive exploration. However, learning-based models are usually vulnerable to adversarial examples so as the UWIE models. To the best of our knowledge, there is no comprehensive study on the adversarial robustness of UWIE models, which indicates that UWIE models are potentially under the threat of adversarial attacks. In this paper, we propose a general adversarial attack protocol. We make a first attempt to conduct adversarial attacks on five well-designed UWIE models on three common underwater image benchmark datasets. Considering the scattering and absorption of light in the underwater environment, there exists a strong correlation between color correction and underwater image enhancement. On the basis of that, we also design two effective UWIE-oriented adversarial attack methods, Pixel Attack and Color Shift Attack targeting different color spaces. The results show that five models exhibit varying degrees of vulnerability to adversarial attacks and well-designed small perturbations on degraded images are capable of preventing UWIE models from generating enhanced results. In addition, we conduct adversarial training on these models and successfully mitigated the effectiveness of adversarial attacks. In summary, we reveal the adversarial vulnerability of UWIE models and propose a new evaluation dimension of UWIE models.
Siyu Zhai, Zhibo He, Xiaofeng Cong, Junming Hou, Jie Gui, Jian Wei You, Xin Gong 0001, James T. Kwok, Yuan Yan Tang
IEEE Trans. Multim.4
2024 Underwater Organism Color Fine-Tuning via Decomposition and Guidance
abstract
Due to the wavelength dependent light attenuation and scattering, the color of the underwater organism usually appears distorted. The existing underwater image enhancement methods mainly focus on designing networks capable of generating enhanced underwater organisms with fixed color. Due to the complexity of the underwater environment, ground truth labels are difficult to obtain, which results in the non-existence of perfect enhancement effects. Different from the existing methods, this paper proposes an algorithm with color enhancement and color fine-tuning (CECF) capabilities. The color enhancement behavior of CECF is the same as that of existing methods, aiming to restore the color of the distorted underwater organism. Beyond this general purpose, the color fine-tuning behavior of CECF can adjust the color of organisms in a controlled manner, which can generate enhanced organisms with diverse colors. To achieve this purpose, four processes are used in CECF. A supervised enhancement process learns the mapping from a distorted image to an enhanced image by the decomposition of color code. A self reconstruction process and a cross-reconstruction process are used for content-invariant learning. A color fine-tuning process is designed based on the guidance for obtaining various enhanced results with different colors. Experimental results have proven the enhancement ability and color fine-tuning ability of the proposed CECF. The source code is provided in https://github.com/Xiaofeng-life/CECF.
Xiaofeng Cong, Jie Gui, Junming Hou
AAAI3
2024 A Semi-Supervised Nighttime Dehazing Baseline with Spatial-Frequency Aware and Realistic Brightness Constraint
abstract
Existing research based on deep learning has extensively explored the problem of daytime image dehazing. However, few studies have considered the characteristics of nighttime hazy scenes. There are two distinctions between nighttime and daytime haze. First, there may be multiple active col-ored light sources with lower illumination intensity in night-time scenes, which may cause haze, glow and noise with localized, coupled and frequency inconsistent characteris-tics. Second, due to the domain discrepancy between simulated and real-world data, unrealistic brightness may occur when applying a dehazing model trained on simulated data to real-world data. To address the above two issues, we propose a semi-supervised model for real-world nighttime dehazing. First, the spatial attention and frequency spectrum filtering are implemented as a spatial-frequency do-main information interaction module to handle the first is-sue. Second, a pseudo-label-based retraining strategy and a local window-based brightness loss for semi-supervised training process is designed to suppress haze and glow while achieving realistic brightness. Experiments on public benchmarks validate the effectiveness of the proposed method and its superiority over state-of-the-art methods. The source code and Supplementary Materials are placed in the https://github.com/Xiaofeng-life/SFSNiD.
Xiaofeng Cong, Jie Gui, Jing Zhang 0037, Junming Hou, Hao Shen 0006
CVPR4
2024 Probing Synergistic High-Order Interaction in Infrared and Visible Image Fusion
abstract
Infrared and visible image fusion aims to generate a fused image by integrating and distinguishing complementary information from multiple sources. While the cross-attention mechanism with global spatial interactions appears promising, it only capture second-order spatial inter-actions, neglecting higher-order interactions in both spatial and channel dimensions. This limitation hampers the ex-ploitation of synergies between multi-modalities. To bridge this gap, we introduce a Synergistic High-order Interaction Paradigm (SHIP), designed to systematically investigate the spatial fine-grained and global statistics collaborations between infrared and visible images across two fundamental dimensions: 1) Spatial dimension: we construct spatial fine-grained interactions through element-wise multiplication, mathematically equivalent to global interactions, and then foster high-order formats by iteratively aggregating and evolving complementary information, enhancing both efficiency andflexibility; 2) Channel dimension: expanding on channel interactions with first-order statistics (mean), we devise high-order channel interactions to facilitate the discernment of inter-dependencies between source images based on global statistics. Harnessing high-order interactions significantly enhances our model's ability to exploit multi-modal synergies, leading to superior performance over state-of-the-art alternatives, as shown through comprehensive experiments across various benchmarks. Code is available at https://github.com/zheng980629/SHIP.
Naishan Zheng, Man Zhou 0003, Jie Huang 0017, Junming Hou, Haoying Li, Feng Zhao 0004
CVPR4
2024 Linearly-evolved Transformer for Pan-sharpening
Junming Hou, Zihan Cao, Naishan Zheng, Xuan Li 0012, Xiaofeng Cong, Danfeng Hong, Man Zhou 0003
ACM Multimedia1
2024 NC process information mining based optimization method of roughing tool sequence selection for pocket features
Changhong Xu, Shusheng Zhang, Jiachen Liang, Bian Rong, Junming Hou
Adv. Eng. Informatics5
2024 SSCAConv: Self-Guided Spatial-Channel Adaptive Convolution for Image Fusion
abstract
Pansharpening, which attempts to obtain a high-resolution multispectral (HR-MS) image by fusing a panchromatic (PAN) image with a low-resolution multispectral (LR-MS) image, is a critical yet difficult remote sensing image processing task. In this study, we present a novel convolution operation, self-guided spatial-channel adaptive convolution (SSCAConv), for pansharpening. Unlike the reported adaptive convolutions that only focus on spatial details, our SSCAConv also considers channel specificity by generating an individual convolution kernel for each channel patch according to its own content and supplements the interchannel information by introducing a global bias. We further apply the designed SSCAConv to a simple residual network architecture to construct the image fusion network (SSCANet). Experimental results show that SSCANet outperforms state-of-the-art (SOTA) pansharpening algorithms and achieves better generalization ability with fewer parameters. In addition, our network also yields the best results when extended to the hyperspectral image super-resolution (HISR) problem. The code is available athttps://github.com/Pluto-wei/SSCAConv.
Xiaoya Lu, Yu-Wei Zhuo, Hongming Chen 0003, Liang-Jian Deng, Junming Hou
IEEE Geosci. Remote. Sens. Lett.5
2024 Rethinking Pan-Sharpening via Spectral-Band Modulation
abstract
Pan-sharpening aims to super-resolve the low-resolution (LR) multispectral (MS) image under the guidance of a high-resolution (HR) panchromatic (PAN) image. Existing deep learning (DL)-based pan-sharpening methods usually adhere to a common philosophy of learning complementary information between MS and PAN images. Despite remarkable advances, few studies consider the band-private characteristics which differ greatly from band to band. An ideal MS image, however, is jointly determined by its diverse spectral bands, thus the accurate restoration of every band will benefit the pan-sharpening performance. In this work, we propose a novel yet effective solution to reconstruct the HRMS image by explicitly modulating every spectral band under the conditions of the PAN image. As a result, we design a spatially-adaptive spectral modulation network, dubbed SSMNet, which consists of three core designs: source-aware spectral modulator (SSM), cross-band information aggregation (CBIA) module, and cross-stage feature integration (CSFI) module. The first predicts a series of spatially-adaptive kernels to capture the local information of every spectral band. Followed by, the second is responsible for facilitating the information communication among various bands to guarantee continuous spectral representations. Furthermore, the third attends to integrate the cross-stage output features to produce the pan-sharpened result. In addition, we also introduce the histogram loss to constrain the band-wise distribution of the final fused products. Extensive experiments demonstrate that our SSMNet achieves favorable performance against other state-of-the-art (SOTA) methods on multiple satellite datasets. The code is available athttps://github.com/ez4lionky/SSMNet/.
Junming Hou, Xiaofeng Cong, Hao Shen 0006, Zhuochen Lou, Liang-Jian Deng, Jian Wei You
IEEE Trans. Geosci. Remote. Sens.2
2023 Bidomain Modeling Paradigm for Pansharpening
abstract
Pansharpening is a challenging low-level vision task whose aim is to learn the complementary representation between spectral information and spatial detail. Despite the remarkable progress, existing deep neural network (DNN) based pansharpening algorithms are still confronted with common limitations. 1) These methods rarely consider the local specificity of different spectral bands; 2) They often extract the global detail in the spatial domain, which ignore the task-related degradation, e.g., the down-sampling process of MS image, and also suffer from limited receptive field. In this work, we propose a novel bidomain modeling paradigm for pansharpening problem (dubbed as BiMPan), which takes into both local spectral specificity and global spatial detail. More specifically, we first customize the specialized source-discriminative adaptive convolution (SDAConv) for every spectral band instead of sharing the identical kernels across all bands like prior works. Then, we devise a novel Fourier global modeling module (FGMM), which is capable of embracing global information while benefiting the disentanglement of image degradation. By integrating the band-aware local feature and Fourier global detail from these two functional designs, we can fuse a texture-rich while visually pleasing high-resolution MS image. Extensive experiments demonstrate that the proposed framework achieves favorable performance against current state-of-the-art pansharpening methods. The code is available at https://github.com/coder-qicao/BiMPan.
Junming Hou, Ran Ran 0001, Che Liu 0004, Junling Li, Liang-Jian Deng
ACM Multimedia1
2019 À la Carte: Turning Historical Menu into Menu Network
Junming Hou, Keven Liu
TPDL2
2008 Research of collaborative process workflow modeling based on stochastic Petri nets
abstract
Collaborative process workflow is able to implement effective control and management for process planning flow and meet the process informationization needs of manufacturing enterprises. With the application of stochastic Petri nets (SPN) to the model of collaborative process workflow, the article proposes the method of collaborative process workflow modeling. The relationship of actual process flow and the elements of SPN is studied. This model is on flow described based on SPN, and also established by use case diagram and sequence diagram of UML to compensate for the lack of SPN. Finally, a prototype system is developed based on the model built by this article combining with the actual process situation of an enterprise.
Tianbiao Yu, Junming Hou, Wanshan Wang
CSCWD3