Chunming He

dblp:251/5104 · DBLP profile ↗
← Back
24ranked-venue papers
9as first author
23since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 15 · 9 first-author · 15 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 2 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 4 since 2021Computer networks · 1
YearPublicationVenuePosition
2026 Vision-based lightweight digital twin for real-time semantic state aggregation and representation in large-scale transportation networks
Meng Feng, Chunming He, Lianbao Yang, Pengfei Hu 0003, Zhiyou Cai, Zhengxiang Yan
Expert Syst. Appl.2
2026 An Unsupervised Image Dehazing With Scene Geometry Prior for Road Traffic Scenarios
abstract
Despite advances in single image dehazing, robust dehazing for real-world road traffic scenes remains challenging due to scarce paired data, traffic-specific geometry, and real-time constraints. To address this issue, we propose a novel image prior for road traffic scenes, termed scene geometry prior (SGP), which leverages depth cues derived from vanishing point (VP) to provide geometry-aware guidance and reduce reliance on paired training data. Our SGP comprises two components: a global SGP (G-SGP) that captures the global geometric distribution and a non-local SGP (NL-SGP) that corrects the errors, among obstructions belonging to the same category, in captured global distribution. Building on the proposed prior, we develop a lightweight and unsupervised road traffic image dehazing network (RTDnet). It consists of a main sub-network guided by the G-SGP to reconstruct the haze-free image, alongside two auxiliary sub-networks that leverage the NL-SGP and VP information to respectively estimate transmission map, and atmospheric light. During training, we introduce an atmospheric scattering model (ASM)-driven mutual-boost learning mechanism (ASM-ML), which is rooted in Bayesian theory and effectively integrates the strengths of different priors without mutual interference while distilling ASM-based physical knowledge into each sub-network. By coupling SGP with ASM-ML, RTDnet can be trained without paired traffic data by exploiting traffic-specific geometry, whose accurate guidance reduces the reliance on large model capacity and enables lightweight real-time deployment. Experiments demonstrate that our RTDnet surpasses state-of-the-art competitors in terms of restoration quality, efficiency, and model size. Moreover, its robust dehazing performance benefits downstream tasks operating in hazy conditions.
Mingye Ju, Tianyi Lyu, Chunming He, Qingshan Liu 0001, Kai-Kuang Ma
IEEE Trans. Image Process.3
2025 MultiBooth: Towards Generating All Your Concepts in an Image from Text
abstract
This paper introduces MultiBooth, a method that generates images from texts containing various concepts from users.Despite diffusion models bringing significant advancements for customized text-to-image generation, existing methods often struggle with multi-concept scenarios due to low concept fidelity and high inference cost. MultiBooth addresses these issues by dividing the multi-concept generation process into two phases: a single-concept learning phase and a multi-concept integration phase. During the single-concept learning phase, we employ a multi-modal image encoder and an efficient concept encoding technique to learn a concise and discriminative representation for each concept. In the multi-concept integration phase, we use bounding boxes to define the generation area for each concept within the cross-attention map. This method enables the creation of individual concepts within their specified regions, thereby facilitating the formation of multi-concept images. This strategy not only improves concept fidelity but also reduces additional inference cost. MultiBooth surpasses various baselines in both qualitative and quantitative evaluations, showcasing its superior performance and computational efficiency.
Chenyang Zhu 0007, Kai Li 0012, Yue Ma 0016, Chunming He, Xiu Li 0001
AAAI4
2025 Reti-Diff: Illumination Degradation Image Restoration with Retinex-based Latent Diffusion Model
abstract
Illumination degradation image restoration (IDIR) techniques aim to improve the visibility of degraded images and mitigate the adverse effects of deteriorated illumination. Among these algorithms, diffusion-based models (DM) have shown promising performance but are often burdened by heavy computational demands and pixel misalignment issues when predicting the image-level distribution. To tackle these problems, we propose to leverage DM within a compact latent space to generate concise guidance priors and introduce a novel solution called Reti-Diff for the IDIR task. Specifically, Reti-Diff comprises two significant components: the Retinex-based latent DM (RLDM) and the Retinex-guided transformer (RGformer). RLDM is designed to acquire Retinex knowledge, extracting reflectance and illumination priors to facilitate detailed reconstruction and illumination correction. RGformer subsequently utilizes these compact priors to guide the decomposition of image features into their respective reflectance and illumination components. Following this, RGformer further enhances and consolidates these decomposed features, resulting in the production of refined images with consistent content and robustness to handle complex degradation scenarios. Extensive experiments demonstrate that Reti-Diff outperforms existing methods on three IDIR tasks, as well as downstream applications.
Chunming He, Chengyu Fang 0001, Yulun Zhang 0001, Longxiang Tang, Jinfa Huang, Kai Li 0012, Zhenhua Guo 0001, Xiu Li 0001, Sina Farsiu
ICLR1
2025 RUN: Reversible Unfolding Network for Concealed Object Segmentation
abstract
Concealed object segmentation (COS) is a challenging problem that focuses on identifying objects that are visually blended into their background. Existing methods often employ reversible strategies to concentrate on uncertain regions but only focus on the mask level, overlooking the valuable of the RGB domain. To address this, we propose a Reversible Unfolding Network (RUN) in this paper. RUN formulates the COS task as a foreground-background separation process and incorporates an extra residual sparsity constraint to minimize segmentation uncertainties. The optimization solution of the proposed model is unfolded into a multistage network, allowing the original fixed parameters to become learnable. Each stage of RUN consists of two reversible modules: the Segmentation-Oriented Foreground Separation (SOFS) module and the Reconstruction-Oriented Background Extraction (ROBE) module. SOFS applies the reversible strategy at the mask level and introduces Reversible State Space to capture non-local information. ROBE extends this to the RGB domain, employing a reconstruction network to address conflicting foreground and background regions identified as distortion-prone areas, which arise from their separate estimation by independent modules. As the stages progress, RUN gradually facilitates reversible modeling of foreground and background in both the mask and RGB domains, reducing false-positive and false-negative regions. Extensive experiments demonstrate the superior performance of RUN and underscore the promise of unfolding-based frameworks for COS and other high-level vision tasks. Code is available at https://github.com/ChunmingHe/RUN.
Chunming He, Rihan Zhang, Fengyang Xiao, Chengyu Fang 0001, Longxiang Tang, Yulun Zhang 0001, Linghe Kong, Deng-Ping Fan, Kai Li 0012, Sina Farsiu
ICML1
2025 Pan-LUT: Efficient Pan-sharpening via Learnable Look-Up Tables
abstract
Recently, deep learning-based pan-sharpening algorithms have achieved notable advancements over traditional methods. However, deep learning-based methods incur substantial computational overhead during inference, especially with large images. This excessive computational demand limits the applicability of these methods in real-world scenarios, particularly in the absence of dedicated computing devices such as GPUs and TPUs. To address these challenges, we propose Pan-LUT, a novel learnable look-up table (LUT) framework for pan-sharpening that strikes a balance between performance and computational efficiency for large remote sensing images. Our method makes it possible to process 15K$\times$15K remote sensing images on a 24GB GPU. To finely control the spectral transformation, we devise the PAN-guided look-up table (PGLUT) for channel-wise spectral mapping. To effectively capture fine-grained spatial details, we introduce the spatial details look-up table (SDLUT). Furthermore, to adaptively aggregate channel information for generating high-resolution multispectral images, we design an adaptive output look-up table (AOLUT). Our model contains fewer than 700K parameters and processes a 9K$\times$9K image in under 1 ms using one RTX 2080 Ti GPU, demonstrating significantly faster performance compared to other methods. Experiments reveal that Pan-LUT efficiently processes large remote sensing images in a lightweight manner, bridging the gap to real-world applications. Furthermore, our model surpasses SOTA methods in full-resolution scenes under real-world conditions, highlighting its effectiveness and efficiency. We also extend our method to general image fusion tasks.
Zhongnan Cai, Yingying Wang 0005, Hui Zheng 0003, Panwang Pan, Zixu Lin, Ge Meng, Chenxin Li, Chunming He, Jiaxin Xie, Yunlong Lin, Junbin Lu, Yue Huang 0001, Xinghao Ding
NeurIPS8
2025 Segment Concealed Objects With Incomplete Supervision
abstract
Incompletely-Supervised Concealed Object Segmentation (ISCOS) involves segmenting objects that seamlessly blend into their surrounding environments, utilizing incompletely annotated data, such as weak and semi-annotations, for model training. This task remains highly challenging due to (1) the limited supervision provided by the incompletely annotated training data, and (2) the difficulty of distinguishing concealed objects from the background, which arises from the intrinsic similarities in concealed scenarios. In this paper, we introduce the first unified method for ISCOS to address these challenges. To tackle the issue of incomplete supervision, we propose a unified mean-teacher framework, SEE, that leverages the vision foundation model, "Segment Anything Model (SAM)", to generate pseudo-labels using coarse masks produced by the teacher model as prompts. To mitigate the effect of low-quality segmentation masks, we introduce a series of strategies for pseudo-label generation, storage, and supervision. These strategies aim to produce informative pseudo-labels, store the best pseudo-labels generated, and select the most reliable components to guide the student model, thereby ensuring robust network training. Additionally, to tackle the issue of intrinsic similarity, we design a hybrid-granularity feature grouping module that groups features at different granularities and aggregates these results. By clustering similar features, this module promotes segmentation coherence, facilitating more complete segmentation for both single-object and multiple-object images. We validate the effectiveness of our approach across multiple ISCOS tasks, and experimental results demonstrate that our method achieves state-of-the-art performance. Furthermore, SEE can serve as a plug-and-play solution, enhancing the performance of existing models.
Chunming He, Kai Li 0012, Yachao Zhang 0001, Ziyun Yang, Youwei Pang, Longxiang Tang, Chengyu Fang 0001, Yulun Zhang 0001, Linghe Kong, Xiu Li 0001, Sina Farsiu
IEEE Trans. Pattern Anal. Mach. Intell.1
2025 Diffusion Models in Low-Level Vision: A Survey
abstract
Deep generative models have gained considerable attention in low-level vision tasks due to their powerful generative capabilities. Among these, diffusion model-based approaches, which employ a forward diffusion process to degrade an image and a reverse denoising process for image generation, have become particularly prominent for producing high-quality, diverse samples with intricate texture details. Despite their widespread success in low-level vision, there remains a lack of a comprehensive, insightful survey that synthesizes and organizes the advances in diffusion model-based techniques. To address this gap, this paper presents the first comprehensive review focused on denoising diffusion models applied to low-level vision tasks, covering both theoretical and practical contributions. We outline three general diffusion modeling frameworks and explore their connections with other popular deep generative models, establishing a solid theoretical foundation for subsequent analysis. We then categorize diffusion models used in low-level vision tasks from multiple perspectives, considering both the underlying framework and the target application. Beyond natural image processing, we also summarize diffusion models applied to other low-level vision domains, including medical imaging, remote sensing, and video processing. Additionally, we provide an overview of widely used benchmarks and evaluation metrics in low-level vision tasks. Our review includes an extensive evaluation of diffusion model-based techniques across six representative tasks, with both quantitative and qualitative analysis. Finally, we highlight the limitations of current diffusion models and propose four promising directions for future research. This comprehensive review aims to foster a deeper understanding of the role of denoising diffusion models in low-level vision.
Chunming He, Yuqi Shen, Chengyu Fang 0001, Fengyang Xiao, Longxiang Tang, Yulun Zhang 0001, Wangmeng Zuo, Zhenhua Guo 0001, Xiu Li 0001
IEEE Trans. Pattern Anal. Mach. Intell.1
2025 All-Inclusive Image Enhancement for Degraded Images Exhibiting Low-Frequency Corruption
abstract
In this paper, a novel image enhancement method, called the all-inclusive image enhancement (AIIE), is proposed that can effectively enhance the degraded images for improving the visibility of image content. These imageries were acquired under various types of weather conditions such as haze, low-light, underwater, and sandstorm, etc. One commonality shared by this class of noise is that the resulted degradations on visual quality or visibility are caused by low-frequency interference. Existing image enhancement methods lack the ability to deal with all types of degradations from this class, while our proposed AIIE offers a unified treatment for them. To achieve this goal, a statistical property is obtained from the study of the discrete cosine transform (DCT) of 1,000 high- and 1000 low-quality images on their DCT domains. It shows that the normalized DCT coefficients (between 0 and 1) of high-quality images has about 95% fall in the interval [0, 0.2]; for low-quality images, almost all the coefficients are in the same interval. This fundamental property, called the DCT prior (DCT-P), is instrumental to the development of our AIIE algorithm proposed in this paper. Since the proposed DCT-P delineates the attributes of high- and low-quality images clearly, it becomes a highly effective ‘tool’ to convert low-quality images to its enhanced version. Extensive experimental results have clearly validated the superior performance of the AIIE conducted on different types of deteriorated images in terms of visual quality and efficiency as well as significant advantages on computational complexity, which is essential for real-time applications.
Mingye Ju, Chunming He, Can Ding 0002, Wenqi Ren, Lin Zhang 0014, Kai-Kuang Ma
IEEE Trans. Circuits Syst. Video Technol.2
2024 Mind the Interference: Retaining Pre-trained Knowledge in Parameter Efficient Continual Learning of Vision-Language Models
Longxiang Tang, Zhuotao Tian, Kai Li 0012, Chunming He, Hantao Zhou, Hengshuang Zhao, Xiu Li 0001, Jiaya Jia
ECCV (36)4
2024 Strategic Preys Make Acute Predators: Enhancing Camouflaged Object Detectors by Generating Camouflaged Objects
abstract
Camouflaged object detection (COD) is the challenging task of identifying camouflaged objects visually blended into surroundings. Albeit achieving remarkable success, existing COD detectors still struggle to obtain precise results in some challenging cases. To handle this problem, we draw inspiration from the prey-vs-predator game that leads preys to develop better camouflage and predators to acquire more acute vision systems and develop algorithms from both the prey side and the predator side. On the prey side, we propose an adversarial training framework, Camouflageator, which introduces an auxiliary generator to generate more camouflaged objects that are harder for a COD method to detect. Camouflageator trains the generator and detector in an adversarial way such that the enhanced auxiliary generator helps produce a stronger detector. On the predator side, we introduce a novel COD method, called Internal Coherence and Edge Guidance (ICEG), which introduces a camouflaged feature coherence module to excavate the internal coherence of camouflaged objects, striving to obtain more complete segmentation results. Additionally, ICEG proposes a novel edge-guided separated calibration module to remove false predictions to avoid obtaining ambiguous boundaries. Extensive experiments show that ICEG outperforms existing COD detectors and Camouflageator is flexible to improve various COD detectors, including ICEG, which brings state-of-the-art COD performance.
Chunming He, Kai Li 0012, Yachao Zhang 0001, Yulun Zhang 0001, Chenyu You, Zhenhua Guo 0001, Xiu Li 0001, Martin Danelljan, Fisher Yu 0001
ICLR1
2024 Real-world Image Dehazing with Coherence-based Pseudo Labeling and Cooperative Unfolding Network
abstract
Real-world Image Dehazing (RID) aims to alleviate haze-induced degradation in real-world settings. This task remains challenging due to the complexities in accurately modeling real haze distributions and the scarcity of paired real-world data. To address these challenges, we first introduce a cooperative unfolding network that jointly models atmospheric scattering and image scenes, effectively integrating physical knowledge into deep networks to restore haze-contaminated details. Additionally, we propose the first RID-oriented iterative mean-teacher framework, termed the Coherence-based Label Generator, to generate high-quality pseudo labels for network training. Specifically, we provide an optimal label pool to store the best pseudo-labels during network training, leveraging both global and local coherence to select high-quality candidates and assign weights to prioritize haze-free regions. We verify the effectiveness of our method, with experiments demonstrating that it achieves state-of-the-art performance on RID tasks. Code will be available at https://github.com/cnyvfang/CORUN-Colabator.
Chengyu Fang 0001, Chunming He, Fengyang Xiao, Yulun Zhang 0001, Longxiang Tang, Yuelin Zhang, Kai Li 0012, Xiu Li 0001
NeurIPS2
2024 HQG-Net: Unpaired Medical Image Enhancement With High-Quality Guidance
abstract
Unpaired medical image enhancement (UMIE) aims to transform a low-quality (LQ) medical image into a high-quality (HQ) one without relying on paired images for training. While most existing approaches are based on Pix2Pix/CycleGAN and are effective to some extent, they fail to explicitly use HQ information to guide the enhancement process, which can lead to undesired artifacts and structural distortions. In this article, we propose a novel UMIE approach that avoids the above limitation of existing methods by directly encoding HQ cues into the LQ enhancement process in a variational fashion and thus model the UMIE task under the joint distribution between the LQ and HQ domains. Specifically, we extract features from an HQ image and explicitly insert the features, which are expected to encode HQ cues, into the enhancement network to guide the LQ enhancement with the variational normalization module. We train the enhancement network adversarially with a discriminator to ensure the generated HQ image falls into the HQ domain. We further propose a content-aware loss to guide the enhancement process with wavelet-based pixel-level and multiencoder-based feature-level constraints. Additionally, as a key motivation for performing image enhancement is to make the enhanced images serve better for downstream tasks, we propose a bi-level learning scheme to optimize the UMIE task and downstream tasks cooperatively, helping generate HQ images both visually appealing and favorable for downstream tasks. Experiments on three medical datasets verify that our method outperforms existing techniques in terms of both enhancement quality and downstream task performance. The code and the newly collected datasets are publicly available at https://github.com/ChunmingHe/HQG-Net.
Chunming He, Kai Li 0012, Guoxia Xu, Jiangpeng Yan, Longxiang Tang, Yulun Zhang 0001, Yaowei Wang 0001, Xiu Li 0001
IEEE Trans. Neural Networks Learn. Syst.1
2024 DM-Fusion: Deep Model-Driven Network for Heterogeneous Image Fusion
abstract
Heterogeneous image fusion (HIF) is an enhancement technique for highlighting the discriminative information and textural detail from heterogeneous source images. Although various deep neural network-based HIF methods have been proposed, the most widely used single data-driven manner of the convolutional neural network always fails to give a guaranteed theoretical architecture and optimal convergence for the HIF problem. In this article, a deep model-driven neural network is designed for this HIF problem, which adaptively integrates the merits of model-based techniques for interpretability and deep learning-based methods for generalizability. Unlike the general network architecture as a black box, the proposed objective function is tailored to several domain knowledge network modules to model the compact and explainable deep model-driven HIF network termed DM-fusion. The proposed deep model-driven neural network shows the feasibility and effectiveness of three parts, the specific HIF model, an iterative parameter learning scheme, and data-driven network architecture. Furthermore, the task-driven loss function strategy is proposed to achieve feature enhancement and preservation. Numerous experiments on four fusion tasks and downstream applications illustrate the advancement of DM-fusion compared with the state-of-the-art (SOTA) methods both in fusion quality and efficiency. The source code will be available soon.
Guoxia Xu, Chunming He, Hao Wang 0003, Hu Zhu, Weiping Ding 0001
IEEE Trans. Neural Networks Learn. Syst.2
2023 Camouflaged Object Detection with Feature Decomposition and Edge Reconstruction
abstract
Camouflaged object detection (COD) aims to address the tough issue of identifying camouflaged objects visually blended into the surrounding backgrounds. COD is a challenging task due to the intrinsic similarity of camouflaged objects with the background, as well as their ambiguous boundaries. Existing approaches to this problem have developed various techniques to mimic the human visual system. Albeit effective in many cases, these methods still struggle when camouflaged objects are so deceptive to the vision system. In this paper, we propose the FEature Decomposition and Edge Reconstruction (FEDER) model for COD. The FEDER model addresses the intrinsic similarity of foreground and background by decomposing the features into different frequency bands using learnable wavelets. It then focuses on the most informative bands to mine subtle cues that differentiate foreground and background. To achieve this, a frequency attention module and a guidance-based feature aggregation module are developed. To combat the ambiguous boundary problem, we propose to learn an auxiliary edge reconstruction task alongside the COD task. We design an ordinary differential equation-inspired edge reconstruction module that generates exact edges. By learning the auxiliary task in conjunction with the COD task, the FEDER model can generate precise prediction maps with accurate object boundaries. Experiments show that our FEDER model significantly outperforms state-of-the-art methods with cheaper computational and memory costs. The code will be available at https://github.com/ChunmingHe/FEDER.
Chunming He, Kai Li 0012, Yachao Zhang 0001, Longxiang Tang, Yulun Zhang 0001, Zhenhua Guo 0001, Xiu Li 0001
CVPR1
2023 Towards Realizing the Value of Labeled Target Samples: A Two-Stage Approach for Semi-Supervised Domain Adaptation
abstract
Semi-Supervised Domain Adaptation (SSDA) is a recently emerging research topic that extends from the widely-investigated Unsupervised Domain Adaptation (UDA) by further having a few target samples labeled, i.e., the model is trained with labeled source samples, unlabeled target samples as well as a few labeled target samples. Compared with UDA, the key to SSDA lies how to most effectively utilize the few labeled target samples. Existing SSDA approaches simply merge the few precious labeled target samples into vast labeled source samples or further align them, which dilutes the value of labeled target samples and thus still obtains a biased model. To remedy this, in this paper, we propose to decouple SSDA as an UDA problem and a semi-supervised learning problem where we first learn an UDA model using labeled source and unlabeled target samples and then adapt the learned UDA model in a semi-supervised way using labeled and unlabeled target samples. By utilizing the labeled source samples and target samples separately, the bias problem can be well mitigated. We further propose a consistency learning based mean teacher model to effectively adapt the learned UDA model using labeled and unlabeled target samples. Experiments show our approach outperforms existing methods.
Mengqun Jin, Kai Li 0012, Shuyan Li, Chunming He, Xiu Li 0001
ICASSP4
2023 Degradation-Resistant Unfolding Network for Heterogeneous Image Fusion
abstract
Heterogeneous image fusion (HIF) techniques aim to enhance image quality by merging complementary information from images captured by different sensors. Among these algorithms, deep unfolding network (DUN)-based methods achieve promising performance but still suffer from two issues: they lack a degradation-resistant-oriented fusion model and struggle to adequately consider the structural properties of DUNs, making them vulnerable to degradation scenarios. In this paper, we propose a Degradation-Resistant Unfolding Network (DeRUN) for the HIF task to generate high-quality fused images even in degradation scenarios. Specifically, we introduce a novel HIF model for degradation resistance and derive its optimization procedures. Then, we incorporate the optimization unfolding process into the proposed DeRUN for end-to-end training. To ensure the robustness and efficiency of DeRUN, we employ a joint constraint strategy and a lightweight partial weight sharing module. To train DeRUN, we further propose a gradient direction-based entropy loss with powerful texture representation capacity. Extensive experiments show that DeRUN significantly outperforms existing methods on four HIF tasks, as well as downstream applications, with cheaper computational and memory costs.
Chunming He, Kai Li 0012, Guoxia Xu, Yulun Zhang 0001, Runze Hu, Zhenhua Guo 0001, Xiu Li 0001
ICCV1
2023 Source-Free Domain Adaptive Fundus Image Segmentation with Class-Balanced Mean Teacher
Longxiang Tang, Kai Li 0012, Chunming He, Yulun Zhang 0001, Xiu Li 0001
MICCAI (1)3
2023 Weakly-Supervised Concealed Object Segmentation with SAM-based Pseudo Labeling and Multi-scale Feature Grouping
abstract
Weakly-Supervised Concealed Object Segmentation (WSCOS) aims to segment objects well blended with surrounding environments using sparsely-annotated data for model training. It remains a challenging task since (1) it is hard to distinguish concealed objects from the background due to the intrinsic similarity and (2) the sparsely-annotated training data only provide weak supervision for model learning. In this paper, we propose a new WSCOS method to address these two challenges. To tackle the intrinsic similarity challenge, we design a multi-scale feature grouping module that first groups features at different granularities and then aggregates these grouping results. By grouping similar features together, it encourages segmentation coherence, helping obtain complete segmentation results for both single and multiple-object images. For the weak supervision challenge, we utilize the recently-proposed vision foundation model, ``Segment Anything Model (SAM)'', and use the provided sparse annotations as prompts to generate segmentation masks, which are used to train the model. To alleviate the impact of low-quality segmentation masks, we further propose a series of strategies, including multi-augmentation result ensemble, entropy-based pixel-level weighting, and entropy-based image-level selection. These strategies help provide more reliable supervision to train the segmentation model. We verify the effectiveness of our method on various WSCOS tasks, and experiments demonstrate that our method achieves state-of-the-art performance on these tasks.
Chunming He, Kai Li 0012, Yachao Zhang 0001, Guoxia Xu, Longxiang Tang, Yulun Zhang 0001, Zhenhua Guo 0001, Xiu Li 0001
NeurIPS1
2023 IVF-Net: An Infrared and Visible Data Fusion Deep Network for Traffic Object Enhancement in Intelligent Transportation Systems
abstract
Infrared and visible data fusion (IVF) aims to generate a fused output that simultaneously highlights salient thermal radiation features and preserves texture information, which can not only grasp the necessary information for traffic movement, but also highlight the invisible objects that need to be dodged in intelligent transportation system (ITS). Therefore, IVF is capable of improving the environmental perception ability for various challenging traffic situations, e.g., foggy scenarios, rainy environments, and low-light illumination. However, current available IVF algorithms cannot offer a theoretical manner to integrate a priori knowledge and the network structure into a unified model. Moreover, they always fail to handle infrared and visible data pairs with different resolutions, which is a common occurrence in real ITS scenarios. To this end, this study develops a novel model-inspired unsupervised network termed IVF-Net. Specifically, an enhanced IVF model (IVFM), which pays more attention on detailed texture information and salient objects, is first established. According to proximal gradient theory, then we map this model into a deep network with learnable feature extraction parameters, aiming to draw on the strengths of the fusion model and deep learning to better describe the IVF task. Finally, a multiple task-driven loss function is designed to train the mapped network. Unlike previous work, our IVF-Net is motivated by IVFM, each layer in which has a semantic interpretability and a clear mission, thereby leading to a significantly enhanced fusion effect. Another advantage is that it is only composed of simple convolution-based structures, which ensures its lightweight and efficiency. Experiments demonstrate that IVF-Net can have a stronger ability to capture the key traffic information and highlight the salient feature of imperceptible objects, which makes it an excellent candidate to improve the reliability of subsequent applications in ITS.
Mingye Ju, Chunming He, Juping Liu, Bin Kang, Jian Su 0001, Dengyin Zhang
IEEE Trans. Intell. Transp. Syst.2
2022 Multi-modal sequence learning for Alzheimer's disease progression prediction with incomplete variable-length longitudinal data
Lei Xu 0028, Chunming He, Jun Wang 0024, Changqing Zhang 0002, Feiping Nie 0001, Lei Chen 0011
Medical Image Anal.3
2022 PcGAN: A Noise Robust Conditional Generative Adversarial Network for One Shot Learning
abstract
Traffic sign classification plays a vital role in autonomous vehicles for its powerful capability in information representation. However, the low-quality data of traffic signs captured by in-vehicle cameras often inevitably bring inherent challenges to the one-shot classification task. Apart from the problem of data degradation, learning-based classification techniques of real traffic signs also come across the challenges of intra-class and inter-class data imbalance from the training data. To overcome the aforementioned problems, we propose an end-to-end degradation robust deep model, termed PcGAN, to classify traffic signs in a manner of few-shot learning. The proposed PcGAN models the joint distribution between the degraded traffic signal data and the corresponding prototypes from both degradation removal and generation perspectives by two alternating optimized modules, which ensures the generalization of the learned embedding of latent space for novel tasks. A multi-task loss function is designed to improve the robustness of PcGAN. Numerous experiments comprehensively demonstrate that the accuracy of our proposed PcGAN is improved by 5% compared with other state-of-the-art (SOTA) approaches in few-shot classification.
Lizhen Deng, Chunming He, Guoxia Xu, Hu Zhu, Hao Wang 0003
IEEE Trans. Intell. Transp. Syst.2
2021 Vector co-occurrence morphological edge detection for colour image
abstract
Abstract Morphological edge detection is a principal component in pattern recognition and machine vision. Traditional edge detection operators only take pixel mutual into consideration. However, the edges are influenced not only by pixel mutual but also by the boundary characteristics. Here, the vector co‐occurrence morphological edge detection operator is proposed, which takes the pixel and boundary information both into consideration. The vector co‐occurrence algorithm is exploited to resist the influence of the noise points and detect the edges from the colour image rather than the grey image. And, we lead to define a precise definition of the manner of sorting high‐dimensional data for the colour image. The experiment results always illustrate the advancement and practicability of our methods against the baseline method. In terms of experiments, the BSDS500 dataset is introduced to compare and analyse with other algorithms. Based on the standard benchmark index evaluation in the BSDS500 dataset, the ODS and AP of various algorithms are compared and analysed.
Chunming He, Yu-Feng Yu 0001, Guoxia Xu, Hu Zhu, Lizhen Deng
IET Image Process.2
2020 Software-Defined Edge Computing (SDEC): Principle, Open IoT System Architecture, Applications, and Challenges
abstract
Edge computing is a bridge for realizing the convergence between physical space and cyber space in the Internet of Things (IoT) paradigm. Large numbers of physical objects produce a huge amount of data that needs to be efficiently processed in the edge side. This situation urgently requires novel ideas and framework in the design and management of edge computing to improve and enhance its performance. In this article, we propose an approach and principle of software-defined edge computing (SDEC) from the perspective of cyber-physical mapping, where the ultimate goal is to achieve a highly automatic and intelligent edge computing system. The SDEC can also help realize flexible management and intelligent collaboration among various edge hardware resources and services by way of software. To this end, we design an SDEC-based open IoT system architecture which decouples upper level IoT applications from the underlying physical edge resources and builds dynamically reconfigurable smart edge services. The software-definition mechanism of the SDEC platform is proposed to introduce the detailed processes that the underlying physical devices are defined in the form of software. We also describe an illustrative application case about smart factory to present the practical effectiveness of the proposed scheme. Finally, we outline several challenges which are worthy of in-depth study and research. The SDEC paradigm can share, reuse, recombine, and reconfigure edge resources and services so that the overall service capability of the edge side can be improved.
Pengfei Hu 0003, Wai Chen, Chunming He, Huansheng Ning
IEEE Internet Things J.3