EDBT 2026 Demo / reviewers in the wild / expert
Yuchao Feng
dblp:157/3981
· DBLP profile ↗
28ranked-venue papers
10as first author
28since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 16 · 5 first-author · 16 since 2021Artificial intelligence and machine learning · 9 · 9 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 5 first-author · 8 since 2021Computer networks · 2 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | 2D-Slice and 3D-Cube Mamba Network for Snapshot Spectral Compressive ImagingabstractHyperspectral image (HSI) reconstruction algorithms are fundamental to coded aperture snapshot spectral imaging (CASSI) systems. Recently, deep unfolding networks (DUNs) have emerged as a dominant solution, seamlessly combining traditional optimization frameworks with the strengths of deep learning. Among these, Mamba stands out as a prominent method for modeling long-range dependencies. However, its reliance on one-dimensional (1D) spatial scanning often compromises spectral consistency and spatial coherence, leading to misalignment of neighboring pixels within sequences. To address these limitations, we propose a novel multi-view framework based on 2D-slice modeling, which ensures spatial-spectral continuity in 1D sequences while maintaining computational efficiency. Furthermore, motivated by the need for precise local patch modeling in 2D images, we develop a 3D-cube Mamba model for HSI reconstruction. By integrating the UNet architecture, this model enhances spatial and spectral detail representation through multi-scale receptive field modeling, using fixed cube sizes to dynamically adjust pixel distances. These advancements are incorporated into the A-HQS-accelerated deep unfolding framework, synergistically combining the strengths of 2D-slice and 3D-cube MambaNet to achieve state-of-the-art HSI reconstruction performance. Experimental evaluations on simulated and real-world CASSI datasets demonstrate the efficacy of the proposed approach, achieving superior spectral fidelity and detailed feature representation. The source code is available at: https://github.com/fengyuchao97/SCM-DUN. Yuchao Feng, Zongliang Wu, Yuxiang Yang 0001, Junhua Gao, Xin Yuan 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2025 | Detail Matters: Mamba-Inspired Joint Unfolding Network for Snapshot Spectral Compressive ImagingabstractIn the coded aperture snapshot spectral imaging system, Deep Unfolding Networks (DUNs) have made impressive progress in recovering 3D hyperspectral images (HSIs) from a single 2D measurement. However, the inherent nonlinear and ill-posed characteristics of HSI reconstruction still pose challenges to existing methods in terms of accuracy and stability. To address this issue, we propose a Mamba-inspired Joint Unfolding Network (MiJUN), which integrates physics-embedded DUNs with learning-based HSI imaging. Firstly, leveraging the concept of trapezoid discretization to expand the representation space of unfolding networks, we introduce an accelerated unfolding network scheme. This approach can be interpreted as a generalized accelerated half-quadratic splitting with a second-order differential equation, which reduces the reliance on initial optimization stages and addresses challenges related to long-range interactions. Crucially, within the Mamba framework, we restructure the Mamba-inspired global-to-local attention mechanism by incorporating a selective state space model and an attention mechanism. This effectively reinterprets Mamba as a variant of the Transformer architecture, improving its adaptability and efficiency. Furthermore, we refine the scanning strategy with Mamba by integrating the tensor mode-k unfolding into the Mamba network. This approach emphasizes the low-rank properties of tensors along various modes, while conveniently facilitating 12 scanning directions. Numerical and visual comparisons on both simulation and real datasets demonstrate the superiority of our proposed MiJUN, and achieving overwhelming detail representation. Yuchao Feng, Zongliang Wu, Yulun Zhang 0001, Xin Yuan 0002 |
AAAI | 2 |
| 2025 | Lightweight Accelerated Unfolding Network With Collaborative Attention for Snapshot Spectral Compressive ImagingabstractABSTRACT In coded aperture snapshot spectral imaging (CASSI) systems, deep unfolding networks (DUNs) have made significant strides in recovering 3D hyperspectral images (HSIs) from a single 2D measurement. However, the inherent nonlinearity and ill‐posed nature of HSI reconstruction continue to challenge existing methods in terms of accuracy and stability. To address these challenges, we propose a lightweight collaborative attention‐enhanced accelerated unfolding network (), which integrates a DUN framework with a streamlined prior extractor. Our integrated approach introduces a generically accelerated half‐quadratic splitting algorithm (A‐HQS) for degradation estimation, overcoming the limitations of first‐order optimization and enabling effective long‐range dependency modeling. Within the prior extractor, we introduce cross‐convergence attention, facilitating iterative information exchange between local and non‐local Transformers to capture holistic features and enhance inductive capacity. Notably, the concept of collaborative cross‐convergence is embedded throughout all submodules, ensuring effective information flow. The proposed not only accelerates the convergence of spectral reconstruction, but also fully exploits compressed spatial‐spectral information. Numerical and visual comparisons on both synthetic and real datasets demonstrate the superior performance of this approach. Comparisons on both synthetic and real datasets illustrate the superiority of this approach. The source code is available at https://github.com/Mengjie‐s/CA2UN . Yuchao Feng |
IET Image Process. | 2 |
| 2025 | UP-Diff: Latent Diffusion Model for Remote Sensing Urban PredictionabstractRemote sensing (RS) technology has become essential for monitoring urban development, including applications like population growth analysis, transportation congestion forecasting, and climate change detection (CD). However, its potential for future urban planning (UP), particularly in predicting future urban layouts remains largely unexplored. This study introduces UP-Diff, a novel method leveraging generative models for UP, to address this gap. UP-Diff leverages information from current urban layouts and planned change maps to predict future urban configurations. Key challenges addressed include the integration of urban layouts and change maps into latent diffusion model (LDM) through careful architecture improvements and the mitigation of limited training data by employing a pretrained stable diffusion (SD) model with fixed weights, trainable ConvNeXt, and trainable cross-attention layers. Our method significantly streamlines the UP process by automating layout predictions, thus reducing the time and effort required compared to traditional manual methods. Comprehensive evaluations on the learning, vision, and RS dataset (LEVIR-CD) and Sun Yat-Sen University dataset (SYSU-CD) validate that UP-Diff achieves high-fidelity predictions of future urban layouts, demonstrating its effectiveness and potential for advancing RS-based UP methodologies. Our code and model weights are available athttps://github.com/zeyuwang-zju/UP-Diff. Zeyu Wang 0010, Zecheng Hao, Yuhan Zhang 0006, Yuchao Feng, Yufei Guo 0001 |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2025 | Rethinking Semantic-Level Building Change Detection: Ensemble Learning and Dynamic InteractionabstractBuilding change detection (BCD) of multi-temporal images plays a significant role in urban expansion and area internal change analysis. However, current BCD methods remain stagnant at binary-level predictions due to the scarcity of detectable changes and the imbalance between new constructions and demolitions. To advance semantic-level BCD, we propose a dynamic interaction ensemble learning network (DIELNet) using a collaborative training paradigm across multiple datasets and tasks. Firstly, we create a simulated BCD dataset, Inria-CD, derived from the building segmentation dataset. It features complex structures, large scale, and balanced ratios with both binary- and semantic-level labels. Importantly, we shift the traditional single-dataset and single-task BCD learning paradigm by introducing ensemble learning. This mechanism feeds multiple datasets into the model to obtain binary- and semantic-level predictions through a single training process, accommodating partial samples without semantic-level labels. In addition, our DIELNet incorporates bitemporal dynamic interactions during data processing and feature extraction. The former generates progressive sequences by swapping mutual high-frequency components during the Fourier transformation, while the latter is achieved through Mamba-structure modules, which integrate local convolution with dynamic-static kernels and long-range dependencies via state space models. Numerical and visual comparisons demonstrate the superiority of DIELNet. Moreover, existing algorithms can also benefit significantly from our ensemble learning approach. Datasets and codes are available at: https://github.com/fengyuchao97/DIELNet. Yuchao Feng, Yuxiang Yang 0001, Junhua Gao, Xin Yuan 0002 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2025 | Axial-shunted Spatial-temporal Conversation for Change DetectionabstractBenefitting from the maturing of intelligence techniques and advanced sensors, recent years have witnessed the full flourishing of change detection (CD) on multi-temporal remote sensing images. However, extraneous interference caused by normal temporal evolution and the extreme sparsity of spatial changes still plague the detection accuracy. To counteract this dilemma, a lightweight axial-shunted spatial-temporal conversation network (ASCNet) is proposed, which models the intrinsic representations in dually augmented images with a parallel treatment of convolutions and attentions. Specifically, for the features of weakly augmented bi-temporal image pairs from Siamese CNN, a roundtable attention-based and intra-scale axial-shunted interaction, with linear complexity, is presented. By splitting horizontally or vertically into multiple chunks and then performing axial-squeeze operation, axial-shunted scheme can achieve fine-grained attention while maintaining linear complexity. Moreover, roundtable attention pursues efficient bi-temporal modeling by incorporating both self-attention and cross-attention in a single attentional computation, while imposing change guiding and difference gating for focusing on changes. Simultaneously, a video transformer is introduced for the modeling of strongly augmented sequences, followed by an inter-scale spatial-temporal alignment to recalibrate the feature responses. ASCNet demonstrates state-of-the-art performance on four publicly available CD datasets while maintaining superior computational efficiency. The source code is available at https://github.com/fengyuchao97/ASCNet . Yuchao Feng, Jiawei Jiang 0002, Jintao Lai, Jianwei Zheng 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2024 | SyFormer: Structure-Guided Synergism Transformer for Large-Portion Image InpaintingabstractImage inpainting is in full bloom accompanied by the progress of convolutional neural networks (CNNs) and transformers, revolutionizing the practical management of abnormity disposal, image editing, etc. However, due to the ever-mounting image resolutions and missing areas, the challenges of distorted long-range dependencies from cluttered background distributions and reduced reference information in image domain inevitably rise, which further cause severe performance degradation. To address the challenges, we propose a novel large-portion image inpainting approach, namely the Structure-Guided Synergism Transformer (SyFormer), to rectify the discrepancies in feature representation and enrich the structural cues from limited reference. Specifically, we devise a dual-routing filtering module that employs a progressive filtering strategy to eliminate invalid noise interference and establish global-level texture correlations. Simultaneously, the structurally compact perception module maps an affinity matrix within the introduced structural priors from a structure-aware generator, assisting in matching and filling the corresponding patches of large-proportionally damaged images. Moreover, we carefully assemble the aforementioned modules to achieve feature complementarity. Finally, a feature decoding alignment scheme is introduced in the decoding process, which meticulously achieves texture amalgamation across hierarchical features. Extensive experiments are conducted on two publicly available datasets, i.e., CelebA-HQ and Places2, to qualitatively and quantitatively demonstrate the superiority of our model over state-of-the-arts. Yuchao Feng, Honghui Xu 0002, Chuanmeng Zhu, Jianwei Zheng 0001 |
AAAI | 2 |
| 2024 | Learning Object Placement via Convolution Scoring Attention
Yuchao Feng, Jianwei Zheng 0001 |
BMVC | 2 |
| 2024 | Context-Aware Integration of Language and Visual References for Natural Language TrackingabstractTracking by natural language specification (TNL) aims to consistently localize a target in a video sequence given a linguistic description in the initial frame. Existing methodologies perform language-based and template-based matching for target reasoning separately and merge the matching results from two sources, which suffer from tracking drift when language and visual templates missalign with the dynamic target state and ambiguity in the later merging stage. To tackle the issues, we propose a joint multi-modal tracking framework with 1) a prompt modulation module to leverage the complementarity between temporal visual templates and language expressions, enabling precise and context-aware appearance and linguistic cues, and 2) a unified target decoding module to integrate the multi-modal reference cues and executes the integrated queries on the search image to predict the target location in an end-to-end manner directly. This design ensures spatio-temporal consistency by leveraging historical visual information and introduces an integrated solution, generating predictions in a single step. Extensive experiments conducted on TNL2K, OTB-Lang, LaSOT, and RefCOCOg validate the efficacy of our proposed approach. The results demonstrate competitive performance against state-of-the-art methods for both tracking and grounding. Code is available at https://github.com/twotw02/QueryNLT Yanyan Shao, Shuting He, Qi Ye 0001, Yuchao Feng, Wenhan Luo, Jiming Chen 0001 |
CVPR | 4 |
| 2024 | Multimodal-XAD: Explainable Autonomous Driving Based on Multimodal Environment DescriptionsabstractIn recent years, deep learning-based end-to-end autonomous driving has become increasingly popular. However, deep neural networks are like black boxes. Their outputs are generally not explainable, making them not reliable to be used in real-world environments. To provide a solution to this problem, we propose an explainable deep neural network that jointly predicts driving actions and multimodal environment descriptions of traffic scenes, including bird-eye-view (BEV) maps and natural-language environment descriptions. In this network, both the context information from BEV perception and the local information from semantic perception are considered before producing the driving actions and natural-language environment descriptions. To evaluate our network, we build a new dataset with hand-labelled ground truth for driving actions and multimodal environment descriptions. Experimental results show that the combination of context information and local information enhances the prediction performance of driving action and environment description, thereby improving the safety and explainability of our end-to-end autonomous driving network. Yuchao Feng, Wei Hua 0002, Yuxiang Sun 0002 |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2024 | Contrastive Attention-guided Multi-level Feature Registration for Reference-based Super-resolutionabstractGiven low-quality input and assisted by referential images, reference-based super-resolution (RefSR) strives to enlarge the spatial size with the guarantee of realistic textures, for which sophisticated feature-matching strategies are naturally demanded. However, the miserable transformation gap between inputs and references, e.g., texture rotation and scaling within patches, often yields distorted textures and terrible ghosting artifacts, which seriously hampers the visual senses and their further investigation. To circumvent this challenge, we propose a contrastive attention-guided multi-level feature registration for RefSR, explicitly tapping the potential of interacting between inputs and references. Specifically, we develop a multi-level feature warping scheme, involving patch-level coarse feature swapping and pixel-level deformable alignment, to model generalized spatial transformation correspondences steered by contrastive attention. Notably, a spatial registration module is embedded for further calibration against the potential misalignment issue and inter-feature distribution difference. In addition, aiming at suppressing the impacts of irrelevant or superfluous information on cross-scale features, we incorporate a multi-residual feature fusion module to strive for visually plausible textures. Experimental results on four publicly available datasets demonstrate that our method outperforms most state-of-the-art approaches in terms of both efficiency and perceptual effectiveness. Jianwei Zheng 0001, Yu Liu 0151, Yuchao Feng, Honghui Xu 0002 |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2023 | GCN-based Autism Spectrum Disorder Diagnosis via Convolutional Restructuring Attention
Qianwei Zhou, Yuchao Feng, Jianwei Zheng 0001 |
CogSci | 4 |
| 2023 | Building Change Detection Using Cross-Temporal Feature Interaction NetworkabstractBuilding change detection of remote sensing images is in full flourishing accompanied by the prosperity of convolutional neural networks. For spatial-temporal context modeling, existing solutions disregard the inter-image interactions, albeit their positive contribution to the acquisition of differences. To fill the gap, we propose a cross-temporal feature interaction network to effectively derive the change representations. Specifically, we propose a linearized cross-attention, which motivates each counterpart to glimpse the representation of another image while preserving its own features. In addition, to circumvent the misalignment caused by step-down sampling in the backbone, we introduce multi-level feature alignment using learnable affine transformation and stepwise aggregation. Based on a naive backbone (ResNet18) without sophisticated structures, our model outperforms other state-of-the-art methods on three datasets in terms of both efficiency and effectiveness. Yuchao Feng, Jiawei Jiang 0002, Honghui Xu 0002, Jianwei Zheng 0001 |
ICASSP | 1 |
| 2023 | Low-Dose CT Reconstruction Via Optimization-Inspired GANabstractMost research on Low-dose Computed Tomography (LDCT) reconstruction is designed as a black box, lacking controllability and interpretability. In this paper, a Proximal Linear ADMM framework-based Generative Adversarial Network (PLA-GAN) is proposed. Specifically, without loss of interpretability, channel attention blocks and NonLocal Sparse Attention (NLSA) modules are embedded into two regularizers respectively and iterated alternately, driving the network to cope with real and complex CT image degradation through a multi-scale and adaptive way. To further promote the visual quality, a discriminator containing NLSA module is also introduced. The comparisons with state-of-the-arts on the Mayo dataset validate the superiority of our proposed algorithm both numerically and visually. The advantages of generalizability and interpretability are also evident. Jiawei Jiang 0002, Yuchao Feng, Honghui Xu 0002, Jianwei Zheng 0001 |
ICASSP | 2 |
| 2023 | Compact Intertemporal Coupling Network for Remote Sensing Change DetectionabstractChange detection of multi-temporal remote sensing images is in full flourishing accompanied by the popularity and prosperity of deep learning. The prominent challenge lies in the crude distribution of the newly constructed and demolished changes, the interferences of massive irrelevant objects, and the spatial-temporal changes from the passage of time. For context modeling, existing solutions waste massive attention on task-irrelevant features, spotlighting insufficiently on the genuinely changed regions. To fill the gap, we propose a compact intertemporal coupling network (CICNet) to derive the change representations. Specifically, to underpin the interaction of spatial-temporal differences in a global perspective, we detach and innovate the solo-head self-attention into a lightweight intertemporal-attention, favorably bridging the intra-level features. In addition, to circumvent the misalignment imposed by spatial sampling, lightweight global channel- and spatial- attentions are globally incorporated for stepwise calibration between localization seduced low-level information and semantics abundant high-level features. Based on a naive backbone (ResNet18/34) without sophisticated structures, our model outperforms other state-of-the-art methods on four datasets in terms of both efficiency and effectiveness. Yuchao Feng, Honghui Xu 0002, Jiawei Jiang 0002, Jianwei Zheng 0001 |
ICME | 1 |
| 2023 | GA-HQS: MRI reconstruction via a generically accelerated unfolding approachabstractDeep unfolding networks (DUNs) are the foremost methods in the realm of compressed sensing MRI, as they can employ learnable networks to facilitate interpretable forward-inference operators. However, several daunting issues still exist, including the heavy dependency on first-order optimization algorithms, the insufficient information fusion mechanisms, and the limitation of capturing long-range relationships. To address the issues, we propose a Generically Accelerated Half-Quadratic Splitting (GA-HQS) algorithm that incorporates second-order gradient information and pyramid attention modules for the delicate fusion of inputs at the pixel level. Moreover, a multi-scale split transformer is also designed to enhance the global feature representation. Comprehensive experiments demonstrate that our method surpasses previous ones on single-coil MRI acceleration tasks. Jiawei Jiang 0002, Honghui Xu 0002, Yuchao Feng, Jianwei Zheng 0001 |
ICME | 4 |
| 2023 | CA-GAN: Object Placement via Coalescing Attention based Generative Adversarial NetworkabstractLearning to posit a foreground object over a background scene is an intriguing yet challenging problem, which frequently emerges in applications such as image editing and scene parsing. To date, most existing studies are fed up with knotty issues, including the deficiency of harnessing the interaction between the object and the scene, the astriction of involving little prior knowledge during training, etc. To break the shackles, we propose a novel end-to-end framework dubbed Coalescing Attention based Generative Adversarial Network (CA-GAN). Specifically, in our synthesizer, a feature polymerizer is designed to distill multi-scale information from both background and foreground. On that basis, a dual-branch coalescing attention module is proposed for a better exploration of the global feature-interaction relationships between object and scene. In addition, we add a supervised trail to learn the prior knowledge from the positive composite image, which further guides the synthesizer to discover a credible placement for the foreground object. With extensive experiments conducted on the OPA dataset, our proposal presents superiority in both rationality and diversity compared with other state-of-the-art methods. Our code is available at https://github.com/ZhengJianwei2/CA-GAN. Yuchao Feng, Honghui Xu 0002, Jianwei Zheng 0001 |
ICME | 2 |
| 2023 | Adaptive-Mask Fusion Network for Segmentation of Drivable Road and Negative Obstacle With Untrustworthy FeaturesabstractSegmentation of drivable roads and negative obstacles is critical to the safe driving of autonomous vehicles. Currently, many multi-modal fusion methods have been proposed to improve segmentation accuracy, such as fusing RGB and depth images. However, we find that when fusing two modals of data with untrustworthy features, the performance of multi-modal networks could be degraded, even lower than those using a single modality. In this paper, the untrustworthy features refer to those extracted from regions (e.g., far objects that are beyond the depth measurement range) with invalid depth data (i.e., 0 pixel value) in depth images. The untrustworthy features can confuse the segmentation results, and hence lead to inferior results. To provide a solution to this issue, we propose the adaptive-mask fusion Network (AMFNet) by introducing adaptive-weight masks in the fusion module to fuse features from RGB and depth images with inconsistency. In addition, we release a large-scale RGB-depth dataset with manually-labeled ground truth based on the NPO dataset for drivable roads and negative obstacles segmentation. Extensive experimental results demonstrate that our network achieves state-of-the-art performance compared with other networks. Our code and dataset are available at: https://github.com/lab-sun/AMFNet. Yuchao Feng, Yanning Guo, Yuxiang Sun 0002 |
IV | 2 |
| 2023 | A Lightweight Collective-attention Network for Change DetectionabstractChange detection of multi-temporal remote sensing images is mushrooming with the innovations of neural networks, whose daunting challenge lies in locating sporadically distributed spatial-temporal changes given sophisticated scenes and various imaging conditions. Unfortunately, instead of devoting full attention to changes, most existing solutions often expend unnecessary resources yet derive task-irrelevant features. To relieve this issue, we propose a collective-attention network, which enjoys lightweight model architecture yet guarantees high performance. Specifically, an inter-temporal collective-attention module is developed for efficient interaction of bi-temporal features, in which a shared attention distribution is derived via the multiplication of temporal-concatenated queries and spatial-subtracted keys. Additionally, we present a non-change consistency-constraint, enforcing a change-oriented attention distribution and a noise-suppressed treatment. With the learned interaction features, bi-temporal differences are captured simply using the operations of spatial absolute error and temporal concatenation. Finally, decoding multi-scale differences is accomplished by lightweight temporal self-attention and spatial self-attention. Experiments on four datasets demonstrate that our model achieves state-of-the-art performance, yet requires only 1.71M parameters and 1.98G FLOPs. Yuchao Feng, Yanyan Shao, Honghui Xu 0002, Jinshan Xu, Jianwei Zheng 0001 |
ACM Multimedia | 1 |
| 2023 | Latent-space Unfolding for MRI ReconstructionabstractTo circumvent the problems caused by prolonged acquisition periods, compressed sensing MRI enjoys a high usage profile to accelerate the recovery of high-quality images from under-sampled k-space data. Most current solutions dedicate to solving this issue with the pursuit of certain prior properties, yet the treatments are all enforced in the original space, resulting in limited feature information. To achieve a performance promotion yet with the guarantee of running efficiency, in this work, we propose a latent-space unfolding network (LsUNet). Specifically, by an elaborately designed reversible network, the inputs are first mapped to a channel-lifted latent space, which taps the potential of capturing spatial-invariant features sufficiently. Within the latent space, we then unfold an accelerated optimization algorithm to iterate an efficient and feasible solution, in which a parallelly dual-domain update is equipped for better feature fusion. Finally, an inverse embedding transformation of the recovered high-dimensional representation is applied to achieve the expected estimation. LsUNet enjoys high interpretability due to the physically induced modules, which not only facilitates an intuitive understanding of the internal operating mechanism but also endows it with high generalization ability. Comprehensive experiments on different datasets and various sampling rates/patterns demonstrate the advantages of our proposal over the latest methods both visually and numerically. Jiawei Jiang 0002, Yuchao Feng, Dongyan Guo, Jianwei Zheng 0001 |
ACM Multimedia | 2 |
| 2023 | Tensor completion via hybrid shallow-and-deep priors
Honghui Xu 0002, Jiawei Jiang 0002, Yuchao Feng, Yiting Jin, Jianwei Zheng 0001 |
Appl. Intell. | 3 |
| 2023 | Change Detection on Remote Sensing Images Using Dual-Branch Multilevel Intertemporal NetworkabstractChange detection (CD) of remote sensing (RS) images is mushrooming up accompanied by the on-going innovation of convolutional neural networks (CNNs). Yet with the high-speed technology upgrade, the obstacle that identifies unbalanced variations in foreground–background categories still lies on the table, especially in cases with limited samples and massive interference such as seasonal turnover, illumination intensity, and building reformation. Moreover, to date, neither of the off-the-shelf methods probes the feasibility of direct interaction between bitemporal images before accessing difference features. In this article, we propose a dual-branch multilevel intertemporal network (DMINet) to efficiently and effectively derive the change representations. Specifically, by unifying self-attention (SelfAtt) and cross-attention (CrossAtt) in a single module, we present an intertemporal joint-attention (JointAtt) block to steer the global feature distribution of each input, motivating information coupling between intralevel representations and meanwhile suppressing the task-irrelevant interferences. In addition, centering more on the detection of difference features, a reliable architecture is designed by spotlighting two concerns, i.e., the difference acquisition using subtraction and concatenation as well as the multilevel difference aggregation using incremental feature alignment. Based on a naive backbone without sophisticated structures, i.e., ResNet18, our model outperforms other state-of-the-art (SOTA) methods on four CD datasets, especially in cases with rarely samples. Moreover, the achievement is attained with light overheads. Yuchao Feng, Jiawei Jiang 0002, Honghui Xu 0002, Jianwei Zheng 0001 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2023 | NLE-DM: Natural-Language Explanations for Decision Making of Autonomous Driving Based on Semantic Scene UnderstandingabstractIn recent years, the advancement of deep-learning technologies has greatly promoted the research progress of autonomous driving. However, deep neural network is like a black box. Given a specific input, it is difficult to explain the output of the network. Without explainable results, it would be unsafe to deploy deep networks in unseen environments or environments with potential unexpected situations. Especially for decision-making networks, inappropriate outputs could lead to severe traffic accidents. To provide a solution to this problem, we propose a deep neural network that jointly predicts the decision-making actions and corresponding natural-language explanations based on semantic scene understanding. Two types of explanations, the reasons of driving actions and the surrounding environment descriptions of the ego-vehicle, are designed. Both the reasons and descriptions are in the form of natural language. The decision-making actions could be explained by the corresponding reasons or the environment descriptions. We also release a large-scale dataset with hand-labelled ground truth including driving actions and environment descriptions. The superiority of our network over other methods is demonstrated on both our dataset and a public dataset. Yuchao Feng, Wei Hua 0002, Yuxiang Sun 0002 |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2022 | Boosting Feature-Aware Network for Salient Object Detection
Jianwei Zheng 0001, Yubin Gu, Yuchao Feng, Jinshan Xu |
ICANN (4) | 3 |
| 2022 | Fast Tensor Nuclear Norm for Structured Low-Rank Visual InpaintingabstractLow-rank modeling has achieved great success in visual data completion. However, the low-rank assumption of original visual data may be in approximate mode, which leads to suboptimality for the recovery of underlying details, especially when the missing rate is extremely high. In this paper, we go further by providing a detailed analysis about the rank distributions in Hankel structured and clustered cases, and figure out both non-local similarity and patch-based structuralization play a positive role. This motivates us to develop a new Hankel low-rank tensor recovery method that is competent to truthfully capture the underlying details with sacrifice of slightly more computational burden. First, benefiting from the correlation of different spectral bands and the smoothness of local spatial neighborhood, we divide the visual data into overlapping 3D patches and group the similar ones into individual clusters exploring the non-local similarity. Second, the 3D patches are individually mapped to the structured Hankel tensors for better revealing low-rank property of the image. Finally, we solve the tensor completion model via the well-known alternating direction method of multiplier (ADMM) optimization algorithm. Due to the fact that size expansion happens inevitably in Hankelization operation, we further propose a fast randomized skinny tensor singular value decomposition (rst-SVD) to accelerate the per-iteration running efficiency. Extensive experimental results on real world datasets verify the superiority of our method compared to the state-of-the-art visual inpainting approaches. Honghui Xu 0002, Jianwei Zheng 0001, Xiaomin Yao, Yuchao Feng, Shengyong Chen |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2022 | ICIF-Net: Intra-Scale Cross-Interaction and Inter-Scale Feature Fusion Network for Bitemporal Remote Sensing Images Change DetectionabstractChange detection (CD) of remote sensing (RS) images has enjoyed remarkable success by virtue of convolutional neural networks (CNNs) with promising discriminative capabilities. However, CNNs lack the capability of modeling long-range dependencies in bitemporal image pairs, resulting in inferior identifiability against the same semantic targets yet with varying features. The recently thriving Transformer, on the contrary, is warranted, for practice, with global receptive fields. To jointly harvest the local-global features and circumvent the misalignment issues caused by step-by-step downsampling operations in traditional backbone networks, we propose an intra-scale cross-interaction and inter-scale feature fusion network (ICIF-Net), explicitly tapping the potential of integrating CNN and Transformer. In particular, the local features and global features, respectively, extracted by CNN and Transformer, are interactively communicated at the same spatial resolution using a linearized Conv Attention module, which motivates the counterpart to glimpse the representation of another branch while preserving its own features. In addition, with the introduction of two attention-based inter-scale fusion schemes, including mask-based aggregation and spatial alignment (SA), information integration is enforced at different resolutions. Finally, the integrated features are fed into a conventional change prediction head to generate the output. Extensive experiments conducted on four CD datasets of bitemporal (RS) images demonstrate that our ICIF-Net surpasses the other state-of-the-art (SOTA) approaches. Yuchao Feng, Honghui Xu 0002, Jiawei Jiang 0002, Jianwei Zheng 0001 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2021 | Tensor completion using patch-wise high order Hankelization and randomized tensor ring initialization
Jianwei Zheng 0001, Honghui Xu 0002, Yuchao Feng, Peijun Chen, Shengyong Chen |
Eng. Appl. Artif. Intell. | 4 |
| 2021 | Hyperspectral Image Classification Using Mixed Convolutions and Covariance PoolingabstractRecently, convolution neural network (CNN)-based hyperspectral image (HSI) classification has enjoyed high popularity due to its appealing performance. However, using 2-D or 3-D convolution in a standalone mode may be suboptimal in real applications. On the one hand, the 2-D convolution overlooks the spectral information in extracting feature maps. On the other hand, the 3-D convolution suffers from heavy computation in practice and seems to perform poorly in scenarios having analogous textures along with consecutive spectral bands. To solve these problems, we propose a mixed CNN with covariance pooling for HSI classification. Specifically, our network architecture starts with spectral-spatial 3-D convolutions that followed by a spatial 2-D convolution. Through this mixture operation, we fuse the feature maps generated by 3-D convolutions along the spectral bands for providing complementary information and reducing the dimension of channels. In addition, the covariance pooling technique is adopted to fully extract the second-order information from spectral-spatial feature maps. Motivated by the channel-wise attention mechanism, we further propose two principal component analysis (PCA)-involved strategies, channel-wise shift and channel-wise weighting, to highlight the importance of different spectral bands and recalibrate channel-wise feature response, which can effectively improve the classification accuracy and stability, especially in the case of limited sample size. To verify the effectiveness of the proposed model, we conduct classification experiments on three well-known HSI data sets, Indian Pines, University of Pavia, and Salinas Scene. The experimental results show that our proposal, although with less parameters, achieves better accuracy than other state-of-the-art methods. Jianwei Zheng 0001, Yuchao Feng, Cong Bai, Jinglin Zhang 0003 |
IEEE Trans. Geosci. Remote. Sens. | 2 |