VLDB 2026 Research / reviewers in the wild / expert
Yuanfei Huang
dblp:207/5397
· DBLP profile ↗
23ranked-venue papers
9as first author
13since 2021 · last 2026
0000-0002-5242-9904ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 11 · 4 first-author · 5 since 2021Artificial intelligence and machine learning · 6 · 3 first-author · 5 since 2021Systems, architecture and hardware · 4 · 2 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Optical-to-SAR domain adaptation with inversion regularization for unsupervised ship detection
Yuanfei Huang, Ping Wang 0009, Hua Huang 0001 |
Knowl. Based Syst. | 2 |
| 2026 | Tackling Ill-Posedness of Reversible Image Conversion With Well-Posed Invertible NetworkabstractReversible image conversion (RIC) suffers from ill-posedness issues due to its forward conversion process being considered an underdetermined system. Despite employing invertible neural networks (INN), existing RIC methods intrinsically remain ill-posed as inevitably introducing uncertainty by incorporating randomly sampled variables. To tackle the ill-posedness dilemma, we focus on developing a reliable approximate left inverse for the underdetermined system by constructing an overdetermined system with a non-zero Gram determinant, thus ensuring a well-posed solution. Based on this principle, we propose a well-posed invertible $1\times 1$1×1 convolution (WIC), which eliminates the reliance on random variable sampling and enables the development of well-posed invertible networks. Furthermore, we design two innovative networks, WIN-Naïve and WIN, with the latter incorporating advanced skip-connections to enhance long-term memory. Our methods are evaluated across diverse RIC tasks, including reversible image hiding, image rescaling, and image decolorization, consistently achieving state-of-the-art performance. Extensive experiments validate the effectiveness of our approach, demonstrating its ability to overcome the bottlenecks of existing RIC solutions and setting a new benchmark in the field. Yuanfei Huang, Hua Huang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2025 | DeflareMamba: Hierarchical Vision Mamba for Contextually Consistent Lens Flare RemovalabstractLens flare removal remains an information confusion challenge in the underlying image background and the optical flares, due to the complex optical interactions between light sources and camera lens. While recent solutions have shown promise in decoupling the flare corruption from image, they often fail to maintain contextual consistency, leading to incomplete and inconsistent flare removal. To eliminate this limitation, we propose DeflareMamba, which leverages the efficient sequence modeling capabilities of state space models while maintains the ability to capture local-global dependencies. Particularly, we design a hierarchical framework that establishes long-range pixel correlations through varied stride sampling patterns, and utilize local-enhanced state space models that simultaneously preserves local details. To the best of our knowledge, this is the first work that introduces state space models to the flare removal task. Extensive experiments demonstrate that our method effectively removes various types of flare artifacts, including scattering and reflective flares, while maintaining the natural appearance of non-flare regions. Further downstream applications demonstrate the capacity of our method to improve visual object recognition and cross-modal semantic understanding. Code is available at https://github.com/BNU-ERC-ITEA/DeflareMamba. Yuanfei Huang, Junhui Lin, Hua Huang 0001 |
ACM Multimedia | 2 |
| 2025 | Beyond Image Prior: Embedding Noise Prior into Latent Space of Conditional Denoising Transformer
Yuanfei Huang, Hua Huang 0001 |
Int. J. Comput. Vis. | 1 |
| 2025 | Infrared Video Dynamic Range Compression Based on Global and Local Temporal CoherenceabstractDynamic range compression (DRC) for infrared (IR) video aims to compress high dynamic range (HDR) IR video into low dynamic range (LDR) IR video, facilitating display on common devices and enhancing visual perception. This research requires balancing the preservation of spatial and temporal information, particularly detail preservation and temporal brightness coherence. However, existing methods fail to simultaneously maintain both global and local temporal coherence due to the lack of distinction between global and local motion, which inevitably introduces flickering and reduces the visual experience. To address this issue, this paper proposes an IR video DRC (vDRC) method based on global and local temporal coherence (GLTC). Specifically, a motion mask strategy based on structural similarity is introduced to differentiate between global motion regions, local motion regions, and static regions. A motion estimation strategy using different block-matching scales is then applied to estimate HDR motion information between consecutive frames, which is used to constrain the LDR of the previous frame and construct a temporal constraint term to preserve both global and local temporal coherence. Extensive experiments conducted on two public IR video datasets demonstrate that the proposed method outperforms state-of-the-art methods both quantitatively and qualitatively, offering a more effective solution for IR vDRC and enhancing visualization. Jinyi Qiu, Yuanfei Huang, Hua Huang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2024 | Deep Convolution Modulation for Image Super-ResolutionabstractRecently, deep-learning-based super-resolution methods have achieved excellent performances, but mainly focus on training a single generalized deep network by feeding numerous samples. Yet intuitively, each image has its specific representation, and is expected to acquire an adaptive model. For this issue, we propose a novel convolution modulation (CoMo) mechanism to build image-specific deep networks, by exploiting the principal information of the feature to generate a modulation weight, and thereby adaptively modulating the kernel weights of convolution without any additional parameters, which outperforms the vanilla convolution and several existing attention mechanisms when embedding into the state-of-the-art architectures. To optimize the modulated convolutions in mini-batch training, we introduce an image-specific optimization (IsO) algorithm, which tackles the infeasibility of the conventional optimization algorithms on this issue. Furthermore, we investigate the effect of CoMo on state-of-the-art architectures and design a new CoMoNet architecture by employing the U-style residual learning and hourglass dense block learning, which is an appropriate architecture to utmost improve the effectiveness of CoMo theoretically. Extensive experiments on benchmarks show that the proposed methods achieve superior performances and higher flexibility against the state-of-the-art SISR and blind SR methods. The code is available at github.com/YuanfeiHuang/CoMoNet. Yuanfei Huang, Jie Li 0001, Yanting Hu, Hua Huang 0001, Xinbo Gao 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2024 | Infrared Image Dynamic Range Compression Based on Adaptive Contrast Adjustment and Structure PreservationabstractThe infrared (IR) image dynamic range compression (DRC) technology involves compressing high dynamic range (HDR) IR images into low dynamic range (LDR) images for display on common devices. To facilitate human observation, DRC methods should preserve the structural information as much as possible while adjusting the contrast of HDR IR images. However, existing DRC methods struggle to adapt to various highly dynamic IR scenes when using fixed parameter settings. To address this limitation, a novel gradient domain-based DRC method with adaptive contrast adjustment and structure preservation (ACASP) is proposed. Our ACASP adapts local contrast and gradients by analyzing local features of HDR IR images, effectively handling different HDR IR scenes. We introduce local contrast and variance to enhance visibility in low-contrast areas and preserve details in high-contrast areas. Specifically, we design a contrast-adaptive mapping curve and a gradient-adaptive modulation factor (GMF) to optimize both contrast and structure in the LDR image. Extensive experiments on three public HDR IR datasets demonstrate that the proposed method can outperform state-of-the-art DRC methods in both quantitative and qualitative analyses. This work contributes to the field by offering a more adaptive and robust approach to IR image DRC. Jinyi Qiu, Zhan Wang 0007, Yuanfei Huang, Hua Huang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2023 | Multi-scale information distillation network for efficient image super-resolution
Yanting Hu, Yuanfei Huang, Kaibing Zhang |
Knowl. Based Syst. | 2 |
| 2023 | Transitional Learning: Exploring the Transition States of Degradation for Blind Super-resolutionabstractBeing extremely dependent on iterative estimation of the degradation prior or optimization of the model from scratch, the existing blind super-resolution (SR) methods are generally time-consuming and less effective, as the estimation of degradation proceeds from a blind initialization and lacks interpretable representation of degradations. To address it, this article proposes a transitional learning method for blind SR using an end-to-end network without any additional iterations in inference, and explores an effective representation for unknown degradation. To begin with, we analyze and demonstrate the transitionality of degradations as interpretable prior information to indirectly infer the unknown degradation model, including the widely used additive and convolutive degradations. We then propose a novel Transitional Learning method for blind Super-Resolution (TLSR), by adaptively inferring a transitional transformation function to solve the unknown degradations without any iterative operations in inference. Specifically, the end-to-end TLSR network consists of a degree of transitionality (DoT) estimation network, a homogeneous feature extraction network, and a transitional learning module. Quantitative and qualitative evaluations on blind SR tasks demonstrate that the proposed TLSR achieves superior performances and costs fewer complexities against the state-of-the-art blind SR methods. The code is available at github.com/YuanfeiHuang/TLSR. Yuanfei Huang, Jie Li 0001, Yanting Hu, Xinbo Gao 0001, Hua Huang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2022 | Image Super-Resolution With Self-Similarity Prior Guided Network and Sample-Discriminating LearningabstractThe nonlocal self-similarity in natural image provides an effective prior for single image super-resolution (SISR), which is beneficial to contextual information capture and performance improvement, as demonstrated by conventional SISR methods. However, it is little explored to utilize this property in deep neural networks. In this paper, we propose a self-similarity prior guided (SSPG) network to incorporate self-similarity-based nonlocal operation into deep neural network for SISR. Specifically, we design a cross-scale nearest-neighbor residual (CSNNR) block via introducing cross-scale$k$-nearest neighbors (KNN) matching into a residual block, which can be flexibly integrated into deep networks to capture long-range correlations among multi-scale and multi-level features. Meanwhile, by stacking a CSNNR block and a sequence of wide-activated residual blocks with a local skip-connection, a multi-level residual self-similarity (MRSS) module is developed to effectively employ local and nonlocal information for detail recovery. Thus, through cascading multiple MRSS modules, the proposed SSPG network performs both self-similarity-based nonlocal operation and convolution-based local operation on multi-level features to reconstruct informative features for accurate SISR. In addition, for pursuing visually pleasing results, we apply our SSPG network to the perception-oriented SISR field by following the framework of generative adversarial networks. In particular, we explore a sample-discriminating learning mechanism based on the statistical descriptions of training samples, and include it in optimization procedure to automatically tune the contributions of different samples according to their characteristics and then focus the network on creating realistic results. Extensive quantitative and qualitative evaluations on benchmark datasets illustrate the superiority of our proposed models over the state-of-the-art methods for both distortion-oriented and perception-oriented image super-resolution tasks. Yanting Hu, Jie Li 0001, Yuanfei Huang, Xinbo Gao 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2021 | RF Energy Harvesting and Management for Near-Zero Power Passive DevicesabstractWe present RF energy harvester and management strategy tailored for the passive near-zero power devices. Radio- less RF-powered backscattering tags that have the ability to recognize and localize activities in the surrounding environment are example of such devices. We propose a management strategy that determines the operation regime of the harvester based on the input power level at which harvester provides the instantaneous supply voltage for device operation. As the input power exceeds this level, the storage of the excess energy is managed by an adaptive capacitor charging circuit that keeps the voltage at the input of voltage regulator constant. We demonstrate that backscatter-based RF tag in the listening mode of operation can instantaneously operate with an input power of -34.4 dBm. Due to the adaptive capacitor charging circuit, the power efficiency of the energy harvester is higher than 50% over a range of input powers from -25 dBm up to -5 dBm. Yuanfei Huang, Akshay Athalye, Samir Ranjan Das, Petar M. Djuric, Milutin Stanacevic |
ISCAS | 1 |
| 2021 | Single image super-resolution with multi-scale information cross-fusion network
Yanting Hu, Xinbo Gao 0001, Jie Li 0001, Yuanfei Huang, Hanzi Wang |
Signal Process. | 4 |
| 2021 | Interpretable Detail-Fidelity Attention Network for Single Image Super-ResolutionabstractBenefiting from the strong capabilities of deep CNNs for feature representation and nonlinear mapping, deep-learning-based methods have achieved excellent performance in single image super-resolution. However, most existing SR methods depend on the high capacity of networks that are initially designed for visual recognition, and rarely consider the initial intention of super-resolution for detail fidelity. To pursue this intention, there are two challenging issues that must be solved: (1) learning appropriate operators which is adaptive to the diverse characteristics of smoothes and details; (2) improving the ability of the model to preserve low-frequency smoothes and reconstruct high-frequency details. To solve these problems, we propose a purposeful and interpretable detail-fidelity attention network to progressively process these smoothes and details in a divide-and-conquer manner, which is a novel and specific prospect of image super-resolution for the purpose of improving detail fidelity. This proposed method updates the concept of blindly designing or using deep CNNs architectures for only feature representation in local receptive fields. In particular, we propose a Hessian filtering for interpretable high-profile feature representation for detail inference, along with a dilated encoder-decoder and a distribution alignment cell to improve the inferred Hessian features in a morphological manner and statistical manner respectively. Extensive experiments demonstrate that the proposed method achieves superior performance compared to the state-of-the-art methods both quantitatively and qualitatively. The code is available at github.com/YuanfeiHuang/DeFiAN. Yuanfei Huang, Jie Li 0001, Xinbo Gao 0001, Yanting Hu, Wen Lu 0004 |
IEEE Trans. Image Process. | 1 |
| 2020 | On Measuring Doppler Shifts between Tags in a Backscattering Tag-to-Tag Network with Applications in TrackingabstractIn this paper, we present a technique whereby passive tags can track each other in a backscattering tag-to-tag network (BTTN). In such a network, passive tags without any on-board radio transceivers communicate directly with each other by backscattering an external excitation signal. First, we explain how the tags determine their distances to other communicating tags in their proximity and then how they can track nearby tags. Our technique is based on multiphase backscattering, more specifically, on the ability of backscattering tags to systematically change the phase offset of the signal that is being backscattered. A passive receiving tag with an envelope detector can then examine the received signal amplitude over the multiple backscattering phases and can draw inferences about the inter-tag distance. We demonstrate our method and show its accuracy on tags that we have built in our lab. Experiments show that our passive tags can measure Doppler shifts with approximately the same accuracy as that achieved by active conventional RFID readers. Our median tracking error based on data from two tags is only about 2.5 cm. Abeer Ahmad, Yuanfei Huang, Xiao Sha, Akshay Athalye, Milutin Stanacevic, Samir Ranjan Das, Petar M. Djuric |
ICASSP | 2 |
| 2020 | A Self-Biased Low Modulation Index ASK Demodulator for Implantable DevicesabstractFree floating sub-mm and mm sized brain implants can communicate through a backscatter-based link in a presence of the EM field generated by the external coil. This link reduces the bandwidth requirement in the uplink communication of these implants to the external coil and enables a close-loop operation of the distributed implant system through reduced latency. The critical challenge in the link design stems from the low modulation index in the incident signal at the receiving coil. This calls for the design of the ASK demodulator that can resolve signals with low modulation index. We propose a demodulator design comprising a self-biased common-source based envelope detector that provides sufficient conversion gain and at the same time operates with a low power consumption. With 90 MHz carrier frequency and 50-kbps data rate, the ASK demodulator, implemented in 65 nm CMOS technology, resolves input RF signal with 1% modulation index consuming less than 100 nW when amplitude of the input RF signal is 200 mV. Xiao Sha, Yuanfei Huang, Tutu Wan, Yasha Karimi, Samir Ranjan Das, Petar M. Djuric, Milutin Stanacevic |
ISCAS | 2 |
| 2020 | Channel-Wise and Spatial Feature Modulation Network for Single Image Super-ResolutionabstractThe performance of single image super-resolution has achieved significant improvement by utilizing deep convolutional neural networks (CNNs). The features in deep CNN contain different types of information which make different contributions to image reconstruction. However, the most CNN-based models lack discriminative ability for different types of information and deal with them equally, which results in the representational capacity of the models being limited. On the other hand, as the depth of neural network grows, the long-term information coming from preceding layers is easy to be weaken or lost at later layers, which is adverse to super-resolving image. To capture more informative features and maintain long-term information for image super-resolution, we propose a channel-wise and spatial feature modulation (CSFM) network in which a series of feature modulation memory (FMM) modules are cascaded with a densely connected structure to transform shallow features to high informative features. In each FMM module, we construct a set of channel-wise and spatial attention residual (CSAR) blocks and stack them in a chain structure to dynamically modulate the multi-level features in global and local manners. This feature modulation strategy enables the valuable information to be enhanced and the redundant information to be suppressed. Meanwhile, for long-term information persistence, a gated fusion (GF) node is attached at the end of the FMM module to adaptively fuse hierarchical features and distill more effective information via the dense skip connections and the gating mechanism. The extensive quantitative and qualitative evaluations on benchmark datasets illustrate the superiority of our proposed method over the state-of-the-art methods. Yanting Hu, Jie Li 0001, Yuanfei Huang, Xinbo Gao 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2019 | Multi-scale Spatial-temporal Network for Person Re-identificationabstractVideo-based person re-identification (ReID) is an important task, which has received much attention in recent years due to its efficiency in the field of surveillance. Researchers have employed many effective approaches for video-based person ReID, but there are still two problems. Firstly, the same pedestrian in the video sequences differs in size. Secondly, traditional RNNs can only process one-dimension features, which are not suitable for dealing with video sequences. To solve above problems, we propose a new network called Multi-scale Spatial-Temporal Network (MSTN), which combines multi-scale feature extractor and CLSTM together to tackle the discrepant sizes of pedestrians and extract more representative temporal information for the video sequences. We conduct the experiments on the iLIDS-VID, PRID-2011 and MARS datasets, and our approach outperforms state-of-the-art methods by a large margin. Zhikang Wang, Lihuo He, Xinbo Gao 0001, Yuanfei Huang |
ICASSP | 4 |
| 2019 | Improving Image Super-Resolution via Feature Re-Balancing FusionabstractRecently, benefiting from the strong ability of feature representation, deep-learning-based methods have achieved excellent performance in single image super-resolution (SR). Furthermore, skip connection and feature fusion have been demonstrated to be a commendable strategy to deal with various features in different depth for informative reconstruction. Nevertheless, cross-layer features show diverse characteristic in detail representation, blindly fusion then introduces unavoidable interference of features. In this paper, we propose a novel feature fusion unit by utilizing alternative dilated convolutions for re-balancing diverse cross-layer features, named Feature Re-Balancing Fusion Network (RBFNet), which is theoretically and experimentally demonstrated to be robust to the interference in feature fusion for SR. Extensive experiments show that the proposed method achieves excellent performances quantitatively and qualitatively against the state-of-the-art methods. Yuanfei Huang, Jie Li 0001, Xinbo Gao 0001, Wen Lu 0004, Yanting Hu |
ICME | 1 |
| 2019 | Signal Shaping at Interface of Wireless Power Harvesting and AC Computational LogicabstractThe wirelessly powered adiabatic logic has introduced significant power savings in the design of the computational logic. We explore energy-efficient interfacing of one of the most efficient adiabatic logic families, pass-transistor adiabatic logic (PAL), with RF harvested signal. The interface circuit, signal shaper, transforms the bipolar sinusoidal input voltage to nonnegative unipolar sinusoidal output that serves as the power clock signal for PAL. A theoretical analysis of the operation of the signal shaper is presented and verified using simulations in 65 nm CMOS technology. The designed shaper, when interfaced with 8-bit multiplier implemented using PAL, demonstrates the settling time of a few clock periods and high power conversion efficiency as high as 90%. Yuanfei Huang, Tutu Wan, Emre Salman, Milutin Stanacevic |
ISCAS | 1 |
| 2019 | Passive Wireless Channel Estimation in RF Tag NetworkabstractWe envision a future where every object in our living and working environment will carry one or more RF tags. Based on the backscattering tag-to-tag communication link, these RF tags will be connected in a network without the need for the central interrogating device. We present a novel tag architecture that enables estimation of the parameters of wireless tag-to-tag channel by a passive receiver. Sampling the received baseband signal at different reflecting phases at the backscattering tag enables estimation of amplitude and phase of the tag-to-tag channel. The low-power implementation of the channel estimator, after envelope detection, integrates amplification and filtering of the baseband signal that is followed by analog-to-digital conversion. The channel estimator, implemented in 65 nm CMOS technology, has sensitivity of -45 dBm at 2.5% modulation index and consumes 122 nW. Yasha Karimi, Yuanfei Huang, Akshay Athalye, Samir Ranjan Das, Petar M. Djuric, Milutin Stanacevic |
ISCAS | 2 |
| 2019 | Densely convolutional attention network for image super-resolution
Furui Bai, Wen Lu 0004, Yuanfei Huang, Lin Zha |
Neurocomputing | 3 |
| 2018 | Single Image Super Resolution Based on Deep Residual Network via Lateral ModulesabstractRecently, convolutional neural networks have demonstrated high-quality reconstruction for single image super resolution (SISR). In this paper, we propose a Deep Residual Network via lateral modules (DRNLM), DRNLM is the structure with lateral modules, progressive and symmetric residual blocks (convolutional residual blocks and deconvolutional residual blocks). First, DRNLM introduces lateral modules, which are used to transmit low-level features (coarse residue) into high-level features (fine residue) effectively, thus finer residue can be obtained for better image reconstruction. Second, considering more channels can stack more details, progressive channels that vary from 64 to 256 are utilized in DRNLM through residual blocks. Third, symmetric residual blocks have same dimensions of input and output, which can ensure the gradient ranging within certain limits when the network goes deeper. Extensive experiments demonstrate that the proposed method outperforms the existing methods in accuracy and visual impression. Rui Wang 0173, Wen Lu 0004, Yuanfei Huang, Xinbo Gao 0001, Lihuo He |
ICIP | 4 |
| 2018 | Single Image Super-Resolution via Multiple Mixture Prior ModelsabstractExample learning-based single image super-resolution (SR) is a promising method for reconstructing a high-resolution (HR) image from a single-input low-resolution (LR) image. Lots of popular SR approaches are more likely either time-or space-intensive, which limit their practical applications. Hence, some research has focused on a subspace view and delivered state-of-the-art results. In this paper, we utilize an effective way with mixture prior models to transform the large nonlinear feature space of LR images into a group of linear subspaces in the training phase. In particular, we first partition image patches into several groups by a novel selective patch processing method based on difference curvature of LR patches, and then learning the mixture prior models in each group. Moreover, different prior distributions have various effectiveness in SR, and in this case, we find that student-t prior shows stronger performance than the well-known Gaussian prior. In the testing phase, we adopt the learned multiple mixture prior models to map the input LR features into the appropriate subspace, and finally reconstruct the corresponding HR image in a novel mixed matching way. Experimental results indicate that the proposed approach is both quantitatively and qualitatively superior to some state-of-the-art SR methods. Yuanfei Huang, Jie Li 0001, Xinbo Gao 0001, Lihuo He, Wen Lu 0004 |
IEEE Trans. Image Process. | 1 |