VLDB 2026 Research / reviewers in the wild / expert
Feng Li 0037
dblp:92/2954-37
· DBLP profile ↗
40ranked-venue papers
12as first author
34since 2021 · last 2026
0000-0001-9862-0432ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 25 · 5 first-author · 22 since 2021Artificial intelligence and machine learning · 13 · 3 first-author · 10 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 3 first-author · 5 since 2021Software engineering, systems software and programming languages · 4 · 2 first-author · 3 since 2021Computer networks · 1 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | BlindDiff: empowering degradation modeling in diffusion models for blind image super-resolution
Feng Li 0037, Zichao Liang, Runmin Cong, Huihui Bai 0001, Yao Zhao 0001, Meng Wang 0001 |
Sci. China Inf. Sci. | 1 |
| 2026 | You can mask more for extremely low-bitrate image compression
Feng Li 0037, Jiaxin Han, Runmin Cong, Yunchao Wei, Weisi Lin, Yao Zhao 0001, Huihui Bai 0001 |
Pattern Recognit. | 2 |
| 2026 | DiffLLFace: Learning Alternate Illumination-Diffusion Adaptation for Low-Light Face Super-Resolution and BeyondabstractFacial image acquisition under constrained illumination and with limited-resolution imaging devices often results in coupled photometric and geometric degradations, manifesting as low-light and low-resolution (LLR) conditions. Prevailing research predominantly follows fragmented optimization paradigms that address low-light image enhancement (LLIE) and face super-resolution (FSR) as isolated tasks. This approach overlooks the compound nature of the degradations, thereby significantly limiting their applicability in practical scenarios. To bridge this gap, we present DiffLLFace, a unified framework that harnesses diffusive generative capabilities with illumination-aware trajectories to achieve robust FSR from LLR observations. The core of our method lies in its alternate illumination-diffusion adaptation, which operates throughout the generation process. This mechanism not only captures degradation patterns in both brightness and structure to harmonize latent representations but also dynamically calibrates the illumination prior with the generative knowledge inherent to diffusion models. As such, DiffLLFace attains precise control over conditional adaptation and illumination rectification. We further devise a simple yet effective non-parametric Fourier enhancement strategy, which provides structural appearance clues that work in concert with the alternate adaptation to ensure texture and color consistency. Extensive experiments demonstrate the superiority of DiffLLFace over existing methods and remarkable generalizability on complex natural scenes. Code is available at https://github.com/KaishengPang/DiffLLFace. Runmin Cong, Kaisheng Pang, Feng Li 0037, Hua Li 0012, Huihui Bai 0001, Sam Kwong, Wei Zhang 0021 |
IEEE Trans. Image Process. | 3 |
| 2026 | Bridging Component Learning With Degradation Modelling for Blind Image Super-ResolutionabstractConvolutional Neural Network (CNN)-based image super-resolution (SR) has exhibited impressive success on known degraded low-resolution (LR) images. However, this type of approach is hard to hold its performance in practical scenarios when the degradation process (i.e.blur and downsampling) is unknown. Despite existing blind SR methods proposed to solve this problem using blur kernel estimation, the perceptual quality and reconstruction accuracy are still unsatisfactory. In this paper, we analyze the degradation of a high-resolution (HR) image from image intrinsic components according to a degradation-based formulation model. We propose a components decomposition and co-optimization network (CDCN) for blind SR. Firstly, CDCN decomposes the input LR image into structure and detail components in feature space. Then, the mutual collaboration block (MCB) is presented to exploit the relationship between both two components. In this way, the detail component can provide informative features to enrich the structural context and the structure component can carry structural context for better detail revealing via a mutual complementary manner. After that, we present a degradation-driven learning strategy to jointly supervise the HR image detail and structure restoration process. Finally, a multi-scale fusion module followed by an upsampling layer is designed to fuse the structure and detail features and perform SR reconstruction. Empowered by such degradation-based components decomposition, collaboration, and mutual optimization, we can bridge the correlation between component learning and degradation modelling for blind SR, thereby producing SR results with more accurate textures. Extensive experiments on both synthetic SR datasets and real-world images show that the proposed method achieves the state-of-the-art performance compared to existing methods. Feng Li 0037, Huihui Bai 0001, Weisi Lin, Runmin Cong, Yao Zhao 0001 |
IEEE Trans. Multim. | 2 |
| 2025 | Attend and Enrich: Enhanced Visual Prompt for Zero-Shot LearningabstractZero-shot learning (ZSL) endeavors to transfer knowledge from the seen categories to recognize unseen categories, which mostly relies on the semantic-visual interactions between image and attribute tokens. Recently, the prompt learning has emerged in ZSL and demonstrated significant potential as it allows the zero-shot transfer of diverse visual concepts to downstream tasks. However, current methods explore the fixed adaptation of the learnable prompt on the seen domains, which make them over-emphasize the primary visual features observed during training, limiting their generalization capabilities to the unseen domains. In this work, we propose AENet, which endows semantic information into the visual prompt to distill semantic-enhanced prompt for visual representation enrichment, enabling effective knowledge transfer for ZSL. AENet comprises two key steps: 1) exploring the concept-harmonized tokens for the visual and attribute modalities, grounded on the modal-sharing token that represents consistent visual-semantic concepts; and 2) yielding the semantic-enhanced prompt via the visual residual refinement unit with attribute consistency supervision. It is further integrated with primary visual features to attend to semantic-related information for visual enhancement, thus strengthening transferable ability. Experimental results on three benchmarks show that our AENet outperforms existing state-of-the-art ZSL methods. Man Liu 0003, Huihui Bai 0001, Feng Li 0037, Chunjie Zhang 0001, Yunchao Wei, Tat-Seng Chua, Yao Zhao 0001 |
AAAI | 3 |
| 2025 | Sign-IDD: Iconicity Disentangled Diffusion for Sign Language ProductionabstractSign Language Production (SLP) aims to generate semantically consistent sign videos from textual statements, where the conversion from textual glosses to sign poses (G2P) is a crucial step. Existing G2P methods typically treat sign poses as discrete three-dimensional coordinates and directly fit them, which overlooks the relative positional relationships among joints. To this end, we provide a new perspective, constraining joint associations and gesture details by modeling the limb bones to improve the accuracy and naturalness of the generated poses. In this work, we propose a pioneering iconicity disentangled diffusion framework, termed Sign-IDD, specifically designed for SLP. Sign-IDD incorporates a novel Iconicity Disentanglement (ID) module to bridge the gap between relative positions among joints. The ID module disentangles the conventional 3D joint representation into a 4D bone representation, comprising the 3D spatial direction vector and 1D spatial distance vector between adjacent joints. Additionally, an Attribute Controllable Diffusion (ACD) module is introduced to further constrain joint associations, in which the attribute separation layer aims to separate the bone direction and length attributes, and the attribute control layer is designed to guide the pose generation by leveraging the above attributes. The ACD module utilizes the gloss embeddings as semantic conditions and finally generates sign poses from noise embeddings. Extensive experiments on PHOENIX14T and USTC-CSL datasets validate the effectiveness of our method. Shengeng Tang, Dan Guo 0001, Yanyan Wei, Feng Li 0037, Richang Hong |
AAAI | 5 |
| 2025 | EvEnhancer: Empowering Effectiveness, Efficiency and Generalizability for Continuous Space-Time Video Super-Resolution with EventsabstractContinuous space-time video super-resolution (C-STVSR) endeavors to upscale videos simultaneously at arbitrary spatial and temporal scales, which has recently garnered increasing interest. However, prevailing methods struggle to yield satisfactory videos at out-of-distribution spatial and temporal scales. On the other hand, event streams characterized by high temporal resolution and high dynamic range, exhibit compelling promise in vision tasks. This paper presents EvEnhancer, an innovative approach that marries the unique advantages of event streams to elevate effectiveness, efficiency, and generalizability for C-STVSR. Our approach hinges on two pivotal components: 1) Event-adapted synthesis capitalizes on the spatiotemporal correlations between frames and events to discern and learn long-term motion trajectories, enabling the adaptive interpolation and fusion of informative spatiotemporal features; 2) Local implicit video transformer integrates local implicit video neural function with cross-scale spatiotemporal attention to learn continuous video representations utilized to generate plausible videos at arbitrary resolutions and frame rates. Experiments show that EvEnhancer achieves superiority on synthetic and real-world datasets and preferable generalizability on out-of-distribution scales against state-of-the-art methods. Code is available at https://github.com/W-Shuoyan/EvEnhancer. Shuoyan Wei, Feng Li 0037, Shengeng Tang, Yao Zhao 0001, Huihui Bai 0001 |
CVPR | 2 |
| 2025 | Once-for-All: Controllable Generative Image Compression with Dynamic Granularity AdaptationabstractAlthough recent generative image compression methods have demonstrated impressive potential in optimizing the rate-distortion-perception trade-off, they still face the critical challenge of flexible rate adaptation to diverse compression necessities and scenarios. To overcome this challenge, this paper proposes a $\textbf{Control}$lable $\textbf{G}$enerative $\textbf{I}$mage $\textbf{C}$ompression framework, $\textbf{Control-GIC}$, the first capable of fine-grained bitrate adaptation across a broad spectrum while ensuring high-fidelity and generality compression. We base Control-GIC on a VQGAN framework representing an image as a sequence of variable-length codes ($\textit{i.e.}$ VQ-indices), which can be losslessly compressed and exhibits a direct positive correlation with bitrates. Drawing inspiration from the classical coding principle, we correlate the information density of local image patches with their granular representations. Hence, we can flexibly determine a proper allocation of granularity for the patches to achieve dynamic adjustment for VQ-indices, resulting in desirable compression rates. We further develop a probabilistic conditional decoder capable of retrieving historic encoded multi-granularity representations according to transmitted codes, and then reconstruct hierarchical granular features in the formalization of conditional probability, enabling more informative aggregation to improve reconstruction realism. Our experiments show that Control-GIC allows highly flexible and controllable bitrate adaptation where the results demonstrate its superior performance over recent state-of-the-art methods. Feng Li 0037, Yuxi Liu 0020, Runmin Cong, Yao Zhao 0001, Huihui Bai 0001 |
ICLR | 2 |
| 2025 | SUEDE: Shared Unified Experts for Physical- Digital Face Attack Detection EnhancementabstractFace recognition systems are vulnerable to physical attacks (e.g., printed photos) and digital threats (e.g., DeepFake), which are currently being studied as independent visual tasks, such as Face Anti-Spoofing and Forgery Detection. The inherent differences among various attack types present significant challenges in identifying a common feature space, making it difficult to develop a unified framework for detecting data from both attack modalities simultaneously. Inspired by the efficacy of Mixture-of-Experts (MoE) in learning across diverse domains, we explore utilizing multiple experts to learn the distinct features of various attack types. However, the feature distributions of physical and digital attacks overlap and differ. This suggests that relying solely on distinct experts to learn the unique features of each attack type may overlook shared knowledge between them. To address these issues, we propose SUEDE, the Shared Unified Experts for Physical-Digital Face Attack Detection Enhancement. SUEDE combines a shared expert (always activated) to capture common features for both attack types and multiple routed experts (selectively activated) for specific attack types. Further, we integrate CLIP as the base network to ensure the shared expert benefits from prior visual knowledge and align visual-text representations in a unified space. Extensive results demonstrate SUEDE achieves superior performance compared to state-of-the-art unified detection methods. Zuying Xie, Changtao Miao, Ajian Liu 0001, Jiabao Guo, Feng Li 0037, Dan Guo 0001, Yunfeng Diao |
ICME | 5 |
| 2025 | MM-Prompt: Multi-modality and Multi-granularity Prompts for Few-Shot SegmentationabstractDespite the effectiveness of Segment Anything Model (SAM) based methods in Few-Shot Segmentation (FSS) tasks, our closer examination of their prompt encoding mechanism reveals that these methods rely solely on visual information to generate a single type of prompt. Consequently, they suffer from semantic granularity representation bias and a loss of spatial information. To address these limitations, this paper introduces an innovative multi-modal prompt encoder, enabling SAM to leverage both annotated reference images and textual descriptions of class names as segmentation prompts. This approach generates text prompts, dense visual prompts, and sparse visual prompts, spanning multiple modalities and granularities. These prompts provide enhanced representations of the target class, capturing both abstract semantics and specific details, while ensuring granularity appropriateness. When our multi-modal prompt encoder is integrated with SAM's image encoder and mask decoder, the overall model is referred to as MM-Prompt. To validate its effectiveness, we conducted extensive empirical studies on the PASCAL-5^i and COCO-20^i datasets. The experimental results demonstrate that MM-Prompt achieves state-of-the-art performance in FSS tasks, highlighting its substantial potential and value in this domain. Runmin Cong, Jinpeng Chen 0003, Chen Zhang 0013, Feng Li 0037, Huihui Bai 0001, Sam Kwong |
ACM Multimedia | 5 |
| 2025 | Exploring Identifiable Tokens for Text-Based Person Re-identification
Wentao Ma 0003, Pengfei Yuan, Feng Li 0037, Shan Zhao 0002 |
PRCV (18) | 5 |
| 2025 | SRConvNet: A Transformer-Style ConvNet for Lightweight Image Super-Resolution
Feng Li 0037, Runmin Cong, Jingjing Wu 0001, Huihui Bai 0001, Meng Wang 0001, Yao Zhao 0001 |
Int. J. Comput. Vis. | 1 |
| 2025 | Coverage-Based Fault Localization in Haskell
Feng Li 0037, Guo-Qing Wang, Meng Wang 0001, Dan Hao 0001 |
J. Comput. Sci. Technol. | 1 |
| 2025 | PSVMA+: Exploring Multi-Granularity Semantic-Visual Adaption for Generalized Zero-Shot LearningabstractGeneralized zero-shot learning (GZSL) endeavors to identify the unseen categories using knowledge from the seen domain, necessitating the intrinsic interactions between the visual features and attribute semantic features. However, GZSL suffers from insufficient visual-semantic correspondences due to the attribute diversity and instance diversity. Attribute diversity refers to varying semantic granularity in attribute descriptions, ranging from low-level (specific, directly observable) to high-level (abstract, highly generic) characteristics. This diversity challenges the collection of adequate visual cues for attributes under a uni-granularity. Additionally, diverse visual instances corresponding to the same sharing attributes introduce semantic ambiguity, leading to vague visual patterns. To tackle these problems, we propose a multi-granularity progressive semantic-visual mutual adaption (PSVMA+) network, where sufficient visual elements across granularity levels can be gathered to remedy the granularity inconsistency. PSVMA+ explores semantic-visual interactions at different granularity levels, enabling awareness of multi-granularity in both visual and semantic elements. At each granularity level, the dual semantic-visual transformer module (DSVTM) recasts the sharing attributes into instance-centric attributes and aggregates the semantic-related visual regions, thereby learning unambiguous visual features to accommodate various instances. Given the diverse contributions of different granularities, PSVMA+ employs selective cross-granularity learning to leverage knowledge from reliable granularities and adaptively fuses multi-granularity features for comprehensive representations. Experimental results demonstrate that PSVMA+ consistently outperforms state-of-the-art methods. Man Liu 0003, Huihui Bai 0001, Feng Li 0037, Chunjie Zhang 0001, Yunchao Wei, Meng Wang 0001, Tat-Seng Chua, Yao Zhao 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2025 | Reference-Based Iterative Interaction With P2-Matching for Stereo Image Super-ResolutionabstractStereo Image Super-Resolution (SSR) holds great promise in improving the quality of stereo images by exploiting the complementary information between left and right views. Most SSR methods primarily focus on the inter-view correspondences in low-resolution (LR) space. The potential of referencing a high-quality SR image of one view benefits the SR for the other is often overlooked, while those with abundant textures contribute to accurate correspondences. Therefore, we propose Reference-based Iterative Interaction (RIISSR), which utilizes reference-based iterative pixel-wise and patch-wise matching, dubbed $P^{2}$ -Matching, to establish cross-view and cross-resolution correspondences for SSR. Specifically, we first design the information perception block (IPB) cascaded in parallel to extract hierarchical contextualized features for different views. Pixel-wise matching is embedded between two parallel IPBs to exploit cross-view interaction in LR space. Iterative patch-wise matching is then executed by utilizing the SR stereo pair as another mutual reference, capitalizing on the cross-scale patch recurrence property to learn high-resolution (HR) correspondences for SSR performance. Moreover, we introduce the supervised side-out modulator (SSOM) to re-weight local intra-view features and produce intermediate SR images, which seamlessly bridge two matching mechanisms. Experimental results demonstrate the superiority of RIISSR against existing state-of-the-art methods. Runmin Cong, Rongxin Liao, Feng Li 0037, Ronghui Sheng, Huihui Bai 0001, Renjie Wan, Sam Kwong, Wei Zhang 0021 |
IEEE Trans. Image Process. | 3 |
| 2025 | Prompt to Restore, Restore to Prompt: Cyclic Prompting for Universal Adverse Weather RemovalabstractUniversal adverse weather removal (UAWR) seeks to address various weather degradations within a unified framework. Recent methods are inspired by prompt learning using pre-trained vision-language models (e.g., CLIP), leveraging degradation-aware prompts to facilitate weather-free image restoration, yielding significant improvements. In this work, we propose CyclicPrompt, an innovative cyclic prompt approach designed to enhance the effectiveness, adaptability, and generalizability of UAWR. CyclicPrompt comprises two key components: 1) a composite context prompt that integrates weather-related information and context-aware representations into the network to guide restoration. This prompt differs from previous methods by marrying learnable input-conditional vectors with weather-specific knowledge, thereby improving adaptability across various degradations and 2) the erase-and-paste mechanism, after the initial guided restoration, substitutes weather-specific knowledge with constrained restoration priors, inducing high-quality weather-free concepts into the composite prompt to further fine-tune the restoration process. Therefore, we can form a cyclic "Prompt-Restore-Prompt" pipeline that adeptly harnesses weather-specific knowledge, textual contexts, and reliable textures. Extensive experiments on synthetic and real-world datasets validate the superior performance of CyclicPrompt. The code is available at: https://github.com/RongxinL/CyclicPrompt. Rongxin Liao, Feng Li 0037, Yanyan Wei, Zenglin Shi, Le Zhang 0001, Huihui Bai 0001, Meng Wang 0001 |
IEEE Trans. Image Process. | 2 |
| 2025 | Enhancing Light Field Salient Object Detection With Variance-Maximized Key Focal Slice SelectionabstractLight field saliency object detection (LF SOD) methods have made significant progress recently. Most of them explore abundant multi-modal information from the all-focus image and the focal stacks at all focal planes to enrich scene details and depth perception. However, in light-field images, the spatial and depth information varies slightly across different slices, raising redundancy within focal stacks. Besides, the noise can appear repeatedly in multiple images of the focal stacks, which brings interference. To address these issues, in this work, we propose VMKNet, an effective approach that leverages innovative variance-maximized key slice selection and interacts with the all-focus image, to improve LF SOD. Specifically, we measure consistency differences between the all-focus image and each focal slice in the salient region as saliency scores. Then, we randomly assemble sets of them, where each score corresponds to a certain slice. The one exhibiting the highest variance is singled out to determine key focal slices as they reveal the diversity of salient objects. Then, the bidirectional guidance module (BGM) is presented to learn attentive features of all-focus and selected key slices in a mutual guidance manner, thus producing enhanced and holistic features. With hierarchical BGMs, our model can progressively aggregate common salient semantics and meaningful contextual details, generating more discriminative representations. Moreover, we introduce the edge enhancement module in conjunction with BGM to improve the sharpness of saliency maps. Extensive experiments on common light field datasets demonstrate that our method, termed VMKNet, outperforms recent state-of-the-art LF, RGB-D, and RGB methods. Our code is available athttps://github.com/Han-jiaxin/VMKNet. Jiaxin Han, Feng Li 0037, Mengmeng Zhang 0008, Huihui Bai 0001, Jimin Xiao, Yao Zhao 0001 |
IEEE Trans. Multim. | 2 |
| 2024 | DeepEnhancer: Temporally Consistent Focal Transformer for Comprehensive Video EnhancementabstractRestoring and colorizing old films is a comprehensive video enhancement task, marked by the presence of heterogeneous and structured degradations. Our DeepEnhancer addresses this challenge through a unified pipeline that combines restoration and colorization. In this workflow, we implement a bidirectional propagation strategy. Specifically, we incorporate second-order feature alignment to reduce the accumulation of inaccuracies in optical flow estimation. Simultaneously, we utilize cross-scale long-term attention mechanisms to model correlations within hidden states, thereby ensuring spatial and temporal consistency. To address the notable content loss in aging films, we introduce a temporally consistent focal transformer guided by global information. This transformer utilizes various window levels with distinct sub-window sizes to seamlessly integrate fine-grained and coarse-grained features. Comprehensive experimental results conclusively demonstrate the superiority of our model in both quantitative and qualitative comparisons when compared to existing approaches. Lihua Chi, Wentao Ma 0003, Feng Li 0037, Jie Liu 0002 |
ICMR | 5 |
| 2024 | What can we learn from quality assurance badges in open-source software?
Feng Li 0037, Yiling Lou, Xin Tan 0003, Zhenpeng Chen 0001, Jinhao Dong, Xuanzhi Wang, Dan Hao 0001, Lu Zhang 0023 |
Sci. China Inf. Sci. | 1 |
| 2024 | Exploring Resolution Fields for Scalable Image Compression With Uncertainty GuidanceabstractRecently, there are significant advancements in learning-based image compression methods surpassing traditional coding standards. Most of them prioritize achieving the best rate-distortion performance for a particular compression rate, which limits their flexibility and adaptability in various applications with complex and varying constraints. In this work, we explore the potential of resolution fields in scalable image compression and propose the reciprocal pyramid network (RPN) that fulfills the need for more adaptable and versatile compression. Specifically, RPN first builds a compression pyramid and generates the resolution fields at different levels in a top-down manner. The key design lies in the cross-resolution context mining module between adjacent levels, which performs feature enriching and distillation to mine meaningful contextualized information and remove unnecessary redundancy, producing informative resolution fields as residual priors. The scalability is achieved by progressive bitstream reusing and resolution field incorporation varying at different levels. Furthermore, between adjacent compression levels, we explicitly quantify the aleatoric uncertainty from the bottom decoded representations and develop an uncertainty-guided loss to update the upper-level compression parameters, forming a reverse pyramid process that enforces the network to focus on the textured pixels with high variance for more reliable and accurate reconstruction. Combining resolution field exploration and uncertainty guidance in a pyramid manner, RPN can effectively achieve spatial and quality scalable image compression. Experiments show the superiority of RPN against existing classical and deep learning-based scalable codecs. Code will be available athttps://github.com/JGIroro/RPNSIC. Dongyi Zhang, Feng Li 0037, Man Liu 0003, Runmin Cong, Huihui Bai 0001, Meng Wang 0001, Yao Zhao 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2024 | Remote Sensing Image-Text Retrieval With Implicit-Explicit Relation ReasoningabstractRemote sensing image-text retrieval (RSITR) has become a research hotspot in recent years for its wide application. Existing methods in this context, based either on local or global feature matching, overlook the sensing variation-leaded visual deviation and geographically nearby image-text mismatching problems of remote sensing (RS) images. This work notes that this would limit the retrieval accuracy for RSITR. To handle this, we present IERR, an implicit-explicit relation reasoning framework that learns relations between local visual-textual tokens and enhances global image-text matching without requiring additional prior supervision. Specifically, masked image modeling (MIM) and masked language modeling (MLM) are used for symmetric mask reasoning consistency alignment. Meanwhile, masked features (i.e., implicit relation) and unmasked features (i.e., explicit relation) are fed into a multimodal interaction encoder to enhance the representations of the textual-visual features. Extensive experimental results on the RSICD and RSITMD datasets demonstrate the superiority of IERR compared with 17 baselines. Lingling Yang, Tongqing Zhou, Wentao Ma 0003, Mengze Du, Lu Liu 0023, Feng Li 0037, Shan Zhao 0002, Yuwei Wang 0002 |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2024 | Progressive Complementary Knowledge Aggregation for CdZnTe Defect SegmentationabstractAutomatic quality inspection of industrial products is an indispensable part of modern manufacturing. Cadmium zinc telluride (CdZnTe) crystal is an important industrial raw material, but the special photosensitive properties of CdZnTe make it show different defect boundaries under different lighting angles, which poses challenges for quality inspection. In this article, we propose progressive complementary knowledge aggregation (PCKA) for CdZnTe defect segmentation, which is model-agnostic. First, the 12 images of CdZnTe crystal with different lighting angles are fed into the preliminary aggregation net to aggregate unique pixel-level clues. Second, we use a latent aggregation net to acquire the feature-level complementary clues under the guidance of the pixel-level clues within latent space. Such a learning paradigm is an effective solution for the special photosensitive properties of CdZnTe crystal. Extensive experiments on self-collected dataset demonstrate the effectiveness and efficiency of our PCKA compared with other solutions. Feng Li 0037, Man Liu 0003, Huihui Bai 0001, Yunchao Wei, Anhong Wang, Shijie Ma, Yao Zhao 0001 |
IEEE Trans. Ind. Informatics | 2 |
| 2024 | Reflection Intensity Guided Single Image Reflection Removal and Transmission RecoveryabstractSingle image reflection removal (SIRR) aims at eliminating unwanted interference caused by the reflection of transparent or smooth surfaces and obtaining an estimation of a clear transmission layer. Existing data-driven methods typically rely on decomposing the observed image into transmission and reflection layers, which neglects the physical generation principles of an image with reflections, thus leading to unsatisfactory results, especially in strong reflection regions. To address this issue, in this work, we analyze the imaging process of reflection image from the physical perspective and derive a conclusion that the physical quantity: illuminance of the reflection layer determines the reflection intensity. Then a two-stage reflection intensity-guided network (RINet) is proposed for reflection removal and transmission recovery. The key lies in the first stage are the parallel modules that generate the reflection intensity map and transmission layer. In the second stage, besides utilizing such intensity map as the guidance, we additionally calculate the gradient field as the other prior to facilitate the final reflection removal. Specifically, we design a dual-flow joint learning module (JLM) comprised of a transmission recovery branch and a gradient optimization branch that jointly optimizes image structures and details by exploiting the interactions between transmission and gradient features. In particular, guided by the reflection intensity map, the transmission recovery branch can dynamically focus on removing reflections. Equipped with the two-stage framework, our RINet constitutes a divide-and-conquer process to achieve effective transmission recovery and reflection removal. Experimental results on public datasets demonstrate the superiority of the proposed method over recent state-of-the-art methods. Lingzhi He, Feng Li 0037, Runmin Cong, Yao Zhao 0001 |
IEEE Trans. Multim. | 2 |
| 2024 | SYRER: Synergistic Relational Reasoning for RGB-D Cross-Modal Re-IdentificationabstractRGB-D cross-modal person re-identification is designed to match the people across the RGB and depth image modalities, where the large modality discrepancy makes this task intractable to tackle. To alleviate the negative effect brought by the discrepancy, this paper proposes a novel SYnergistic RElational Reasoning (SYRER) method, which targets at exploring the synergy between hetero-modalities for recognizing persons. We design a heterogeneous relationship contrast branch to establish intra-class and inter-class cross-modal relationships, which implements the cross-modal relation contrast learning to cope with imperceptible cross-modal inter-class differences and large cross-modal intra-class discrepancy. Additionally, in order to adequately represent the irregular depth images, we propose a point-wise depth extractor to extract non-uniform discriminative point features from depth images. Experimental results on two public datasets indicate the proposed SYRER surpasses the state-of-the-arts. And we also perform a series of analytic experiments to verify the effectiveness of each submodule of our SYRER. Hao Liu 0003, Jingjing Wu 0001, Feng Li 0037, Richang Hong |
IEEE Trans. Multim. | 3 |
| 2024 | Enhanced Video Super-Resolution Network towards Compressed DataabstractVideo super-resolution (VSR) algorithms aim at recovering a temporally consistent high-resolution (HR) video from its corresponding low-resolution (LR) video sequence. Due to the limited bandwidth during video transmission, most available videos on the internet are compressed. Nevertheless, few existing algorithms consider the compression factor in practical applications. In this paper, we propose an enhanced VSR model towards compressed videos, termed as ECVSR, to simultaneously achieve compression artifacts reduction and SR reconstruction end-to-end. ECVSR contains a motion-excited temporal adaption network (METAN) and a multi-frame SR network (SRNet). The METAN takes decoded LR video frames as input and models inter-frame correlations via bidirectional deformable alignment and motion-excited temporal adaption, where temporal differences are calculated as motion prior to excite the motion-sensitive regions of temporal features. In SRNet, cascaded recurrent multi-scale blocks (RMSB) are employed to learn deep spatio-temporal representations from adapted multi-frame features. Then, we build a reconstruction module for spatio-temporal information integration and HR frame reconstruction, which is followed by a detail refinement module for texture and visual quality enhancement. Extensive experimental results on compressed videos demonstrate the superiority of our method for compressed VSR. Code will be available at https://github.com/lifengcs/ECVSR . Feng Li 0037, Huihui Bai 0001, Runmin Cong, Yao Zhao 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2023 | Progressive Semantic-Visual Mutual Adaption for Generalized Zero-Shot LearningabstractGeneralized Zero-Shot Learning (GZSL) identifies unseen categories by knowledge transferred from the seen domain, relying on the intrinsic interactions between visual and semantic information. Prior works mainly localize regions corresponding to the sharing attributes. When various visual appearances correspond to the same attribute, the sharing attributes inevitably introduce semantic ambiguity, hampering the exploration of accurate semantic-visual interactions. In this paper, we deploy the dual semantic-visual transformer module (DSVTM) to progressively model the correspondences between attribute prototypes and visual features, constituting a progressive semantic-visual mutual adaption (PSVMA) network for semantic disambiguation and knowledge transferability improvement. Specifically, DSVTM devises an instance-motivated semantic encoder that learns instance-centric prototypes to adapt to different images, enabling the recast of the unmatched semantic-visual pair into the matched one. Then, a semantic-motivated instance decoder strengthens accurate cross-domain interactions between the matched pair for semantic-related instance adaption, en-couraging the generation of unambiguous visual representations. Moreover, to mitigate the bias towards seen classes in GZSL, a debiasing loss is proposed to pursue response consistency between seen and unseen predictions. The PSVMA consistently yields superior performances against other state-of-the-art methods. Code will be available at: https://github.com/ManLiuCoder/PSVMA. Man Liu 0003, Feng Li 0037, Chunjie Zhang 0001, Yunchao Wei, Huihui Bai 0001, Yao Zhao 0001 |
CVPR | 2 |
| 2023 | Learning Detail-Structure Alternative Optimization for Blind Super-ResolutionabstractExisting convolutional neural networks (CNN) based image super-resolution (SR) methods have achieved impressive performance on bicubic kernel, which is not valid to handle unknown degradations in real-world applications. Recent blind SR methods suggest to reconstruct SR images relying on blur kernel estimation. However, their results still remain visible artifacts and detail distortion due to the estimation errors. To alleviate these problems, in this paper, we propose an effective and kernel-free network, namely DSSR, which enables recurrent detail-structure alternative optimization without blur kernel prior incorporation for blind SR. Specifically, in our DSSR, a detail-structure modulation module (DSMM) is built to exploit the interaction and collaboration of image details and structures. The DSMM consists of two components: a detail restoration unit (DRU) and a structure modulation unit (SMU). The former aims at regressing the intermediate HR detail reconstruction from LR structural contexts, and the latter performs structural contexts modulation conditioned on the learned detail maps at both HR and LR spaces. Besides, we use the output of DSMM as the hidden state and design our DSSR architecture from a recurrent convolutional neural network (RCNN) view. In this way, the network can alternatively optimize the image details and structural contexts, achieving co-optimization across time. Moreover, equipped with the recurrent connection, our DSSR allows low- and high-level feature representations complementary by observing previous HR details and contexts at every unrolling time. Extensive experiments on synthetic datasets and real-world images demonstrate that our method achieves the state-of-the-art against existing methods. Feng Li 0037, Huihui Bai 0001, Weisi Lin, Runmin Cong, Yao Zhao 0001 |
IEEE Trans. Multim. | 1 |
| 2022 | Bridging the Gap between Different Programming Paradigms in Coverage-based Fault LocalizationabstractFault localization is to identify faulty program elements. Among the large number of fault localization approaches in the literature, coverage-based fault localization, especially spectrum-based fault localization has been intensively studied due to its effectiveness and lightweightness. Despite the rich literature, almost all existing fault localization approaches and studies are conducted on imperative programming languages such as Java and C, leaving a gap in other programming paradigms. In this paper, we aim to study fault localization approaches for the functional programming paradigm, using Haskell language as a representation. We build up the first dataset on real Haskell projects including both real and seeded faults, which enables the research of fault localization for functional languages. With this dataset, we explore fault localization techniques for Haskell. In particular, as typically for SBFL approaches, we study methods for coverage collection as well as formulae for suspiciousness scores computation, and carefully adapt these two components to Haskell considering the language features and characteristics, resulting in a series of adaption approaches and a learning-based approach, which are evaluated on the dataset to demonstrate the promises of the direction. Feng Li 0037, Meng Wang 0002, Dan Hao 0001 |
Internetware | 1 |
| 2022 | Adaptive feature fusion network based on boosted attention mechanism for single image dehazing
Feng Li 0037, Runmin Cong, Huihui Bai 0001, Yao Zhao 0001 |
Multim. Tools Appl. | 2 |
| 2022 | AGA: An Accelerated Greedy Additional Algorithm for Test Case PrioritizationabstractIn recent years, many test case prioritization (TCP) techniques have been proposed to speed up the process of fault detection. However, little work has taken the efficiency problem of these techniques into account. In this paper, we target the Greedy Additional (GA) algorithm, which has been widely recognized to be effective but less efficient, and try to improve its efficiency while preserving effectiveness. In our Accelerated GA (AGA) algorithm, we use some extra data structures to reduce redundant data accesses in the GA algorithm and thus the time complexity is reduced from O(m2n) to O(kmn) when n > m, where m is the number of test cases, n is the number of program elements, and k is the iteration number. Moreover, we observe the impact of iteration numbers on prioritization efficiency on our dataset and propose to use a specific iteration number in the AGA algorithm to further improve the efficiency. We conducted experiments on 55 open-source subjects. In particular, we implemented each TCP algorithm with two kinds of widely-used input formats, adjacency matrix and adjacency list. Since a TCP algorithm with adjacency matrix is less efficient than the algorithm with adjacency list, the result analysis is mainly conducted based on TCP algorithms with adjacency list. The results show that AGA achieves 5.95X speedup ratio over GA on average, while it achieves the same average effectiveness as GA in terms of Average Percentage of Fault Detected (APFD). Moreover, we conducted an industrial case study on 22 subjects, collected from Baidu, and find that the average speedup ratio of AGA over GA is 44.27X, which indicates the practical usage of AGA in real-world scenarios. Feng Li 0037, Yinzhu Li, Dan Hao 0001, Lu Zhang 0023 |
IEEE Trans. Software Eng. | 1 |
| 2021 | Towards Fast and Accurate Real-World Depth Super-Resolution: Benchmark Dataset and BaselineabstractDepth maps obtained by commercial depth sensors are always in low-resolution, making it difficult to be used in various computer vision tasks. Thus, depth map super-resolution (SR) is a practical and valuable task, which up-scales the depth map into high-resolution (HR) space. However, limited by the lack of real-world paired low-resolution (LR) and HR depth maps, most existing methods use down-sampling to obtain paired training samples. To this end, we first construct a large-scale dataset named "RGB-D-D", which can greatly promote the study of depth map SR and even more depth-related real-world tasks. The "D-D" in our dataset represents the paired LR and HR depth maps captured from mobile phone and Lucid Helios respectively ranging from indoor scenes to challenging outdoor scenes. Besides, we provide a fast depth map super-resolution (FDSR) baseline, in which the high-frequency component adaptively decomposed from RGB image to guide the depth map SR. Extensive experiments on existing public datasets demonstrate the effectiveness and efficiency of our network compared with the state-of-the-art methods. Moreover, for the real-world LR depth maps, our algorithm can produce more accurate HR depth maps with clearer boundaries and to some extent correct the depth value errors. Lingzhi He, Hongguang Zhu, Feng Li 0037, Huihui Bai 0001, Runmin Cong, Chunjie Zhang 0001, Chunyu Lin, Meiqin Liu 0002, Yao Zhao 0001 |
CVPR | 3 |
| 2021 | Towards Complete Scene and Regular Shape for Distortion Rectification by Curve-Aware ExtrapolationabstractThe wide-angle lens gains increasing attention since it can capture a wide field-of-view (FoV) scene. However, the obtained image is contaminated with radial distortion, making the scene not realistic. Previous distortion rectification methods rectify the image in a rectangle or invagination, failing to display the complete content and regular shape simultaneously. In this paper, we rethink the representation of rectification results and present a Rectification OutPainting (ROP) method, aiming to extrapolate the coherent semantics to the blank area and create a wider FoV beyond the original wide-angle lens. To address the specific challenges such as the variable painting region and curve boundary, a rectification module is designed to rectify the image with geometry supervision, and the extrapolated results are generated using a dual conditional expansion strategy. In terms of the spatially discounted correlation, a curve-aware correlation measurement is proposed to focus on the generated region to enforce the local consistency. To our knowledge, we are the first to tackle the challenging rectification via outpainting, and our curve-aware strategy can reach a rectification construction with complete content and regular shape. Extensive experiments well demonstrate the superiority of our ROP over other state-of-the-art solutions. Kang Liao, Chunyu Lin, Yunchao Wei, Feng Li 0037, Shangrong Yang, Yao Zhao 0001 |
ICCV | 4 |
| 2021 | Cross-modality Discrepant Interaction Network for RGB-D Salient Object DetectionabstractThe popularity and promotion of depth maps have brought new vigor and vitality into salient object detection (SOD), and a mass of RGB-D SOD algorithms have been proposed, mainly concentrating on how to better integrate cross-modality features from RGB image and depth map. For the cross-modality interaction in feature encoder, existing methods either indiscriminately treat RGB and depth modalities, or only habitually utilize depth cues as auxiliary information of the RGB branch. Different from them, we reconsider the status of two modalities and propose a novel Cross-modality Discrepant Interaction Network (CDINet) for RGB-D SOD, which differentially models the dependence of two modalities according to the feature representations of different layers. To this end, two components are designed to implement the effective cross-modality interaction: 1) the RGB-induced Detail Enhancement (RDE) module leverages RGB modality to enhance the details of the depth features in low-level encoder stage. 2) the Depth-induced Semantic Enhancement (DSE) module transfers the object positioning and internal consistency of depth features to the RGB branch in high-level encoder stage. Furthermore, we also design a Dense Decoding Reconstruction (DDR) structure, which constructs a semantic block by combining multi-level encoder features to upgrade the skip connection in the feature decoding. Extensive experiments on five benchmark datasets demonstrate that our network outperforms $15$ state-of-the-art methods both quantitatively and qualitatively. Our code is publicly available at:https://rmcong.github.io/proj_CDINet.html. Chen Zhang 0013, Runmin Cong, Qinwei Lin, Lin Ma 0002, Feng Li 0037, Yao Zhao 0001, Sam Kwong |
ACM Multimedia | 5 |
| 2021 | A Study of Bug Resolution Characteristics in Popular Programming LanguagesabstractThis paper presents a large-scale study that investigates the bug resolution characteristics among popular Github projects written in different programming languages. We explore correlations but, of course, we cannot infer causation. Specifically, we analyse bug resolution data from approximately 70 million Source Line of Code, drawn from 3 million commits to 600 GitHub projects, primarily written in 10 programming languages. We find notable variations in apparent bug resolution time and patch (fix) size. While interpretation of results from such large-scale empirical studies is inherently difficult, we believe that the differences in medians are sufficiently large to warrant further investigation, replication, re-analysis and follow up research. For example, in our corpus, the median apparent bug resolution time (elapsed time from raise to resolve) for Ruby was 4X that for Go and 2.5X for Java. We also found that patches tend to touch more files for the corpus of strongly typed and for statically typed programs. However, we also found evidence for alowerelapsed resolution time for bug resolution committed to projects constructed from statically typed languages. These findings, if replicated in subsequent follow on studies, may shed further empirical light on the debate about the importance of static typing. Jie Zhang 0050, Feng Li 0037, Dan Hao 0001, Meng Wang 0002, Lu Zhang 0023, Mark Harman |
IEEE Trans. Software Eng. | 2 |
| 2020 | Deep Interleaved Network for Single Image Super-Resolution with Asymmetric Co-AttentionabstractRecently, Convolutional Neural Networks (CNN) based image super-resolution (SR) have shown significant success in the literature. However, these methods are implemented as single-path stream to enrich feature maps from the input for the final prediction, which fail to fully incorporate former low-level features into later high-level features. In this paper, to tackle this problem, we propose a deep interleaved network (DIN) to learn how information at different states should be combined for image SR where shallow information guides deep representative features prediction. Our DIN follows a multi-branch pattern allowing multiple interconnected branches to interleave and fuse at different states. Besides, the asymmetric co-attention (AsyCA) is proposed and attacked to the interleaved nodes to adaptively emphasize informative features from different states and improve the discriminative ability of networks. Extensive experiments demonstrate the superiority of our proposed DIN in comparison with the state-of-the-art SR methods. Feng Li 0037, Runming Cong, Huihui Bai 0001 |
IJCAI | 1 |
| 2020 | Cost-Effective Testing of a Deep Learning Model through Input ReductionabstractWith the increasing adoption of Deep Learning (DL) models in various applications, testing DL models is vitally important. However, testing DL models is costly and expensive, e.g., manual labelling is widely-recognized to be costly. To reduce testing cost, we propose to select only a subset of testing data, which is small but representative enough for a quick estimation of the performance of DL models. Our approach, DeepReduce, adopts a two-phase strategy. At first, our approach selects testing data for the purpose of satisfying testing adequacy. Then, it selects more testing data to approximate the distribution between the whole testing data and the selected data by leveraging relative entropy minimization. We evaluate DeepReduce on four widely-used datasets (with 15 models in total). We find that DeepReduce reduces the whole testing data to 7.5% on average and can reliably estimate the performance of DL models. Feng Li 0037, Jinhao Dong, Hongyu Zhang 0002, Dan Hao 0001 |
ISSRE | 2 |
| 2020 | FilterNet: Adaptive Information Filtering Network for Accurate and Fast Image Super-ResolutionabstractDeep convolutional neural network (CNN) approaches have achieved impressive performance for image super-resolution (SR). The main issue of image SR is to effectively recover the high-frequency detail of low-resolution (LR) input. However, existing CNN methods often inevitably exhibit a large amount of memory consumption and computational cost. In addition, in most SR networks, the low-frequency and high-frequency components of the LR features are treated equally in the training process, which can ignore the local detailed information and hinder the representational capacity of networks. To solve these issues, in this paper, we propose a deep adaptive information filtering network (FilterNet) for accurate and fast image SR. In contrast to the existing methods that adopt fully CNN methods to directly predict the HR images, the proposed FilterNet concentrates on more useful features and adaptively filters the redundant low-frequency information. In general, we present the dilated residual group (DRG), which consists of multiple dilated residual units. The DRGs can directly expand the receptive field of the network to efficiently exploit the contextual information of the LR input. In the dilated residual unit, a gated selective mechanism is proposed to adaptively learn more high-frequency information and filter the low-frequency information. Besides, we introduce a novel adaptive information fusion structure, which builds long scaling skip connections among the DRGs to rescale the hierarchical features and fuse more detailed information. The scaling weights can be deemed as the part parameters of our network and trained adaptively. The Extensive evaluations on benchmark datasets demonstrate that our FilterNet achieves superior performance both on accuracy and speed compared with recent state-of-the-art methods. Feng Li 0037, Huihui Bai 0001, Yao Zhao 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2020 | Learning a Deep Dual Attention Network for Video Super-ResolutionabstractRecently, deep learning based video super-resolution (SR) methods combine the convolutional neural networks (CNN) with motion compensation to estimate a high-resolution (HR) video from its low-resolution (LR) counterpart. However, most previous methods conduct downscaling motion estimation to handle large motions, which can lead to detrimental effects on the accuracy of motion estimation due to the reduction of spatial resolution. Besides, these methods usually treat different types of intermediate features equally, which lack flexibility to emphasize meaningful information for revealing the high-frequency details. In this paper, to solve above issues, we propose a deep dual attention network (DDAN), including a motion compensation network (MCNet) and a SR reconstruction network (ReconNet), to fully exploit the spatio-temporal informative features for accurate video SR. The MCNet progressively learns the optical flow representations to synthesize the motion information across adjacent frames in a pyramid fashion. To decrease the mis-registration errors caused by the optical flow based motion compensation, we extract the detail components of original LR neighboring frames as complementary information for accurate feature extraction. In the ReconNet, we implement dual attention mechanisms on a residual unit and form a residual attention unit to focus on the intermediate informative features for high-frequency details recovery. Extensive experimental results on numerous datasets demonstrate the proposed method can effectively achieve superior performance in terms of quantitative and qualitative assessments compared with state-of-the-art methods. Feng Li 0037, Huihui Bai 0001, Yao Zhao 0001 |
IEEE Trans. Image Process. | 1 |
| 2019 | Detail-preserving image super-resolution via recursively dilated residual network
Feng Li 0037, Huihui Bai 0001, Yao Zhao 0001 |
Neurocomputing | 1 |
| 2012 | Collaborative visual modeling for automatic image annotation via sparse model coding
Feng Li 0037, Meng Wang 0001 |
Neurocomputing | 2 |