VLDB 2026 Research / reviewers in the wild / expert
Kui Jiang
dblp:124/2034
· DBLP profile ↗
129ranked-venue papers
23as first author
115since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 90 · 17 first-author · 82 since 2021Artificial intelligence and machine learning · 64 · 9 first-author · 58 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 2 first-author · 6 since 2021Computer networks · 4 · 4 since 2021Systems, architecture and hardware · 2 · 1 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Learning Depth from Past Selves: Self-Evolution Contrast for Robust Depth EstimationabstractSelf-supervised depth estimation has gained significant attention in autonomous driving and robotics. However, existing methods exhibit substantial performance degradation under adverse weather conditions such as rain and fog, where reduced visibility critically impairs depth prediction. To address this issue, we propose a novel self-evolution contrastive learning framework called SEC-Depth for self-supervised robust depth estimation tasks. Our approach leverages intermediate parameters generated during training to construct temporally evolving latency models. Using these, we design a self-evolution contrastive scheme to mitigate performance loss under challenging conditions. Concretely, we first design a dynamic update strategy of latency models for the depth estimation task to capture optimization states across training stages. To effectively leverage latency models, we introduce a self-evolution contrastive Loss (SECL) that treats outputs from historical latency models as negative samples. This mechanism adaptively adjusts learning objectives while implicitly sensing weather degradation severity, reducing the needs for manual intervention. Experiments show that our method integrates seamlessly into diverse baseline models and significantly enhances robustness in zero-shot evaluations. Kui Jiang, Shenyi Li |
AAAI | 2 |
| 2026 | MambaOVSR: Multiscale Fusion with Global Motion Modeling for Chinese Opera Video Super-ResolutionabstractChinese opera is celebrated for preserving classical art. However, early filming equipment limitations have degraded videos of last-century performances by renowned artists (e.g., low frame rates and resolution), hindering archival efforts. Although space-time video super-resolution (STVSR) has advanced significantly, applying it directly to opera videos remains challenging. The scarcity of datasets impedes the recovery of high-frequency details, and existing STVSR methods lack global modeling capabilities—compromising visual quality when handling opera’s characteristic large motions. To address these challenges, we pioneer a large-scale Chinese Opera Video Clip (COVC) dataset and propose the Mamba-based multiscale fusion network for space-time Opera Video Super-Resolution (MambaOVSR). Specifically, MambaOVSR involves three novel components: the Global Fusion Module (GFM) for motion modeling through a multiscale alternating scanning mechanism, and the Multiscale Synergistic Mamba Module (MSMM) for alignment across different sequence lengths. Additionally, our MambaVR block resolves feature artifacts and positional information loss during alignment. Experimental results on the COVC dataset show that MambaOVSR significantly outperforms the SOTA STVSR method by an average of 1.86 dB in terms of PSNR. Hua Chang, Xin Xu 0007, Wei Liu 0183, Wei Wang 0170, Xin Yuan 0009, Kui Jiang |
AAAI | 6 |
| 2026 | ICLR: Inter-Chrominance and Luminance Interaction for Natural Color Restoration in Low-Light Image EnhancementabstractLow-Light Image Enhancement (LLIE) task aims at improving contrast while restoring details and textures for images captured in low-light conditions. HVI color space has made significant progress in this task by enabling precise decoupling of chrominance and luminance. However, for the interaction of chrominance and luminance branches, substantial distributional differences between the two branches prevalent in natural images limit complementary feature extraction, and luminance errors are propagated to chrominance channels through the nonlinear parameter. Furthermore, for interaction between different chrominance branches, images with large homogeneous-color regions usually exhibit weak correlation between chrominance branches due to concentrated distributions. Traditional pixel-wise losses exploit strong inter-branch correlations for co-optimization, causing gradient conflicts in weakly correlated regions. Therefore, we propose an Inter-Chrominance and Luminance Interaction (ICLR) framework including a Dual-stream Interaction Enhancement Module (DIEM) and a Covariance Correction Loss (CCL). The DIEM improves the extraction of complementary information from two dimensions, fusion and enhancement, respectively. The CCL utilizes luminance residual statistics to penalize chrominance errors and balances gradient conflicts by constraining chrominance branches covariance. Experimental results on multiple datasets show that the proposed ICLR framework outperforms state-of-the-art methods. Xin Xu 0007, Wei Liu 0183, Wei Wang 0170, Kui Jiang |
AAAI | 6 |
| 2026 | Semantics and Content Matter: Towards Multi-Prior Hierarchical Mamba for Image DerainingabstractRain significantly degrades the performance of computer vision systems, particularly in applications like autonomous driving and video surveillance. While existing deraining methods have made considerable progress, they often struggle with fidelity of semantic and spatial details. To address these limitations, we propose the Multi-Prior Hierarchical Mamba (MPHM) network for image deraining. This novel architecture synergistically integrates macro-semantic textual priors (CLIP) for task-level semantic guidance and micro-structural visual priors (DINOv2) for scene-aware structural information. To alleviate potential conflicts between heterogeneous priors, we devise a progressive Priors Fusion Injection (PFI) that strategically injects complementary cues at different decoder levels. Meanwhile, we equip the backbone network with an elaborate Hierarchical Mamba Module (HMM) to facilitate robust feature representation, featuring a Fourier-enhanced dual-path design that concurrently addresses global context modeling and local detail recovery. Comprehensive experiments demonstrate MPHM's state-of-the-art performance, achieving a 0.57 dB PSNR gain on the Rain200H dataset while delivering superior generalization on real-world rainy scenarios. Zhaocheng Yu, Kui Jiang, Junjun Jiang, Xianming Liu 0005, Guanglu Sun, Yi Xiao 0003 |
AAAI | 2 |
| 2026 | MCMC with adaptive principal-component transformation: rotation-invariant universal samplers for bayesian structural system identification
Xianghao Meng, James L. Beck, Kui Jiang, Hui Li 0035 |
Adv. Eng. Informatics | 4 |
| 2026 | SfMamba: Efficient source-free domain adaptation via selective scan modeling
Xi Chen 0110, Hongxun Yao, Sicheng Zhao, Jiankun Zhu, Kui Jiang |
Expert Syst. Appl. | 6 |
| 2026 | Learning from History: Task-agnostic Model Contrastive Learning for Image Restoration
Gang Wu 0010, Junjun Jiang, Kui Jiang, Xianming Liu 0005, Wangmeng Zuo |
Int. J. Comput. Vis. | 3 |
| 2026 | CMPF: Harmonizing Cross-Model Prior Fusion for Open-Vocabulary Segmentation
Sicheng Zhao, Xi Chen 0110, Hongxun Yao, Haosen Yang 0003, Yanhao Zhang 0001, Sheng Jin 0002, Xiatian Zhu, Haonan Lu, Kui Jiang, Guiguang Ding |
Int. J. Comput. Vis. | 9 |
| 2026 | A Natural Language Guided Approach for Blind Face Restoration: Methodology and DatasetabstractBlind Face Restoration (BFR) aims to reconstruct high-quality face images from low-quality inputs without any prior knowledge of the specific degradation types or levels. In recent years, remarkable progress has been achieved, particularly through GAN- and diffusion-based approaches, which have greatly improved perceptual realism and reconstruction fidelity. However, existing approaches typically rely solely on visual cues from degraded images. This often results in inaccurate reconstruction of facial details and noticeable identity distortion, particularly under severe or complex degradations. To address these limitations, we incorporate auxiliary textual information into BFR to enable the recovery of subtle facial attributes, such as wrinkles, moles, and skin marks that are often overlooked or hard to reconstruct by conventional visual priors. To support this idea, we first construct a large-scale dataset containing 30,000 detailed textual descriptions paired with CelebA-HQ face images, explicitly designed to capture fine-grained facial semantics. To effectively bridge the gap between visual data and natural language, we further propose FaceCLIP, a fine-tuned vision-language model specifically tailored to the human face. FaceCLIP enables more accurate alignment between face images and their corresponding textual descriptions by effectively capturing nuanced semantic cues critical for faithful face reconstruction. Built upon these foundations, we propose Text-guided Blind Face Restoration (TBFR), a novel diffusion-based framework that explicitly integrates textual guidance into the face restoration pipeline. Within TBFR, a text-guided hybrid attention block is designed to effectively fuse visual and textual features, while a text-aware loss is employed to enforce semantic consistency between the generated images and their associated textual descriptions. Extensive experimental results show that TBFR outperforms state-of-the-art BFR methods in terms of both quantitative metrics and subjective perceptual quality, establishing a new benchmark for BFR tasks. Wenjie An, Chenyang Wang 0002, Junjun Jiang, Kui Jiang, Xianming Liu 0005, Liqiang Nie |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2026 | Beyond Degradation Redundancy: Contrastive Prompt Learning for All-in-One Image RestorationabstractAll-in-one image restoration, addressing diverse degradation types with a unified model, presents significant challenges in designing task-aware prompts that effectively guide restoration across multiple degradation scenarios. While adaptive prompt learning enables end-to-end optimization, it often yields overlapping or redundant task representations. Conversely, explicit prompts derived from pretrained classifiers enhance discriminability but may discard critical visual information for reconstruction. To address these limitations, we introduce Contrastive Prompt Learning (CPL), a novel framework that fundamentally enhances prompt-task alignment through two complementary innovations: a Sparse Prompt Module (SPM) that efficiently captures degradation-specific features while minimizing redundancy, and a Contrastive Prompt Regularization (CPR) that explicitly strengthens task boundaries by incorporating negative prompt samples across different degradation types. Unlike previous approaches that focus primarily on degradation classification, CPL optimizes the critical interaction between prompts and the restoration model itself. Extensive experiments across comprehensive benchmarks demonstrate that CPL consistently enhances state-of-the-art all-in-one restoration models, achieving significant improvements in both standard multi-task scenarios and challenging composite degradation settings. Our framework establishes new state-of-the-art performance while maintaining parameter efficiency, offering a principled solution for unified image restoration. The code is available at https://github.com/Aitical/CPLIR. Gang Wu 0010, Junjun Jiang, Kui Jiang, Xianming Liu 0005, Liqiang Nie |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2026 | DSwinIR: Rethinking Window-Based Attention for Image RestorationabstractImage restoration has witnessed significant advancements with the development of deep learning models. Transformer-based models, particularly those using window-based self-attention, have become a dominant force. However, their performance is constrained by the rigid, non-overlapping window partitioning scheme, which leads to insufficient feature interaction across windows and limited receptive fields. This highlights the need for more adaptive and flexible attention mechanisms. In this paper, we propose the Deformable Sliding Window Transformer for Image Restoration (DSwinIR), a new attention mechanism: the Deformable Sliding Window (DSwin) Attention. This mechanism introduces a token-centric and content-aware paradigm that moves beyond the grid and fixed window partition. It comprises two complementary components. First, it replaces the rigid partitioning with a token-centric sliding window paradigm, making it effective at eliminating boundary artifacts. Second, it incorporates a content-aware deformable sampling strategy, which allows the attention mechanism to learn data-dependent offsets and actively shape its receptive field to focus on the most informative image regions. Extensive experiments show that DSwinIR achieves strong results, including state-of-the-art performance on several evaluated benchmarks. For instance, in all-in-one image restoration, our DSwinIR surpasses the most recent backbone GridFormer by 0.53 dB on the three-task benchmark and 0.87 dB on the five-task benchmark. Gang Wu 0010, Junjun Jiang, Kui Jiang, Xianming Liu 0005, Liqiang Nie |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2026 | Image dehazing via RGB-FIR multimodal fusion and collaborative learning
Ruolin Du, Han Wang 0018, Wenjie Liu 0004, Guangcheng Wang, Kui Jiang, Hanseok Ko |
Pattern Recognit. | 5 |
| 2026 | AM40: Enhancing action recognition through matting-driven interaction analysis
Wenxuan Liu 0008, Kui Jiang, Siyuan Yang 0001, Chia-Wen Lin, Xian Zhong |
Pattern Recognit. | 4 |
| 2026 | A database and model for the PM2.5 concentration measurement with visible and infrared imaging
Hongxing Jiang, Guangcheng Wang, Kui Jiang |
Pattern Recognit. | 5 |
| 2026 | M2DAO-Talker: Harmonizing Multi-Granular Motion Decoupling and Alternating Optimization for Talking-Head GenerationabstractAudio-driven talking head generation holds significant potential for film production. While existing 3D methods have advanced motion modeling and content synthesis, they often produce rendering artifacts, such as motion blur, temporal jitter, and local penetration, due to limitations in representing stable, fine-grained motion fields. Through systematic analysis, we reformulate talking head generation into a unified framework comprising three steps: video preprocessing, motion representation, and rendering reconstruction. This framework underpins our proposed M2DAO-Talker, which addresses current limitations via multi-granular motion decoupling and alternating optimization. Specifically, we devise a novel 2D portrait preprocessing pipeline to extract frame-wise deformation control conditions (motion region segmentation masks, and camera parameters) to facilitate motion representation. To ameliorate motion modeling, we elaborate a multi-granular motion decoupling strategy, which independently models non-rigid (oral and facial) and rigid (head) motions for improved reconstruction accuracy. Meanwhile, a motion consistency constraint is developed to ensure head-torso kinematic consistency, thereby mitigating penetration artifacts caused by motion aliasing. In addition, an alternating optimization strategy is designed to iteratively refine facial and oral motion parameters, enabling more realistic video generation. Experiments across multiple datasets show that M2DAO-Talker achieves state-of-the-art performance, with the 2.43 dB PSNR improvement in generation quality and 0.64 gain in user-evaluated video realness versus TalkingGaussian while with 150 FPS inference speed. Our project homepage is https://m2dao-talker.github.io/M2DAO-Talk.github.io. Kui Jiang, Junjun Jiang, Hongxun Yao, Xiaopeng Fan 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2026 | SGCNeRF: Few-Shot Neural Rendering via Sparse Geometric Consistency GuidanceabstractNeural Radiance Field (NeRF) technology has made significant strides in creating novel viewpoints. However, its effectiveness is hampered when working with sparsely available views, often leading to performance dips due to overfitting. FreeNeRF attempts to overcome this limitation by integrating implicit geometry regularization, which incrementally improves both geometry and textures. Nonetheless, an initial low positional encoding bandwidth results in the exclusion of high-frequency elements. The quest for a holistic approach that simultaneously addresses overfitting and the preservation of high-frequency details remains ongoing. This study presents a novel feature-matching-based sparse geometry regularization module, enhanced by a spatially consistent geometry filtering mechanism and a frequency-guided geometric regularization strategy. This module excels at accurately identifying high-frequency keypoints, effectively preserving fine structural details. Through progressive refinement of geometry and textures across NeRF iterations, we unveil an effective few-shot neural rendering architecture, designated as SGCNeRF, for enhanced novel view synthesis. Our experiments demonstrate that SGCNeRF not only achieves superior geometry-consistent outcomes but also surpasses FreeNeRF, with improvements of 0.7 dB in PSNR on LLFF and DTU. Yuru Xiao, Xianming Liu 0005, Deming Zhai, Kui Jiang, Junjun Jiang, Xiangyang Ji |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2026 | TH-Mamba: Spatial-Temporal Correlation Learning for Mamba-Based Talking Head GenerationabstractTalking head generation aims to synthesize high-quality and lip-synchronized talking head videos from the given portrait images and audio. However, previous methods directly learn the alignment between lip movements and the driven audio, barely focusing on the fidelity and continuity of the generated videos, suffering from visual distortions and jitter. To deal with this issue, we propose to promote the consistency of audio and image by exploring their spatiotemporal relations, and construct a Mamba-based spatiotemporal fusion scheme. Specifically, we devise an Intra-frame Mamba module to characterize facial features from the source image, which encourages the content consistence between the generated frame and the current source frame. Meanwhile, an Inter-frame Mamba module is designed to excavate the complementary information across sequential frames, which provides clues for better motion simulation. The aggregated spatiotemporal representation with audio features are then aligned with a deformation network to alleviate visual distortions and jitter. In addition, we investigate the practical composite constraints on the structure, details, and motion aspects, involving the keypoint constraint, multi-scale content constraint, and displacement constraint to promote the training stability and model performance. With the above strategies, we construct a novel Talking Head Mamba network, termed as TH-Mamba for high-quality talking head generation. Extensive experiments on the HDTF and Mead-Neutral datasets verify the superiority of our proposed TH-Mamba, which significantly outperforms the current state-of-the-art method by 0.78dB and 1.55dB in PSNR, respectively. The demo is available at https://github.com/YZX-codesky/TH-Mamba. Xin Xu 0007, Zhixi Yu, Kui Jiang, Chia-Wen Lin |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2026 | PH-Mamba: Enhancing Mamba With Position Encoding and Harmonized Attention for Image Deraining and BeyondabstractMamba and its variants excel at modeling long-range dependencies with linear computational complexity, making them effective for diverse vision tasks. However, Mamba's reliance on unfolding 1D sequential representations necessitates multiple directional scans to recover lost spatial dependencies. This introduces significant computational overhead, redundant token traversal, and inefficiencies that compromise accuracy in real-world applications. To this end, we propose PH-Mamba, a novel framework integrating position encoding and harmonized attention for image deraining and beyond. PH-Mamba transforms Mamba's scanning process into a position-guided, unidirectional scanning that selectively prioritizes degradation-relevant tokens. Specifically, we devise a position-guided hybrid Mamba module (PHMM) that jointly encodes perturbation features alongside their spatial coordinates and harmonized representation to model consistent degradation patterns. Within PHMM, a harmonized Transformer is developed to focus on uncertain regions while suppressing noise interference, thereby improving spatial modeling fidelity. Additionally, we employ a vector decomposition and synthesis strategy to enable the unified representation layout to global degradation by directional scanning while minimizing redundancy. By cascading multiple PHMM blocks, PH-Mamba combines global positional guidance with local differential features to strengthen contextual learning. Extensive experiments demonstrate the superiority of PH-Mamba across low-level image restoration benchmarks. For example, compared to NeRD, PH-Mamba achieves a 0.60 dB PSNR improvement while requiring 88.9% fewer parameters, 36.2% less computation, and 63.0% faster inference time. Kui Jiang, Junjun Jiang, Xianming Liu 0005, Hongxun Yao, Chia-Wen Lin |
IEEE Trans. Image Process. | 1 |
| 2026 | VDMamba: Vector Decomposition in Vision Mamba for Image Deraining and BeyondabstractImage deraining aims to remove rain perturbations from rainy images and restore clear backgrounds. Recent research has employed the Mamba technique for image restoration, achieving exceptional results due to its effectiveness and efficiency in modeling long-range sequence relationships. However, a significant challenge remains: developing a comprehensive framework that considers the intrinsic coupling characteristics between image deraining and the Mamba architecture is largely unexplored. We propose that introducing a 1D sequential representation of Mamba could enhance image deraining by characterizing the direction-aware distribution of rain perturbations. This motivates us to introduce a new vector decomposition-based vision Mamba approach (VDMamba). This method investigates vector decomposition within the context of vision Mamba, addressing the challenging task of image deraining and beyond in the frequency embedding space. The key innovation of VDMamba is the Mamba-based vector decomposition and synthesis module (VDSM). This module derives 1D basic vectors (vertical and horizontal) from the frequency components via vector decomposition and employs the single-direction scanning of Mamba to eliminate the direction-specific degradation perturbation. This transformation allows the incipient Mamba to explore directionspecific global relationships for accurate perturbation learning, without requiring an elaborate design of the Mamba scanning. Additionally, the vertical and horizontal components in VDSM are encoded jointly in a bidirectional coupling manner, enabling the exploration of complementary and redundant components for refinement. Experiments on various image enhancement tasks, including image deraining, raindrop removal, rain haze removal, image dehazing, low-light image enhancement, and underwater image enhancement, demonstrate that VDMamba delivers competitive performance compared to the NeRD method. Specifically, it achieves a 0.58 dB improvement in PSNR for the image deraining task while reducing model parameters by 94.3%, computational cost by 88.3%, and inference time by 77.5%. Kui Jiang, Junjun Jiang, Shiqi Wang 0001, Wenqi Ren, Chia-Wen Lin, Zhengguo Li |
IEEE Trans. Multim. | 1 |
| 2026 | S2ML: Spatio-Spectral Mutual Learning for Depth CompletionabstractThe raw depth images captured by RGB-D cameras using Time-of-Flight (TOF) or structured light often suffer from incomplete depth values due to weak reflections, boundary shadows, and artifacts, which limit their applications in downstream vision tasks. Existing methods address this problem through depth completion in the image domain, but they overlook the physical characteristics of raw depth images. It has been observed that the presence of invalid depth areas alters the frequency distribution pattern. In this work, we propose a Spatio-Spectral Mutual Learning framework (S2ML) to harmonize the advantages of both spatial and frequency domains for depth completion. Specifically, we consider the distinct properties of amplitude and phase spectra and devise a dedicated spectral fusion module. Meanwhile, the local and global correlations between spatial-domain and frequency-domain features are calculated in a unified embedding space. The gradual mutual representation and refinement encourage the network to fully explore complementary physical characteristics and priors for more accurate depth completion. Extensive experiments demonstrate the effectiveness of our proposed S2ML method, outperforming the state-of-the-art method CFormer by 0.828 dB and 0.834 dB on the NYU-Depth V2 and SUN RGB-D datasets, respectively. Zihui Zhao, Zheng Wang 0007, Yang Li 0104, Kui Jiang, Zihan Geng, Chia-Wen Lin |
IEEE Trans. Multim. | 5 |
| 2026 | Edge-Enhanced Calcaneal Fracture Segmentation Using a Multi-Task CNN-Transformer Hybrid NetworkabstractTraditional fracture diagnosis relies heavily on the experience of clinicians and the interpretation of medical imaging. In complex cases, the inefficiency of manual interpretation often leads to misdiagnosis or missed detection, underscoring the need for automated segmentation techniques. A major challenge in calcaneal fracture image segmentation lies in the blurred and irregular boundaries of fractures, coupled with the scarcity of high-quality annotated data. To address these issues, this study independently constructs the first dataset specifically dedicated to Calcaneal Fracture segmentation, termed CalFrac. This dataset, collected from Ruijin Hospital in Shanghai, comprises CT scans of calcaneal fractures from 139 patients, along with corresponding pixel-level annotated ground truth segmentation masks. In addition, we propose the Calcaneal Fracture segmentation-Edge detection Network (CFE-Net), a multi-task CNN-Transformer hybrid architecture that employs a dual-branch structure to jointly perform fracture segmentation and edge detection. The main segmentation network adopts an encoder–decoder design to localize the fracture region, while the edge detection branch extracts boundary information and refines the segmentation via cross-branch feature interaction. Experiments on the CalFrac dataset compare CFE-Net with eight state-of-the-art methods. CFE-Net achieves superior performance across all evaluation metrics, demonstrating its advantages in both region integrity and boundary delineation. We have released the dataset and code at https://github.com/esdszdx0/CalFrac-Dataset . Xinfan Zhu, Guangcheng Wang, Lijuan Tang, Kui Jiang |
ACM Trans. Multim. Comput. Commun. Appl. | 6 |
| 2025 | Disentangle Nighttime Lens Flares: Self-supervised Generation-based Lens Flare RemovalabstractLens flares arise from light reflection and refraction within sensor arrays, whose diverse types include glow, veiling glare, reflective flare and so on. Existing methods are specialized for one specific type only, and overlook the simultaneous occurrence of multiple typed lens flares, which is common in the real-world, e.g. coexistence of glow and displacement reflections from the same light source. These co-occurring lens flares cannot be effectively resolved by the simple combination of individual flare removal methods, since these coexisting flares originates from the same light source and are generated simultaneously within the same sensor array, exhibit a complex interdependence rather than simple additive relation. To model this interdependent flares’ relationship, our Nighttime Lens Flare Formation model is the first attempt to learn the intrinsic physical relationship between flares on the imaging plane. Building on this physical model, we introduce a solution to this joint flare removal task named Self-supervised Generation-based Lens Flare Removal Network (SGLFR-Net), which is self-supervised without pre-training. Specifically, the nighttime glow is detangled in PSF Rendering Network(PSFR-Net) based on PSF Rendering Prior, while the reflective flare is modelled in Texture Prior Based Reflection Flare Removal Network (TPRR-Net). Empirical evaluations demonstrate the effectiveness of the proposed method in both joint and individual glare removal tasks. Yuwen He, Wanyu Wu, Kui Jiang |
AAAI | 4 |
| 2025 | The Parables of the Mustard Seed and the Yeast: Extremely Low-Budget, High-Performance Nighttime Semantic SegmentationabstractNighttime Semantic Segmentation (NSS) is essential to many cutting-edge vision applications. However, existing technologies overly rely on massive labeled data, whose annotation is time-consuming and laborious. In this paper, we pioneer a new task focusing on exploring the potential of training strategy and framework design with limited annotation to achieve high-performance NSS. Insufficient information at very low labeling budgets can easily lead to under-optimization or overfitting of the model. Our solution comprises two main components: i) a novel region-based active sampling strategy called Contextual-Aware Region Query (CARQ), which identifies highly informative target nighttime regions for labeling; and ii) an innovative Fragmentation Synergy Active Domain Adaptation framework (FS-ADA), which progressively broadcasts the limited annotation to the unlabeled regions, achieving high performance with a minimal annotation budget. Extensive experiments demonstrate that our method outperforms state-of-the-art UDA-NSS & ADA-SS methods across four day-to-nighttime benchmarks, and generalizes well to foggy, rainy, & snowy scenes. In particular only with 1% target nighttime data annotation, our method is on par with the mainstream fully-supervised methods on the BDD100K-Night val dataset. Shiqin Wang, Xin Xu 0007, Kui Jiang, Zheng Wang 0007 |
AAAI | 4 |
| 2025 | Debiased All-in-one Image Restoration with Task Uncertainty RegularizationabstractAll-in-one image restoration is a fundamental low-level vision task with significant real-world applications. The primary challenge lies in addressing diverse degradations within a single model. While current methods primarily exploit task prior information to guide the restoration models, they typically employ uniform multi-task learning, overlooking the heterogeneity in model optimization across different degradation tasks. To eliminate the bias, we propose a task-aware optimization strategy, that introduces adaptive task-specific regularization for multi-task image restoration learning. Specifically, our method dynamically weights and balances losses for different restoration tasks during training, encouraging the implementation of the most reasonable optimization route. In this way, we can achieve more robust and effective model training. Notably, our approach can serve as a plug-and-play strategy to enhance existing models without requiring modifications during inference. Extensive experiments in diverse all-in-one restoration settings demonstrate the superiority and generalization of our approach. For example, AirNet retrained with TUR achieves average improvements of 1.16 dB on three distinct tasks and 1.81 dB on five distinct all-in-one tasks. These results underscore TUR's effectiveness in advancing the SOTAs in all-in-one image restoration, paving the way for more robust and versatile image restoration. Gang Wu 0010, Junjun Jiang, Kui Jiang, Xianming Liu 0005 |
AAAI | 4 |
| 2025 | Spatial Annealing for Efficient Few-shot Neural RenderingabstractNeural Radiance Fields (NeRF) with hybrid representations have shown impressive capabilities for novel view synthesis, delivering high efficiency. Nonetheless, their performance significantly drops with sparse input views. Various regularization strategies have been devised to address these challenges. However, these strategies either require additional rendering costs or involve complex pipeline designs, leading to a loss of training efficiency. Although FreeNeRF has introduced an efficient frequency annealing strategy, its operation on frequency positional encoding is incompatible with the efficient hybrid representations. In this paper, we introduce an accurate and efficient few-shot neural rendering method named Spatial Annealing regularized NeRF (SANeRF), which adopts the pre-filtering design of a hybrid representation. We initially establish the analytical formulation of the frequency band limit for a hybrid architecture by deducing its filtering process. Based on this analysis, we propose a universal form of frequency annealing in the spatial domain, which can be implemented by modulating the sampling kernel to exponentially shrink from an initial one with a narrow grid tangent kernel spectrum. This methodology is crucial for stabilizing the early stages of the training phase and significantly contributes to enhancing the subsequent process of detail refinement. Our extensive experiments reveal that, by adding merely one line of code, SANeRF delivers superior rendering quality and much faster reconstruction speed compared to current few-shot neural rendering methods. Notably, SANeRF outperforms FreeNeRF on the Blender dataset, achieving 700X faster reconstruction speed. Yuru Xiao, Deming Zhai, Wenbo Zhao 0004, Kui Jiang, Junjun Jiang, Xianming Liu 0005 |
AAAI | 4 |
| 2025 | OODML: Whole Slide Image Classification Meets Online Pseudo-Supervision and Dynamic Mutual LearningabstractBag-label-based multi-instance learning (MIL) has demonstrated significant performance in whole slide image (WSI) analysis, particularly in pseudo-label-based learning schemes. However, due to inaccurate feature representation and interference, existing MIL methods often yield unreliable pseudo-labels, which spawn undesired predictions. To address these issues, we propose an Online Pseudo-Supervision and Dynamic Mutual Learning (OODML) framework that enhances pseudo-label generation and feature representation while exploring their mutual learning to improve bag-level prediction. Specifically, we design an Adaptive Memory Bank (AMB) to collect the most informative components of the current WSI. We also introduce a Self-Progressive Feature Fusion (SPFF) module that integrates label-related historical information from the AMB with current semantic variations, thereby enhancing the representation of pseudo-bag tokens. Furthermore, we propose a Decision Revision Pseudo-Label (DRPL) generation scheme to explore intrinsic connections between pseudo-bag representations and bag-label predictions, resulting in more reliable pseudo-label generation. To alleviate redundant and ambiguous representations, the class-wise prior of pseudo-label prediction is borrowed to facilitate label-related feature learning and to update the AMB, forming a mutual refinement between feature representation and pseudo-label generation. Additionally, a Dynamic Decision-Making (DDM) module is developed to harmonize explicit and implicit representations of bag information for more robust decision-making. Extensive experiments on four datasets demonstrate that our OODML surpasses the state-of-the-art by 3.3% and 6.9% on the CAMELYON16 and TCGA Lung datasets. Kui Jiang, Hongxun Yao, Yi Xiao 0003, Zhongyuan Wang 0001 |
AAAI | 2 |
| 2025 | CLAIM: Mitigating Multilingual Object Hallucination in Large Vision-Language Models with Cross-Lingual Attention InterventionabstractLarge Vision-Language Models (LVLMs) have demonstrated impressive multimodal abilities but remain prone to multilingual object hallucination, with a higher likelihood of generating responses inconsistent with the visual input when utilizing queries in non-English languages compared to English. Most existing approaches to address these rely on pretraining or fine-tuning, which are resource-intensive. In this paper, inspired by observing the disparities in cross-modal attention patterns across languages, we propose Cross-Lingual Attention Intervention for Mitigating multilingual object hallucination (CLAIM) in LVLMs, a novel near training-free method by aligning attention patterns. CLAIM first identifies language-specific cross-modal attention heads, then estimates language shift vectors from English to the target language, and finally intervenes in the attention outputs during inference to facilitate cross-lingual visual perception capability alignment. Extensive experiments demonstrate that CLAIM achieves an average improvement of 13.56% (up to 30% in Spanish) on the POPE and 21.75% on the hallucination subsets of the MME benchmark across various languages. Further analysis reveals that multilingual attention divergence is most prominent in intermediate layers, highlighting their critical role in multilingual scenarios. Zekai Ye, Libo Qin 0001, Yichong Huang, Baohang Li, Kui Jiang, Yang Xiang 0003, Zhirui Zhang, Yunfei Lu, Duyu Tang, Dandan Tu, Bing Qin 0001 |
ACL (1) | 7 |
| 2025 | DashGaussian: Optimizing 3D Gaussian Splatting in 200 Secondsabstract3D Gaussian Splatting (3DGS) renders pixels by rasterizing Gaussian primitives, where the rendering resolution and the primitive number, concluded as the optimization complexity, dominate the time cost in primitive optimization. In this paper, we propose DashGaussian, a scheduling scheme over the optimization complexity of 3DGS that strips redundant complexity to accelerate 3DGS optimization. Specifically, we formulate 3DGS optimization as progressively fitting 3DGS to higher levels of frequency components in the training views, and propose a dynamic rendering resolution scheme that largely reduces the optimization complexity based on this formulation. Besides, we argue that a specific rendering resolution should cooperate with a proper primitive number for a better balance between computing redundancy and fitting quality, where we schedule the growth of the primitives to synchronize with the rendering resolution. Extensive experiments show that our method accelerates the optimization of various 3DGS backbones by 45.7% on average while preserving the rendering quality. Project page is available at dashgaussian.github.io. Youyu Chen, Junjun Jiang, Kui Jiang, Xianming Liu 0005, Yinyu Nie |
CVPR | 3 |
| 2025 | Fast and Accurate Gigapixel Pathological Image Classification with Hierarchical Distillation Multi-Instance LearningabstractAlthough multi-instance learning (MIL) has succeeded in pathological image classification, it faces the challenge of high inference costs due to processing numerous patches from gigapixel whole slide images (WSIs). To address this, we propose HDMIL, a hierarchical distillation multi-instance learning framework that achieves fast and accurate classification by eliminating irrelevant patches. HD-MIL consists of two key components: the dynamic multi-instance network (DMIN) and the lightweight instance pre-screening network (LIPN). DMIN operates on high-resolution WSIs, while LIPN operates on the corresponding low-resolution counterparts. During training, DMIN are trained for WSI classification while generating attention-score-based masks that indicate irrelevant patches. These masks then guide the training of LIPN to predict the relevance of each low-resolution patch. During testing, LIPN first determines the useful regions within low-resolution WSIs, which indirectly enables us to eliminate irrelevant regions in high-resolution WSIs, thereby reducing inference time without causing performance degradation. In addition, we further design the first Chebyshev-polynomials-based Kolmogorov-Arnold classifier in computational pathology, which enhances the performance of HDMIL through learnable activation layers. Extensive experiments on three public datasets demonstrate that HDMIL outperforms previous state-of-the-art methods, e.g., achieving improvements of 3.13% in AUC while reducing inference time by 28.6% on the Camelyon16 dataset. The project is available at https://github.com/JiuyangDong/HDMIL. Jiuyang Dong, Junjun Jiang, Kui Jiang, Jiahan Li, Yongbing Zhang 0002 |
CVPR | 3 |
| 2025 | COB-GS: Clear Object Boundaries in 3DGS Segmentation Based on Boundary-Adaptive Gaussian SplittingabstractAccurate object segmentation is crucial for high-quality scene understanding in the 3D vision domain. However, 3D segmentation based on 3D Gaussian Splatting (3DGS) struggles with accurately delineating object boundaries, as Gaussian primitives often span across object edges due to their inherent volume and the lack of semantic guidance during training. In order to tackle these challenges, we introduce Clear Object Boundaries for 3DGS Segmentation (COB-GS), which aims to improve segmentation accuracy by clearly delineating blurry boundaries of interwoven Gaussian primitives within the scene. Unlike existing approaches that remove ambiguous Gaussians and sacrifice visual quality, COB-GS, as a 3DGS refinement method, jointly optimizes semantic and visual information, allowing the two different levels to cooperate with each other effectively. Specifically, for the semantic guidance, we introduce a boundary-adaptive Gaussian splitting technique that leverages semantic gradient statistics to identify and split ambiguous Gaussians, aligning them closely with object boundaries. For the visual optimization, we rectify the degraded suboptimal texture of the 3DGS scene, particularly along the refined boundary structures. Experimental results show that COB-GS substantially improves segmentation accuracy and robustness against inaccurate masks from pre-trained model, yielding clear boundaries while preserving high visual quality. Code is available at https://github.com/ZestfulJX/COB-GS. Junjun Jiang, Youyu Chen, Kui Jiang, Xianming Liu 0005 |
CVPR | 4 |
| 2025 | M3amba: Memory Mamba is All You Need for Whole Slide Image ClassificationabstractMulti-instance learning (MIL) has demonstrated impressive performance in whole slide image (WSI) analysis. However, existing approaches struggle with undesirable results and unbearable computational overhead due to the quadratic complexity of Transformers. Recently, Mamba has offered a feasible solution for modeling long-range dependencies with linear complexity. However, vanilla Mamba inherently suffers from contextual forgetting issues, making it ill-suited for capturing global dependencies across instances in large-scale WSIs. To address this, we propose a memory-driven Mamba network, dubbed M3amba, to fully explore the global latent relations among instances. Specifically, M3amba retains and iteratively updates historical information with a dynamic memory bank (DMB), thus overcoming the catastrophic forgetting defects of Mamba for long-term context representation. For better feature representation, M3amba involves an intra-group bidirectional Mamba (BiMamba) block to refine local interactions within groups. Meanwhile, we additionally perform cross-attention fusion to incorporate relevant historical information across groups, facilitating richer inter-group connections. The joint learning of inter- and intra-group representations with memory merits enables M3amba with a more powerful capability for achieving accurate and comprehensive WSI representation. Extensive experiments on four datasets demonstrate that M3amba outperforms the state-of-the-art by 6.2% and 7.0% in accuracy on the TCGA BRCA and TCGA Lung datasets while maintaining low computational costs. Kui Jiang, Yi Xiao 0003, Sicheng Zhao, Hongxun Yao |
CVPR | 2 |
| 2025 | Balancing Task-Invariant Interaction and Task-Specific Adaptation for Unified Image FusionabstractUnified image fusion aims to integrate complementary information from multi-source images, enhancing image quality through a unified framework applicable to diverse fusion tasks. While treating all fusion tasks as a unified problem facilitates task-invariant knowledge sharing, it often overlooks task-specific characteristics, thereby limiting the overall performance. Existing general image fusion methods incorporate explicit task identification to enable adaptation to different fusion tasks. However, this dependence during inference restricts the model's generalization to unseen fusion tasks. To address these issues, we propose a novel unified image fusion framework named "TITA", which dynamically balances both Task-invariant Interaction and Task-specific Adaptation. For task-invariant interaction, we introduce the Interaction-enhanced Pixel Attention (IPA) module to enhance pixel-wise interactions for better multi-source complementary information extraction. For task-specific adaptation, the Operation-based Adaptive Fusion (OAF) module dynamically adjusts operation weights based on task properties. Additionally, we incorporate the Fast Adaptive Multitask Optimization (FAMO) strategy to mitigate the impact of gradient conflicts across tasks during joint training. Extensive experiments demonstrate that TITA not only achieves competitive performance compared to specialized methods across three image fusion scenarios but also exhibits strong generalization to unseen fusion tasks. The source codes are released at https://github.com/huxingyuabc/TITA. Junjun Jiang, Chenyang Wang 0002, Kui Jiang, Xianming Liu 0005, Jiayi Ma 0001 |
ICCV | 4 |
| 2025 | GMMamba: Group Masking Mamba for Whole Slide Image Classification
Hongxun Yao, Kui Jiang, Yi Xiao 0003, Sicheng Zhao |
ICCV | 3 |
| 2025 | Always Clear Depth: Robust Monocular Depth Estimation Under Adverse WeatherabstractMonocular depth estimation is critical for applications such as autonomous driving and scene reconstruction. While existing methods perform well under normal scenarios, their performance declines in adverse weather, due to challenging domain shifts and difficulties in extracting scene information. To address this issue, we present a robust monocular depth estimation method called ACDepth from the perspective of high-quality training data generation and domain adaptation. Specifically, we introduce a one-step diffusion model for generating samples that simulate adverse weather conditions, constructing a multi-tuple degradation dataset during training. To ensure the quality of the generated degradation samples, we employ LoRA adapters to fine-turn the generation weights of diffusion model. Additionally, we integrate circular consistency loss and adversarial training to guarantee the fidelity and naturalness of the scene contents. Furthermore, we elaborate on a multi-granularity knowledge distillation strategy (MKD) that encourages the student network to absorb knowledge from both the teacher model and pretrained Depth Anything V2. This strategy guides the student model in learning degradation-agnostic scene information from various degradation inputs. In particular, we introduce an ordinal guidance distillation mechanism (OGD) that encourages the network to focus on uncertain regions through differential ranking, leading to a more precise depth estimation. Experimental results demonstrate that our ACDepth surpasses md4all-DD by 2.50% for night scene and 2.61% for rainy scene on the nuScenes dataset in terms of the absRel metric. Kui Jiang, Zhaocheng Yu, Junjun Jiang, Jingchun Zhou |
IJCAI | 1 |
| 2025 | FUSE: Label-Free Image-Event Joint Monocular Depth Estimation via Frequency-Decoupled Alignment and Degradation-Robust FusionabstractImage-event joint depth estimation methods leverage complementary modalities for robust perception, yet face challenges in generalizability stemming from two factors: 1) limited annotated image-event-depth datasets causing insufficient cross-modal supervision, and 2) inherent frequency mismatches between static images and dynamic event streams with distinct spatiotemporal patterns, leading to ineffective feature fusion. To address this dual challenge, we propose Frequency-decoupled Unified Self-supervised Encoder (FUSE) with two synergistic components: The Parameter-efficient Self-supervised Transfer (PST) leverages image foundation models for cross-modal knowledge transfer, effectively mitigating data scarcity by enabling joint encoding without depth ground truth. Complementing this, the Frequency-Decoupled Fusion module (FreDFuse) resolves modality-specific frequency mismatches by decoupling features into high- and low-frequency bands and then performing a guided cross-attention fusion, where the modality dominant in each band steers the integration. This combined approach enables FUSE to construct a universal image-event encoder that only requires lightweight decoder adaptation for target datasets. Extensive experiments demonstrate state-of-the-art performance with 14% and 24.9% improvements in Abs.Rel on MVSEC and DENSE datasets. The framework exhibits remarkable robustness and generalization in challenging scenarios, including extreme lighting and motion blur, significantly advancing its real-world deployment capabilities. The source code for our method is publicly available at: https://github.com/sunpihai-up/FUSE. Pihai Sun, Junjun Jiang, Yuanqi Yao, Youyu Chen, Wenbo Zhao 0004, Kui Jiang, Xianming Liu 0005 |
IROS | 6 |
| 2025 | Positive Style Accumulation: A Style Screening and Continuous Utilization Framework for Federated DG-ReIDabstractThe Federated Domain Generalization for Person re-identification (FedDG-ReID) aims to learn a global server model that can be effectively generalized to source and target domains through distributed source domain data. Existing methods mainly improve the diversity of samples through style transformation, which to some extent enhances the generalization performance of the model. However, we discover that not all styles contribute to the generalization performance. Therefore, we define styles that are beneficial/harmful to the model's generalization performance as positive/negative styles. Based on this, new issues arise: How to effectively screen and continuously utilize the positive styles. To solve these problems, we propose a Style Screening and Continuous Utilization (SSCU) framework. Firstly, we design a Generalization Gain-guided Dynamic Style Memory (GGDSM) for each client model to screen and accumulate generated positive styles. Specifically, the memory maintains a prototype initialized from raw data for each category, then screens positive styles that enhance the global model during training, and updates these positive styles into the memory using a momentum-based approach. Meanwhile, we propose a style memory recognition loss to fully leverage the positive styles memorized by GGDSM. Furthermore, we propose a Collaborative Style Training (CST) strategy to make full use of positive styles. Unlike traditional learning strategies, our approach leverages both newly generated styles and the accumulated positive styles stored in memory to train client models on two distinct branches. This training strategy is designed to effectively promote the rapid acquisition of new styles by the client models, ensuring that they can quickly adapt to and integrate novel stylistic variations. Simultaneously, this strategy guarantees the continuous and thorough utilization of positive styles, which is highly beneficial for the model's generalization performance. Extensive experimental results demonstrate that our method outperforms existing methods in both the source domain and the target domain. Xin Xu 0007, Chaoyue Ren, Wei Liu 0183, Wenke Huang 0003, Bin Yang 0026, Zhixi Yu, Kui Jiang |
ACM Multimedia | 7 |
| 2025 | Reframing Gaussian Splatting Densification with Complexity-Density Consistency of PrimitivesabstractThe essence of 3D Gaussian Splatting (3DGS) training is to smartly allocate Gaussian primitives, expressing complex regions with more primitives and vice versa.
Prior researches typically mark out under-reconstructed regions in a rendering-loss-driven manner.
However, such a loss-driven strategy is often dominated by low-frequency regions, which leads to insufficient modeling of high-frequency details in texture-rich regions. As a result, it yields a suboptimal spatial allocation of Gaussian primitives.
This inspires us to excavate the loss-agnostic visual prior in training views to identify complex regions that need more primitives to model.
Based on this insight, we propose Complexity-Density Consistent Gaussian Splatting (CDC-GS), which allocates primitives based on the consistency between visual complexity of training views and the density of primitives.
Specifically, primitives involved in rendering high visual complexity areas are categorized as modeling high complexity regions, where we leverage the high frequency wavelet components of training views to measure the visual complexity.
And the density of a primitive is computed with the inverse of geometric mean of its distance to the neighboring primitives.
Guided by the positive correlation between primitive complexity and density, we determine primitives to be densified as well as pruned.
Extensive experiments demonstrate that our CDC-GS surpasses the baseline methods in rendering quality by a large margin using the same amount of Gaussians.
And we provide insightful analysis to reveal that our method serves perpendicularly to rendering loss in guiding Gaussian primitive allocation. Zhemeng Dong, Junjun Jiang, Youyu Chen, Kui Jiang, Xianming Liu 0005 |
NeurIPS | 5 |
| 2025 | Spiking Meets Attention: Efficient Remote Sensing Image Super-Resolution with Attention Spiking Neural NetworksabstractSpiking neural networks (SNNs) are emerging as a promising alternative to traditional artificial neural networks (ANNs), offering biological plausibility and energy efficiency. Despite these merits, SNNs are frequently hampered by limited capacity and insufficient representation power, yet remain underexplored in remote sensing image (RSI) super-resolution (SR) tasks. In this paper, we first observe that spiking signals exhibit drastic intensity variations across diverse textures, highlighting an active learning state of the neurons. This observation motivates us to apply SNNs for efficient SR of RSIs. Inspired by the success of attention mechanisms in representing salient information, we devise the spiking attention block (SAB), a concise yet effective component that optimizes membrane potentials through inferred attention weights, which, in turn, regulates spiking activity for superior feature representation. Our key contributions include: 1) we bridge the independent modulation between temporal and channel dimensions, facilitating joint feature correlation learning, and 2) we access the global self-similar patterns in large-scale remote sensing imagery to infer spatial attention weights, incorporating effective priors for realistic and faithful reconstruction. Building upon SAB, we proposed SpikeSR, which achieves state-of-the-art performance across various remote sensing benchmarks such as AID, DOTA, and DIOR, while maintaining high computational efficiency. Code of SpikeSR will be available at https://github.com/XY-boy/SpikeSR. Yi Xiao 0003, Qiangqiang Yuan, Kui Jiang, Wenke Huang 0003, Qiang Zhang 0011, Chia-Wen Lin, Liangpei Zhang 0001 |
NeurIPS | 3 |
| 2025 | GDPS: A general distillation architecture for end-to-end person search
Shichang Fu, Tao Lu 0001, Jiaming Wang 0001, Jiayi Cai, Kui Jiang |
J. Vis. Commun. Image Represent. | 6 |
| 2025 | Collaborative optimization for whole slide image classification
Hongxun Yao, Kui Jiang, Yi Xiao 0003 |
Knowl. Based Syst. | 3 |
| 2025 | A Survey on All-in-One Image Restoration: Taxonomy, Evaluation and Future TrendsabstractImage restoration (IR) seeks to recover high-quality images from degraded observations caused by a wide range of factors, including noise, blur, compression, and adverse weather. While traditional IR methods have made notable progress by targeting individual degradation types, their specialization often comes at the cost of generalization, leaving them ill-equipped to handle the multifaceted distortions encountered in real-world applications. In response to this challenge, the all-in-one image restoration (AiOIR) paradigm has recently emerged, offering a unified framework that adeptly addresses multiple degradation types. These innovative models enhance the convenience and versatility by adaptively learning degradation-specific features while simultaneously leveraging shared knowledge across diverse corruptions. In this survey, we provide the first in-depth and systematic overview of AiOIR, delivering a structured taxonomy that categorizes existing methods by architectural designs, learning paradigms, and their core innovations. We systematically categorize current approaches and assess the challenges these models encounter, outlining research directions to propel this rapidly evolving field. To facilitate the evaluation of existing methods, we also consolidate widely-used datasets, evaluation protocols, and implementation practices, and compare and summarize the most advanced open-source models. As the first comprehensive review dedicated to AiOIR, this paper aims to map the conceptual landscape, synthesize prevailing techniques, and ignite further exploration toward more intelligent, unified, and adaptable visual restoration systems. Junjun Jiang, Zengyuan Zuo, Gang Wu 0010, Kui Jiang, Xianming Liu 0005 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2025 | Dynamic and static mutual fitting for action recognition
Wenxuan Liu 0008, Xuemei Jia, Xian Zhong, Kui Jiang, Xiaohan Yu 0001, Mang Ye |
Pattern Recognit. | 4 |
| 2025 | GraphMamba: Whole slide image classification meets graph-driven selective state space model
Hongxun Yao, Sicheng Zhao, Kui Jiang, Yi Xiao 0003 |
Pattern Recognit. | 4 |
| 2025 | Multiple Pedestrian Tracking Under Occlusion: A Survey and OutlookabstractAs an intermediate task in computer vision, multiple pedestrian tracking (MPT) aiming at tracking the pedestrians from a given video, has attracted attention due to its potential academic and commercial value. However, pedestrians commonly suffer from occlusion due to diverse and complex scenarios, which increases the challenge of this task. This survey provides comprehensive review in terms of occlusion scenarios encountered during MPT, and investigates the model robustness of the existing methods in this scenarios. Firstly, this survey introduces the various and states of occlusion. Secondly, the related occlusion datasets are introduced. Subsequently, we categorize existing occlusion handling methods according to the tracking process and detail their pros and cons. In addition, occlusion handling precision (OHP) metric is proposed to evaluate the ability of a tracker in handling occlusion in this survey. Moreover, comprehensive analyzes and discussions in several public datasets are provided to verify the effectiveness of these methods. Finally, the existing issues and future directions for occlusion handling methods are discussed. In doing so, this work serves as a foundation for future research by providing researchers with information about the occlusion handling method of MPT. Guoheng Wei, Mang Ye, Kui Jiang, Chao Liang 0001, Mithun Mukherjee 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2025 | For Overall Nighttime Visibility: Integrate Irregular Glow Removal With Glow-Aware EnhancementabstractCurrent low-light image enhancement (LLIE) techniques truly enhance luminance but have limited exploration on another harmful factor of nighttime visibility, the glow effects with multiple shapes in the real world. The presence of glow is inevitable due to widespread artificial light sources, and direct enhancement can cause further glow diffusion. In the pursuit of Overall Nighttime Visibility Enhancement (ONVE), we propose a physical model guided framework ONVE to derive a Nighttime Imaging Model with Near-Field Light Sources (NIM-NLS), whose APSF prior generator is validated efficiently in six categories of glow shapes. Guided by this physical-world model as domain knowledge, we subsequently develop an extensible Light-aware Blind Deconvolution Network (LBDN) to face the blind decomposition challenge on direct transmission map D and light source map G based on APSF. Then, an innovative Glow-guided Retinex-based progressive Enhancement module (GRE) is introduced as a further optimization on reflection R from D to harmonize the conflict of glow removal and brightness boost. Notably, ONVE is an unsupervised framework based on a zero-shot learning strategy and uses physical domain knowledge to form the overall pipeline and network. Empirical evaluations on multiple datasets validate the remarkable efficacy of the proposed ONVE in improving nighttime visibility and performance of high-level vision tasks. Wanyu Wu, Wei Wang 0170, Zheng Wang 0007, Kui Jiang, Zhengguo Li |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2025 | Rep-Mamba: Re-Parameterization in Vision Mamba for Lightweight Remote Sensing Image Super-ResolutionabstractThe selective space model (Mamba) has recently demonstrated great potential in remote sensing image super-resolution (RSISR) tasks due to its capability for long-range dependency modeling with linear computational complexity. Despite these merits, existing Mamba architectures face two critical challenges in large-scale remote sensing scenarios: 1) neglecting the local semantic integrity due to the unfolding 1-D sequential representations and 2) facing the dilemma between effectiveness and efficiency. To address these issues, we propose Rep-Mamba, a lightweight progressive multiscale feature fusion architecture based on the state-space model (SSM) for RSISR. Specifically, we innovatively design a cross-scale state propagation (CSSP) mechanism and construct a lightweight progressive fusion module (LPFM) to dynamically capture hierarchical spatial dependencies in remote sensing scenes while maintaining high computational efficiency. Moreover, to achieve synergistic optimization between local semantic structure preservation and global context modeling, we introduce differentiable re-parameterization convolution (RepConv), which significantly enhances reconstruction accuracy and visual quality without compromising computational efficiency. Extensive experiments across multiple benchmarks demonstrate that Rep-Mamba achieves a superior tradeoff between accuracy and complexity, highlighting its effectiveness and scalability. The code is available athttps://github.com/meigeni0929/Rep-Mambahttps://github.com/meigeni0929/Rep-Mamba Kui Jiang, Mengru Yang, Yi Xiao 0003, Guangcheng Wang, Junjun Jiang |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2025 | Learning Dynamic Prompts for All-in-One Image RestorationabstractAll-in-one image restoration, which seeks to handle multiple types of degradation within a unified model, has become a prominent research topic in computer vision. While existing deep learning models have achieved remarkable success in specific restoration tasks, extending these models to heterogenous degradations presents significant challenges. Current all-in-one methods predominantly concentrate on extracting degradation priors, often employing learned and fixed task prompts to guide the restoration process. However, these static prompts are inclined to generate an average distribution characteristics of degradations, unable to accurately depict the unique attribute of the given input, consequently providing suboptimal restoration results. To tackle these challenges, we propose a novel dynamic prompt approach called Degradation Prototype Assignment and Prompt Distribution Learning (DPPD). Our approach decouples the degradation prior extraction into two novel components: Degradation Prototype Assignment (DPA) and Prompt Distribution Learning (PDL). DPA anchors the degradation representations to predefined prototypes, providing discriminative and scalable representations. In addition, PDL models prompts as distributions rather than fixed parameters, facilitating dynamic and adaptive prompt sampling. Extensive experiments demonstrate that our DPPD framework can achieve significant performance improvement on different image restoration tasks. Codes are available at our project page https://github.com/Aitical/DPPD. Gang Wu 0010, Junjun Jiang, Kui Jiang, Xianming Liu 0005, Liqiang Nie |
IEEE Trans. Image Process. | 3 |
| 2025 | Multi-Axis Feature Diversity Enhancement for Remote Sensing Video Super-ResolutionabstractHow to aggregate spatial-temporal information plays an essential role in video super-resolution (VSR) tasks. Despite the remarkable success, existing methods adopt static convolution to encode spatial-temporal information, which lacks flexibility in aggregating information in large-scale remote sensing scenes, as they often contain heterogeneous features (e.g., diverse textures). In this paper, we propose a spatial feature diversity enhancement module (SDE) and channel diversity enhancement module (CDE), which explore the diverse representation of different local patterns while aggregating the global response with compactly channel-wise embedding representation. Specifically, SDE introduces multiple learnable filters to extract representative spatial variants and encodes them to generate a dynamic kernel for enriched spatial representation. To explore the diversity in the channel dimension, CDE exploits the discrete cosine transform to transform the feature into the frequency domain. This enriches the channel representation while mitigating massive frequency loss caused by pooling operation. Based on SDE and CDE, we further devise a multi-axis feature diversity enhancement (MADE) module to harmonize the spatial, channel, and pixel-wise features for diverse feature fusion. These elaborate strategies form a novel network for satellite VSR, termed MADNet, which achieves favorable performance against state-of-the-art method BasicVSR++ in terms of average PSNR by 0.14 dB on various video satellites, including JiLin-1, Carbonite-2, SkySat-1, and UrtheCast. Code will be available at https://github.com/XY-boy/MADNet. Yi Xiao 0003, Qiangqiang Yuan, Kui Jiang, Yuzeng Chen, Shiqi Wang 0001, Chia-Wen Lin |
IEEE Trans. Image Process. | 3 |
| 2025 | Disentangled Pseudo-Bag Augmentation for Whole Slide Image Multiple Instance LearningabstractAs the predominant approach for pathological whole slide image (WSI) classification, multiple instance learning (MIL) methods struggle with limited labeled WSIs. Although MIL has achieved notable progress with pseudo-bag-oriented augmentation methods, their effectiveness is often constrained by noisy pseudo-labels and low-quality pseudo-bags. To overcome these problems, we revisit the use of pseudo-bags for WSI data augmentation and propose a new pseudo-bag generation paradigm, dubbed DPBAug. Its distinctive features can be summarized as: i) We develop an intra-slide pseudo-bag generation module, which separates the heterogeneous instances within each slide through phenotype partitioning. Moreover, to ensure accurate label inheritance when generating pseudo-bags, we propose an instance sampling algorithm with replacement. ii) An inter-slide pseudo-bag fusion module is designed to integrate heterogeneous information across multiple WSIs, producing high-quality training samples that better leverage the potential of neural networks. iii) A pseudo-bag memory update module prioritizes valuable synthetic pseudo-bags. This further enhances the network's classification performance. Extensive experiments demonstrate that DPBAug surpasses existing augmentation methods, enhancing the classification performance and reliability of multiple MIL baselines across various public datasets. DPBAug also improves the generalization and data efficiency of existing MIL methods, facilitating their adoption in clinical practice and rare cancer research The project is available at: https://github.com/JiuyangDong/DPBAug. Jiuyang Dong, Junjun Jiang, Kui Jiang, Jiahan Li, Linghan Cai, Yongbing Zhang 0002 |
IEEE Trans. Medical Imaging | 3 |
| 2025 | Does Adding a Modality Really Make Positive Impacts in Incomplete Multi-Modal Brain Tumor Segmentation?abstractPrevious incomplete multi-modal brain tumor segmentation technologies, while effective in integrating diverse modalities, commonly deliver under-expected performance gains. The reason lies in that the new modality may cause confused predictions due to uncertain and inconsistent patterns and quality in some positions, where the direct fusion consequently raises the negative gain for the final decision. In this paper, considering the potentially negative impacts within a modality, we propose multi-modal Positive-Negative impact region Double Calibration pipeline, called PNDC, to mitigate misinformation transfer of modality fusion. Concretely, PNDC involves two elaborate pipelines, Reverse Audit and Forward Checksum. The former is to identify negative regions impacts of each modality. The latter calibrates whether the fusion prediction is reliable in these regions by integrating the positive impacts regions of each modality. Finally, the negative impacts region from each modality and miss-match reliable fusion predictions are utilized to enhance the learning of individual modalities and fusion process. It is noted that PNDC adopts the standard training strategy without specific architectural choices and does not introduce any learning parameters, and thus can be easily plugged into existing network training for incomplete multi-modal brain tumor segmentation. Extensive experiments confirm that our PNDC greatly alleviates the performance degradation of current state-of-the-art incomplete medical multi-modal methods, arising from overlooking the positive/negative impacts regions of the modality. The code is released at PNDC. Yansheng Qiu, Kui Jiang, Hongdou Yao, Zheng Wang 0007, Shin'ichi Satoh 0001 |
IEEE Trans. Medical Imaging | 2 |
| 2025 | DALFace: Dynamic Association Learning for Face RecognitionabstractFace recognition owes its success to the availability of large-scale training data. Recent adaptive margin-based loss functions pay more attention to hard (misclassified) samples, resulting in more discriminative face embeddings. However, large-scale datasets inevitably include open-set noise samples, which are usually mistaken for hard samples by mining-based methods and thus mislead the training of the model. In this work, we redefine hard samples and further design a dynamic association learning strategy for mining hard samples while ignoring noise. We argue that the difficulty of recognizing a sample depends on both identity-related and objective factors. On one hand, intrinsic attributes such as facial structure and face shape inherently influence the ease of identity recognition. On the other hand, external factors, including pose, occlusion, and resolution, directly affect the recognizability of a sample. Particularly in the case of noise samples, although they pose challenges for the deep network similar to hard samples, should not be regarded as hard samples. To this end, we propose an associated prototype learning method to achieve an approximation of face identity difficulty by exploring the fitting trends of identity prototype. Furthermore, we design a dynamic sample learning method to distinguish noise samples from hard samples by observing the distance fluctuation from the class center during sample learning. All observations are integrated into the loss function through adaptive margins and sample weights. Extensive experiments and visualizations on several datasets demonstrate that our method significantly outperforms state-of-the-art counterparts. Baojin Huang, Guangcheng Wang, Kui Jiang, Zhongyuan Wang 0001 |
IEEE Trans. Multim. | 3 |
| 2025 | Frequency-Assisted Mamba for Remote Sensing Image Super-ResolutionabstractRecent progress in remote sensing image (RSI) super-resolution (SR) has exhibited remarkable performance using deep neural networks, e.g., Convolutional Neural Networks and Transformers. However, existing SR methods often suffer from either a limited receptive field or quadratic computational overhead, resulting in sub-optimal global representation and unacceptable computational costs in large-scale RSI. To alleviate these issues, we develop the first attempt to integrate the Vision State Space Model (Mamba) for RSI-SR, which specializes in processing large-scale RSI by capturing long-range dependency with linear complexity. To achieve better SR reconstruction, building upon Mamba, we devise a Frequency-assisted Mamba framework, dubbed FMSR, to explore the spatial and frequent correlations. In particular, our FMSR features a multi-level fusion architecture equipped with the Frequency Selection Module (FSM), Vision State Space Module (VSSM), and Hybrid Gate Module (HGM) to grasp their merits for effective spatial-frequency fusion. Considering that global and local dependencies are complementary and both beneficial for SR, we further recalibrate these multi-level features for accurate feature fusion via learnable scaling adaptors. Extensive experiments on AID, DOTA, and DIOR benchmarks demonstrate that our FMSR outperforms state-of-the-art Transformer-based methods HAT-L in terms of PSNR by 0.11 dB on average, while consuming only 28.05% and 19.08% of its memory consumption and complexity, respectively. Yi Xiao 0003, Qiangqiang Yuan, Kui Jiang, Yuzeng Chen, Qiang Zhang 0011, Chia-Wen Lin |
IEEE Trans. Multim. | 3 |
| 2025 | DAWN+: Wavelet-Based Image Deraining Meets Direction-Aware Attention and Mutual RepresentationabstractThe single-image deraining aims to restore clean scenes from rainy inputs by eliminating precipitation artifacts. Current methods often neglect the directional nature of rain streaks-a critical oversight that causes heterogeneous degradation, particularly in texture regions aligned with rain orientations. To address this issue and advance image deraining, we propose a novel direction-aware attention wavelet network (DAWN) for rain streaks removal. DAWN has several key distinctions and innovative features compared with existing wavelet transform-based methods: 1) introducing vector decomposition to parameterize rain distribution through vertical (V) and horizontal (H) component decomposition, enabling explicit direction-aware representation; 2) devising a novel direction-aware attention module (DAM) to learn projection/transformation parameters via coordinate attention mechanisms for precise rain removal and texture preservation; and 3) exploring practical composite constraints to jointly optimize structural coherence, detail fidelity, and chrominance accuracy. Building upon the conference version (DAWN), we devise DAWN+ with enhanced capabilities: 1) decoupling diagonal coefficient learning to eliminate frequency aliasing by characterizing diagonal components with dedicated projection parameters; 2) dividing vector decomposition and parameter fitting into multiple stages to reduce error accumulation; and 3) applying cross-frequency mutual representation to boost training and performance. Experiments across six tasks (deraining, raindrop/rainhaze removal, dehazing, and low-light/underwater enhancement) demonstrate the portability and reusability of these strategies. Meanwhile, DAWN+ delivers significant performance gains over DAWN, achieving an average peak signal to noise ratio (PSNR) increase of 1.17 dB with an acceptable complexity increase. Meanwhile, DAWN+ achieves the competitive performance to the state-of-the-art DRSformer (gaining 0.15 dB in PSNR) while saving 94.4% and 95% model parameters and inference time, respectively. Kui Jiang, Junjun Jiang, Zheng Wang 0007, Zihan Geng, Xianming Liu 0005 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2025 | Fragrant: frequency-auxiliary guided relational attention network for low-light action recognition
Wenxuan Liu 0008, Xuemei Jia, Yihao Ju, Yakun Ju, Kui Jiang, Shifeng Wu, Luo Zhong, Xian Zhong |
Vis. Comput. | 5 |
| 2025 | Context-aware target texture perturbation attack for concealed object detection
Kui Jiang, Nan Mu |
Vis. Comput. | 4 |
| 2024 | Learning from History: Task-agnostic Model Contrastive Learning for Image RestorationabstractContrastive learning has emerged as a prevailing paradigm for high-level vision tasks, which, by introducing properly negative samples, has also been exploited for low-level vision tasks to achieve a compact optimization space to account for their ill-posed nature. However, existing methods rely on manually predefined and task-oriented negatives, which often exhibit pronounced task-specific biases. To address this challenge, our paper introduces an innovative method termed 'learning from history', which dynamically generates negative samples from the target model itself. Our approach, named Model Contrastive Learning for Image Restoration (MCLIR), rejuvenates latency models as negative models, making it compatible with diverse image restoration tasks. We propose the Self-Prior guided Negative loss (SPN) to enable it. This approach significantly enhances existing models when retrained with the proposed model contrastive paradigm. The results show significant improvements in image restoration across various tasks and architectures. For example, models retrained with SPN outperform the original FFANet and DehazeFormer by 3.41 and 0.57 dB on the RESIDE indoor dataset for image dehazing. Similarly, they achieve notable improvements of 0.47 dB on SPA-Data over IDT for image deraining and 0.12 dB on Manga109 for a 4x scale super-resolution over lightweight SwinIR, respectively. Code and retrained models are available at https://github.com/Aitical/MCLIR. Gang Wu 0010, Junjun Jiang, Kui Jiang, Xianming Liu 0005 |
AAAI | 3 |
| 2024 | FMRNet: Image Deraining via Frequency Mutual RevisionabstractThe wavelet transform has emerged as a powerful tool in deciphering structural information within images. And now, the latest research suggests that combining the prowess of wavelet transform with neural networks can lead to unparalleled image deraining results. By harnessing the strengths of both the spatial domain and frequency space, this innovative approach is poised to revolutionize the field of image processing. The fascinating challenge of developing a comprehensive framework that takes into account the intrinsic frequency property and the correlation between rain residue and background is yet to be fully explored. In this work, we propose to investigate the potential relationships among rain-free and residue components at the frequency domain, forming a frequency mutual revision network (FMRNet) for image deraining. Specifically, we explore the mutual representation of rain residue and background components at frequency domain, so as to better separate the rain layer from clean background while preserving structural textures of the degraded images. Meanwhile, the rain distribution prediction from the low-frequency coefficient, which can be seen as the degradation prior is used to refine the separation of rain residue and background components. Inversely, the updated rain residue is used to benefit the low-frequency rain distribution prediction, forming the multi-layer mutual learning. Extensive experiments demonstrate that our proposed FMRNet delivers significant performance gains for seven datasets on image deraining task, surpassing the state-of-the-art method ELFormer by 1.14 dB in PSNR on the Rain100L dataset, while with similar computation cost. Code and retrained models are available at https://github.com/kuijiang94/FMRNet. Kui Jiang, Junjun Jiang, Xianming Liu 0005, Xin Xu 0007, Xianzheng Ma |
AAAI | 1 |
| 2024 | Low-Light Face Super-resolution via Illumination, Structure, and Texture Associated RepresentationabstractHuman face captured at night or in dimly lit environments has become a common practice, accompanied by complex low-light and low-resolution degradations. However, the existing face super-resolution (FSR) technologies and derived cascaded schemes are inadequate to recover credible textures. In this paper, we propose a novel approach that decomposes the restoration task into face structural fidelity maintaining and texture consistency learning. The former aims to enhance the quality of face images while improving the structural fidelity, while the latter focuses on eliminating perturbations and artifacts caused by low-light degradation and reconstruction. Based on this, we develop a novel low-light low-resolution face super-resolution framework. Our method consists of two steps: an illumination correction face super-resolution network (IC-FSRNet) for lighting the face and recovering the structural information, and a detail enhancement model (DENet) for improving facial details, thus making them more visually appealing and easier to analyze. As the relighted regions could provide complementary information to boost face super-resolution and vice versa, we introduce the mutual learning to harness the informative components from relighted regions and reconstruction, and achieve the iterative refinement. In addition, DENet equipped with diffusion probabilistic model is built to further improve face image visual quality. Experiments demonstrate that the proposed joint optimization framework achieves significant improvements in reconstruction quality and perceptual quality over existing two-stage sequential solutions. Code is available at https://github.com/wcy-cs/IC-FSRDENet. Chenyang Wang 0002, Junjun Jiang, Kui Jiang, Xianming Liu 0005 |
AAAI | 3 |
| 2024 | IQ-VFI: Implicit Quadratic Motion Estimation for Video Frame InterpolationabstractAdvanced video frame interpolation (VFI) algorithms approximate intermediate motions between two input frames to synthesize intermediate frame. However, they struggle to handle complex scenarios with curvilinear motions since they overlook the latent acceleration information between the input frames. Moreover, the supervision of predicted motions is tricky because ground-truth motions are not available. To this end, we propose a novel frame-work for implicit quadratic video frame interpolation (IQ-VFI), which explores latent acceleration information and accurate intermediate motions via knowledge distillation. Specifically, the proposed IQ-VFI consists of an implicit acceleration estimation network (IANet) and a VFI back-bone, the former fully leverages spatio-temporal information to explore latent acceleration priors between two input frames, which is then used to progressively modulate linear motions from the latter into quadratic motions in coarse-to-fine manner. Furthermore, to encourage both components to distill more acceleration and motion cues oriented towards VFI, we propose a knowledge distillation strategy in which implicit acceleration distillation loss and implicit motion distillation loss are employed to adaptively guide latent acceleration priors and intermediate motions learning, respectively. Extensive experiments show that our proposed IQ-VFI can achieve state-of-the-art performances on various benchmark datasets. Mengshun Hu, Kui Jiang, Zhihang Zhong, Zheng Wang 0007, Yinqiang Zheng |
CVPR | 2 |
| 2024 | OpticalDR: A Deep Optical Imaging Model for Privacy-Protective Depression RecognitionabstractDepression Recognition (DR) poses a considerable chal-lenge, especially in the context of the growing concerns surrounding privacy. Traditional automatic diagnosis of DR technology necessitates the use of facial images, un-doubtedly expose the patient identity features and poses privacy risks. In order to mitigate the potential risks as-sociated with the inappropriate disclosure of patient fa-cial images, we design a new imaging system to erase the identity information of captured facial images while re-tain disease-relevant features. It is irreversible for identity information recovery while preserving essential disease-related characteristics necessary for accurate DR. More specifically, we try to record a de-identified facial image (erasing the identifiable features as much as possible) by a learnable lens, which is optimized in conjunction with the following DR task as well as a range of face analy-sis related auxiliary tasks in an end-to-end manner. These aforementioned strategies form our final Optical deep De-pression Recognition network (OpticalDR). Experiments on CelebA, AVEC 2013, and AVEC 2014 datasets demonstrate that our OpticalDR has achieved state-of-the-art privacy protection performance with an average AUC of 0.51 on popular facial recognition models, and competitive results for DR with MAEIRMSE of 7.5318.48 on AVEC 2013 and 7.8918.82 on AVEC 2014, respectively. Code is available at https://github.com/divertingPanIOpticalDR. Junjun Jiang, Kui Jiang, Keyuan Yu, Xianming Liu 0005 |
CVPR | 3 |
| 2024 | Dynamic Policy-Driven Adaptive Multi-Instance Learning for Whole Slide Image ClassificationabstractMulti-Instance Learning (MIL) has shown impressive performance for histopathology whole slide image (WSI) analysis using bags or pseudo-bags. It involves instance sampling, feature representation, and decision-making. However, existing MIL-based technologies at least suffer from one or more of the following problems: 1) requiring high storage and intensive preprocessing for numerous instances (sampling); 2) potential over-fitting with limited knowledge to predict bag labels (feature representation); 3) pseudo-bag counts and prior biases affect model robustness and generalizability (decision-making). Inspired by clinical diagnostics, using the past sampling instances can facili-tate the final WSI analysis, but it is barely explored in prior technologies. To break free these limitations, we integrate the dynamic instance sampling and reinforcement learning into a unified framework to improve the instance selection and feature aggregation, forming a novel Dynamic Policy Instance Selection (DPIS) scheme for better and more cred-ible decision-making. Specifically, the measurement of feature distance and reward function are employed to boost continuous instance sampling. To alleviate the over-fitting, we explore the latent global relations among instances for more robust and discriminative feature representation while establishing reward and punishment mechanisms to correct biases in pseudo-bags using contrastive learning. These strategies form the final Dynamic Policy-Driven Adaptive Multi-Instance Learning (PAMIL) method for WSI tasks. Extensive experiments reveal that our PAMIL method outperforms the state-of-the-art by 3.8% on CAMELYON16 and 4.4% on TCGA lung cancer datasets. Kui Jiang, Hongxun Yao |
CVPR | 2 |
| 2024 | Improving Domain Generalization in Self-supervised Monocular Depth Estimation via Stabilized Adversarial Training
Yuanqi Yao, Gang Wu 0010, Kui Jiang, Siao Liu, Jian Kuai, Xianming Liu 0005, Junjun Jiang |
ECCV (24) | 3 |
| 2024 | Mutuality Attribute Makes Better Video Anomaly DetectionabstractVideo anomaly detection (VAD) is an essential but challenging task. Existing prevalent methods focus on analyzing the reconstruction or prediction difference between normal and abnormal patterns through multiple deep features, e.g., optic flow. However, these approaches independently use deep features to characterize attributes, ignore the mutuality among multiple deep features. Therefore, the constructed representation is limited to indirectly representing the anomaly from isolated attributes, and makes the network difficult to capture the high-level causes of anomaly. In this paper, we proposed a novel Mutuality Attribute-based Representation framework (MAR-VAD) for the VAD task, which absorbs the mutuality among deep features to characterize the mutuality attribute. Specifically, the mutuality attribute encapsulates high-level semantic information, such as the specific abnormal object or action, which mutually utilizes information from multiple deep features. In this way, the system is able to directly capture the high-level causes of anomaly, thus providing a more comprehensive perspective to accurately detect anomaly events. Following a process-transparent density estimation, we produce the final anomaly scores. Experiments show that MAR-VAD achieves state-of-the-art performance on ShanghaiTech and Avenue. Xingshuo Han, Xiao Wang 0029, Kui Jiang, Wei Liu 0183, Ruimin Hu, Xuefeng Pan, Xin Xu 0007 |
ICASSP | 3 |
| 2024 | Exploiting Self-Supervised Constraints in image Super-ResolutionabstractRecent advances in self-supervised learning, predominantly studied in high-level visual tasks, have been explored in low-level image processing. This paper introduces a novel self-supervised constraint for single image super-resolution, termed SSC-SR. SSC-SR uniquely addresses the divergence in image complexity by employing a dual asymmetric paradigm and a target model updated via exponential moving average to enhance stability. The proposed SSC-SR framework works as a plug-and-play paradigm and can be easily applied to existing SR models. Empirical evaluations reveal that our SSC-SR framework delivers substantial enhancements on a variety of benchmark datasets, achieving an average increase of 0.1 dB over EDSR and 0.06 dB over SwinIR. In addition, extensive ablation studies corroborate the effectiveness of each component in our SSC-SR framework. Codes are available at https://github.com/Aitical/SSCSR. Gang Wu 0010, Junjun Jiang, Kui Jiang, Xianming Liu 0005 |
ICME | 3 |
| 2024 | Learning a Spiking Neural Network for Efficient Image Deraining
Tianyu Song 0003, Guiyue Jin, Pengpeng Li 0001, Kui Jiang, Xiang Chen 0015, Jiyu Jin |
IJCAI | 4 |
| 2024 | Harmony in Diversity: Improving All-in-One Image Restoration via Multi-Task CollaborationabstractDeep learning-based all-in-one image restoration methods have garnered significant attention in recent years due to capable of addressing multiple degradation tasks. These methods focus on extracting task-oriented information to guide the unified model and have achieved promising results through elaborate architecture design. They commonly adopt a simple mix training paradigm, and the proper optimization strategy for all-in-one tasks has been scarcely investigated. This oversight neglects the intricate relationships and potential conflicts among various restoration tasks, consequently leading to inconsistent optimization rhythms. In this paper, we extend and redefine the conventional all-in-one image restoration task as a multi-task learning problem and propose a straightforward yet effective active-reweighting strategy, dubbed Art, to harmonize the optimization of multiple degradation tasks. Art is a plug-and-play optimization strategy designed to mitigate hidden conflicts among multi-task optimization processes. Through extensive experiments on a diverse range of all-in-one image restoration settings, Art has been demonstrated to substantially enhance the performance of existing methods. When incorporated into the AirNet and TransWeather models, it achieves average improvements of 1.16 dB and 1.21 dB on PSNR, respectively. We hope this work will provide a principled framework for collaborating multiple tasks in all-in-one image restoration and pave the way for more efficient and effective restoration models, ultimately advancing the state-of-the-art in this critical research domain. Code and pre-trained models are available at our project page https://github.com/Aitical/Art. Gang Wu 0010, Junjun Jiang, Kui Jiang, Xianming Liu 0005 |
ACM Multimedia | 3 |
| 2024 | Disentangled-Multimodal Privileged Knowledge Distillation for Depression Recognition with Incomplete Multimodal DataabstractDepression recognition (DR) using facial images, audio signals, or language text recordings has achieved remarkable performance. Recently, multimodal DR has shown improved performance over single-modal methods by leveraging information from a combination of these modalities. However, collecting high-quality data containing all modalities poses a challenge. In particular, these methods often encounter performance degradation when certain modalities are either missing or degraded. To tackle this issue, we present a generalizable multimodal framework for DR by aggregating feature disentanglement and privileged knowledge distillation. In detail, our approach aims to disentangle homogeneous and heterogeneous features within multimodal signals while suppressing noise, thereby adaptively aggregating the most informative components for high-quality DR. Subsequently, we leverage knowledge distillation to transfer privileged knowledge from complete modalities to the observed input with limited information, thereby significantly improving the tolerance and compatibility. These strategies form our novel Feature Disentanglement and Privileged knowledge Distillation Network for DR, dubbed Dis2DR. Experimental evaluations on AVEC 2013, AVEC 2014, AVEC 2017, and AVEC 2019 datasets demonstrate the effectiveness of our Dis2DR method. Remarkably, Dis2DR achieves superior performance even when only a single modality is available, surpassing existing state-of-the-art multimodal DR approaches AVA-DepressNet by up to 9.8% on the AVEC 2013 dataset. Junjun Jiang, Kui Jiang, Xianming Liu 0005 |
ACM Multimedia | 3 |
| 2024 | F4SR: A Feed-Forward Regression Approach for Few-Shot Face Super-Resolution
Jican Fu, Kui Jiang |
PRCV (8) | 2 |
| 2024 | Development and validation of an explainable machine learning model for predicting multidimensional frailty in hospitalized patients with cirrhosisabstractWe sought to develop and validate a machine learning (ML) model for predicting multidimensional frailty based on clinical and laboratory data. Moreover, an explainable ML model utilizing SHapley Additive exPlanations (SHAP) was constructed. This study enrolled 622 patients hospitalized due to decompensating episodes at a tertiary hospital. The cohort data were randomly divided into training and test sets. External validation was carried out using 131 patients from other tertiary hospitals. The frail phenotype was defined according to a self-reported questionnaire (Frailty Index). The area under the receiver operating characteristics curve was adopted to compare the performance of five ML models. The importance of the features and interpretation of the ML models were determined using the SHAP method. The proportions of cirrhotic patients with nonfrail and frail phenotypes in combined training and test sets were 87.8% and 12.2%, respectively, while they were 88.5% and 11.5% in the external validation dataset. Five ML algorithms were used, and the random forest (RF) model exhibited substantially predictive performance. Regarding the external validation, the RF algorithm outperformed other ML models. Moreover, the SHAP method demonstrated that neutrophil-to-lymphocyte ratio, age, lymphocyte-to-monocyte ratio, ascites, and albumin served as the most important predictors for frailty. At the patient level, the SHAP force plot and decision plot exhibited a clinically meaningful explanation of the RF algorithm. We constructed an ML model (RF) providing accurate prediction of frail phenotype in decompensated cirrhosis. The explainability and generalizability may foster clinicians to understand contributors to this physiologically vulnerable situation and tailor interventions. Yumei He, Kui Jiang |
Briefings Bioinform. | 6 |
| 2024 | Contour Counts: Restricting Deformation for Accurate Animation InterpolationabstractAnimated videos with low frame rates commonly degrade the visual experience with choppy motion. Recently, video frame interpolation has been growing rapidly which can increase frame rates. However, the sparse texture information and complex motion scenes of animated videos make the objects in the frames generated by existing video frame interpolation methods appear significantly deformation, causing distortion in the content of generated frames. To address this issue, the Restricting Deformation by Contour Network (RDC-Net) is proposed to repair and fill the content by optimizing object contours and leveraging context features for high-quality animation interpolation. Specifically, the RDC-Net was proposed to learn optical flow maps utilized to effectively capture the spatial shifts in the motion subject's contours across time intervals. Furthermore, contour information is used to refine the structure of the object in optical flow estimation and moderate the object deformation scale in the generated frame. In addition, the context information is explored to characterize the motion between adjacent frames via bidirectional optical flow learning, enabling the filling of distortion in the content of generated frames by feature-filling technology. Experiments on commonly used benchmarks show our state-of-the-art performance. Xin Xu 0007, Kui Jiang, Wei Liu 0183, Zheng Wang 0007 |
IEEE Signal Process. Lett. | 3 |
| 2024 | Local-Global Temporal Difference Learning for Satellite Video Super-ResolutionabstractOptical-flow-based and kernel-based approaches have been extensively explored for temporal compensation in satellite Video Super-Resolution (VSR). However, these techniques are less generalized in large-scale or complex scenarios, especially in satellite videos. In this paper, we propose to exploit the well-defined temporal difference for efficient and effective temporal compensation. To fully utilize the local and global temporal information within frames, we systematically modeled the short-term and long-term temporal discrepancies since we observe that these discrepancies offer distinct and mutually complementary properties. Specifically, we devise a Short-term Temporal Difference Module (S-TDM) to extract local motion representations from RGB difference maps between adjacent frames, which yields more clues for accurate texture representation. To explore the global dependency in the entire frame sequence, a Long-term Temporal Difference Module (L-TDM) is proposed, where the differences between forward and backward segments are incorporated and activated to guide the modulation of the temporal feature, leading to a holistic global compensation. Moreover, we further propose a Difference Compensation Unit (DCU) to enrich the interaction between the spatial distribution of the target frame and temporal compensated results, which helps maintain spatial consistency while refining the features to avoid misalignment. Rigorous objective and subjective evaluations conducted across five mainstream video satellites demonstrate that our method performs favorably against state-of-the-art approaches. Code will be available athttps://github.com/XY-boy/LGTD. Yi Xiao 0003, Qiangqiang Yuan, Kui Jiang, Xianyu Jin, Liangpei Zhang 0001, Chia-Wen Lin |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2024 | EDiffSR: An Efficient Diffusion Probabilistic Model for Remote Sensing Image Super-ResolutionabstractRecently, convolutional networks have achieved remarkable development in remote sensing image (RSI) super-resolution (SR) by minimizing the regression objectives, e.g., MSE loss. However, despite achieving impressive performance, these methods often suffer from poor visual quality with oversmooth issues. Generative adversarial networks (GANs) have the potential to infer intricate details, but they are easy to collapse, resulting in undesirable artifacts. To mitigate these issues, in this article, we first introduce diffusion probabilistic model (DPM) for efficient RSI SR, dubbed efficient diffusion model for RSI SR (EDiffSR). EDiffSR is easy to train and maintains the merits of DPM in generating perceptual-pleasant images. Specifically, different from previous works using heavy UNet for noise prediction, we develop an efficient activation network (EANet) to achieve favorable noise prediction performance by simplified channel attention and simple gate operation, which dramatically reduces the computational budget. Moreover, to introduce more valuable prior knowledge into the proposed EDiffSR, a practical conditional prior enhancement module (CPEM) is developed to help extract an enriched condition. Unlike most DPM-based SR models that directly generate conditions by amplifying LR images, the proposed CPEM helps to retain more informative cues for accurate SR. Extensive experiments on four remote sensing datasets demonstrate that EDiffSR can restore visual-pleasant images on simulated and real-world RSIs, both quantitatively and qualitatively. The code of EDiffSR will be available athttps://github.com/XY-boy/EDiffSR. Yi Xiao 0003, Qiangqiang Yuan, Kui Jiang, Xianyu Jin, Liangpei Zhang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2024 | Coarse- and Fine-Grained Fusion Hierarchical Network for Hole Filling in View SynthesisabstractDepth image-based rendering (DIBR) techniques play an essential role in free-viewpoint videos (FVVs), which generate the virtual views from a reference 2D texture video and its associated depth information. However, the background regions occluded by the foreground in the reference view will be exposed in the synthesized view, resulting in obvious irregular holes in the synthesized view. To this end, this paper proposes a novel coarse and fine-grained fusion hierarchical network (CFFHNet) for hole filling, which fills the irregular holes produced by view synthesis using the spatial contextual correlations between the visible and hole regions. CFFHNet adopts recurrent calculation to learn the spatial contextual correlation, while the hierarchical structure and attention mechanism are introduced to guide the fine-grained fusion of cross-scale contextual features. To promote texture generation while maintaining fidelity, we equip CFFHNet with a two-stage framework involving an inference sub-network to generate the coarse synthetic result and a refinement sub-network for refinement. Meanwhile, to make the learned hole-filling model better adaptable and robust to the "foreground penetration" distortion, we trained CFFHNet by generating a batch of training samples by adding irregular holes to the foreground and background connection regions of high-quality images. Extensive experiments show the superiority of our CFFHNet over the current state-of-the-art DIBR methods. The source code will be available at https://github.com/wgc-vsfm/view-synthesis-CFFHNet. Guangcheng Wang, Kui Jiang, Ke Gu 0001, Hongyan Liu 0004, Hantao Liu, Wenjun Zhang 0001 |
IEEE Trans. Image Process. | 2 |
| 2024 | Multi-Scale Fusion and Decomposition Network for Single Image DerainingabstractConvolutional neural networks (CNNs) and self-attention (SA) have demonstrated remarkable success in low-level vision tasks, such as image super-resolution, deraining, and dehazing. The former excels in acquiring local connections with translation equivariance, while the latter is better at capturing long-range dependencies. However, both CNNs and Transformers suffer from individual limitations, such as limited receptive field and weak diversity representation of CNNs during low efficiency and weak local relation learning of SA. To this end, we propose a multi-scale fusion and decomposition network (MFDNet) for rain perturbation removal, which unifies the merits of these two architectures while maintaining both effectiveness and efficiency. To achieve the decomposition and association of rain and rain-free features, we introduce an asymmetrical scheme designed as a dual-path mutual representation network that enables iterative refinement. Additionally, we incorporate high-efficiency convolutions throughout the network and use resolution rescaling to balance computational complexity with performance. Comprehensive evaluations show that the proposed approach outperforms most of the latest SOTA deraining methods and is versatile and robust in various image restoration tasks, including underwater image enhancement, image dehazing, and low-light image enhancement. The source codes and pretrained models are available at https://github.com/qwangg/MFDNet. Kui Jiang, Zheng Wang 0007, Wenqi Ren, Chia-Wen Lin |
IEEE Trans. Image Process. | 2 |
| 2024 | TTST: A Top-k Token Selective Transformer for Remote Sensing Image Super-ResolutionabstractTransformer-based method has demonstrated promising performance in image super-resolution tasks, due to its long-range and global aggregation capability. However, the existing Transformer brings two critical challenges for applying it in large-area earth observation scenes: (1) redundant token representation due to most irrelevant tokens; (2) single-scale representation which ignores scale correlation modeling of similar ground observation targets. To this end, this paper proposes to adaptively eliminate the interference of irreverent tokens for a more compact self-attention calculation. Specifically, we devise a Residual Token Selective Group (RTSG) to grasp the most crucial token by dynamically selecting the top- k keys in terms of score ranking for each query. For better feature aggregation, a Multi-scale Feed-forward Layer (MFL) is developed to generate an enriched representation of multi-scale feature mixtures during feed-forward process. Moreover, we also proposed a Global Context Attention (GCA) to fully explore the most informative components, thus introducing more inductive bias to the RTSG for an accurate reconstruction. In particular, multiple cascaded RTSGs form our final Top- k Token Selective Transformer (TTST) to achieve progressive representation. Extensive experiments on simulated and real-world remote sensing datasets demonstrate our TTST could perform favorably against state-of-the-art CNN-based and Transformer-based methods, both qualitatively and quantitatively. In brief, TTST outperforms the state-of-the-art approach (HAT-L) in terms of PSNR by 0.14 dB on average, but only accounts for 47.26% and 46.97% of its computational cost and parameters. The code and pre-trained TTST will be available at https://github.com/XY-boy/TTST for validation. Yi Xiao 0003, Qiangqiang Yuan, Kui Jiang, Chia-Wen Lin, Liangpei Zhang 0001 |
IEEE Trans. Image Process. | 3 |
| 2024 | Omniscient Video Super-Resolution with Explicit-Implicit AlignmentabstractWhen considering the temporal relationships, most previous video super-resolution (VSR) methods follow the iterative or recurrent framework. The iterative framework adopts neighboring low-resolution (LR) frames from a sliding window, while the recurrent framework utilizes the output generated in the previous SR procedure. The hybrid framework combines them but still cannot fully leverage the temporal relationships. Meanwhile, the existing methods are limited in the receptive field of the optical flow or lack semantic constrains on motion information. In this work, we propose an omniscient framework to fully explore the temporal relationships in the video, which encompasses both LR frames and SR outputs from the past, present, and future. The omniscient framework is more generic because the iterative, recurrent, and hybrid frameworks can be regarded as its special cases. Besides, when addressing the motion information, most previous VSR methods adopt the explicit motion estimation and compensation, while many recent methods turn to implicit alignment. In implicit alignment methods, because basic non-local means suffers from heavy computational costs, we improve it by capturing the non-local correlations in a relatively local manner to reduce the complexity. Moreover, we integrate the explicit and implicit methods into an explicit-implicit alignment module to better utilize motion information. We have conducted extensive experiments on public datasets, which show that our method is superior over the state-of-the-art methods in objective metrics, subjective visual quality, and complexity. In particular, on datasets of Vid4 and UDM10, our method improves PSNR by 0.19 dB, 0.49 dB against the most advanced method BasicVSR++, respectively. Peng Yi 0002, Zhongyuan Wang 0002, Laigan Luo, Kui Jiang, Zheng He 0001, Junjun Jiang, Tao Lu 0001, Jiayi Ma 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2023 | Store and Fetch Immediately: Everything Is All You Need for Space-Time Video Super-resolutionabstractExisting space-time video super-resolution (ST-VSR) methods fail to achieve high-quality reconstruction since they fail to fully explore the spatial-temporal correlations, long-range components in particular. Although the recurrent structure for ST-VSR adopts bidirectional propagation to aggregate information from the entire video, collecting the temporal information between the past and future via one-stage representations inevitably loses the long-range relations. To alleviate the limitation, this paper proposes an immediate storeand-fetch network to promote long-range correlation learning, where the stored information from the past and future can be refetched to help the representation of the current frame. Specifically, the proposed network consists of two modules: a backward recurrent module (BRM) and a forward recurrent module (FRM). The former first performs backward inference from future to past, while storing future super-resolution (SR) information for each frame. Following that, the latter performs forward inference from past to future to super-resolve all frames, while storing past SR information for each frame. Since FRM inherits SR information from BRM, therefore, spatial and temporal information from the entire video sequence is immediately stored and fetched, which allows drastic improvement for ST-VSR. Extensive experiments both on ST-VSR and space video super-resolution (S-VSR) as well as time video super-resolution (T-VSR) have demonstrated the effectiveness of our proposed method over other state-of-the-art methods on public datasets. Code is available https://github.com/hhhhhumengshun/SFI-STVR Mengshun Hu, Kui Jiang, Zhixiang Nie, Jiahuan Zhou, Zheng Wang 0007 |
AAAI | 2 |
| 2023 | Refined Semantic Enhancement towards Frequency Diffusion for Video CaptioningabstractVideo captioning aims to generate natural language sentences that describe the given video accurately. Existing methods obtain favorable generation by exploring richer visual representations in encode phase or improving the decoding ability. However, the long-tailed problem hinders these attempts at low-frequency tokens, which rarely occur but carry critical semantics, playing a vital role in the detailed generation. In this paper, we introduce a novel Refined Semantic enhancement method towards Frequency Diffusion (RSFD), a captioning model that constantly perceives the linguistic representation of the infrequent tokens. Concretely, a Frequency-Aware Diffusion (FAD) module is proposed to comprehend the semantics of low-frequency tokens to break through generation limitations. In this way, the caption is refined by promoting the absorption of tokens with insufficient occurrence. Based on FAD, we design a Divergent Semantic Supervisor (DSS) module to compensate for the information loss of high-frequency tokens brought by the diffusion process, where the semantics of low-frequency tokens is further emphasized to alleviate the long-tailed problem. Extensive experiments indicate that RSFD outperforms the state-of-the-art methods on two benchmark datasets, i.e., MSR-VTT and MSVD, demonstrate that the enhancement of low-frequency tokens semantics can obtain a competitive generation effect. Code is available at https://github.com/lzp870/RSFD. Xian Zhong, Shuqin Chen, Kui Jiang, Chen Chen 0001, Mang Ye |
AAAI | 4 |
| 2023 | Implicit Attention-Based Cross-Modal Collaborative Learning for Action RecognitionabstractHuman action recognition is an active research topic in recent years. Multiple modalities often convey heterogeneous but potentially complementary action information that single modality does not hold. Some efforts have been resoted to explore cross-modal representation to promote the modeling capability, but with limited improvement due to the simple fusion of different modalities. To this end, we propose an impliCit attention-based Cross-modal Collaborative Learning (C3L) for action recognition. Specifically, we apply a Modality Generalization network with Grayscale enhancement (MGG) to learn specific modality representation and interaction (infrared and RGB). Then, we construct a unified representation space through the Uniform Modality Representation module (UMR), which preserves the modality information while enhancing the overall representation ability. Finally, feature extractors adaptively leverage modality-specific knowledge to realize cross-modal collaborative learning. Extensive experiments conducted on three widely-used public benchmarks InfAR, HMDB51, and UCF101, demonstrate the effectiveness and strength of our proposed method. Jianghao Zhang, Xian Zhong, Wenxuan Liu 0008, Kui Jiang, Zhengwei Yang 0001, Zheng Wang 0007 |
ICIP | 4 |
| 2023 | HPCNet: A Hybrid Progressive Coupled Network for Image DerainingabstractConvolutional neural networks (CNNs) and Transformers have significantly succeeded in low-level vision tasks. Although prominent complementary characteristics exist regarding the larger receptive field and better convergence, only some efforts have compacted them efficiently due to their individual and nonnegligible weakness. In this paper, we propose a hybrid progressive coupled network (HPCNet) for rain perturbation removal, which integrates the advantages of these two architectures while maintaining both effectiveness and efficiency. In particular, we achieve the progressive decomposition and association of rain-free and rain features, designed as an asymmetrical dual-path mutual representation network to alleviate the computational cost. Meanwhile, we equip the network with high-efficiency convolutions and resolution rescaling strategy to trade off the computational complexity. Extensive experiments show that our method outperforms MPRNet on average while saving 87.2% and 61.1% of the computational cost and parameters. Kui Jiang, Jinyi Lai, Zheng Wang 0007 |
ICME | 2 |
| 2023 | From Generation to Suppression: Towards Effective Irregular Glow Removal for Nighttime Visibility EnhancementabstractMost existing Low-Light Image Enhancement (LLIE) methods are primarily designed to improve brightness in dark regions, which suffer from severe degradation in nighttime images. However, these methods have limited exploration in another major visibility damage, the glow effects in real night scenes. Glow effects are inevitable in the presence of artificial light sources and cause further diffused blurring when directly enhanced. To settle this issue, we innovatively consider the glow suppression task as learning physical glow generation via multiple scattering estimation according to the Atmospheric Point Spread Function (APSF). In response to the challenges posed by uneven glow intensity and varying source shapes, an APSF-based Nighttime Imaging Model with Near-field Light Sources (NIM-NLS) is specifically derived to design a scalable Light-aware Blind Deconvolution Network (LBDN). The glow-suppressed result is then brightened via a Retinex-based Enhancement Module (REM). Remarkably, the proposed glow suppression method is based on zero-shot learning and does not rely on any paired or unpaired training data. Empirical evaluations demonstrate the effectiveness of the proposed method in both glow suppression and low-light enhancement tasks. Wanyu Wu, Wei Wang 0170, Zheng Wang 0007, Kui Jiang, Xin Xu 0007 |
IJCAI | 4 |
| 2023 | DAWN: Direction-aware Attention Wavelet Network for Image DerainingabstractSingle image deraining aims to remove rain perturbation while restoring the clean background scene from a rain image. However, existing methods tend to produce blurry and over-smooth outputs, lacking some textural details. Wavelet transform can depict the contextual and textural information of an image at different levels, showing impressive capability of learning structural information in the images to avoid artifacts, and thus has been recently explored to consider the inherent overlap of background and rain perturbation in both the pixel domain and the frequency embedding space. However, the existing wavelet-based methods ignore the heterogeneous degradation for different coefficients due to the inherent directional characteristics of rain streaks, leading to inter-frequency conflicts and compromised deraining results. To address this issue, we propose a novel Direction-aware Attention Wavelet Network (DAWN) for rain streaks removal. DAWN has several key distinctions from existing wavelet transform-based methods: 1) introducing the vector decomposition to parameterize the learning procedure, where the rain streaks are derived into the vertical (V) and horizontal (H) components to learn the specific representation; 2) a novel direction-aware attention module (DAM) to fit the projection and transformation parameters to characterize the direction-specific rain components, which helps accurate texture restoration; 3) exploring practical composite constraints on the structure, details, and chrominance aspects for high-quality background restoration. Our proposed DAWN delivers significant performance gains on nine datasets across image deraining and object detection tasks, exceeding the state-of-the-art method MPRNet by 0.88 dB in PSNR on the Test1200 dataset with only 35.5% computation cost. Kui Jiang, Wenxuan Liu 0008, Zheng Wang 0007, Xian Zhong, Junjun Jiang, Chia-Wen Lin |
ACM Multimedia | 1 |
| 2023 | Informative Classes Matter: Towards Unsupervised Domain Adaptive Nighttime Semantic SegmentationabstractUnsupervised Domain Adaptive Nighttime Semantic Segmentation (UDA-NSS) aims to adapt a robust model from a labeled daytime domain to an unlabeled nighttime domain. However, current advanced segmentation methods ignore the illumination effect and class discrepancies of different semantic classes during domain adaptation, showing an uneven prediction phenomenon. It is the completely ignored and underexplored issues of ''hard-to-adapt'' classes that some classes have a large performance gap between existing UDA-NSS methods and supervised learning counterparts while others have a very low performance gap. To realize ''hard-to-adapt'' classes' more sufficient learning and facilitate the UDA-NSS task, we present an Online Informative Class Sampling (OICS) strategy to adaptively mine informative classes from the target nighttime domain according to the corresponding spectrogram mean and the class frequency via our Informative Mixture of Experts. Furthermore, an Informativeness-based cross-domain Mixed Sampling (InforMS) framework is designed to focus on informative classes from the target nighttime domain by vesting their higher sampling probabilities when cross-domain mixing sampling and achieves better performance in UDA-NSS tasks. Consequently, our method outperforms state-of-the-art UDA-NSS methods by large margins on three widely-used benchmarks (e.g., ACDC, Dark Zurich, and Nighttime Driving). Notably, our method achieves state-of-the-art performance with 65.1% mIoU on ACDC-night-test and 55.4% mIoU on ACDC-night-val. Shiqin Wang, Xin Xu 0007, Xianzheng Ma, Kui Jiang, Zheng Wang 0007 |
ACM Multimedia | 4 |
| 2023 | CycMuNet+: Cycle-Projected Mutual Learning for Spatial-Temporal Video Super-ResolutionabstractSpatial-Temporal Video Super-Resolution (ST-VSR) aims to generate high-quality videos with higher resolution (HR) and higher frame rate (HFR). Quite intuitively, pioneering two-stage based methods complete ST-VSR by directly combining two sub-tasks: Spatial Video Super-Resolution (S-VSR) and Temporal Video Super-Resolution (T-VSR) but ignore the reciprocal relations among them. 1) T-VSR to S-VSR: temporal correlations help accurate spatial detail representation; 2) S-VSR to T-VSR: abundant spatial information contributes to the refinement of temporal prediction. To this end, we propose a one-stage based Cycle-projected Mutual learning network (CycMuNet) for ST-VSR, which makes full use of spatial-temporal correlations via the mutual learning between S-VSR and T-VSR. Specifically, we propose to exploit the mutual information among them via iterative up- and down projections, where spatial and temporal features are fully fused and distilled, helping high-quality video reconstruction. In addition, we also show interesting extensions for efficient network design (CycMuNet+), such as parameter sharing and dense connection on projection units and feedback mechanism in CycMuNet. Besides extensive experiments on benchmark datasets, we also compare our proposed CycMuNet (+) with S-VSR and T-VSR tasks, demonstrating that our method significantly outperforms the state-of-the-art methods. Mengshun Hu, Kui Jiang, Zheng Wang 0007, Xiang Bai, Ruimin Hu |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2023 | PLFace: Progressive Learning for Face Recognition with Mask Bias
Baojin Huang, Zhongyuan Wang 0001, Guangcheng Wang, Kui Jiang, Zhen Han 0002, Tao Lu 0001, Chao Liang 0001 |
Pattern Recognit. | 4 |
| 2023 | Dual-Recommendation Disentanglement Network for View Fuzz in Action RecognitionabstractMulti-view action recognition aims to identify action categories from given clues. Existing studies ignore the negative influences of fuzzy views between view and action in disentangling, commonly arising the mistaken recognition results. To this end, we regard the observed image as the composition of the view and action components, and give full play to the advantages of multiple views via the adaptive cooperative representation among these two components, forming a Dual-Recommendation Disentanglement Network (DRDN) for multi-view action recognition. Specifically, 1) For the action, we leverage a multi-level Specific Information Recommendation (SIR) to enhance the interaction among intricate activities and views. SIR offers a more comprehensive representation of activities, measuring the trade-off between global and local information. 2) For the view, we utilize a Pyramid Dynamic Recommendation (PDR) to learn a complete and detailed global representation by transferring features from different views. It is explicitly restricted to resist the fuzzy noise influence, focusing on positive knowledge from other views. Our DRDN aims for complete action and view representation, where PDR directly guides action to disentangle with view features and SIR considers mutual exclusivity of view and action clues. Extensive experiments have indicated that the multi-view action recognition method DRDN we proposed achieves state-of-the-art performance over powerful competitors on several standard benchmarks. The code will be available at https://github.com/51cloud/DRDN. Wenxuan Liu 0008, Xian Zhong, Zhuo Zhou, Kui Jiang, Zheng Wang 0007, Chia-Wen Lin |
IEEE Trans. Image Process. | 4 |
| 2023 | Visual Exposes You: Pedestrian Trajectory Prediction Meets Visual IntentionabstractPedestrian trajectory prediction in multiple scenarios is of immense importance in autonomous driving and disentanglement of human behavior but is limited in catching human intention and initiative. Most previous works tend to predict the trajectory using only 2D coordinates, which generally cause two common problems: a) Overlooking the subjective initiative, including sudden swerve and erratic movement; b) A potential challenge called abnormal collision caused by unlabeled pedestrians on dataset is not being identified and resolved, which would ruin the model prediction. To break those limitations, we introduce visual localization and orientation as Visual Intention Knowledge to help the trajectory prediction, which is learned directly from visual scenarios. It benefits to comprehend human intention and formulates decision-making processes. Moreover, by learning from the visual information and decision-making policy, we construct the Visual Intention Knowledge associated spatio-temporal Transformer (VIKT) to predict human trajectory by combining the intention knowledge with the novel Transformer. Extensive experimental results demonstrate that our VIKT model could achieve competitive performance by the Visual Intention Knowledge through optimizing the model prediction compared with state-of-the-art methods in terms of prediction accuracy on ETH/UCY and SDD benchmarks. Xian Zhong, Zhengwei Yang 0001, Wenxin Huang, Kui Jiang, Ryan Wen Liu, Zheng Wang 0007 |
IEEE Trans. Intell. Transp. Syst. | 5 |
| 2023 | Joint Segmentation and Identification Feature Learning for Occlusion Face RecognitionabstractThe existing occlusion face recognition algorithms almost tend to pay more attention to the visible facial components. However, these models are limited because they heavily rely on existing face segmentation approaches to locate occlusions, which is extremely sensitive to the performance of mask learning. To tackle this issue, we propose a joint segmentation and identification feature learning framework for end-to-end occlusion face recognition. More particularly, unlike employing an external face segmentation model to locate the occlusion, we design an occlusion prediction module supervised by known mask labels to be aware of the mask. It shares underlying convolutional feature maps with the identification network and can be collaboratively optimized with each other. Furthermore, we propose a novel channel refinement network to cast the predicted single-channel occlusion mask into a multi-channel mask matrix with each channel owing a distinct mask map. Occlusion-free feature maps are then generated by projecting multi-channel mask probability maps onto original feature maps. Thus, it can suppress the representation of occlusion elements in both the spatial and channel dimensions under the guidance of the mask matrix. Moreover, in order to avoid misleading aggressively predicted mask maps and meanwhile actively exploit usable occlusion-robust features, we aggregate the original and occlusion-free feature maps to distill the final candidate embeddings by our proposed feature purification module. Lastly, to alleviate the scarcity of real-world occlusion face recognition datasets, we build large-scale synthetic occlusion face datasets, totaling up to 980193 face images of 10574 subjects for the training dataset and 36721 face images of 6817 subjects for the testing dataset, respectively. Extensive experimental results on the synthetic and real-world occlusion face datasets show that our approach significantly outperforms the state-of-the-art in both 1:1 face verification and 1:N face identification. Baojin Huang, Zhongyuan Wang 0001, Kui Jiang, Qin Zou 0001, Xin Tian 0006, Tao Lu 0001, Zhen Han 0002 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2023 | Multi-Scale Hybrid Fusion Network for Single Image DerainingabstractDeep learning models have been able to generate rain-free images effectively, but the extension of these methods to complex rain conditions where rain streaks show various blurring degrees, shapes, and densities has remained an open problem. Among the major challenges are the capacity to encode the rain streaks and the sheer difficulty of learning multi-scale context features that preserve both global color coherence and exactness of detail. To address the first problem, we design a non-local fusion module (NFM) and an attention fusion module (AFM), and construct the multi-level pyramids' architecture to explore the local and global correlations of rain information from the rain image pyramid. More specifically, we apply the non-local operation to fully exploit the self-similarity of rain streaks and perform the fusion of multi-scale features along the image pyramid. To address the latter challenge, we additionally design a residual learning branch that is capable of adaptively bridging the gaps (e.g., texture and color information) between the predicted rain-free image and the clean background via a hybrid embedding representation. Extensive results have demonstrated that our proposed method is able to generate much better rain-free images on several benchmark datasets than the state-of-the-art algorithms. Moreover, we conduct the joint evaluation experiments with respect to deraining performance and the detection/segmentation accuracy to further verify the effectiveness of our deraining method for downstream vision tasks/applications. The source code is available at https://github.com/kuihua/MSHFN. Kui Jiang, Zhongyuan Wang 0001, Peng Yi 0002, Chen Chen 0001, Guangcheng Wang, Zhen Han 0002, Junjun Jiang, Zixiang Xiong |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2023 | Local Eyebrow Feature Attention Network for Masked Face RecognitionabstractDuring the COVID-19 coronavirus epidemic, wearing masks has become increasingly popular. Traditional occlusion face recognition algorithms are almost ineffective for such heavy mask occlusion. Therefore, it is urgent to improve the recognition performance of the existing face recognition technology on masked faces. Due to the limited visible feature points of the masked face image relative to the normal face image, we have to exploit the identification potential of eyebrow (referring to eyes and brows) features. This article proposes a local eyebrow feature attention network for masked face recognition, which consists of feature extraction, eyebrow region pooling, and feature fusion. To highlight the eyebrow region, we first use the eyebrow region pooling to separate the local features of eyebrows from the learned overall facial features. We then make full use of the symmetry of left and right eyebrows to emphasize their discriminant ability, due to the inadequate fine information of the low-resolution eyebrows. In particular, in view of the symmetrical similarity between eyebrow pairs and the subordinate relationship between facial components and the whole, we propose a feature fusion model based on graph convolutional network (GCN) to learn the feature association structure of eye features, brow features, and global facial features. We construct the benchmark datasets for masked face recognition to validate our approach, including real-world masked face recognition dataset (RMFRD) and synthetic masked face recognition dataset (SMFRD). Extensive experimental results on both public datasets and our built masked face datasets show that our approach significantly outperforms the state-of-the-arts. Baojin Huang, Zhongyuan Wang 0001, Guangcheng Wang, Zhen Han 0002, Kui Jiang |
ACM Trans. Multim. Comput. Commun. Appl. | 5 |
| 2023 | Beyond the Parts: Learning Coarse-to-Fine Adaptive Alignment Representation for Person SearchabstractPerson search is a time-consuming computer vision task that entails locating and recognizing query people in scenic pictures. Body components are commonly mismatched during matching due to position variation, occlusions, and partially absent body parts, resulting in unsatisfactory person search results. Existing approaches for extracting local characteristics of the human body using keypoint information are unable to handle the search job when distinct body parts are misaligned, ignoring to exploit multiple granularities, which is crucial in the person search process. Moreover, the alignment learning methods learn body part features with fixed and equal weights, ignoring the beneficial contextual information, e.g., the umbrella carried by the pedestrian, which supplements compelling clues for identifying the person. In this paper, we propose a Coarse-to-Fine Adaptive Alignment Representation (CFA 2 R) network for learning multiple granular features in misaligned person search in the coarse-to-fine perspective. To exploit more beneficial body parts and related context of the cropped pedestrians, we design a Part-Attentional Progressive Module (PAPM) to guide the network to focus on informative body parts and positive accessorial regions. Besides, we propose a Re-weighting Alignment Module (RAM) shedding light on more contributive parts instead of treating them equally. Specifically, adaptive re-weighted but not fixed part features are reconstructed by Re-weighting Reconstruction module, considering that different parts serve unequally during image matching. Extensive experiments conducted on CUHK-SYSU and PRW datasets demonstrate competitive performance of our proposed method. Wenxin Huang, Xuemei Jia, Xian Zhong, Xiao Wang 0029, Kui Jiang, Zheng Wang 0007 |
ACM Trans. Multim. Comput. Commun. Appl. | 5 |
| 2022 | Degrade Is Upgrade: Learning Degradation for Low-Light Image EnhancementabstractLow-light image enhancement aims to improve an image's visibility while keeping its visual naturalness. Different from existing methods, which tend to accomplish the relighting task directly, we investigate the intrinsic degradation and relight the low-light image while refining the details and color in two steps. Inspired by the color image formulation (diffuse illumination color plus environment illumination color), we first estimate the degradation from low-light inputs to simulate the distortion of environment illumination color, and then refine the content to recover the loss of diffuse illumination color. To this end, we propose a novel Degradation-to-Refinement Generation Network (DRGN). Its distinctive features can be summarized as 1) A novel two-step generation network for degradation learning and content refinement. It is not only superior to one-step methods, but also capable of synthesizing sufficient paired samples to benefit the model training; 2) A multi-resolution fusion network to represent the target information (degradation or contents) in a multi-scale cooperative manner, which is more effective to address the complex unmixing problems. Extensive experiments on both the enhancement task and the joint detection task have verified the effectiveness and efficiency of our proposed method, surpassing the SOTA by 1.59dB on average and 3.18\% in mAP on the ExDark dataset. The code will be available soon. Kui Jiang, Zhongyuan Wang 0001, Zheng Wang 0007, Chen Chen 0001, Peng Yi 0002, Tao Lu 0001, Chia-Wen Lin |
AAAI | 1 |
| 2022 | Unpaired Deep Image Deraining Using Dual Contrastive LearningabstractLearning single image deraining (SID) networks from an unpaired set of clean and rainy images is practical and valuable as acquiring paired real-world data is almost infeasible. However, without the paired data as the supervision, learning a SID network is challenging. Moreover, simply using existing unpaired learning methods (e.g., unpaired adversarial learning and cycle-consistency constraints) in the SID task is insufficient to learn the underlying relationship from rainy inputs to clean outputs as there exists significant domain gap between the rainy and clean images. In this paper, we develop an effective unpaired SID adversarial framework which explores mutual properties of the unpaired exemplars by a dual contrastive learning manner in a deep feature space, named as DCD-GAN. The proposed method mainly consists of two cooperative branches: Bidirectional Translation Branch (BTB) and Contrastive Guidance Branch (CGB). Specifically, BTB exploits full advantage of the circulatory architecture of adversarial consistency to generate abundant exemplar pairs and excavates latent feature distributions between two domains by equipping it with bidirectional mapping. Simultaneously, CGB implicitly constrains the embeddings of different exemplars in the deep feature space by encouraging the similar feature distributions closer while pushing the dissimilar further away, in order to better facilitate rain removal and help image restoration. Extensive experiments demonstrate that our method performs favorably against existing unpaired deraining approaches on both synthetic and real-world datasets, and generates comparable results against several fully-supervised or semi-supervised models. Xiang Chen 0015, Jinshan Pan, Kui Jiang, Yufeng Li 0001, Caihua Kong, Longgang Dai, Zhentao Fan |
CVPR | 3 |
| 2022 | Spatial-Temporal Space Hand-in-Hand: Spatial-Temporal Video Super-Resolution via Cycle-Projected Mutual LearningabstractSpatial-Temporal Video Super-Resolution (ST-VSR) aims to generate super-resolved videos with higher resolution (HR) and higher frame rate (HFR). Quite intuitively, pioneering two-stage based methods complete ST-VSR by directly combining two sub-tasks: Spatial Video Super-Resolution (S-VSR) and Temporal Video Super-Resolution (T-VSR) but ignore the reciprocal relations among them. Specifically, 1) T-VSR to S-VSR: temporal correlations help accurate spatial detail representation with more clues; 2) S-VSR to T-VSR: abundant spatial information contributes to the refinement of temporal prediction. To this end, we propose a one-stage based Cycle-projected Mutual learning network (CycMu-Net) for ST-VSR, which makes full use of spatial-temporal correlations via the mutual learning between S-VSR and T-VSR. Specifically, we propose to exploit the mutual information among them via iterative up-and-down projections, where the spatial and temporal features are fully fused and distilled, helping the high-quality video reconstruction. Besides extensive experiments on benchmark datasets, we also compare our proposed CycMu-Net with S-VSR and T-VSR tasks, demonstrating that our method significantly outperforms state-of-the-art methods. Codes are publicly available at: https://github.com/hhhhhumengshun/CycMuNet. Mengshun Hu, Kui Jiang, Jing Xiao 0004, Junjun Jiang, Zheng Wang 0007 |
CVPR | 2 |
| 2022 | Self-Supervised Learning on A Lightweight Low-Light Image Enhancement Model with Curve RefinementabstractDeep learning networks with deeper layers become a trend for their good performance but lacks the potential for real-time mobile deployment. Another challenge for paired training networks is the limited generalization capacity caused by the sample bias. To overcome these two challenges, we propose a lightweight self-supervised low-light image enhancement method, that trains with low light images only. Specifically, our method consists of a low-resolution dense CNN network stream and a full-resolution guidance stream, responsible for image-to-curve transformation with refinement and spatial guidance fusion, respectively. Then, a new self-supervised loss function is introduced to measure the restored patch-based color deviations among color channels. Experimental results show that our method gives competitive performance to the full-supervised approaches. Wanyu Wu, Wei Wang 0170, Kui Jiang, Xin Xu 0007, Ruimin Hu |
ICASSP | 3 |
| 2022 | VCD: View-Constraint Disentanglement for Action RecognitionabstractAction recognition is a hot topic in computer vision due to its wide range of applications in urban surveillance. Although some methods are more advanced from an invariant view perspective, those approaches do not perform well for the viewpoint change. To address this issue, one possible solution is tantamount to track the view-invariant representation as it evolves with the performed action. However, the views’ and actions’ performance always complement each other, once simply looking for the view-invariant representation may cause some behavior information to be lost. In this paper, we propose the View-Constraint Disentanglement (VCD) framework for cross-view action recognition. Specifically, Constraint Disentanglement Module (CDM) is utilized to learn an action-invariant representation by discretizing view-specific representation and its normal distribution, which resolves the entangled relationship between view and action. Moreover, a novel Adaptive Distribution Module (ADM) is intended to befit enhance the high-correlation viewpoint variation information and refine the suitable weight. Extensive experiments are conducted on public benchmarks, indicating that our approach achieves better performance than other state-of-the-art approaches. Xian Zhong, Zhuo Zhou, Wenxuan Liu 0008, Kui Jiang, Xuemei Jia, Wenxin Huang, Zheng Wang 0007 |
ICASSP | 4 |
| 2022 | Dual-Scale Alignment-Based Transformer on Linguistic Skeleton Tags for Non-Autoregressive Video CaptioningabstractDue to the characteristic of one-time parallel generation of a caption, non-autoregressive video captioning lacks strong dependencies between words. Although using guideline of scene-related visual words can promote caption generation, the semantic relations among visual words are barely explored, limiting the accurate representation. To this end, we propose a Dual-Scale Alignment-based transformer on Linguistic Skeleton Tags (DSA-LST), which alleviates the defect above in the form of visual words group (several words representing a video frame). Different groups represent different semantic dependencies by attention. We utilize linguistic skeleton tags (i.e., several groups) as sentence-level supervision for visual words sequence. For visual words group to accurately express a specific frame, we further design dual scales of visual-language bi-direction alignment to achieve internal relevance of the tags. Extensive experiments conducted on widely used datasets: MSVD and MSR-VTT demonstrate the effectiveness of our method when compared with existing approaches. Xian Zhong, Shuqin Chen, Zhixin Sun, Huantao Zheng, Kui Jiang |
ICME | 6 |
| 2022 | DANet: Image Deraining via Dynamic Association LearningabstractRain streaks and background components in a rainy input are highly correlated, making the deraining task a composition of the rain streak removal and background restoration. However, the correlation of these two components is barely considered, leading to unsatisfied deraining results. To this end, we propose a dynamic associated network (DANet) to achieve the association learning between rain streak removal and background recovery. There are two key aspects to fulfill the association learning: 1) DANet unveils the latent association knowledge between rain streak prediction and background texture recovery, and leverages it as an extra prior via an associated learning module (ALM) to promote the texture recovery. 2) DANet introduces the parametric association constraint for enhancing the compatibility of deraining model with background reconstruction, enabling it to be automatically learned from the training data. Moreover, we observe that the sampled rainy image enjoys the similar distribution to the original one. We thus propose to learn the rain distribution at the sampling space, and exploit super-resolution to reconstruct high-frequency background details for computation and memory reduction. Our proposed DANet achieves the approximate deraining performance to the state-of-the-art MPRNet but only requires 52.6\% and 23\% inference time and computational cost, respectively. Kui Jiang, Zhongyuan Wang 0001, Zheng Wang 0007, Peng Yi 0002, Junjun Jiang, Jinsheng Xiao, Chia-Wen Lin |
IJCAI | 1 |
| 2022 | Rainy WCity: A Real Rainfall Dataset with Diverse Conditions for Semantic Driving Scene UnderstandingabstractScene understanding in adverse weather conditions (e.g. rainy and foggy days) has drawn increasing attention, arising some specific benchmarks and algorithms. However, scene segmentation under rainy weather is still challenging and under-explored due to the following limitations on the datasets and methods: 1) Manually synthetic rainy samples with empirically settings and human subjective assumptions; 2) Limited rainy conditions, including the rain patterns, intensity, and degradation factors; 3) Separated training manners for image deraining and semantic segmentation. To break these limitations, we pioneer a real, comprehensive, and well-annotated scene understanding dataset under rainy weather, named Rainy WCity. It covers various rain patterns and their bring-in negative visual effects, covering wiper, droplet, reflection, refraction, shadow, windshield-blurring, etc. In addition, to alleviate dependence on paired training samples, we design an unsupervised contrastive learning network for real image deraining and the final rainy scene semantic segmentation via multi-task joint optimization. A comprehensive comparison analysis is also provided, which shows that scene understanding in rainy weather is a largely open problem. Finally, we summarize our general observations, identify open research challenges, and point out future directions. Xian Zhong, Shidong Tu, Xianzheng Ma, Kui Jiang, Wenxin Huang, Zheng Wang 0007 |
IJCAI | 4 |
| 2022 | Progressive Spatial-temporal Collaborative Network for Video Frame InterpolationabstractMost video frame interpolation (VFI) algorithms infer the intermediate frame with the help of adjacent frames through the cascaded motion estimation and content refinement.However, the intrinsic correlations between motion and content are barely investigated, commonly producing interpolated results with inconsistency and blurry contents.Specifically, we first discover a simple yet essential domain knowledge that contents and motions characteristics should be homogeneous to a certain degree from the same objects, and formulate the consistency into the loss function for model optimization. Based on this, we propose to learn the collaborative representation between motions and contents, and construct a novel progressive spatial-temporal Collaborative network (Prost-Net) for video frame interpolation.Specifically, we develop a content-guided motion module (CGMM) and a motion-guided content module (MGCM) for individual content and motion representation. In particular, the predicted motion in CGMM is used to guide the fusion and distillation of contents for intermediate frame interpolation, and vice versa. Furthermore, by considering collaborative strategy in a multi-scale framework, our Prost-Net progressively optimizes motions and contents in a coarse-to-fine manner, making it robust to various challenging scenarios (occlusion and large motions) in VFI. Extensive experiments on the benchmark datasets demonstrate that our method significantly outperforms state-of-the-art methods. Mengshun Hu, Kui Jiang, Zhixiang Nie, Jing Xiao 0004, Zheng Wang 0007 |
ACM Multimedia | 2 |
| 2022 | You Only Align Once: Bidirectional Interaction for Spatial-Temporal Video Super-ResolutionabstractSpatial-Temporal Video Super-Resolution (ST-VSR) technology generates high-quality videos with higher resolution and higher frame rates. Existing advanced methods accomplish ST-VSR tasks through the association of Spatial and Temporal video super-resolution (S-VSR and T-VSR). These methods require two alignments and fusions in S-VSR and T-VSR, which is obviously redundant and fails to sufficiently explore the information flow of consecutive spatial LR frames. Although bidirectional learning (future-to-past and past-to-future) was introduced to cover all input frames, the direct fusion of final predictions fails to sufficiently exploit intrinsic correlations of bidirectional motion learning and spatial information from all frames. We propose an effective yet efficient recurrent network with bidirectional interaction for ST-VSR, where only one alignment and fusion is needed. Specifically, it first performs backward inference from future to past, and then follows forward inference to super-resolve intermediate frames. The backward and forward inferences are assigned to learn structures and details to simplify the learning task with joint optimizations. Furthermore, a Hybrid Fusion Module (HFM) is designed to aggregate and distill information to refine spatial information and reconstruct high-quality video frames. Extensive experiments on two public datasets demonstrate that our method outperforms state-of-the-art methods in efficiency, and reduces calculation cost by about 22%. Mengshun Hu, Kui Jiang, Zhixiang Nie, Zheng Wang 0007 |
ACM Multimedia | 2 |
| 2022 | Magic ELF: Image Deraining Meets Association Learning and TransformerabstractConvolutional neural network (CNN) and Transformer have achieved great success in multimedia applications. However, little effort has been made to effectively and efficiently harmonize these two architectures to satisfy image deraining. This paper aims to unify these two architectures to take advantage of their learning merits for image deraining. In particular, the local connectivity and translation equivariance of CNN and the global aggregation ability of self-attention (SA) in Transformer are fully exploited for specific local context and global structure representations. Based on the observation that rain distribution reveals the degradation location and degree, we introduce degradation prior to help background recovery and accordingly present the association refinement deraining scheme. A novel multi-input attention module (MAM) is proposed to associate rain perturbation removal and background recovery. Moreover, we equip our model with effective depth-wise separable convolutions to learn the specific feature representations and trade off computational complexity. Extensive experiments show that our proposed method (dubbed as ELF) outperforms the state-of-the-art approach (MPRNet) by 0.25 dB on average, but only accounts for 11.7% and 42.1% of its computational cost and parameters. Kui Jiang, Zhongyuan Wang 0001, Chen Chen 0001, Zheng Wang 0007, Laizhong Cui, Chia-Wen Lin |
ACM Multimedia | 1 |
| 2022 | Multimedia Content Understanding in Harsh EnvironmentsabstractMultimedia content understanding methods often encounter a severe performance degradation under harsh environments. This tutorial covers several important components of multimedia content understanding in harsh environments. It introduces some multimedia enhancement methods, presents recent advances in 2D and 3D visual scene understanding, shows strategies to estimate the prediction uncertainty, provides a brief summary, and shows some typical applications. Zheng Wang 0007, Zhedong Zheng, Kui Jiang |
ACM Multimedia | 4 |
| 2022 | Two-stage unsupervised facial image quality measurement
Guangcheng Wang, Zhongyuan Wang 0001, Baojin Huang, Kui Jiang, Zheng He 0001, Hancheng Zhu, Jinsheng Xiao, Xin Tian 0006 |
Inf. Sci. | 4 |
| 2022 | A Progressive Fusion Generative Adversarial Network for Realistic and Consistent Video Super-ResolutionabstractHow to effectively fuse temporal information from consecutive frames remains to be a non-trivial problem in video super-resolution (SR), since most existing fusion strategies (direct fusion, slow fusion, or 3D convolution) either fail to make full use of temporal information or cost too much calculation. To this end, we propose a novel progressive fusion network for video SR, in which frames are processed in a way of progressive separation and fusion for the thorough utilization of spatio-temporal information. We particularly incorporate multi-scale structure and hybrid convolutions into the network to capture a wide range of dependencies. We further propose a non-local operation to extract long-range spatio-temporal correlations directly, taking place of traditional motion estimation and motion compensation (ME&MC). This design relieves the complicated ME&MC algorithms, but enjoys better performance than various ME&MC schemes. Finally, we improve generative adversarial training for video SR to avoid temporal artifacts such as flickering and ghosting. In particular, we propose a frame variation loss with a single-sequence training method to generate more realistic and temporally consistent videos. Extensive experiments on public datasets show the superiority of our method over state-of-the-art methods in terms of performance and complexity. Our code is available at https://github.com/psychopa4/MSHPFNL. Peng Yi 0002, Zhongyuan Wang 0001, Kui Jiang, Junjun Jiang, Tao Lu 0001, Jiayi Ma 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2022 | Actor-Aware Alignment Network for Action RecognitionabstractAction recognition has attracted growing interest recently. It suffers from the problem that complex and diverse environments may disturb the extraction of action features. Existing methods propose to explore the temporal associations to alleviate the issue. However, they cannot handle long-range frames, and the rigid techniques are powerless against the differences caused by the deformation of the actors. To this end, we propose the Actor-Aware Alignment Network (A$^{3}$Net), which helps locate the action region. Specifically, through the intra-snippet correction, we afford the local segment alignment frames. The inter-snippet is designed to rectify the results, avoiding the occlusion situation that may appear in the local snippet. In addition, we consider intra-alignment short-range adjustive frames and long-range context frames between different snippets, which allows our A$^{3}$Net network to achieve the effect of focusing on long-range frame information. Multiple Reasoning Attention (MRA) modules are introduced to integrate features along the temporal dimension to keep the video spatio-temporal consistent. Extensive experiments conducted on three widely-used public benchmarks,UCF101,HMDB51, andInfAR, indicate that the excellence of our approach over other state-of-the-art models in wild scenarios. Wenxuan Liu 0008, Xian Zhong, Xuemei Jia, Kui Jiang, Chia-Wen Lin |
IEEE Signal Process. Lett. | 4 |
| 2022 | Reference-Free DIBR-Synthesized Video Quality Metric in Spatial and Temporal DomainsabstractDepth image-based rendering (DIBR) techniques play an important role in free viewpoint videos (FVVs), which have a wide range of applications including immersive entertainment, remote monitoring, education, etc. FVVs are usually synthesized by DIBR techniques in a “blind” environment (without a reference video). Thus, an effective reference-free synthesized video quality assessment (VQA) metric is vital. At present, many image quality assessment (IQA) algorithms for DIBR-synthesized images have been proposed, but limited researches have been concerned about the quality assessment of DIBR-synthesized videos. To this end, this paper proposes a novel reference-free VQA method for synthesized videos, which operates in Spatial and Temporal Domains, dubbed as STD. The design fundamental of the proposed STD metric considers the effects of two major distortions introduced by DIBR techniques on the visual quality of synthesized videos. First, considering the geometric distortion introduced by DIBR technologies can increase high-frequency contents of the synthesized frame, the influence of the geometric distortion on the visual quality of a synthesized video can be effectively evaluated by estimating high-frequency energies of each synthesized frame in spatial domain. Second, temporal inconsistency caused by DIBR techniques brings the temporal flicker distortion, which is one of the most annoying artifacts in DIBR-synthesized videos. In temporal domain, we quantify temporal inconsistency by measuring motion differences between consecutive frames. Specifically, optical flow method is first used to estimate the motion field between adjacent frames. Then, we calculate the structural similarity of adjacent optical flow fields and further adopt the structural similarity value to weight the pixel differences of adjacent optical flow fields. Experiments show that the above two features are able to well perceive the visual quality of DIBR-synthesized videos. Furthermore, since the two features are extracted from spatial and temporal domains, respectively, we integrate them using a linear weighting strategy to obtain our STD metric, which proves advantageous over two components and the competing state-of-the-art I/VQA methods. The source code is available athttps://github.com/wgc-vsfm/DIBR-video-quality-assessment. Guangcheng Wang, Zhongyuan Wang 0001, Ke Gu 0001, Kui Jiang, Zheng He 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2022 | Dual-Path Deep Fusion Network for Face Image HallucinationabstractAlong with the performance improvement of deep-learning-based face hallucination methods, various face priors (facial shape, facial landmark heatmaps, or parsing maps) have been used to describe holistic and partial facial features, making the cost of generating super-resolved face images expensive and laborious. To deal with this problem, we present a simple yet effective dual-path deep fusion network (DPDFN) for face image super-resolution (SR) without requiring additional face prior, which learns the global facial shape and local facial components through two individual branches. The proposed DPDFN is composed of three components: a global memory subnetwork (GMN), a local reinforcement subnetwork (LRN), and a fusion and reconstruction module (FRM). In particular, GMN characterize the holistic facial shape by employing recurrent dense residual learning to excavate wide-range context across spatial series. Meanwhile, LRN is committed to learning local facial components, which focuses on the patch-wise mapping relations between low-resolution (LR) and high-resolution (HR) space on local regions rather than the entire image. Furthermore, by aggregating the global and local facial information from the preceding dual-path subnetworks, FRM can generate the corresponding high-quality face image. Experimental results of face hallucination on public face data sets and face recognition on real-world data sets (VGGface and SCFace) show the superiority both on visual effect and objective indicators over the previous state-of-the-art methods. Kui Jiang, Zhongyuan Wang 0001, Peng Yi 0002, Tao Lu 0001, Junjun Jiang, Zixiang Xiong |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2021 | When Face Recognition Meets Occlusion: A New BenchmarkabstractThe existing face recognition datasets usually lack occlusion samples, which hinders the development of face recognition. Especially during the COVID-19 coronavirus epidemic, wearing a mask has become an effective means of preventing the virus spread. Traditional CNN-based face recognition models trained on existing datasets are almost ineffective for heavy occlusion. To this end, we pioneer a simulated occlusion face recognition dataset. In particular, we first collect a variety of glasses and masks as occlusion, and randomly combine the occlusion attributes (occlusion objects, textures,and colors) to achieve a large number of more realistic occlusion types. We then cover them in the proper position of the face image with the normal occlusion habit. Furthermore, we reasonably combine original normal face images and occluded face images to form our final dataset, termed as Webface-OCC. It covers 804,704 face images of 10,575 subjects, with diverse occlusion types to ensure its diversity and stability. Extensive experiments on public datasets show that the ArcFace retrained by our dataset significantly outperforms the state-of-the-arts. Webface-OCC is available at https://github.com/Baojin-Huang/Webface-OCC. Baojin Huang, Zhongyuan Wang 0001, Guangcheng Wang, Kui Jiang, Kangli Zeng, Zhen Han 0002, Xin Tian 0006, Yuhong Yang 0001 |
ICASSP | 4 |
| 2021 | Omniscient Video Super-ResolutionabstractMost recent video super-resolution (SR) methods either adopt an iterative manner to deal with low-resolution (LR) frames from a temporally sliding window, or leverage the previously estimated SR output to help reconstruct the current frame recurrently. A few studies try to combine these two structures to form a hybrid framework but have failed to give full play to it. In this paper, we propose an omniscient framework to not only utilize the preceding SR output, but also leverage the SR outputs from the present and future. The omniscient framework is more generic because the iterative, recurrent and hybrid frameworks can be regarded as its special cases. The proposed omniscient framework enables a generator to behave better than its counterparts under other frameworks. Abundant experiments on public datasets show that our method is superior to the state-of-the-art methods in objective metrics, subjective visual effects and complexity. Peng Yi 0002, Zhongyuan Wang 0001, Kui Jiang, Junjun Jiang, Tao Lu 0001, Xin Tian 0006, Jiayi Ma 0001 |
ICCV | 3 |
| 2021 | PCNET: Progressive Coupled Network for Real-Time Image DerainingabstractImage deraining is an effective solution to avoid performance drop of vision-oriented tasks in rainy weather. Most existing image deraining approaches either fail to produce satisfactory restoration results or cost too much computation. In this paper, we propose a low-complexity and high-performance coupled representation module (CRM), designed to learn the joint features of rain-free contents and rain information as well as their blending correlations. To promote the computation efficiency, we employ depth-wise separable convolutions, and construct CRM in an asymmetric U-shaped architecture to reduce model parameters and memory footprint. Our final model–PCNet achieves the progressive separation of rain-free contents and rain streaks using cascaded residual learning. Extensive experiments are conducted to evaluate the efficacy of the proposed PCNet on several synthetic and real-world rain datasets. Kui Jiang, Zhongyuan Wang 0001, Peng Yi 0002, Chen Chen 0001, Zheng Wang 0007, Chia-Wen Lin |
ICIP | 1 |
| 2021 | Unsupervised Vehicle Search in the Wild: A New BenchmarkabstractIn urban surveillance systems, finding a specific vehicle in video frames efficiently and accurately has always been an essential part of traffic supervision and criminal investigation. Existing studies focus on vehicle re-identification (re-ID), but vehicle search is still underexploited. These methods depend on the locations of many vehicles (bounding boxes) that are not available in most real-world applications. Therefore, the unsupervised joint study of vehicle location and identification for the observed scene is a pressing need. Inspired by person search, we conduct a study on the vehicle search while considering four main discrepancies among them, summarized as: 1) It is challenging to select the candidate regions for the observed vehicle due to the perspective differences (front or side); 2) The sides of the same type of vehicles are almost the same, resulting in smaller inter-class; 3) Lacking satisfied dataset for vehicle search to meet the practical scenarios; 4) Supervised search publishing methods rely on datasets with expensive annotations. To address these issues, we have established a new vehicle search dataset. We design an unsupervised framework on this benchmark dataset to generate pseudo labels for further training existing vehicle re-ID or person search models. Experimental results reveal that these methods turn less effective on vehicle search tasks. Therefore, the vehicle search task needs to be further developed, and this dataset can advance the research of vehicle search. Https://github.com/zsl1997/VSW. Xian Zhong, Xiao Wang 0029, Kui Jiang, Wenxuan Liu 0008, Wenxin Huang, Zheng Wang 0007 |
ACM Multimedia | 4 |
| 2021 | Silicone mask face anti-spoofing detection based on visual saliency and facial motion
Guangcheng Wang, Zhongyuan Wang 0001, Kui Jiang, Baojin Huang, Zheng He 0001, Ruimin Hu |
Neurocomputing | 3 |
| 2021 | Decomposition Makes Better Rain Removal: An Improved Attention-Guided Deraining NetworkabstractRain streaks in the air show diverse characteristics with different shapes, directions, densities, even the complex overlapped phenomenon, causing great challenges for the deraining task. Recently, deep learning based image deraining methods have been extensively investigated due to their excellent performance. However, most of the existing algorithms still have limitations in removing rain streaks while preserving rich textural details under complicated rain conditions. To this end, we propose to decompose rain streaks into multiple rain layers and individually estimate each of them along the network stages to cope with the increasing abstracts. To better characterize rain layers, an improved non-local block is designed to exploit the self-similarity of rain information by learning the holistic spatial feature correlations while reducing the calculation complexity. Moreover, a mixed attention mechanism is applied to guide the fusion of rain layers by focusing on the local and global overlaps among these rain layers. Extensive experiments on both synthetic rainy/rain-haze/raindrop datasets, real-world samples, the haze, and low-light scenarios show substantial improvements both on quantitative indicators and visual effects over the current state-of-the-art technologies. The source code is available athttps://github.com/kuihua/IADN. Kui Jiang, Zhongyuan Wang 0001, Peng Yi 0002, Chen Chen 0001, Zhen Han 0002, Tao Lu 0001, Baojin Huang, Junjun Jiang |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2021 | Rain-Free and Residue Hand-in-Hand: A Progressive Coupled Network for Real-Time Image DerainingabstractRainy weather is a challenge for many vision-oriented tasks (e.g., object detection and segmentation), which causes performance degradation. Image deraining is an effective solution to avoid performance drop of downstream vision tasks. However, most existing deraining methods either fail to produce satisfactory restoration results or cost too much computation. In this work, considering both effectiveness and efficiency of image deraining, we propose a progressive coupled network (PCNet) to well separate rain streaks while preserving rain-free details. To this end, we investigate the blending correlations between them and particularly devise a novel coupled representation module (CRM) to learn the joint features and the blending correlations. By cascading multiple CRMs, PCNet extracts the hierarchical features of multi-scale rain streaks, and separates the rain-free content and rain streaks progressively. To promote computation efficiency, we employ depth-wise separable convolutions and a U-shaped structure, and construct CRM in an asymmetric architecture to reduce model parameters and memory footprint. Extensive experiments are conducted to evaluate the efficacy of the proposed PCNet in two aspects: (1) image deraining on several synthetic and real-world rain datasets and (2) joint image deraining and downstream vision tasks (e.g., object detection and segmentation). Furthermore, we show that the proposed CRM can be easily adopted to similar image restoration tasks including image dehazing and low-light enhancement with competitive performance. The source code is available at https://github.com/kuijiang0802/PCNet. Kui Jiang, Zhongyuan Wang 0001, Peng Yi 0002, Chen Chen 0001, Zheng Wang 0007, Xiao Wang 0029, Junjun Jiang, Chia-Wen Lin |
IEEE Trans. Image Process. | 1 |
| 2020 | Multi-Scale Progressive Fusion Network for Single Image DerainingabstractRain streaks in the air appear in various blurring degrees and resolutions due to different distances from their positions to the camera. Similar rain patterns are visible in a rain image as well as its multi-scale (or multi-resolution) versions, which makes it possible to exploit such complementary information for rain streak representation. In this work, we explore the multi-scale collaborative representation for rain streaks from the perspective of input image scales and hierarchical deep features in a unified framework, termed multi-scale progressive fusion network (MSPFN) for single image rain streak removal. For the similar rain streaks at different positions, we employ recurrent calculation to capture the global texture, thus allowing to explore the complementary and redundant information at the spatial dimension to characterize target rain streaks. Besides, we construct multi-scale pyramid structure, and further introduce the attention mechanism to guide the fine fusion of these correlated information from different scales. This multi-scale progressive fusion strategy not only promotes the cooperative representation, but also boosts the end-to-end training. Our proposed method is extensively evaluated on several benchmark datasets and achieves the state-of-the-art results. Moreover, we conduct experiments on joint deraining, detection, and segmentation tasks, and inspire a new research direction of vision task driven image deraining. The source code is available at https://github.com/kuihua/MSPFN. Kui Jiang, Zhongyuan Wang 0001, Peng Yi 0002, Chen Chen 0001, Baojin Huang, Yimin Luo, Jiayi Ma 0001, Junjun Jiang |
CVPR | 1 |
| 2020 | Attention-Guided Deraining Network Via Stage-Wise LearningabstractDue to diverse rain shapes, directions, densities as well as different distances to cameras, rain streaks in the air are interweaved and overlapped. However, most existing deraining methods are inherently oblivious this phenomenon and tend to learn a single rain streak layer to simulate this complex distribution, consequently failing to restore high-quality rain-free images. To solve this problem, along with the stage-wise learning, we propose a novel attention-guided deraining network (ADN) for rain streak removal. Specially, we decompose the rain streaks into multiple rain streak layers, and individually model them along the stages of the network to match the increasing abstracts. Moreover, the attention mechanism is utilized to guide the fusion of these rain streak layers by handling the overlaps between them. Extensive experiments on several benchmark datasets and real-world scenarios show substantial improvements both on quantitative indicators and visual effects over the current top-performing methods. Kui Jiang, Zhongyuan Wang 0001, Peng Yi 0002, Chen Chen 0001, Yuhong Yang 0001, Xin Tian 0006, Junjun Jiang |
ICASSP | 1 |
| 2020 | Lightweight Progressive Residual Clique Network for Image Super-ResolutionabstractDeeper and wider convolutional neural networks (CNN) hava been widely applied to the single image super-resolution (SR) task for its appealing performance. However, enormous parametric memory footprint hinders its real-time application on mobile devices, especially in the energy-sensitive environment. In this work, we take both the reconstruction performance and efficiency into consideration and propose a lightweight progressive residual clique network (PRCN) for image SR. PRCN is built on the two-stage residual channel separation block (RCSB) and long-skip connections. First, we divide the input into four channel groups to differently learn texture details, immediately followed by a primary fusion to establish cross-channel correspondence in the first stage. Then we perform a further fusion on the outputs of the first stage to constitute a clique for the refinement in the second stage. Meanwhile, we employ SENet to improve the outputs of the second stage with the separate features of the first stage. This design not only enforces the correlation across channels, but also allows fewer densely connected blocks. Experimental results on public datasets show that PRCN outperforms state-of-the-art methods in terms of performance and complexity. Baojin Huang, Zheng He 0001, Zhongyuan Wang 0001, Kui Jiang, Guangcheng Wang |
ICTAI | 4 |
| 2020 | Single image de-raining via clique recursive feedback mechanism
Jun Chen 0001, Kui Jiang, Zhen Han 0002, Weijian Ruan, Zhongyuan Wang 0001, Chao Liang 0001 |
Neurocomputing | 3 |
| 2020 | Ultra-dense GAN for satellite imagery super-resolution
Zhongyuan Wang 0001, Kui Jiang, Peng Yi 0002, Zhen Han 0002, Zheng He 0001 |
Neurocomputing | 2 |
| 2020 | Hierarchical dense recursive network for image super-resolution
Kui Jiang, Zhongyuan Wang 0001, Peng Yi 0002, Junjun Jiang |
Pattern Recognit. | 1 |
| 2020 | Multi-Temporal Ultra Dense Memory Network for Video Super-ResolutionabstractVideo super-resolution (SR) aims to reconstruct the corresponding high-resolution (HR) frames from consecutive low-resolution (LR) frames. It is crucial for video SR to harness both inter-frame temporal correlations and intra-frame spatial correlations among frames. Previous video SR methods based on convolutional neural network (CNN) mostly adopt a single-channel structure and a single memory module, so they are unable to fully exploit inter-frame temporal correlations specific for video. To this end, this paper proposes a multi-temporal ultra-dense memory (MTUDM) network for video super-resolution. Particularly, we embed convolutional long-short-term memory (ConvLSTM) into ultra-dense residual block (UDRB) to construct an ultra-dense memory block (UDMB) for extracting and retaining spatio-temporal correlations. This design also reduces the layer depth by expanding the width, thus avoiding training difficulties, such as gradient exploding and vanishing under a large model. We further adopt multi-temporal information fusion (MTIF) strategy to merge the extracted temporal feature maps in consecutive frames, improving the accuracy without requiring much extra computational cost. The experimental results on extensive public datasets demonstrate that our method outperforms the state-of-the-art methods by a large margin. Peng Yi 0002, Zhongyuan Wang 0001, Kui Jiang, Jiayi Ma 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2020 | ATMFN: Adaptive-Threshold-Based Multi-Model Fusion Network for Compressed Face HallucinationabstractAlthough tremendous strides have been recently made in face hallucination, exiting methods based on a single deep learning framework can hardly satisfactorily provide fine facial features from tiny faces under complex degradation. This article advocates an adaptive-threshold-based multi-model fusion network (ATMFN) for compressed face hallucination, which unifies different deep learning models to take advantages of their respective learning merits. First of all, we construct CNN-, GAN- and RNN-based underlying super-resolvers to produce candidate SR results. Further, the attention subnetwork is proposed to learn the individual fusion weight matrices capturing the most informative components of the candidate SR faces. Particularly, the hyper-parameters of the fusion matrices and the underlying networks are optimized together in an end-to-end manner to drive them for collaborative learning. Finally, a threshold-based fusion and reconstruction module is employed to exploit the candidates' complementarity and thus generate high-quality face images. Extensive experiments on benchmark face datasets and real-world samples show that our model outperforms the state-of-the-art SR methods in terms of quantitative indicators and visual effects. The code and configurations are released at https://github.com/kuihua/ATMFN. Kui Jiang, Zhongyuan Wang 0001, Peng Yi 0002, Guangcheng Wang, Ke Gu 0001, Junjun Jiang |
IEEE Trans. Multim. | 1 |
| 2019 | Progressive Fusion Video Super-Resolution Network via Exploiting Non-Local Spatio-Temporal CorrelationsabstractMost previous fusion strategies either fail to fully utilize temporal information or cost too much time, and how to effectively fuse temporal information from consecutive frames plays an important role in video super-resolution (SR). In this study, we propose a novel progressive fusion network for video SR, which is designed to make better use of spatio-temporal information and is proved to be more efficient and effective than the existing direct fusion, slow fusion or 3D convolution strategies. Under this progressive fusion framework, we further introduce an improved non-local operation to avoid the complex motion estimation and motion compensation (ME&MC) procedures as in previous video SR approaches. Extensive experiments on public datasets demonstrate that our method surpasses state-of-the-art with 0.96 dB in average, and runs about 3 times faster, while requires only about half of the parameters. Peng Yi 0002, Zhongyuan Wang 0001, Kui Jiang, Junjun Jiang, Jiayi Ma 0001 |
ICCV | 3 |
| 2019 | GAN-Based Multi-level Mapping Network for Satellite Imagery Super-ResolutionabstractAlthough many deep-learning-based image super-resolution (SR) methods have been proposed, most of them assume that all hierarchical features share the unified mapping equations. They ignore the differences between mapping equations at different feature levels, and create an average effect of mapping prediction, thus poorly building the mapping relations between low resolution (LR) and high resolution (HR) spaces. In this paper, we propose a multi-level mapping framework along with the adversarial learning strategy, namely MMGAN, for satellite imageries SR reconstruction. We also construct a feature extraction and tuning block (FETB) for fine feature expression. In particular, a novel two-dimension dense unit (DU) and a mapping attention unit (MAU) are constructed for building multi-level mappings in different stages. With our strategies, an HR image is reconstructed directly from the input image using multi-level mappings. Extensive experiments on Kaggle Open Source Dataset and Jilin-1 video satellite images exhibit superior reconstruction performance when compared with the state-of-the-art SR approaches. Kui Jiang, Zhongyuan Wang 0001, Peng Yi 0002, Junjun Jiang, Guangcheng Wang, Zhen Han 0002, Tao Lu 0001 |
ICME | 1 |
| 2019 | Edge-Enhanced GAN for Remote Sensing Image SuperresolutionabstractThe current superresolution (SR) methods based on deep learning have shown remarkable comparative advantages but remain unsatisfactory in recovering the high-frequency edge details of the images in noise-contaminated imaging conditions, e.g., remote sensing satellite imaging. In this paper, we propose a generative adversarial network (GAN)-based edge-enhancement network (EEGAN) for robust satellite image SR reconstruction along with the adversarial learning strategy that is insensitive to noise. In particular, EEGAN consists of two main subnetworks: an ultradense subnetwork (UDSN) and an edge-enhancement subnetwork (EESN). In UDSN, a group of 2-D dense blocks is assembled for feature extraction and to obtain an intermediate high-resolution result that looks sharp but is eroded with artifacts and noises as previous GAN-based methods do. Then, EESN is constructed to extract and enhance the image contours by purifying the noise-contaminated components with mask processing. The recovered intermediate image and enhanced edges can be combined to generate the result that enjoys high credibility and clear contents. Extensive experiments on Kaggle Open Source Data set, Jilin-1 video satellite images, and Digitalglobe show superior reconstruction performance compared to the state-of-the-art SR approaches. Kui Jiang, Zhongyuan Wang 0001, Peng Yi 0002, Guangcheng Wang, Tao Lu 0001, Junjun Jiang |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2019 | Multi-Memory Convolutional Neural Network for Video Super-ResolutionabstractVideo super-resolution (SR) is focused on reconstructing high-resolution (HR) frames from consecutive lowresolution (LR) frames. Most previous video SR methods based on convolutional neural network (CNN) use a direct connection and single-memory module within the network, and they thus fail to make full use of spatio-temporal complementary information from LR observed frames. To fully exploit spatio-temporal correlations between adjacent LR frames and reveal more realistic details, this paper proposes a multi-memory convolutional neural network (MMCNN) for video SR, cascading an optical flow network and an image-reconstruction network. A serial of residual blocks engaged in utilizing intra-frame spatial correlations are proposed for feature extraction and reconstruction. Particularly, instead of using single-memory module, we embed convolutional long short-term memory (ConvLSTM) into the residual block, thus form a multi-memory residual block to progressively extract and retain inter-frame temporal correlations between consecutive LR frames. We conduct extensive experiments on numerous testing datasets with respect to different scaling factors. Our proposed MMCNN shows superiority over the state-of-the-art methods in terms of PSNR and visual quality and surpasses the best counterpart method 1 dB at most. The code and datasets are available at https://github.com/psychopa4/MMCNN. Zhongyuan Wang 0001, Peng Yi 0002, Kui Jiang, Junjun Jiang, Zhen Han 0002, Tao Lu 0001, Jiayi Ma 0001 |
IEEE Trans. Image Process. | 3 |
| 2018 | A Progressively Enhanced Network for Video Satellite Imagery SuperresolutionabstractDeep convolutional neural networks (CNNs) have been extensively applied to image or video processing and analysis tasks. For single-image superresolution (SR) processing, previous CNN-based methods have led to significant improvements, when compared to the shallow learning-based methods. However, these CNN-based algorithms with simply direct or skip connections are not suitable for satellite imagery SR because of complex imaging conditions and unknown degradation process. More importantly, they ignore the extraction and utilization of the structural information in satellite images, which is very unfavorable for video satellite imagery SR with such characteristics as small ground targets, weak textures, and over-compression distortion. To this end, this letter proposes a novel progressively enhanced network for satellite image SR called PECNN, which is composed of a pretraining CNN-based network and an enhanced dense connection network. The pretraining part is used to extract the low-level feature maps and reconstructs a basic high-resolution image from the low-resolution input. In particular, we propose a transition unit to obtain the structural information from the base output. Then, the obtained structural information and the extracted low-level feature maps are transmitted to the enhanced network for further extraction to enforce the feature expression. Finally, a residual image with enhanced fine details obtained from the dense connection network is used to enrich the basic image for the ultimate SR output. Experiments on real-world Jilin-1 video satellite images and Kaggle Open Source Dataset show that the proposed PECNN outperforms the state-of-the-art methods both in visual effects and quantitative metrics. Code is available at https://github.com/kuihua/PECNN. Kui Jiang, Zhongyuan Wang 0001, Peng Yi 0002, Junjun Jiang |
IEEE Signal Process. Lett. | 1 |
| 2015 | A new control strategy of ironless stator axial-flux PM motor fed by inverter with output LC filterabstractAxial-flux ironless stator PM motors are becoming a promising choice for aero applications due to their advantages of light weight, strong overload capacity, high efficiency and high reliability. However, it comes with extremely small inductance resulting in a challenging difficulty to the inverter that drives the motors. Considering the limitation on increasing the switching frequency of high-power device with high DC-link voltage, high-power level and high fundamental frequency, a practical and cost-effective power circuit for the driving operation of a 50kW axial-flux ironless stator PM motors is developed. A new and simple control strategy of the inverter with output LC filter is studied to achieve a good steady and dynamic performance for driving the axial flux ironless stator PM motors. The dual closed-loop of the phase-current of the motor and the branch current of capacitors ensures the stability and torque control of the system. Due to the dual-loop current regulator, the maximum torque per ampere (MTPA) control is achieved. The development of the power circuit topology and effective control method provides an effective solution for ironless stator motors or high-speed motor with small inductance. Weiwei Geng, Zhuoran Zhang 0002, Kui Jiang |
IECON | 3 |