EDBT 2026 Demo / reviewers in the wild / expert
Xianming Liu 0005
dblp:89/5820-5
· DBLP profile ↗
189ranked-venue papers
35as first author
96since 2021 · last 2026
0000-0002-8857-1785ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 144 · 33 first-author · 63 since 2021Artificial intelligence and machine learning · 67 · 4 first-author · 57 since 2021Databases, data management, data science and information retrieval · 9 · 4 first-author · 1 since 2021Systems, architecture and hardware · 6 · 1 first-author · 3 since 2021Security and privacy · 4 · 1 since 2021Computer networks · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Variation-Bounded Loss for Noise-Tolerant LearningabstractMitigating the negative impact of noisy labels has been a perennial issue in supervised learning. Robust loss functions have emerged as a prevalent solution to this problem. In this work, we introduce the Variation Ratio as a novel property related to the robustness of loss functions, and propose a new family of robust loss functions, termed Variation-Bounded Loss (VBL), which is characterized by a bounded variation ratio. We provide theoretical analyses of the variation radio, proving that a smaller variation ratio would lead to better robustness. Furthermore, we reveal that the variation ratio provides a feasible method to relax the symmetric condition and offers a more concise path to achieve the asymmetric condition. Based on the variation ratio, we reformulate several commonly used loss functions into a variation-bounded form for pract ical applications. Positive experiments on various datasets exhibit the effectiveness and flexibility of our approach. Jialiang Wang 0003, Xianming Liu 0005, Gangfeng Hu, Deming Zhai, Junjun Jiang, Haoliang Li |
AAAI | 3 |
| 2026 | Semantics and Content Matter: Towards Multi-Prior Hierarchical Mamba for Image DerainingabstractRain significantly degrades the performance of computer vision systems, particularly in applications like autonomous driving and video surveillance. While existing deraining methods have made considerable progress, they often struggle with fidelity of semantic and spatial details. To address these limitations, we propose the Multi-Prior Hierarchical Mamba (MPHM) network for image deraining. This novel architecture synergistically integrates macro-semantic textual priors (CLIP) for task-level semantic guidance and micro-structural visual priors (DINOv2) for scene-aware structural information. To alleviate potential conflicts between heterogeneous priors, we devise a progressive Priors Fusion Injection (PFI) that strategically injects complementary cues at different decoder levels. Meanwhile, we equip the backbone network with an elaborate Hierarchical Mamba Module (HMM) to facilitate robust feature representation, featuring a Fourier-enhanced dual-path design that concurrently addresses global context modeling and local detail recovery. Comprehensive experiments demonstrate MPHM's state-of-the-art performance, achieving a 0.57 dB PSNR gain on the Rain200H dataset while delivering superior generalization on real-world rainy scenarios. Zhaocheng Yu, Kui Jiang, Junjun Jiang, Xianming Liu 0005, Guanglu Sun, Yi Xiao 0003 |
AAAI | 4 |
| 2026 | Learning from History: Task-agnostic Model Contrastive Learning for Image Restoration
Gang Wu 0010, Junjun Jiang, Kui Jiang, Xianming Liu 0005, Wangmeng Zuo |
Int. J. Comput. Vis. | 4 |
| 2026 | A Natural Language Guided Approach for Blind Face Restoration: Methodology and DatasetabstractBlind Face Restoration (BFR) aims to reconstruct high-quality face images from low-quality inputs without any prior knowledge of the specific degradation types or levels. In recent years, remarkable progress has been achieved, particularly through GAN- and diffusion-based approaches, which have greatly improved perceptual realism and reconstruction fidelity. However, existing approaches typically rely solely on visual cues from degraded images. This often results in inaccurate reconstruction of facial details and noticeable identity distortion, particularly under severe or complex degradations. To address these limitations, we incorporate auxiliary textual information into BFR to enable the recovery of subtle facial attributes, such as wrinkles, moles, and skin marks that are often overlooked or hard to reconstruct by conventional visual priors. To support this idea, we first construct a large-scale dataset containing 30,000 detailed textual descriptions paired with CelebA-HQ face images, explicitly designed to capture fine-grained facial semantics. To effectively bridge the gap between visual data and natural language, we further propose FaceCLIP, a fine-tuned vision-language model specifically tailored to the human face. FaceCLIP enables more accurate alignment between face images and their corresponding textual descriptions by effectively capturing nuanced semantic cues critical for faithful face reconstruction. Built upon these foundations, we propose Text-guided Blind Face Restoration (TBFR), a novel diffusion-based framework that explicitly integrates textual guidance into the face restoration pipeline. Within TBFR, a text-guided hybrid attention block is designed to effectively fuse visual and textual features, while a text-aware loss is employed to enforce semantic consistency between the generated images and their associated textual descriptions. Extensive experimental results show that TBFR outperforms state-of-the-art BFR methods in terms of both quantitative metrics and subjective perceptual quality, establishing a new benchmark for BFR tasks. Wenjie An, Chenyang Wang 0002, Junjun Jiang, Kui Jiang, Xianming Liu 0005, Liqiang Nie |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2026 | Beyond Degradation Redundancy: Contrastive Prompt Learning for All-in-One Image RestorationabstractAll-in-one image restoration, addressing diverse degradation types with a unified model, presents significant challenges in designing task-aware prompts that effectively guide restoration across multiple degradation scenarios. While adaptive prompt learning enables end-to-end optimization, it often yields overlapping or redundant task representations. Conversely, explicit prompts derived from pretrained classifiers enhance discriminability but may discard critical visual information for reconstruction. To address these limitations, we introduce Contrastive Prompt Learning (CPL), a novel framework that fundamentally enhances prompt-task alignment through two complementary innovations: a Sparse Prompt Module (SPM) that efficiently captures degradation-specific features while minimizing redundancy, and a Contrastive Prompt Regularization (CPR) that explicitly strengthens task boundaries by incorporating negative prompt samples across different degradation types. Unlike previous approaches that focus primarily on degradation classification, CPL optimizes the critical interaction between prompts and the restoration model itself. Extensive experiments across comprehensive benchmarks demonstrate that CPL consistently enhances state-of-the-art all-in-one restoration models, achieving significant improvements in both standard multi-task scenarios and challenging composite degradation settings. Our framework establishes new state-of-the-art performance while maintaining parameter efficiency, offering a principled solution for unified image restoration. The code is available at https://github.com/Aitical/CPLIR. Gang Wu 0010, Junjun Jiang, Kui Jiang, Xianming Liu 0005, Liqiang Nie |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2026 | DSwinIR: Rethinking Window-Based Attention for Image RestorationabstractImage restoration has witnessed significant advancements with the development of deep learning models. Transformer-based models, particularly those using window-based self-attention, have become a dominant force. However, their performance is constrained by the rigid, non-overlapping window partitioning scheme, which leads to insufficient feature interaction across windows and limited receptive fields. This highlights the need for more adaptive and flexible attention mechanisms. In this paper, we propose the Deformable Sliding Window Transformer for Image Restoration (DSwinIR), a new attention mechanism: the Deformable Sliding Window (DSwin) Attention. This mechanism introduces a token-centric and content-aware paradigm that moves beyond the grid and fixed window partition. It comprises two complementary components. First, it replaces the rigid partitioning with a token-centric sliding window paradigm, making it effective at eliminating boundary artifacts. Second, it incorporates a content-aware deformable sampling strategy, which allows the attention mechanism to learn data-dependent offsets and actively shape its receptive field to focus on the most informative image regions. Extensive experiments show that DSwinIR achieves strong results, including state-of-the-art performance on several evaluated benchmarks. For instance, in all-in-one image restoration, our DSwinIR surpasses the most recent backbone GridFormer by 0.53 dB on the three-task benchmark and 0.87 dB on the five-task benchmark. Gang Wu 0010, Junjun Jiang, Kui Jiang, Xianming Liu 0005, Liqiang Nie |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2026 | Federated domain generalization via data-centric flatness optimization
Chenyang Wang 0002, Junjun Jiang, Xianming Liu 0005, Xiangyang Ji |
Pattern Recognit. | 4 |
| 2026 | SGCNeRF: Few-Shot Neural Rendering via Sparse Geometric Consistency GuidanceabstractNeural Radiance Field (NeRF) technology has made significant strides in creating novel viewpoints. However, its effectiveness is hampered when working with sparsely available views, often leading to performance dips due to overfitting. FreeNeRF attempts to overcome this limitation by integrating implicit geometry regularization, which incrementally improves both geometry and textures. Nonetheless, an initial low positional encoding bandwidth results in the exclusion of high-frequency elements. The quest for a holistic approach that simultaneously addresses overfitting and the preservation of high-frequency details remains ongoing. This study presents a novel feature-matching-based sparse geometry regularization module, enhanced by a spatially consistent geometry filtering mechanism and a frequency-guided geometric regularization strategy. This module excels at accurately identifying high-frequency keypoints, effectively preserving fine structural details. Through progressive refinement of geometry and textures across NeRF iterations, we unveil an effective few-shot neural rendering architecture, designated as SGCNeRF, for enhanced novel view synthesis. Our experiments demonstrate that SGCNeRF not only achieves superior geometry-consistent outcomes but also surpasses FreeNeRF, with improvements of 0.7 dB in PSNR on LLFF and DTU. Yuru Xiao, Xianming Liu 0005, Deming Zhai, Kui Jiang, Junjun Jiang, Xiangyang Ji |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2026 | Practical Lossless Volumetric Medical Image Compression via Tri-Plane Context Tree LearningabstractLossless compression of volumetric medical images is of paramount importance for clinical and research applications where data fidelity is essential. Traditional compression methods are often limited in efficiency due to rigid, handcrafted models. Conversely, deep neural network (DNN)-based compression methods, while effective, demand substantial computational resources, hindering deployment in resource-constrained settings. To address these challenges, we propose a novel tri-plane context tree (TCT)-based method for lossless volumetric medical image compression that delivers high performance without relying on DNNs or external training data. To exploit intra-slice and inter-slice redundancies, we introduce a compact tri-plane context representation that decomposes complex 3D context modeling into efficient 2D modeling on three orthogonal planes. By integrating this representation with a context tree framework, we develop an input-specific TCT model employing an adaptive binary tree structure. At each tree node, the model dynamically selects from a suite of tri-plane based predictors and contextual feature extractors, enabling data-adaptive context modeling tailored to local structural characteristics. Instead of offline training, we sample a subset of the input volume to learn the TCT model by optimizing the minimum description length (MDL) through iterative construction and pruning. With the learned TCT model, each pixel retrieves its corresponding context, computes the prediction residual using the predictor dictated by the context, and performs entropy encoding based on the associated histograms. Experimental results demonstrate that the proposed method achieves compression performance on par with recent DNN-based methods on multiple datasets, while maintaining low computational cost and fast coding speeds, making it highly applicable in practice. Yuanchao Bai, Kai Wang 0070, Yuanbo Du, Jie Chen 0001, Teng Fang, Xianming Liu 0005, Wen Gao 0001 |
IEEE Trans. Image Process. | 7 |
| 2026 | PH-Mamba: Enhancing Mamba With Position Encoding and Harmonized Attention for Image Deraining and BeyondabstractMamba and its variants excel at modeling long-range dependencies with linear computational complexity, making them effective for diverse vision tasks. However, Mamba's reliance on unfolding 1D sequential representations necessitates multiple directional scans to recover lost spatial dependencies. This introduces significant computational overhead, redundant token traversal, and inefficiencies that compromise accuracy in real-world applications. To this end, we propose PH-Mamba, a novel framework integrating position encoding and harmonized attention for image deraining and beyond. PH-Mamba transforms Mamba's scanning process into a position-guided, unidirectional scanning that selectively prioritizes degradation-relevant tokens. Specifically, we devise a position-guided hybrid Mamba module (PHMM) that jointly encodes perturbation features alongside their spatial coordinates and harmonized representation to model consistent degradation patterns. Within PHMM, a harmonized Transformer is developed to focus on uncertain regions while suppressing noise interference, thereby improving spatial modeling fidelity. Additionally, we employ a vector decomposition and synthesis strategy to enable the unified representation layout to global degradation by directional scanning while minimizing redundancy. By cascading multiple PHMM blocks, PH-Mamba combines global positional guidance with local differential features to strengthen contextual learning. Extensive experiments demonstrate the superiority of PH-Mamba across low-level image restoration benchmarks. For example, compared to NeRD, PH-Mamba achieves a 0.60 dB PSNR improvement while requiring 88.9% fewer parameters, 36.2% less computation, and 63.0% faster inference time. Kui Jiang, Junjun Jiang, Xianming Liu 0005, Hongxun Yao, Chia-Wen Lin |
IEEE Trans. Image Process. | 3 |
| 2026 | 3D-SLARM: Practical Lossless Volumetric Image Compression via a 3D-Scanning Lightweight Autoregressive ModelabstractVolumetric images often encapsulate critical information, making it essential to employ lossless compression to preserve data integrity. Although various learned methods have demonstrated effective lossless compression for volumetric images, balancing high compression ratios with rapid coding speeds and lightweight architectures remains challenging. In this paper, we propose a 3D-scanning lightweight autoregressive model (3D-SLARM) for practical lossless volumetric image compression. 3D-SLARM integrates a novel 3D plane scanning module, a lightweight feature extraction (FE) module, and a lightweight distribution parameter and adaptive range predictor (DPARP) module. Initially, 3D-SLARM leverages a 3D plane scanning module to determine the scanning order of each voxel, allowing parallel coding of voxels within the same plane. Next, the lightweight FE module captures both intra-slice and inter-slice dependencies in the receptive field defined by the 3D plane scanning module. By incorporating our proposed serial re-parameterization (SerRep) technology alongside non-centric masked convolution (NCMC), the FE module attains a lightweight design while effectively capturing complex dependencies. Finally, 3D-SLARM employs a lightweight DPARP module to compute distribution parameters for both 8-bit and high bit-depth volumetric images. For high bit-depth images, the module further generates an adaptive probability range for each voxel, resulting in compact, voxel-specific PMF tables that facilitate efficient compression. Extensive experiments demonstrate that our 3D-SLARM achieves state-of-the-art lossless compression performance on majority volumetric image datasets and maintains fast coding speed with a lightweight design, underscoring its practical applicability. Kai Wang 0070, Yuanchao Bai, Daxin Li, Deming Zhai, Junjun Jiang, Xianming Liu 0005 |
IEEE Trans. Image Process. | 6 |
| 2025 | CALLIC: Content Adaptive Learning for Lossless Image CompressionabstractLearned lossless image compression has achieved significant advancements in recent years. However, existing methods often rely on training amortized generative models on massive datasets, resulting in sub-optimal probability distribution estimation for specific testing images during encoding process. To address this challenge, we explore the connection between the Minimum Description Length (MDL) principle and Parameter-Efficient Transfer Learning (PETL), leading to the development of a novel content-adaptive approach for learned lossless image compression, dubbed CALLIC. Specifically, we first propose a content-aware autoregressive self-attention mechanism by leveraging convolutional gating operations, termed Masked Gated ConvFormer (MGCF), and pretrain MGCF on training dataset. Cache then Crop Inference (CCI) is proposed to accelerate the coding process. During encoding, we decompose pretrained layers, including depth-wise convolutions, using low-rank matrices and then adapt the incremental weights on testing image by Rate-guided Progressive Fine-Tuning (RPFT). RPFT fine-tunes with gradually increasing patches that are sorted in descending order by estimated entropy, optimizing learning process and reducing adaptation time. Extensive experiments across diverse datasets demonstrate that CALLIC sets a new state-of-the-art (SOTA) for learned lossless image compression. Daxin Li, Yuanchao Bai, Kai Wang 0070, Junjun Jiang, Xianming Liu 0005, Wen Gao 0001 |
AAAI | 5 |
| 2025 | Debiased All-in-one Image Restoration with Task Uncertainty RegularizationabstractAll-in-one image restoration is a fundamental low-level vision task with significant real-world applications. The primary challenge lies in addressing diverse degradations within a single model. While current methods primarily exploit task prior information to guide the restoration models, they typically employ uniform multi-task learning, overlooking the heterogeneity in model optimization across different degradation tasks. To eliminate the bias, we propose a task-aware optimization strategy, that introduces adaptive task-specific regularization for multi-task image restoration learning. Specifically, our method dynamically weights and balances losses for different restoration tasks during training, encouraging the implementation of the most reasonable optimization route. In this way, we can achieve more robust and effective model training. Notably, our approach can serve as a plug-and-play strategy to enhance existing models without requiring modifications during inference. Extensive experiments in diverse all-in-one restoration settings demonstrate the superiority and generalization of our approach. For example, AirNet retrained with TUR achieves average improvements of 1.16 dB on three distinct tasks and 1.81 dB on five distinct all-in-one tasks. These results underscore TUR's effectiveness in advancing the SOTAs in all-in-one image restoration, paving the way for more robust and versatile image restoration. Gang Wu 0010, Junjun Jiang, Kui Jiang, Xianming Liu 0005 |
AAAI | 5 |
| 2025 | Spatial Annealing for Efficient Few-shot Neural RenderingabstractNeural Radiance Fields (NeRF) with hybrid representations have shown impressive capabilities for novel view synthesis, delivering high efficiency. Nonetheless, their performance significantly drops with sparse input views. Various regularization strategies have been devised to address these challenges. However, these strategies either require additional rendering costs or involve complex pipeline designs, leading to a loss of training efficiency. Although FreeNeRF has introduced an efficient frequency annealing strategy, its operation on frequency positional encoding is incompatible with the efficient hybrid representations. In this paper, we introduce an accurate and efficient few-shot neural rendering method named Spatial Annealing regularized NeRF (SANeRF), which adopts the pre-filtering design of a hybrid representation. We initially establish the analytical formulation of the frequency band limit for a hybrid architecture by deducing its filtering process. Based on this analysis, we propose a universal form of frequency annealing in the spatial domain, which can be implemented by modulating the sampling kernel to exponentially shrink from an initial one with a narrow grid tangent kernel spectrum. This methodology is crucial for stabilizing the early stages of the training phase and significantly contributes to enhancing the subsequent process of detail refinement. Our extensive experiments reveal that, by adding merely one line of code, SANeRF delivers superior rendering quality and much faster reconstruction speed compared to current few-shot neural rendering methods. Notably, SANeRF outperforms FreeNeRF on the Blender dataset, achieving 700X faster reconstruction speed. Yuru Xiao, Deming Zhai, Wenbo Zhao 0004, Kui Jiang, Junjun Jiang, Xianming Liu 0005 |
AAAI | 6 |
| 2025 | DashGaussian: Optimizing 3D Gaussian Splatting in 200 Secondsabstract3D Gaussian Splatting (3DGS) renders pixels by rasterizing Gaussian primitives, where the rendering resolution and the primitive number, concluded as the optimization complexity, dominate the time cost in primitive optimization. In this paper, we propose DashGaussian, a scheduling scheme over the optimization complexity of 3DGS that strips redundant complexity to accelerate 3DGS optimization. Specifically, we formulate 3DGS optimization as progressively fitting 3DGS to higher levels of frequency components in the training views, and propose a dynamic rendering resolution scheme that largely reduces the optimization complexity based on this formulation. Besides, we argue that a specific rendering resolution should cooperate with a proper primitive number for a better balance between computing redundancy and fitting quality, where we schedule the growth of the primitives to synchronize with the rendering resolution. Extensive experiments show that our method accelerates the optimization of various 3DGS backbones by 45.7% on average while preserving the rendering quality. Project page is available at dashgaussian.github.io. Youyu Chen, Junjun Jiang, Kui Jiang, Xianming Liu 0005, Yinyu Nie |
CVPR | 6 |
| 2025 | COB-GS: Clear Object Boundaries in 3DGS Segmentation Based on Boundary-Adaptive Gaussian SplittingabstractAccurate object segmentation is crucial for high-quality scene understanding in the 3D vision domain. However, 3D segmentation based on 3D Gaussian Splatting (3DGS) struggles with accurately delineating object boundaries, as Gaussian primitives often span across object edges due to their inherent volume and the lack of semantic guidance during training. In order to tackle these challenges, we introduce Clear Object Boundaries for 3DGS Segmentation (COB-GS), which aims to improve segmentation accuracy by clearly delineating blurry boundaries of interwoven Gaussian primitives within the scene. Unlike existing approaches that remove ambiguous Gaussians and sacrifice visual quality, COB-GS, as a 3DGS refinement method, jointly optimizes semantic and visual information, allowing the two different levels to cooperate with each other effectively. Specifically, for the semantic guidance, we introduce a boundary-adaptive Gaussian splitting technique that leverages semantic gradient statistics to identify and split ambiguous Gaussians, aligning them closely with object boundaries. For the visual optimization, we rectify the degraded suboptimal texture of the 3DGS scene, particularly along the refined boundary structures. Experimental results show that COB-GS substantially improves segmentation accuracy and robustness against inaccurate masks from pre-trained model, yielding clear boundaries while preserving high visual quality. Code is available at https://github.com/ZestfulJX/COB-GS. Junjun Jiang, Youyu Chen, Kui Jiang, Xianming Liu 0005 |
CVPR | 5 |
| 2025 | Balancing Task-Invariant Interaction and Task-Specific Adaptation for Unified Image FusionabstractUnified image fusion aims to integrate complementary information from multi-source images, enhancing image quality through a unified framework applicable to diverse fusion tasks. While treating all fusion tasks as a unified problem facilitates task-invariant knowledge sharing, it often overlooks task-specific characteristics, thereby limiting the overall performance. Existing general image fusion methods incorporate explicit task identification to enable adaptation to different fusion tasks. However, this dependence during inference restricts the model's generalization to unseen fusion tasks. To address these issues, we propose a novel unified image fusion framework named "TITA", which dynamically balances both Task-invariant Interaction and Task-specific Adaptation. For task-invariant interaction, we introduce the Interaction-enhanced Pixel Attention (IPA) module to enhance pixel-wise interactions for better multi-source complementary information extraction. For task-specific adaptation, the Operation-based Adaptive Fusion (OAF) module dynamically adjusts operation weights based on task properties. Additionally, we incorporate the Fast Adaptive Multitask Optimization (FAMO) strategy to mitigate the impact of gradient conflicts across tasks during joint training. Extensive experiments demonstrate that TITA not only achieves competitive performance compared to specialized methods across three image fusion scenarios but also exhibits strong generalization to unseen fusion tasks. The source codes are released at https://github.com/huxingyuabc/TITA. Junjun Jiang, Chenyang Wang 0002, Kui Jiang, Xianming Liu 0005, Jiayi Ma 0001 |
ICCV | 5 |
| 2025 | Joint Asymmetric Loss for Learning with Noisy LabelsabstractLearning with noisy labels is a crucial task for training accurate deep neural networks. To mitigate label noise, prior studies have proposed various robust loss functions, particularly symmetric losses. Nevertheless, symmetric losses usually suffer from the underfitting issue due to the overly strict constraint. To address this problem, the Active Passive Loss (APL) jointly optimizes an active and a passive loss to mutually enhance the overall fitting ability. Within APL, symmetric losses have been successfully extended, yielding advanced robust loss functions. Despite these advancements, emerging theoretical analyses indicate that asymmetric losses, a new class of robust loss functions, possess superior properties compared to symmetric losses. However, existing asymmetric losses are not compatible with advanced optimization frameworks such as APL, limiting their potential and applicability. Motivated by this theoretical gap and the prospect of asymmetric losses, we extend the asymmetric loss to the more complex passive loss scenario and propose the Asymetric Mean Square Error (AMSE), a novel asymmetric loss. We rigorously establish the necessary and sufficient condition under which AMSE satisfies the asymmetric condition. By substituting the traditional symmetric passive loss in APL with our proposed AMSE, we introduce a novel robust loss framework termed Joint Asymmetric Loss (JAL). Extensive experiments demonstrate the effectiveness of our method in mitigating label noise. Code available at: https://github.com/cswjl/joint-asymmetric-loss Jialiang Wang 0003, Xianming Liu 0005, Gangfeng Hu, Deming Zhai, Junjun Jiang, Xiangyang Ji |
ICCV | 2 |
| 2025 | Unveiling the Depths: A Multi-Modal Fusion Framework for Challenging ScenariosabstractMonocular depth estimation from RGB images plays a pivotal role in 3D vision. However, its accuracy can deteriorate in challenging environments such as nighttime or adverse weather conditions. While long-wave infrared cameras offer stable imaging in such challenging conditions, they are inherently low-resolution, lacking rich texture and semantics as delivered by the RGB image. Current methods focus solely on a single modality due to the difficulties to identify and integrate faithful depth cues from both sources. To address these issues, this paper presents a novel approach that identifies and integrates dominant cross-modality depth features with a learning-based framework. Concretely, we independently compute the coarse depth maps with separate networks by fully utilizing the individual depth cues from each modality. As the advantageous depth spreads across both modalities, we propose a novel confidence loss steering a confidence predictor network to yield a confidence map specifying latent potential depth areas. With the resulting confidence map, we propose a multi-modal fusion network that fuses the final depth in an end-to-end manner. Harnessing the proposed pipeline, our method demonstrates the ability of robust depth estimation in a variety of difficult scenarios. Experimental results on the challenging$\text{MS}^{2}$and ViViD++ datasets demonstrate the effectiveness and robustness of our method. Jialei Xu, Rui Li 0013, Junjun Jiang, Xianming Liu 0005 |
ICRA | 5 |
| 2025 | FUSE: Label-Free Image-Event Joint Monocular Depth Estimation via Frequency-Decoupled Alignment and Degradation-Robust FusionabstractImage-event joint depth estimation methods leverage complementary modalities for robust perception, yet face challenges in generalizability stemming from two factors: 1) limited annotated image-event-depth datasets causing insufficient cross-modal supervision, and 2) inherent frequency mismatches between static images and dynamic event streams with distinct spatiotemporal patterns, leading to ineffective feature fusion. To address this dual challenge, we propose Frequency-decoupled Unified Self-supervised Encoder (FUSE) with two synergistic components: The Parameter-efficient Self-supervised Transfer (PST) leverages image foundation models for cross-modal knowledge transfer, effectively mitigating data scarcity by enabling joint encoding without depth ground truth. Complementing this, the Frequency-Decoupled Fusion module (FreDFuse) resolves modality-specific frequency mismatches by decoupling features into high- and low-frequency bands and then performing a guided cross-attention fusion, where the modality dominant in each band steers the integration. This combined approach enables FUSE to construct a universal image-event encoder that only requires lightweight decoder adaptation for target datasets. Extensive experiments demonstrate state-of-the-art performance with 14% and 24.9% improvements in Abs.Rel on MVSEC and DENSE datasets. The framework exhibits remarkable robustness and generalization in challenging scenarios, including extreme lighting and motion blur, significantly advancing its real-world deployment capabilities. The source code for our method is publicly available at: https://github.com/sunpihai-up/FUSE. Pihai Sun, Junjun Jiang, Yuanqi Yao, Youyu Chen, Wenbo Zhao 0004, Kui Jiang, Xianming Liu 0005 |
IROS | 7 |
| 2025 | Reframing Gaussian Splatting Densification with Complexity-Density Consistency of PrimitivesabstractThe essence of 3D Gaussian Splatting (3DGS) training is to smartly allocate Gaussian primitives, expressing complex regions with more primitives and vice versa.
Prior researches typically mark out under-reconstructed regions in a rendering-loss-driven manner.
However, such a loss-driven strategy is often dominated by low-frequency regions, which leads to insufficient modeling of high-frequency details in texture-rich regions. As a result, it yields a suboptimal spatial allocation of Gaussian primitives.
This inspires us to excavate the loss-agnostic visual prior in training views to identify complex regions that need more primitives to model.
Based on this insight, we propose Complexity-Density Consistent Gaussian Splatting (CDC-GS), which allocates primitives based on the consistency between visual complexity of training views and the density of primitives.
Specifically, primitives involved in rendering high visual complexity areas are categorized as modeling high complexity regions, where we leverage the high frequency wavelet components of training views to measure the visual complexity.
And the density of a primitive is computed with the inverse of geometric mean of its distance to the neighboring primitives.
Guided by the positive correlation between primitive complexity and density, we determine primitives to be densified as well as pruned.
Extensive experiments demonstrate that our CDC-GS surpasses the baseline methods in rendering quality by a large margin using the same amount of Gaussians.
And we provide insightful analysis to reveal that our method serves perpendicularly to rendering loss in guiding Gaussian primitive allocation. Zhemeng Dong, Junjun Jiang, Youyu Chen, Kui Jiang, Xianming Liu 0005 |
NeurIPS | 6 |
| 2025 | A Wavelet-based Image Coding Framework for Data Storage on DNAabstractIn the face of the exponential growth of digital data, DNA is expected to become a new storage medium. Image data makes up a large proportion of digital data. However, existing DNA data storage models are mainly designed for general files. To address this issue, we propose a novel image encoding method for DNA data storage. We employ discrete wavelet transform to decompose the image and utilize an improved exponent-mantissa representation for numerical data. Subsequently, we achieve enhanced compression performance through context-adaptive arithmetic coding. Additionally, we construct a dictionary between ternary sequences and oligonucleotides to generate nucleotide sequences that meet the specified constraints. Experiments show that our method outperforms JPEG-DNA and BioCoder in compression performance and generates higher-quality nucleotide sequences. Chen Qin, Yuanchao Bai, Wenbo Zhao 0004, Xianming Liu 0005 |
VCIP | 4 |
| 2025 | Enhancing consistency and mitigating bias: A data replay approach for incremental learning
Chenyang Wang 0002, Junjun Jiang, Xianming Liu 0005, Xiangyang Ji |
Neural Networks | 4 |
| 2025 | A Survey on All-in-One Image Restoration: Taxonomy, Evaluation and Future TrendsabstractImage restoration (IR) seeks to recover high-quality images from degraded observations caused by a wide range of factors, including noise, blur, compression, and adverse weather. While traditional IR methods have made notable progress by targeting individual degradation types, their specialization often comes at the cost of generalization, leaving them ill-equipped to handle the multifaceted distortions encountered in real-world applications. In response to this challenge, the all-in-one image restoration (AiOIR) paradigm has recently emerged, offering a unified framework that adeptly addresses multiple degradation types. These innovative models enhance the convenience and versatility by adaptively learning degradation-specific features while simultaneously leveraging shared knowledge across diverse corruptions. In this survey, we provide the first in-depth and systematic overview of AiOIR, delivering a structured taxonomy that categorizes existing methods by architectural designs, learning paradigms, and their core innovations. We systematically categorize current approaches and assess the challenges these models encounter, outlining research directions to propel this rapidly evolving field. To facilitate the evaluation of existing methods, we also consolidate widely-used datasets, evaluation protocols, and implementation practices, and compare and summarize the most advanced open-source models. As the first comprehensive review dedicated to AiOIR, this paper aims to map the conceptual landscape, synthesize prevailing techniques, and ignite further exploration toward more intelligent, unified, and adaptable visual restoration systems. Junjun Jiang, Zengyuan Zuo, Gang Wu 0010, Kui Jiang, Xianming Liu 0005 |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2025 | Pixel2Pixel: A Pixelwise Approach for Zero-Shot Single Image DenoisingabstractWe propose Pixel2Pixel, a novel zero-shot image denoising framework that leverages the non-local self-similarity of images to generate a large number of training samples using only the input noisy image. This framework employs a compact convolutional neural network architecture to achieve high-quality image denoising. Given a single observed noisy image, we first aim to obtain multiple images with different noise versions. We ensure that the content remains as consistent as possible with the true signal of the noisy image while keeping the noise independent. Specifically, we construct a pixel bank tensor, where each pixel consists of the most similar pixels from the non-local region of the noisy image. Then, multiple training samples, also known as pseudo instances, can be derived from the pixel bank by randomly pixel sampling. By harnessing pixel-wise random sampling, Pixel2Pixel generates a large number of training pseudo instances, thus avoiding reliance on specific training data. In addition, this non-local pixel selection and random sampling strategy helps to break down the spatial correlation of real-world noise as well. Since the proposed method does not require accurate priors on the noise distribution and clean training images, it is suitable for a wide range of noise types and different noise levels, exhibiting strong generalization ability, especially in real noisy scenes. Extensive experiments across various noise types show that Pixel2Pixel outperforms existing methods. Junjun Jiang, Pengwei Liang, Xianming Liu 0005, Jiayi Ma 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2025 | Dual-Level Cross-Modality Neural Architecture Search for Guided Image Super-ResolutionabstractGuided image super-resolution (GISR) aims to reconstruct a high-resolution (HR) target image from its low-resolution (LR) counterpart with the guidance of a HR image from another modality. Existing learning-based methods typically employ symmetric two-stream networks to extract features from both the guidance and target images, and then fuse these features at either an early or late stage through manually designed modules to facilitate joint inference. Despite significant performance, these methods still face several issues: i) the symmetric architectures treat images from different modalities equally, which may overlook the inherent differences between them; ii) lower-level features contain detailed information while higher-level features capture semantic structures. However, determining which layers should be fused and which fusion operations should be selected remain unresolved; iii) most methods achieve performance gains at the cost of increased computational complexity, so balancing the trade-off between computational complexity and model performance remains a critical issue. To address these issues, we propose a Dual-level Cross-modality Neural Architecture Search (DCNAS) framework to automatically design efficient GISR models. Specifically, we propose a dual-level search space that enables the NAS algorithm to identify effective architectures and optimal fusion strategies. Moreover, we propose a supernet training strategy that employs a pairwise ranking loss trained performance predictor to guide the supernet training process. To the best of our knowledge, this is the first attempt to introduce the NAS algorithm into GISR tasks. Extensive experiments demonstrate that the discovered model family, DCNAS-Tiny and DCNAS, achieve significant improvements on several GISR tasks, including guided depth map super-resolution, guided saliency map super-resolution, guided thermal image super-resolution, and pan-sharpening. Furthermore, we analyze the architectures searched by our method and provide some new insights for future research. Zhiwei Zhong 0001, Xianming Liu 0005, Junjun Jiang, Debin Zhao, Shiqi Wang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2025 | Zero6DOT: Zero-Shot 6D Object Pose Tracking With Monocular RGB Videoabstract6D object tracking plays an important role in various applications, including robotic manipulation and virtual reality. While current methodologies have achieved significant advancements through the use of CAD models, multi-modal sensor data, and category-level assumptions, such resources are often inaccessible in open-world scenarios. Consequently, tracking 6D object poses using only RGB data in such scenarios remains a challenging task. In this paper, we introduce Zero6DOT, an innovative and efficient method for real-time tracking of unknown 6D object poses in monocular RGB video sequences at 8Hz. Our approach requires only the mask of the initial frame, eliminating the need for additional data. The core of Zero6DOT lies in its ability to establish high-quality correspondences across images, from which accurate poses are derived. To achieve this, we employ a transformer-based neural network to predict initial long-term correspondences across frames and integrate a robust Dynamic Units System to refine these predictions. This combination facilitates precise pose tracking while maintaining both efficiency and robustness, even under challenging conditions such as object disappearance, reappearance, and handheld motion. The effectiveness of our approach has been rigorously evaluated through both qualitative and quantitative analyses on the OnePose, YCB-V, and RBOT datasets. The results demonstrate the potential of our proposed Zero6DOT to redefine 6D object pose tracking for real-world scenarios. Deming Zhai, Jianan Zhen, Guofeng Zhang 0001, Xianming Liu 0005 |
IEEE Trans. Circuits Syst. Video Technol. | 7 |
| 2025 | FusionINV: A Diffusion-Based Approach for Multimodal Image FusionabstractInfrared images exhibit a significantly different appearance compared to visible counterparts. Existing infrared and visible image fusion (IVF) methods fuse features from both infrared and visible images, producing a new "image" appearance not inherently captured by any existing device. From an appearance perspective, infrared, visible, and fused images belong to different data domains. This difference makes it challenging to apply fused images because their domain-specific appearance may be difficult for downstream systems, e.g., pre-trained segmentation models. Therefore, accurately assessing the quality of the fused image is challenging. To address those problem, we propose a novel IVF method, FusionINV, which produces fused images with an appearance similar to visible images. FusionINV employs the pre-trained Stable Diffusion (SD) model to invert infrared images into the noise feature space. To inject visible-style appearance information into the infrared features, we leverage the inverted features from visible images to guide this inversion process. In this way, we can embed all the information of infrared and visible images in the noise feature space, and then use the prior of the pre-trained SD model to generate visually friendly images that align more closely with the RGB distribution. Specially, to generate the fused image, we design a tailored fusion rule within the denoising process that iteratively fuses visible-style infrared and visible features. In this way, the fused image falls into the visible domain and can be directly applied to existing downstream machine systems. Thanks to advancements in image inversion, FusionINV can directly produce fused images in a training-free manner. Extensive experiments demonstrate that FusionINV achieves outstanding performance in both human visual evaluation and machine perception tasks. The code is available at https://github.com/erfect2020/FusionINV. Pengwei Liang, Junjun Jiang, Chenyang Wang 0002, Xianming Liu 0005, Jiayi Ma 0001 |
IEEE Trans. Image Process. | 5 |
| 2025 | Learning Lossless Compression for High Bit-Depth Volumetric Medical ImageabstractRecent advances in learning-based methods have markedly enhanced the capabilities of image compression. However, these methods struggle with high bit-depth volumetric medical images, facing issues such as degraded performance, increased memory demand, and reduced processing speed. To address these challenges, this paper presents the Bit-Division based Lossless Volumetric Image Compression (BD-LVIC) framework, which is tailored for high bit-depth medical volume compression. The BD-LVIC framework skillfully divides the high bit-depth volume into two lower bit-depth segments: the Most Significant Bit-Volume (MSBV) and the Least Significant Bit-Volume (LSBV). The MSBV concentrates on the most significant bits of the volumetric medical image, capturing vital structural details in a compact manner. This reduction in complexity greatly improves compression efficiency using traditional codecs. Conversely, the LSBV deals with the least significant bits, which encapsulate intricate texture details. To compress this detailed information effectively, we introduce an effective learning-based compression model equipped with a Transformer-Based Feature Alignment Module, which exploits both intra-slice and inter-slice redundancies to accurately align features. Subsequently, a Parallel Autoregressive Coding Module merges these features to precisely estimate the probability distribution of the least significant bit-planes. Our extensive testing demonstrates that the BD-LVIC framework not only sets new performance benchmarks across various datasets but also maintains a competitive coding speed, highlighting its significant potential and practical utility in the realm of volumetric medical image compression. Kai Wang 0070, Yuanchao Bai, Daxin Li, Deming Zhai, Junjun Jiang, Xianming Liu 0005 |
IEEE Trans. Image Process. | 6 |
| 2025 | Learning Dynamic Prompts for All-in-One Image RestorationabstractAll-in-one image restoration, which seeks to handle multiple types of degradation within a unified model, has become a prominent research topic in computer vision. While existing deep learning models have achieved remarkable success in specific restoration tasks, extending these models to heterogenous degradations presents significant challenges. Current all-in-one methods predominantly concentrate on extracting degradation priors, often employing learned and fixed task prompts to guide the restoration process. However, these static prompts are inclined to generate an average distribution characteristics of degradations, unable to accurately depict the unique attribute of the given input, consequently providing suboptimal restoration results. To tackle these challenges, we propose a novel dynamic prompt approach called Degradation Prototype Assignment and Prompt Distribution Learning (DPPD). Our approach decouples the degradation prior extraction into two novel components: Degradation Prototype Assignment (DPA) and Prompt Distribution Learning (PDL). DPA anchors the degradation representations to predefined prototypes, providing discriminative and scalable representations. In addition, PDL models prompts as distributions rather than fixed parameters, facilitating dynamic and adaptive prompt sampling. Extensive experiments demonstrate that our DPPD framework can achieve significant performance improvement on different image restoration tasks. Codes are available at our project page https://github.com/Aitical/DPPD. Gang Wu 0010, Junjun Jiang, Kui Jiang, Xianming Liu 0005, Liqiang Nie |
IEEE Trans. Image Process. | 4 |
| 2025 | Self-Supervised Multi-Camera Collaborative Depth Prediction With Latent Diffusion ModelsabstractDepth map estimation from images is a crucial task in self-driving applications. Existing methods can be categorized into two groups: multi-view stereo and monocular depth estimation. The former requires cameras to have large overlapping areas and a sufficient baseline between them, while the latter that processes each image independently can hardly guarantee the structure consistency between cameras. In this paper, we propose a novel self-supervised multi-camera collaborative depth prediction method with latent diffusion models, which does not require large overlapping areas while maintaining structure consistency between cameras. Specifically, we introduce MCDP, a new generative foundation model for estimating depth attributes for multi-cameras. We formulate the depth estimation as a weighted combination of depth bases, in which the weights are updated iteratively by the recurrent refinement strategy. During the iterative update, the results of depth estimation are compared across cameras, and the information of overlapping areas is propagated to the whole depth maps with the help of basis formulation in diffusion process. We integrate the GRU-based Weight Net into the diffusion process, allowing the refined hidden state to serve as a conditional input to accurately control the next iterative denoising step. Furthermore, by incorporating the proposed depth consistency loss, we ensure structural consistency across cameras, even in regions with minimal overlap. Experimental results on DDAD, NuScenes, Cityscapes, and Waymo Open Datasets demonstrate the superior performance of our method, and show great help for the downstream task. Jialei Xu, Xianming Liu 0005, Yuanchao Bai, Junjun Jiang, Xiangyang Ji |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2025 | DAWN+: Wavelet-Based Image Deraining Meets Direction-Aware Attention and Mutual RepresentationabstractThe single-image deraining aims to restore clean scenes from rainy inputs by eliminating precipitation artifacts. Current methods often neglect the directional nature of rain streaks-a critical oversight that causes heterogeneous degradation, particularly in texture regions aligned with rain orientations. To address this issue and advance image deraining, we propose a novel direction-aware attention wavelet network (DAWN) for rain streaks removal. DAWN has several key distinctions and innovative features compared with existing wavelet transform-based methods: 1) introducing vector decomposition to parameterize rain distribution through vertical (V) and horizontal (H) component decomposition, enabling explicit direction-aware representation; 2) devising a novel direction-aware attention module (DAM) to learn projection/transformation parameters via coordinate attention mechanisms for precise rain removal and texture preservation; and 3) exploring practical composite constraints to jointly optimize structural coherence, detail fidelity, and chrominance accuracy. Building upon the conference version (DAWN), we devise DAWN+ with enhanced capabilities: 1) decoupling diagonal coefficient learning to eliminate frequency aliasing by characterizing diagonal components with dedicated projection parameters; 2) dividing vector decomposition and parameter fitting into multiple stages to reduce error accumulation; and 3) applying cross-frequency mutual representation to boost training and performance. Experiments across six tasks (deraining, raindrop/rainhaze removal, dehazing, and low-light/underwater enhancement) demonstrate the portability and reusability of these strategies. Meanwhile, DAWN+ delivers significant performance gains over DAWN, achieving an average peak signal to noise ratio (PSNR) increase of 1.17 dB with an acceptable complexity increase. Meanwhile, DAWN+ achieves the competitive performance to the state-of-the-art DRSformer (gaining 0.15 dB in PSNR) while saving 94.4% and 95% model parameters and inference time, respectively. Kui Jiang, Junjun Jiang, Zheng Wang 0007, Zihan Geng, Xianming Liu 0005 |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2025 | Image Deblurring by Exploring In-Depth Properties of TransformerabstractImage deblurring continues to achieve impressive performance with the development of generative models. Nonetheless, there still remains a displeasing problem if one wants to improve perceptual quality and quantitative scores of recovered image at the same time. In this study, drawing inspiration from the research of transformer properties, we introduce the pretrained transformers to address this problem. In particular, we leverage deep features extracted from a pretrained vision transformer (ViT) to encourage recovered images to be sharp without sacrificing the performance measured by the quantitative metrics. The pretrained transformer can capture the global topological relations (i.e., self-similarity) of image, and we observe that the captured topological relationships about the sharp image will change when blur occurs. By comparing the transformer features between recovered image and target one, the pretrained transformer provides high-resolution blur-sensitive semantic information, which is critical in measuring the sharpness of the deblurred image. On the basis of the advantages, we present two types of novel perceptual losses to guide image deblurring. One regards the features as vectors and computes the discrepancy between representations extracted from recovered image and target one in Euclidean space. The other type considers the features extracted from an image as a distribution and compares the distribution discrepancy between recovered image and target one. We demonstrate the effectiveness of transformer properties in improving the perceptual quality while not sacrificing the quantitative scores peak signal-to-noise ratio (PSNR) over the most competitive models, such as Uformer, Restormer, and NAFNet, on defocus deblurring and motion deblurring tasks. The code is available at https://github. com/erfect2020/TransformerPerceptualLoss. Pengwei Liang, Junjun Jiang, Xianming Liu 0005, Jiayi Ma 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2025 | Fast and Accurate 6-D Object Pose Refinement via Implicit Surface OptimizationabstractAligning a point cloud to a fixed 3D model is a crucial task in many applications, such as 6D pose estimation for robotic grasping. Typically, an initial pose is estimated by analyzing both the point cloud and the 3D model, after which the Iterative Closest Point (ICP) algorithm is used to refine the pose, reducing large errors and improving accuracy. In this paper, we propose an accurate and efficient alternative to ICP. Our method encodes the fixed 3D model into an implicit neural network, which is trained offline as a one-time process in just a few minutes, requiring only the CAD model of the object. The network takes the point cloud and pose as inputs and outputs the signed distance field (SDF) value. By minimizing the absolute SDF value with the fixed point cloud and network weights, while optimizing the pose, we obtain the final, precise alignment. The key advantage of our method is that it eliminates the need to explicitly establish one-to-one correspondences between the point cloud and the 3D model, a necessary step in ICP and its variants. This enables our framework to avoid local optima and makes it more robust to challenging conditions such as large initial pose gaps, noisy data, variations in scale, occlusions, and reflections. Furthermore, the end-to-end network of our framework offers significant runtime efficiency. We validate the superior performance of our approach through extensive comparisons with various ICP variants on both synthetic and real-world datasets.The source code of the proposed method is available athttps://github.com/pangbo1997/SDFR. Deming Zhai, Jianan Zhen, Xianming Liu 0005 |
IEEE Trans. Robotics | 5 |
| 2024 | Learning from History: Task-agnostic Model Contrastive Learning for Image RestorationabstractContrastive learning has emerged as a prevailing paradigm for high-level vision tasks, which, by introducing properly negative samples, has also been exploited for low-level vision tasks to achieve a compact optimization space to account for their ill-posed nature. However, existing methods rely on manually predefined and task-oriented negatives, which often exhibit pronounced task-specific biases. To address this challenge, our paper introduces an innovative method termed 'learning from history', which dynamically generates negative samples from the target model itself. Our approach, named Model Contrastive Learning for Image Restoration (MCLIR), rejuvenates latency models as negative models, making it compatible with diverse image restoration tasks. We propose the Self-Prior guided Negative loss (SPN) to enable it. This approach significantly enhances existing models when retrained with the proposed model contrastive paradigm. The results show significant improvements in image restoration across various tasks and architectures. For example, models retrained with SPN outperform the original FFANet and DehazeFormer by 3.41 and 0.57 dB on the RESIDE indoor dataset for image dehazing. Similarly, they achieve notable improvements of 0.47 dB on SPA-Data over IDT for image deraining and 0.12 dB on Manga109 for a 4x scale super-resolution over lightweight SwinIR, respectively. Code and retrained models are available at https://github.com/Aitical/MCLIR. Gang Wu 0010, Junjun Jiang, Kui Jiang, Xianming Liu 0005 |
AAAI | 4 |
| 2024 | FMRNet: Image Deraining via Frequency Mutual RevisionabstractThe wavelet transform has emerged as a powerful tool in deciphering structural information within images. And now, the latest research suggests that combining the prowess of wavelet transform with neural networks can lead to unparalleled image deraining results. By harnessing the strengths of both the spatial domain and frequency space, this innovative approach is poised to revolutionize the field of image processing. The fascinating challenge of developing a comprehensive framework that takes into account the intrinsic frequency property and the correlation between rain residue and background is yet to be fully explored. In this work, we propose to investigate the potential relationships among rain-free and residue components at the frequency domain, forming a frequency mutual revision network (FMRNet) for image deraining. Specifically, we explore the mutual representation of rain residue and background components at frequency domain, so as to better separate the rain layer from clean background while preserving structural textures of the degraded images. Meanwhile, the rain distribution prediction from the low-frequency coefficient, which can be seen as the degradation prior is used to refine the separation of rain residue and background components. Inversely, the updated rain residue is used to benefit the low-frequency rain distribution prediction, forming the multi-layer mutual learning. Extensive experiments demonstrate that our proposed FMRNet delivers significant performance gains for seven datasets on image deraining task, surpassing the state-of-the-art method ELFormer by 1.14 dB in PSNR on the Rain100L dataset, while with similar computation cost. Code and retrained models are available at https://github.com/kuijiang94/FMRNet. Kui Jiang, Junjun Jiang, Xianming Liu 0005, Xin Xu 0007, Xianzheng Ma |
AAAI | 3 |
| 2024 | Low-Light Face Super-resolution via Illumination, Structure, and Texture Associated RepresentationabstractHuman face captured at night or in dimly lit environments has become a common practice, accompanied by complex low-light and low-resolution degradations. However, the existing face super-resolution (FSR) technologies and derived cascaded schemes are inadequate to recover credible textures. In this paper, we propose a novel approach that decomposes the restoration task into face structural fidelity maintaining and texture consistency learning. The former aims to enhance the quality of face images while improving the structural fidelity, while the latter focuses on eliminating perturbations and artifacts caused by low-light degradation and reconstruction. Based on this, we develop a novel low-light low-resolution face super-resolution framework. Our method consists of two steps: an illumination correction face super-resolution network (IC-FSRNet) for lighting the face and recovering the structural information, and a detail enhancement model (DENet) for improving facial details, thus making them more visually appealing and easier to analyze. As the relighted regions could provide complementary information to boost face super-resolution and vice versa, we introduce the mutual learning to harness the informative components from relighted regions and reconstruction, and achieve the iterative refinement. In addition, DENet equipped with diffusion probabilistic model is built to further improve face image visual quality. Experiments demonstrate that the proposed joint optimization framework achieves significant improvements in reconstruction quality and perceptual quality over existing two-stage sequential solutions. Code is available at https://github.com/wcy-cs/IC-FSRDENet. Chenyang Wang 0002, Junjun Jiang, Kui Jiang, Xianming Liu 0005 |
AAAI | 4 |
| 2024 | OpticalDR: A Deep Optical Imaging Model for Privacy-Protective Depression RecognitionabstractDepression Recognition (DR) poses a considerable chal-lenge, especially in the context of the growing concerns surrounding privacy. Traditional automatic diagnosis of DR technology necessitates the use of facial images, un-doubtedly expose the patient identity features and poses privacy risks. In order to mitigate the potential risks as-sociated with the inappropriate disclosure of patient fa-cial images, we design a new imaging system to erase the identity information of captured facial images while re-tain disease-relevant features. It is irreversible for identity information recovery while preserving essential disease-related characteristics necessary for accurate DR. More specifically, we try to record a de-identified facial image (erasing the identifiable features as much as possible) by a learnable lens, which is optimized in conjunction with the following DR task as well as a range of face analy-sis related auxiliary tasks in an end-to-end manner. These aforementioned strategies form our final Optical deep De-pression Recognition network (OpticalDR). Experiments on CelebA, AVEC 2013, and AVEC 2014 datasets demonstrate that our OpticalDR has achieved state-of-the-art privacy protection performance with an average AUC of 0.51 on popular facial recognition models, and competitive results for DR with MAEIRMSE of 7.5318.48 on AVEC 2013 and 7.8918.82 on AVEC 2014, respectively. Code is available at https://github.com/divertingPanIOpticalDR. Junjun Jiang, Kui Jiang, Keyuan Yu, Xianming Liu 0005 |
CVPR | 6 |
| 2024 | Improving Domain Generalization in Self-supervised Monocular Depth Estimation via Stabilized Adversarial Training
Yuanqi Yao, Gang Wu 0010, Kui Jiang, Siao Liu, Jian Kuai, Xianming Liu 0005, Junjun Jiang |
ECCV (24) | 6 |
| 2024 | Context-Adaptive Entropy Model With Adapters For Lossless Point Cloud Geometry CompressionabstractLearning-based point cloud compression has achieved tremendous progress in recent years. However, existing methods often train an optimal occupancy distribution predictor for the entire train dataset in an amortization sense, which struggles to handle point clouds with unique characteristics. In this work, we focus on the lossless point cloud compression, and propose a novel context-adaptive entropy model to achieve adaptive occupancy prediction. Specifically, given a baseline entropy model and a point cloud, we firstly integrate adapters into diverse feature extraction modules. These adapters are then trained to be specifically attuned to the input cloud. Finally, the trained adapter parameters are encoded and transmitted along with the point cloud bitstream, which allow us to recover the integrated model in decoder. The experimental results demonstrate that our method can enhance the performance of the entropy model, especially improving the compression performance of data that performs poorly in conventional methods. Wenbo Zhao 0004, Daxin Li, Junjun Jiang, Xianming Liu 0005 |
ICIP | 5 |
| 2024 | Zero-Mean Regularized Spectral Contrastive Learning: Implicitly Mitigating Wrong Connections in Positive-Pair GraphsabstractContrastive learning has emerged as a popular paradigm of self-supervised learning that learns representations by encouraging representations of positive pairs to be similar while representations of negative pairs to be far apart. The spectral contrastive loss, in synergy with the notion of positive-pair graphs, offers valuable theoretical insights into the empirical successes of contrastive learning. In this paper, we propose incorporating an additive factor into the term of spectral contrastive loss involving negative pairs. This simple modification can be equivalently viewed as introducing a regularization term that enforces the mean of representations to be zero, which thus is referred to as *zero-mean regularization*. It intuitively relaxes the orthogonality of representations between negative pairs and implicitly alleviates the adverse effect of wrong connections in the positive-pair graph, leading to better performance and robustness. To clarify this, we thoroughly investigate the role of zero-mean regularized spectral contrastive loss in both unsupervised and supervised scenarios with respect to theoretical analysis and quantitative evaluation. These results highlight the potential of zero-mean regularized spectral contrastive learning to be a promising approach in various tasks. Xianming Liu 0005, Feilong Zhang 0002, Gang Wu 0010, Deming Zhai, Junjun Jiang, Xiangyang Ji |
ICLR | 2 |
| 2024 | Variance-enlarged Poisson Learning for Graph-based Semi-Supervised Learning with Extremely Sparse Labeled DataabstractGraph-based semi-supervised learning, particularly in the context of extremely sparse labeled data, often suffers from degenerate solutions where label functions tend to be nearly constant across unlabeled data. In this paper, we introduce Variance-enlarged Poisson Learning (VPL), a simple yet powerful framework tailored to alleviate the issues arising from the presence of degenerate solutions. VPL incorporates a variance-enlarged regularization term, which induces a Poisson equation specifically for unlabeled data. This intuitive approach increases the dispersion of labels from their average mean, effectively reducing the likelihood of degenerate solutions characterized by nearly constant label functions. We subsequently introduce two streamlined algorithms, V-Laplace and V-Poisson, each intricately designed to enhance Laplace and Poisson learning, respectively. Furthermore, we broaden the scope of VPL to encompass graph neural networks, introducing Variance-enlarged Graph Poisson Networks (V-GPN) to facilitate improved label propagation. To achieve a deeper understanding of VPL's behavior, we conduct a comprehensive theoretical exploration in both discrete and variational cases. Our findings elucidate that VPL inherently amplifies the importance of connections within the same class while concurrently tempering those between different classes. We support our claims with extensive experiments, demonstrating the effectiveness of VPL and showcasing its superiority over existing methods. The code is available at https://github.com/hitcszx/VPL. Xianming Liu 0005, Jialiang Wang 0003, Zeke Xie, Junjun Jiang, Xiangyang Ji |
ICLR | 2 |
| 2024 | Exploiting Self-Supervised Constraints in image Super-ResolutionabstractRecent advances in self-supervised learning, predominantly studied in high-level visual tasks, have been explored in low-level image processing. This paper introduces a novel self-supervised constraint for single image super-resolution, termed SSC-SR. SSC-SR uniquely addresses the divergence in image complexity by employing a dual asymmetric paradigm and a target model updated via exponential moving average to enhance stability. The proposed SSC-SR framework works as a plug-and-play paradigm and can be easily applied to existing SR models. Empirical evaluations reveal that our SSC-SR framework delivers substantial enhancements on a variety of benchmark datasets, achieving an average increase of 0.1 dB over EDSR and 0.06 dB over SwinIR. In addition, extensive ablation studies corroborate the effectiveness of each component in our SSC-SR framework. Codes are available at https://github.com/Aitical/SSCSR. Gang Wu 0010, Junjun Jiang, Kui Jiang, Xianming Liu 0005 |
ICME | 4 |
| 2024 | SDGE: Stereo Guided Depth Estimation for 360°Camera SetsabstractDepth estimation is a critical technology in autonomous driving, and multi-camera systems are often used to achieve a 360° perception. These 360° camera sets often have limited or low-quality overlap regions, making multi-view stereo methods infeasible for the entire image. Alternatively, monocular methods may not produce consistent cross-view predictions. To address these issues, we propose the Stereo Guided Depth Estimation (SGDE) method, which enhances depth estimation of the full image by explicitly utilizing multi-view stereo results on the overlap. We suggest building virtual pinhole cameras to resolve the distortion problem of fisheye cameras and unify the processing for the two types of 360° cameras. For handling the varying noise on camera poses caused by unstable movement, the approach employs a self-calibration method to obtain highly accurate relative poses of the adjacent cameras with minor overlap. These enable the use of robust stereo methods to obtain a high-quality depth prior in the overlap region. This prior serves not only as an additional input but also as pseudo-labels that enhance the accuracy of depth estimation methods and improve cross-view prediction consistency. The effectiveness of SGDE is evaluated on one fisheye camera dataset, Synthetic Urban, and two pinhole camera datasets, DDAD and nuScenes. Our experiments demonstrate that SGDE is effective for both supervised and self-supervised depth estimation, and highlight the potential of our method for advancing autonomous driving technology. Our project page is at https://github.com/JialeiXu/SGDE. Jialei Xu, Dong Gong, Junjun Jiang, Xianming Liu 0005 |
IROS | 5 |
| 2024 | Harmony in Diversity: Improving All-in-One Image Restoration via Multi-Task CollaborationabstractDeep learning-based all-in-one image restoration methods have garnered significant attention in recent years due to capable of addressing multiple degradation tasks. These methods focus on extracting task-oriented information to guide the unified model and have achieved promising results through elaborate architecture design. They commonly adopt a simple mix training paradigm, and the proper optimization strategy for all-in-one tasks has been scarcely investigated. This oversight neglects the intricate relationships and potential conflicts among various restoration tasks, consequently leading to inconsistent optimization rhythms. In this paper, we extend and redefine the conventional all-in-one image restoration task as a multi-task learning problem and propose a straightforward yet effective active-reweighting strategy, dubbed Art, to harmonize the optimization of multiple degradation tasks. Art is a plug-and-play optimization strategy designed to mitigate hidden conflicts among multi-task optimization processes. Through extensive experiments on a diverse range of all-in-one image restoration settings, Art has been demonstrated to substantially enhance the performance of existing methods. When incorporated into the AirNet and TransWeather models, it achieves average improvements of 1.16 dB and 1.21 dB on PSNR, respectively. We hope this work will provide a principled framework for collaborating multiple tasks in all-in-one image restoration and pave the way for more efficient and effective restoration models, ultimately advancing the state-of-the-art in this critical research domain. Code and pre-trained models are available at our project page https://github.com/Aitical/Art. Gang Wu 0010, Junjun Jiang, Kui Jiang, Xianming Liu 0005 |
ACM Multimedia | 4 |
| 2024 | Disentangled-Multimodal Privileged Knowledge Distillation for Depression Recognition with Incomplete Multimodal DataabstractDepression recognition (DR) using facial images, audio signals, or language text recordings has achieved remarkable performance. Recently, multimodal DR has shown improved performance over single-modal methods by leveraging information from a combination of these modalities. However, collecting high-quality data containing all modalities poses a challenge. In particular, these methods often encounter performance degradation when certain modalities are either missing or degraded. To tackle this issue, we present a generalizable multimodal framework for DR by aggregating feature disentanglement and privileged knowledge distillation. In detail, our approach aims to disentangle homogeneous and heterogeneous features within multimodal signals while suppressing noise, thereby adaptively aggregating the most informative components for high-quality DR. Subsequently, we leverage knowledge distillation to transfer privileged knowledge from complete modalities to the observed input with limited information, thereby significantly improving the tolerance and compatibility. These strategies form our novel Feature Disentanglement and Privileged knowledge Distillation Network for DR, dubbed Dis2DR. Experimental evaluations on AVEC 2013, AVEC 2014, AVEC 2017, and AVEC 2019 datasets demonstrate the effectiveness of our Dis2DR method. Remarkably, Dis2DR achieves superior performance even when only a single modality is available, surpassing existing state-of-the-art multimodal DR approaches AVA-DepressNet by up to 9.8% on the AVEC 2013 dataset. Junjun Jiang, Kui Jiang, Xianming Liu 0005 |
ACM Multimedia | 4 |
| 2024 | AFBench: A Large-scale Benchmark for Airfoil DesignabstractData-driven generative models have emerged as promising approaches towards achieving efficient mechanical inverse design. However, due to prohibitively high cost in time and money, there is still lack of open-source and large-scale benchmarks in this field. It is mainly the case for airfoil inverse design, which requires to generate and edit diverse geometric-qualified and aerodynamic-qualified airfoils following the multimodal instructions, \emph{i.e.,} dragging points and physical parameters. This paper presents the open-source endeavors in airfoil inverse design, \emph{AFBench}, including a large-scale dataset with 200 thousand airfoils and high-quality aerodynamic and geometric labels, two novel and practical airfoil inverse design tasks, \emph{i.e.,} conditional generation on multimodal physical parameters, controllable editing, and comprehensive metrics to evaluate various existing airfoil inverse design methods. Our aim is to establish \emph{AFBench} as an ecosystem for training and evaluating airfoil inverse design methods, with a specific focus on data-driven controllable inverse design models by multimodal instructions capable of bridging the gap between ideas and execution, the academic research and industrial applications. We have provided baseline models, comprehensive experimental observations, and analysis to accelerate future research. Our baseline model is trained on an RTX 3090 GPU within 16 hours. The codebase, datasets and benchmarks will be available at \url{https://hitcslj.github.io/afbench/}. Jian Liu 0036, Hairun Xie, Wei Liu 0123, Wanli Ouyang, Junjun Jiang, Xianming Liu 0005, Shixiang Tang |
NeurIPS | 9 |
| 2024 | $\epsilon$-Softmax: Approximating One-Hot Vectors for Mitigating Label NoiseabstractNoisy labels pose a common challenge for training accurate deep neural networks. To mitigate label noise, prior studies have proposed various robust loss functions to achieve noise tolerance in the presence of label noise, particularly symmetric losses. However, they usually suffer from the underfitting issue due to the overly strict symmetric condition. In this work, we propose a simple yet effective approach for relaxing the symmetric condition, namely **$\epsilon$-softmax**, which simply modifies the outputs of the softmax layer to approximate one-hot vectors with a controllable error $\epsilon$. Essentially, ***$\epsilon$-softmax** not only acts as an alternative for the softmax layer, but also implicitly plays the crucial role in modifying the loss function.* We prove theoretically that **$\epsilon$-softmax** can achieve noise-tolerant learning with controllable excess risk bound for almost any loss function. Recognizing that **$\epsilon$-softmax**-enhanced losses may slightly reduce fitting ability on clean datasets, we further incorporate them with one symmetric loss, thereby achieving a better trade-off between robustness and effective learning. Extensive experiments demonstrate the superiority of our method in mitigating synthetic and real-world label noise. Jialiang Wang 0003, Deming Zhai, Junjun Jiang, Xiangyang Ji, Xianming Liu 0005 |
NeurIPS | 6 |
| 2024 | Semantic Ensemble Loss and Latent Refinement for High-Fidelity Neural Image CompressionabstractRecent advancements in neural compression have surpassed traditional codecs in PSNR and MS-SSIM measurements. However, at low bit-rates, these methods can introduce visually displeasing artifacts, such as blurring, color shifting, and texture loss, thereby compromising perceptual quality of images. To address these issues, this study presents an enhanced neural compression method designed for optimal visual fidelity. We have trained our model with a sophisticated semantic ensemble loss, integrating Charbonnier loss, perceptual loss, style loss, and a non-binary adversarial loss, to enhance the perceptual quality of image reconstructions. Additionally, we have implemented a latent refinement process to generate content-aware latent codes. These codes adhere to bit-rate constraints, and prioritize bit allocation to regions of greater importance. Our empirical findings demonstrate that this approach significantly improves the statistical fidelity of neural image compression. Daxin Li, Yuanchao Bai, Kai Wang 0070, Junjun Jiang, Xianming Liu 0005 |
VCIP | 5 |
| 2024 | Enhancing Privacy-Utility Tradeoff with Few-Round Strategy in Heterogeneous Federated LearningabstractFederated learning inherently provides a certain level of privacy protection, which however is often inadequate in many real-world scenarios. Existing privacy-preserving methods frequently incur unbearable time overheads or result in non-negligible deterioration to model performance, thus suffering from the tradeoff between performance and privacy. In this work, we propose a novel Federated Privacy-Preserving Knowledge Transfer framework, namely FedPPKT, which employs data-free knowledge distillation in a meta-learning manner to rapidly generates pseudo data and performs privacy-preserving knowledge transfer. FedPPKT establishes a protective barrier between the original private data and the federated model, thereby ensuring user privacy. Furthermore, leveraging the few-round strategy of FedPPKT, it has the capability to reduce the number of communication rounds, further mitigating the risk of privacy exposure for user data. With the help of the meta generator, the problem of uneven local label distribution on clients is alleviated, mitigating data heterogeneity and improving model performance. Experiments show that FedPPKT outperforms the state-of-the-art privacy-preserving federated learning methods. Our code is publicly available at https://github.com/HIT-weiqb/FedPPKT. Qingbin Wei, Feilong Zhang 0002, Yuanchao Bai, Deming Zhai, Junjun Jiang, Xianming Liu 0005 |
VCIP | 6 |
| 2024 | Deep Lossy Plus Residual Coding for Lossless and Near-Lossless Image CompressionabstractLossless and near-lossless image compression is of paramount importance to professional users in many technical fields, such as medicine, remote sensing, precision engineering and scientific research. But despite rapidly growing research interests in learning-based image compression, no published method offers both lossless and near-lossless modes. In this paper, we propose a unified and powerful deep lossy plus residual (DLPR) coding framework for both lossless and near-lossless image compression. In the lossless mode, the DLPR coding system first performs lossy compression and then lossless coding of residuals. We solve the joint lossy and residual compression problem in the approach of VAEs, and add autoregressive context modeling of the residuals to enhance lossless compression performance. In the near-lossless mode, we quantize the original residuals to satisfy a given ℓ∞error bound, and propose a scalable near-lossless compression scheme that works for variable ℓ∞bounds instead of training multiple networks. To expedite the DLPR coding, we increase the degree of algorithm parallelization by a novel design of coding context, and accelerate the entropy coding with adaptive residual interval. Experimental results demonstrate that the DLPR coding system achieves both the state-of-the-art lossless and near-lossless image compression performance with competitive coding speed. Yuanchao Bai, Xianming Liu 0005, Kai Wang 0070, Xiangyang Ji, Xiaolin Wu 0001, Wen Gao 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2024 | GroupedMixer: An Entropy Model With Group-Wise Token-Mixers for Learned Image CompressionabstractTransformer-based entropy models have gained prominence in recent years due to their superior ability to capture long-range dependencies in probability distribution estimation compared to convolution-based methods. However, previous transformer-based entropy models suffer from a sluggish coding process due to pixel-wise autoregression or duplicated computation during inference. In this paper, we propose a novel transformer-based entropy model called GroupedMixer, which enjoys both faster coding speed and better compression performance than previous transformer-based methods. Specifically, our approach builds upon group-wise autoregression by first partitioning the latent variables into groups along spatial-channel dimensions, and then entropy coding the groups with the proposed transformer-based entropy model. The global causal self-attention is decomposed into more efficient group-wise interactions, implemented using inner-group and cross-group token-mixers. The inner-group token-mixer incorporates contextual elements within a group while the cross-group token-mixer interacts with previously decoded groups. Alternate arrangement of two token-mixers enables global contextual reference. To further expedite the network inference, we introduce context cache optimization to GroupedMixer, which caches attention activation values in cross-group token-mixers and avoids complex and duplicated computation. Experimental results demonstrate that the proposed GroupedMixer yields the state-of-the-art rate-distortion performance with fast compression speed. Daxin Li, Yuanchao Bai, Kai Wang 0070, Junjun Jiang, Xianming Liu 0005, Wen Gao 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2024 | Incrementally Adapting Pretrained Model Using Network Prior for Multi-Focus Image FusionabstractMulti-focus image fusion can fuse the clear parts of two or more source images captured at the same scene with different focal lengths into an all-in-focus image. On the one hand, previous supervised learning-based multi-focus image fusion methods relying on synthetic datasets have a clear distribution shift with real scenarios. On the other hand, unsupervised learning-based multi-focus image fusion methods can well adapt to the observed images but lack the general knowledge of defocus blur that can be learned from paired data. To avoid the problems of existing methods, this paper presents a novel multi-focus image fusion model by considering both the general knowledge brought by the supervised pretrained backbone and the extrinsic priors optimized on specific testing sample to improve the performance of image fusion. To be specific, the Incremental Network Prior Adaptation (INPA) framework is proposed to incrementally integrate features extracted from the pretrained strong baselines into a tiny prior network (6.9% parameters of the backbone network) to boost the performance for test samples. We evaluate our method on both synthetic and real-world public datasets (Lytro, MFI-WHU, and Real-MFF) and show that our method outperforms existing supervised learning-based methods and unsupervised learning based methods. Junjun Jiang, Chenyang Wang 0002, Xianming Liu 0005, Jiayi Ma 0001 |
IEEE Trans. Image Process. | 4 |
| 2024 | BinsFormer: Revisiting Adaptive Bins for Monocular Depth EstimationabstractMonocular depth estimation (MDE) is a fundamental task in computer vision and has drawn increasing attention. Recently, some methods reformulate it as a classification-regression task to boost the model performance, where continuous depth is estimated via a linear combination of predicted probability distributions and discrete bins. In this paper, we present a novel framework called BinsFormer, tailored for the classification-regression-based depth estimation. It mainly focuses on two crucial components in the specific task: 1) proper generation of adaptive bins; and 2) sufficient interaction between probability distribution and bins predictions. To specify, we employ a Transformer decoder to generate bins, novelly viewing it as a direct set-to-set prediction problem. We further integrate a multi-scale decoder structure to achieve a comprehensive understanding of spatial geometry information and estimate depth maps in a coarse-to-fine manner. Moreover, an extra scene understanding query is proposed to improve the estimation accuracy, which turns out that models can implicitly learn useful information from the auxiliary environment classification task. Extensive experiments on the KITTI, NYU, and SUN RGB-D datasets demonstrate that BinsFormer surpasses state-of-the-art MDE methods with prominent margins. Code and pretrained models are made publicly available at https://github.com/zhyever/ Monocular-Depth-Estimation-Toolbox/tree/main/configs/ binsformer. Zhenyu Li 0007, Xianming Liu 0005, Junjun Jiang |
IEEE Trans. Image Process. | 3 |
| 2024 | Transforming Image Super-Resolution: A ConvFormer-Based Efficient ApproachabstractRecent progress in single-image super-resolution (SISR) has achieved remarkable performance, yet the computational costs of these methods remain a challenge for deployment on resource-constrained devices. In particular, transformer-based methods, which leverage self-attention mechanisms, have led to significant breakthroughs but also introduce substantial computational costs. To tackle this issue, we introduce the Convolutional Transformer layer (ConvFormer) and propose a ConvFormer-based Super-Resolution network (CFSR), offering an effective and efficient solution for lightweight image super-resolution. The proposed method inherits the advantages of both convolution-based and transformer-based approaches. Specifically, CFSR utilizes large kernel convolutions as a feature mixer to replace the self-attention module, efficiently modeling long-range dependencies and extensive receptive fields with minimal computational overhead. Furthermore, we propose an edge-preserving feed-forward network (EFN) designed to achieve local feature aggregation while effectively preserving high-frequency information. Extensive experiments demonstrate that CFSR strikes an optimal balance between computational cost and performance compared to existing lightweight SR methods. When benchmarked against state-of-the-art methods such as ShuffleMixer, the proposed CFSR achieves a gain of 0.39 dB on the Urban100 dataset for the x2 super-resolution task while requiring 26% and 31% fewer parameters and FLOPs, respectively. The code and pre-trained models are available at https://github.com/Aitical/CFSR. Gang Wu 0010, Junjun Jiang, Junpeng Jiang, Xianming Liu 0005 |
IEEE Trans. Image Process. | 4 |
| 2024 | ReSmooth: Detecting and Utilizing OOD Samples When Training With Data AugmentationabstractData augmentation (DA) is a widely used technique for enhancing the training of deep neural networks. Recent DA techniques which achieve state-of-the-art performance always meet the need for diversity in augmented training samples. However, an augmentation strategy that has a high diversity usually introduces out-of-distribution (OOD) augmented samples and these samples consequently impair the performance. To alleviate this issue, we propose ReSmooth, a framework that first detects OOD samples in augmented samples and then leverages them. To be specific, we first use a Gaussian mixture model (GMM) to fit the loss distribution of both the original and augmented samples and accordingly split these samples into in-distribution (ID) samples and OOD samples. Then we start a new training where ID and OOD samples are incorporated with different smooth labels. By treating ID samples and OOD samples unequally, we can make better use of the diverse augmented data. Furthermore, we incorporate our ReSmooth framework with negative DA (NDA) strategies. By properly handling their intentionally created OOD samples, the classification performance of NDAs is largely ameliorated. Experiments on several classification benchmarks show that ReSmooth can be easily extended to the existing augmentation strategies [such as RandAugment (RA), rotate, and jigsaw] and improve on them. Our code is available at https://github.com/Chenyang4/ReSmooth. Chenyang Wang 0002, Junjun Jiang, Xianming Liu 0005 |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2024 | A Practical Contrastive Learning Framework for Single-Image Super-ResolutionabstractContrastive learning has achieved remarkable success on various high-level tasks, but there are fewer contrastive learning-based methods proposed for low-level tasks. It is challenging to adopt vanilla contrastive learning technologies proposed for high-level visual tasks to low-level image restoration problems straightly. Because the acquired high-level global visual representations are insufficient for low-level tasks requiring rich texture and context information. In this article, we investigate the contrastive learning-based single-image super-resolution (SISR) from two perspectives: positive and negative sample construction and feature embedding. The existing methods take naive sample construction approaches (e.g., considering the low-quality input as a negative sample and the ground truth as a positive sample) and adopt a prior model (e.g., pretrained very deep convolutional networks proposed by visual geometry group (VGG) model) to obtain the feature embedding. To this end, we propose a practical contrastive learning framework for SISR (PCL-SR). We involve the generation of many informative positive and hard negative samples in frequency space. Instead of utilizing an additional pretrained network, we design a simple but effective embedding network inherited from the discriminator network, which is more task-friendly. Compared with the existing benchmark methods, we retrain them by our proposed PCL-SR framework and achieve superior performance. Extensive experiments have been conducted to show the effectiveness and technical contributions of our proposed PCL-SR thorough ablation studies. The code and resulting models will be released via https://github.com/Aitical/PCL-SISR. Gang Wu 0010, Junjun Jiang, Xianming Liu 0005 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2024 | Deep Attentional Guided Image FilteringabstractGuided filter is a fundamental tool in computer vision and computer graphics, which aims to transfer structure information from the guide image to the target image. Most existing methods construct filter kernels from the guidance itself without considering the mutual dependency between the guidance and the target. However, since there typically exist significantly different edges in two images, simply transferring all structural information from the guide to the target would result in various artifacts. To cope with this problem, we propose an effective framework named deep attentional guided image filtering, the filtering process of which can fully integrate the complementary information contained in both images. Specifically, we propose an attentional kernel learning module to generate dual sets of filter kernels from the guidance and the target and then adaptively combine them by modeling the pixelwise dependency between the two images. Meanwhile, we propose a multiscale guided image filtering module to progressively generate the filtering result with the constructed kernels in a coarse-to-fine manner. Correspondingly, a multiscale fusion strategy is introduced to reuse the intermediate results in the coarse-to-fine process. Extensive experiments show that the proposed framework compares favorably with the state-of-the-art methods in a wide range of guided image filtering applications, such as guided super-resolution (SR), cross-modality restoration, and semantic segmentation. Moreover, our scheme achieved the first place in the real depth map SR challenge held in ACM ICMR 2021. The codes can be found at https://github.com/zhwzhong/DAGF. Zhiwei Zhong 0001, Xianming Liu 0005, Junjun Jiang, Debin Zhao, Xiangyang Ji |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2024 | Illumination-Aware Low-Light Image Enhancement with Transformer and Auto-Knee CurveabstractImages captured under low-light conditions suffer from several combined degradation factors, including low brightness, low contrast, noise, and color bias. Many learning-based techniques attempt to learn the low-to-clear mapping between low-light and normal-light images. However, they often fall short when applied to low-light images taken in wide-contrast scenes because uneven illumination brings illumination-varying noise and the enhanced images are easily over-saturated in highlight areas. In this article, we present a novel two-stage method to tackle the problem of uneven illumination distribution in low-light images. Under the assumption that noise varies with illumination, we design an illumination-aware transformer network for the first stage of image restoration. In this stage, we introduce the Illumination-aware Attention Block featured with Illumination-aware Multi-head Self-attention, which incorporates different scales of illumination features to guide the attention module, thereby enhancing the denoising and reconstruction capabilities of the restoration network. In the second stage, we innovatively introduce a cubic auto-knee curve transfer with a global parameter predictor to alleviate the over-exposure caused by uneven illumination. We also adopt a white balance correction module to address color bias issues at this stage. Extensive experiments on various benchmarks demonstrate the advantages of our method over state-of-the-art methods qualitatively and quantitatively. Jinwang Pan, Xianming Liu 0005, Yuanchao Bai, Deming Zhai, Junjun Jiang, Debin Zhao |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2023 | Spatial-Frequency Mutual Learning for Face Super-ResolutionabstractFace super-resolution (FSR) aims to reconstruct high-resolution (HR) face images from the low-resolution (LR) ones. With the advent of deep learning, the FSR technique has achieved significant breakthroughs. However, existing FSR methods either have a fixed receptive field or fail to maintain facial structure, limiting the FSRperformance. To circumvent this problem, Fourier transform is introduced, which can capture global facial structure information and achieve image-size receptive field. Relying on the Fourier transform, we devise a spatial-frequency mutual network (SFMNet) for FSR, which is the first FSR method to explore the correlations between spatial and frequency domains as far as we know. To be specific, our SFMNet is a two-branch network equipped with a spatial branch and a frequency branch. Benefiting from the property of Fourier transform, the frequency branch can achieve image-size receptive field and capture global dependency while the spatial branch can extract local dependency. Considering that these dependencies are complementary and both favorable for FSR, we further develop a frequency-spatial interaction block (FSIB) which mutually amalgamates the complementary spatial and frequency information to enhance the capability of the model. Quantitative and qualitative experimental results show that the proposed method out-performs state-of-the-art FSR methods in recovering face images. The implementation and model will be released at https://github.com/wcy-cs/SFMNet. Chenyang Wang 0002, Junjun Jiang, Zhiwei Zhong 0001, Xianming Liu 0005 |
CVPR | 4 |
| 2023 | SynFacePAD 2023: Competition on Face Presentation Attack Detection Based on Privacy-aware Synthetic Training DataabstractThis paper presents a summary of the Competition on Face Presentation Attack Detection Based on Privacy-aware Synthetic Training Data (SynFacePAD 2023) held at the 2023 International Joint Conference on Biometrics (IJCB 2023). The competition attracted a total of 8 participating teams with valid submissions from academia and industry. The competition aimed to motivate and attract solutions that target detecting face presentation attacks while considering synthetic-based training data motivated by privacy, legal and ethical concerns associated with personal data. To achieve that, the training data used by the participants was limited to synthetic data provided by the organizers. The submitted solutions presented innovations and novel approaches that led to outperforming the considered baseline in the investigated benchmarks. Meiling Fang, Marco Huber, Julian Fierrez, Ramachandra Raghavendra, Naser Damer, Alhasan Alkhaddour, Maksim Kasantcev, Vasiliy Pryadchenko, Ziyuan Yang 0001, Huijie Huangfu, Yi Zhang 0018, Junjun Jiang, Xianming Liu 0005, Xianyun Sun, Caiyong Wang, Zhaohua Chang, Guangzhe Zhao, Juan E. Tapia, Lázaro J. González Soler, Carlos M. Aravena, Daniel Schulz |
IJCB | 15 |
| 2023 | Learning Lossless Compression for High Bit-Depth Medical ImagingabstractWe propose a learned lossless image compression method for high bit-depth medical imaging (up to 16 bit-depths). Instead of compressing a high bit-depth medical image as a whole, we split it into two low bit-depth subimages, i.e., the most significant bytes (MSB) subimage and the least significant bytes (LSB) subimage, respectively. The MSB subimage depicts piece-wise smooth structure information that is relatively easy to compress. We thus use traditional lossless codecs for low complexity. The LSB subimage depicts the complementary texture information that is more challenging to compress. We design an autoregressive entropy model conditioned on the MSB subimage that models the probability distribution of the LSB subimage and effectively reduces the redundancy between the MSB and LSB subimages. We then encode the LSB subimage to bitstreams based on the learned entropy model. The compressed high bit-depth medical image is finally stored including the bitstreams of the MSB and LSB subimages. Experimental results demonstrate the state-of-the-art compression performance of the proposed method on high bit-depth medical images, compared with both existing traditional and learned lossless image codecs. Kai Wang 0070, Yuanchao Bai, Deming Zhai, Daxin Li, Junjun Jiang, Xianming Liu 0005 |
ICME | 6 |
| 2023 | No One Idles: Efficient Heterogeneous Federated Learning with Parallel Edge and Server ComputationabstractFederated learning suffers from a latency bottleneck induced by network stragglers, which hampers the training efficiency significantly. In addition, due to the heterogeneous data distribution and security requirements, simple and fast averaging aggregation is not feasible anymore. Instead, complicated aggregation operations, such as knowledge distillation, are required. The time cost for complicated aggregation becomes a new bottleneck that limits the computational efficiency of FL. In this work, we claim that the root cause of training latency actually lies in the aggregation-then-broadcasting workflow of the server. By swapping the computational order of aggregation and broadcasting, we propose a novel and efficient parallel federated learning (PFL) framework that unlocks the edge nodes during global computation and the central server during local computation. This fully asynchronous and parallel pipeline enables handling complex aggregation and network stragglers, allowing flexible device participation as well as achieving scalability in computation. We theoretically prove that synchronous and asynchronous PFL can achieve a similar convergence rate as vanilla FL. Extensive experiments empirically show that our framework brings up to $5.56\times$ speedup compared with traditional FL. Code is available at: https://github.com/Hypervoyager/PFL. Feilong Zhang 0002, Xianming Liu 0005, Gang Wu 0010, Junjun Jiang, Xiangyang Ji |
ICML | 2 |
| 2023 | On the Dynamics Under the Unhinged Loss and BeyondabstractRecent works have studied implicit biases in deep learning, especially the behavior of last-layer features and classifier weights. However, they usually need to simplify the intermediate dynamics under gradient flow or gradient descent due to the intractability of loss functions and model architectures. In this paper, we introduce the unhinged loss, a concise loss function, that offers more mathematical opportunities to analyze the closed-form dynamics while requiring as few simplifications or assumptions as possible. The unhinged loss allows for considering more practical techniques, such as time-vary learning rates and feature normalization. Based on the layer-peeled model that views last-layer features as free optimization variables, we conduct a thorough analysis in the unconstrained, regularized, and spherical constrained cases, as well as the case where the neural tangent kernel remains invariant. To bridge the performance of the unhinged loss to that of Cross-Entropy (CE), we investigate the scenario of fixing classifier weights with a specific structure, (e.g., a simplex equiangular tight frame). Our analysis shows that these dynamics converge exponentially fast to a solution depending on the initialization of features and classifier weights. These theoretical results not only offer valuable insights, including explicit feature regularization and rescaled learning rates for enhancing practical training with the unhinged loss, but also extend their applicability to other loss functions. Finally, we empirically demonstrate these theoretical results and insights through extensive experiments. Xianming Liu 0005, Deming Zhai, Junjun Jiang, Xiangyang Ji |
J. Mach. Learn. Res. | 2 |
| 2023 | Self-Supervised Arbitrary-Scale Implicit Point Clouds UpsamplingabstractPoint clouds upsampling (PCU), which aims to generate dense and uniform point clouds from the captured sparse input of 3D sensor such as LiDAR, is a practical yet challenging task. It has potential applications in many real-world scenarios, such as autonomous driving, robotics, AR/VR, etc. Deep neural network based methods achieve remarkable success in PCU. However, most existing deep PCU methods either take the end-to-end supervised training, where large amounts of pairs of sparse input and dense ground-truth are required to serve as the supervision; or treat up-scaling of different factors as independent tasks, where multiple networks are required for different scaling factors, leading to significantly increased model complexity and training time. In this article, we propose a novel method that achieves self-supervised and magnification-flexible PCU simultaneously. No longer explicitly learning the mapping between sparse and dense point clouds, we formulate PCU as the task of seeking nearest projection points on the implicit surface for seed points. We then define two implicit neural functions to estimate projection direction and distance respectively, which can be trained by the pretext learning tasks. Moreover, the projection rectification strategy is tailored to remove outliers so as to keep the shape of object clear and sharp. Experimental results demonstrate that our self-supervised learning based scheme achieves competitive or even better performance than state-of-the-art supervised methods. Wenbo Zhao 0004, Xianming Liu 0005, Deming Zhai, Junjun Jiang, Xiangyang Ji |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2023 | Asymmetric Loss Functions for Noise-Tolerant Learning: Theory and ApplicationsabstractSupervised deep learning has achieved tremendous success in many computer vision tasks, which however is prone to overfit noisy labels. To mitigate the undesirable influence of noisy labels, robust loss functions offer a feasible approach to achieve noise-tolerant learning. In this work, we systematically study the problem of noise-tolerant learning with respect to both classification and regression. Specifically, we propose a new class of loss function, namelyasymmetric loss functions(ALFs), which are tailored to satisfy the Bayes-optimal condition and thus are robust to noisy labels. For classification, we investigate general theoretical properties of ALFs on categorical noisy labels, and introduce the asymmetry ratio to measure the asymmetry of a loss function. We extend several commonly-used loss functions, and establish the necessary and sufficient conditions to make them asymmetric and thus noise-tolerant. For regression, we extend the concept of noise-tolerant learning for image restoration with continuous noisy labels. We theoretically prove that$\ell _{p}$loss ($p>0$) is noise-tolerant for targets with the additive white Gaussian noise. For targets with general noise, we introduce two losses as surrogates of$\ell _{0}$loss that seeks the mode when clean pixels keep dominant. Experimental results demonstrate that ALFs can achieve better or comparative performance compared with the state-of-the-arts. The source code of our method is available at:https://github.com/hitcszx/ALFs. Xianming Liu 0005, Deming Zhai, Junjun Jiang, Xiangyang Ji |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2023 | Unsupervised Deep Exemplar Colorization via Pyramid Dual Non-Local AttentionabstractExemplar-based colorization is a challenging task, which attempts to add colors to the target grayscale image with the aid of a reference color image, so as to keep the target semantic content while with the reference color style. In order to achieve visually plausible chromatic results, it is important to sufficiently exploit the global color style and the semantic color information of the reference color image. However, existing methods are either clumsy in exploiting the semantic color information, or lack of the dedicated fusion mechanism to decorate the target grayscale image with the reference semantic color information. Besides, these methods usually use a single-stage encoder-decoder architecture, which results in the loss of spatial details. To remedy these problems, we propose an effective exemplar colorization strategy based on pyramid dual non-local attention network to exploit the long-range dependency as well as multi-scale correlation. Specifically, two symmetrical branches of pyramid non-local attention block are tailored to achieve alignments from the target feature to the reference feature and from the reference feature to the target feature respectively. The bidirectional non-local fusion strategy is further applied to get a sufficient fusion feature that achieves full semantic consistency between multi-modal information. To train the network, we propose an unsupervised learning manner, which employs the hybrid supervision including the pseudo paired supervision from the reference color images and unpaired supervision from both the target grayscale and reference color images. Extensive experimental results are provided to demonstrate that our method achieves better photo-realistic colorization performance than the state-of-the-art methods. Deming Zhai, Xianming Liu 0005, Junjun Jiang, Wen Gao 0001 |
IEEE Trans. Image Process. | 3 |
| 2022 | Towards End-to-End Image Compression and Analysis with TransformersabstractWe propose an end-to-end image compression and analysis model with Transformers, targeting to the cloud-based image classification application. Instead of placing an existing Transformer-based image classification model directly after an image codec, we aim to redesign the Vision Transformer (ViT) model to perform image classification from the compressed features and facilitate image compression with the long-term information from the Transformer. Specifically, we first replace the patchify stem (i.e., image splitting and embedding) of the ViT model with a lightweight image encoder modelled by a convolutional neural network. The compressed features generated by the image encoder are injected convolutional inductive bias and are fed to the Transformer for image classification bypassing image reconstruction. Meanwhile, we propose a feature aggregation module to fuse the compressed features with the selected intermediate features of the Transformer, and feed the aggregated features to a deconvolutional neural network for image reconstruction. The aggregated features can obtain the long-term information from the self-attention mechanism of the Transformer and improve the compression performance. The rate-distortion-accuracy optimization problem is finally solved by a two-step training strategy. Experimental results demonstrate the effectiveness of the proposed model in both the image compression and the classification tasks. Yuanchao Bai, Xianming Liu 0005, Junjun Jiang, Yaowei Wang 0001, Xiangyang Ji, Wen Gao 0001 |
AAAI | 3 |
| 2022 | SimIPU: Simple 2D Image and 3D Point Cloud Unsupervised Pre-training for Spatial-Aware Visual RepresentationsabstractPre-training has become a standard paradigm in many computer vision tasks. However, most of the methods are generally designed on the RGB image domain. Due to the discrepancy between the two-dimensional image plane and the three-dimensional space, such pre-trained models fail to perceive spatial information and serve as sub-optimal solutions for 3D-related tasks. To bridge this gap, we aim to learn a spatial-aware visual representation that can describe the three-dimensional space and is more suitable and effective for these tasks. To leverage point clouds, which are much more superior in providing spatial information compared to images, we propose a simple yet effective 2D Image and 3D Point cloud Unsupervised pre-training strategy, called SimIPU. Specifically, we develop a multi-modal contrastive learning framework that consists of an intra-modal spatial perception module to learn a spatial-aware representation from point clouds and an inter-modal feature interaction module to transfer the capability of perceiving spatial information from the point cloud encoder to the image encoder, respectively. Positive pairs for contrastive losses are established by the matching algorithm and the projection matrix. The whole framework is trained in an unsupervised end-to-end fashion. To the best of our knowledge, this is the first study to explore contrastive learning pre-training strategies for outdoor multi-modal datasets, containing paired camera images and LIDAR point clouds. Zhenyu Li 0007, Liangji Fang, Qinhong Jiang, Xianming Liu 0005, Junjun Jiang, Bolei Zhou, Hang Zhao 0021 |
AAAI | 6 |
| 2022 | Local Surface Descriptor for Geometry and Feature Preserved Mesh Denoisingabstract3D meshes are widely employed to represent geometry structure of 3D shapes. Due to limitation of scanning sensor precision and other issues, meshes are inevitably affected by noise, which hampers the subsequent applications. Convolultional neural networks (CNNs) achieve great success in image processing tasks, including 2D image denoising, and have been proven to own the capacity of modeling complex features at different scales, which is also particularly useful for mesh denoising. However, due to the nature of irregular structure, CNNs-based denosing strategies cannot be trivially applied for meshes. To circumvent this limitation, in the paper, we propose the local surface descriptor (LSD), which is able to transform the local deformable surface around a face into 2D grid representation and thus facilitates the deployment of CNNs to generate denoised face normals. To verify the superiority of LSD, we directly feed LSD into the classical Resnet without any complicated network design. The extensive experimental results show that, compared to the state-of-the-arts, our method achieves encouraging performance with respect to both objective and subjective evaluations. Wenbo Zhao 0004, Xianming Liu 0005, Junjun Jiang, Debin Zhao, Ge Li 0002, Xiangyang Ji |
AAAI | 2 |
| 2022 | Self-Supervised Arbitrary-Scale Point Clouds Upsampling via Implicit Neural RepresentationabstractPoint clouds upsampling is a challenging issue to gener-ate dense and uniform point clouds from the given sparse input. Most existing methods either take the end-to-end su-pervised learning based manner, where large amounts of pairs of sparse input and dense ground-truth are exploited as supervision information; or treat up-scaling of different scale factors as independent tasks, and have to build multiple networks to handle upsampling with varying factors. In this paper, we propose a novel approach that achieves self-supervised and magnification-flexible point clouds upsampling simultaneously. We formulate point clouds upsampling as the task of seeking nearest projection points on the implicit surface for seed points. To this end, we define two implicit neural functions to estimate projection direction and distance respectively, which can be trained by two pretext learning tasks. Experimental results demonstrate that our self-supervised learning based scheme achieves competitive or even better performance than supervised learning based state-of-the-art methods. The source code is publicly available at https://github.com/xnowbzhaolsapcu. Wenbo Zhao 0004, Xianming Liu 0005, Zhiwei Zhong 0001, Junjun Jiang, Wei Gao 0003, Ge Li 0002, Xiangyang Ji |
CVPR | 2 |
| 2022 | Shadows can be Dangerous: Stealthy and Effective Physical-world Adversarial Attack by Natural PhenomenonabstractEstimating the risk level of adversarial examples is essential for safely deploying machine learning models in the real world. One popular approach for physical-world attacks is to adopt the “sticker-pasting” strategy, which however suffers from some limitations, including difficulties in access to the target or printing by valid colors. A new type of non-invasive attacks emerged recently, which attempt to cast perturbation onto the target by optics based tools, such as laser beam and projector. However, the added optical patterns are artificial but not natural. Thus, they are still conspicuous and attention-grabbed, and can be easily noticed by humans. In this paper, we study a new type of optical adversarial examples, in which the perturbations are generated by a very common natural phenomenon, shadow, to achieve naturalistic and stealthy physical-world adversarial attack under the black-box setting. We extensively evaluate the effectiveness of this new attack on both simulated and real-world environments. Experimental results on traffic sign recognition demonstrate that our algorithm can generate adversarial examples effectively, reaching 98.23% and 90.47% success rates on LISA and GTSRB test sets respectively, while continuously misleading a moving camera over 95% of the time in real-world scenarios. We also offer discussions about the limitations and the defense mechanism of this attack11Our code is available at https://github.com/hncszyq/ShadowAttack. Yiqi Zhong, Xianming Liu 0005, Deming Zhai, Junjun Jiang, Xiangyang Ji |
CVPR | 2 |
| 2022 | Unsupervised Domain Adaptation for Monocular 3D Object Detection via Self-training
Zhenyu Li 0007, Liangji Fang, Qinhong Jiang, Xianming Liu 0005, Junjun Jiang |
ECCV (9) | 6 |
| 2022 | Fusion from Decomposition: A Self-Supervised Decomposition Approach for Image Fusion
Pengwei Liang, Junjun Jiang, Xianming Liu 0005, Jiayi Ma 0001 |
ECCV (18) | 3 |
| 2022 | Learning Towards The Largest Margins
Xianming Liu 0005, Deming Zhai, Junjun Jiang, Xiangyang Ji |
ICLR | 2 |
| 2022 | Prototype-Anchored Learning for Learning with Imperfect AnnotationsabstractThe success of deep neural networks greatly relies on the availability of large amounts of high-quality annotated data, which however are difficult or expensive to obtain. The resulting labels may be class imbalanced, noisy or human biased. It is challenging to learn unbiased classification models from imperfectly annotated datasets, on which we usually suffer from overfitting or underfitting. In this work, we thoroughly investigate the popular softmax loss and margin-based loss, and offer a feasible approach to tighten the generalization error bound by maximizing the minimal sample margin. We further derive the optimality condition for this purpose, which indicates how the class prototypes should be anchored. Motivated by theoretical analysis, we propose a simple yet effective method, namely prototype-anchored learning (PAL), which can be easily incorporated into various learning-based classification schemes to handle imperfect annotation. We verify the effectiveness of PAL on class-imbalanced learning and noise-tolerant learning by extensive experiments on synthetic and real-world datasets. Xianming Liu 0005, Deming Zhai, Junjun Jiang, Xiangyang Ji |
ICML | 2 |
| 2022 | ChebyLighter: Optimal Curve Estimation for Low-light Image EnhancementabstractLow-light enhancement aims to recover a high contrast normal light image from a low-light image with bad exposure and low contrast. Inspired by curve adjustment in photo editing software and Chebyshev approximation, this paper presents a novel model for brightening low-light images. The proposed model, ChebyLighter, learns to estimate pixel-wise adjustment curves for a low-light image recurrently to reconstruct an enhanced output. In ChebyLighter, Chebyshev image series are first generated. Then pixel-wise coefficient matrices are estimated with Triple Coefficient Estimation (TCE) modules and the final enhanced image is recurrently reconstructed by Chebyshev Attention Weighted Summation (CAWS). The TCE module is specifically designed based on dual attention mechanism with three necessary inputs. Our method can achieve ideal performance because adjustment curves can be obtained with numerical approximation by our model. With extensive quantitative and qualitative experiments on diverse test images, we demonstrate that the proposed method performs favorably against state-of-the-art low-light image enhancement algorithms. Jinwang Pan, Deming Zhai, Yuanchao Bai, Junjun Jiang, Debin Zhao, Xianming Liu 0005 |
ACM Multimedia | 6 |
| 2022 | Hybrid Conditional Deep Inverse Tone MappingabstractEmerging modern displays are capable to render ultra-high definition (UHD) media contents with high dynamic range (HDR) and wide color gamut (WCG). Although more and more native contents as such have been getting produced, the total amount is still in severe lack. Considering the massive amount of legacy contents with standard dynamic range (SDR) which may be exploitable, the urgent demand for proper conversion techniques thus springs up. In this paper, we try to tackle the conversion task from SDR to HDR-WCG for media contents and consumer displays. We propose a deep learning based SDR-to-HDR solution, Hybrid Conditional Deep Inverse Tone Mapping (HyCondITM), which is an end-to-end trainable framework including global transform, local adjustment, and detail refinement in a single unified pipeline. We present a hybrid condition network that can simultaneously extract both global and local priors for guidance to achieve scene-adaptive and spatially-variant manipulations. Experiments show that our method achieves state-of-the-art performance in both quantitative comparisons and visual quality, out-performing the previous methods. Tong Shao, Deming Zhai, Junjun Jiang, Xianming Liu 0005 |
ACM Multimedia | 4 |
| 2022 | Multi-Camera Collaborative Depth Prediction via Consistent Structure EstimationabstractDepth map estimation from images is an important task in robotic systems. Existing methods can be categorized into two groups including multi-view stereo and monocular depth estimation. The former requires cameras to have large overlapping areas and sufficient baseline between cameras, while the latter that processes each image independently can hardly guarantee the structure consistency between cameras. In this paper, we propose a novel multi-camera collaborative depth prediction method that does not require large overlapping areas while maintaining structure consistency between cameras. Specifically, we formulate the depth estimation as a weighted combination of depth basis, in which the weights are updated iteratively by a refinement network driven by the proposed consistency loss. During the iterative update, the results of depth estimation are compared across cameras and the information of overlapping areas is propagated to the whole depth maps with the help of basis formulation. Experimental results on DDAD and NuScenes datasets demonstrate the superior performance of our method. Jialei Xu, Xianming Liu 0005, Yuanchao Bai, Junjun Jiang, Xiaozhi Chen, Xiangyang Ji |
ACM Multimedia | 2 |
| 2022 | Propagating Facial Prior Knowledge for Multitask Learning in Face Super-ResolutionabstractExisting face hallucination methods always achieve improved performance through regularizing the model with facial prior. Most of them always estimate facial prior information first and then leverage it to help the prediction of the target high-resolution face image. However, the accuracy of prior estimation is difficult to guarantee, especially for the low-resolution face image. Once the estimated prior is inaccurate or wrong, the following face super-resolution performance is unavoidably influenced. A natural question that arises: how to incorporate facial prior effectively and efficiently without prior estimation? To achieve this goal, we propose to learn facial prior knowledge at training stage, but test only with low-resolution face image, which can overcome the difficulty of estimating accurate prior. In addition, instead of estimating facial prior, we directly explore the potential of high-quality facial prior in the training phase and progressively propagate the facial prior knowledge from the teacher network (trained with the low-resolution face/high-quality facial prior and high-resolution face image pairs) to the student network (trained with the low-resolution face and high-resolution face image pairs). Quantitative and qualitative comparisons on benchmark face datasets demonstrate that our method outperforms the state-of-the-art face super-resolution methods. The source codes of the proposed method will be available athttps://github.com/wcy-cs/KDFSRNet. Chenyang Wang 0002, Junjun Jiang, Zhiwei Zhong 0001, Xianming Liu 0005 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2022 | Multi-Task Interaction Learning for Spatiospectral Image Super-ResolutionabstractHigh spatial resolution and high spectral resolution images (HR-HSIs) are widely applied in geosciences, medical diagnosis, and beyond. However, how to get images with both high spatial resolution and high spectral resolution is still a problem to be solved. In this paper, we present a deep spatial-spectral feature interaction network (SSFIN) for reconstructing an HR-HSI from a low-resolution multispectral image (LR-MSI), e.g., RGB image. In particular, we introduce two auxiliary tasks, i.e., spatial super-resolution (SR) and spectral SR to help the network recover the HR-HSI better. Since higher spatial resolution can provide more detailed information about image texture and structure, and richer spectrum can provide more attribute information, we propose a spatial-spectral feature interaction block (SSFIB) to make the spatial SR task and the spectral SR task benefit each other. Therefore, we can make full use of the rich spatial and spectral information extracted from the spatial SR task and spectral SR task, respectively. Moreover, we use a weight decay strategy (for the spatial and spectral SR tasks) to train the SSFIN, so that the model can gradually shift attention from the auxiliary tasks to the primary task. Both quantitative and visual results on three widely used HSI datasets demonstrate that the proposed method achieves a considerable gain compared to other state-of-the-art methods. Source code is available at https://github.com/junjun-jiang/SSFIN. Junjun Jiang, Xianming Liu 0005, Jiayi Ma 0001 |
IEEE Trans. Image Process. | 3 |
| 2022 | High-Resolution Depth Maps Imaging via Attention-Based Hierarchical Multi-Modal FusionabstractDepth map records distance between the viewpoint and objects in the scene, which plays a critical role in many real-world applications. However, depth map captured by consumer-grade RGB-D cameras suffers from low spatial resolution. Guided depth map super-resolution (DSR) is a popular approach to address this problem, which attempts to restore a high-resolution (HR) depth map from the input low-resolution (LR) depth and its coupled HR RGB image that serves as the guidance. The most challenging issue for guided DSR is how to correctly select consistent structures and propagate them, and properly handle inconsistent ones. In this paper, we propose a novel attention-based hierarchical multi-modal fusion (AHMF) network for guided DSR. Specifically, to effectively extract and combine relevant information from LR depth and HR guidance, we propose a multi-modal attention based fusion (MMAF) strategy for hierarchical convolutional layers, including a feature enhancement block to select valuable features and a feature recalibration block to unify the similarity metrics of modalities with different appearance characteristics. Furthermore, we propose a bi-directional hierarchical feature collaboration (BHFC) module to fully leverage low-level spatial information and high-level structure information among multi-scale features. Experimental results show that our approach outperforms state-of-the-art methods in terms of reconstruction accuracy, running speed and memory efficiency. Zhiwei Zhong 0001, Xianming Liu 0005, Junjun Jiang, Debin Zhao, Zhiwen Chen 0002, Xiangyang Ji |
IEEE Trans. Image Process. | 2 |
| 2022 | Graph Signal Processing for Geometric Data and Beyond: Theory and ApplicationsabstractGeometric data acquired from real-world scenes,e.g., 2D depth images, 3D point clouds, and 4D dynamic point clouds, have found a wide range of applications including immersive telepresence, autonomous driving, surveillance,etc. Due to irregular sampling patterns of most geometric data, traditional image/video processing methodologies are limited, while Graph Signal Processing (GSP)—a fast-developing field in the signal processing community—enables processing signals that reside on irregular domains and plays a critical role in numerous applications of geometric data from low-level processing to high-level analysis. To further advance the research in this field, we provide the first timely and comprehensive overview of GSP methodologies for geometric data in a unified manner by bridging the connections between geometric data and graphs, among the various geometric data modalities, and with spectral/nodal graph filtering techniques. We also discuss the recently developed Graph Neural Networks (GNNs) and interpret the operation of these networks from the perspective of GSP. We conclude with a brief discussion of open problems and challenges. Wei Hu 0003, Jiahao Pang, Xianming Liu 0005, Dong Tian, Chia-Wen Lin, Anthony Vetro |
IEEE Trans. Multim. | 3 |
| 2022 | Multilayer Spectral-Spatial Graphs for Label Noisy Robust Hyperspectral Image ClassificationabstractIn hyperspectral image (HSI) analysis, label information is a scarce resource and it is unavoidably affected by human and nonhuman factors, resulting in a large amount of label noise. Although most of the recent supervised HSI classification methods have achieved good classification results, their performance drastically decreases when the training samples contain label noise. To address this issue, we propose a label noise cleansing method based on spectral-spatial graphs (SSGs). In particular, an affinity graph is constructed based on spectral and spatial similarity, in which pixels in a superpixel segmentation-based homogeneous region are connected, and their similarities are measured by spectral feature vectors. Then, we use the constructed affinity graph to regularize the process of label noise cleansing. In this manner, we transform label noise cleansing to an optimization problem with a graph constraint. To fully utilize spatial information, we further develop multiscale segmentation-based multilayer SSGs (MSSGs). It can efficiently merge the complementary information of multilayer graphs and thus provides richer spatial information compared with any single-layer graph obtained from isolation segmentation. Experimental results show that MSSG reduces the level of label noise. Compared with the state of the art, the proposed MSSG method exhibits significantly enhanced classification accuracy toward the training data with noisy labels. The significant advantages of the proposed method over four major classifiers are also demonstrated. The source code is available at https://github.com/junjun-jiang/MSSG. Junjun Jiang, Jiayi Ma 0001, Xianming Liu 0005 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2022 | Fully Unsupervised Person Re-Identification via Selective Contrastive LearningabstractPerson re-identification (ReID) aims at searching the same identity person among images captured by various cameras. Existing fully supervised person ReID methods usually suffer from poor generalization capability caused by domain gaps. Unsupervised person ReID has attracted a lot of attention recently, because it works without intensive manual annotation and thus shows great potential in adapting to new conditions. Representation learning plays a critical role in unsupervised person ReID. In this work, we propose a novel selective contrastive learning framework for fully unsupervised feature learning. Specifically, different from traditional contrastive learning strategies, we propose to use multiple positives and adaptively selected negatives for defining the contrastive loss, enabling to learn a feature embedding model with stronger identity discriminative representation. Moreover, we propose to jointly leverage global and local features to construct three dynamic memory banks, among which the global and local ones are used for pairwise similarity computation and the mixture memory bank are used for contrastive loss definition. Experimental results demonstrate the superiority of our method in unsupervised person ReID compared with the state of the art. Our code is available at https://github.com/pangbo1997/Unsup_ReID.git . Deming Zhai, Junjun Jiang, Xianming Liu 0005 |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2022 | Rectified Meta-learning from Noisy Labels for Robust Image-based Plant Disease ClassificationabstractPlant diseases serve as one of main threats to food security and crop production. It is thus valuable to exploit recent advances of artificial intelligence to assist plant disease diagnosis. One popular approach is to transform this problem as a leaf image classification task, which can be then addressed by the powerful convolutional neural networks (CNNs). However, the performance of CNN-based classification approach depends on a large amount of high-quality manually labeled training data, which inevitably introduce noise on labels in practice, leading to model overfitting and performance degradation. To overcome this problem, we propose a novel framework that incorporates rectified meta-learning module into common CNN paradigm to train a noise-robust deep network without using extra supervision information. The proposed method enjoys the following merits: (i) A rectified meta-learning is designed to pay more attention to unbiased samples, leading to accelerated convergence and improved classification accuracy. (ii) Our method is free on assumption of label noise distribution, which works well on various kinds of noise. (iii) Our method serves as a plug-and-play module, which can be embedded into any deep models optimized by gradient descent-based method. Extensive experiments are conducted to demonstrate the superior performance of our algorithm over the state-of-the-arts. Deming Zhai, Ruifeng Shi, Junjun Jiang, Xianming Liu 0005 |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2021 | Learning Scalable lY=-Constrained Near-Lossless Image Compression via Joint Lossy Image and Residual CompressionabstractWe propose a novel joint lossy image and residual compression framework for learning ℓ∞-constrained near-lossless image compression. Specifically, we obtain a lossy reconstruction of the raw image through lossy image compression and uniformly quantize the corresponding residual to satisfy a given tight ℓ∞error bound. Suppose that the error bound is zero, i.e., lossless image compression, we formulate the joint optimization problem of compressing both the lossy image and the original residual in terms of variational auto-encoders and solve it with end-to-end training. To achieve scalable compression with the error bound larger than zero, we derive the probability model of the quantized residual by quantizing the learned probability model of the original residual, instead of training multiple networks. We further correct the bias of the derived probability model caused by the context mismatch between training and inference. Finally, the quantized residual is encoded according to the bias-corrected probability model and is concatenated with the bitstream of the compressed lossy image. Experimental results demonstrate that our near-lossless codec achieves the state-of-the-art performance for lossless and near-lossless image compression, and achieves competitive PSNR while much smaller ℓ∞error compared with lossy image codecs at high bit rates. Yuanchao Bai, Xianming Liu 0005, Wangmeng Zuo, Yaowei Wang 0001, Xiangyang Ji |
CVPR | 2 |
| 2021 | Physics-Based Iterative Projection Complex Neural Network for Phase Retrieval in Lensless Microscopy ImagingabstractPhase retrieval from intensity-only measurements plays a central role in many real-world imaging tasks. In recent years, deep neural networks based methods emerge and show promising performance for phase retrieval. However, their interpretability and generalization still remain a major challenge. In this paper, we propose to combine the advantages of both model-based alternative projection method and deep neural network for phase retrieval, so as to achieve network interpretability and inference effectiveness simultaneously. Specifically, we unfold the iterative process of the alternative projection phase retrieval into a feed-forward neural network, whose layers mimic the processing flow. The physical model of the imaging process is then naturally embedded into the neural network structure. Moreover, a complex-valued U-Net is proposed for defining image priori for forward and backward projection in dual planes. Finally, we designate physics-based formulation as an untrained deep neural network, whose weights are enforced to fit to the given intensity measurements. In summary, our scheme for phase retrieval is effective, interpretable, physics-based and unsupervised. Experimental results demonstrate that our method achieves superior performance compared with the state-of-the-arts in a practical phase retrieval application—lensless microscopy imaging. Feilong Zhang 0002, Xianming Liu 0005, Cheng Guo 0008, Junjun Jiang, Xiangyang Ji |
CVPR | 2 |
| 2021 | Learning with Noisy Labels via Sparse RegularizationabstractLearning with noisy labels is an important and challenging task for training accurate deep neural networks. Some commonly-used loss functions, such as Cross Entropy (CE), suffer from severe overfitting to noisy labels. Robust loss functions that satisfy the symmetric condition were tailored to remedy this problem, which however encounter the underfitting effect. In this paper, we theoretically prove that any loss can be made robust to noisy labels by restricting the network output to the set of permutations over a fixed vector. When the fixed vector is one-hot, we only need to constrain the output to be one-hot, which however produces zero gradients almost everywhere and thus makes gradient-based optimization difficult. In this work, we introduce the sparse regularization strategy to approximate the one-hot constraint, which is composed of network output sharpening operation that enforces the output distribution of a net-work to be sharp and the ℓp-norm (p ≤ 1) regularization that promotes the network output to be sparse. This simple approach guarantees the robustness of arbitrary loss functions while not hindering the fitting ability. Experimental results demonstrate that our method can significantly improve the performance of commonly-used loss functions in the presence of noisy labels and class imbalance, and out-perform the state-of-the-art methods. The code is available at https://github.com/hitcszx/lnl_sr. Xianming Liu 0005, Chenyang Wang 0002, Deming Zhai, Junjun Jiang, Xiangyang Ji |
ICCV | 2 |
| 2021 | Zero-Shot Multi-Focus Image FusionabstractMulti-focus image fusion (MFIF) is an effective way to eliminate the out-of-focus blur generated in the imaging process. The difficulties in focus level estimation and the lack of real training set for supervised learning make MFIF remain a challenging task after decades of research. According to DIP [1], a neural network can capture the low-level statistics of a single image and can be used as a prior for solving many low-level problems. Based on this idea, we propose a novel architecture named IM-Net comprised of I-Net to model the deep prior of the fused image and M-Net to model the deep prior of the focus map. Without any large scale training set, our method achieves zero-shot learning through the extracted prior information. Experiments on extensively used dataset demonstrate the effectiveness of our approach. Junjun Jiang, Xianming Liu 0005, Jiayi Ma 0001 |
ICME | 3 |
| 2021 | Heatmap-Aware Pyramid Face HallucinationabstractRecent deep-learning-based face hallucination methods have achieved great success. Due to the parameter sharing characteristics of convolutional neural network, most existing deep-learning-based methods essentially use the same kernel for different regions of the entire face image in a convolution layer. This scheme of treating the face image as a whole will lead to the neglect of important facial details. To address this problem, we design a novel heatmap-aware convolution with spatially variant kernels rather than a spatially sharing kernel in the standard convolution to recover different regions. Based on this, we propose a heatmap-aware pyramid face super-resolution network (HaPSR) that embeds our heatmap-aware convolution into a two-branch network for both face super-resolution and facial heatmap estimation. The facial heatmap estimation branch can not only be used as an auxiliary to regularize face super-resolution reconstruction, but also provide an important basis for spatially variant kernels. Quantitative and qualitative experimental results demonstrate that our method outperforms state-of-the-arts. Chenyang Wang 0002, Junjun Jiang, Xianming Liu 0005 |
ICME | 3 |
| 2021 | Asymmetric Loss Functions for Learning with Noisy LabelsabstractRobust loss functions are essential for training deep neural networks with better generalization power in the presence of noisy labels. Symmetric loss functions are confirmed to be robust to label noise. However, the symmetric condition is overly restrictive. In this work, we propose a new class of loss functions, namely asymmetric loss functions, which are robust to learning from noisy labels for arbitrary noise type. Subsequently, we investigate general theoretical properties of asymmetric loss functions, including classification-calibration, excess risk bound, and noise-tolerance. Meanwhile, we introduce the asymmetry ratio to measure the asymmetry of a loss function, and the empirical results show that a higher ratio will provide better robustness. Moreover, we modify several common loss functions, and establish the necessary and sufficient conditions for them to be asymmetric. Experiments on benchmark datasets demonstrate that asymmetric loss functions can outperform state-of-the-art methods. Xianming Liu 0005, Junjun Jiang, Xiangyang Ji |
ICML | 2 |
| 2021 | Target-guided Adaptive Base Class Reweighting for Few-Shot LearningabstractFor few-shot learning, minimizing the empirical risk cannot reach the optimal hypothesis from image to its label due to the effect of overfitting. Therefore, most of the existing work leverages a set of base classes with sufficient labeled samples to pre-train a general encoder for feature representation, which is then applied for all few-shot classification tasks without considering the uniqueness of the target task. We suppose that different base classes help solve a target task in varying degrees, and some classes even introduce a negative effect. To this end, we propose a Target-guided Base Class Reweighting (TBR) approach, which uses a reweighting-in-the-loop optimization algorithm to assign a set of weights for base classes adaptively given a target task. Specifically, TBR learns the parameter of the encoder via minimizing weighted empirical risk on base class data, then optimizes the weights according to the the encoder's performance on support set of the target task. Such an alternating optimization procedure brings reweighting into the loop which makes the encoder more sensitive to the novel classes of the target task. Extensive experiments demonstrate that the proposed method can improve the performance of model-based approaches on two few-shot classification benchmarks. Jiliang Yan, Deming Zhai, Junjun Jiang, Xianming Liu 0005 |
ACM Multimedia | 4 |
| 2021 | Improving Hyperspectral Super-Resolution via Heterogeneous Knowledge DistillationabstractHyperspectral images (HSI) contains rich spectrum information but their spatial resolution is often limited by imaging system. Super-resolution (SR) reconstruction becomes a hot topic aiming to increase spatial resolution without extra hardware cost. The fusion-based hyperspectral image super-resolution (FHSR) methods use supplementary high-resolution multispectral images (HR-MSI) to recover spatial details, but well co-registered HR-MSI is hard to collect. Recently, single hyperspectral image super-resolution (SHSR) methods based on deep learning have made great progress. However, lack of HR-MSI input makes these SHSR methods difficult to exploit the spatial information. To take advantages of FHSR and SHSR methods, in this paper we propose a new pipeline treating HR-MSI as privilege information and try to improve our SHSR model with knowledge distillation. That is, our model uses paired MSI-HSI data to train and only needs LR-HSI as input during inference. Specifically, we combine SHSR and spectral super-resolution (SSR) and design a novel architecture, Distillation-Oriented Dual-branch Net (DODN), to make the SHSR model fully employ transferred knowledge from the SSR model. Since the main stream of SSR model are 2D CNNs and full 2D CNN causes spectral disorder in SHSR task, a new mixed 2D/3D block, called Distillation-Oriented Dual-branch Block (DODB) is proposed, where the 3D branch extracts spectral-spatial correlation while the 2D branch accepts information from the SSR model through knowledge distillation. The main idea is to distill the knowledge of spatial information from HR-MSI to the SHSR model without changing its network architecture. Extensive experiments on two benchmark datasets, CAVE and NTIRE2020, demonstrate that our proposed DODN outperforms the state-of-the-art SHSR methods, in terms of both quantitative and qualitative analysis. Junjun Jiang, Xianming Liu 0005 |
MMAsia | 4 |
| 2021 | NormalNet: Learning-Based Mesh Normal Denoising via Local Partition NormalizationabstractMesh denoising is a critical technology in geometry processing that aims to recover high-fidelity 3D mesh models of objects from noise-corrupted versions. In this work, we propose a learning-based mesh normal denoising scheme, calledNormalNet, which employs deep networks to find the correlation between the volumetric representation and denoised face normal. Overall,NormalNetfollows the iterative framework of filtering-based mesh denoising. During each iteration, firstly, a local partition normalization strategy is applied to split the local structure around each face into dense voxels, in which both the structure and face normal information can be preserved during this transformation. Benefiting from the thorough information preservation, we can use simple residual networks, which employ the volumetric representation as the input and produce the learned denoised face normal, to achieve satisfactory results. Finally, the vertex positions are updated according to the denoised normals. Besides introducing normalization into mesh denoising, our main contributions include a classification-based training faces selection strategy for balancing the training set and a mismatched-faces rejection strategy for removing the mismatched faces between noisy mesh and ground truth. Compared to state-of-the-art works,NormalNetcan effectively remove noise while preserving the original features and avoiding pseudo-features. Wenbo Zhao 0004, Xianming Liu 0005, Yongsen Zhao, Xiaopeng Fan 0001, Debin Zhao |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2021 | Graph-Based Feature-Preserving Mesh Normal FilteringabstractDistinguishing between geometric features and noise is of paramount importance for mesh denoising. In this paper, a graph-based feature-preserving mesh normal filtering scheme is proposed, which includes two stages: graph-based feature detection and feature-aware guided normal filtering. In the first stage, faces in the input noisy mesh are represented by patches, which are then modelled as weighted graphs. In this way, feature detection can be cast as a graph-cut problem. Subsequently, an iterative normalized cut algorithm is applied on each patch to separate the patch into smooth regions according to the detected features. In the second stage, a feature-aware guidance normal is constructed for each face, and guided normal filtering is applied to achieve robust feature-preserving mesh denoising. The results of experiments on synthetic and real scanned models indicate that the proposed scheme outperforms state-of-the-art mesh denoising works in terms of both objective and subjective evaluations. Wenbo Zhao 0004, Xianming Liu 0005, Shiqi Wang 0001, Xiaopeng Fan 0001, Debin Zhao |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2020 | Parsing Map Guided Multi-Scale Attention Network For Face HallucinationabstractFace hallucination that aims to transform a low-resolution (LR) face image to a high-resolution (HR) one is an active domain-specific image super-resolution problem. The performance of existing methods is usually not satisfactory, especially when the upscaling factor is large, such as 8×. In this paper, we propose an effective two- step face hallucination method based on a deep neural network with multi-scale channel and spatial attention mechanism. Specifically, we develop a ParsingNet to extract the prior knowledge of an input LR face, which is then fed into a carefully designed FishSRNet to recover the target HR face. Experimental results demonstrate that our method outperforms the state-of-the-arts in terms of quantitative metrics and visual quality. Chenyang Wang 0002, Zhiwei Zhong 0001, Junjun Jiang, Deming Zhai, Xianming Liu 0005 |
ICASSP | 5 |
| 2020 | ADRN: Attention-Based Deep Residual Network for Hyperspectral Image DenoisingabstractHyperspectral image (HSI) denoising is of crucial importance for many subsequent applications, such as HSI classification and interpretation. In this paper, we propose an attention-based deep residual network to directly learn a mapping from noisy HSI to the clean one. To jointly utilize the spatial-spectral information, the current band and its K adjacent bands are simultaneously exploited as the input. Then, we adopt convolution layer with different filter sizes to fuse the multi-scale feature, and use shortcut connection to incorporate the multi-level information for better noise removal. In addition, the channel attention mechanism is employed to make the network concentrate on the most relevant auxiliary information and features that are beneficial to the de-noising process best. To ease the training procedure, we reconstruct the output through a residual mode rather than a straightforward prediction. Experimental results demonstrate that our proposed ADRN scheme outperforms the state-of-the-art methods both in quantitative and visual evaluations. Yongsen Zhao, Deming Zhai, Junjun Jiang, Xianming Liu 0005 |
ICASSP | 4 |
| 2020 | Semi-Supervised Graph Convolutional Hashing Network For Large-Scale Cross-Modal RetrievalabstractCross-modal retrieval aims to provide flexible retrieval results across different types of multimedia data. To confront with scalability issue, binary codes learning (a.k.a. hash technique) is advocated since it permits exact top-K retrieval with sub-linear time complexity. In this paper, we propose a new method called Semi-supervised Graph Convolutional Hashing network (SGCH), which tries to learn a common hamming space by preserving both intra-modality and intermodality similarities via an end-to-end neural network. On one hand, graph convolutional network is utilized to explore high-order intra-modality similarity, and simultaneously propagate the semantic information from labeled samples to unlabeled data. On the other hand, a siamese network is connected to project the learnt features into a common hamming space. To bridge the inter-modality gap, adversarial loss which aims to learn modality-independent features by confusing a modality classifier is incorporated into the overall loss function. Experimental evaluations on cross-media retrieval tasks demonstrate that SGCH performs competitively against the state-of-the-art methods. Zhanjian Shen, Deming Zhai, Xianming Liu 0005, Junjun Jiang |
ICIP | 3 |
| 2020 | FasterSeg: Searching for Faster Real-time Semantic Segmentation
Wuyang Chen 0001, Xinyu Gong, Xianming Liu 0005, Zhangyang Wang |
ICLR | 3 |
| 2020 | Single Image Deraining via Scale-space Invariant Attention Neural NetworkabstractImage enhancement from degradation of rainy artifacts plays a critical role in outdoor visual computing systems. In this paper, we tackle the notion of scale that deals with visual changes in appearance of rain steaks with respect to the camera. Specifically, we revisit multi-scale representation by scale-space theory, and propose to represent the multi-scale correlation in convolutional feature domain, which is more compact and robust than that in pixel domain. Moreover, to improve the modeling ability of the network, we do not treat the extracted multi-scale features equally, but design a novel scale-space invariant attention mechanism to help the network focus on parts of the features. In this way, we summarize the most activated presence of feature maps as the salient features. Extensive experiments results on synthetic and real rainy scenes demonstrate the superior performance of our scheme over the state-of-the-arts. The source code of our method can be found in: https://github.com/pangbo1997/RainRemoval. Deming Zhai, Junjun Jiang, Xianming Liu 0005 |
ACM Multimedia | 4 |
| 2020 | Single-Image Blind Deblurring Using Multi-Scale Latent Structure PriorabstractBlind image deblurring is a challenging problem in computer vision, which aims to restore both the blur kernel and the latent sharp image from only a blurry observation. Inspired by the prevalent self-example prior in image super-resolution, in this paper, we observe that a coarse enough image down-sampled from a blurry observation is approximately a low-resolution version of the latent sharp image. We prove this phenomenon theoretically and define the coarse enough image as a latent structure prior of the unknown sharp image. Starting from this prior, we propose to restore sharp images from the coarsest scale to the finest scale on a blurry image pyramid and progressively update the prior image using the newly restored sharp image. These coarse-to-fine priors are referred to as multi-scale latent structures (MSLSs). Leveraging the MSLS prior, our algorithm comprises two phases: 1) we first preliminarily restore sharp images in the coarse scales and 2) we then apply a refinement process in the finest scale to obtain the final deblurred image. In each scale, to achieve lower computational complexity, we alternately perform a sharp image reconstruction with fast local self-example matching, an accelerated kernel estimation with error compensation, and a fast non-blind image deblurring, instead of computing any computationally expensive non-convex priors. We further extend the proposed algorithm to solve more challenging non-uniform blind image deblurring problem. The extensive experiments demonstrate that our algorithm achieves the competitive results against the state-of-the-art methods with much faster running speed. Yuanchao Bai, Huizhu Jia, Ming Jiang 0001, Xianming Liu 0005, Wen Gao 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2020 | Contrast Enhancement via Dual Graph Total Variation-Based Image DecompositionabstractImages captured in low lighting environment suffer from both low luminance contrast and noise corruption. However, most existing contrast enhancement algorithms only consider contrast boosting, which tends to reveal or amplify noise that is originally not visible in the dark areas. In this paper, we propose a joint contrast enhancement and denoising algorithm, which is based on structure/texture layer decomposition via minimization of dual forms of graph total variation (GTV). Specifically, the structure layer is expected to be generally smoothing but with sharp edges at the foreground background boundaries, for which we propose a quadratic form of GTV (QGTV) as the prior that promotes signal smoothness along graph structure. For the texture layer, a re-weighted GTV (RGTV) is tailored to noise removal while preserving true image details. We provide theoretical analysis about the filtering behavior of these two priors. Furthermore, a boost factor is derived per patch via optimal contrast-tone mapping to improve the overall brightness level of the patch. Finally, an optimization objective function is formulated, which casts image decomposition, brightness boosting, and noise reduction into a unified optimization framework. We further propose a fast approach to efficiently solve the optimization and provide analysis about the convergency. The experimental results show that the proposed method outperforms the state-of-the-art works in subjective, objective, and statistical quality evaluation. Xianming Liu 0005, Deming Zhai, Yuanchao Bai, Xiangyang Ji, Wen Gao 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2020 | Color-Guided Depth Image Recovery With Adaptive Data Fidelity and Transferred Graph Laplacian RegularizationabstractDepth images play an important role and are prevalently used in many computer vision and computational imaging tasks. However, due to the limitation of active sensing technology, the captured depth images in practice usually suffer from low resolution and noise, which prevents its further applications. To remedy this problem, in this paper, we first propose an adaptive data fidelity formulation to optimally generate each depth pixel from a mixture probability distribution, characterizing the similarity both in the depth map and the corresponding high-resolution guided color image. The proposed method is able to fit the distribution of the input depth signal as an optimization problem by maximizing the mixture probability. Furthermore, to promote the piecewise property that depth images exhibit, we propose a transferred graph Laplacian model as a regularization term, which is general and able to handle various depth recovery tasks such as super-resolution and denoising well. Specifically, each pixel within the recovered depth image is represented as a vertex in a graph with weights in connected edges representing the similarity between vertices. By minimizing the squared variations of the image signal, the task of depth image recovery can be converted to the problem of graph-based image filtering. Since the proposed graph Laplacian regularization model is able to fully exploit a priori information about the depth image, a much more accurate and robust estimation of the underlying depth can be obtained. Extensive experiment evaluations verify that the proposed method obtains recovered depth with higher quality in terms of both objective and subjective criteria, compared with most of the state-of-the-art methods. Yongbing Zhang 0002, Yihui Feng, Xianming Liu 0005, Deming Zhai, Xiangyang Ji, Haoqian Wang, Qionghai Dai |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2020 | Connecting Image Denoising and High-Level Vision Tasks via Deep LearningabstractImage denoising and high-level vision tasks are usually handled independently in the conventional practice of computer vision, and their connection is fragile. In this paper, we cope with the two jointly and explore the mutual influence between them with the focus on two questions, namely (1) how image denoising can help improving high-level vision tasks, and (2) how the semantic information from high-level vision tasks can be used to guide image denoising. First for image denoising we propose a convolutional neural network in which convolutions are conducted in various spatial resolutions via downsampling and upsampling operations in order to fuse and exploit contextual information on different scales. Second we propose a deep neural network solution that cascades two modules for image denoising and various high-level tasks, respectively, and use the joint loss for updating only the denoising network via backpropagation. We experimentally show that on one hand, the proposed denoiser has the generality to overcome the performance degradation of different high-level vision tasks. On the other hand, with the guidance of high-level vision information, the denoising network produces more visually appealing results. Extensive experiments demonstrate the benefit of exploiting image semantics simultaneously for image denoising and highlevel vision tasks via deep learning. The code is available online: https://github.com/Ding-Liu/DeepDenoising. Ding Liu 0001, Bihan Wen, Jianbo Jiao, Xianming Liu 0005, Zhangyang Wang, Thomas S. Huang |
IEEE Trans. Image Process. | 4 |
| 2019 | Reconstruction-cognizant Graph Sampling Using Gershgorin Disc AlignmentabstractGraph sampling with noise is a fundamental problem in graph signal processing (GSP). Previous works assume an unbiased least square (LS) signal reconstruction scheme and select samples greedily via expensive extreme eigenvector computation. A popular biased scheme using graph Laplacian regularization (GLR) solves a system of linear equations for its reconstruction. Assuming this GLR-based scheme, we propose a reconstruction-cognizant sampling strategy to maximize the numerical stability of the linear system-i.e., minimize the condition number of the coefficient matrix. Specifically, we maximize the eigenvalue lower bounds of the matrix, represented by left-ends of Gershgorin discs of the coefficient matrix. To accomplish this efficiently, we propose an iterative algorithm to traverse the graph nodes via Breadth First Search (BFS) and align the left-ends of all corresponding Gershgorin discs at lower-bound threshold T using two basic operations: disc shifting and scaling. We then perform binary search to maximize T given a sample budget K. Experiments on real graph data show that the proposed algorithm can effectively promote large eigenvalue lower bounds, and the reconstruction MSE is the same or smaller than existing sampling methods for different budget K at much lower complexity. Yuanchao Bai, Gene Cheung, Xianming Liu 0005, Wen Gao 0001 |
ICASSP | 4 |
| 2019 | Hyperspectral Image Classification in the Presence of Noisy LabelsabstractLabel information plays an important role in a supervised hyperspectral image classification problem. However, current classification methods all ignore an important and inevitable problem-labels may be corrupted and collecting clean labels for training samples is difficult and often impractical. Therefore, how to learn from the database with noisy labels is a problem of great practical importance. In this paper, we study the influence of label noise on hyperspectral image classification and develop a random label propagation algorithm (RLPA) to cleanse the label noise. The key idea of RLPA is to exploit knowledge (e.g., the superpixel-based spectral-spatial constraints) from the observed hyperspectral images and apply it to the process of label propagation. Specifically, the RLPA first constructs a spectral-spatial probability transform matrix (SSPTM) that simultaneously considers the spectral similarity and superpixel-based spatial information. It then randomly chooses some training samples as “clean” samples and sets the rest as unlabeled samples, and propagates the label information from the “clean” samples to the rest unlabeled samples with the SSPTM. By repeating the random assignment (of “clean” labeled samples and unlabeled samples) and propagation, we can obtain multiple labels for each training sample. Therefore, the final propagated label can be calculated by a majority vote algorithm. Experimental studies show that the RLPA can reduce the level of noisy label and demonstrates the advantages of our proposed method over four major classifiers with a significant margin-the gains in terms of the average overall accuracy, average accuracy, and kappa are impressive, e.g., 9.18%, 9.58%, and 0.1043. The MATLAB source code is available at https://github.com/junjun-jiang/RLPA. Junjun Jiang, Jiayi Ma 0001, Zheng Wang 0007, Chen Chen 0001, Xianming Liu 0005 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2019 | Graph-Based Blind Image Deblurring From a Single PhotographabstractBlind image deblurring, i.e., deblurring without knowledge of the blur kernel, is a highly ill-posed problem. The problem can be solved in two parts: i) estimate a blur kernel from the blurry image, and ii) given an estimated blur kernel, de-convolve the blurry input to restore the target image. In this paper, we propose a graph-based blind image deblurring algorithm by interpreting an image patch as a signal on a weighted graph. Specifically, we first argue that a skeleton image-a proxy that retains the strong gradients of the target but smooths out the details-can be used to accurately estimate the blur kernel and has a unique bi-modal edge weight distribution. Then, we design a reweighted graph total variation (RGTV) prior that can efficiently promote a bi-modal edge weight distribution given a blurry patch. Further, to analyze RGTV in the graph frequency domain, we introduce a new weight function to represent RGTV as a graph l1-Laplacian regularizer. This leads to a graph spectral filtering interpretation of the prior with desirable properties, including robustness to noise and blur, strong piecewise smooth (PWS) filtering and sharpness promotion. Minimizing a blind image deblurring objective with RGTV results in a non-convex non-differentiable optimization problem. Leveraging the new graph spectral interpretation for RGTV, we design an efficient algorithm that solves for the skeleton image and the blur kernel alternately. Specifically for Gaussian blur, we propose a further speedup strategy for blind Gaussian deblurring using accelerated graph spectral filtering. Finally, with the computed blur kernel, recent non-blind image deblurring algorithms can be applied to restore the target image. Experimental results demonstrate that our algorithm successfully restores latent sharp images and outperforms state-of-the-art methods quantitatively and qualitatively. Yuanchao Bai, Gene Cheung, Xianming Liu 0005, Wen Gao 0001 |
IEEE Trans. Image Process. | 3 |
| 2019 | Graph-Regularized Locality-Constrained Joint Dictionary and Residual Learning for Face Sketch SynthesisabstractFace sketch synthesis is a crucial issue in digital entertainment and law enforcement. It can bridge the considerable texture discrepancy between face photos and sketches. Most of the current face sketch synthesis approaches directly to learn the relationship between the photos and sketches, and it is very difficult for them to generate the individual specific features, which we call rare characteristics. In this paper, we propose a novel face sketch synthesis approach through residual learning. In contrast to traditional approaches, which aim to reconstruct a sketch image directly (i.e., learn the mapping relationship between the photo and sketch), we aim to predict the residual image by learning the mapping relationship between the photo and residual, i.e., the difference between the photo and sketch, given an observed photo. This technique will render optimizing the residual mapping easier than optimizing the original mapping and deriving rare characteristic information. We also introduce a joint dictionary learning algorithm by preserving the local geometry structure of a data space. Through the learned joint dictionary, we transform the face sketch synthesis from an image space to a new and compact space; the new and compact space is spanned by learned dictionary atoms, where the manifold assumption can be further guaranteed. Results show that the proposed method demonstrates an impressive performance in the face sketch synthesis task on three public face sketch datasets and various real-world photos. These results are derived by comparing the proposed method with several state-of-the-art techniques, including certain recently proposed deep learning-based approaches. Junjun Jiang, Yi Yu 0001, Zheng Wang 0007, Xianming Liu 0005, Jiayi Ma 0001 |
IEEE Trans. Image Process. | 4 |
| 2019 | Graph-Based Joint Dequantization and Contrast Enhancement of Poorly Lit JPEG ImagesabstractJPEG images captured in poor lighting conditions suffer from both low luminance contrast and coarse quantization artifacts due to lossy compression. Performing dequantization and contrast enhancement in separate back-to-back steps would amplify the residual compression artifacts, resulting in low visual quality. Leveraging on recent development in graph signal processing (GSP), we propose to jointly dequantize and contrast-enhance such images in a single graph-signal restoration framework. Specifically, we separate each observed pixel patch into illumination and reflectance via Retinex theory, where we define generalized smoothness prior and signed graph smoothness prior according to their respective unique signal characteristics. Given only a transform-coded image patch, we compute robust edge weights for each graph via low-pass filtering in the dual graph domain. We compute the illumination and reflectance components for each patch alternately, adopting accelerated proximal gradient (APG) algorithms in the transform domain, with backtracking line search for further speedup. Experimental results show that our generated images outperform the state-of-the-art schemes noticeably in the subjective quality evaluation. Xianming Liu 0005, Gene Cheung, Xiangyang Ji, Debin Zhao, Wen Gao 0001 |
IEEE Trans. Image Process. | 1 |
| 2019 | Depth Restoration From RGB-D Data via Joint Adaptive Regularization and Thresholding on ManifoldsabstractIn this paper, we propose a novel depth restoration algorithm from RGB-D data through combining characteristics of local and non-local manifolds, which provide low-dimensional parameterizations of the local and non-local geometry of depth maps. Specifically, on the one hand, a local manifold model is defined to favor local neighboring relationship of pixels in depth, according to which, manifold regularization is introduced to promote smoothing along the manifold structure. On the other hand, the non-local characteristics of the patch-based manifold can be used to build highly data-adaptive orthogonal bases to extract elongated image patterns, accounting for self-similar structures in the manifold. We further define a manifold thresholding operator in 3D adaptive orthogonal spectral bases-eigenvectors of the discrete Laplacian of local and non-local manifolds-to retain only low graph frequencies for depth maps restoration. Finally, we propose a unified alternating direction method of multipliers optimization framework, which elegantly casts the adaptive manifold regularization and thresholding jointly to regularize the inverse problem of depth maps recovery. Experimental results demonstrate that our method achieves superior performance compared with the state-of-the-art works with respect to both objective and subjective quality evaluations. Xianming Liu 0005, Deming Zhai, Xiangyang Ji, Debin Zhao, Wen Gao 0001 |
IEEE Trans. Image Process. | 1 |
| 2019 | Depth Super-Resolution via Joint Color-Guided Internal and External RegularizationsabstractDepth information is being widely used in many real-world applications. However, due to the limitation of depth sensing technology, the captured depth map in practice usually has much lower resolution than that of color image counterpart. In this paper, we propose to combine the internal smoothness prior and external gradient consistency constraint in graph domain for depth super-resolution. On one hand, a new graph Laplacian regularizer is proposed to preserve the inherent piecewise smooth characteristic of depth, which has desirable filtering properties. A specific weight matrix of the respect graph is defined to make full use of information of both depth and the corresponding guidance image. On the other hand, inspired by an observation that the gradient of depth is small except at edge separating regions, we introduce a graph gradient consistency constraint to enforce that the graph gradient of depth is close to the thresholded gradient of guidance. We reinterpret the gradient thresholding model as variational optimization with sparsity constraint. In this way, we remedy the problem of structure discrepancy between depth and guidance. Finally, the internal and external regularizations are casted into a unified optimization framework, which can be efficiently addressed by ADMM. Experimental results demonstrate that our method outperforms the state-of-the-art with respect to both objective and subjective quality evaluations. Xianming Liu 0005, Deming Zhai, Xiangyang Ji, Debin Zhao, Wen Gao 0001 |
IEEE Trans. Image Process. | 1 |
| 2018 | Compressed Image Restoration via External-Image Assisted Band Adaptive PCA Model LearningabstractVisually annoying compression artifacts frequently appear in block-based transform coding at low bit rates, due to coarse and independent quantization of transform coefficients in coding blocks. This paper presents a subband adaptive modeling framework for reducing quantization artifacts. In this framework, each patch is jointly regularized by bandwise distribution priors adaptively learned in its PCA transform domain together with a quantization constraint prior in the DCT domain. Since the compression artifacts influence the covariance statistics of coded image patches remarkably, external images are utilized to provide more robust PCA domains for patch sparse modeling. Instead of using a global distribution model for all patches, the distribution prior of each patch is adaptively learned from similar patches within the compressed image itself to address the non-stationarity of image signals. The coefficients in different PCA bands are regularized unequally according to the learned priors. Experimental results show that the proposed scheme outperforms existing schemes in terms of both the objective and the perceptual qualities. Ruiqin Xiong, Xiaopeng Fan 0001, Xianming Liu 0005, Tiejun Huang 0001, Wen Gao 0001 |
DCC | 4 |
| 2018 | Blind Image Deblurring Via Reweighted Graph Total VariationabstractBlind image deblurring, i.e., deblurring without knowledge of the blur kernel, is a highly ill-posed problem. The problem can be solved in two parts: i) estimate a blur kernel from the blurry image, and ii) given estimated blur kernel, de-convolve blurry input to restore the target image. In this paper, by interpreting an image patch as a signal on a weighted graph, we first argue that a skeleton image-a proxy that retains the strong gradients of the target but smooths out the details-can be used to accurately estimate the blur kernel and has a unique bi-modal edge weight distribution. We then design a reweighted graph total variation (RGTV) prior that can efficiently promote bi-modal edge weight distribution given a blurry patch. However, minimizing a blind image deblurring objective with RGTV results in a non-convex non-differentiable optimization problem. We propose a fast algorithm that solves for the skeleton image and the blur kernel alternately. Finally with the computed blur kernel, recent non-blind image deblurring algorithms can be applied to restore the target image. Experimental results show that our algorithm can robustly estimate the blur kernel with large kernel size, and the reconstructed sharp image is competitive against the state-of-the-art methods. Yuanchao Bai, Gene Cheung, Xianming Liu 0005, Wen Gao 0001 |
ICASSP | 3 |
| 2018 | A Blind Quality Measure for Industrial 2D Matrix Symbols Using Shallow Convolutional Neural NetworkabstractIndustrial two-dimensional (2D) matrix symbols are ubiquitous throughout the automatic assembly lines. Most industrial 2D symbols are corrupted by various inevitable artifacts. State-of-the-art decoding algorithms are not able to directly handle low-quality symbols irrespective of problematic artifacts. Degraded symbols require appropriate preprocessing methods, such as morphology filtering, median filtering, or sharpening filtering, according to specific distortion type. In this paper, we first establish a database including 3000 industrial 2D symbols which are degraded by 6 types of distortions. Second, we utilize a shallow convolutional neural network (CNN) to identify the distortion type and estimate the quality grade for 2D symbols. Finally, we recommend an appropriate preprocessing method for low-quality symbol according to its distortion type and quality grade. Experimental results indicate that the proposed method outperforms state-of-the-art methods in terms of PLCC, SRCC and RMSE. It also promotes decoding efficiency at the cost of low extra time spent. Zhaohui Che, Guangtao Zhai, Jing Liu 0002, Ke Gu 0001, Patrick Le Callet, Jiantao Zhou 0001, Xianming Liu 0005 |
ICIP | 7 |
| 2018 | When Image Denoising Meets High-Level Vision Tasks: A Deep Learning ApproachabstractConventionally, image denoising and high-level vision tasks are handled separately in computer vision. In this paper, we cope with the two jointly and explore the mutual influence between them. First we propose a convolutional neural network for image denoising which achieves the state-of-the-art performance. Second we propose a deep neural network solution that cascades two modules for image denoising and various high-level tasks, respectively, and use the joint loss for updating only the denoising network via back-propagation. We demonstrate that on one hand, the proposed denoiser has the generality to overcome the performance degradation of different high-level vision tasks. On the other hand, with the guidance of high-level vision information, the denoising network can generate more visually appealing results. To the best of our knowledge, this is the first work investigating the benefit of exploiting image semantics simultaneously for image denoising and high-level vision tasks via deep learning. Ding Liu 0001, Bihan Wen, Xianming Liu 0005, Zhangyang Wang, Thomas S. Huang |
IJCAI | 3 |
| 2018 | Adaptive Screen Content Image Enhancement Strategy using Layer-based SegmentationabstractThe ubiquitous screen content images (SCIs) play a significant role in various scenarios currently. However, most SCIs captured by consumer devices are frequently corrupted with distortions, especially contrast distortion. Unlike the natural images, SCIs are composed of text, graphics and natural scene pictures so that traditional image enhancement methods are not suitable for these compound images. Therefore, we innovatively proposed an adaptive strategy for enhancing SCIs in this paper. Firstly, we devised a segmentation method to divide SCI into text and pictorial regions. Next, the famous guided image filter (GIF) with big and small kernel sizes served as unsharpness masking for processing different regions adaptively. For verifying performance, the proposed method was tested on recently prevalent SCI datasets including SIQAD, and Webpage Dataset. Experimental results indicate that the proposed approach outperforms state-of-the-art methods in most SCIs with flat background. Zhaohui Che, Guangtao Zhai, Ke Gu 0001, Patrick Le Callet, Xianming Liu 0005, Deming Zhai, Xiao Gu 0001 |
ISCAS | 5 |
| 2018 | Parametric local multiview hamming distance metric learning
Deming Zhai, Xianming Liu 0005, Hong Chang 0001, Yi Zhen, Xilin Chen 0001, Maozu Guo 0001, Wen Gao 0001 |
Pattern Recognit. | 2 |
| 2018 | Prior-Based Quantization Bin Matching for Cloud Storage of JPEG ImagesabstractMillions of user-generated images are uploaded to social media sites like Facebook daily, which translate to a large storage cost. However, there exists an asymmetry in upload and download data: only a fraction of the uploaded images are subsequently retrieved for viewing. In this paper, we propose a cloud storage system that reduces the storage cost of all uploaded JPEG photos, at the expense of a controlled increase in computation mainly during download of requested image subset. Specifically, the system first selectively re-encodes code blocks of uploaded JPEG images using coarser quantization parameters for smaller storage sizes. Then during download, the system exploits known signal priors-sparsity prior and graph-signal smoothness prior-for reverse mapping to recover original fine quantization bin indices, with either deterministic guarantee (lossless mode) or statistical guarantee (near-lossless mode). For fast reverse mapping, we use small dictionaries and sparse graphs that are tailored for specific clusters of similar blocks, which are classified via tree-structured vector quantizer. During image upload, cluster indices identifying the appropriate dictionaries and graphs for the re-quantized blocks are encoded as side information using a differential distributed source coding scheme to facilitate reverse mapping during image download. Experimental results show that our system can reap significant storage savings (up to 12.05%) at roughly the same image PSNR (within 0.18 dB). Xianming Liu 0005, Gene Cheung, Chia-Wen Lin, Debin Zhao, Wen Gao 0001 |
IEEE Trans. Image Process. | 1 |
| 2018 | Learning Temporal Dynamics for Video Super-Resolution: A Deep Learning ApproachabstractVideo super-resolution (SR) aims at estimating a high-resolution (HR) video sequence from a low-resolution (LR) one. Given that deep learning has been successfully applied to the task of single image SR, which demonstrates the strong capability of neural networks for modeling spatial relation within one single image, the key challenge to conduct video SR is how to efficiently and effectively exploit the temporal dependency among consecutive LR frames other than the spatial relation. However, this remains challenging because complex motion is difficult to model and can bring detrimental effects if not handled properly. We tackle the problem of learning temporal dynamics from two aspects. First, we propose a temporal adaptive neural network that can adaptively determine the optimal scale of temporal dependency. Inspired by the Inception module in GoogLeNet [1], filters of various temporal scales are applied to the input LR sequence before their responses are adaptively aggregated, in order to fully exploit the temporal relation among consecutive LR frames. Second, we decrease the complexity of motion among neighboring frames using a spatial alignment network that can be end-to-end trained with the temporal adaptive network and has the merit of increasing the robustness to complex motion and the efficiency compared to competing image alignment methods. We provide a comprehensive evaluation of the temporal adaptation and the spatial alignment modules. We show the temporal adaptive design considerably improve SR quality over its plain counterparts, and the spatial alignment network is able to attain comparable SR performance with the sophisticated optical flow based approach, but requires much less running time. Overall our proposed model with learned temporal dynamics is shown to achieve state-of-the-art SR results in terms of not only spatial consistency but also temporal coherence on public video datasets. More information can be found in. Ding Liu 0001, Yuchen Fan 0001, Xianming Liu 0005, Zhangyang Wang, Shiyu Chang, Xinchao Wang, Thomas S. Huang |
IEEE Trans. Image Process. | 4 |
| 2018 | Reduced-Reference Image Quality Assessment in Free-Energy Principle and Sparse RepresentationabstractThe free-energy principle in recent studies of brain theory and neuroscience models the perception and understanding of the outside scene as an active inference process, in which the brain tries to account for the visual scene with an internal generative model. Specifically, with the internal generative model, the brain yields corresponding predictions for its encountered visual scenes. Then, the discrepancy between the visual input and its brain prediction should be closely related to the quality of perceptions. On the other hand, sparse representation has been evidenced to resemble the strategy of the primary visual cortex in the brain for representing natural images. With the strong neurobiological support for sparse representation, in this paper, we approximate the internal generative model with sparse representation and propose an image quality metric accordingly, which is named FSI (free-energy principle and sparse representation-based index for image quality assessment). In FSI, the reference and distorted images are, respectively, predicted by the sparse representation at first. Then, the difference between the entropies of the prediction discrepancies is defined to measure the image quality. Experimental results on four large-scale image databases confirm the effectiveness of the FSI and its superiority over representative image quality assessment methods. The FSI belongs to reduced-reference methods, and it only needs a single number from the reference image for quality estimation. Yutao Liu 0002, Guangtao Zhai, Ke Gu 0001, Xianming Liu 0005, Debin Zhao, Wen Gao 0001 |
IEEE Trans. Multim. | 4 |
| 2018 | Supervised Distributed Hashing for Large-Scale Multimedia RetrievalabstractRecent years have witnessed the growing popularity of hashing for large-scale multimedia retrieval. Extensive hashing methods have been designed for data stored in a single machine, that is, centralized hashing . In many real-world applications, however, the large-scale data are often distributed across different locations, servers, or sites. Although hashing for distributed data can be implemented by assembling all distributed data together as a whole dataset in theory, it usually leads to prohibitive computation, communication, and storage costs in practice. Up to now, only a few methods were tailored for distributed hashing, which are all unsupervised approaches. In this paper, we propose an efficient and effective method called supervised distributed hashing (SupDisH), which learns discriminative hash functions by leveraging the semantic label information in a distributed manner. Specifically, we cast the distributed hashing problem into the framework of classification, where the learned binary codes are expected to be distinct enough for semantic retrieval. By introducing auxiliary variables, the distributed model is then separated into a set of decentralized subproblems with consistency constraints, which can be solved in parallel on each vertex of the distributed network. As such, we can obtain high-quality distinctive unbiased binary codes and consistent hash functions with low computational complexity, which facilitate tackling large-scale multimedia retrieval tasks involving distributed datasets. Experimental evaluations on three large-scale datasets show that SupDisH is competitive to centralized hashing methods and outperforms the state-of-the-art unsupervised distributed method significantly. Deming Zhai, Xianming Liu 0005, Xiangyang Ji, Debin Zhao, Shin'ichi Satoh 0001, Wen Gao 0001 |
IEEE Trans. Multim. | 2 |
| 2017 | Robust Video Super-Resolution with Learned Temporal DynamicsabstractVideo super-resolution (SR) aims to generate a high-resolution (HR) frame from multiple low-resolution (LR) frames in a local temporal window. The inter-frame temporal relation is as crucial as the intra-frame spatial relation for tackling this problem. However, how to utilize temporal information efficiently and effectively remains challenging since complex motion is difficult to model and can introduce adverse effects if not handled properly. We address this problem from two aspects. First, we propose a temporal adaptive neural network that can adaptively determine the optimal scale of temporal dependency. Filters on various temporal scales are applied to the input LR sequence before their responses are adaptively aggregated. Second, we reduce the complexity of motion between neighboring frames using a spatial alignment network which is much more robust and efficient than competing alignment methods and can be jointly trained with the temporal adaptive network in an end-to-end manner. Our proposed models with learned temporal dynamics are systematically evaluated on public video datasets and achieve state-of-the-art SR results compared with other recent video SR approaches. Both of the temporal adaptation and the spatial alignment modules are demonstrated to considerably improve SR quality over their plain counterparts. Ding Liu 0001, Yuchen Fan 0001, Xianming Liu 0005, Zhangyang Wang, Shiyu Chang, Thomas S. Huang |
ICCV | 4 |
| 2017 | Single depth image super-resolution and denoising based on sparse graphs via structure tensorabstractThe existing single depth image super-resolution (SR) methods suppose that the image to be interpolated is noise free. However, the supposition is invalid in practice because noise will be inevitably introduced in the depth image acquisition process. In this paper, we address the problem of image denoising and SR jointly based on designing sparse graphs that are useful for describing the geometric structures of data domains. In our method, we first cluster similar patches in a noisy depth image and compute an average patch. Different from the majority of the graph Fourier transform (GFT) that assumed an underlying 4-connected graph structure with vertical and horizontal edges only, we select more general sparse graph structures and edges weights based on the difference of the blocks' structure tensors. For the average patch, a graph template with edges orthogonal to the principal gradient is designed. Finally, the graph based transform (GBT) dictionary is learned from the derived correlation graph for signal representation. As shown in our experimental results, the proposed method obtains a lot of improvement in performance. Yihui Feng, Xianming Liu 0005, Yongbing Zhang 0002, Qionghai Dai |
ICIP | 2 |
| 2017 | Dynamic backlight scaling considering ambient luminance for mobile energy savingabstractThe mobile video playback involves many subsystems of the devices such as computing, rendering and displaying subsystems. Among all subsystems, the displaying subsystem accounts for at least 38% of all consumed power, and it can be up to 68% with the maximum backlight brightness. What is more, lots of people watch videos via mobile devices in various situations, where the ambient luminance condition is different. Therefore, how to save mobile energy and improve the Quality of Experience (QoE) in different situations become significant problems. In this paper, we try to maximally enhance the battery power performance under various ambient luminance conditions through backlight magnitude adjusting, while without negatively impacting users' QoE. In particular, we conduct a series of subject quality assessment experiments to uncover the quantitative relationship among QoE, ambient luminance, video content luminance and backlight level. We first study whether the continuous playback of backlight-scaled shots using the proposed scaling magnitude would cause flicker effect or not. Then motivated by the findings of these subject studies, we implement a Dynamic Backlight Scaling (DBS) strategy. The experiment results demonstrate that the DBS strategy can save more than 40% power at most and can also save 10% power even at a very high ambient luminance. Wei Sun 0029, Guangtao Zhai, Xiongkuo Min, Yutao Liu 0002, Siwei Ma 0001, Jing Liu 0002, Jiantao Zhou 0001, Xianming Liu 0005 |
ICME | 8 |
| 2017 | Quality assessment for real out-of-focus blurred images
Yutao Liu 0002, Ke Gu 0001, Guangtao Zhai, Xianming Liu 0005, Debin Zhao, Wen Gao 0001 |
J. Vis. Commun. Image Represent. | 4 |
| 2017 | Random Walk Graph Laplacian-Based Smoothness Prior for Soft Decoding of JPEG ImagesabstractGiven the prevalence of joint photographic experts group (JPEG) compressed images, optimizing image reconstruction from the compressed format remains an important problem. Instead of simply reconstructing a pixel block from the centers of indexed discrete cosine transform (DCT) coefficient quantization bins (hard decoding), soft decoding reconstructs a block by selecting appropriate coefficient values within the indexed bins with the help of signal priors. The challenge thus lies in how to define suitable priors and apply them effectively. In this paper, we combine three image priors-Laplacian prior for DCT coefficients, sparsity prior, and graph-signal smoothness prior for image patches-to construct an efficient JPEG soft decoding algorithm. Specifically, we first use the Laplacian prior to compute a minimum mean square error initial solution for each code block. Next, we show that while the sparsity prior can reduce block artifacts, limiting the size of the overcomplete dictionary (to lower computation) would lead to poor recovery of high DCT frequencies. To alleviate this problem, we design a new graph-signal smoothness prior (desired signal has mainly low graph frequencies) based on the left eigenvectors of the random walk graph Laplacian matrix (LERaG). Compared with the previous graph-signal smoothness priors, LERaG has desirable image filtering properties with low computation overhead. We demonstrate how LERaG can facilitate recovery of high DCT frequencies of a piecewise smooth signal via an interpretation of low graph frequency components as relaxed solutions to normalized cut in spectral clustering. Finally, we construct a soft decoding algorithm using the three signal priors with appropriate prior weights. Experimental results show that our proposal outperforms the state-of-the-art soft decoding algorithms in both objective and subjective evaluations noticeably. Xianming Liu 0005, Gene Cheung, Xiaolin Wu 0001, Debin Zhao |
IEEE Trans. Image Process. | 1 |
| 2017 | Sparsity-Based Image Error Concealment via Adaptive Dual Dictionary Learning and RegularizationabstractIn this paper, we propose a novel sparsity-based image error concealment (EC) algorithm through adaptive dual dictionary learning and regularization. We define two feature spaces: the observed space and the latent space, corresponding to the available regions and the missing regions of image under test, respectively. We learn adaptive and complete dictionaries individually for each space, where the training data are collected via an adaptive template matching mechanism. Based on the piecewise stationarity of natural images, a local correlation model is learned to bridge the sparse representations of the aforementioned dual spaces, allowing us to transfer the knowledge of the available regions to the missing regions for EC purpose. Eventually, the EC task is formulated as a unified optimization problem, where the sparsity of both spaces and the learned correlation model are incorporated. Experimental results show that the proposed method outperforms the state-of-the-art techniques in terms of both objective and perceptual metrics. Xianming Liu 0005, Deming Zhai, Jiantao Zhou 0001, Shiqi Wang 0001, Debin Zhao, Huijun Gao |
IEEE Trans. Image Process. | 1 |
| 2017 | Greedy Batch-Based Minimum-Cost Flows for Tracking Multiple ObjectsabstractMinimum-cost flow algorithms have recently achieved state-of-the-art results in multi-object tracking. However, they rely on the whole image sequence as input. When deployed in real-time applications or in distributed settings, these algorithms first operate on short batches of frames and then stitch the results into full trajectories. This decoupled strategy is prone to errors because the batch-based tracking errors may propagate to the final trajectories and cannot be corrected by other batches. In this paper, we propose a greedy batch-based minimum-cost flow approach for tracking multiple objects. Unlike existing approaches that conduct batch-based tracking and stitching sequentially, we optimize consecutive batches jointly so that the tracking results on one batch may benefit the results on the other. Specifically, we apply a generalized minimum-cost flows (MCF) algorithm on each batch and generate a set of conflicting trajectories. These trajectories comprise the ones with high probabilities, but also those with low probabilities potentially missed by detectors and trackers. We then apply the generalized MCF again to obtain the optimal matching between trajectories from consecutive batches. Our proposed approach is simple, effective, and does not require training. We demonstrate the power of our approach on data sets of different scenarios. Xinchao Wang, Bin Fan 0001, Shiyu Chang, Zhangyang Wang, Xianming Liu 0005, Dacheng Tao, Thomas S. Huang |
IEEE Trans. Image Process. | 5 |
| 2017 | Utility-Driven Adaptive Preprocessing for Screen Content Video CompressionabstractIn this work, we propose a utility-driven preprocessing technique for high-efficiency screen content video (SCV) compression based on the temporal masking effect, which was found to be a fundamental attribute that plays an important role in human visual perception of video quality, but has not been fully exploited in the context of SCV coding. Specifically, we investigate the temporal masking effect from the perspective of perceived utility, which allows us to preserve the quality of the high utility content and substitute the low utility regions with the corresponding smooth version. To distinguish the regional utilities, a specifically designed block type identification algorithm for screen content is employed to measure the local properties. Subsequently, the Gaussian filter is applied to smooth out the high-frequency components in the detected low utility regions to save consumption bits. Validations based on subjective testings show that the proposed approach is capable of achieving significant bitrate savings with little sacrifice on the final utility compared with the conventional SCV coding scheme. Shiqi Wang 0001, Xinfeng Zhang 0001, Xianming Liu 0005, Jian Zhang 0018, Siwei Ma 0001, Wen Gao 0001 |
IEEE Trans. Multim. | 3 |
| 2016 | Quantization bin matching for cloud storage of JPEG imagesabstractSocial media sites like Facebook are obligated to store all photos uploaded by an ever growing user base-which translates to an increasingly expensive storage cost-but only a fraction of uploaded images are revisited thereafter. In this paper, we propose a cloud storage system that trades off computation of a small fraction of requested images with storage of all photos. The key idea is to re-encode uploaded JPEG photos with coarser quantization parameters (QP) for permanent storage, then exploit a signal sparsity prior during inverse mapping to recover fine quantization bin indices via a maximum a posteriori (MAP) formulation. Because by design the system guarantees recovery of an original compressed image (either with exactly the same input fine quantization bin indices or has visual quality indistinguishable by human eyes), from the user's viewpoint it is a normal cloud storage, while from the operator's viewpoint there is pure compression gain and hence lower storage cost. Experimental results show that our storage system can reap significant storage savings (up to 20%) at roughly the same image PSNR (within 0.13dB). Xianming Liu 0005, Gene Cheung, Chia-Wen Lin, Debin Zhao |
ICASSP | 1 |
| 2016 | Blind quality assessment of compressed images via pseudo structural similarityabstractBlock-based compression causes severe pseudo structures. We find that the pseudo structures of images compressed by different levels show some degree of similarity. So we propose to evaluate the quality of compressed images via the similarity between pseudo structures of two images. To obtain a “reference” image, we introduce the most distorted image (MDI), which is derived from the distorted image and suffers from the highest degree of compression. The proposed pseudo structural similarity (PSS) model calculates the similarity between pseudo structures of the distorted image and MDI. Pseudo structures of the distorted image become similar to the MDI's under the condition of severe compression. Via comparative tests, the proposed PSS model, on one hand, is shown to be comparable to state-of-the-art competitors, and on the other hand, it is not only good at assessing natural scene images but also performs the best in the hotly-researched screen content image (SCI) database. It deserves to mention that PSS is able to boost the performance of mainstream general-purpose no-reference (NR) quality measures. Xiongkuo Min, Guangtao Zhai, Ke Gu 0001, Yuming Fang 0001, Xiaokang Yang 0001, Xiaolin Wu 0001, Jiantao Zhou 0001, Xianming Liu 0005 |
ICME | 8 |
| 2016 | Perceptual image quality assessment combining free-energy principle and sparse representationabstractSince the purpose of objective image quality assessment is to be consistent with subjective image quality assessment as highly as possible, the understanding of the mechanisms of human visual system will certainly benefit the study of objective image quality assessment. Recent developments in brain theory and neuroscience, particularly the free-energy principle, account for the perception and understanding of visual scenes. As the free-energy principle conjectures, the brain tries to generate the corresponding prediction for its encountered scene by an internal generative model. On the other hand, sparse representation is evidenced to resemble the neural response properties of simple cells in the primary visual cortex. Conjunctively, in this paper, we suppose the prediction manner of the internal generative model in free-energy principle follows sparse representation and propose an image quality metric accordingly. Experiments on LIVE, TID2008 and CSIQ image databases demonstrate the effectiveness of the proposed image quality metric. Noteworthily, our metric needs little information (only a single scalar) of the reference image and is training-free. Yutao Liu 0002, Guangtao Zhai, Xianming Liu 0005, Debin Zhao |
ISCAS | 3 |
| 2016 | Secure Reversible Image Data Hiding Over Encrypted Domain via Key ModulationabstractThis paper proposes a novel reversible image data hiding scheme over encrypted domain. Data embedding is achieved through a public key modulation mechanism, in which access to the secret encryption key is not needed. At the decoder side, a powerful two-class SVM classifier is designed to distinguish encrypted and nonencrypted image patches, allowing us to jointly decode the embedded message and the original image signal. Compared with the state-of-the-art methods, the proposed approach provides higher embedding capacity and is able to perfectly reconstruct the original image as well as the embedded message. Extensive experimental results are provided to validate the superior performance of our scheme. Jiantao Zhou 0001, Weiwei Sun 0009, Li Dong 0006, Xianming Liu 0005, Oscar C. Au, Yuan Yan Tang |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2016 | SIFT Keypoint Removal and Injection via Convex RelaxationabstractScale invariant feature transform (SIFT), as one of the most popular local feature extraction algorithms, has been widely employed in many computer vision and multimedia security applications. Although SIFT has been extensively investigated from various perspectives, its security against malicious attacks has rarely been discussed. In this paper, we show that the SIFT keypoints can be effectively removed with minimized distortion on the processed image. The SIFT keypoint removal is formulated as a constrained optimization problem, where the constraints are carefully designed to suppress the existence of local extrema and prevent generating new keypoints within a local cuboid in the scale space. To hide the traces of performing SIFT keypoint removal, we then propose to inject a large number of fake SIFT keypoints into the previously cleaned image with minimized distortion. As demonstrated experimentally, our proposed SIFT removal and injection algorithms significantly outperform the state-of-the-art techniques. Furthermore, it is shown that the combined SIFT keypoint removal and injection attack strategy is capable of defeating the most powerful forensic detector designed for SIFT keypoint removal. Our results suggest that an authorization mechanism is required for SIFT-based systems to verify the validity of the input data, so as to achieve high reliability. Yuanman Li, Jiantao Zhou 0001, An Cheng, Xianming Liu 0005, Yuan Yan Tang |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2016 | Data-Driven Soft Decoding of Compressed Images in Dual Transform-Pixel DomainabstractIn the large body of research literature on image restoration, very few papers were concerned with compression-induced degradations, although in practice, the most common cause of image degradation is compression. This paper presents a novel approach to restoring JPEG-compressed images. The main innovation is in the approach of exploiting residual redundancies of JPEG code streams and sparsity properties of latent images. The restoration is a sparse coding process carried out jointly in the DCT and pixel domains. The prowess of the proposed approach is directly restoring DCT coefficients of the latent image to prevent the spreading of quantization errors into the pixel domain, and at the same time, using online machine-learned local spatial features to regulate the solution of the underlying inverse problem. Experimental results are encouraging and show the promise of the new approach in significantly improving the quality of DCT-coded images. Xianming Liu 0005, Xiaolin Wu 0001, Jiantao Zhou 0001, Debin Zhao |
IEEE Trans. Image Process. | 1 |
| 2016 | Compressive Sampling-Based Image Coding for Resource-Deficient Visual CommunicationabstractIn this paper, a new compressive sampling-based image coding scheme is developed to achieve competitive coding efficiency at lower encoder computational complexity, while supporting error resilience. This technique is particularly suitable for visual communication with resource-deficient devices. At the encoder, compact image representation is produced, which is a polyphase down-sampled version of the input image; but the conventional low-pass filter prior to down-sampling is replaced by a local random binary convolution kernel. The pixels of the resulting down-sampled pre-filtered image are local random measurements and placed in the original spatial configuration. The advantages of the local random measurements are two folds: 1) preserve high-frequency image features that are otherwise discarded by low-pass filtering and 2) remain a conventional image and can therefore be coded by any standardized codec to remove the statistical redundancy of larger scales. Moreover, measurements generated by different kernels can be considered as the multiple descriptions of the original image and therefore the proposed scheme has the advantage of multiple description coding. At the decoder, a unified sparsity-based soft-decoding technique is developed to recover the original image from received measurements in a framework of compressive sensing. Experimental results demonstrate that the proposed scheme is competitive compared with existing methods, with a unique strength of recovering fine details and sharp edges at low bit-rates. Xianming Liu 0005, Deming Zhai, Jiantao Zhou 0001, Xinfeng Zhang 0001, Debin Zhao, Wen Gao 0001 |
IEEE Trans. Image Process. | 1 |
| 2016 | Low-Rank Decomposition-Based Restoration of Compressed Images via Adaptive Noise EstimationabstractImages coded at low bit rates in real-world applications usually suffer from significant compression noise, which significantly degrades the visual quality. Traditional denoising methods are not suitable for the content-dependent compression noise, which usually assume that noise is independent and with identical distribution. In this paper, we propose a unified framework of content-adaptive estimation and reduction for compression noise via low-rank decomposition of similar image patches. We first formulate the framework of compression noise reduction based upon low-rank decomposition. Compression noises are removed by soft thresholding the singular values in singular value decomposition of every group of similar image patches. For each group of similar patches, the thresholds are adaptively determined according to compression noise levels and singular values. We analyze the relationship of image statistical characteristics in spatial and transform domains, and estimate compression noise level for every group of similar patches from statistics in both domains jointly with quantization steps. Finally, quantization constraint is applied to estimated images to avoid over-smoothing. Extensive experimental results show that the proposed method not only improves the quality of compressed images obviously for post-processing, but are also helpful for computer vision tasks as a pre-processing method. Xinfeng Zhang 0001, Weisi Lin, Ruiqin Xiong, Xianming Liu 0005, Siwei Ma 0001, Wen Gao 0001 |
IEEE Trans. Image Process. | 4 |
| 2016 | Guided Image Contrast Enhancement Based on Retrieved Images in CloudabstractWe propose a guided image contrast enhancement framework based on cloud images, in which the context- sensitive and context-free contrast is jointly improved via solving a multi-criteria optimization problem. In particular, the context-sensitive contrast is improved by performing advanced unsharp masking on the input and edge-preserving filtered images, while the context-free contrast enhancement is achieved by the sigmoid transfer mapping. To automatically determine the contrast enhancement level, the parameters in the optimization process are estimated by taking advantages of the retrieved images with similar content. For the purpose of automatically avoiding the involvement of low-quality retrieved images as the guidance, a recently developed no-reference image quality metric is adopted to rank the retrieved images from the cloud. The image complexity from the free-energy-based brain theory and the surface quality statistics in salient regions are collaboratively optimized to infer the parameters. Experimental results confirm that the proposed technique can efficiently create visually-pleasing enhanced images which are better than those produced by the classical techniques in both subjective and objective comparisons. Shiqi Wang 0001, Ke Gu 0001, Siwei Ma 0001, Weisi Lin, Xianming Liu 0005, Wen Gao 0001 |
IEEE Trans. Multim. | 5 |
| 2015 | Understanding image structure via hierarchical shape parsingabstractExploring image structure is a long-standing yet important research subject in the computer vision community. In this paper, we focus on understanding image structure inspired by the “simple-to-complex” biological evidence. A hierarchical shape parsing strategy is proposed to partition and organize image components into a hierarchical structure in the scale space. To improve the robustness and flexibility of image representation, we further bundle the image appearances into hierarchical parsing trees. Image descriptions are subsequently constructed by performing a structural pooling, facilitating efficient matching between the parsing trees. We leverage the proposed hierarchical shape parsing to study two exemplar applications including edge scale refinement and unsupervised “objectness” detection. We show competitive parsing performance comparing to the state-of-the-arts in above scenarios with far less proposals, which thus demonstrates the advantage of the proposed parsing scheme. Xianming Liu 0005, Rongrong Ji, Changhu Wang, Wei Liu 0005, Bineng Zhong 0001, Thomas S. Huang |
CVPR | 1 |
| 2015 | Data-driven sparsity-based restoration of JPEG-compressed images in dual transform-pixel domainabstractArguably the most common cause of image degradation is compression. This papers presents a novel approach to restoring JPEG-compressed images. The main innovation is in the approach of exploiting residual redundancies of JPEG code streams and sparsity properties of latent images. The restoration is a sparse coding process carried out jointy in the DCT and. pixel domains. The prowess of the proposed approach is directly restoring DCT coefficients of the latent image to prevent the spreading of quantization errors into the pixel domain, and at the same time using on-line machine-learnt local spatial features to regulate the solution of the underlying inverse problem. Experimental results are encouraging and show the promise of the new approach in significantly improving the quality of DCT-coded images. Xianming Liu 0005, Xiaolin Wu 0001, Jiantao Zhou 0001, Debin Zhao |
CVPR | 1 |
| 2015 | Joint denoising and contrast enhancement of images using graph laplacian operatorabstractImages and videos are often captured in poor light conditions, resulting in low-contrast images that are corrupted by acquisition noise. To recreate a high-quality image for visual observation, the captured image must be denoised and contrastenhanced. Conventional methods perform these two tasks in two separate stages: an image is first denoised, followed by an enhancement procedure. In this paper, we propose to jointly denoise and enhance an image in one unified optimization framework. The crux of the optimization rests on the definition of the enhancement operator, described by a graph Laplacian matrix H. The operator must enhance the high frequency details of the original image without amplifying additive noise. We propose a graph-based low-pass filtering approach to denoise edge weights in the graph, resulting in a more robust estimate of H. Experimental results show that our proposed joint approach can outperform the separate approach in demonstrable image quality. Xianming Liu 0005, Gene Cheung, Xiaolin Wu 0001 |
ICASSP | 1 |
| 2015 | Look and Think Twice: Capturing Top-Down Visual Attention with Feedback Convolutional Neural NetworksabstractWhile feedforward deep convolutional neural networks (CNNs) have been a great success in computer vision, it is important to note that the human visual cortex generally contains more feedback than feedforward connections. In this paper, we will briefly introduce the background of feedbacks in the human visual cortex, which motivates us to develop a computational feedback mechanism in deep neural networks. In addition to the feedforward inference in traditional neural networks, a feedback loop is introduced to infer the activation status of hidden layer neurons according to the "goal" of the network, e.g., high-level semantic labels. We analogize this mechanism as "Look and Think Twice." The feedback networks help better visualize and understand how deep neural networks work, and capture visual attention on expected objects, even in images with cluttered background and multiple objects. Experiments on ImageNet dataset demonstrate its effectiveness in solving tasks such as image classification and object localization. Chunshui Cao, Xianming Liu 0005, Yi Yang 0007, Yinan Yu, Jiang Wang 0001, Zilei Wang, Yongzhen Huang, Liang Wang 0001, Chang Huang, Wei Xu 0017, Deva Ramanan, Thomas S. Huang |
ICCV | 2 |
| 2015 | Inter-block consistent soft decoding of JPEG images with sparsity and graph-signal smoothness priorsabstractGiven the prevalence of JPEG compressed images on the Internet, image reconstruction from the compressed format remains an important and practical problem. Instead of simply reconstructing a pixel block from the centers of assigned DCT coefficient quantization bins (hard decoding), we propose to jointly reconstruct a neighborhood group of pixel patches using two image priors while satisfying the quantization bin constraints. First, we assume that a pixel patch can be approximated as a sparse linear combination of atoms from an offline-learned over-complete dictionary. Second, we assume that a patch, when interpreted as a graph-signal, is smooth with respect to an appropriately defined graph that captures the estimated structure of the target image. Finally, neighboring patches in the optimization have sufficient overlaps and are forced to be consistent, so that blocking artifacts typical in JPEG decoded images are avoided. To find the optimal group of patches, we formulate a constrained optimization problem and propose a fast alternating algorithm to find locally optimal solutions. Experimental results show that our proposed algorithm outperforms state-of-the-art soft decoding algorithms by up to 1.47dB in PSNR. Xianming Liu 0005, Gene Cheung, Xiaolin Wu 0001, Debin Zhao |
ICIP | 1 |
| 2015 | Quality assessment for out-of-focus blurred imagesabstractDuring the process of image acquisition, images are often subject to out-of-focus or defocus blur because of the improper adjustment of the camera's focal length, this image blur will degrade the image quality. However, in the literature, image quality assessment (IQA) methods dedicated to evaluating the quality of images with out-of-focus blur remain few. Therefore, in this paper, we focus our attention on the quality assessment of images that suffer from out-of-focus blur and propose an objective quality assessment method accordingly. Concretely, we construct a dedicated out-of-focus blurred image dataset, which is composed of 150 images subjected to different degrees of out-of-focus blur and the mean opinion scores (MOSs). Then, we propose a specific objective quality metric for the blurred images, which combines image sharpness assessment and saliency-guided pooling strategy. Experimental results demonstrate the proposed metric highly correlates with human judgements of image quality. Yutao Liu 0002, Guangtao Zhai, Xianming Liu 0005, Debin Zhao |
VCIP | 3 |
| 2015 | Model-based low bit-rate video coding for resource-deficient wireless visual communication
Xianming Liu 0005, Xinwei Gao, Debin Zhao, Jiantao Zhou 0001, Guangtao Zhai, Wen Gao 0001 |
Neurocomputing | 1 |
| 2015 | Localizing web videos using social images
Liujuan Cao, Xianming Liu 0005, Wei Liu 0005, Rongrong Ji, Thomas S. Huang |
Inf. Sci. | 2 |
| 2014 | Multiple Description Image Coding with Local Random MeasurementsabstractIn this paper, an effective multiple description image coding technique is developed to achieve competitive coding efficiency at low encoder complexity, while being standard compliant. The new technique is particularly suitable for visual communication over packet-switched networks and with resource-deficient wireless devices. To keep the encoder simple and standard compliant, multiple descriptions are produced by quincunx spatial multiplexing. Each side description is a polyphase down sampled version of the input image, but the conventional low-pass filter prior to downsampling is replaced by a local random binary convolution kernel. The pixels of each resulting side description are local random measurements and placed in the original spatial configuration. The advantages of local random measurements are two folds: 1) preservation of high-frequency image features that are otherwise discarded by low-pass filtering, 2) each side description remains a conventional image and can therefore be coded by any standardized codec to remove statistical redundancy of larger scales. The decoder performs joint upsampling of received description(s) and recovers the image from local random measurements in a framework of compressive sensing. Experimental results demonstrate that the proposed multiple description image codec is competitive in rate-distortion performance compared with existing methods, with a unique strength of recovering fine details and sharp edges at low bit rates. Xianming Liu 0005, Xiaolin Wu 0001, Debin Zhao |
DCC | 1 |
| 2014 | Estimation of capacity parameters for dynamic histogram shifting (DHS)-based reversible image watermarkingabstractDynamic histogram shifting (DHS) is a generation of the conventional histogram shifting (HS) technique for reversible image watermarking. Its superior embedding performance is achieved at the cost of significantly increased computational burden incurred by estimating the capacity parameters via multi-rounds of embedding iterations. In this work, we propose an analytical framework on estimating the optimal capacity parameters for DHS-based reversible image watermarking. We demonstrate that such parameter estimation can be cast as a convex optimization problem, which can be numerically solved in an efficient manner. The estimated values can then be utilized to facilitate a local search algorithm to obtain the truly optimal ones with much lowered complexity. Experimental results are provided to verify the validity of our findings. Li Dong 0006, Jiantao Zhou 0001, Yuan Yan Tang, Xianming Liu 0005 |
ICME | 4 |
| 2014 | Where should I stand? Learning based human position recommendation for mobile photographing
Pengfei Xu 0001, Hongxun Yao, Rongrong Ji, Xianming Liu 0005, Xiaoshuai Sun |
Multim. Tools Appl. | 4 |
| 2014 | Scalable Compression of Stream Cipher Encrypted Images Through Context-Adaptive SamplingabstractThis paper proposes a novel scalable compression method for stream cipher encrypted images, where stream cipher is used in the standard format. The bit stream in the base layer is produced by coding a series of nonoverlapping patches of the uniformly down-sampled version of the encrypted image. An off-line learning approach can be exploited to model the reconstruction error from pixel samples of the original image patch, based on the intrinsic relationship between the local complexity and the length of the compressed bit stream. This error model leads to a greedy strategy of adaptively selecting pixels to be coded in the enhancement layer. At the decoder side, an iterative, multiscale technique is developed to reconstruct the image from all the available pixel samples. Experimental results demonstrate that the proposed scheme outperforms the state-of-the-arts in terms of both rate-distortion performance and visual quality of the reconstructed images at low and medium rate regions. Jiantao Zhou 0001, Oscar C. Au, Guangtao Zhai, Yuan Yan Tang, Xianming Liu 0005 |
IEEE Trans. Inf. Forensics Secur. | 5 |
| 2014 | Designing an Efficient Image Encryption-Then-Compression System via Prediction Error Clustering and Random PermutationabstractIn many practical scenarios, image encryption has to be conducted prior to image compression. This has led to the problem of how to design a pair of image encryption and compression algorithms such that compressing the encrypted images can still be efficiently performed. In this paper, we design a highly efficient image encryption-then-compression (ETC) system, where both lossless and lossy compression are considered. The proposed image encryption scheme operated in the prediction error domain is shown to be able to provide a reasonably high level of security. We also demonstrate that an arithmetic coding-based approach can be exploited to efficiently compress the encrypted images. More notably, the proposed compression approach applied to encrypted images is only slightly worse, in terms of compression efficiency, than the state-of-the-art lossless/lossy image coders, which take original, unencrypted images as inputs. In contrast, most of the existing ETC solutions induce significant penalty on the compression efficiency. Jiantao Zhou 0001, Xianming Liu 0005, Oscar C. Au, Yuan Yan Tang |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2014 | Image Interpolation via Graph-Based Bayesian Label PropagationabstractIn this paper, we propose a novel image interpolation algorithm via graph-based Bayesian label propagation. The basic idea is to first create a graph with known and unknown pixels as vertices and with edge weights encoding the similarity between vertices, then the problem of interpolation converts to how to effectively propagate the label information from known points to unknown ones. This process can be posed as a Bayesian inference, in which we try to combine the principles of local adaptation and global consistency to obtain accurate and robust estimation. Specially, our algorithm first constructs a set of local interpolation models, which predict the intensity labels of all image samples, and a loss term will be minimized to keep the predicted labels of the available low-resolution (LR) samples sufficiently close to the original ones. Then, all of the losses evaluated in local neighborhoods are accumulated together to measure the global consistency on all samples. Moreover, a graph-Laplacian-based manifold regularization term is incorporated to penalize the global smoothness of intensity labels, such smoothing can alleviate the insufficient training of the local models and make them more robust. Finally, we construct a unified objective function to combine together the global loss of the locally linear regression, square error of prediction bias on the available LR samples, and the manifold regularization term. It can be solved with a closed-form solution as a convex optimization problem. Experimental results demonstrate that the proposed method achieves competitive performance with the state-of-the-art image interpolation algorithms. Xianming Liu 0005, Debin Zhao, Jiantao Zhou 0001, Wen Gao 0001, Huifang Sun |
IEEE Trans. Image Process. | 1 |
| 2014 | Progressive Image Denoising Through Hybrid Graph Laplacian Regularization: A Unified FrameworkabstractRecovering images from corrupted observations is necessary for many real-world applications. In this paper, we propose a unified framework to perform progressive image recovery based on hybrid graph Laplacian regularized regression. We first construct a multiscale representation of the target image by Laplacian pyramid, then progressively recover the degraded image in the scale space from coarse to fine so that the sharp edges and texture can be eventually recovered. On one hand, within each scale, a graph Laplacian regularization model represented by implicit kernel is learned, which simultaneously minimizes the least square error on the measured samples and preserves the geometrical structure of the image data space. In this procedure, the intrinsic manifold structure is explicitly considered using both measured and unmeasured samples, and the nonlocal self-similarity property is utilized as a fruitful resource for abstracting a priori knowledge of the images. On the other hand, between two successive scales, the proposed model is extended to a projected high-dimensional feature space through explicit kernel mapping to describe the interscale correlation, in which the local structure regularity is learned and propagated from coarser to finer scales. In this way, the proposed algorithm gradually recovers more and more image details and edges, which could not been recovered in previous scale. We test our algorithm on one typical image recovery task: impulse noise removal. Experimental results on benchmark test images demonstrate that the proposed method achieves better performance than state-of-the-art algorithms. Xianming Liu 0005, Deming Zhai, Debin Zhao, Guangtao Zhai, Wen Gao 0001 |
IEEE Trans. Image Process. | 1 |
| 2014 | Toward Statistical Modeling of Saccadic Eye-Movement and Visual SaliencyabstractIn this paper, we present a unified statistical framework for modeling both saccadic eye movements and visual saliency. By analyzing the statistical properties of human eye fixations on natural images, we found that human attention is sparsely distributed and usually deployed to locations with abundant structural information. This observations inspired us to model saccadic behavior and visual saliency based on super-Gaussian component (SGC) analysis. Our model sequentially obtains SGC using projection pursuit, and generates eye movements by selecting the location with maximum SGC response. Besides human saccadic behavior simulation, we also demonstrated our superior effectiveness and robustness over state-of-the-arts by carrying out dense experiments on synthetic patterns and human eye fixation benchmarks. Multiple key issues in saliency modeling research, such as individual differences, the effects of scale and blur, are explored in this paper. Based on extensive qualitative and quantitative experimental results, we show promising potentials of statistical approaches for human behavior research. Xiaoshuai Sun, Hongxun Yao, Rongrong Ji, Xianming Liu 0005 |
IEEE Trans. Image Process. | 4 |
| 2013 | Image Super-Resolution via Hierarchical and Collaborative Sparse RepresentationabstractIn this paper, we propose an efficient image super-resolution algorithm based on hierarchical and collaborative sparse representation (HCSR). Motivated by the observation that natural images typically exhibit multi-modal statistics, we propose a hierarchical sparse coding model which includes two layers: the first layer encodes individual patches, and the second layer jointly encodes the set of patches that belong to the same homogeneous subset of image space. We further present a simple alternative to achieve such target by identifying optimal sparse representation that is adaptive to specific statistics of images. Specially, we cluster images from the offline training set into regions of similar geometric structure, and model each region (cluster) by learning adaptive bases describing the patches within that cluster using principal component analysis (PCA). This cluster-specific dictionary is then exploited to optimally estimate the underlying HR pixel values using the idea of collaborative sparse coding, in which the similarity between patches in the same cluster is further considered. It conceptually and computationally remedies the limitation of many existing algorithms based on standard sparse coding, in which patches are independently encoded. Experimental results demonstrate the proposed method appears to be competitive with state-of-the-art algorithms. Xianming Liu 0005, Deming Zhai, Debin Zhao, Wen Gao 0001 |
DCC | 1 |
| 2013 | Progressive Image Restoration through Hybrid Graph Laplacian RegularizationabstractIn this paper, we propose a unified framework to perform progressive image restoration based on hybrid graph Laplacian regularized regression. We first construct a multi-scale representation of the target image by Laplacian pyramid, then progressively recover the degraded image in the scale space from coarse to fine so that the sharp edges and texture can be eventually recovered. On one hand, within each scale, a graph Laplacian regularization model represented by implicit kernel is learned which simultaneously minimizes the least square error on the measured samples and preserves the geometrical structure of the image data space by exploring non-local self-similarity. In this procedure, the intrinsic manifold structure is considered by using both measured and unmeasured samples. On the other hand, between two scales, the proposed model is extended to the parametric manner through explicit kernel mapping to model the inter-scale correlation, in which the local structure regularity is learned and propagated from coarser to finer scales. Experimental results on benchmark test images demonstrate that the proposed method achieves better performance than state-of-the-art image restoration algorithms. Deming Zhai, Xianming Liu 0005, Debin Zhao, Hong Chang 0001, Wen Gao 0001 |
DCC | 2 |
| 2013 | On the design of an efficient encryption-then-compression systemabstractIn many practical scenarios, image encryption has to be conducted prior to image compression. This has led to the problem of how to design a pair of encryption and compression algorithms such that compressing the encrypted image can still be efficiently performed. In this work, we propose a permutation-based image encryption method conducted over the prediction error domain. We also design an arithmetic coding (AC)-based approach to efficiently compress the encrypted image. It can be shown that the proposed scheme can provide reasonably high level of security. More notably, the compression performance on the encrypted image is only slightly degraded, compared with that of compressing the original, un-encrypted one. In contrast, most of the existing approaches induce significant penalty on the compression performance. Jiantao Zhou 0001, Xianming Liu 0005, Oscar C. Au |
ICASSP | 2 |
| 2013 | Sparsity-based soft decoding of compressed images in transform domainabstractWe propose a sparsity-based soft decoding approach to restore compressed images directly in the transform domain of compression (DCT domain specifically examined in this paper). Restoring transform coefficients rather than pixel values prevents the propagation of quantization errors in the image domain. As natural images are statistically non-stationary with spatially varying sparse representations, we develop an adaptive block-wise sparsity-based restoration method that learns and exploits local statistics. Specially, for each DCT block, we collect sample blocks via non-local patch grouping to learn a compact dictionary based on principal component analysis. The resulting block-specific dictionary is used to estimate the corresponding DCT coefficients by a technique of collaborative sparse coding, in which the similarity between sample DCT patches used in dictionary construction is further considered. Experimental results are encouraging and demonstrate that the proposed soft decoding approach performs competitively on restoring compressed images against existing methods. Xianming Liu 0005, Xiaolin Wu 0001, Debin Zhao |
ICIP | 1 |
| 2013 | Structured Textons for texture representationabstractIn this paper, we propose a novel texture descriptor, Structured Texton, to extract and characterize meaningful texture patterns in images. Structured Textons are constructed by grouping local extremum regions connected by the nesting relationship. To further improve the discriminative ability, high order texton words are generated from the Structured Textons, preserving both the appearance information and the spatial information. Finally, a semantic ranking criterion is proposed for selecting the discriminative high order texton words by means of finding informative patterns from images. The proposed Structured Texton is more discriminative than the single texton-based representation. Experimental results of texture classification and scene classification on public datasets demonstrate the effectiveness and discrimination of the proposed Structured Texton. Pengfei Xu 0001, Xianming Liu 0005, Hongxun Yao, Yanhao Zhang 0001, Shaopeng Tang |
ICIP | 2 |
| 2013 | Parametric Local Multimodal Hashing for Cross-View Similarity Search
Deming Zhai, Hong Chang 0001, Yi Zhen, Xianming Liu 0005, Xilin Chen 0001, Wen Gao 0001 |
IJCAI | 4 |
| 2013 | Low bit-rate image coding via local random down-samplingabstractA common practice in low bit-rate image/video compression is uniform spatial down-sampling at the encoder and upsampling at the decoder. The down-sampling is performed in conjunction with deterministic low-pass filtering (e.g., Gaussian or the alike) to prevent aliasing. The down-sampled image is compressed and decompressed as usual; the upsampling is treated as an image restoration problem. In this paper, we show that the rate-distortion performance of the above low bit-rate image coding system can be improved, if the deterministic low-pass down-sampling filter is replaced by a random convolution kernel. The resulting down-sampled image is a two-dimensional array of local random measurements; this smaller image is still compressible in most cases. Accordingly, the decoder recovers the image from these local random measurements in the framework of compressive sensing. Theoretical analysis is conducted to support the superior performance of the proposed new method over its predecessors, and it is corroborated by our simulation results. At low to medium bit rates, the new method outperforms not only JPEG 2000 but also our earlier low bit-rate image codec CADU, with clear advantages over the competing methods in the reconstruction of high frequency features. In addition, the new method retains the system advantages of low encoder complexity and standard compliance as in CADU. Reza Pournaghi, Xiaolin Wu 0001, Xianming Liu 0005 |
PCS | 3 |
| 2013 | Bidirectional-isomorphic manifold learning at image semantic understanding & representation
Xianming Liu 0005, Hongxun Yao, Rongrong Ji, Pengfei Xu 0001, Xiaoshuai Sun |
Multim. Tools Appl. | 1 |
| 2012 | The scale of edgesabstractAlthough the scale of isotropic visual elements such as blobs and interest points, e.g. SIFT[12], has been well studied and adopted in various applications, how to determine the scale of anisotropic elements such as edges is still an open problem. In this paper, we study the scale of edges, and try to answer two questions: 1) what is the scale of edges, and 2) how to calculate it. From the points of human cognition and physical interpretation, we illustrate the existence of the scale of edges and provide a quantitative definition. Then, an automatic edge scale selection approach is proposed. Finally, a cognitive experiment is conducted to validate the rationality of the detected scales. Moreover, the importance of identifying the scale of edges is also shown in applications such as boundary detection and hierarchical edge parsing. Xianming Liu 0005, Changhu Wang, Hongxun Yao, Lei Zhang 0001 |
CVPR | 1 |
| 2012 | Multi-scale Spatial Error Concealment via Hybrid Bayesian RegressionabstractIn this paper, we propose a novel multi-scale spatial error concealment algorithm to combine the modeling strengthes of the parametric and nonparametric Bayesian regression. We progressively recover missing blocks in the scale space from coarse to fine so that the sharp edges and texture in the finest scale can be eventually recovered. On one hand, in each scale, the nonparametric part of the methodology is used to exploit the intra-scale correlation, which relies on the data itself to dictate the structure of the model. In this procedure, the non-local self-similarity property is utilized as a fruitful resource for abstracting a priori knowledge of images. On the other hand, the parametric part is used to explicitly model the inter-scale correlation, in which the local structure regularity is thoroughly explored to recover the sharp edges and major texture features of images. It is not respected if only the nonparametric modeling is considering. We achieve the best of both worlds within a multi-scale framework. Experimental results on benchmark test images demonstrate that the proposed method achieves very competitive performance with the state-of-the-art error concealment algorithms. Xianming Liu 0005, Deming Zhai, Guangtao Zhai, Debin Zhao, Ruiqin Xiong, Wen Gao 0001 |
DCC | 1 |
| 2012 | Web image interpolation via weighted total least squares regressionabstractAlthough ordinary least squares (OLS) regression achieves great success in clean image interpolation, its effectiveness is questionable in the scenario of web images which are usually compressed beforehand. The inherent flaw of OLS is that it is asymmetric, the perturbation is only confined on the right side of the linear system. It is not reasonable for web images. Considering the drawback of OLS, in this paper, we propose an efficient web image interpolation algorithm based on total least squares (TLS) regression. In the proposed method, small perturbations are allowed in both side of the system, which are optimized by TLS in a patch-based manner. Furthermore, we develop a weighted version of TLS to consider contribution diversity of different samples and patches in model estimation, which can efficiently remove the influence of outliers in regression. Experimental results on benchmark test images demonstrate the efficiency of our method. Xianming Liu 0005, Deming Zhai, Guangtao Zhai, Debin Zhao, Wen Gao 0001 |
ICASSP | 1 |
| 2012 | Low bit-rate video coding via mode-dependent adaptive regression for wireless visual communicationsabstractIn this paper, a practical video coding scheme is developed to realize state-of-the-art video coding efficiency with lower encoder complexity at low bit-rate, while supporting standard compliance and error resilience. Such an architecture is particularly attractive for wireless visual communications. At the encoder, multiple descriptions of a video sequence are generated in the spatio-temporal domain by temporal multiplexing and spatial adaptive downsampling. The resulting side descriptions are interleaved with each other in temporal domain, and still with conventional square sample grids in spatial domain. As such, each side description can be compressed without any change to existing video coding standards. At the decoder, each side description is first decompressed, and then reconstructed to original resolution with the help of the other side description. In this procedure, the decoder recover the original video sequence in a constrained least squares regression process, using 2D or 3D piecewise autoregressive model according to different prediction modes. In this way, the spatial and temporal correlation is sufficiently explored to achieve superior quality. Experiment results demonstrate the proposed video coding scheme outperforms H.264 in rate-distortion performance at low bit-rates and achieves superior visual quality at medium bit-rates as well. Xianming Liu 0005, Xiaolin Wu 0001, Xinwei Gao, Debin Zhao, Wen Gao 0001 |
VCIP | 1 |
| 2012 | Context-Aware Semi-Local Feature DetectorabstractHow can interest point detectors benefit from contextual cues? In this articles, we introduce a context-aware semi-local detector (CASL) framework to give a systematic answer with three contributions: (1) We integrate the context of interest points to recurrently refine their detections. (2) This integration boosts interest point detectors from the traditionally local scale to a semi-local scale to discover more discriminative salient regions. (3) Such context-aware structure further enables us to bring forward category learning (usually in the subsequent recognition phase) into interest point detection to locate category-aware, meaningful salient regions. Our CASL detector consists of two phases. The first phase accumulates multiscale spatial correlations of local features into a difference of contextual Gaussians (DoCG) field. DoCG quantizes detector context to highlight contextually salient regions at a semi-local scale, which also reveals visual attentions to a certain extent. The second phase locates contextual peaks by mean shift search over the DoCG field, which subsequently integrates contextual cues into feature description. This phase enables us to integrate category learning into mean shift search kernels. This learning-based CASL mechanism produces more category-aware features, which substantially benefits the subsequent visual categorization process. We conducted experiments in image search, object characterization, and feature detector repeatability evaluations, which reported superior discriminability and comparable repeatability to state-of-the-art works. Rongrong Ji, Hongxun Yao, Qi Tian 0001, Pengfei Xu 0001, Xiaoshuai Sun, Xianming Liu 0005 |
ACM Trans. Intell. Syst. Technol. | 6 |
| 2011 | Transductive Regression with Local and Global Consistency for Image Super-ResolutionabstractIn this paper, we propose a novel image super-resolution algorithm, referred to as interpolation based on transductive regression with local and global consistency (TRLGC). Our algorithm first constructs a set of local interpolation models which can predict the intensity labels of all image samples, and a loss term will be minimized to keep the predicted labels of available low-resolution (LR) samples sufficiently close to the original ones. Then, all of the losses evaluated in local neighborhoods are accumulated together to measure the global consistency on all samples. Furthermore, a graph-Laplacian based manifold regularization term is incorporated to penalize the global smoothness of intensity labels, such smoothing can alleviate the insufficient training of the local models and make them more robust. Finally, we construct a unified objective function to combine together the accumulated loss of the locally linear regression, square error of prediction bias on the available LR samples and the manifold regularization term, which could be solved with a closed-form solution as a convex optimization problem. In this way, a transductive regression algorithm with local and global consistency is developed. Experimental results on benchmark test images demonstrate that the proposed image super-resolution method achieves very competitive performance with the state-of-the-art algorithms. Xianming Liu 0005, Debin Zhao, Ruiqin Xiong, Siwei Ma 0001, Wen Gao 0001, Huifang Sun |
DCC | 1 |
| 2011 | Sparse representation based visual element analysisabstractModern clothes are designed based on various visual elements of different fashion styles. Traditional vision-based clothes recommendation methods focused on searching clothes which are similar with user preferred samples in the aspects of colors and partial shape elements. In this paper, we propose a method of recommending clothes by mining visual elements of different fashion styles. Independent Component Analysis (ICA) is employed to extract sparse features, and then Term-Frequency (TF) analysis is applied to discover visual elements from these independent components. Finally, we test three ranking metrics for clothes recommendation including Euclidian distance of TFs, Cosine distance of TFs and Minimum TF. Experimental results based on web commercial images demonstrate the effectiveness of the proposed method. Hongxun Yao, Xiaoshuai Sun, Rongrong Ji, Xianming Liu 0005, Pengfei Xu 0001 |
ICIP | 5 |
| 2011 | Side information extrapolation with temporal and spatial consistencyabstractIn this paper, we present an efficient side information extrapolation scheme with temporal and spatial consistency for low delay Wyner-Ziv video coding. Our method is based on the regularized local linear regression (RLLR) model, in which each pixel in SI is approximated as a linear weighted combination of samples within a local temporal neighborhood. The optimal model parameters are estimated by projecting the transformation function onto the temporal training samples to exploit motion-related dependency. During this procedure, moving weights are incorporated into the objective function to express the relative importance of training samples in estimating parameters of the model. Furthermore, spatial correlation is explored by imposing an additional local smoothness penalty, which does good to estimate the occluded regions and complex motion regions. The learned function is smooth and locally linear, and can be obtained with a closed-form solution by solving a convex optimization problem. Experimental results demonstrate that the RLLR method achieves very competitive SI extrapolation performance compared with the state-of-the-art methods. Xianming Liu 0005, Deming Zhai, Debin Zhao, Ruiqin Xiong, Siwei Ma 0001, Wen Gao 0001 |
ISCAS | 1 |
| 2011 | Learning heterogeneous data for hierarchical web video classificationabstractWeb videos such as YouTube are hard to obtain sufficient precisely labeled training data and analyze due to the complex ontology. To deal with these problems, we present a hierarchical web video classification framework by learning heterogeneous web data, and construct a bottom-up semantic forest of video concepts by learning from meta-data. The main contributions are two-folds: firstly, analysis about middle-level concepts' distribution is taken based on data collected from web communities, and a concepts redistribution assumption is made to build effective transfer learning algorithm. Furthermore, an AdaBoost-Like transfer learning algorithm is proposed to transfer the knowledge learned from Flickr images to YouTube video domain and thus it facilitates video classification. Secondly, a group of hierarchical taxonomies named Semantic Forest are mined from YouTube and Flickr tags which reflect better user intention on the semantic level. A bottom-up semantic integration is also constructed with the help of semantic forest, in order to analyze video content hierarchically in a novel perspective. A group of experiments are performed on the dataset collected from Flickr and YouTube. Compared with state-of-the-arts, the proposed framework is more robust and tolerant to web noise. Xianming Liu 0005, Hongxun Yao, Rongrong Ji, Pengfei Xu 0001, Xiaoshuai Sun, Qi Tian 0001 |
ACM Multimedia | 1 |
| 2011 | Unsupervised fast anomaly detection in crowdsabstractIn this paper, we proposed a fast and robust unsupervised framework for anomaly detection and localization in crowed scenes. Our method avoids modeling the normal state of the crowds which is a very complex task due to the large within class variance of the normal target appearance and motion patterns. For each video frame, we extract the spatial temporal features of 3D blocks and generate the saliency map using a block-based center-surround difference operator. Then, motion vector matrix is obtained by adaptive rood pattern search block-matching algorithm and distance normalization. Attractive motion disorder descriptor is proposed to measure the global intensity of anomalies in the scene. Finally, we classify the frames into normal and anomalous ones by a binary classifier. In the experiments, we compared our method against several state-of-the-art approaches on UCSD dataset which is a widely used anomaly detection and localization benchmark. As the only unsupervised approach, our method outputs competitive results with near real-time processing speed Xiaoshuai Sun, Hongxun Yao, Rongrong Ji, Xianming Liu 0005, Pengfei Xu 0001 |
ACM Multimedia | 4 |
| 2011 | Video indexing and recommendation based on affective analysis of viewersabstractMost previous works on video indexing and recommendation were only based on the content of video itself, without considering the affective analysis of viewers, which is an efficient and important way to reflect viewers' attitudes, feelings and evaluations of videos. In this paper, we propose a novel method to index and recommend videos based on affective analysis, mainly on facial expression recognition of viewers. We first build a facial expression recognition classifier by embedding the process of building compositional Haar-like features into hidden conditional random fields (HCRFs). Then we extract viewers' facial expressions frame by frame through the videos, collected from the camera when viewers are watching videos, to obtain the affections of viewers. Finally, we draw the affective curve which tells the process of affection changes. Through the curve, we segment each video into affective sections, give the indexing result of the videos, and list recommendation points from views' aspect. Experiments on our collected database from the web show that the proposed method has a promising performance. Sicheng Zhao, Hongxun Yao, Xiaoshuai Sun, Pengfei Xu 0001, Xianming Liu 0005, Rongrong Ji |
ACM Multimedia | 5 |
| 2011 | Image Interpolation Via Regularized Local Linear RegressionabstractThe linear regression model is a very attractive tool to design effective image interpolation schemes. Some regression-based image interpolation algorithms have been proposed in the literature, in which the objective functions are optimized by ordinary least squares (OLS). However, it is shown that interpolation with OLS may have some undesirable properties from a robustness point of view: even small amounts of outliers can dramatically affect the estimates. To address these issues, in this paper we propose a novel image interpolation algorithm based on regularized local linear regression (RLLR). Starting with the linear regression model where we replace the OLS error norm with the moving least squares (MLS) error norm leads to a robust estimator of local image structure. To keep the solution stable and avoid overfitting, we incorporate the l(2)-norm as the estimator complexity penalty. Moreover, motivated by recent progress on manifold-based semi-supervised learning, we explicitly consider the intrinsic manifold structure by making use of both measured and unmeasured data points. Specifically, our framework incorporates the geometric structure of the marginal probability distribution induced by unmeasured samples as an additional local smoothness preserving constraint. The optimal model parameters can be obtained with a closed-form solution by solving a convex optimization problem. Experimental results on benchmark test images demonstrate that the proposed method achieves very competitive performance with the state-of-the-art interpolation algorithms, especially in image edge structure preservation. Xianming Liu 0005, Debin Zhao, Ruiqin Xiong, Siwei Ma 0001, Wen Gao 0001, Huifang Sun |
IEEE Trans. Image Process. | 1 |
| 2010 | Exploring statistical properties for semantic annotation: sparse distributed and convergent assumptions for keywordsabstractDoes there exist a compact set of visual topics in form of keyword clusters capable to represent all images visual content within an acceptable error? In this paper, we answer this question by analyzing distribution laws for keywords from image descriptions and comparing with traditional techniques in NLP, thereby propose three assumptions: Sparse Distribution Attribute, Local Convergent Assumption and Global Convergent Conjecture. They are essential for keywords selection and image content understanding to overcome the semantic gap. Experiments are performed on a 60,000 web crawled images, and the correctness is validated by the performance. Xianming Liu 0005, Hongxun Yao, Rongrong Ji |
ICASSP | 1 |
| 2010 | Visual saliency as sequential eye fixation probabilityabstractHuman vision system acquires essential information from the environment by sequentially sampling visual contents at important locations under the control of selective attention mechanism. We propose that bottom-up saliency is not based on global statistics but on information sampled at prior eye fixations. Our model calculates visual saliency using sequential eye fixation probability. However, the proposed model needs fixation priors, which are hard to simulate given current fixation data and experimental conditions. An approximation is proposed to generate a single saliency map by fusing all possible conditions of fixation prior. Our method outperforms all state-of-the-art models in predicting eye fixations, and shows reasonable response to various psychological patterns. Xiaoshuai Sun, Hongxun Yao, Rongrong Ji, Pengfei Xu 0001, Xianming Liu 0005, Shaohui Liu |
ICIP | 5 |
| 2010 | Saliency detection based on short-term sparse representationabstractRepresentation and measurement are two important issues for saliency models. Different with previous works that learnt sparse features from large scale natural statistics, we propose to learn features from short-term statistics of single images. For saliency measurement, we define background firing rate (BFR) for each sparse feature, and then we propose to use feature activation rate (FAR) to measure the bottom-up visual saliency. The proposed FAR measure is biological plausible and easy to compute, also with satisfied performance. Experiments on human eye fixations and psychological patterns demonstrate the effectiveness and robustness of our proposed method. Xiaoshuai Sun, Hongxun Yao, Rongrong Ji, Pengfei Xu 0001, Xianming Liu 0005, Shaohui Liu |
ICIP | 5 |
| 2010 | Image interpolation via regularized local linear regressionabstractIn this paper, we present an efficient image interpolation scheme by using regularized local linear regression (RLLR). On one hand, we introduce a robust estimator of local image structure based on moving least squares, which can efficiently handle the statistical outliers compared with ordinary least squares based methods. On the other hand, motivated by recent progress on manifold based semi-supervise learning, the intrinsic manifold structure is explicitly considered by making use of both measured and unmeasured data points. In particular, the geometric structure of the marginal probability distribution induced by unmeasured samples is incorporated as an additional locality preserving constraint. The optimal model parameters can be obtained with a closed-form solution by solving a convex optimization problem. Experimental results demonstrate that our method outperform the existing methods in both objective and subjective visual quality over a wide range of test images. Xianming Liu 0005, Debin Zhao, Ruiqin Xiong, Siwei Ma 0001, Wen Gao 0001 |
PCS | 1 |
| 2010 | Side information enhancement via texture and motion activity analysis in distributed video codingabstractThis paper investigates how to exploit the artifacts constraint to enhance the quality of side information (SI) in distributed video coding (DVC). The idea originates from the observation that there are usually some regions shown as artifacts in the interpolated SI, which seriously degrade the overall performance of the DVC system, especially when the GOP size is large. Encouraged by the previous work about the backward channel in the literature, we propose a simple and effective DVC scheme based on artifact detection and removal. On one hand, texture analysis via dynamic range and local entropy is utilized to determine artifact regions in spatial domain. On the other hand, motion activity evaluation is used to determine artifact regions in temporal domain, and temporal local entropy difference is further employed to eliminate the effect of other factors like brightness variance. After the artifact blocks are detected, we re-code them in intra mode and replace the co-located ones at the decoder side so that the quality of SI is enhanced. Experimental results demonstrate that the proposed SI quality enhancement technique is effective. Xianming Liu 0005, Debin Zhao, Siwei Ma 0001, Wen Gao 0001 |
VCIP | 1 |
| 2010 | A rotation and scale invariant texture description approachabstractThis paper presents a novel texture description approach, which is robust to variances in rotation, scale and illumination in images, to classify the texture of images. A limitation with traditional methods is that they are more or less sensitive to the mentioned changes in images. To overcome this problem, we propose a novel Local Haar Binary Pattern (LHBP) based framework to ensure invariance in global rotation, scale, and light change. Our method consists of two components: feature extraction and scale self-adaptive classification. The global rotation invariant LHBP histogram features are extracted against the variances of illumination and global rotation, and the scale self-adaptive strategy is used for optimizing the classification of different scale textures. Evaluation results on Outex and Brodatz databases illustrate the significant advantages of the proposed approach over existing algorithms. Pengfei Xu 0001, Hongxun Yao, Rongrong Ji, Xiaoshuai Sun, Xianming Liu 0005 |
VCIP | 5 |
| 2009 | Joint learning for side information and correlation model based on linear regression model in distributed video codingabstractThe coding efficiency of distributed video coding system is significantly determined by the side information quality and correlation model. Motivated by theoretical analysis of the maximum likelihood treatment for linear regression model, we propose a novel joint online learning model for side information generation and correlation model estimation in this paper. In our proposed scheme, each pixel in the side information is approximated as the linear weighted combination of samples within a local spatio-temporal neighboring space. Weights are trained in a self-feedback fashion, during which the correlation model parameters can also be achieved. The efficiency of the proposed joint learning model is confirmed experimentally. Xianming Liu 0005, Debin Zhao, Yongbing Zhang 0002, Siwei Ma 0001, Qingming Huang, Wen Gao 0001 |
ICIP | 1 |
| 2009 | What is a complete set of keywords for image description & annotation on the webabstractDoes there exist a compact set of keywords that can completely and effectively cover the image annotation problem by expanding from it? In this paper, we answer this question by presenting a complete set framework for image annotation, which is motivated by the existence of semantic ontology. To generate this set, we propose a cross model optimization strategy from both textual and visual information for topic decomposition, based on a so-called Bipartite LSA model, which minimize multimodal error energy functions in a probabilistic Latent Semantic Analysis model. To achieve complete set based annotation, we present a Gaussian-Kernel-Generative process based keyword generation procedure, which analogizes keyword annotation in a probabilistic generative manner. A group of experiments is performed on Washington University image database and 80,000 Flickr images with comparisons to the state-of-the-arts. Finally, potential advantages and future improvements of our framework are discussed outside the scope of topic modeling. Xianming Liu 0005, Hongxun Yao, Rongrong Ji, Pengfei Xu 0001, Xiaoshuai Sun |
ACM Multimedia | 1 |
| 2009 | Multi-hypothesis based multi-view distributed video codingabstractThis paper proposes a multi-hypothesis based Wyner-Ziv (WZ) decoder for the multi-view distributed video coding (MDVC). Two hypotheses, the intra-view SI and the interview SI, are fed together into the WZ decoder in the proposed scheme. A multi-hypothesis based correlation model (MHBCM) is presented to fully exploit the redundancy between these two SI frames and the original frame. The MHBCM is also applied on the optimal minimum mean-square error reconstruction of the quantized samples. The simulation results show that the proposed algorithms are able to significantly improve the coding efficiency of the MDVC system. Yongpeng Li, Hongbin Liu 0004, Xianming Liu 0005, Siwei Ma 0001, Debin Zhao, Wen Gao 0001 |
PCS | 3 |
| 2009 | Two-pass reconstruction in distributed video codingabstractIn this paper, we propose a novel two-pass reconstruction algorithm for the Wyner-Ziv (WZ) frames in distributed video coding (DVC), in which the traditional reconstructed WZ frame is utilized to perform motion estimation to obtain a more accurate motion field. During the motion estimation, the block, as well as its neighboring pixels are concerned. An overlapped block motion compensation is subsequently performed with the help of the motion field, consequently, an enhanced prediction for the WZ frame can be obtained, based on which an improved reconstruction can be achieved. Simulation results show that both the objective and subjective quality of WZ frames can be improved significantly. Hongbin Liu 0004, Yongpeng Li, Xianming Liu 0005, Siwei Ma 0001, Debin Zhao, Wen Gao 0001 |
PCS | 3 |
| 2009 | Improved low delay distributed video codingabstractThis paper proposes an image partition based approach to enhance side information quality in low delay distributed video coding (DVC). The proposed method employs a checkerboard pattern to group blocks of the Wyner-Ziv frame into two sets, where one set is DPCM encoded and the other set is DVC encoded. These two sets are encoded independently and decoded successively. At decoder, DPCM set will be first reconstructed. Then the temporal concealment tool, such as boundary matching algorithm, is performed to conceal blocks in the DVC set. An improved side information is subsequently obtained for DVC set, based on which a higher compression can be achieved. Simulation results indicate that a more promising performance can be achieved when compared with existing motion extrapolated approach. Hongbin Liu 0004, Yongpeng Li, Xianming Liu 0005, Siwei Ma 0001, Debin Zhao, Wen Gao 0001 |
PCS | 3 |
| 2009 | Local adaptive learning and fusion for side information interpolation in distributed video codingabstractMotivated by theoretical analysis of the curve fitting problem based on equivalent kernel, in this paper we propose a local adaptive learning and fusion model for side information interpolation in distributed video coding. In the proposed model, each pixel in the interpolated frame is approximated as the linear combination of samples within a local spatio-temporal window using kernel parameters as weight. The size of training window can be adaptive to the motion characteristic of video, from samples in which the kernel parameters can be locally learned. In order to further improve the quality of interpolated frames, we introduce a belief-projection based fusion strategy with adaptive weights for multiple interpolated results which are with the same time index. Experimental results demonstrate that the proposed learning and fusion model is effective in performance for side information interpolation in distributed video coding. Xianming Liu 0005, Yongbing Zhang 0002, Yongpeng Li, Hongbin Liu 0004, Siwei Ma 0001, Debin Zhao |
PCS | 1 |
| 2008 | Clustering-based subspace SVM ensemble for relevance feedback learningabstractThis paper presents a subspace SVM ensemble algorithm for adaptive relevance feedback (RF) learning. Our method deals with the case that user’s relevance feedback examples are usually insufficient and overlapped together in feature space, which decreases the learning effectiveness of RF classifiers. To enhance classification efficiency in such case, multiple SVMs are learned by clustering-based training set partition, each of which fits its cluster-specific sample distribution and gives labeling regressions to test samples that fall within this cluster. To adapt features to sample distribution within each cluster, AdaBoost feature selection is conducted onto pyramid Haar of H&I bands in HSI space. In AdaBoost, we evaluate the feature discriminative ability by an entropy-based uncertainty criterion, based on which an Eigen feature subspace is constructed in cluster-specific SVM training. Finally, regression results of multiple SVMs are probabilistic assembled to give the final labeling prediction for test image. We compare our cluster-based cascade SVMs (CSS) RF method in COREL 5,000 database with: 1. Single SVM; 2. Active Learning SVM [5]; 3. Bootstrap Sampling SVM [7]. The superior experimental results demonstrate the efficiency of our algorithm. Rongrong Ji, Hongxun Yao, Pengfei Xu 0001, Xianming Liu 0005 |
ICME | 5 |
| 2008 | Attention-driven action retrieval with DTW-based 3d descriptor matchingabstractFrom visual perception viewpoint, actions in videos can capture high-level semantics for video content understanding and retrieval. However, action-level video retrieval meets great challenges, due to the interferences from global motions or concurrent actions, and the difficulties in robust action describing and matching. This paper presents a content-based action retrieval framework to enable effective search of near-duplicated actions in large-scale video database. Firstly, we present an attention shift model to distill and partition human-concerned saliency actions from global motions and concurrent actions. Secondly, to characterize each saliency action, we extract 3D-SIFT descriptor within its spatial-temporal region, which is robust against rotation, scale, and view point variances. Finally, action similarity is measured using Dynamic Time Warping (DTW) distance to offer tolerance for action duration variance and partial motion missing. Search efficiency in large-scale dataset is achieved by hierarchical descriptor indexing and approximate nearest-neighbor search. In validation, we present a prototype system VILAR to facilitate action search within "Friends" soap operas with excellent accuracy, efficiency, and human perception revealing ability. Rongrong Ji, Xiaoshuai Sun, Hongxun Yao, Pengfei Xu 0001, Tianqiang Liu, Xianming Liu 0005 |
ACM Multimedia | 6 |