EDBT 2026 Demo / reviewers in the wild / expert
Lei Zhang 0006
dblp:64/5666-6
· DBLP profile ↗
510ranked-venue papers
32as first author
177since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 322 · 12 first-author · 136 since 2021Graphics, computer vision, multimedia, augmented reality and games · 321 · 20 first-author · 112 since 2021Applied, interdisciplinary, general and emerging computing · 25 · 2 first-author · 7 since 2021Databases, data management, data science and information retrieval · 12 · 5 since 2021Human-computer interaction and ubiquitous computing · 10 · 1 first-author · 1 since 2021Security and privacy · 9 · 5 since 2021Systems, architecture and hardware · 2 · 2 first-authorSoftware engineering, systems software and programming languages · 2 · 2 since 2021Theory of computation · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Fast Multi-view Consistent 3D Editing with Video PriorsabstractText-driven 3D editing enables user-friendly 3D object or scene editing with text instructions. Due to the lack of multi-view consistency priors, existing methods typically resort to employ 2D generation or editing models to process per-view individually, followed by iterative 2D-3D-2D updating. However, these methods are not only time-consuming but also prone to yielding over-smoothed results, since iterative process averages the different editing signals gathered from different views. In this paper, we propose, an early and pioneering work of generative Video Prior based 3D Editing, ViP3DE in short, to repurpose the temporal consistency priors from pre-trained video generation models to achieve consistent 3D editing within a single forward pass. Our key insight is to condition the video generation model on a single edited view to generate other consistent edited views for 3D updating directly, thereby bypassing iterative editing paradigm. First, 3D updating requires edited views to be paired with specific camera poses. To this end, we propose \textit{motion-preserved noise blending} for the video model to generate edited views at predefined camera poses. In addition, we introduce \textit{geometrically aware denoising} to further enhance multi-view consistency by integrating 3D geometric priors into video models. Extensive experiments demonstrate that our proposed ViP3DE can achieve high-quality 3D editing results even within a single forward pass, significantly outperforming existing methods in both editing quality and editing time cost. Liyi Chen 0002, Ruihuang Li, Guowen Zhang, Pengfei Wang 0012, Lei Zhang 0006 |
AAAI | 5 |
| 2026 | AlignCVC: Aligning Cross-View Consistency for Single-Image-to-3D GenerationabstractSingle-image-to-3D models typically follow a sequential generation and reconstruction workflow. However, intermediate multi-view images synthesized by pre-trained generation models often lack cross-view consistency (CVC), significantly degrading 3D reconstruction performance. While recent methods attempt to refine CVC by feeding reconstruction results back into the multi-view generator, these approaches struggle with noisy and unstable reconstruction outputs that limit effective CVC improvement. We introduce AlignCVC, a novel framework that fundamentally re-frames single-image-to-3D generation through distribution alignment rather than relying on strict regression losses. Our key insight is to align both generated and reconstructed multi-view distributions toward the ground-truth multi-view distribution, establishing a principled foundation for improved CVC. Observing that generated images exhibit weak CVC while reconstructed images display strong CVC due to explicit rendering, we propose a soft-hard alignment strategy with distinct objectives for generation and reconstruction models. This approach not only enhances generation quality but also dramatically accelerates inference to as few as 4 steps. As a plug-and-play paradigm, our method, namely AlignCVC, seamlessly integrates various combinations of multiview generation models with 3D reconstruction models. Extensive experiments demonstrate the effectiveness and efficiency of AlignCVC for single-image-to-3D generation. Zhiyuan Ma 0002, Lingchen Sun, Lei Zhang 0006 |
AAAI | 5 |
| 2026 | BEVDilation: LiDAR-Centric Multi-Modal Fusion for 3D Object DetectionabstractIntegrating LiDAR and camera information in the bird's eye view (BEV) representation has demonstrated its effectiveness in 3D object detection. However, because of the fundamental disparity in geometric accuracy between these sensors, indiscriminate fusion in previous methods often leads to degraded performance. In this paper, we propose BEVDilation, a novel LiDAR-centric framework that prioritizes LiDAR information in the fusion. By formulating image BEV features as implicit guidance rather than naive concatenation, our strategy effectively alleviates the spatial misalignment caused by image depth estimation errors. Furthermore, the image guidance can effectively help the LiDAR-centric paradigm to address the sparsity and semantic limitations of point clouds. Specifically, we propose a Sparse Voxel Dilation Block that mitigates the inherent point sparsity by densifying foreground voxels through image priors. Moreover, we introduce a Semantic-Guided BEV Dilation Block to enhance the LiDAR feature diffusion processing with image semantic guidance and long-range context capture. On the challenging nuScenes benchmark, BEVDilation achieves better performance than state-of-the-art methods while maintaining competitive computational efficiency. Importantly, our LiDAR-centric strategy demonstrates greater robustness to depth noise compared to naive fusion. Guowen Zhang, Chenhang He, Liyi Chen 0002, Lei Zhang 0006 |
AAAI | 4 |
| 2026 | Concretely Efficient Correlated Oblivious Permutation
Xiao Lan, Lei Zhang 0006, Hao Ren 0001, Lin Qu, Yuan Hong 0001 |
AsiaCCS | 4 |
| 2026 | Restoration Adaptation for Semantic Segmentation on Low Quality ImagesabstractAbstract In real-world scenarios, the performance of semantic segmentation often deteriorates when processing low-quality (LQ) images, which may lack clear semantic structures and high-frequency details. Although image restoration techniques offer a promising direction for enhancing degraded visual content, conventional real-world image restoration (Real-IR) models primarily focus on pixel-level fidelity and often fail to recover task-relevant semantic cues, limiting their effectiveness when directly applied to downstream vision tasks. Conversely, existing segmentation models trained on high-quality data lack robustness under real-world degradations. In this paper, we propose Restoration Adaptation for Semantic Segmentation (RASS), which effectively integrates semantic image restoration into the segmentation process, enabling high-quality semantic segmentation on the LQ images directly. Specifically, we first propose a Semantic-Constrained Restoration (SCR) model, which injects segmentation priors into the restoration model by aligning its cross-attention maps with segmentation masks, encouraging semantically faithful image reconstruction. Then, RASS transfers semantic restoration knowledge into segmentation through LoRA-based module merging and task-specific fine-tuning, thereby enhancing the model’s robustness to LQ images. To validate the effectiveness of our framework, we construct a real-world LQ image segmentation dataset with high-quality annotations, and conduct extensive experiments on both synthetic and real-world LQ benchmarks. The results show that SCR and RASS significantly outperform state-of-the-art methods in segmentation and restoration tasks. Code, models, and datasets will be available at https://github.com/Ka1Guan/RASS.git . Rongyuan Wu, Shuai Li 0014, Wentao Zhu 0001, Wenjun Zeng 0001, Lei Zhang 0006 |
Int. J. Comput. Vis. | 6 |
| 2026 | ConSept: Continual semantic segmentation via adapter-based vision transformer
Bowen Dong 0001, Guanglei Yang, Lei Zhang 0006, Wangmeng Zuo |
Pattern Recognit. Lett. | 3 |
| 2025 | MARS: Mixture of Auto-Regressive Models for Fine-grained Text-to-image SynthesisabstractAuto-regressive models have made significant progress in the realm of text-to-image synthesis, yet devising an appropriate model architecture and training strategy to achieve a satisfactory level remains an important avenue of exploration. In this work, we introduce MARS, a novel framework for T2I generation that incorporates a specially designed Semantic Vision-Language Integration Expert (SemVIE). This innovative component integrates pre-trained LLMs by independently processing linguistic and visual information—freezing the textual component while fine-tuning the visual component. This methodology preserves the NLP capabilities of LLMs while imbuing them with exceptional visual understanding. Building upon the powerful base of the pre-trained Qwen-7B, MARS stands out with its bilingual generative capabilities corresponding to both English and Chinese language prompts and the capacity for joint image and text generation. The flexibility of this framework lends itself to migration towards any-to-any task adaptability. Furthermore, MARS employs a multi-stage training strategy that first establishes robust image-text alignment through complementary bidirectional tasks and subsequently concentrates on refining the T2I generation process, significantly augmenting text-image synchrony and the granularity of image details. Notably, MARS requires only 9% of the GPU days needed by SD1.5, yet it achieves remarkable results across a variety of benchmarks, illustrating the training efficiency and the potential for swift deployment in various applications. Wanggui He, Siming Fu, Mushui Liu, Xierui Wang, Wenyi Xiao, Fangxun Shu, Yi Wang 0068, Lei Zhang 0006, Zhelun Yu, Haoyuan Li 0002, Ziwei Huang 0005, Leilei Gan, Hao Jiang 0014 |
AAAI | 8 |
| 2025 | SyncNoise: Geometrically Consistent Noise Prediction for Instruction-based 3D EditingabstractText-based 2D diffusion models have demonstrated impressive capabilities in image generation and editing. Meanwhile, the 2D diffusion models also exhibit substantial potentials for 3D editing tasks. However, how to achieve consistent edits across multiple viewpoints remains a challenge. While the iterative dataset update method is capable of achieving global consistency, it suffers from slow convergence and over-smoothed textures. We propose SyncNoise, a novel geometry-guided multi-view consistent noise editing approach for high-fidelity 3D scene editing. SyncNoise synchronously edits multiple views with 2D diffusion models while enforcing multi-view noise predictions to be geometrically consistent, which ensures global consistency in both semantic structure and low-frequency appearance. To further enhance local consistency in high-frequency details, we set a group of anchor views and propagate them to their neighboring frames through cross-view reprojection. To improve the reliability of multi-view correspondences, we introduce depth supervision during training to enhance the reconstruction of precise geometries. Our method achieves high-quality 3D editing results respecting the textual instructions, especially in scenes with complex textures, by enhancing geometric consistency at the noise and pixel levels. Ruihuang Li, Liyi Chen 0002, Zhengqiang Zhang, Varun Jampani, Vishal M. Patel, Lei Zhang 0006 |
AAAI | 6 |
| 2025 | Progressive Rendering Distillation: Adapting Stable Diffusion for Instant Text-to-Mesh Generation without 3D DataabstractIt is highly desirable to obtain a model that can generate high-quality 3D meshes from text prompts in just seconds. While recent attempts have adapted pre-trained text-to-image diffusion models, such as Stable Diffusion (SD), into generators of 3D representations (e.g., Triplane), they often suffer from poor quality due to the lack of sufficient high-quality 3D training data. Aiming at overcoming the data shortage, we propose a novel training scheme, termed as Progressive Rendering Distillation (PRD), eliminating the need for 3D ground-truths by distilling multi-view diffusion models and adapting SD into a native 3D generator. In each iteration of training, PRD uses the U-Net to progressively denoise the latent from random noise for a few steps, and in each step it decodes the denoised latent into 3D output. Multi-view diffusion models, including MVDream and RichDreamer, are used in joint with SD to distill text-consistent textures and geometries into the 3D outputs through score distillation. Since PRD supports training without 3D ground-truths, we can easily scale up the training data and improve generation quality for challenging text prompts with creative concepts. Meanwhile, PRD can accelerate the inference speed of the generation model in just a few steps. With PRD, we train a Triplane generator, namely TriplaneTurbo, which adds only 2.5% trainable parameters to adapt SD for Triplane generation. TriplaneTurbo outperforms previous text-to-3D generators in both efficiency and quality. Specifically, it can produce high-quality 3D meshes in 1.2 seconds and generalize well for challenging text input. The code is available at github.com/theEricMa/TriplaneTurbo. Zhiyuan Ma 0002, Rongyuan Wu, Xiangyu Zhu 0001, Zhen Lei 0001, Lei Zhang 0006 |
CVPR | 6 |
| 2025 | Toward Generalized Image Quality Assessment: Relaxing the Perfect Reference Quality AssumptionabstractFull-Reference image quality assessment (FR-IQA) generally assumes that reference images are of perfect quality. However, this assumption is flawed due to the sensor and optical limitations of modern imaging systems. Moreover, recent generative enhancement methods are capable of producing images of higher quality than their original. All of these challenge the effectiveness and applicability of current FR-IQA models. To relax the assumption of perfect reference image quality, we build a large-scale IQA database, namely DiffIQA, containing approximately 180,000 images generated by a diffusion-based image enhancer with adjustable hyper-parameters. Each image is annotated by human subjects as either worse, similar, or better quality compared to its reference. Building on this, we present a generalized FR-IQA model, namely Adaptive Fidelity-Naturalness Evaluator (A-FINE), to accurately assess and adaptively combine the fidelity and naturalness of a test image. A-FINE aligns well with standard FR-IQA when the reference image is much more natural than the test image. We demonstrate by extensive experiments that A-FINE surpasses standard FR-IQA models on well-established IQA datasets and our newly created DiffIQA. To further validate A-FINE, we additionally construct a super-resolution IQA benchmark (SRIQA-Bench), encompassing test images derived from ten state-of-the-art SR methods with reliable human quality annotations. Tests on SRIQA-Bench re-affirm the advantages of A-FINE. The code and dataset are available at https://tianhewu.github.io/A-FINEpage.github.io/. Tianhe Wu, Kede Ma, Lei Zhang 0006 |
CVPR | 4 |
| 2025 | RORem: Training a Robust Object Remover with Human-in-the-LoopabstractDespite the significant advancements, existing object removal methods struggle with incomplete removal, incorrect content synthesis and blurry synthesized regions, resulting in low success rates. Such issues are mainly caused by the lack of high-quality paired training data, as well as the self-supervised training paradigm adopted in these methods, which forces the model to in-paint the masked regions, leading to ambiguity between synthesizing the masked objects and restoring the background. To address these issues, we propose a semi-supervised learning strategy with human-in-the-loop to create high-quality paired training data, aiming to train a Robust Object Remover (RORem). We first collect 60K training pairs from open-source datasets to train an initial object removal model for generating removal samples, and then utilize human feedback to select a set of high-quality object removal pairs, with which we train a discriminator to automate the following training data generation process. By iterating this process for several rounds, we finally obtain a substantial object removal dataset with over 200K pairs. Fine-tuning the pre-trained stable diffusion model with this dataset, we obtain our RORem, which demonstrates state-of-the-art object removal performance in terms of both reliability and image quality. Particularly, RORem improves the object removal success rate over previous methods by more than 18%. The dataset, source code and trained model are available at https://github.com/leeruibin/RORem. Ruibin Li, Tao Yang 0042, Song Guo 0001, Lei Zhang 0006 |
CVPR | 4 |
| 2025 | Pixel-level and Semantic-level Adjustable Super-resolution: A Dual-LoRA ApproachabstractDiffusion prior-based methods have shown impressive results in real-world image super-resolution (SR). However, most existing methods entangle pixel-level and semantic-level SR objectives in the training process, struggling to balance pixel-wise fidelity and perceptual quality. Meanwhile, users have varying preferences on SR results, thus it is demanded to develop an adjustable SR model that can be tailored to different fidelity-perception preferences during inference without re-training. We present Pixel-level and Semantic-level Adjustable SR (PiSA-SR), which learns two LoRA modules upon the pre-trained stable-diffusion (SD) model to achieve improved and adjustable SR results. We first formulate the SD-based SR problem as learning the residual between the low-quality input and the high-quality output, then show that the learning objective can be decoupled into two distinct LoRA weight spaces: one is characterized by the ℓ2-loss for pixel-level regression, and another is characterized by the LPIPS and classifier score distillation losses to extract semantic information from pre-trained classification and SD models. In its default setting, PiSA-SR can be performed in a single diffusion step, achieving leading real-world SR results in both quality and efficiency. By introducing two adjustable guidance scales on the two LoRA modules to control the strengths of pixel-wise fidelity and semantic-level details during inference, PiSA-SR can offer flexible SR results according to user preference without re-training. The source code of our method can be found at https://github.com/csslc/PiSA-SR. Lingchen Sun, Rongyuan Wu, Zhiyuan Ma 0002, Shuaizheng Liu, Qiaosi Yi, Lei Zhang 0006 |
CVPR | 6 |
| 2025 | Generalized and Efficient 2D Gaussian Splatting for Arbitrary-Scale Super-Resolution
Liyi Chen 0002, Zhengqiang Zhang, Lei Zhang 0006 |
ICCV | 4 |
| 2025 | Integrating Visual Interpretation and Linguistic Reasoning for Geometric Problem Solving
Zixian Guo, Ming Liu 0018, Qilong Wang 0001, Zhilong Ji, Jinfeng Bai, Lei Zhang 0006, Wangmeng Zuo |
ICCV | 6 |
| 2025 | FiVE-Bench: A Fine-Grained Video Editing Benchmark for Evaluating Emerging Diffusion and Rectified Flow Models
Minghan Li 0001, Lei Zhang 0006, Mengyu Wang 0001 |
ICCV | 4 |
| 2025 | InsViE-1M: Effective Instruction-Based Video Editing with Elaborate Dataset ConstructionabstractInstruction-based video editing allows effective and interactive editing of videos using only instructions without extra inputs such as masks or attributes. However, collecting high-quality training triplets (source video, edited video, instruction) is a challenging task. Existing datasets mostly consist of low-resolution, short duration, and limited amount of source videos with unsatisfactory editing quality, limiting the performance of trained editing models. In this work, we present a high-quality Instruction-based Video Editing dataset with 1M triplets, namely InsViE-1M. We first curate high-resolution and high-quality source videos and images, then design an effective editing-filtering pipeline to construct high-quality editing triplets for model training. For a source video, we generate multiple edited samples of its first frame with different intensities of classifier-free guidance, which are automatically filtered by GPT-4o with carefully crafted guidelines. The edited first frame is propagated to subsequent frames to produce the edited video, followed by another round of filtering for frame quality and motion evaluation. We also generate and filter a variety of video editing triplets from high-quality images. With the InsViE-1M dataset, we propose a multi-stage learning strategy to train our InsViE model, progressively enhancing its instruction following and editing ability. Extensive experiments demonstrate the advantages of our InsViE-1M dataset and the trained model over state-of-the-art works. Codes are available at \href{https://github.com/langmanbusi/InsViE}{InsViE}. Yuhui Wu 0001, Liyi Chen 0002, Ruibin Li, Lei Zhang 0006 |
ICCV | 6 |
| 2025 | Fine-Structure Preserved Real-World Image Super-Resolution Via Transfer Vae Training
Qiaosi Yi, Shuai Liu 0009, Rongyuan Wu, Lingchen Sun, Yuhui Wu 0001, Lei Zhang 0006 |
ICCV | 6 |
| 2025 | Toward Generalizing Visual Brain Decoding to Unseen SubjectsabstractVisual brain decoding aims to decode visual information from human brain activities. Despite the great progress, one critical limitation of current brain decoding research lies in the lack of generalization capability to unseen subjects. Prior work typically focuses on decoding brain activity of individuals based on the observation that different subjects exhibit different brain activities, while it remains unclear whether brain decoding can be generalized to unseen subjects. This study aims to answer this question. We first consolidate an image-fMRI dataset consisting of stimulus-image and fMRI-response pairs, involving 177 subjects in the movie-viewing task of the Human Connectome Project (HCP). This dataset allows us to investigate the brain decoding performance with the increase of participants. We then present a learning paradigm that applies uniform processing across all subjects, instead of employing different network heads or tokenizers for individuals as in previous methods, so that we can accommodate a large number of subjects to explore the generalization capability across different subjects. A series of experiments are conducted and we have the following findings. First, the network exhibits clear generalization capabilities with the increase of training subjects. Second, the generalization capability is common to popular network architectures (MLP, CNN and Transformer). Third, the generalization performance is affected by the similarity between subjects. Our findings reveal the inherent similarities in brain activities across individuals. With the emergence of larger and more comprehensive datasets, it is possible to train a brain decoding foundation model in the future. Codes and models can be found at https://github.com/Xiangtaokong/TGBD}{https://github.com/Xiangtaokong/TGBD. Xiangtao Kong, Lei Zhang 0006 |
ICLR | 4 |
| 2025 | Visual-O1: Understanding Ambiguous Instructions via Multi-modal Multi-turn Chain-of-thoughts ReasoningabstractAs large-scale models evolve, language instructions are increasingly utilized in multi-modal tasks. Due to human language habits, these instructions often contain ambiguities in real-world scenarios, necessitating the integration of visual context or common sense for accurate interpretation. However, even highly intelligent large models exhibit observable performance limitations on ambiguous instructions, where weak reasoning abilities of disambiguation can lead to catastrophic errors. To address this issue, this paper proposes Visual-O1, a multi-modal multi-turn chain-of-thought reasoning framework. It simulates human multi-modal multi-turn reasoning, providing instantial experience for highly intelligent models or empirical experience for generally intelligent models to understand ambiguous instructions. Unlike traditional methods that require models to possess high intelligence to understand long texts or perform lengthy complex reasoning, our framework does not notably increase computational overhead and is more general and effective, even for generally intelligent models. Experiments show that our method not only enhances the performance of models of different intelligence levels on ambiguous instructions but also improves their performance on general datasets. Our work highlights the potential of artificial intelligence to work like humans in real-world scenarios with uncertainty and ambiguity. We release our data and code at https://github.com/kodenii/Visual-O1. Minheng Ni, Yutao Fan, Lei Zhang 0006, Wangmeng Zuo |
ICLR | 3 |
| 2025 | LLaVA-MoD: Making LLaVA Tiny via MoE-Knowledge DistillationabstractWe introduce LLaVA-MoD, a novel framework designed to enable the efficient training of small-scale Multimodal Language Models ($s$-MLLM) distilling knowledge from large-scale MLLM ($l$-MLLM). Our approach tackles two fundamental challenges in MLLM distillation. First, we optimize the network structure of $s$-MLLM by integrating a sparse Mixture of Experts (MoE) architecture into the language model, striking a balance between computational efficiency and model expressiveness. Second, we propose a progressive knowledge transfer strategy for comprehensive knowledge transfer. This strategy begins with mimic distillation, where we minimize the Kullback-Leibler (KL) divergence between output distributions to enable $s$-MLLM to emulate $s$-MLLM's understanding. Following this, we introduce preference distillation via Preference Optimization (PO), where the key lies in treating $l$-MLLM as the reference model. During this phase, the $s$-MLLM's ability to discriminate between superior and inferior examples is significantly enhanced beyond $l$-MLLM, leading to a better $s$-MLLM that surpasses $l$-MLLM, particularly in hallucination benchmarks.
Extensive experiments demonstrate that LLaVA-MoD surpasses existing works across various benchmarks while maintaining a minimal activated parameters and low computational costs. Remarkably, LLaVA-MoD-2B surpasses Qwen-VL-Chat-7B with an average gain of 8.8\%, using merely $0.3\%$ of the training data and 23\% trainable parameters. The results underscore LLaVA-MoD's ability to effectively distill comprehensive knowledge from its teacher model, paving the way for developing efficient MLLMs. Fangxun Shu, Yue Liao, Lei Zhang 0006, Le Zhuo, Chenning Xu, Long Chan, Zhelun Yu, Wanggui He, Siming Fu, Haoyuan Li 0002, Si Liu 0001, Hongsheng Li 0001, Hao Jiang 0062 |
ICLR | 3 |
| 2025 | Spatial-Mamba: Effective Visual State Space Models via Structure-Aware State FusionabstractSelective state space models (SSMs), such as Mamba, highly excel at capturing long-range dependencies in 1D sequential data, while their applications to 2D vision tasks still face challenges. Current visual SSMs often convert images into 1D sequences and employ various scanning patterns to incorporate local spatial dependencies. However, these methods are limited in effectively capturing the complex image spatial structures and the increased computational cost caused by the lengthened scanning paths. To address these limitations, we propose Spatial-Mamba, a novel approach that establishes neighborhood connectivity directly in the state space. Instead of relying solely on sequential state transitions, we introduce a structure-aware state fusion equation, which leverages dilated convolutions to capture image spatial structural dependencies, significantly enhancing the flow of visual contextual information. Spatial-Mamba proceeds in three stages: initial state computation in a unidirectional scan, spatial context acquisition through structure-aware state fusion, and final state computation using the observation equation. Our theoretical analysis shows that Spatial-Mamba unifies the original Mamba and linear attention under the same matrix multiplication framework, providing a deeper understanding of our method. Experimental results demonstrate that Spatial-Mamba, even with a single scan, attains or surpasses the state-of-the-art SSM-based models in image classification, detection and segmentation. Source codes and trained models can be found at \url{ https://github.com/EdwardChasel/Spatial-Mamba }. Chaodong Xiao, Minghan Li 0001, Zhengqiang Zhang, Deyu Meng, Lei Zhang 0006 |
ICLR | 5 |
| 2025 | FreCaS: Efficient Higher-Resolution Image Generation via Frequency-aware Cascaded SamplingabstractWhile image generation with diffusion models has achieved a great success, generating images of higher resolution than the training size remains a challenging task due to the high computational cost. Current methods typically perform the entire sampling process at full resolution and process all frequency components simultaneously, contradicting with the inherent coarse-to-fine nature of latent diffusion models and wasting computations on processing premature high-frequency details at early diffusion stages. To address this issue, we introduce an efficient $\textbf{Fre}$quency-aware $\textbf{Ca}$scaded $\textbf{S}$ampling framework, $\textbf{FreCaS}$ in short, for higher-resolution image generation. FreCaS decomposes the sampling process into cascaded stages with gradually increased resolutions, progressively expanding frequency bands and refining the corresponding details. We propose an innovative frequency-aware classifier-free guidance (FA-CFG) strategy to assign different guidance strengths for different frequency components, directing the diffusion model to add new details in the expanded frequency domain of each stage. Additionally, we fuse the cross-attention maps of previous and current stages to avoid synthesizing unfaithful layouts. Experiments demonstrate that FreCaS significantly outperforms state-of-the-art methods in image quality and generation speed. In particular, FreCaS is about 2.86$\times$ and 6.07$\times$ faster than ScaleCrafter and DemoFusion in generating a 2048$\times$2048 image using a pretrained SDXL model and achieves an $\text{FID}_b$ improvement of 11.6 and 3.7, respectively. FreCaS can be easily extended to more complex models such as SD3. The source code of FreCaS can be found at https://github.com/xtudbxk/FreCaS. Zhengqiang Zhang, Ruihuang Li, Lei Zhang 0006 |
ICLR | 3 |
| 2025 | Algernon: A Flag-Guided Hybrid Fuzzer for Unlocking Hidden Program PathsabstractFuzz testing is a widely used method for finding security issues in software. However, certain code paths can only be explored under specific program states. Flag variables, which represent internal states, are crucial in influencing program behavior through flag-guarded branches. Unfortunately, existing fuzzing tools struggle to efficiently explore them due to the implicit data dependency between flag variables and the input. As a result, they commonly lack awareness of the dependency between program input and the assignments of critical flag variables, leading to a blind or random approach to satisfy flag-checking constraints, which greatly impacts the fuzzing efficiency.To address this issue, this paper proposes a dynamic flag-guided hybrid fuzzing approach, which automates the identification of flag variables and provides guidance for fuzz testing. Specifically, we first design a pre-fuzzing program analysis to recognize flag variables and a novel data structure to present how flag variables guard code branches. Then, we propose a new constraint-solving approach by separating complex flag-checking constraints into a set of atomic ones and sequentially solving them by traversing our FDG to locate execution paths that could assign the flag variables with the desired values.We implement a prototype tool, called Algernon, and evaluate it on 20 popular open-source programs. Across all tested programs, Algernon outperforms QSYM, Angora, AFL++, and INVSCOV in terms of both code coverage and vulnerability discovery, demonstrating the effectiveness of our approach. During our experiments, Algernon successfully found 30 zero-day vulnerabilities with 11 CVE IDs assigned. Lei Zhang 0006, Jingqi Long, Wenzheng Hong, Zhemin Yang, Yuan Zhang 0009, Donglai Zhu, Min Yang 0002 |
ASE | 2 |
| 2025 | Exploring Static Taint Analysis in LLMs: A Dynamic Benchmarking Framework for Measurement and EnhancementabstractLLMs offer a promising avenue to overcome the limitations of traditional taint analysis techniques, with a growing number of studies leveraging LLMs for taint analysis and its downstream applications. However, these studies lack a systematic understanding of LLMs’ taint analysis capabilities, limiting their transferability and reliability. To bridge this gap and better apply LLMs to static taint analysis, we aim to comprehensively measure and understand LLMs’ taint analysis capabilities.Using existing benchmarks is a straightforward approach, but they are unsuitable due to issues such as training data leakage, not accounting for LLMs’ features, and improper assessment criteria. Manually constructing new benchmarks is not only labor-intensive but also struggles to remain effective as LLMs evolve. To address these, we propose LLMCapLens, a dynamic benchmark generation framework to systematically measure and enhance LLMs’ capabilities. LLMCapLens models influencing factors of LLMs’ taint analysis capabilities, employing a Basic Unit-Based generation method and a lightweight dynamic taint analysis-based verification method to implement the automated generation of targeted benchmarks, ensuring both diversity and correctness. Furthermore, LLMCapLens proposes a measurement-driven, training-free, model-specific enhancement approach.We apply LLMCapLens to 10 mainstream LLMs, revealing how they perform under various influencing factors and identifying unique characteristics, such as the underlying error causes for each model. Notably, our enhancement approach significantly improves LLM performance—GPT-4 Turbo, for instance, achieved improvements across 16 out of 19 factors, with an average True Negative Rate increase of 21.29%. Finally, we validate the real-world impact of our method by applying enhanced LLMs to vulnerability detection, demonstrating a substantial improvement over prior approaches. Lei Zhang 0006, Keke Lian, Fute Sun, Bofei Chen, Yongheng Liu, Zhiyu Wu, Yuan Zhang 0009, Min Yang 0002 |
ASE | 2 |
| 2025 | Knowledge Regularized Negative Feature Tuning of Vision-Language Models for Out-of-Distribution DetectionabstractOut-of-distribution (OOD) detection is crucial for building reliable machine learning models. Although negative prompt tuning has enhanced the OOD detection capabilities of vision-language models, these tuned models often suffer from reduced generalization performance on unseen classes and styles. To address this challenge, we propose a novel method called Knowledge Regularized Negative Feature Tuning (KR-NFT), which integrates an innovative adaptation architecture termed Negative Feature Tuning (NFT) and a corresponding knowledge-regularization (KR) optimization strategy. Specifically, NFT applies distribution-aware transformations to pre-trained text features, effectively separating positive and negative features into distinct spaces. This separation maximizes the distinction between in-distribution (ID) and OOD images. Additionally, we introduce image-conditional learnable factors through a lightweight meta-network, enabling dynamic adaptation to individual images and mitigating sensitivity to class and style shifts. Compared to traditional negative prompt tuning, NFT demonstrates superior efficiency and scalability. To optimize this adaptation architecture, the KR optimization strategy is designed to enhance the discrimination between ID and OOD sets while mitigating pre-trained knowledge forgetting. This enhances OOD detection performance on trained ID classes while simultaneously improving OOD detection on unseen ID datasets. Notably, when trained with few-shot samples from ImageNet dataset, KR-NFT not only improves ID classification accuracy and OOD detection but also significantly reduces the FPR95 by 5.44% under an unexplored generalization setting with unseen ID categories. Codes can be found at https://github.com/ZhuWenjie98/KRNFT. Wenjie Zhu 0003, Yabin Zhang 0001, Xin Jin 0014, Wenjun Zeng 0001, Lei Zhang 0006 |
ACM Multimedia | 5 |
| 2025 | MIRAGE: Assessing Hallucination in Multimodal Reasoning Chains of MLLMabstractMultimodal hallucination in multimodal large language models (MLLMs) restricts the correctness of MLLMs. However, multimodal hallucinations are multi-sourced and arise from diverse causes. Existing benchmarks fail to adequately distinguish between perception-induced hallucinations and reasoning-induced hallucinations. This failure constitutes a significant issue and hinders the diagnosis of multimodal reasoning failures within MLLMs. To address this, we propose the MIRAGE benchmark, which isolates reasoning hallucinations by constructing questions where input images are correctly perceived by MLLMs yet reasoning errors persist. MIRAGE introduces multi-granular evaluation metrics: accuracy, factuality, and LLMs hallucination score for hallucination quantification. Our analysis reveals strong correlations between question types and specific hallucination patterns, particularly systematic failures of MLLMs in spatial reasoning involving complex relationships (\emph{e.g.}, complex geometric patterns across images). This highlights a critical limitation in the reasoning capabilities of current MLLMs and provides targeted insights for hallucination mitigation on specific types. To address these challenges, we propose Logos, a method that combines curriculum reinforcement fine-tuning to encourage models to generate logic-consistent reasoning chains by stepwise reducing learning difficulty, and collaborative hint inference to reduce reasoning complexity. Logos establishes a baseline on MIRAGE, and reduces the logical hallucinations in original base models. Link: \url{https://bit.ly/25mirage}. Bowen Dong 0001, Minheng Ni, Zitong Huang, Guanglei Yang, Wangmeng Zuo, Lei Zhang 0006 |
NeurIPS | 6 |
| 2025 | Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and VideosabstractWe present Perceive Anything Model (PAM), a conceptually straightforward and efficient framework for comprehensive region-level visual understanding in images and videos. Our approach extends the powerful segmentation model SAM 2 by integrating Large Language Models (LLMs), enabling simultaneous object segmentation with the generation of diverse, region-specific semantic outputs, including categories, label definition, functional explanations, and detailed captions. A key component, Semantic Perceiver, is introduced to efficiently transform SAM 2's rich visual features, which inherently carry general vision, localization, and semantic priors into multi-modal tokens for LLM comprehension. To support robust multi-granularity understanding, we also develop a dedicated data refinement and augmentation pipeline, yielding a high-quality dataset of 1.5M image and 0.6M video region-semantic annotations, including novel region-level streaming video caption data. PAM is designed for lightweightness and efficiency, while also demonstrates strong performance across a diverse range of region understanding tasks. It runs 1.2$-$2.4$\times$ faster and consumes less GPU memory than prior approaches, offering a practical solution for real-world applications. We believe that our effective approach will serve as a strong baseline for future research in region-level visual understanding. Weifeng Lin, Ruichuan An, Tianhe Ren, Renrui Zhang, Wentao Zhang 0001, Lei Zhang 0006, Hongsheng Li 0001 |
NeurIPS | 9 |
| 2025 | InstructRestore: Region-Customized Image Restoration with Human InstructionsabstractDespite the significant progress in diffusion prior-based image restoration for real-world scenarios, most existing methods apply uniform processing to the entire image, lacking the capability to perform region-customized image restoration according to user preferences. In this work, we propose a new framework, namely InstructRestore, to perform region-adjustable image restoration following human instructions. To achieve this, we first develop a data generation engine to produce training triplets, each consisting of a high-quality image, the target region description, and the corresponding region mask. With this engine and careful data screening, we construct a comprehensive dataset comprising 536,945 triplets to support the training and evaluation of this task. We then examine how to integrate the low-quality image features under the ControlNet architecture to adjust the degree of image details enhancement. Consequently, we develop a ControlNet-like model to identify the target region and allocate different integration scales to the target and surrounding regions, enabling region-customized image restoration that aligns with user instructions. Experimental results demonstrate that our proposed InstructRestore approach enables effective human-instructed image restoration, including restoration with controllable bokeh blur effects and region-specific restoration with continuous intensity control. Our work advances the investigation of interactive image restoration and enhancement techniques. Data, code, and models are publicly available at https://github.com/shuaizhengliu/InstructRestore.git. Shuaizheng Liu, Jianqi Ma, Lingchen Sun, Xiangtao Kong, Lei Zhang 0006 |
NeurIPS | 5 |
| 2025 | BurstDeflicker: A Benchmark Dataset for Flicker Removal in Dynamic ScenesabstractFlicker artifacts in short-exposure images are caused by the interplay between the row-wise exposure mechanism of rolling shutter cameras and the temporal intensity variations of alternating current (AC)-powered lighting. These artifacts typically appear as uneven brightness distribution across the image, forming noticeable dark bands. Beyond compromising image quality, this structured noise also affects high-level tasks, such as object detection and tracking, where reliable lighting is crucial. Despite the prevalence of flicker, the lack of a large-scale, realistic dataset has been a significant barrier to advancing research in flicker removal. To address this issue, we present BurstDeflicker, a scalable benchmark constructed using three complementary data acquisition strategies. First, we develop a Retinex-based synthesis pipeline that redefines the goal of flicker removal and enables controllable manipulation of key flicker-related attributes (e.g., intensity, area, and frequency), thereby facilitating the generation of diverse flicker patterns. Second, we capture 4,000 real-world flicker images from different scenes, which help the model better understand the spatial and temporal characteristics of real flicker artifacts and generalize more effectively to wild scenarios. Finally, due to the non-repeatable nature of dynamic scenes, we propose a green-screen method to incorporate motion into image pairs while preserving real flicker degradation. Comprehensive experiments demonstrate the effectiveness of our dataset and its potential to advance research in flicker removal. Lishen Qu, Shihao Zhou 0003, Yaqi Luo, Jie Liang 0007, Hui Zeng 0001, Lei Zhang 0006, Jufeng Yang |
NeurIPS | 7 |
| 2025 | One-Step Diffusion for Detail-Rich and Temporally Consistent Video Super-ResolutionabstractIt is a challenging problem to reproduce rich spatial details while maintaining temporal consistency in real-world video super-resolution (Real-VSR), especially when we leverage pre-trained generative models such as stable diffusion (SD) for realistic details synthesis. Existing SD-based Real-VSR methods often compromise spatial details for temporal coherence, resulting in suboptimal visual quality.
We argue that the key lies in how to effectively extract the degradation-robust temporal consistency priors from the low-quality (LQ) input video and enhance the video details while maintaining the extracted consistency priors.
To achieve this, we propose a Dual LoRA Learning (DLoRAL) paradigm to train an effective SD-based one-step diffusion model, achieving realistic frame details and temporal consistency simultaneously.
Specifically, we introduce a Cross-Frame Retrieval (CFR) module to aggregate complementary information across frames, and train a Consistency-LoRA (C-LoRA) to learn robust temporal representations from degraded inputs.
After consistency learning, we fix the CFR and C-LoRA modules and train a Detail-LoRA (D-LoRA) to enhance spatial details while aligning with the temporal space defined by C-LoRA to keep temporal coherence.
The two phases alternate iteratively for optimization, collaboratively delivering consistent and detail-rich outputs. During inference, the two LoRA branches are merged into the SD model, allowing efficient and high-quality video restoration in a single diffusion step. Experiments show that DLoRAL achieves strong performance in both accuracy and speed. Code and models will be released. Lingchen Sun, Shuaizheng Liu, Rongyuan Wu, Zhengqiang Zhang, Lei Zhang 0006 |
NeurIPS | 6 |
| 2025 | DP²O-SR: Direct Perceptual Preference Optimization for Real-World Image Super-ResolutionabstractBenefiting from pre-trained text-to-image (T2I) diffusion models, real-world image super-resolution (Real-ISR) methods can synthesize rich and realistic details. However, due to the inherent stochasticity of T2I models, different noise inputs often lead to outputs with varying perceptual quality. Although this randomness is sometimes seen as a limitation, it also introduces a wider perceptual quality range, which can be exploited to improve Real-ISR performance. To this end, we introduce Direct Perceptual Preference Optimization for Real-ISR (DP²O-SR), a framework that aligns generative models with perceptual preferences without requiring costly human annotations. We construct a hybrid reward signal by combining full-reference and no-reference image quality assessment (IQA) models trained on large-scale human preference datasets. This reward encourages both structural fidelity and natural appearance. To better utilize perceptual diversity, we move beyond the standard best-vs-worst selection and construct multiple preference pairs from outputs of the same model. Our analysis reveals that the optimal selection ratio depends on model capacity: smaller models benefit from broader coverage, while larger models respond better to stronger contrast in supervision. Furthermore, we propose hierarchical preference optimization, which adaptively weights training pairs based on intra-group reward gaps and inter-group diversity, enabling more efficient and stable learning. Extensive experiments across both diffusion- and flow-based T2I backbones demonstrate that DP²O-SR significantly improves perceptual quality and generalizes well to real-world benchmarks. Rongyuan Wu, Lingchen Sun, Zhengqiang Zhang, Tianhe Wu, Qiaosi Yi, Shuai Li 0014, Lei Zhang 0006 |
NeurIPS | 8 |
| 2025 | VisualQuality-R1: Reasoning-Induced Image Quality Assessment via Reinforcement Learning to RankabstractDeepSeek-R1 has demonstrated remarkable effectiveness in incentivizing reasoning and generalization capabilities of large language models (LLMs) through reinforcement learning. Nevertheless, the potential of reasoning-induced computation has not been thoroughly explored in the context of image quality assessment (IQA), a task depending critically on visual reasoning. In this paper, we introduce VisualQuality-R1, a reasoning-induced no-reference IQA (NR-IQA) model, and we train it with reinforcement learning to rank, a learning algorithm tailored to the intrinsically relative nature of visual quality. Specifically, for a pair of images, we employ group relative policy optimization to generate multiple quality scores for each image. These estimates are used to compute comparative probabilities of one image having higher quality than the other under the Thurstone model. Rewards for each quality estimate are defined using continuous fidelity measures rather than discretized binary labels. Extensive experiments show that the proposed VisualQuality-R1 consistently outperforms discriminative deep learning-based NR-IQA models as well as a recent reasoning-induced quality regression method. Moreover, VisualQuality-R1 is capable of generating contextually rich, human-aligned quality descriptions, and supports multi-dataset training without requiring perceptual scale realignment. These features make VisualQuality-R1 especially well-suited for reliably measuring progress in a wide range of image processing tasks like super-resolution and image generation. Tianhe Wu, Jie Liang 0007, Lei Zhang 0006, Kede Ma |
NeurIPS | 4 |
| 2025 | DNAEdit: Direct Noise Alignment for Text-Guided Rectified Flow EditingabstractLeveraging the powerful generation capability of large-scale pretrained text-to-image models, training-free methods have demonstrated impressive image editing results. Conventional diffusion-based methods, as well as recent rectified flow (RF)-based methods, typically reverse synthesis trajectories by gradually adding noise to clean images, during which the noisy latent at the current timestep is used to approximate that at the next timesteps, introducing accumulated drift and degrading reconstruction accuracy. Considering the fact that in RF the noisy latent is estimated through direct interpolation between Gaussian noises and clean images at each timestep, we propose Direct Noise Alignment (DNA), which directly refines the desired Gaussian noise in the noise domain, significantly reducing the error accumulation in previous methods. Specifically, DNA estimates the velocity field of the interpolated noised latent at each timestep and adjusts the Gaussian noise by computing the difference between the predicted and expected velocity field. We validate the effectiveness of DNA and reveal its relationship with existing RF-based inversion methods. Additionally, we introduce a Mobile Velocity Guidance (MVG) to control the target prompt-guided generation process, balancing image background preservation and target object editability. DNA and MVG collectively constitute our proposed method, namely DNAEdit. Finally, we introduce DNA-Bench, a long-prompt benchmark, to evaluate the performance of advanced image editing models. Experimental results demonstrate that our DNAEdit achieves superior performance to state-of-the-art text-guided editing methods. Our code, model, and benchmark will be made publicly available. Minghan Li 0001, Shuai Li 0014, Yuhui Wu 0001, Qiaosi Yi, Lei Zhang 0006 |
NeurIPS | 6 |
| 2025 | Registration is a Powerful Rotation-Invariance Learner for 3D Anomaly Detectionabstract3D anomaly detection in point-cloud data is critical for industrial quality control, aiming to identify structural defects with high reliability. However, current memory bank-based methods often suffer from inconsistent feature transformations and limited discriminative capacity, particularly in capturing local geometric details and achieving rotation invariance. These limitations become more pronounced when registration fails, leading to unreliable detection results. We argue that point-cloud registration plays an essential role not only in aligning geometric structures but also in guiding feature extraction toward rotation-invariant and locally discriminative representations. To this end, we propose a registration-induced, rotation-invariant feature extraction framework that integrates the objectives of point-cloud registration and memory-based anomaly detection. Our key insight is that both tasks rely on modeling local geometric structures and leveraging feature similarity across samples. By embedding feature extraction into the registration learning process, our framework jointly optimizes alignment and representation learning. This integration enables the network to acquire features that are both robust to rotations and highly effective for anomaly detection. Extensive experiments on the Anomaly-ShapeNet and Real3D-AD datasets demonstrate that our method consistently outperforms existing approaches in effectiveness and generalizability. Yuyang Yu, Zhengwei Chen, Xuemiao Xu, Lei Zhang 0006, Haoxin Yang, Yongwei Nie, Shengfeng He |
NeurIPS | 4 |
| 2025 | GPSToken: Gaussian Parameterized Spatially-adaptive Tokenization for Image Representation and GenerationabstractEffective and efficient tokenization plays an important role in image representation and generation. Conventional methods, constrained by uniform 2D/1D grid tokenization, are inflexible to represent regions with varying shapes and textures and at different locations, limiting their efficacy of feature representation. In this work, we propose **GPSToken**, a novel **G**aussian **P**arameterized **S**patially-adaptive **Token**ization framework, to achieve non-uniform image tokenization by leveraging parametric 2D Gaussians to dynamically model the shape, position, and textures of different image regions. We first employ an entropy-driven algorithm to partition the image into texture-homogeneous regions of variable sizes. Then, we parameterize each region as a 2D Gaussian (mean for position, covariance for shape) coupled with texture features. A specialized transformer is trained to optimize the Gaussian parameters, enabling continuous adaptation of position/shape and content-aware feature extraction. During decoding, Gaussian parameterized tokens are reconstructed into 2D feature maps through a differentiable splatting-based renderer, bridging our adaptive tokenization with standard decoders for end-to-end training. GPSToken disentangles spatial layout (Gaussian parameters) from texture features to enable efficient two-stage generation: structural layout synthesis using lightweight networks, followed by structure-conditioned texture generation. Experiments demonstrate the state-of-the-art performance of GPSToken, which achieves rFID and FID scores of 0.65 and 1.50 on image reconstruction and generation tasks using 128 tokens, respectively. Codes and models of GPSToken can be found at https://github.com/xtudbxk/GPSToken. Zhengqiang Zhang, Rongyuan Wu, Lingchen Sun, Lei Zhang 0006 |
NeurIPS | 4 |
| 2025 | Polyline Path Masked Attention for Vision TransformerabstractGlobal dependency modeling and spatial position modeling are two core issues of the foundational architecture design in current deep learning frameworks. Recently, Vision Transformers (ViTs) have achieved remarkable success in computer vision, leveraging the powerful global dependency modeling capability of the self-attention mechanism. Furthermore, Mamba2 has demonstrated its significant potential in natural language processing tasks by explicitly modeling the spatial adjacency prior through the structured mask. In this paper, we propose Polyline Path Masked Attention (PPMA) that integrates the self-attention mechanism of ViTs with an enhanced structured mask of Mamba2, harnessing the complementary strengths of both architectures. Specifically, we first ameliorate the traditional structured mask of Mamba2 by introducing a 2D polyline path scanning strategy and derive its corresponding structured mask, polyline path mask, which better preserves the adjacency relationships among image tokens. Notably, we conduct a thorough theoretical analysis on the structural characteristics of the proposed polyline path mask and design an efficient algorithm for the computation of the polyline path mask. Next, we embed the polyline path mask into the self-attention mechanism of ViTs, enabling explicit modeling of spatial adjacency prior. Extensive experiments on standard benchmarks, including image classification, object detection, and segmentation, demonstrate that our model outperforms previous state-of-the-art approaches based on both state-space models and Transformers. For example, our proposed PPMA-T/S/B models achieve 48.7%/51.1%/52.3% mIoU on the ADE20K semantic segmentation task, surpassing RMT-T/S/B by 0.7%/1.3%/0.3%, respectively. Code is available at https://github.com/zhongchenzhao/PPMA. Zhongchen Zhao, Chaodong Xiao, Qi Xie 0002, Lei Zhang 0006, Deyu Meng |
NeurIPS | 5 |
| 2025 | Practical Keyword Private Information Retrieval from Key-to-Index Mappings
Meng Hao 0001, Liqiang Peng, Pengfei Wu 0003, Lei Zhang 0006, Hongwei Li 0001, Robert H. Deng |
USENIX Security Symposium | 6 |
| 2025 | Deep attribute graph clustering based on bisymmetric network information fusion and mutual influence
Shuqiu Tan, Lei Zhang 0006 |
Appl. Intell. | 2 |
| 2025 | EvidenceMap: Learning evidence analysis to unleash the power of small language models for biomedical question answering
Chang Zong, Siliang Tang, Lei Zhang 0006 |
Artif. Intell. Medicine | 4 |
| 2025 | Personalized Image Generation with Deep Generative Models: A Decade SurveyabstractRecent advances in generative models have significantly facilitated the development of personalized content creation. Given a small set of images containing a user-specific concept, personalized image generation allows the user to create images that incorporate that concept while adhering to provided text descriptions. The technologies used for personalization have evolved alongside the development of generative models, with their distinct and interrelated components. In this survey, we present a comprehensive review of generalized personalized image generation across various generative models, including traditional GANs, contemporary text-to-image diffusion models, and emerging multi-modal autoregressive (AR) models. We first define a unified framework that standardizes the personalization process across different generative models, encompassing three key components: inversion spaces, inversion methods, and personalization schemes. This unified framework offers a structured approach to dissecting and comparing personalization techniques across different generative architectures. Building upon our framework, we provide an in-depth analysis of personalization techniques within each generative model, highlighting their unique contributions and innovations. Through comparative analysis, we elucidate the current landscape of personalized image generation, identifying commonalities and distinguishing features of existing methods. Finally, we discuss open challenges in the field and propose potential directions for future research. We keep a bibliography of related works at https://github.com/csyxwei/Awesome-Personalized-Image-Generation. Yuxiang Wei 0001, Yiheng Zheng, Yabo Zhang, Ming Liu 0018, Zhilong Ji, Lei Zhang 0006, Wangmeng Zuo |
Comput. Vis. Media | 6 |
| 2025 | Adaptive network combination for single-image reflection removal: a domain generalization perspective
Ming Liu 0018, Jianan Pan, Zifei Yan, Wangmeng Zuo, Lei Zhang 0006 |
Frontiers Comput. Sci. | 5 |
| 2025 | TokenPacker: Efficient Visual Projector for Multimodal LLM
Wentong Li 0001, Yuqian Yuan, Jian Liu 0012, Dongqi Tang, Song Wang 0019, Jie Qin 0004, Jianke Zhu, Lei Zhang 0006 |
Int. J. Comput. Vis. | 8 |
| 2025 | EMBANet: A flexible efficient multi-branch attention networkabstractRecent advances in the design of convolutional neural networks have shown that performance can be enhanced by improving the ability to represent multi-scale features. However, most existing methods either focus on designing more sophisticated attention modules, which leads to higher computational costs, or fail to effectively establish long-range channel dependencies, or neglect the extraction and utilization of structural information. This work introduces a novel module, the Multi-Branch Concatenation (MBC), designed to process input tensors and extract multi-scale feature maps. The MBC module introduces new degrees of freedom (DoF) in the design of attention networks by allowing for flexible adjustments to the types of transformation operators and the number of branches. This study considers two key transformation operators: multiplexing and splitting, both of which facilitate a more granular representation of multi-scale features and enhance the receptive field range. By integrating the MBC with an attention module, a Multi-Branch Attention (MBA) module is developed to capture channel-wise interactions within feature maps, thereby establishing long-range channel dependencies. Replacing the 3x3 convolutions in the bottleneck blocks of ResNet with the proposed MBA yields a new block, the Efficient Multi-Branch Attention (EMBA), which can be seamlessly integrated into state-of-the-art backbone CNN models. Furthermore, a new backbone network, named EMBANet, is constructed by stacking EMBA blocks. The proposed EMBANet has been thoroughly evaluated across various computer vision tasks, including classification, detection, and segmentation, consistently demonstrating superior performance compared to popular backbones. Keke Zu, Lei Zhang 0006, Jian Lu 0002, Chen Xu 0004, Hongyang Chen 0001, Yu Zheng 0004 |
Neural Networks | 3 |
| 2025 | Reliable and Private Utility Signaling for Data MarketsabstractThe explosive growth of data has highlighted its critical role in driving economic growth through data marketplaces, which enable extensive data sharing and access to high-quality datasets. To support effective trading, signaling mechanisms provide participants with information about data products before transactions, enabling informed decisions and facilitating trading. However, due to the inherent free-duplication nature of data, commonly practiced signaling methods face a dilemma between privacy and reliability, undermining the effectiveness of signals in guiding decision-making. To address this, this paper explores the benefits and develops a non-TCP-based construction for a desirable signaling mechanism that simultaneously ensures privacy and reliability. We begin by formally defining the desirable utility signaling mechanism and proving its ability to prevent suboptimal decisions for both participants and facilitate informed data trading. To design a protocol to realize its functionality, we propose leveraging maliciously secure multi-party computation (MPC) to ensure the privacy and robustness of signal computation and introduce an MPC-based hash verification scheme to ensure input reliability. In multi-seller scenarios requiring fair data valuation, we further explore the design and optimization of the MPC-based KNN-Shapley method with improved efficiency. Rigorous experiments demonstrate the efficiency and practicality of our approach. Jiayao Zhang 0006, Yihang Wu, Jinfei Liu, Zheng Yan 0002, Kui Ren 0001, Lei Zhang 0006, Lin Qu |
Proc. ACM Manag. Data | 8 |
| 2025 | Charge Your Clients: Payable Secure Computation and Its ApplicationsabstractThe online realm has witnessed a surge in the buying and selling of data, prompting the emergence of dedicated data marketplaces. These platforms cater to servers (sellers), enabling them to set prices for access to their data, and clients (buyers), who can subsequently purchase these data, thereby streamlining and facilitating such transactions. However, the current data market is primarily confronted with the following issues. Firstly, they fail to protect client privacy, presupposing that clients submit their queries in plaintext. Secondly, these models are susceptible to being impacted by malicious client behavior, for example, enabling clients to potentially engage in arbitrage activities. To address the aforementioned issues, we propose payable secure computation, a novel secure computation paradigm specifically designed for data pricing scenarios. It grants the server the ability to securely procure essential pricing information while protecting the privacy of client queries. Additionally, it fortifies the server’s privacy against potential malicious client activities. As specific applications, we have devised customized payable protocols for two distinct secure computation scenarios: Keyword Private Information Retrieval (KPIR) and Private Set Intersection (PSI). We implement our two payable protocols and compare them with the state-of-the-art related protocols that do not support pricing as a baseline. Since our payable protocols are more powerful in the data pricing setting, the experiment results show that they do not introduce much overhead over the baseline protocols. Our payable KPIR achieves the same online cost as baseline, while the setup is about 1.3−1.6× slower than it. Our payable PSI needs about 2× more communication cost than that of baseline protocol, while the runtime is 1.5−3.2× slower than it depending on the network setting. Liqiang Peng, Meng Hao 0001, Lei Zhang 0006, Dongdai Lin |
IEEE Trans. Inf. Forensics Secur. | 6 |
| 2025 | Improving the Stability and Efficiency of Diffusion Models for Content Consistent Super-ResolutionabstractThe generative priors of pre-trained latent diffusion models (DMs) have demonstrated great potential to enhance the visual quality of image super-resolution (SR) results. However, the noise sampling process in DMs introduces randomness in the SR outputs, and the generated contents can differ a lot with different noise samples. The multi-step diffusion process can be accelerated by distilling methods, but the generative capacity is difficult to control. To address these issues, we analyze the respective advantages of DMs and generative adversarial networks (GANs) and propose to partition the generative SR process into two stages, where the DM is employed for reconstructing image structures and the GAN is employed for improving fine-grained details. Specifically, we propose a non-uniform timestep sampling strategy in the first stage. A single timestep sampling is first applied to extract the coarse information from the input image, then a few reverse steps are used to reconstruct the main structures. In the second stage, we finetune the decoder of the pre-trained variational auto-encoder by adversarial GAN training for deterministic detail enhancement. Once trained, our proposed method, namely content consistent super-resolution (CCSR), allows flexible use of different diffusion steps in the inference stage without re-training. Extensive experiments show that with 2 or even 1 diffusion step, CCSR can significantly improve the content consistency of SR outputs while keeping high perceptual quality. Codes and models can be found at https://github.com/csslc/CCSR. Lingchen Sun, Rongyuan Wu, Jie Liang 0007, Zhengqiang Zhang, Hongwei Yong, Lei Zhang 0006 |
IEEE Trans. Image Process. | 6 |
| 2025 | IVAC-$\mathbf {P^{2}L}$: Leveraging Irregular Repetition Priors for Improving Video Action CountingabstractThe quantification of repetitive actions in videos, a task commonly referred to as Video Action Counting (VAC), is a critical challenge in understanding and analyzing content in sports, fitness, and daily activities. Traditional approaches to VAC have largely overlooked the nuanced irregularities inherent in action repetitions, such as interruptions and variable lengths between cycles. Addressing this gap, our study introduces a novel perspective on VAC, focusing on Irregular Video Action Counting (IVAC), which emphasizes the importance of modeling the irregular repetition priors present in video content. We conceptualize these priors through two key aspects:Inter-cycle ConsistencyandCycle-interval Inconsistency. Inter-cycle Consistency ensures that thespatiotemporalrepresentations across all cycle segments in a videoremainhomogeneous, thereby reflecting the uniformity of actions betweendifferent cycle segments. In contrast, Cycle-interval Inconsistency mandates a clear semantic distinction between the representations of cycle segments and intervals, acknowledging the inherent dissimilarities in content. To effectively encapsulate these priors, we introduce a novel methodology consisting of consistency and inconsistency modules, underpinned by a tailored pull-push loss ($\mathrm {P^{2}~L}$) mechanism. This approach employs a pull loss to enhance the cohesion among cycle segment features and a push loss to distinctly differentiate between cycle and interval segment features. Empirical evaluations on the RepCount dataset illustrate that our IVAC-$\mathrm {P^{2}~L}$model sets a new benchmark in state-of-the-art performance for the VAC task. Moreover, our model demonstrates adaptability and generalization across diverse video content, achieving superior performance on two additional datasets, UCFRep and Countix, without necessitating dataset-specific fine-tuning. These findings not only validate the effectiveness of our approach in addressing the complexities of irregular repetitions in videos but also open new avenues for future research in video understanding and analysis. Zhi-Qi Cheng, Youtian Du, Lei Zhang 0006 |
IEEE Trans. Multim. | 4 |
| 2025 | Point-DAE: Denoising Autoencoders for Self-Supervised Point Cloud LearningabstractMasked autoencoder (MAE) has demonstrated its effectiveness in self-supervised point cloud learning. Considering that masking is a kind of corruption, in this work we explore a more general denoising autoencoder for point cloud learning (Point-DAE) by investigating more types of corruptions beyond masking. Specifically, we degrade the point cloud with certain corruptions as input, and learn an encoder-decoder model to reconstruct the original point cloud from its corrupted version. Three corruption families (i.e., density/masking, noise, and affine transformation) and a total of 14 corruption types are investigated with traditional non-Transformer encoders. Besides the popular masking corruption, we identify another effective corruption family, i.e., affine transformation. The affine transformation disturbs all points globally, which is complementary to the masking corruption where some local regions are dropped. We also validate the effectiveness of affine transformation corruption with the Transformer backbones, where we decompose the reconstruction of the complete point cloud into the reconstructions of detailed local patches and rough global shape, alleviating the position leakage problem in the reconstruction. Extensive experiments on tasks of object classification, few-shot learning, robustness testing, part segmentation, and 3-D object detection validate the effectiveness of the proposed method. The codes are available at https://github.com/YBZh/Point-DAE. Yabin Zhang 0001, Jiehong Lin, Ruihuang Li, Kui Jia, Lei Zhang 0006 |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2024 | Dual Memory Networks: A Versatile Adaptation Approach for Vision-Language ModelsabstractWith the emergence of pre-trained vision-language models like CLIP, how to adapt them to various downstream classification tasks has garnered significant attention in re-cent research. The adaptation strategies can be typically categorized into three paradigms: zero-shot adaptation, few-shot adaptation, and the recently-proposed training-free few-shot adaptation. Most existing approaches are tai-lored for a specific setting and can only cater to one or two of these paradigms. In this paper, we introduce a versa-tile adaptation approach that can effectively work under all three settings. Specifically, we propose the dual memory networks that comprise dynamic and static memory components. The static memory caches training data knowledge, enabling training-free few-shot adaptation, while the dynamic memory preserves historical test features online during the testing process, allowing for the exploration of additional data insights beyond the training set. This novel capability enhances model performance in the few-shot setting and enables model usability in the absence of training data. The two memory networks employ the same flexible memory interactive strategy, which can operate in a training-free mode and can be further enhanced by in-corporating learnable projection layers. Our approach is tested across 11 datasets under the three task settings. Re-markably, in the zero-shot scenario, it outperforms existing methods by over 3% and even shows superior results against methods utilizing external training data. Addition-ally, our method exhibits robust performance against nat-ural distribution shifts. Codes are available at https://github.com/YBZh/DMN. Yabin Zhang 0001, Wenjie Zhu 0003, Zhiyuan Ma 0002, Kaiyang Zhou, Lei Zhang 0006 |
CVPR | 6 |
| 2024 | Neural Super-Resolution for Real-Time Rendering with Radiance DemodulationabstractIt is time-consuming to render high-resolution images in applications such as video games and virtual reality, and thus super-resolution technologies become increasingly popular for real-time rendering. However, it is challenging to preserve sharp texture details, keep the temporal stability and avoid the ghosting artifacts in real-time super-resolution rendering. To address this issue, we introduce radiance demodulation to separate the rendered image or radiance into a lighting component and a material component, considering the fact that the light component is smoother than the rendered image so that the high-resolution material component with detailed textures can be easily obtained. We perform the super-resolution on the lighting component only and re-modulate it with the high-resolution material component to obtain the final super-resolution image with more texture details. A reliable warping module is proposed by explicitly marking the occluded regions to avoid the ghosting artifacts. To further enhance the temporal stability, we design a frame-recurrent neural network and a temporal loss to aggregate the previous and current frames, which can better capture the spatial-temporal consistency among reconstructed frames. As a result, our method is able to produce temporally stable results in real-time rendering with high-quality details, even in the challenging 4 × 4 super-resolution scenarios. Code is available at: https://github.com/Riga2/NSRD. Ziling Chen, Lu Wang 0007, Beibei Wang 0002, Lei Zhang 0006 |
CVPR | 6 |
| 2024 | UniVS: Unified and Universal Video Segmentation with Prompts as QueriesabstractDespite the recent advances in unified image segmentation (IS), developing a unified video segmentation (VS) model remains a challenge. This is mainly because generic category-specified VS tasks need to detect all objects and track them across consecutive frames, while prompt-guided VS tasks require re-identifying the target with visual/text prompts throughout the entire video, making it hard to handle the different tasks with the same architecture. We make an attempt to address these issues and present a novel unified VS architecture, namely UniVS, by using prompts as queries. UniVS averages the prompt features of the target from previous frames as its initial query to explicitly decode masks, and introduces a target-wise prompt crossattention layer in the mask decoder to integrate prompt features in the memory pool. By taking the predicted masks of entities from previous frames as their visual prompts, UniVS converts different VS tasks into prompt-guided target segmentation, eliminating the heuristic inter-frame matching process. Our framework not only unifies the different VS tasks but also naturally achieves universal training and testing, ensuring robust performance across different scenarios. UniVS shows a commendable balance between performance and universality on 10 challenging VS benchmarks, covering video instance, semantic, panoptic, object, and referring segmentation tasks. Code can be found at https://github.com/MinghanLi/UniVS. Minghan Li 0001, Shuai Li 0014, Lei Zhang 0006 |
CVPR | 4 |
| 2024 | SeeSR: Towards Semantics-Aware Real-World Image Super-ResolutionabstractOwe to the powerful generative priors, the pretrained text-to-image (T2I) diffusion models have become increasingly popular in solving the real-world image super-resolution problem. However, as a consequence of the heavy quality degradation of input low-resolution (LR) images, the destruction of local structures can lead to ambiguous image semantics. As a result, the content of reproduced high-resolution image may have semantic errors, deteriorating the super-resolution performance. To address this issue, we present a semantics-aware approach to better preserve the semantic fidelity of generative real-world image super-resolution. First, we train a degradation-aware prompt extractor, which can generate accurate soft and hard semantic prompts even under strong degradation. The hard semantic prompts refer to the image tags, aiming to enhance the local perception ability of the T2I model, while the soft semantic prompts compensate for the hard ones to provide additional representation information. These semantic prompts encourage the T2I model to generate detailed and semantically accurate results. Further-more, during the inference process, we integrate the LR images into the initial sampling noise to mitigate the diffusion model's tendency to generate excessive random details. The experiments show that our method can reproduce more realistic image details and hold better the semantics. The source code of our method can be found at https://github.com/cswry/SeeSR. Rongyuan Wu, Tao Yang 0042, Lingchen Sun, Zhengqiang Zhang, Shuai Li 0014, Lei Zhang 0006 |
CVPR | 6 |
| 2024 | Osprey: Pixel Understanding with Visual Instruction TuningabstractMultimodal large language models (MLLMs) have recently achieved impressive general-purpose vision-language capabilities through visual instruction tuning. However, current MLLMs primarily focus on image-level or box-level understanding, falling short in achieving fine-grained vision-language alignment at pixel level. Besides, the lack of mask-based instruction data limits their ad-vancements. In this paper, we propose Osprey, a mask-text instruction tuning approach, to extend MLLMs by incor-porating fine-grained mask regions into language instruction, aiming at achieving pixel-wise visual understanding. To achieve this goal, we first meticulously curate a mask-based region-text dataset with 724K samples, and then design a vision-language model by injecting pixel-level representation into LLM. Specifically, Osprey adopts a convolutional CLIP backbone as the vision encoder and employs a mask-aware visual extractor to extract precise visual mask features from high resolution input. Experimen-tal results demonstrate Osprey's superiority in various region understanding tasks, showcasing its new capability for pixel-level instruction tuning. In particular, Osprey can be integrated with Segment Anything Model (SAM) seamlessly to obtain multi-granularity semantics. The source code, dataset and demo can be found at https://github.com/CircleRadon/Osprey. Yuqian Yuan, Wentong Li 0001, Jian Liu 0012, Dongqi Tang, Xinjie Luo, Chi Qin, Lei Zhang 0006, Jianke Zhu |
CVPR | 7 |
| 2024 | ScatterFormer: Efficient Voxel Transformer with Scattered Linear Attention
Chenhang He, Ruihuang Li, Guowen Zhang, Lei Zhang 0006 |
ECCV (29) | 4 |
| 2024 | Source Prompt Disentangled Inversion for Boosting Image Editability with Diffusion Models
Ruibin Li, Ruihuang Li, Song Guo 0001, Lei Zhang 0006 |
ECCV (26) | 4 |
| 2024 | Dense Multimodal Alignment for Open-Vocabulary 3D Scene Understanding
Ruihuang Li, Zhengqiang Zhang, Chenhang He, Zhiyuan Ma 0002, Vishal M. Patel, Lei Zhang 0006 |
ECCV (49) | 6 |
| 2024 | ScaleDreamer: Scalable Text-to-3D Synthesis with Asynchronous Score Distillation
Zhiyuan Ma 0002, Yuxiang Wei 0001, Yabin Zhang 0001, Xiangyu Zhu 0001, Zhen Lei 0001, Lei Zhang 0006 |
ECCV (7) | 6 |
| 2024 | Responsible Visual Editing
Minheng Ni, Yeli Shen, Lei Zhang 0006, Wangmeng Zuo |
ECCV (22) | 3 |
| 2024 | Open Vocabulary 3D Scene Understanding via Geometry Guided Self-Distillation
Pengfei Wang 0012, Yuxi Wang 0001, Shuai Li 0014, Zhaoxiang Zhang 0001, Zhen Lei 0001, Lei Zhang 0006 |
ECCV (15) | 6 |
| 2024 | MasterWeaver: Taming Editability and Face Identity for Personalized Text-to-Image Generation
Yuxiang Wei 0001, Zhilong Ji, Jinfeng Bai, Lei Zhang 0006, Wangmeng Zuo |
ECCV (51) | 5 |
| 2024 | A Comprehensive Study of Multimodal Large Language Models for Image Quality Assessment
Tianhe Wu, Kede Ma, Jie Liang 0007, Yujiu Yang 0001, Lei Zhang 0006 |
ECCV (74) | 5 |
| 2024 | Self-Supervised Video Desmoking for Laparoscopic Surgery
Renlong Wu, Zhilu Zhang 0001, Shuohao Zhang, Longfei Gou, Lei Zhang 0006, Hao Chen 0003, Wangmeng Zuo |
ECCV (72) | 6 |
| 2024 | Pixel-Aware Stable Diffusion for Realistic Image Super-Resolution and Personalized Stylization
Tao Yang 0042, Rongyuan Wu, Peiran Ren, Xuansong Xie, Lei Zhang 0006 |
ECCV (11) | 5 |
| 2024 | General Geometry-Aware Weakly Supervised 3D Object Detection
Guowen Zhang, Junsong Fan, Liyi Chen 0002, Zhaoxiang Zhang 0001, Zhen Lei 0001, Lei Zhang 0006 |
ECCV (51) | 6 |
| 2024 | LAPT: Label-Driven Automated Prompt Tuning for OOD Detection with Vision-Language Models
Yabin Zhang 0001, Wenjie Zhu 0003, Chenhang He, Lei Zhang 0006 |
ECCV (72) | 4 |
| 2024 | PA-SAM: Prompt Adapter SAM for High-Quality Image SegmentationabstractThe Segment Anything Model (SAM) has exhibited outstanding performance in various image segmentation tasks. Despite being trained with over a billion masks, SAM faces challenges in mask prediction quality in numerous scenarios, especially in real-world contexts. In this paper, we introduce a novel prompt-driven adapter into SAM, namely Prompt Adapter Segment Anything Model (PA-SAM), aiming to enhance the segmentation mask quality of the original SAM. By exclusively training the prompt adapter, PA-SAM extracts detailed information from images and optimizes the mask decoder feature at both sparse and dense prompt levels, improving the segmentation performance of SAM to produce high-quality masks. Experimental results demonstrate that our PA-SAM outperforms other SAM-based methods in high-quality, zero-shot, and open-set segmentation. We’re making the source code and models available at https://github.com/xzz2/pa-sam. Zhaozhi Xie, Bochen Guan, Muyang Yi, Yue Ding 0001, Hongtao Lu 0001, Lei Zhang 0006 |
ICME | 7 |
| 2024 | LIDIA: Precise Liver Tumor Diagnosis on Multi-Phase Contrast-Enhanced CT via Iterative Fusion and Asymmetric Contrastive Learning
Wei Liu 0127, Xiaoming Zhang 0008, Xiaoli Yin, Xu Han 0023, Chunli Li, Yuan Gao 0017, Le Lu 0001, Ling Zhang 0002, Lei Zhang 0006, Ke Yan 0006 |
MICCAI (9) | 11 |
| 2024 | SSL: A Self-similarity Loss for Improving Generative Image Super-resolutionabstractGenerative adversarial networks (GAN) and generative diffusion models (DM) have been widely used in real-world image super-resolution (Real-ISR) to enhance the image perceptual quality. However, these generative models are prone to generating visual artifacts and false image structures, resulting in unnatural Real-ISR results. Based on the fact that natural images exhibit high self-similarities, i.e., a local patch can have many similar patches to it in the whole image, in this work we propose a simple yet effective self-similarity loss (SSL) to improve the performance of generative Real-ISR models, enhancing the hallucination of structural and textural details while reducing the unpleasant visual artifacts. Specifically, we compute a self-similarity graph (SSG) of the ground-truth image, and enforce the SSG of Real-ISR output to be close to it. To reduce the training cost and focus on edge areas, we generate an edge mask from the ground-truth image, and compute the SSG only on the masked pixels. The proposed SSL serves as a general plug-and-play penalty, which could be easily applied to the off-the-shelf Real-ISR models. Our experiments demonstrate that, by coupling with SSL, the performance of many state-of-the-art Real-ISR models, including those GAN and DM based ones, can be largely improved, reproducing more perceptually realistic image details and eliminating many false reconstructions and visual artifacts. Codes and supplementary material are available at https://github.com/ChrisDud0257/SSL Zhengqiang Zhang, Jie Liang 0007, Lei Zhang 0006 |
ACM Multimedia | 4 |
| 2024 | TAPTRv2: Attention-based Position Update Improves Tracking Any PointabstractIn this paper, we present TAPTRv2, a Transformer-based approach built upon TAPTR for solving the Tracking Any Point (TAP) task. TAPTR borrows designs from DEtection TRansformer (DETR) and formulates each tracking point as a point query, making it possible to leverage well-studied operations in DETR-like algorithms. TAPTRv2 improves TAPTR by addressing a critical issue regarding its reliance on cost-volume, which contaminates the point query’s content feature and negatively impacts both visibility prediction and cost-volume computation. In TAPTRv2, we propose a novel attention-based position update (APU) operation and use key-aware deformable attention to realize. For each query, this operation uses key-aware attention weights to combine their corresponding deformable sampling positions to predict a new query position. This design is based on the observation that local attention is essentially the same as cost-volume, both of which are computed by dot-production between a query and its surrounding features. By introducing this new operation, TAPTRv2 not only removes the extra burden of cost-volume computation, but also leads to a substantial performance improvement. TAPTRv2 surpasses TAPTR and achieves state-of-the-art performance on many challenging datasets, demonstrating the effectiveness of our approach. Hongyang Li 0003, Hao Zhang 0097, Shilong Liu 0004, Zhaoyang Zeng, Feng Li 0040, Tianhe Ren, Lei Zhang 0006 |
NeurIPS | 8 |
| 2024 | One-Step Effective Diffusion Network for Real-World Image Super-ResolutionabstractThe pre-trained text-to-image diffusion models have been increasingly employed to tackle the real-world image super-resolution (Real-ISR) problem due to their powerful generative image priors. Most of the existing methods start from random noise to reconstruct the high-quality (HQ) image under the guidance of the given low-quality (LQ) image. While promising results have been achieved, such Real-ISR methods require multiple diffusion steps to reproduce the HQ image, increasing the computational cost. Meanwhile, the random noise introduces uncertainty in the output, which is unfriendly to image restoration tasks. To address these issues, we propose a one-step effective diffusion network, namely OSEDiff, for the Real-ISR problem.
We argue that the LQ image contains rich information to restore its HQ counterpart, and hence the given LQ image can be directly taken as the starting point for diffusion, eliminating the uncertainty introduced by random noise sampling. We finetune the pre-trained diffusion network with trainable layers to adapt it to complex image degradations. To ensure that the one-step diffusion model could yield HQ Real-ISR output, we apply variational score distillation in the latent space to conduct KL-divergence regularization. As a result, our OSEDiff model can efficiently and effectively generate HQ images in just one diffusion step.
Our experiments demonstrate that OSEDiff achieves comparable or even better Real-ISR results, in terms of both objective metrics and subjective evaluations, than previous diffusion model-based Real-ISR methods that require dozens or hundreds of steps. The source codes are released at https://github.com/cswry/OSEDiff. Rongyuan Wu, Lingchen Sun, Zhiyuan Ma 0002, Lei Zhang 0006 |
NeurIPS | 4 |
| 2024 | Voxel Mamba: Group-Free State Space Models for Point Cloud based 3D Object DetectionabstractSerialization-based methods, which serialize the 3D voxels and group them into multiple sequences before inputting to Transformers, have demonstrated their effectiveness in 3D object detection. However, serializing 3D voxels into 1D sequences will inevitably sacrifice the voxel spatial proximity. Such an issue is hard to be addressed by enlarging the group size with existing serialization-based methods due to the quadratic complexity of Transformers with feature sizes. Inspired by the recent advances of state space models (SSMs), we present a Voxel SSM, termed as Voxel Mamba, which employs a group-free strategy to serialize the whole space of voxels into a single sequence. The linear complexity of SSMs encourages our group-free design, alleviating the loss of spatial proximity of voxels. To further enhance the spatial proximity, we propose a Dual-scale SSM Block to establish a hierarchical structure, enabling a larger receptive field in the 1D serialization curve, as well as more complete local regions in 3D space. Moreover, we implicitly apply window partition under the group-free framework by positional encoding, which further enhances spatial proximity by encoding voxel positional information. Our experiments on Waymo Open Dataset and nuScenes dataset show that Voxel Mamba not only achieves higher accuracy than state-of-the-art methods, but also demonstrates significant advantages in computational efficiency. The source code is available at https://github.com/gwenzhang/Voxel-Mamba. Guowen Zhang, Lue Fan, Chenhang He, Zhen Lei 0001, Zhaoxiang Zhang 0001, Lei Zhang 0006 |
NeurIPS | 6 |
| 2024 | Efficient participating media rendering with differentiable regularizationabstractHighly scattering media, such as milk, skin, and clouds, are common in the real world. Rendering participating media is challenging, especially for high-order scattering dominant media, because the light may undergo a large number of scattering events before leaving the surface. Monte Carlo-based methods typically require a long time to produce noise-free results. Based on the observation that low-albedo media contain less noise than high-albedo media, we propose reducing the variance of the rendered results using differentiable regularization. We first render an image with low-albedo participating media together with the gradient with respect to the albedo, and then predict the final rendered image with a low-albedo image and gradient image via a novel prediction function. To achieve high quality, we also consider the gradients of neighboring frames to provide a noise-free gradient image. Ultimately, our method can produce results with much less overall error than equal-time path tracing methods. Wenshi Wu, Beibei Wang 0002, Milos Hasan, Lei Zhang 0006, Zhong Jin, Lingqi Yan 0001 |
Comput. Vis. Media | 4 |
| 2024 | Towards Diverse Binary Segmentation via a Simple yet General Gated Network
Xiaoqi Zhao 0003, Youwei Pang, Lihe Zhang, Huchuan Lu, Lei Zhang 0006 |
Int. J. Comput. Vis. | 5 |
| 2024 | Local Differentially Private Heavy Hitter Detection in Data Streams with Bounded MemoryabstractTop-k frequent items detection is a fundamental task in data stream mining. Many promising solutions are proposed to improve memory efficiency while still maintaining high accuracy for detecting the Top-k items. Despite the memory efficiency concern, the users could suffer from privacy loss if participating in the task without proper protection, since their contributed local data streams may continually leak sensitive individual information. However, most existing works solely focus on addressing either the memory-efficiency problem or the privacy concerns but seldom jointly, which cannot achieve a satisfactory tradeoff between memory efficiency, privacy protection, and detection accuracy. In this paper, we present a novel framework HG-LDP to achieve accurate Top-k item detection at bounded memory expense, while providing rigorous local differential privacy (LDP) protection. Specifically, we identify two key challenges naturally arising in the task, which reveal that directly applying existing LDP techniques will lead to an inferior "accuracy-privacy-memory efficiency" tradeoff. Therefore, we instantiate three advanced schemes under the framework by designing novel LDP randomization methods, which address the hurdles caused by the large size of the item domain and by the limited space of the memory. We conduct comprehensive experiments on both synthetic and real-world datasets to show that the proposed advanced schemes achieve a superior "accuracy-privacy-memory efficiency" tradeoff, saving 2300× memory over baseline methods when the item domain size is 41,270. Our code is anonymously open-sourced via the link. Jian Lou 0001, Yuan Hong 0001, Lei Zhang 0006, Zhan Qin, Kui Ren 0001 |
Proc. ACM Manag. Data | 5 |
| 2024 | Box2Mask: Box-Supervised Instance Segmentation via Level-Set EvolutionabstractIn contrast to fully supervised methods using pixel-wise mask labels, box-supervised instance segmentation takes advantage of simple box annotations, which has recently attracted increasing research attention. This paper presents a novel single-shot instance segmentation approach, namely Box2Mask, which integrates the classical level-set evolution model into deep neural network learning to achieve accurate mask prediction with only bounding box supervision. Specifically, both the input image and its deep features are employed to evolve the level-set curves implicitly, and a local consistency module based on a pixel affinity kernel is used to mine the local context and spatial relations. Two types of single-stage frameworks, i.e., CNN-based and transformer-based frameworks, are developed to empower the level-set evolution for box-supervised instance segmentation, and each framework consists of three essential components: instance-aware decoder, box-level matching assignment and level-set evolution. By minimizing the level-set energy function, the mask map of each instance can be iteratively optimized within its bounding box annotation. The experimental results on five challenging testbeds, covering general scenes, remote sensing, medical and scene text images, demonstrate the outstanding performance of our proposed Box2Mask approach for box-supervised instance segmentation. In particular, with the Swin-Transformer large backbone, our Box2Mask obtains 42.4% mask AP on COCO, which is on par with the recently developed fully mask-supervised methods. Wentong Li 0001, Wenyu Liu 0005, Jianke Zhu, Miaomiao Cui, Risheng Yu, Xian-Sheng Hua 0001, Lei Zhang 0006 |
IEEE Trans. Pattern Anal. Mach. Intell. | 7 |
| 2024 | Deep Variational Network Toward Blind Image RestorationabstractBlind image restoration (IR) is a common yet challenging problem in computer vision. Classical model-based methods and recent deep learning (DL)-based methods represent two different methodologies for this problem, each with their own merits and drawbacks. In this paper, we propose a novel blind image restoration method, aiming to integrate both the advantages of them. Specifically, we construct a general Bayesian generative model for the blind IR, which explicitly depicts the degradation process. In this proposed model, a pixel-wise non-i.i.d. Gaussian distribution is employed to fit the image noise. It is with more flexibility than the simple i.i.d. Gaussian or Laplacian distributions as adopted in most of conventional methods, so as to handle more complicated noise types contained in the image degradation. To solve the model, we design a variational inference algorithm where all the expected posteriori distributions are parameterized as deep neural networks to increase their model capability. Notably, such an inference algorithm induces a unified framework to jointly deal with the tasks of degradation estimation and image restoration. Further, the degradation information estimated in the former task is utilized to guide the latter IR process. Experiments on two typical blind IR tasks, namely image denoising and super-resolution, demonstrate that the proposed method achieves superior performance over current state-of-the-arts. Zongsheng Yue, Hongwei Yong, Qian Zhao 0002, Lei Zhang 0006, Deyu Meng, Kwan-Yee Kenneth Wong |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2024 | Perception-Distortion Balanced Super-Resolution: A Multi-Objective Optimization PerspectiveabstractHigh perceptual quality and low distortion degree are two important goals in image restoration tasks such as super-resolution (SR). Most of the existing SR methods aim to achieve these goals by minimizing the corresponding yet conflicting losses, such as the$\ell _{1}$loss and the adversarial loss. Unfortunately, the commonly used gradient-based optimizers, such as Adam, are hard to balance these objectives due to the opposite gradient decent directions of the contradictory losses. In this paper, we formulate the perception-distortion trade-off in SR as a multi-objective optimization problem and develop a new optimizer by integrating the gradient-free evolutionary algorithm (EA) with gradient-based Adam, where EA and Adam focus on the divergence and convergence of the optimization directions respectively. As a result, a population of optimal models with different perception-distortion preferences is obtained. We then design a fusion network to merge these models into a single stronger one for an effective perception-distortion trade-off. Experiments demonstrate that with the same backbone network, the perception-distortion balanced SR model trained by our method can achieve better perceptual quality than its competitors while attaining better reconstruction fidelity. Codes and models can be found athttps://github.com/csslc/EA-Adam. Lingchen Sun, Jie Liang 0007, Shuaizheng Liu, Hongwei Yong, Lei Zhang 0006 |
IEEE Trans. Image Process. | 5 |
| 2024 | TMP: Temporal Motion Propagation for Online Video Super-ResolutionabstractOnline video super-resolution (online-VSR) highly relies on an effective alignment module to aggregate temporal information, while the strict latency requirement makes accurate and efficient alignment very challenging. Though much progress has been achieved, most of the existing online-VSR methods estimate the motion fields of each frame separately to perform alignment, which is computationally redundant and ignores the fact that the motion fields of adjacent frames are correlated. In this work, we propose an efficient Temporal Motion Propagation (TMP) method, which leverages the continuity of motion field to achieve fast pixel-level alignment among consecutive frames. Specifically, we first propagate the offsets from previous frames to the current frame, and then refine them in the neighborhood, significantly reducing the matching space and speeding up the offset estimation process. Furthermore, to enhance the robustness of alignment, we perform spatial-wise weighting on the warped features, where the positions with more precise offsets are assigned higher importance. Experiments on benchmark datasets demonstrate that the proposed TMP method achieves leading online-VSR accuracy as well as inference speed. The source code of TMP can be found at https://github.com/xtudbxk/TMP. Zhengqiang Zhang, Ruihuang Li, Shi Guo, Yang Cao 0017, Lei Zhang 0006 |
IEEE Trans. Image Process. | 5 |
| 2024 | Domain Adaptation Transformer for Unsupervised Driving-Scene Segmentation in Adverse ConditionsabstractSemantic segmentation in driving scenarios is important for modern autonomous driving technology. While the existing methods have shown promising results in segmenting normal-condition images, their performance in adverse scenes remains unsatisfactory due to limited visual field and lack of annotation. To address this issue, we propose an unsupervised domain adaptation semantic segmentation method with the transformer architecture, namely ACSegFormer, for driving-scene adverse conditions, aiming at mining image features in visually restricted scenes. Three effective training strategies are proposed in ACSegFormer to learn the latent image context relations and to reduce the gaps between different domains: an entropy-based pseudo label correction scheme that refines the target domain predictions with the normal reference predictions, an optimal transport-based inter-domain alignment module that performs domain alignment on the outputs of transformer encoder, and a masked context learning module that enhances the model’s ability to perceive the missing information of target domain image. Our ACSegFormer has no additional training parameters on top of the existing transformer segmentation framework, which can be easily used for self-training-based unsupervised domain adaptation approaches. The experimental results show that our ACSegFormer achieves state-of-the-art performance on driving-scene segmentation benchmarks in adverse conditions, including Dark Zurich and ACDC. Codes and models are available athttps://github.com/wenyyu/ACSegFormer. Wenyu Liu 0005, Song Wang 0019, Jianke Zhu, Xuansong Xie, Lei Zhang 0006 |
IEEE Trans. Intell. Transp. Syst. | 5 |
| 2024 | Confusion-Based Metric Learning for Regularizing Zero-Shot Image Retrieval and ClusteringabstractDeep metric learning turns to be attractive in zero-shot image retrieval and clustering (ZSRC) task in which a good embedding/metric is requested such that the unseen classes can be distinguished well. Most existing works deem this "good" embedding just to be the discriminative one and race to devise the powerful metric objectives or the hard-sample mining strategies for learning discriminative deep metrics. However, in this article, we first emphasize that the generalization ability is also a core ingredient of this "good" metric and it largely affects the metric performance in zero-shot settings as a matter of fact. Then, we propose the confusion-based metric learning (CML) framework to explicitly optimize a robust metric. It is mainly achieved by introducing two interesting regularization terms, i.e., the energy confusion (EC) and diversity confusion (DC) terms. These terms daringly break away from the traditional deep metric learning idea of designing discriminative objectives and instead seek to "confuse" the learned model. These two confusion terms focus on local and global feature distribution confusions, respectively. We train these confusion terms together with the conventional deep metric objective in an adversarial manner. Although it seems weird to "confuse" the model learning, we show that our CML indeed serves as an efficient regularization framework for deep metric learning and it is applicable to various conventional metric methods. This article empirically and experimentally demonstrates the importance of learning an embedding/metric with good generalization, achieving the state-of-the-art performances on the popular CUB, CARS, Stanford Online Products, and In-Shop datasets for ZSRC tasks. Binghui Chen, Weihong Deng, Lei Zhang 0006 |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2024 | TensoSDF: Roughness-aware Tensorial Representation for Robust Geometry and Material ReconstructionabstractReconstructing objects with realistic materials from multi-view images is problematic, since it is highly ill-posed. Although the neural reconstruction approaches have exhibited impressive reconstruction ability, they are designed for objects with specific materials (e.g., diffuse or specular materials). To this end, we propose a novel framework for robust geometry and material reconstruction, where the geometry is expressed with the implicit signed distance field (SDF) encoded by a tensorial representation, namely TensoSDF. At the core of our method is the roughness-aware incorporation of the radiance and reflectance fields, which enables a robust reconstruction of objects with arbitrary reflective materials. Furthermore, the tensorial representation enhances geometry details in the reconstructed surface and reduces the training time. Finally, we estimate the materials using an explicit mesh for efficient intersection computation and an implicit SDF for accurate representation. Consequently, our method can achieve more robust geometry reconstruction, outperform the previous works in terms of relighting quality, and reduce 50% training times and 70% inference time. Codes and datasets are available at https://github.com/Riga2/TensoSDF. Lu Wang 0007, Lei Zhang 0006, Beibei Wang 0002 |
ACM Trans. Graph. | 3 |
| 2023 | Inferring and Leveraging Parts from Object Shape for Improving Semantic Image SynthesisabstractDespite the progress in semantic image synthesis, it remains a challenging problem to generate photo-realistic parts from input semantic map. Integrating part segmentation map can undoubtedly benefit image synthesis, but is bothersome and inconvenient to be provided by users. To improve part synthesis, this paper presents to infer Parts from Object ShapE (iPOSE) and leverage it for improving semantic image synthesis. However, albeit several part segmentation datasets are available, part annotations are still not provided for many object categories in semantic image synthesis. To circumvent it, we resort to few-shot regime to learn a PartNet for predicting the object part map with the guidance of pre-defined support part maps. PartNet can be readily generalized to handle a new object category when a small number (e.g., 3) of support part maps for this category are provided. Furthermore, part semantic modulation is presented to incorporate both inferred part map and semantic map for image synthesis. Experiments show that our iPOSE not only generates objects with rich part details, but also enables to control the image synthesis flexibly. And our iPOSE performs favorably against the state-of-the-art methods in terms of quantitative and qualitative evaluation. Our code will be publicly available at https://github.com/csyxwei/iPOSE. Yuxiang Wei 0001, Zhilong Ji, Xiaohe Wu, Jinfeng Bai, Lei Zhang 0006, Wangmeng Zuo |
CVPR | 5 |
| 2023 | Human Guided Ground-Truth Generation for Realistic Image Super-ResolutionabstractHow to generate the ground-truth (GT) image is a critical issue for training realistic image super-resolution (Real-ISR) models. Existing methods mostly take a set of high-resolution (HR) images as GTs and apply various degradations to simulate their low-resolution (LR) counterparts. Though great progress has been achieved, such an LR-HR pair generation scheme has several limitations. First, the perceptual quality of HR images may not be high enough, limiting the quality of Real-ISR outputs. Second, existing schemes do not consider much human perception in GT generation, and the trained models tend to produce over-smoothed results or unpleasant artifacts. With the above considerations, we propose a human guided GT generation scheme. We first elaborately train multiple image enhancement models to improve the perceptual quality of HR images, and enable one LR image having multiple HR counterparts. Human subjects are then involved to annotate the high quality regions among the enhanced HR images as GTs, and label the regions with unpleasant artifacts as negative samples. A human guided GT image dataset with both positive and negative samples is then constructed, and a loss function is proposed to train the Real-ISR models. Experiments show that the Real-ISR models trained on our dataset can produce perceptually more realistic results with less artifacts. Dataset and codes can be found at https://github.com/ChrisDud0257/HGGT Jie Liang 0007, Ming Liu 0018, Hui Zeng 0001, Lei Zhang 0006 |
CVPR | 6 |
| 2023 | MSF: Motion-guided Sequential Fusion for Efficient 3D Object Detection from Point Cloud SequencesabstractPoint cloud sequences are commonly used to accurately detect 3D objects in applications such as autonomous driving. Current top-performing multi-frame detectors mostly follow a Detect-and-Fuse framework, which extracts features from each frame of the sequence and fuses them to detect the objects in the current frame. However, this inevitably leads to redundant computation since adjacent frames are highly correlated. In this paper, we propose an efficient Motion-guided Sequential Fusion (MSF) method, which exploits the continuity of object motion to mine useful sequential contexts for object detection in the current frame. We first generate 3D proposals on the current frame and propagate them to preceding frames based on the estimated velocities. The points-of-interest are then pooled from the sequence and encoded as proposal features. A novel Bidi-rectional Feature Aggregation (BiFA) module is further proposed to facilitate the interactions of proposal features across frames. Besides, we optimize the point cloud pooling by a voxel-based sampling technique so that millions of points can be processed in several milliseconds. The proposed MSF method achieves not only better efficiency than other multi-frame detectors but also leading accuracy, with 83.12% and 78.30% mAP on the LEVEL1 and LEVEL2 test sets of Waymo Open Dataset, respectively. Codes can be found at https://github.com/skyhehe123/MSF. Chenhang He, Ruihuang Li, Yabin Zhang 0001, Shuai Li 0014, Lei Zhang 0006 |
CVPR | 5 |
| 2023 | MDQE: Mining Discriminative Query Embeddings to Segment Occluded Instances on Challenging VideosabstractWhile impressive progress has been achieved, video instance segmentation (VIS) methods with per-clip input often fail on challenging videos with occluded objects and crowded scenes. This is mainly because instance queries in these methods cannot encode well the discriminative embeddings of instances, making the query-based segmenter difficult to distinguish those ‘hard’ instances. To address these issues, we propose to mine discriminative query embeddings (MDQE) to segment occluded instances on challenging videos. First, we initialize the positional embeddings and content features of object queries by considering their spatial contextual information and the inter-frame object motion. Second, we propose an inter-instance mask repulsion loss to distance each instance from its nearby non-target instances. The proposed MDQE is the first VIS method with per-clip input that achieves state-of-the-art results on challenging videos and competitive performance on simple videos. In specific, MDQE with ResNet50 achieves 33.0% and 44.5% mask AP on OVIS and YouTube- Vis 2021, respectively. Code of MDQE can be found at https://github.com/MinghanLi/MDQE_CVPR2023. Minghan Li 0001, Shuai Li 0014, Wangmeng Xiang, Lei Zhang 0006 |
CVPR | 4 |
| 2023 | DynaMask: Dynamic Mask Selection for Instance SegmentationabstractThe representative instance segmentation methods mostly segment different object instances with a mask of the fixed resolution, e.g., 28 × 28 grid. However, a low-resolution mask loses rich details, while a high-resolution mask incurs quadratic computation overhead. It is a challenging task to predict the optimal binary mask for each instance. In this paper, we propose to dynamically select suitable masks for different object proposals. First, a dual-level Feature Pyramid Network (FPN) with adaptive feature aggregation is developed to gradually increase the mask grid resolution, ensuring high-quality segmentation of objects. Specifically, an efficient region-level top-down path (r-FPN) is introduced to incorporate complementary contextual and detailed information from different stages of image-level FPN (i-FPN). Then, to alleviate the increase of computation and memory costs caused by using large masks, we develop a Mask Switch Module (MSM) with negligible computational cost to select the most suitable mask resolution for each instance, achieving high efficiency while maintaining high segmentation accuracy. Without bells and whistles, the proposed method, namely DynaMask, brings consistent and noticeable performance improvements over other state-of-the-arts at a moderate computation overhead. The source code: https://github.com/lslrh/DynaMask. Ruihuang Li, Chenhang He, Shuai Li 0014, Yabin Zhang 0001, Lei Zhang 0006 |
CVPR | 5 |
| 2023 | SIM: Semantic-aware Instance Mask Generation for Box-Supervised Instance SegmentationabstractWeakly supervised instance segmentation using only bounding box annotations has recently attracted much re-search attention. Most of the current efforts leverage low-level image features as extra supervision without explicitly exploiting the high-level semantic information of the objects, which will become ineffective when the foreground objects have similar appearances to the background or other objects nearby. We propose a new box-supervised instance segmentation approach by developing a Semantic-aware In-stance Mask (SIM) generation paradigm. Instead of heavily relying on local pair-wise affinities among neighboring pixels, we construct a group of category-wise feature cen-troids as prototypes to identify foreground objects and as-sign them semantic-level pseudo labels. Considering that the semantic-aware prototypes cannot distinguish different instances of the same semantics, we propose a self-correction mechanism to rectify the falsely activated regions while enhancing the correct ones. Furthermore, to handle the occlusions between objects, we tailor the Copy-Paste operation for the weakly-supervised instance segmentation task to augment challenging training data. Extensive experimental results demonstrate the superiority of our proposed SIM approach over other state-of-the-art methods. The source code: https://github.com/1slrh/SIM. Ruihuang Li, Chenhang He, Yabin Zhang 0001, Shuai Li 0014, Liyi Chen 0002, Lei Zhang 0006 |
CVPR | 6 |
| 2023 | One-to-Few Label Assignment for End-to-End Dense DetectionabstractOne-to-one (o2o) label assignment plays a key role for transformer based end-to-end detection, and it has been recently introduced in fully convolutional detectors for end-to-end dense detection. However, o2o can degrade the feature learning efficiency due to the limited number of positive samples. Though extra positive samples are introduced to mitigate this issue in recent DETRs, the computation of self- and cross- attentions in the decoder limits its practical application to dense and fully convolutional detectors. In this work, we propose a simple yet effective one-to-few (o2f) label assignment strategy for end-to-end dense detection. Apart from defining one positive and many negative anchors for each object, we define several soft anchors, which serve as positive and negative samples simultaneously. The positive and negative weights of these soft anchors are dynamically adjusted during training so that they can contribute more to “representation learning” in the early training stage, and contribute more to “duplicated prediction removal” in the later stage. The detector trained in this way can not only learn a strong feature representation but also perform end-to-end dense detection. Experiments on COCO and CrowdHuman datasets demonstrate the effectiveness of the o2f scheme. Code is available at https://github.com/strongwolf/o2f Shuai Li 0014, Minghan Li 0001, Ruihuang Li, Chenhang He, Lei Zhang 0006 |
CVPR | 5 |
| 2023 | Joint HDR Denoising and Fusion: A Real-World Mobile HDR Image DatasetabstractMobile phones have become a ubiquitous and indispensable photographing device in our daily life, while the small aperture and sensor size make mobile phones more susceptible to noise and over-saturation, resulting in low dynamic range (LDR) and low image quality. It is thus crucial to develop high dynamic range (HDR) imaging techniques for mobile phones. Unfortunately, the existing HDR image datasets are mostly constructed by DSLR cameras in daytime, limiting their applicability to the study of HDR imaging for mobile phones. In this work, we develop, for the first time to our best knowledge, an HDR image dataset by using mobile phone cameras, namely Mobile-HDR dataset. Specifically, we utilize three mobile phone cameras to collect paired LDR-HDR images in the raw image domain, covering both daytime and night-time scenes with different noise levels. We then propose a transformer based model with a pyramid cross-attention alignment module to aggregate highly correlated features from different exposure frames to perform joint HDR denoising and fusion. Experiments validate the advantages of our dataset and our method on mobile HDR imaging. Dataset and codes are available at https://github.com/shuaizhengliu/Joint-HDRDN. Shuaizheng Liu, Lingchen Sun, Zhetong Liang, Hui Zeng 0001, Lei Zhang 0006 |
CVPR | 6 |
| 2023 | OTAvatar: One-Shot Talking Face Avatar with Controllable Tri-Plane RenderingabstractControllability, generalizability and efficiency are the major objectives of constructing face avatars represented by neural implicit field. However, existing methods have not managed to accommodate the three requirements simultaneously. They either focus on static portraits, restricting the representation ability to a specific subject, or suffer from substantial computational cost, limiting their flexibility. In this paper, we propose One-shot Talking face Avatar (OTAvatar), which constructs face avatars by a generalized controllable tri-plane rendering solution so that each personalized avatar can be constructed from only one portrait as the reference. Specifically, OTAvatar first inverts a portrait image to a motion-free identity code. Second, the identity code and a motion code are utilized to modulate an efficient CNN to generate a tri-plane formulated volume, which encodes the subject in the desired motion. Finally, volume rendering is employed to generate an image in any view. The core of our solution is a novel decoupling-by-inverting strategy that disentangles identity and motion in the latent code via optimization-based inversion. Benefiting from the efficient tri-plane representation, we achieve controllable rendering of generalized face avatar at 35 FPS on AIOO. Experiments show promising performance of crossidentity reenactment on subjects out of the training set and better 3D consistency. The code is available at https://github.com/theEricMaIOTAvatar. Zhiyuan Ma 0002, Xiangyu Zhu 0001, Guo-Jun Qi, Zhen Lei 0001, Lei Zhang 0006 |
CVPR | 5 |
| 2023 | Sharpness-Aware Gradient Matching for Domain GeneralizationabstractThe goal of domain generalization (DG) is to enhance the generalization capability of the model learned from a source domain to other unseen domains. The recently developed Sharpness-Aware Minimization (SAM) method aims to achieve this goal by minimizing the sharpness measure of the loss landscape. Though SAM and its variants have demonstrated impressive DG performance, they may not always converge to the desired flat region with a small loss value. In this paper, we present two conditions to ensure that the model could converge to a flat minimum with a small loss, and present an algorithm, named Sharpness-Aware Gradient Matching (SAGM), to meet the two conditions for improving model generalization capability. Specifically, the optimization objective of SAGM will simultaneously minimize the empirical risk, the perturbed loss (i.e., the maximum loss within a neighborhood in the parameter space), and the gap between them. By implicitly aligning the gradient directions between the empirical risk and the perturbed loss, SAGM improves the generalization capability over SAM and its variants without increasing the computational cost. Extensive experimental results show that our proposed SAGM method consistently outperforms the state-of-the-art methods on five DG benchmarks, including PACS, VLCS, OfficeHome, TerraIncognita, and DomainNet. Codes are available at https://github.com/Wang-pengfei/SAGM. Pengfei Wang 0009, Zhaoxiang Zhang 0001, Zhen Lei 0001, Lei Zhang 0006 |
CVPR | 4 |
| 2023 | A General Regret Bound of Preconditioned Gradient Method for DNN TrainingabstractWhile adaptive learning rate methods, such as Adam, have achieved remarkable improvement in optimizing Deep Neural Networks (DNNs), they consider only the diagonal elements of the full preconditioned matrix. Though the full-matrix preconditioned gradient methods theoretically have a lower regret bound, they are impractical for use to train DNNs because of the high complexity. In this paper, we present a general regret bound with a constrained full-matrix preconditioned gradient, and show that the updating formula of the preconditioner can be derived by solving a cone-constrained optimization problem. With the block-diagonal and Kronecker-factorized constraints, a specific guide function can be obtained. By minimizing the upper bound of the guide function, we develop a new DNN optimizer, termed AdaBK. A series of techniques, including statistics updating, dampening, efficient matrix inverse root computation, and gradient amplitude preservation, are developed to make AdaBK effective and efficient to implement. The proposed AdaBK can be readily embedded into many existing DNN optimizers, e.g., SGDM and Adam W, and the corresponding SGDM _BK and Adam W _BK algorithms demonstrate significant improvements over existing DNN optimizers on benchmark vision tasks, including image classification, object detection and segmentation. The code is publicly available at https://github.com/Yonghongwei/AdaBK. Hongwei Yong, Lei Zhang 0006 |
CVPR | 3 |
| 2023 | FPR: False Positive Rectification for Weakly Supervised Semantic SegmentationabstractMany weakly supervised semantic segmentation (WSSS) methods employ the class activation map (CAM) to generate the initial segmentation results. However, CAM often fails to distinguish the foreground from its co-occurred background (e.g., train and railroad), resulting in inaccurate activation from the background. Previous endeavors address this co-occurrence issue by introducing external supervision and human priors. In this paper, we present a False Positive Rectification (FPR) approach to tackle the co-occurrence problem by leveraging the false positives of CAM. Based on the observation that the CAM-activated regions of absent classes contain class-specific co-occurred background cues, we collect these false positives and utilize them to guide the training of CAM network by proposing a region-level contrast loss and a pixel-level rectification loss. Without introducing any external supervision and human priors, the proposed FPR effectively suppresses wrong activations from the background objects. Extensive experiments on the PASCAL VOC 2012 and MS COCO 2014 demonstrate that FPR brings significant improvements for off-the-shelf methods and achieves state-of-the-art performance. Code is available at https://github.com/mt-cly/FPR. Liyi Chen 0002, Chenyang Lei, Ruihuang Li, Shuai Li 0014, Zhaoxiang Zhang 0001, Lei Zhang 0006 |
ICCV | 6 |
| 2023 | Point2Mask: Point-supervised Panoptic Segmentation via Optimal TransportabstractWeakly-supervised image segmentation has recently attracted increasing research attentions, aiming to avoid the expensive pixel-wise labeling. In this paper, we present an effective method, namely Point2Mask, to achieve high-quality panoptic prediction using only a single random point annotation per target for training. Specifically, we formulate the panoptic pseudo-mask generation as an Optimal Transport (OT) problem, where each ground-truth (gt) point label and pixel sample are defined as the label supplier and consumer, respectively. The transportation cost is calculated by the introduced task-oriented maps, which focus on the category-wise and instance-wise differences among the various thing and stuff targets. Furthermore, a centroid-based scheme is proposed to set the accurate unit number for each gt point supplier. Hence, the pseudo-mask generation is converted into finding the optimal transport plan at a globally minimal transportation cost, which can be solved via the Sinkhorn-Knopp Iteration. Experimental results on Pascal VOC and COCO demonstrate the promising performance of our proposed Point2Mask approach to point-supervised panoptic segmentation. Source code is available at: https://github.com/LiWentomng/Point2Mask. Wentong Li 0001, Yuqian Yuan, Song Wang 0019, Jianke Zhu, Jianshu Li, Jian Liu 0012, Lei Zhang 0006 |
ICCV | 7 |
| 2023 | A Benchmark for Chinese-English Scene Text Image Super-resolutionabstractScene Text Image Super-resolution (STISR) aims to recover high-resolution (HR) scene text images with visually pleasant and readable text content from the given low-resolution (LR) input. Most existing works focus on recovering English texts, which have relatively simple character structures, while little work has been done on the more challenging Chinese texts with diverse and complex character structures. In this paper, we propose a real-world Chinese-English benchmark dataset, namely Real-CE, for the task of STISR with the emphasis on restoring structurally complex Chinese characters. The benchmark provides 1,935/783 real-world LR-HR text image pairs (contains 33,789 text lines in total) for training/testing in 2× and 4× zooming modes, complemented by detailed annotations, including detection boxes and text transcripts. Moreover, we design an edge-aware learning method, which provides structural supervision in image and feature domains, to effectively reconstruct the dense structures of Chinese characters. We conduct experiments on the proposed Real-CE benchmark and evaluate the existing STISR models with and without our edge-aware loss. The benchmark, including data and source code, is available at https://github.com/mjq11302010044/Real-CE. Jianqi Ma, Zhetong Liang, Wangmeng Xiang, Xi Yang 0022, Lei Zhang 0006 |
ICCV | 5 |
| 2023 | Core: Cooperative Reconstruction for Multi-Agent PerceptionabstractThis paper presents Core, a conceptually simple, effective and communication-efficient model for multi-agent cooperative perception. It addresses the task from a novel perspective of cooperative reconstruction, based on two key insights: 1) cooperating agents together provide a more holistic observation of the environment, and 2) the holistic observation can serve as valuable supervision to explicitly guide the model learning how to reconstruct the ideal observation based on collaboration. Core instantiates the idea with three major components: a compressor for each agent to create more compact feature representation for efficient broadcasting, a lightweight attentive collaboration component for cross-agent message aggregation, and a reconstruction module to reconstruct the observation based on aggregated feature representations. This learning-to-reconstruct idea is task-agnostic, and offers clear and reasonable supervision to inspire more effective collaboration, eventually promoting perception tasks. We validate Core on two large-scale multi-agent percetion dataset, OPV2V and V2X-Sim, in two tasks, i.e., 3D object detection and semantic segmentation. Results demonstrate that Core achieves state-of-the-art performance, and is more communication-efficient. Binglu Wang, Lei Zhang 0006, Zhaozhong Wang, Yongqiang Zhao 0001, Tianfei Zhou |
ICCV | 2 |
| 2023 | ELITE: Encoding Visual Concepts into Textual Embeddings for Customized Text-to-Image GenerationabstractIn addition to the unprecedented ability in imaginary creation, large text-to-image models are expected to take customized concepts in image generation. Existing works generally learn such concepts in an optimization-based manner, yet bringing excessive computation or memory burden. In this paper, we instead propose a learning-based encoder, which consists of a global and a local mapping networks for fast and accurate customized text-to-image generation. In specific, the global mapping network projects the hierarchical features of a given image into multiple "new" words in the textual word embedding space, i.e., one primary word for well-editable concept and other auxiliary words to exclude irrelevant disturbances (e.g., background). In the meantime, a local mapping network injects the encoded patch features into cross attention layers to provide omitted details, without sacrificing the editability of primary concepts. We compare our method with existing optimization-based approaches on a variety of user-defined concepts, and demonstrate that our method enables high-fidelity inversion and more robust editability with a significantly faster encoding process. Our code is publicly available at https://github.com/csyxwei/ELITE. Yuxiang Wei 0001, Yabo Zhang, Zhilong Ji, Jinfeng Bai, Lei Zhang 0006, Wangmeng Zuo |
ICCV | 5 |
| 2023 | Generative Action Description Prompts for Skeleton-based Action RecognitionabstractSkeleton-based action recognition has recently received considerable attention. Current approaches to skeleton-based action recognition are typically formulated as one-hot classification tasks and do not fully exploit the semantic relations between actions. For example, "make victory sign" and "thumb up" are two actions of hand gestures, whose major difference lies in the movement of hands. This information is agnostic from the categorical one-hot encoding of action classes but could be unveiled from the action description. Therefore, utilizing action description in training could potentially benefit representation learning. In this work, we propose a Generative Action-description Prompts (GAP) approach for skeleton-based action recognition. More specifically, we employ a pre-trained large-scale language model as the knowledge engine to automatically generate text descriptions for body parts movements of actions, and propose a multi-modal training scheme by utilizing the text encoder to generate feature vectors for different body parts and supervise the skeleton encoder for action representation learning. Experiments show that our proposed GAP method achieves noticeable improvements over various baseline models without extra computation cost at inference. GAP achieves new state-of-the-arts on popular skeleton-based action recognition benchmarks, including NTU RGB+D, NTU RGB+D 120 and NW-UCLA. The source code is available at https://github.com/MartinXM/GAP. Wangmeng Xiang, Chao Li 0034, Yuxuan Zhou 0004, Lei Zhang 0006 |
ICCV | 5 |
| 2023 | Isomer: Isomerous Transformer for Zero-shot Video Object SegmentationabstractRecent leading zero-shot video object segmentation (ZVOS) works devote to integrating appearance and motion information by elaborately designing feature fusion modules and identically applying them in multiple feature stages. Our preliminary experiments show that with the strong long-range dependency modeling capacity of Transformer, simply concatenating the two modality features and feeding them to vanilla Transformers for feature fusion can distinctly benefit the performance but at a cost of heavy computation. Through further empirical analysis, we find that attention dependencies learned in Transformer in different stages exhibit completely different properties: global query-independent dependency in the low-level stages and semantic-specific dependency in the high-level stages. Motivated by the observations, we propose two Transformer variants: i) Context-Sharing Transformer (CST) that learns the global-shared contextual information within image frames with a lightweight computation. ii) Semantic Gathering-Scattering Transformer (SGST) that models the semantic correlation separately for the foreground and background and reduces the computation cost with a soft token merging mechanism. We apply CST and SGST for low-level and high-level feature fusions, respectively, formulating a level-isomerous Transformer framework for ZVOS task. Compared with the baseline that uses vanilla Transformers for multi-stage fusion, ours significantly increase the speed by 13× and achieves new state-of-the-art ZVOS performance. Code is available at https://github.com/DLUT-yyc/Isomer. Yifan Wang 0004, Lijun Wang 0001, Xiaoqi Zhao 0003, Huchuan Lu, Yu Wang 0108, Weibo Su, Lei Zhang 0006 |
ICCV | 8 |
| 2023 | Towards Fairness-aware Adversarial Network PruningabstractNetwork pruning aims to compress models while minimizing loss in accuracy. With the increasing focus on bias in AI systems, the bias inheriting or even magnification nature of traditional network pruning methods has raised a new perspective towards fairness-aware network pruning. Straightforward pruning plus debias methods and recent designs for monitoring disparities of demographic attributes during pruning have endeavored to enhance fairness in pruning. However, neither simple assembling of two tasks nor specifically designed pruning strategies could achieve the optimal trade-off among pruning ratio, accuracy, and fairness. This paper proposes an end-to-end learnable framework for fairness-aware network pruning, which optimizes both pruning and debias tasks jointly by adversarial training against those final evaluation metrics like accuracy for pruning, and disparate impact (DI) and equalized odds (DEO) for fairness. In other words, our fairness-aware adversarial pruning method would learn to prune without any handcraft rules. Therefore, our approach could flexibly adapt to variate network structures. Exhaustive experimentation demonstrates the generalization capacity of our approach, as well as superior performance on pruning and debias simultaneously. To highlight, the proposed method could preserve the SOTA pruning performance while significantly improving fairness by around 50% as compared to traditional pruning methods. Lei Zhang 0006, Zhibo Wang 0001, Xiaowei Dong, Yunhe Feng, Xiaoyi Pang, Kui Ren 0001 |
ICCV | 1 |
| 2023 | Flow-Guided Deformable Attention Network for Fast Online Video Super-ResolutionabstractReal-time online video super-resolution (VSR) on resource limited applications is a very challenging problem due to the constraints on complexity, latency and memory foot-print, etc. Recently, a series of fast online VSR methods have been proposed to tackle this issue. In particular, attention based methods have achieved much progress by adaptively aligning or aggregating the information in preceding frames. However, these methods are still limited in network design to effectively and efficiently propagate the useful features in temporal domain. In this work, we propose a new fast online VSR algorithm with a flow-guided deformable attention propagation module, which leverages corresponding priors provided by a fast optical flow network in deformable attention computation and consequently helps propagating recurrent state information effectively and efficiently. The proposed algorithm achieves state-of-the-art results on widely-used benchmarking VSR datasets in terms of effectiveness and efficiency. Code can be found at https://github.com/IanYeung/FastOnlineVSR. Xi Yang 0022, Lei Zhang 0006 |
ICIP | 3 |
| 2023 | Label-efficient Segmentation via Affinity PropagationabstractWeakly-supervised segmentation with label-efficient sparse annotations has attracted increasing research attention to reduce the cost of laborious pixel-wise labeling process, while the pairwise affinity modeling techniques play an essential role in this task. Most of the existing approaches focus on using the local appearance kernel to model the neighboring pairwise potentials. However, such a local operation fails to capture the long-range dependencies and ignores the topology of objects. In this work, we formulate the affinity modeling as an affinity propagation process, and propose a local and a global pairwise affinity terms to generate accurate soft pseudo labels. An efficient algorithm is also developed to reduce significantly the computational cost. The proposed approach can be conveniently plugged into existing segmentation networks. Experiments on three typical label-efficient segmentation tasks, i.e. box-supervised instance segmentation, point/scribble-supervised semantic segmentation and CLIP-guided semantic segmentation, demonstrate the superior performance of the proposed approach. Wentong Li 0001, Yuqian Yuan, Song Wang 0019, Wenyu Liu 0005, Dongqi Tang, Jian Liu 0012, Jianke Zhu, Lei Zhang 0006 |
NeurIPS | 8 |
| 2023 | NKFAC: A Fast and Stable KFAC Optimizer for Deep Neural Networks
Hongwei Yong, Lei Zhang 0006 |
ECML/PKDD (4) | 3 |
| 2023 | Uncovering User Interest from Biased and Noised Watch Time in Video RecommendationabstractIn the video recommendation, watch time is commonly adopted as an indicator of user interest. However, watch time is not only influenced by the matching of users’ interests but also by other factors, such as duration bias and noisy watching. Duration bias refers to the tendency for users to spend more time on videos with longer durations, regardless of their actual interest level. Noisy watching, on the other hand, describes users taking time to determine whether they like a video or not, which can result in users spending time watching videos they do not like. Consequently, the existence of duration bias and noisy watching make watch time an inadequate label for indicating user interest. Furthermore, current methods primarily address duration bias and ignore the impact of noisy watching, which may limit their effectiveness in uncovering user interest from watch time. In this study, we first analyze the generation mechanism of users’ watch time from a unified causal viewpoint. Specifically, we considered the watch time as a mixture of the user’s actual interest level, the duration-biased watch time, and the noisy watch time. To mitigate both the duration bias and noisy watching, we propose Debiased and Denoised watch time Correction (D2Co), which can be divided into two steps: First, we employ a duration-wise Gaussian Mixture Model plus frequency-weighted moving average for estimating the bias and noise terms; then we utilize a sensitivity-controlled correction function to separate the user interest from the watch time, which is robust to the estimation error of bias and noise terms. The experiments on two public video recommendation datasets and online A/B testing indicate the effectiveness of the proposed method. Haiyuan Zhao, Lei Zhang 0006, Jun Xu 0001, Guohao Cai, Zhenhua Dong, Ji-Rong Wen |
RecSys | 2 |
| 2023 | Multiple-bounce Smith Microfacet BRDFs using the Invariance PrincipleabstractSmith microfacet models are widely used in computer graphics to represent materials. Traditional microfacet models do not consider the multiple bounces on microgeometries, leading to visible energy missing, especially on rough surfaces. Later, as the equivalence between the microfacets and volume has been revealed, random walk solutions have been proposed to introduce multiple bounces, but at the cost of high variance. Recently, the position-free property has been introduced into the multiple-bounce model, resulting in much less noise, but also bias or a complex derivation.In this paper, we propose a simple way to derive the multiple-bounce Smith microfacet bidirectional reflectance distribution functions (BRDFs) using the invariance principle. At the core of our model is a shadowing-masking function for a path consisting of direction collections, rather than separated bounces. Our model ensures unbiasedness and can produce less noise compared to the previous work with equal time, thanks to the simple formulation. Furthermore, we also propose a novel probability density function (PDF) for BRDF multiple importance sampling, which has a better match with the multiple-bounce BRDFs, producing less noise than previous naive approximations. Yuang Cui, Gaole Pan, Jian Yang 0003, Lei Zhang 0006, Lingqi Yan 0001, Beibei Wang 0002 |
SIGGRAPH Asia | 4 |
| 2023 | Conditional counterfactual causal effect for individual attributionabstractIdentifying the causes of an event, also termed as causal attribution, is a commonly encountered task in many application problems. Available methods, mostly in Bayesian or causal inference literature, suffer from two main drawbacks: 1) cannot attribute for individuals, and 2) attributing one single cause at a time and cannot deal with the interaction effect among multiple causes. In this paper, based on our proposed new measurement, called conditional counterfactual causal effect (CCCE), we introduce an individual causal attribution method, which is able to utilize the individual observation as the evidence and consider common influence and interaction effect of multiple causes simultaneously. We discuss the identifiability of CCCE and also give the identification formulas under proper assumptions. Finally, we conduct experiments on simulated and real data to illustrate the effectiveness of CCCE and the results show that our proposed method outperforms significantly state-of-the-art methods. Lei Zhang 0006, Shengyu Zhu 0001, Zitong Lu, Zhenhua Dong, Chaoliang Zhang, Jun Xu 0001, Zhi Geng, Yangbo He |
UAI | 2 |
| 2023 | A Novel Exponential Time Delayed Fractional Grey Model and Its Application in Forecasting Oil Production and Consumption of ChinaabstractBased on particle swarm optimization (PSO), a new exponential time delay fraction order grey prediction model is proposed in this paper. Firstly, the original data is preprocessed by fractional-order accumulation; on the basis of fractional-order accumulation, it is proved that the initial value of the original sequence satisfies the fixed point theorem. On the basis of GM(1,1) model, a new model is established by adding exponential time delay term. The model is discretized by integral, the least square estimation of the linear parameters and the approximate time response equation are obtained. Finally, PSO is used to search the optimal parameters of the model and the experimental results are verified by Wilcoxon rank sum test. In order to test the good adaptability and strong prediction ability of the new model, two groups of data are verified for the production and consumption of oil and electricity in China. The results show that the new model has better prediction accuracy and adaptability than the other existing six grey models. Yong Wang 0032, Lei Zhang 0006, Xinbo He, Xin Ma 0004, Wenqing Wu 0001, Pei Chi |
Cybern. Syst. | 2 |
| 2023 | Survey on leveraging pre-trained generative adversarial networks for image editing and restorationabstractGenerative adversarial networks (GANs) have drawn enormous attention due to their simple yet effective training mechanism and superior image generation quality. With the ability to generate photorealistic high-resolution (e.g., 1024 × 1024) images, recent GAN models have greatly narrowed the gaps between the generated images and the real ones. Therefore, many recent studies show emerging interest to take advantage of pre-trained GAN models by exploiting the well-disentangled latent space and the learned GAN priors. In this study, we briefly review recent progress on leveraging pre-trained large-scale GAN models from three aspects, i.e., (1) the training of large-scale generative adversarial networks, (2) exploring and understanding the pre-trained GAN models, and (3) leveraging these models for subsequent tasks like image restoration and editing. Ming Liu 0018, Yuxiang Wei 0001, Xiaohe Wu, Wangmeng Zuo, Lei Zhang 0006 |
Sci. China Inf. Sci. | 5 |
| 2023 | Learning Dual Memory Dictionaries for Blind Face RestorationabstractBlind face restoration is a challenging task due to the unknown, unsynthesizable and complex degradation, yet is valuable in many practical applications. To improve the performance of blind face restoration, recent works mainly treat the two aspects, i.e., generic and specific restoration, separately. In particular, generic restoration attempts to restore the results through general facial structure prior, while on the one hand, cannot generalize to real-world degraded observations due to the limited capability of direct CNNs' mappings in learning blind restoration, and on the other hand, fails to exploit the identity-specific details. On the contrary, specific restoration aims to incorporate the identity features from the reference of the same identity, in which the requirement of proper reference severely limits the application scenarios. Generally, it is a challenging and intractable task to improve the photo-realistic performance of blind restoration and adaptively handle the generic and specific restoration scenarios with a single unified model. Instead of implicitly learning the mapping from a low-quality image to its high-quality counterpart, this paper suggests a DMDNet by explicitly memorizing the generic and specific features through dual dictionaries. First, the generic dictionary learns the general facial priors from high-quality images of any identity, while the specific dictionary stores the identity-belonging features for each person individually. Second, to handle the degraded input with or without specific reference, dictionary transform module is suggested to read the relevant details from the dual dictionaries which are subsequently fused into the input features. Finally, multi-scale dictionaries are leveraged to benefit the coarse-to-fine restoration. The whole framework including the generic and specific dictionaries is optimized in an end-to-end manner and can be flexibly plugged into different application scenarios. Moreover, a new high-quality dataset, termed CelebRef-HQ, is constructed to promote the exploration of specific face restoration in the high-resolution space. Experimental results demonstrate that the proposed DMDNet performs favorably against the state of the arts in both quantitative and qualitative evaluation, and generates more photo-realistic results on the real-world low-quality images. The codes, models and the CelebRef-HQ dataset will be publicly available at https://github.com/csxmli2016/DMDNet. Xiaoming Li 0002, Shiguang Zhang, Shangchen Zhou, Lei Zhang 0006, Wangmeng Zuo |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2023 | Improving Nighttime Driving-Scene Segmentation via Dual Image-Adaptive Learnable FiltersabstractSemantic segmentation on driving-scene images is vital for autonomous driving. Although encouraging performance has been achieved on daytime images, the performance on nighttime images are less satisfactory due to the insufficient exposure and the lack of labeled data. To address these issues, we present an add-on module called dual image-adaptive learnable filters (DIAL-Filters) to improve the semantic segmentation in nighttime driving conditions, aiming at exploiting the intrinsic features of driving-scene images under different illuminations. DIAL-Filters consist of two parts, including an image-adaptive processing module (IAPM) and a learnable guided filter (LGF). With DIAL-Filters, we design both unsupervised and supervised frameworks for nighttime driving-scene segmentation, which can be trained in an end-to-end manner. Specifically, the IAPM module consists of a small convolutional neural network with a set of differentiable image filters, where each image can be adaptively enhanced for better segmentation with respect to the different illuminations. The LGF is employed to enhance the output of segmentation network to get the final segmentation result. The DIAL-Filters are light-weight and efficient and they can be readily applied for both daytime and nighttime images. Our experiments show that DAIL-Filters can significantly improve the supervised segmentation performance on ACDC_Night and NightCity datasets, while it demonstrates the state-of-the-art performance on unsupervised nighttime semantic segmentation on Dark Zurich and Nighttime Driving testbeds. Codes and models are available athttps://github.com/wenyyu/IA-Seg. Wenyu Liu 0005, Wentong Li 0001, Jianke Zhu, Miaomiao Cui, Xuansong Xie, Lei Zhang 0006 |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2023 | VirFace∞: A Semi-Supervised Method for Enhancing Face Recognition via Unlabeled Shallow DataabstractThe semi-supervised face recognition problem has become a popular research topic in recent years. However, one common and important situation, in which the unlabeled data is shallow, has rarely been considered in most existing works. In this paper, shallow data means there are only few images per identity. In the unlabeled shallow situation, the existing semi-supervised face recognition methods generally do not work well. Thus, how to effectively utilize the unlabeled shallow face data for improving face recognition performance is an important issue. In this paper, we propose a novel semi-supervised face recognition method, namely VirFace$^{\infty} $, to enhance the face recognition performance effectively with the unlabeled shallow data. VirFace$^{\infty} $consists of VirClass and VirDistribution components. In VirClass, we inject the unlabeled data as virtual classes into the feature space to enlarge the inter-class distance. In VirDistribution, we predict the distribution of each virtual class, namely virtual distribution, and then enhance the inter-class discriminativeness by enlarging the distances between the labeled features and the virtual distributions. To the best of our knowledge, we are among the first to tackle the face recognition problem on unlabeled shallow face data. Extensive experiments demonstrate the superiority of our proposed method. Tianchu Guo, Binghui Chen, Wangmeng Zuo, Lei Zhang 0006 |
IEEE Trans. Inf. Forensics Secur. | 7 |
| 2023 | Text Prior Guided Scene Text Image Super-ResolutionabstractScene text image super-resolution (STISR) aims to improve the resolution and visual quality of low-resolution (LR) scene text images, while simultaneously boost the performance of text recognition. However, most of the existing STISR methods regard text images as natural scene images, ignoring the categorical information of text. In this paper, we make an inspiring attempt to embed text recognition prior into STISR model. Specifically, we adopt the predicted character recognition probability sequence as the text prior, which can be obtained conveniently from a text recognition model. The text prior provides categorical guidance to recover high-resolution (HR) text images. On the other hand, the reconstructed HR image can refine the text prior in return. Finally, we present a multi-stage text prior guided super-resolution (TPGSR) framework for STISR. Our experiments on the benchmark TextZoom dataset show that TPGSR can not only effectively improve the visual quality of scene text images, but also significantly improve the text recognition accuracy over existing STISR methods. Our model trained on TextZoom also demonstrates certain generalization capability to the LR images in other datasets. The source code of our work is available at: https://github.com/mjq11302010044/TPGSR. Jianqi Ma, Shi Guo, Lei Zhang 0006 |
IEEE Trans. Image Process. | 3 |
| 2023 | One-Stage Visual Relationship Referring With Transformers and Adaptive Message PassingabstractThere exist a variety of visual relationships among entities in an image. Given a relationship query $\langle subject, predicate, object \rangle $ , the task of visual relationship referring (VRR) aims to disambiguate instances of the same entity category and simultaneously localize the subject and object entities in an image. Previous works of VRR can be generally categorized into one-stage and multi-stage methods. The former ones directly localize a pair of entities from the image but they suffer from low prediction accuracy, while the latter ones perform better but they are indirect to localize only a couple of entities by pre-generating a rich amount of candidate proposals. In this paper, we formulate the task of VRR as an end-to-end bounding box regression problem and propose a novel one-stage approach, called VRR-TAMP, by effectively integrating Transformers and an adaptive message passing mechanism. First, visual relationship queries and images are respectively encoded to generate the basic modality-specific embeddings, which are then fed into a cross-modal Transformer encoder to produce the joint representation. Second, to obtain the specific representation of each entity, we introduce an adaptive message passing mechanism and design an entity-specific information distiller SR-GMP, which refers to a gated message passing (GMP) module that works on the joint representation learned from a single learnable token. The GMP module adaptively distills the final representation of an entity by incorporating the contextual cues regarding the predicate and the other entity. Experiments on VRD and Visual Genome datasets demonstrate that our approach significantly outperforms its one-stage competitors and achieves competitive results with the state-of-the-art multi-stage methods. Youtian Du, Yabin Zhang 0001, Shuai Li 0014, Lei Zhang 0006 |
IEEE Trans. Image Process. | 5 |
| 2022 | Image-Adaptive YOLO for Object Detection in Adverse Weather ConditionsabstractThough deep learning-based object detection methods have achieved promising results on the conventional datasets, it is still challenging to locate objects from the low-quality images captured in adverse weather conditions. The existing methods either have difficulties in balancing the tasks of image enhancement and object detection, or often ignore the latent information beneficial for detection. To alleviate this problem, we propose a novel Image-Adaptive YOLO (IA-YOLO) framework, where each image can be adaptively enhanced for better detection performance. Specifically, a differentiable image processing (DIP) module is presented to take into account the adverse weather conditions for YOLO detector, whose parameters are predicted by a small convolutional neural network (CNN-PP). We learn CNN-PP and YOLOv3 jointly in an end-to-end fashion, which ensures that CNN-PP can learn an appropriate DIP to enhance the image for detection in a weakly supervised manner. Our proposed IA-YOLO approach can adaptively process images in both normal and adverse weather conditions. The experimental results are very encouraging, demonstrating the effectiveness of our proposed IA-YOLO method in both foggy and low-light scenarios. The source code can be found at https://github.com/wenyyu/Image-Adaptive-YOLO. Wenyu Liu 0005, Gaofeng Ren, Runsheng Yu, Shi Guo, Jianke Zhu, Lei Zhang 0006 |
AAAI | 6 |
| 2022 | Efficient Hardware-Aware Neural Architecture Search for Image Super-Resolution on Mobile Devices
Hui Zeng 0001, Lei Zhang 0006 |
ACCV (3) | 3 |
| 2022 | SP-ViT: Learning 2D Spatial Priors for Vision Transformers
Yuxuan Zhou 0004, Wangmeng Xiang, Chao Li 0034, Xihan Wei, Lei Zhang 0006, Margret Keuper, Xian-Sheng Hua 0001 |
BMVC | 6 |
| 2022 | Dense Learning based Semi-Supervised Object DetectionabstractSemi-supervised object detection (SSOD) aims to facilitate the training and deployment of object detectors with the help of a large amount of unlabeled data. Though various self-training based and consistency-regularization based SSOD methods have been proposed, most of them are anchor-based detectors, ignoring the fact that in many real-world applications anchor-free detectors are more demanded. In this paper, we intend to bridge this gap and propose a DenSe Learning (DSL) based anchor-free SSOD algorithm. Specifically, we achieve this goal by introducing several novel techniques, including an Adaptive Filtering strategy for assigning multi-level and accurate dense pixel-wise pseudo-labels, an Aggregated Teacher for producing stable and precise pseudo-labels, and an uncertainty-consistency-regularization term among scales and shuffled patches for improving the generalization capability of the detector. Extensive experiments are conducted on MS-COCO and PASCAL-VOC, and the results show that our proposed DSL method records new state-of-the-art SSOD performance, surpassing existing methods by a large margin. Codes can be found at https://github.com/chenbinghui1/DSL. Binghui Chen, Lei Zhang 0006, Xian-Sheng Hua 0001 |
CVPR | 5 |
| 2022 | A Differentiable Two-stage Alignment Scheme for Burst Image Reconstruction with Large ShiftabstractDenoising and demosaicking are two essential steps to reconstruct a clean full-color image from the raw data. Recently, joint denoising and demosaicking (JDD) for burst images, namely JDD-B, has attracted much attention by using multiple raw images captured in a short time to reconstruct a single high-quality image. One key challenge of JDD-B lies in the robust alignment of image frames. State-of-the-art alignment methods in feature domain cannot effectively utilize the temporal information of burst images, where large shifts commonly exist due to camera and object motion. In addition, the higher resolution (e.g., 4K) of modern imaging devices results in larger displacement between frames. To address these challenges, we design a differentiable two-stage alignment scheme sequentially in patch and pixel level for effective JDD-B. The input burst images are firstly aligned in the patch level by using a differentiable progressive block matching method, which can estimate the offset between distant frames with small computational cost. Then we perform implicit pixel-wise alignment in full-resolution feature domain to refine the alignment results. The two stages are jointly trained in an end-to-end manner. Extensive experiments demonstrate the significant improvement of our method over existing JDD-B methods. Codes are available at https://github.com/GuoShi28/2StageAlign. Shi Guo, Xi Yang 0001, Jianqi Ma, Gaofeng Ren, Lei Zhang 0006 |
CVPR | 5 |
| 2022 | Voxel Set Transformer: A Set-to-Set Approach to 3D Object Detection from Point CloudsabstractTransformer has demonstrated promising performance in many 2D vision tasks. However, it is cumbersome to compute the self-attention on large-scale point cloud data because point cloud is a long sequence and unevenly distributed in 3D space. To solve this issue, existing methods usually compute self-attention locally by grouping the points into clusters of the same size, or perform convolutional self-attention on a discretized representation. However, the former results in stochastic point dropout, while the latter typically has narrow attention fields. In this paper, we propose a novel voxel-based architecture, namely Voxel Set Transformer (VoxSeT), to detect 3D objects from point clouds by means of set-to-set translation. VoxSeT is built upon a voxel-based set attention (VSA) module, which reduces the self-attention in each voxel by two cross-attentions and models features in a hidden space induced by a group of latent codes. With the VSA module, VoxSeT can manage voxelized point clusters with arbitrary size in a wide range, and process them in parallel with linear complexity. The proposed VoxSeT integrates the high performance of transformer with the efficiency of voxel-based model, which can be used as a good alternative to the convolutional and point-based backbones. VoxSeT reports competitive results on the KITTI and Waymo detection benchmarks. The source codes can be found at https://github.com/skyhehe123/VoxSeT. Chenhang He, Ruihuang Li, Shuai Li 0014, Lei Zhang 0006 |
CVPR | 4 |
| 2022 | A Dual Weighting Label Assignment Scheme for Object DetectionabstractLabel assignment (LA), which aims to assign each training sample a positive (pos) and a negative (neg) loss weight, plays an important role in object detection. Existing LA methods mostly focus on the design of pos weighting function, while the neg weight is directly derived from the pos weight. Such a mechanism limits the learning capacity of detectors. In this paper, we explore a new weighting paradigm, termed dual weighting (DW), to specify pos and neg weights separately. We first identify the key influential factors of pos/neg weights by analyzing the evaluation metrics in object detection, and then design the pos and neg weighting functions based on them. Specifically, the pos weight of a sample is determined by the consistency degree between its classification and localization scores, while the neg weight is decomposed into two terms: the probability that it is a neg sample and its importance conditioned on being a neg sample. Such a weighting strategy offers greater flexibility to distinguish between important and less important samples, resulting in a more effective object detector. Equipped with the proposed DW method, a single FCOS-ResNet-50 detector can reach 41.5% mAP on COCO under 1× schedule, outperforming other existing LA methods. It consistently improves the baselines on COCO by a large margin under various backbones without bells and whistles. Code is available at https://github.com/strongwolf/DW. Shuai Li 0014, Chenhang He, Ruihuang Li, Lei Zhang 0006 |
CVPR | 4 |
| 2022 | Class-Balanced Pixel-Level Self-Labeling for Domain Adaptive Semantic SegmentationabstractDomain adaptive semantic segmentation aims to learn a model with the supervision of source domain data, and produce satisfactory dense predictions on unlabeled target domain. One popular solution to this challenging task is self-training, which selects high-scoring predictions on target samples as pseudo labels for training. However, the produced pseudo labels often contain much noise because the model is biased to source domain as well as majority categories. To address the above issues, we propose to di-rectly explore the intrinsic pixel distributions of target do-main data, instead of heavily relying on the source domain. Specifically, we simultaneously cluster pixels and rectify pseudo labels with the obtained cluster assignments. This process is done in an online fashion so that pseudo labels could co-evolve with the segmentation model without extra training rounds. To overcome the class imbalance problem on long-tailed categories, we employ a distribution align-ment technique to enforce the marginal class distribution of cluster assignments to be close to that of pseudo labels. The proposed method, namely Class-balanced Pixel-level Self-Labeling (CPSL), improves the segmentation performance on target domain over state-of-the-arts by a large margin, especially on long-tailed categories. The source code is available at ht tps: / / gi thub. com/lslrh/CPSL. Ruihuang Li, Shuai Li 0014, Chenhang He, Yabin Zhang 0001, Xu Jia 0012, Lei Zhang 0006 |
CVPR | 6 |
| 2022 | Details or Artifacts: A Locally Discriminative Learning Approach to Realistic Image Super-ResolutionabstractSingle image super-resolution (SISR) with generative adversarial networks (GAN) has recently attracted increasing attention due to its potentials to generate rich details. However, the training of GAN is unstable, and it often introduces many perceptually unpleasant artifacts along with the generated details. In this paper, we demonstrate that it is possible to train a GAN-based SISR model which can stably generate perceptually realistic details while inhibiting visual artifacts. Based on the observation that the local statistics (e.g., residual variance) of artifact areas are often different from the areas of perceptually friendly details, we develop a framework to discriminate between GAN-generated artifacts and realistic details, and consequently generate an artifact map to regularize and stabilize the model training process. Our proposed locally discriminative learning (LDL) method is simple yet effective, which can be easily plugged in off-the-shelf SISR methods and boost their performance. Experiments demonstrate that LDL outperforms the state-of-the-art GAN based SISR methods, achieving not only higher reconstruction accuracy but also superior perceptual quality on both synthetic and real-world datasets. Codes and models are available at https://github.com/csjliang/LDL. Jie Liang 0007, Hui Zeng 0001, Lei Zhang 0006 |
CVPR | 3 |
| 2022 | A Text Attention Network for Spatial Deformation Robust Scene Text Image Super-resolutionabstractScene text image super-resolution aims to increase the resolution and readability of the text in low-resolution images. Though significant improvement has been achieved by deep convolutional neural networks (CNNs), it remains difficult to reconstruct high-resolution images for spatially deformed texts, especially rotated and curve-shaped ones. This is because the current CNN-based methods adopt locality-based operations, which are not effective to deal with the variation caused by deformations. In this paper, we propose a CNN based Text ATTention network (TATT) to address this problem. The semantics of the text are firstly extracted by a text recognition module as text prior information. Then we design a novel transformer-based module, which leverages global attention mechanism, to exert the semantic guidance of text prior to the text reconstruction process. In addition, we propose a text structure consistency loss to refine the visual appearance by imposing structural consistency on the reconstructions of regular and deformed texts. Experiments on the benchmark TextZoom dataset show that the proposed TATT not only achieves state-of-the-art performance in terms of PSNR/SSIM metrics, but also significantly improves the recognition accuracy in the downstream text recognition task, particularly for text instances with multi-orientation and curved shapes. Code is available at https://github.com/mjq11302010044/TATT. Jianqi Ma, Zhetong Liang, Lei Zhang 0006 |
CVPR | 3 |
| 2022 | Blind Image Super-resolution with Elaborate Degradation Modeling on Noise and KernelabstractWhile researches on model-based blind single image super-resolution (SISR) have achieved tremendous successes recently, most of them do not consider the image degradation sufficiently. Firstly, they always assume image noise obeys an independent and identically distributed (i.i.d.) Gaussian or Laplacian distribution, which largely underestimates the complexity of real noise. Secondly, previous commonly-used kernel priors (e.g., normalization, sparsity) are not effective enough to guarantee a rational kernel solution, and thus degenerates the performance of subsequent SISR task. To address the above issues, this paper proposes a model-based blind SISR method under the probabilistic framework, which elaborately models image degradation from the perspectives of noise and blur kernel. Specifically, instead of the traditional i.i.d. noise assumption, a patch-based non-i.i.d. noise model is proposed to tackle the complicated real noise, expecting to increase the degrees of freedom of the model for noise representation. As for the blur kernel, we novelly construct a concise yet effective kernel generator, and plug it into the proposed blind SISR method as an explicit kernel prior (EKP). To solve the proposed model, a theoretically grounded Monte Carlo EM algorithm is specifically designed. Comprehensive experiments demonstrate the superiority of our method over current state-of-the-arts on synthetic and real datasets. The source code is available at https://github.com/zsyOAOA/BSRDM. Zongsheng Yue, Qian Zhao 0002, Jianwen Xie, Lei Zhang 0006, Deyu Meng, Kwan-Yee Kenneth Wong |
CVPR | 4 |
| 2022 | Exact Feature Distribution Matching for Arbitrary Style Transfer and Domain GeneralizationabstractArbitrary style transfer (AST) and domain generalization (DG) are important yet challenging visual learning tasks, which can be cast as a feature distribution matching problem. With the assumption of Gaussian feature distribution, conventional feature distribution matching methods usually match the mean and standard deviation of features. However, the feature distributions of real-world data are usually much more complicated than Gaussian, which cannot be accurately matched by using only the first-order and second-order statistics, while it is computationally prohibitive to use high-order statistics for distribution matching. In this work, we, for the first time to our best knowledge, propose to perform Exact Feature Distribution Matching (EFDM) by exactly matching the empirical Cumulative Distribution Functions (eCDFs) of image features, which could be implemented by applying the Exact Histogram Matching (EHM) in the image feature space. Particularly, a fast EHM algorithm, named Sort-Matching, is employed to perform EFDM in a plug-and-play manner with minimal cost. The effectiveness of our proposed EFDM method is verified on a variety of AST and DG tasks, demonstrating new state-of-the-art results. Codes are available at https://github.com/YBZh/EFDM. Yabin Zhang 0001, Minghan Li 0001, Ruihuang Li, Kui Jia, Lei Zhang 0006 |
CVPR | 5 |
| 2022 | From Face to Natural Image: Learning Real Degradation for Blind Image Super-Resolution
Xiaoming Li 0002, Chaofeng Chen, Xianhui Lin, Wangmeng Zuo, Lei Zhang 0006 |
ECCV (18) | 5 |
| 2022 | Box-Supervised Instance Segmentation with Level Set Evolution
Wentong Li 0001, Wenyu Liu 0001, Jianke Zhu, Miaomiao Cui, Xian-Sheng Hua 0001, Lei Zhang 0006 |
ECCV (29) | 6 |
| 2022 | Efficient and Degradation-Adaptive Network for Real-World Image Super-Resolution
Jie Liang 0007, Hui Zeng 0001, Lei Zhang 0006 |
ECCV (18) | 3 |
| 2022 | Spatiotemporal Self-attention Modeling with Temporal Patch Shift for Action Recognition
Wangmeng Xiang, Chao Li 0034, Xihan Wei, Xian-Sheng Hua 0001, Lei Zhang 0006 |
ECCV (3) | 6 |
| 2022 | An Embedded Feature Whitening Approach to Deep Neural Network Optimization
Hongwei Yong, Lei Zhang 0006 |
ECCV (23) | 2 |
| 2022 | Efficient Long-Range Attention Network for Image Super-Resolution
Hui Zeng 0001, Shi Guo, Lei Zhang 0006 |
ECCV (17) | 4 |
| 2022 | Unfolded Deep Kernel Estimation for Blind Image Super-Resolution
Hongyi Zheng, Hongwei Yong, Lei Zhang 0006 |
ECCV (18) | 3 |
| 2022 | A novel fractional time-delayed grey Bernoulli forecasting model and its application for the energy production and consumption prediction
Yong Wang 0032, Xinbo He, Lei Zhang 0006, Xin Ma 0004, Wenqing Wu 0001, Pei Chi |
Eng. Appl. Artif. Intell. | 3 |
| 2022 | A novel self-adaptive fractional multivariable grey model and its application in forecasting energy production and conversion of China
Yong Wang 0032, Lingling Ye, Xin Ma 0004, Wenqing Wu 0001, Zhongsen Yang, Xinbo He, Lei Zhang 0006, Yongxian Luo |
Eng. Appl. Artif. Intell. | 8 |
| 2022 | A novel fractional structural adaptive grey Chebyshev polynomial Bernoulli model and its application in forecasting renewable energy production of China
Yong Wang 0032, Pei Chi, Xin Ma 0004, Wenqing Wu 0001, Binhong Guo, Xinbo He, Lei Zhang 0006 |
Expert Syst. Appl. | 8 |
| 2022 | A novel structure adaptive fractional discrete grey forecasting model and its application in China's crude oil production prediction
Yong Wang 0032, Lingling Ye, Zhongsen Yang, Xin Ma 0004, Wenqing Wu 0001, Xinbo He, Lei Zhang 0006, Yongxian Luo |
Expert Syst. Appl. | 8 |
| 2022 | Mutual consistency learning for semi-supervised medical image segmentation
Yicheng Wu 0001, ZongYuan Ge, Donghao Zhang 0004, Minfeng Xu, Lei Zhang 0006, Yong Xia 0001, Jianfei Cai 0001 |
Medical Image Anal. | 5 |
| 2022 | Learning Image-Adaptive 3D Lookup Tables for High Performance Photo Enhancement in Real-TimeabstractRecent years have witnessed the increasing popularity of learning based methods to enhance the color and tone of photos. However, many existing photo enhancement methods either deliver unsatisfactory results or consume too much computational and memory resources, hindering their application to high-resolution images (usually with more than 12 megapixels) in practice. In this paper, we learn image-adaptive 3-dimensional lookup tables (3D LUTs) to achieve fast and robust photo enhancement. 3D LUTs are widely used for manipulating color and tone of photos, but they are usually manually tuned and fixed in camera imaging pipeline or photo editing tools. We, for the first time to our best knowledge, propose to learn 3D LUTs from annotated data using pairwise or unpaired learning. More importantly, our learned 3D LUT is image-adaptive for flexible photo enhancement. We learn multiple basis 3D LUTs and a small convolutional neural network (CNN) simultaneously in an end-to-end manner. The small CNN works on the down-sampled version of the input image to predict content-dependent weights to fuse the multiple basis 3D LUTs into an image-adaptive one, which is employed to transform the color and tone of source images efficiently. Our model contains less than 600K parameters and takes less than 2 ms to process an image of 4K resolution using one Titan RTX GPU. While being highly efficient, our model also outperforms the state-of-the-art photo enhancement methods by a large margin in terms of PSNR, SSIM and a color difference metric on two publically available benchmark datasets. Code will be released at https://github.com/HuiZeng/Image-Adaptive-3DLUT. Hui Zeng 0001, Jianrui Cai, Lida Li, Zisheng Cao, Lei Zhang 0006 |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2022 | Grid Anchor Based Image Cropping: A New Benchmark and An Efficient ModelabstractImage cropping aims to improve the composition as well as aesthetic quality of an image by removing extraneous content from it. Most of the existing image cropping databases provide only one or several human-annotated bounding boxes as the groundtruths, which can hardly reflect the non-uniqueness and flexibility of image cropping in practice. The employed evaluation metrics such as intersection-over-union cannot reliably reflect the real performance of a cropping model, either. This work revisits the problem of image cropping, and presents a grid anchor based formulation by considering the special properties and requirements (e.g., local redundancy, content preservation, aspect ratio) of image cropping. Our formulation reduces the searching space of candidate crops from millions to no more than ninety. Consequently, a grid anchor based cropping benchmark is constructed, where all crops of each image are annotated and more reliable evaluation metrics are defined. To meet the practical demands of robust performance and high efficiency, we also design an effective and lightweight cropping model. By simultaneously considering the region of interest and region of discard, and leveraging multi-scale information, our model can robustly output visually pleasing crops for images of different scenes. With less than 2.5M parameters, our model runs at a speed of 200 FPS on one single GTX 1080Ti GPU and 12 FPS on one i7-6800K CPU. The code is available at: https://github.com/HuiZeng/Grid-Anchor-based-Image-Cropping-Pytorch. Hui Zeng 0001, Lida Li, Zisheng Cao, Lei Zhang 0006 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2022 | Unsupervised Multi-Class Domain Adaptation: Theory, Algorithms, and PracticeabstractIn this paper, we study the formalism of unsupervised multi-class domain adaptation (multi-class UDA), which underlies a few recent algorithms whose learning objectives are only motivated empirically. Multi-Class Scoring Disagreement (MCSD) divergence is presented by aggregating the absolute margin violations in multi-class classification, and this proposed MCSD is able to fully characterize the relations between any pair of multi-class scoring hypotheses. By using MCSD as a measure of domain distance, we develop a new domain adaptation bound for multi-class UDA; its data-dependent, probably approximately correct bound is also developed that naturally suggests adversarial learning objectives to align conditional feature distributions across source and target domains. Consequently, an algorithmic framework of Multi-class Domain-adversarial learning Networks (McDalNets) is developed, and its different instantiations via surrogate learning objectives either coincide with or resemble a few recently popular methods, thus (partially) underscoring their practical effectiveness. Based on our identical theory for multi-class UDA, we also introduce a new algorithm of Domain-Symmetric Networks (SymmNets), which is featured by a novel adversarial strategy of domain confusion and discrimination. SymmNets affords simple extensions that work equally well under the problem settings of either closed set, partial, or open set UDA. We conduct careful empirical studies to compare different algorithms of McDalNets and our newly introduced SymmNets. Experiments verify our theoretical analysis and show the efficacy of our proposed SymmNets. In addition, we have made our implementation code publicly available. Yabin Zhang 0001, Bin Deng 0003, Lei Zhang 0006, Kui Jia |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2022 | Plug-and-Play Image Restoration With Deep Denoiser PriorabstractRecent works on plug-and-play image restoration have shown that a denoiser can implicitly serve as the image prior for model-based methods to solve many inverse problems. Such a property induces considerable advantages for plug-and-play image restoration (e.g., integrating the flexibility of model-based method and effectiveness of learning-based methods) when the denoiser is discriminatively learned via deep convolutional neural network (CNN) with large modeling capacity. However, while deeper and larger CNN models are rapidly gaining popularity, existing plug-and-play image restoration hinders its performance due to the lack of suitable denoiser prior. In order to push the limits of plug-and-play image restoration, we set up a benchmark deep denoiser prior by training a highly flexible and effective CNN denoiser. We then plug the deep denoiser prior as a modular part into a half quadratic splitting based iterative algorithm to solve various image restoration problems. We, meanwhile, provide a thorough analysis of parameter setting, intermediate results and empirical convergence to better understand the working mechanism. Experimental results on three representative image restoration tasks, including deblurring, super-resolution and demosaicing, demonstrate that the proposed plug-and-play image restoration with deep denoiser prior not only significantly outperforms other state-of-the-art model-based methods but also achieves competitive or even superior performance against state-of-the-art learning-based methods. The source code is available at https://github.com/cszn/DPIR. Kai Zhang 0008, Yawei Li 0001, Wangmeng Zuo, Lei Zhang 0006, Luc Van Gool, Radu Timofte |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2022 | Detachable Second-Order Pooling: Toward High-Performance First-Order NetworksabstractSecond-order pooling has proved to be more effective than its first-order counterpart in visual classification tasks. However, second-order pooling suffers from the high demand for a computational resource, limiting its use in practical applications. In this work, we present a novel architecture, namely a detachable second-order pooling network, to leverage the advantage of second-order pooling by first-order networks while keeping the model complexity unchanged during inference. Specifically, we introduce second-order pooling at the end of a few auxiliary branches and plug them into different stages of a convolutional neural network. During the training stage, the auxiliary second-order pooling networks assist the backbone first-order network to learn more discriminative feature representations. When training is completed, all auxiliary branches can be removed, and only the backbone first-order network is used for inference. Experiments conducted on CIFAR-10, CIFAR-100, and ImageNet data sets clearly demonstrated the leading performance of our network, which achieves even higher accuracy than second-order networks but keeps the low inference complexity of first-order networks. Lida Li, Jiangtao Xie, Peihua Li, Lei Zhang 0006 |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2022 | Directional Deep Embedding and Appearance Learning for Fast Video Object SegmentationabstractMost recent semisupervised video object segmentation (VOS) methods rely on fine-tuning deep convolutional neural networks online using the given mask of the first frame or predicted masks of subsequent frames. However, the online fine-tuning process is usually time-consuming, limiting the practical use of such methods. We propose a directional deep embedding and appearance learning (DDEAL) method, which is free of the online fine-tuning process, for fast VOS. First, a global directional matching module (GDMM), which can be efficiently implemented by parallel convolutional operations, is proposed to learn a semantic pixel-wise embedding as an internal guidance. Second, an effective directional appearance model-based statistics is proposed to represent the target and background on a spherical embedding space for VOS. Equipped with the GDMM and the directional appearance model learning module, DDEAL learns static cues from the labeled first frame and dynamically updates cues of the subsequent frames for object segmentation. Our method exhibits the state-of-the-art VOS performance without using online fine-tuning. Specifically, it achieves a J & F mean score of 74.8% on DAVIS 2017 data set and an overall score G of 71.3% on the large-scale YouTube-VOS data set, while retaining a speed of 25 fps with a single NVIDIA TITAN Xp GPU. Furthermore, our faster version runs 31 fps with only a little accuracy loss. Yingjie Yin, De Xu, Xingang Wang 0003, Lei Zhang 0006 |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2022 | Robust 3D reconstruction from uncalibrated small motion clips
Zhaoxin Li, Wangmeng Zuo, Lei Zhang 0006 |
Vis. Comput. | 4 |
| 2021 | Deep Metric Learning with Graph ConsistencyabstractDeep Metric Learning (DML) has been more attractive and widely applied in many computer vision tasks, in which a discriminative embedding is requested such that the image features belonging to the same class are gathered together and the ones belonging to different classes are pushed apart. Most existing works insist to learn this discriminative embedding by either devising powerful pair-based loss functions or hard-sample mining strategies. However, in this paper, we start from an another perspective and propose Deep Consistent Graph Metric Learning (CGML) framework to enhance the discrimination of the learned embedding. It is mainly achieved by rethinking the conventional distance constraints as a graph regularization and then introducing a Graph Consistency regularization term, which intends to optimize the feature distribution from a global graph perspective. Inspired by the characteristic of our defined ’Discriminative Graph’, which regards DML from another novel perspective, the Graph Consistency regularization term encourages the sub-graphs randomly sampled from the training set to be consistent. We show that our CGML indeed serves as an efficient technique for learning towards discriminative embedding and is applicable to various popular metric objectives, e.g. Triplet, N-Pair and Binomial losses. This paper empirically and experimentally demonstrates the effectiveness of our graph regularization idea, achieving competitive results on the popular CUB, CARS, Stanford Online Products and In-Shop datasets. Binghui Chen, Zhaoyi Yan, Lei Zhang 0006 |
AAAI | 5 |
| 2021 | Category Dictionary Guided Unsupervised Domain Adaptation for Object DetectionabstractUnsupervised domain adaption (UDA) is a promising solution to enhance the generalization ability of a model from a source domain to a target domain without manually annotating labels for target data. Recent works in cross-domain object detection mostly resort to adversarial feature adaptation to match the marginal distributions of two domains. However, perfect feature alignment is hard to achieve and is likely to cause negative transfer due to the high complexity of object detection. In this paper, we propose a category dictionary guided (CDG) UDA model for cross-domain object detection, which learns category-specific dictionaries from the source domain to represent the candidate boxes in target domain. The representation residual can be used for not only pseudo label assignment but also quality (e.g., IoU) estimation of the candidate box. A residual weighted self-training paradigm is then developed to implicitly align source and target domains for detection model training. Compared with decision boundary based classifiers such as softmax, the proposed CDG scheme can select more informative and reliable pseudo-boxes. Experimental results on benchmark datasets show that the proposed CDG significantly exceeds the state-of-the-arts in cross-domain object detection. Shuai Li 0014, Jianqiang Huang 0001, Xian-Sheng Hua 0001, Lei Zhang 0006 |
AAAI | 4 |
| 2021 | Adversarial Pose Regression Network for Pose-Invariant Face Recognitions
Lei Zhang 0006 |
AAAI | 3 |
| 2021 | Unsupervised Domain Adaptation of Black-Box Source Models
Haojian Zhang, Yabin Zhang 0001, Kui Jia, Lei Zhang 0006 |
BMVC | 4 |
| 2021 | Progressive Semantic-Aware Style Transformation for Blind Face RestorationabstractFace restoration is important in face image processing, and has been widely studied in recent years. However, previous works often fail to generate plausible high quality (HQ) results for real-world low quality (LQ) face images. In this paper, we propose a new progressive semantic-aware style transformation framework, named PSFR-GAN, for face restoration. Specifically, instead of using an encoder-decoder framework as previous methods, we formulate the restoration of LQ face images as a multi-scale progressive restoration procedure through semantic-aware style transformation. Given a pair of LQ face image and its corresponding parsing map, we first generate a multi-scale pyramid of the inputs, and then progressively modulate different scale features from coarse-to-fine in a semantic-aware style transfer way. Compared with previous networks, the proposed PSFR-GAN makes full use of the semantic (parsing maps) and pixel (LQ images) space information from different scales of input pairs. In addition, we further introduce a semantic aware style loss which calculates the feature style loss for each semantic region individually to improve the details of face textures. Finally, we pretrain a face parsing network which can generate decent parsing maps from real-world LQ face images. Experiment results show that our model trained with synthetic data can not only produce more realistic high-resolution results for synthetic LQ inputs but also generalize better to natural LQ face images compared with state-of-the-art methods. Chaofeng Chen, Xiaoming Li 0002, Lingbo Yang, Xianhui Lin, Lei Zhang 0006, Kwan-Yee Kenneth Wong |
CVPR | 5 |
| 2021 | VirFace: Enhancing Face Recognition via Unlabeled Shallow DataabstractRecently, how to exploit unlabeled data for training face recognition models has been attracting increasing attention. However, few works consider the unlabeled shallow data1in real-world scenarios. The existing semi-supervised face recognition methods that focus on generating pseudo labels or minimizing softmax classification probabilities of the unlabeled data do not work very well on the unlabeled shallow data. It is still a challenge on how to effectively utilize the unlabeled shallow face data to improve the performance of face recognition. In this paper, we propose a novel face recognition method, named VirFace, to effectively exploit the unlabeled shallow data for face recognition. VirFace consists of VirClass and VirInstance. Specifically, VirClass enlarges the inter-class distance by injecting the unlabeled data as new identities, while VirInstance produces virtual instances sampled from the learned distribution of each identity to further enlarge the inter-class distance. To the best of our knowledge, we are the first to tackle the problem of unlabeled shallow face data. Extensive experiments have been conducted on both the small- and large-scale datasets, e.g. LFW and IJB-C, etc, demonstrating the superiority of the proposed method. Tianchu Guo, Binghui Chen, Wangmeng Zuo, Lei Zhang 0006 |
CVPR | 7 |
| 2021 | Spatial Feature Calibration and Temporal Fusion for Effective One-Stage Video Instance SegmentationabstractModern one-stage video instance segmentation networks suffer from two limitations. First, convolutional features are neither aligned with anchor boxes nor with ground-truth bounding boxes, reducing the mask sensitivity to spatial location. Second, a video is directly divided into individual frames for frame-level instance segmentation, ignoring the temporal correlation between adjacent frames. To address these issues, we propose a simple yet effective one-stage video instance segmentation framework by spatial calibration and temporal fusion, namely STMask. To ensure spatial feature calibration with ground-truth bounding boxes, we first predict regressed bounding boxes around ground-truth bounding boxes, and extract features from them for frame-level instance segmentation. To further explore temporal correlation among video frames, we aggregate a temporal fusion module to infer instance masks from each frame to its adjacent frames, which helps our frame-work to handle challenging videos such as motion blur, partial occlusion and unusual object-to-camera poses. Experiments on the YouTube-VIS valid set show that the proposed STMask with ResNet-50/-101 backbone obtains 33.5 % / 36.8 % mask AP, while achieving 28.6 / 23.4 FPS on video instance segmentation. The code is released online https://github.com/MinghanLi/STMask. Minghan Li 0001, Shuai Li 0014, Lida Li, Lei Zhang 0006 |
CVPR | 4 |
| 2021 | Virtual Fully-Connected Layer: Training a Large-Scale Face Recognition Dataset With Limited Computational ResourcesabstractRecently, deep face recognition has achieved significant progress because of Convolutional Neural Networks (CNNs) and large-scale datasets. However, training CNNs on a large-scale face recognition dataset with limited computational resources is still a challenge. This is because the classification paradigm needs to train a fully-connected layer as the category classifier, and its parameters will be in the hundreds of millions if the training dataset contains millions of identities. This requires many computational resources, such as GPU memory. The metric learning paradigm is an economical computation method, but its performance is greatly inferior to that of the classification paradigm. To address this challenge, we propose a simple but effective CNN layer called the Virtual fully-connected (Virtual FC) layer to reduce the computational consumption of the classification paradigm. Without bells and whistles, the proposed Virtual FC reduces the parameters by more than 100 times with respect to the fully-connected layer and achieves competitive performance on mainstream face recognition evaluation datasets. Moreover, the performance of our Virtual FC layer on the evaluation datasets is superior to that of the metric learning paradigm by a significant margin. Our code will be released in hopes of disseminating our idea to other domains1. Lei Zhang 0006 |
CVPR | 3 |
| 2021 | High-Resolution Photorealistic Image Translation in Real-Time: A Laplacian Pyramid Translation NetworkabstractExisting image-to-image translation (I2IT) methods are either constrained to low-resolution images or long inference time due to their heavy computational burden on the convolution of high-resolution feature maps. In this paper, we focus on speeding-up the high-resolution photorealistic I2IT tasks based on closed-form Laplacian pyramid decomposition and reconstruction. Specifically, we reveal that the attribute transformations, such as illumination and color manipulation, relate more to the low-frequency component, while the content details can be adaptively refined on high-frequency components. We consequently propose a Laplacian Pyramid Translation Network (LPTN) to simultaneously perform these two tasks, where we design a lightweight network for translating the low-frequency component with reduced resolution and a progressive masking strategy to efficiently refine the high-frequency ones. Our model avoids most of the heavy computation consumed by processing high-resolution feature maps and faithfully preserves the image details. Extensive experimental results on various tasks demonstrate that the proposed method can translate 4K images in real-time using one normal GPU while achieving comparable transformation performance against existing methods. Datasets and codes are available: https://github.com/csjliang/LPTN. Jie Liang 0007, Hui Zeng 0001, Lei Zhang 0006 |
CVPR | 3 |
| 2021 | PPR10K: A Large-Scale Portrait Photo Retouching Dataset With Human-Region Mask and Group-Level ConsistencyabstractDifferent from general photo retouching tasks, portrait photo retouching (PPR), which aims to enhance the visual quality of a collection of flat-looking portrait photos, has its special and practical requirements such as human-region priority (HRP) and group-level consistency (GLC). HRP requires that more attention should be paid to human regions, while GLC requires that a group of portrait photos should be retouched to a consistent tone. Models trained on existing general photo retouching datasets, however, can hardly meet these requirements of PPR. To facilitate the research on this high-frequency task, we construct a largescale PPR dataset, namely PPR10K, which is the first of its kind to our best knowledge. PPR10K contains 1, 681 groups and 11, 161 high-quality raw portrait photos in total. High-resolution segmentation masks of human regions are provided. Each raw photo is retouched by three experts, while they elaborately adjust each group of photos to have consistent tones. We define a set of objective measures to evaluate the performance of PPR and propose strategies to learn PPR models with good HRP and GLC performance. The constructed PPR10K dataset provides a good bench-mark for studying automatic PPR methods, and experiments demonstrate that the proposed learning strategies are effective to improve the retouching performance. Datasets and codes are available: https://github.com/csjliang/PPR10K. Jie Liang 0007, Hui Zeng 0001, Miaomiao Cui, Xuansong Xie, Lei Zhang 0006 |
CVPR | 5 |
| 2021 | Learning Parallel Dense Correspondence From Spatio-Temporal Descriptors for Efficient and Robust 4D ReconstructionabstractThis paper focuses on the task of 4D shape reconstruction from a sequence of point clouds. Despite the recent success achieved by extending deep implicit representations into 4D space [29], it is still a great challenge in two respects, i.e. how to design a flexible framework for learning robust spatio-temporal shape representations from 4D point clouds, and develop an efficient mechanism for capturing shape dynamics. In this work, we present a novel pipeline to learn a temporal evolution of the 3D human shape through spatially continuous transformation functions among cross-frame occupancy fields. The key idea is to parallelly establish the dense correspondence between predicted occupancy fields at different time steps via explicitly learning continuous displacement vector fields from robust spatio-temporal shape representations. Extensive comparisons against previous state-of-the-arts show the superior accuracy of our approach for 4D human reconstruction in the problems of 4D shape auto-encoding and completion, and a much faster network inference with about 8 times speedup demonstrates the significant efficiency of our approach. The trained models and implementation code are available at https://github.com/tangjiapeng/LPDC-Net. Jiapeng Tang, Dan Xu 0002, Kui Jia, Lei Zhang 0006 |
CVPR | 4 |
| 2021 | GAN Prior Embedded Network for Blind Face Restoration in the WildabstractBlind face restoration (BFR) from severely degraded face images in the wild is a very challenging problem. Due to the high illness of the problem and the complex unknown degradation, directly training a deep neural network (DNN) usually cannot lead to acceptable results. Existing generative adversarial network (GAN) based methods can produce better results but tend to generate over-smoothed restorations. In this work, we propose a new method by first learning a GAN for high-quality face image generation and embedding it into a U-shaped DNN as a prior decoder, then fine-tuning the GAN prior embedded DNN with a set of synthesized low-quality face images. The GAN blocks are designed to ensure that the latent code and noise input to the GAN can be respectively generated from the deep and shallow features of the DNN, controlling the global face structure, local face details and background of the reconstructed image. The proposed GAN prior embedded network (GPEN) is easy-to-implement, and it can generate visually photo-realistic results. Our experiments demonstrated that the proposed GPEN achieves significantly superior results to state-of-the-art BFR methods both quantitatively and qualitatively, especially for the restoration of severely degraded face images in the wild. The source code and models can be found at https://github.com/yangxy/GPEN. Tao Yang 0042, Peiran Ren, Xuansong Xie, Lei Zhang 0006 |
CVPR | 4 |
| 2021 | Interactive Self-Training With Mean Teachers for Semi-Supervised Object DetectionabstractThe goal of semi-supervised object detection is to learn a detection model using only a few labeled data and large amounts of unlabeled data, thereby reducing the cost of data labeling. Although a few studies have proposed various self-training-based methods or consistency regularization-based methods, they ignore the discrepancies among the detection results in the same image that occur during different training iterations. Additionally, the predicted detection results vary among different detection models. In this paper, we propose an interactive form of self-training using mean teachers for semi-supervised object detection. Specifically, to alleviate the instability among the detection results in different iterations, we propose using nonmaximum suppression to fuse the detection results from different iterations. Simultaneously, we use multiple detection heads that predict pseudo labels for each other to provide complementary information. Furthermore, to avoid different detection heads collapsing to each other, we use a mean teacher model instead of the original detection model to predict the pseudo labels. Thus, the object detection model can be trained on both labeled and unlabeled data. Extensive experimental results verify the effectiveness of our proposed method. Qize Yang, Xihan Wei, Xian-Sheng Hua 0001, Lei Zhang 0006 |
CVPR | 5 |
| 2021 | Deep Convolutional Dictionary Learning for Image DenoisingabstractInspired by the great success of deep neural networks (DNNs), many unfolding methods have been proposed to integrate traditional image modeling techniques, such as dictionary learning (DicL) and sparse coding, into DNNs for image restoration. However, the performance of such methods remains limited for several reasons. First, the unfolded architectures do not strictly follow the image representation model of DicL and lose the desired physical meaning. Second, handcrafted priors are still used in most unfolding methods without effectively utilizing the learning capability of DNNs. Third, a universal dictionary is learned to represent all images, reducing the model representation flexibility. We propose a novel framework of deep convolutional dictionary learning (DCDicL), which follows the representation model of DicL strictly, learns the priors for both representation coefficients and the dictionaries, and can adaptively adjust the dictionary for each input image based on its content. The effectiveness of our DCDicL method is validated on the image denoising problem. DCDicL demonstrates leading denoising performance in terms of both quantitative metrics (e.g., PSNR, SSIM) and visual quality. In particular, it can reproduce the subtle image structures and textures, which are hard to recover by many existing denoising DNNs. The code is available at: https://github.com/natezhenghy/DCDicL_denoising. Hongyi Zheng, Hongwei Yong, Lei Zhang 0006 |
CVPR | 3 |
| 2021 | Clustering-Augmented Multi-instance Learning for Neural Relation Extraction
Qi Zhang 0001, Siliang Tang, Jinquan Sun, Yu Wang 0108, Lei Zhang 0006 |
ECIR (2) | 5 |
| 2021 | Correlated Randomness Teleportation via Semi-trusted Hardware - Enabling Silent Multi-party Computation
Yibiao Lu, Bingsheng Zhang, Hong-Sheng Zhou, Lei Zhang 0006, Kui Ren 0001 |
ESORICS (2) | 5 |
| 2021 | Label-Aware Text Representation for Multi-Label Text ClassificationabstractMulti-label text classification (MLTC) is an important task in natural language processing (NLP), which is appealing to researchers in both academia and industry. However, few of studies have been conducted on the relations among the labels. Most of existing methods tend to neglect the semantic information between labels and words. In this paper, we propose a label-aware network to obtain both the label correlation and text representation. A heterogeneous graph is built from words and labels to learn the label representation by metap-ath2vec, since two nearby labels or words in the graph have similar relation and the graph structure is beneficial for label representation as well. Each part of the text contributes differently to label inference, therefore bidirectional attention flow is exploited for label-aware text representation in two directions: from text to label and from label to text. Experimental evaluations illustrate that the proposed method outperforms various baselines on both offline benchmarks and real-world online systems. Lei Zhang 0006 |
ICASSP | 3 |
| 2021 | HDR Video Reconstruction: A Coarse-to-fine Network and A Real-world Benchmark DatasetabstractHigh dynamic range (HDR) video reconstruction from sequences captured with alternating exposures is a very challenging problem. Existing methods often align low dynamic range (LDR) input sequence in the image space using optical flow, and then merge the aligned images to produce HDR output. However, accurate alignment and fusion in the image space are difficult due to the missing details in the over-exposed regions and noise in the under-exposed regions, resulting in unpleasing ghosting artifacts. To enable more accurate alignment and HDR fusion, we introduce a coarse-to-fine deep learning framework for HDR video reconstruction. Firstly, we perform coarse alignment and pixel blending in the image space to estimate the coarse HDR video. Secondly, we conduct more sophisticated alignment and temporal fusion in the feature space of the coarse HDR video to produce better reconstruction. Considering the fact that there is no publicly available dataset for quantitative and comprehensive evaluation of HDR video reconstruction methods, we collect such a benchmark dataset, which contains 97 sequences of static scenes and 184 testing pairs of dynamic scenes. Extensive experiments show that our method outperforms previous state-of-the-art methods. Our code and dataset can be found at https://guanyingc.github.io/DeepHDRVideo. Guanying Chen, Chaofeng Chen, Shi Guo, Zhetong Liang, Kwan-Yee Kenneth Wong, Lei Zhang 0006 |
ICCV | 6 |
| 2021 | Variational Attention: Propagating Domain-Specific Knowledge for Multi-Domain Learning in Crowd CountingabstractIn crowd counting, due to the problem of laborious labelling, it is perceived intractability of collecting a new large-scale dataset which has plentiful images with large diversity in density, scene, etc. Thus, for learning a general model, training with data from multiple different datasets might be a remedy and be of great value. In this paper, we resort to the multi-domain joint learning and propose a simple but effective Domain-specific Knowledge Propagating Network (DKPNet) for unbiasedly learning the knowledge from multiple diverse data domains at the same time. It is mainly achieved by proposing the novel Variational Attention(VA) technique for explicitly modeling the attention distributions for different domains. And as an extension to VA, Intrinsic Variational Attention(InVA) is proposed to handle the problems of over-lapped domains and sub-domains. Extensive experiments have been conducted to validate the superiority of our DKPNet over several popular datasets, including ShanghaiTech A/B, UCF-QNRF and NWPU. Binghui Chen, Zhaoyi Yan, Ke Li 0004, Wangmeng Zuo, Lei Zhang 0006 |
ICCV | 7 |
| 2021 | SA-ConvONet: Sign-Agnostic Optimization of Convolutional Occupancy NetworksabstractSurface reconstruction from point clouds is a fundamental problem in the computer vision and graphics community. Recent state-of-the-arts solve this problem by individually optimizing each local implicit field during inference. Without considering the geometric relationships between local fields, they typically require accurate normals to avoid the sign conflict problem in overlapped regions of local fields, which severely limits their applicability to raw scans where surface normals could be unavailable. Although SAL breaks this limitation via sign-agnostic learning, further works still need to explore how to extend this technique for local shape modeling. To this end, we propose to learn implicit surface reconstruction by sign-agnostic optimization of convolutional occupancy networks, to simultaneously achieve advanced scalability to large-scale scenes, generality to novel shapes, and applicability to raw scans in a unified framework. Concretely, we achieve this goal by a simple yet effective design, which further optimizes the pre-trained occupancy prediction networks with an unsigned cross-entropy loss during inference. The learning of occupancy fields is conditioned on convolutional features from an hourglass network architecture. Extensive experimental comparisons with previous state-of-the-arts on both object-level and scene-level datasets demonstrate the superior accuracy of our approach for surface reconstruction from un-orientated point clouds. The code is available at https://github.com/tangjiapeng/SA-ConvONet. Jiapeng Tang, Jiabao Lei, Dan Xu 0002, Feiying Ma, Kui Jia, Lei Zhang 0006 |
ICCV | 6 |
| 2021 | Real-world Video Super-resolution: A Benchmark Dataset and A Decomposition based Learning SchemeabstractVideo super-resolution (VSR) aims to improve the spatial resolution of low-resolution (LR) videos. Existing VSR methods are mostly trained and evaluated on synthetic datasets, where the LR videos are uniformly downsampled from their high-resolution (HR) counterparts by some simple operators (e.g., bicubic downsampling). Such simple synthetic degradation models, however, cannot well describe the complex degradation processes in real-world videos, and thus the trained VSR models become ineffective in real-world applications. As an attempt to bridge the gap, we build a real-world video super-resolution (RealVSR) dataset by capturing paired LR-HR video sequences using the multi-camera system of iPhone 11 Pro Max. Since the LR-HR video pairs are captured by two separate cameras, there are inevitably certain misalignment and luminance/color differences between them. To more robustly train the VSR model and recover more details from the LR inputs, we convert the LR-HR videos into YCbCr space and decompose the luminance channel into a Laplacian pyramid, and then apply different loss functions to different components. Experiments validate that VSR models trained on our RealVSR dataset demonstrate better visual quality than those trained on synthetic datasets under real-world settings. They also exhibit good generalization capability in cross-camera tests. The dataset and code can be found at https://github.com/IanYeung/RealVSR. Xi Yang 0001, Wangmeng Xiang, Hui Zeng 0001, Lei Zhang 0006 |
ICCV | 4 |
| 2021 | Semi-supervised Left Atrium Segmentation with Mutual Consistency Training
Yicheng Wu 0001, Minfeng Xu, ZongYuan Ge, Jianfei Cai 0001, Lei Zhang 0006 |
MICCAI (2) | 5 |
| 2021 | Conditional Directed Graph Convolution for 3D Human Pose EstimationabstractGraph convolutional networks have significantly improved 3D human pose estimation by representing the human skeleton as an undirected graph. However, this representation fails to reflect the articulated characteristic of human skeletons as the hierarchical orders among the joints are not explicitly presented. In this paper, we propose to represent the human skeleton as a directed graph with the joints as nodes and bones as edges that are directed from parent joints to child joints. By so doing, the directions of edges can explicitly reflect the hierarchical relationships among the nodes. Based on this representation, we further propose a spatial-temporal conditional directed graph convolution to leverage varying non-local dependence for different poses by conditioning the graph topology on input poses. Altogether, we form a U-shaped network, named U-shaped Conditional Directed Graph Convolutional Network, for 3D human pose estimation from monocular videos. To evaluate the effectiveness of our method, we conducted extensive experiments on two challenging large-scale benchmarks: Human3.6M and MPI-INF-3DHP. Both quantitative and qualitative results show that our method achieves top performance. Also, ablation studies show that directed graphs can better exploit the hierarchy of articulated human skeletons than undirected graphs, and the conditional connections can yield adaptive graph topologies for different poses. Wenbo Hu 0002, Changgong Zhang, Fangneng Zhan, Lei Zhang 0006, Tien-Tsin Wong |
ACM Multimedia | 4 |
| 2021 | Edge-oriented Convolution Block for Real-time Super Resolution on Mobile DevicesabstractEfficient and light-weight super resolution (SR) is highly demanded in practical applications. However, most of the existing studies focusing on reducing the number of model parameters and FLOPs may not necessarily lead to faster running speed on mobile devices. In this work, we propose a re-parameterizable building block, namely Edge-oriented Convolution Block (ECB), for efficient SR design. In the training stage, the ECB extracts features in multiple paths, including a normal 3 x 3 convolution, a channel expanding-and-squeezing convolution, and 1st-order and 2nd-order spatial derivatives from intermediate features. In the inference stage, the multiple operations can be merged into one single 3 3 convolution. ECB can be regarded as a drop-in replacement to improve the performance of normal 3 3 convolution without introducing any additional cost in the inference stage. We then propose an extremely efficient SR network for mobile devices based on ECB, namely ECBSR. Extensive experiments across five benchmark datasets demonstrate the effectiveness and efficiency of ECB and ECBSR. Our ECBSR achieves comparable PSNR/SSIM performance to state-of-the-art light-weight SR models, while it can super resolve images from 270p/540p to 1080p in real-time on commodity mobile devices, e.g., Snapdragon 865 SOC and Dimensity 1000+ SOC. The source code can be found at https://github.com/xindongzhang/ECBSR. Hui Zeng 0001, Lei Zhang 0006 |
ACM Multimedia | 3 |
| 2021 | Simultaneous Fidelity and Regularization Learning for Image RestorationabstractMost existing non-blind restoration methods are based on the assumption that a precise degradation model is known. As the degradation process can only be partially known or inaccurately modeled, images may not be well restored. Rain streak removal and image deconvolution with inaccurate blur kernels are two representative examples of such tasks. For rain streak removal, although an input image can be decomposed into a scene layer and a rain streak layer, there exists no explicit formulation for modeling rain streaks and the composition with scene layer. For blind deconvolution, as estimation error of blur kernel is usually introduced, the subsequent non-blind deconvolution process does not restore the latent image well. In this paper, we propose a principled algorithm within the maximum a posterior framework to tackle image restoration with a partially known or inaccurate degradation model. Specifically, the residual caused by a partially known or inaccurate degradation model is spatially dependent and complexly distributed. With a training set of degraded and ground-truth image pairs, we parameterize and learn the fidelity term for a degradation model in a task-driven manner. Furthermore, the regularization term can also be learned along with the fidelity term, thereby forming a simultaneous fidelity and regularization learning model. Extensive experimental results demonstrate the effectiveness of the proposed model for image deconvolution with inaccurate blur kernels, deconvolution with multiple degradations and rain streak removal. Dongwei Ren, Wangmeng Zuo, David Zhang 0001, Lei Zhang 0006, Ming-Hsuan Yang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2021 | Deep CNNs Meet Global Covariance Pooling: Better Representation and GeneralizationabstractCompared with global average pooling in existing deep convolutional neural networks (CNNs), global covariance pooling can capture richer statistics of deep features, having potential for improving representation and generalization abilities of deep CNNs. However, integration of global covariance pooling into deep CNNs brings two challenges: (1) robust covariance estimation given deep features of high dimension and small sample size; (2) appropriate usage of geometry of covariances. To address these challenges, we propose a global Matrix Power Normalized COVariance (MPN-COV) Pooling. Our MPN-COV conforms to a robust covariance estimator, very suitable for scenario of high dimension and small sample size. It can also be regarded as Power-Euclidean metric between covariances, effectively exploiting their geometry. Furthermore, a global Gaussian embedding network is proposed to incorporate first-order statistics into MPN-COV. For fast training of MPN-COV networks, we implement an iterative matrix square root normalization, avoiding GPU unfriendly eigen-decomposition inherent in MPN-COV. Additionally, progressive 1×1 convolutions and group convolution are introduced to compress covariance representations. The proposed methods are highly modular, readily plugged into existing deep CNNs. Extensive experiments are conducted on large-scale object classification, scene categorization, fine-grained visual recognition and texture classification, showing our methods outperform the counterparts and obtain state-of-the-art performance. Qilong Wang 0001, Jiangtao Xie, Wangmeng Zuo, Lei Zhang 0006, Peihua Li |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2021 | AGUnet: Annotation-guided U-net for fast one-shot video object segmentation
Yingjie Yin, De Xu, Xingang Wang 0003, Lei Zhang 0006 |
Pattern Recognit. | 4 |
| 2021 | Scaled Simplex Representation for Subspace ClusteringabstractThe self-expressive property of data points, that is, each data point can be linearly represented by the other data points in the same subspace, has proven effective in leading subspace clustering (SC) methods. Most self-expressive methods usually construct a feasible affinity matrix from a coefficient matrix, obtained by solving an optimization problem. However, the negative entries in the coefficient matrix are forced to be positive when constructing the affinity matrix via exponentiation, absolute symmetrization, or squaring operations. This consequently damages the inherent correlations among the data. Besides, the affine constraint used in these methods is not flexible enough for practical applications. To overcome these problems, in this article, we introduce a scaled simplex representation (SSR) for the SC problem. Specifically, the non-negative constraint is used to make the coefficient matrix physically meaningful, and the coefficient vector is constrained to be summed up to a scalar to make it more discriminative. The proposed SSR-based SC (SSRSC) model is reformulated as a linear equality-constrained problem, which is solved efficiently under the alternating direction method of multipliers framework. Experiments on benchmark datasets demonstrate that the proposed SSRSC algorithm is very efficient and outperforms the state-of-the-art SC methods on accuracy. The code can be found at https://github.com/csjunxu/SSRSC. Jun Xu 0019, Mengyang Yu, Ling Shao 0001, Wangmeng Zuo, Deyu Meng, Lei Zhang 0006, David Zhang 0001 |
IEEE Trans. Cybern. | 6 |
| 2021 | Joint Denoising and Demosaicking With Green Channel Prior for Real-World Burst ImagesabstractDenoising and demosaicking are essential yet correlated steps to reconstruct a full color image from the raw color filter array (CFA) data. By learning a deep convolutional neural network (CNN), significant progress has been achieved to perform denoising and demosaicking jointly. However, most existing CNN-based joint denoising and demosaicking (JDD) methods work on a single image while assuming additive white Gaussian noise, which limits their performance on real-world applications. In this work, we study the JDD problem for real-world burst images, namely JDD-B. Considering the fact that the green channel has twice the sampling rate and better quality than the red and blue channels in CFA raw data, we propose to use this green channel prior (GCP) to build a GCP-Net for the JDD-B task. In GCP-Net, the GCP features extracted from green channels are utilized to guide the feature extraction and feature upsampling of the whole image. To compensate for the shift between frames, the offset is also estimated from GCP features to reduce the impact of noise. Our GCP-Net can preserve more image structures and details than other JDD methods while removing noise. Experiments on synthetic and real-world noisy images demonstrate the effectiveness of GCP-Net quantitatively and qualitatively. Shi Guo, Zhetong Liang, Lei Zhang 0006 |
IEEE Trans. Image Process. | 3 |
| 2021 | MGSeg: Multiple Granularity-Based Real-Time Semantic Segmentation NetworkabstractRecent works on semantic segmentation witness significant performance improvement by utilizing global contextual information. In this paper, an efficient multi-granularity based semantic segmentation network (MGSeg) is proposed for real-time semantic segmentation, by modeling the latent relevance between multi-scale geometric details and high-level semantics for fine granularity segmentation. In particular, a light-weight backbone ResNet-18 is first adopted to produce the hierarchical features. Hybrid Attention Feature Aggregation (HAFA) is designed to filter the noisy spatial details of features, acquire the scale-invariance representation, and alleviate the gradient vanishing problem of the early-stage feature learning. After aggregating the learned features, Fine Granularity Refinement (FGR) module is employed to explicitly model the relationship between the multi-level features and categories, generating proper weights for fusion. More importantly, to meet the real-time processing, a series of light-weight strategies and simplified structures are applied to accelerate the efficiency, including light-weight backbone, channel compression, narrow neck structure, and so on. Extensive experiments conducted on benchmark datasets Cityscapes and CamVid demonstrate that the proposed method achieves the state-of-the-art performance, 77.8%@50fps and 72.7%@127fps on Cityscapes and CamVid datasets, respectively, having the capability for real-time applications. Jun-Yan He, Shi-Hua Liang, Xiao Wu 0001, Bo Zhao 0032, Lei Zhang 0006 |
IEEE Trans. Image Process. | 5 |
| 2021 | Online Rain/Snow Removal From Surveillance VideosabstractVideo rain/snow removal from surveillance videos is an important task in the computer vision community since rain/snow existed in videos can severely degenerate the performance of many surveillance system. Various methods have been investigated extensively, but most only consider consistent rain/snow under stable background scenes. Rain/snow captured from practical surveillance camera, however, is always highly dynamic in time, and those videos also include occasionally transformed background scenes and background motions caused by waving leaves or water surfaces. To this issue, this paper proposes a novel rain/snow removal approach, which fully considers dynamic statistics of both rain/snow and background scenes taken from a video sequence. Specifically, the rain/snow is encoded as an online multi-scale convolutional sparse coding (OMS-CSC) model, which not only finely delivers the sparse scattering and multi-scale shapes of real rain/snow, but also well distinguish the components of background motion from rain/snow layer. The real-time ameliorated parameters in the model well encodes their temporally dynamic configurations. Furthermore, a transformation operator imposed on the background scenes is further embedded into the proposed model, which finely conveys the background transformations, such as rotations, scalings and distortions, inevitably existed in a real video sequence. The approach so constructed can naturally better adapt to the dynamic rain/snow as well as background changes, and also suitable to deal with the streaming video attributed its online learning mode. The proposed model is formulated in a concise maximum a posterior (MAP) framework and is readily solved by the alternating direction method of multipliers (ADMM). Compared with the state-of-the-art online and offline video rain/snow removal methods, the proposed method achieves best performance on synthetic and real videos datasets both visually and quantitatively. Specifically, our method can be implemented in relatively high efficiency, showing its potential to real-time video rain/snow removal. The code page is at: https://github.com/MinghanLi/OTMSCSC_matlab_2020. Minghan Li 0001, Xiangyong Cao, Qian Zhao 0002, Lei Zhang 0006, Deyu Meng |
IEEE Trans. Image Process. | 4 |
| 2021 | CameraNet: A Two-Stage Framework for Effective Camera ISP LearningabstractTraditional image signal processing (ISP) pipeline consists of a set of cascaded image processing modules onboard a camera to reconstruct a high-quality sRGB image from the sensor raw data. Recently, some methods have been proposed to learn a convolutional neural network (CNN) to improve the performance of traditional ISP. However, in these works usually a CNN is directly trained to accomplish the ISP tasks without considering much the correlation among the different components in an ISP. As a result, the quality of reconstructed images is barely satisfactory in challenging scenarios such as low-light imaging. In this paper, we firstly analyze the correlation among the different tasks in an ISP, and categorize them into two weakly correlated groups: restoration and enhancement. Then we design a two-stage network, called CameraNet, to progressively learn the two groups of ISP tasks. In each stage, a ground truth is specified to supervise the subnetwork learning, and the two subnetworks are jointly fine-tuned to produce the final output. Experiments on three benchmark datasets show that the proposed CameraNet achieves consistently compelling reconstruction quality and outperforms the recently proposed ISP learning methods. Zhetong Liang, Jianrui Cai, Zisheng Cao, Lei Zhang 0006 |
IEEE Trans. Image Process. | 4 |
| 2021 | Generalized Unitarily Invariant Gauge Regularization for Fast Low-Rank Matrix RecoveryabstractSpectral regularization is a widely used approach for low-rank matrix recovery (LRMR) by regularizing matrix singular values. Most of the existing LRMR solvers iteratively compute the singular values via applying singular value decomposition (SVD) on a dense matrix, which is computationally expensive and severely limits their applications to large-scale problems. To address this issue, we present a generalized unitarily invariant gauge (GUIG) function for LRMR. The proposed GUIG function does not act on the singular values; however, we show that it generalizes the well-known spectral functions, including the rank function, the Schatten- p quasi-norm, and logsum of singular values. The proposed GUIG regularization model can be formulated as a bilinear variational problem, which can be efficiently solved without computing SVD. Such a property makes it well suited for large-scale LRMR problems. We apply the proposed GUIG model to matrix completion and robust principal component analysis and prove the convergence of the algorithms. Experimental results demonstrate that the proposed GUIG method is not only more accurate but also much faster than the state-of-the-art algorithms, especially on large-scale problems. Xixi Jia, Xiangchu Feng, Weiwei Wang 0005, Lei Zhang 0006 |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2020 | Part-Aware Attention Network for Person Re-identification
Wangmeng Xiang, Jianqiang Huang 0001, Xian-Sheng Hua 0001, Lei Zhang 0006 |
ACCV (4) | 4 |
| 2020 | Second-Order Camera-Aware Color Transformation for Cross-Domain Person Re-identification
Wangmeng Xiang, Hongwei Yong, Jianqiang Huang 0001, Xian-Sheng Hua 0001, Lei Zhang 0006 |
ACCV (2) | 5 |
| 2020 | Degradation Model Learning for Real-World Single Image Super-Resolution
Hongwei Yong, Lei Zhang 0006 |
ACCV (2) | 3 |
| 2020 | CPR-GCN: Conditional Partial-Residual Graph Convolutional Network in Automated Anatomical Labeling of Coronary ArteriesabstractAutomated anatomical labeling plays a vital role in coronary artery disease diagnosing procedure. The main challenge in this problem is the large individual variability inherited in human anatomy. Existing methods usually rely on the position information and the prior knowledge of the topology of the coronary artery tree, which may lead to unsatisfactory performance when the main branches are confusing. Motivated by the wide application of the graph neural network in structured data, in this paper, we propose a conditional partial-residual graph convolutional network (CPR-GCN), which takes both position and CT image into consideration, since CT image contains abundant information such as branch size and spanning direction. Two majority parts, a Partial-Residual GCN and a conditions extractor, are included in CPR-GCN. The conditions extractor is a hybrid model containing the 3D CNN and the LSTM, which can extract 3D spatial image features along the branches. On the technical side, the Partial-Residual GCN takes the position features of the branches, with the 3D spatial image features as conditions, to predict the label for each branches. While on the mathematical side, our approach twists the partial differential equation (PDE) into the graph modeling. A dataset with 511 subjects is collected from the clinic and annotated by two experts with a two-phase annotation process. According to the five-fold cross-validation, our CPR-GCN yields 95.8% meanRecall, 95.4% meanPrecision and 0.955 meanF1, which outperforms state-of-the-art approaches. Han Yang 0009, Xingjian Zhen, Ying Chi, Lei Zhang 0006, Xian-Sheng Hua 0001 |
CVPR | 4 |
| 2020 | Structure Aware Single-Stage 3D Object Detection From Point Cloudabstract3D object detection from point cloud data plays an essential role in autonomous driving. Current single-stage detectors are efficient by progressively downscaling the 3D point clouds in a fully convolutional manner. However, the downscaled features inevitably lose spatial information and cannot make full use of the structure information of 3D point cloud, degrading their localization precision. In this work, we propose to improve the localization precision of single-stage detectors by explicitly leveraging the structure information of 3D point cloud. Specifically, we design an auxiliary network which converts the convolutional features in the backbone network back to point-level representations. The auxiliary network is jointly optimized, by two point-level supervisions, to guide the convolutional features in the backbone network to be aware of the object structure. The auxiliary network can be detached after training and therefore introduces no extra computation in the inference stage. Besides, considering that single-stage detectors suffer from the discordance between the predicted bounding boxes and corresponding classification confidences, we develop an efficient part-sensitive warping operation to align the confidences to the predicted bounding boxes. Our proposed detector ranks at the top of KITTI 3D/BEV detection leaderboards and runs at 25 FPS for inference. Chenhang He, Hui Zeng 0001, Jianqiang Huang 0001, Xian-Sheng Hua 0001, Lei Zhang 0006 |
CVPR | 5 |
| 2020 | Multi-Domain Learning for Accurate and Few-Shot Color ConstancyabstractColor constancy is an important process in camera pipeline to remove the color bias of captured image caused by scene illumination. Recently, significant improvements in color constancy accuracy have been achieved by using deep neural networks (DNNs). However, existing DNNbased color constancy methods learn distinct mappings for different cameras, which require a costly data acquisition process for each camera device. In this paper, we start a pioneer work to introduce multi-domain learning to color constancy area. For different camera devices, we train a branch of networks which share the same feature extractor and illuminant estimator, and only employ a camera-specific channel re-weighting module to adapt to the camera-specific characteristics. Such a multi-domain learning strategy enables us to take benefit from crossdevice training data. The proposed multi-domain learning color constancy method achieved state-of-the-art performance on three commonly used benchmark datasets. Furthermore, we also validate the proposed method in a fewshot color constancy setting. Given a new unseen device with limited number of training samples, our method is capable of delivering accurate color constancy by merely learning the camera-specific parameters from the few-shot dataset. Our project page is publicly available at https://github.com/msxiaojin/MDLCC. Shuhang Gu, Lei Zhang 0006 |
CVPR | 3 |
| 2020 | Blind Face Restoration via Deep Multi-scale Component Dictionaries
Xiaoming Li 0002, Chaofeng Chen, Shangchen Zhou, Xianhui Lin, Wangmeng Zuo, Lei Zhang 0006 |
ECCV (9) | 6 |
| 2020 | LST-Net: Learning a Convolutional Neural Network with a Learnable Sparse Transform
Lida Li, Kun Wang 0027, Shuai Li 0014, Xiangchu Feng, Lei Zhang 0006 |
ECCV (10) | 5 |
| 2020 | A Decoupled Learning Scheme for Real-World Burst Denoising from Raw Images
Zhetong Liang, Shi Guo, Huaqi Zhang, Lei Zhang 0006 |
ECCV (25) | 5 |
| 2020 | Gradient Centralization: A New Optimization Technique for Deep Neural Networks
Hongwei Yong, Jianqiang Huang 0001, Xian-Sheng Hua 0001, Lei Zhang 0006 |
ECCV (1) | 4 |
| 2020 | Momentum Batch Normalization for Deep Learning with Small Batch Size
Hongwei Yong, Jianqiang Huang 0001, Deyu Meng, Xian-Sheng Hua 0001, Lei Zhang 0006 |
ECCV (12) | 5 |
| 2020 | Dual Adversarial Network: Toward Real-World Noise Removal and Noise Generation
Zongsheng Yue, Qian Zhao 0002, Lei Zhang 0006, Deyu Meng |
ECCV (10) | 3 |
| 2020 | Label Propagation with Augmented Anchors: A Simple Semi-supervised Learning Baseline for Unsupervised Domain Adaptation
Yabin Zhang 0001, Bin Deng 0003, Kui Jia, Lei Zhang 0006 |
ECCV (4) | 4 |
| 2020 | Suppress and Balance: A Simple Gated Network for Salient Object Detection
Xiaoqi Zhao 0003, Youwei Pang, Lihe Zhang, Huchuan Lu, Lei Zhang 0006 |
ECCV (2) | 5 |
| 2020 | A Single Stream Network for Robust and Real-Time RGB-D Salient Object Detection
Xiaoqi Zhao 0003, Lihe Zhang, Youwei Pang, Huchuan Lu, Lei Zhang 0006 |
ECCV (22) | 5 |
| 2020 | Optimizing Filter-bank Canonical Correlation Analysis for fast response SSVEP Brain-Computer Interface (BCI)abstractSteady-State Visual Evoked Potential (SSVEP) BCI brings high accuracy and consistent performance across subjects at the expense of a long stimulus presentation time window. Several recent methods exploited subject-specific features to improve SSVEP recognition performance in a short time window less than 1s. Although the calibration process is tedious and causes inconvenience, small calibration data with short duration resulting in higher performance gains are worth considering. So we propose a method by optimizing Filter-Bank Canonical Correlation Analysis (FBCCA) with subjects' calibrated templates, subject-specific weights and multiple reference types. The proposed method, subject-calibration extended FBCCA (SCEF) leverages independent and distinct discrimination characteristics of multiple references with subject-specific weight-adjusted features to improve SSVEP recognition performance. We tested the proposed method with different parameters compared with FBCCA baseline and state-of-the-art calibration methods on forty targets SSVEP dataset using 0.2s to 4s time windows. Our evaluation results show SCEF with three reference templates and subject-specific weighted features perform significantly better than all FBCCA variants in 0.2 s to 1 s time window (p <; 0.001). SCEF performs marginally, not statistically significant, better than existing methods about 2.69 ± 2.32% mean accuracy across time windows. Including multiple templates and subject-specific weight increases 15.73 ± 5.34% and 8.06 ± 2.06% in mean accuracy resulting the overall performance improvements in short time window. The proposed optimization only requires prior calibration data to create subject-specific templates and weights instead of learning features from calibration data every time. This enables not requiring to repeat the calibration step in every SSVEP session for the same subject while still maintaining accuracy similar to state-of-the-art calibration methods. Aung Aung Phyo Wai, Ying Chi, Lei Zhang 0006, Xian-Sheng Hua 0001, Cuntai Guan |
IJCNN | 4 |
| 2020 | Weakly Supervised Organ Localization with Attention Maps Regularized by Local Area Reconstruction
Minfeng Xu, Ying Chi, Lei Zhang 0006, Xian-Sheng Hua 0001 |
MICCAI (1) | 4 |
| 2020 | Landmarks Detection with Anatomical Constraints for Total Hip Arthroplasty Preoperative Measurements
Wei Liu 0127, Yu Wang 0108, Ying Chi, Lei Zhang 0006, Xian-Sheng Hua 0001 |
MICCAI (4) | 5 |
| 2020 | BS-MCVR: Binary-sensing based Mobile-cloud Visual RecognitionabstractThe mobile-cloud based visual recognition (MCVR) system, in which the low-end mobile sensors are deployed to persistently collect and transmit visual data to the cloud for analysis and recognition, is important for visual monitoring applications such as wildfire detection, wildlife monitoring, etc. However, the current MCVR systems are mostly human-perception-oriented, which consume many computational resources and much energy for data sensing as well as much bandwidth for data transmission, limiting their large-scale deployment. In this work, we present a machine-perception-oriented MCVR system, called BS-MCVR, where the mobile end is designed to efficiently sense highly compact and discriminative features directly from the scene, and the sensed features are analyzed on the cloud for recognition. Particularly, the mobile end is designed to operate with completely binary operations and generate fixed-point feature maps. Experiments on benchmark datasets show that our system only needs to transmit 1/200 the amount of original image data without degrading much the recognition accuracy, while it consumes minimal computational cost in the data sensing process. BS-MCVR provides a highly cost-effective solution for deploying MCVR systems at a large-scale. Hongyi Zheng, Wangmeng Zuo, Lei Zhang 0006 |
ACM Multimedia | 3 |
| 2020 | Two recent advances on normalization methods for deep neural network optimizationabstractThe normalization methods are very important for the effective and efficient optimization of deep neural networks (DNNs). The statistics such as mean and variance can be used to normalize the network activations or weights to make the training process more stable. Among the activation normalization techniques, batch normalization (BN) is the most popular one. However, BN has poor performance when the batch size is small in training. We found that the formulation of BN in the inference stage is problematic, and consequently presented a corrected one. Without any change in the training stage, the corrected BN significantly improves the inference performance when training with small batch size. Lei Zhang 0006 |
VCIP | 1 |
| 2020 | Joint optimizing of interleaving and LDPC decoding for burst errors in PON systems
Lei Zhang 0006, Chuanchuan Yang, Fan Zhang 0080 |
Sci. China Inf. Sci. | 1 |
| 2020 | Learned Dynamic Guidance for Depth Image ReconstructionabstractThe depth images acquired by consumer depth sensors (e.g., Kinect and ToF) usually are of low resolution and insufficient quality. One natural solution is to incorporate a high resolution RGB camera and exploit the statistical correlation of its data and depth. In recent years, both optimization-based and learning-based approaches have been proposed to deal with the guided depth reconstruction problems. In this paper, we introduce a weighted analysis sparse representation (WASR) model for guided depth image enhancement, which can be considered a generalized formulation of a wide range of previous optimization-based models. We unfold the optimization by the WASR model and conduct guided depth reconstruction with dynamically changed stage-wise operations. Such a guidance strategy enables us to dynamically adjust the stage-wise operations that update the depth image, thus improving the reconstruction quality and speed. To learn the stage-wise operations in a task-driven manner, we propose two parameterizations and their corresponding methods: dynamic guidance with Gaussian RBF nonlinearity parameterization (DG-RBF) and dynamic guidance with CNN nonlinearity parameterization (DG-CNN). The network structures of the proposed DG-RBF and DG-CNN methods are designed with the the objective function of our WASR model in mind and the optimal network parameters are learned from paired training data. Such optimization-inspired network architectures enable our models to leverage the previous expertise as well as take benefit from training data. The effectiveness is validated for guided depth image super-resolution and for realistic depth image reconstruction tasks using standard benchmarks. Our DG-RBF and DG-CNN methods achieve the best quantitative results (RMSE) and better visual quality than the state-of-the-art approaches at the time of writing. The code is available at https://github.com/ShuhangGu/GuidedDepthSR. Shuhang Gu, Shi Guo, Wangmeng Zuo, Yunjin Chen, Radu Timofte, Luc Van Gool, Lei Zhang 0006 |
IEEE Trans. Pattern Anal. Mach. Intell. | 7 |
| 2020 | Group Maximum Differentiation Competition: Model Comparison with Few SamplesabstractIn many science and engineering fields that require computational models to predict certain physical quantities, we are often faced with the selection of the best model under the constraint that only a small sample set can be physically measured. One such example is the prediction of human perception of visual quality, where sample images live in a high dimensional space with enormous content variations. We propose a new methodology for model comparison named group maximum differentiation (gMAD) competition. Given multiple computational models, gMAD maximizes the chances of falsifying a "defender" model using the rest models as "attackers". It exploits the sample space to find sample pairs that maximally differentiate the attackers while holding the defender fixed. Based on the results of the attacking-defending game, we introduce two measures, aggressiveness and resistance, to summarize the performance of each model at attacking other models and defending attacks from other models, respectively. We demonstrate the gMAD competition using three examples-image quality, image aesthetics, and streaming video quality-of-experience. Although these examples focus on visually discriminable quantities, the gMAD methodology can be extended to many other fields, and is especially useful when the sample space is large, the physical measurement is expensive and the cost of computational prediction is low. Kede Ma, Zhengfang Duanmu, Zhou Wang 0001, Qingbo Wu 0001, Wentao Liu 0001, Hongwei Yong, Hongliang Li 0001, Lei Zhang 0006 |
IEEE Trans. Pattern Anal. Mach. Intell. | 8 |
| 2020 | A concave optimization algorithm for matching partially overlapping point sets
Wei Lian, Lei Zhang 0006 |
Pattern Recognit. | 2 |
| 2020 | Adversarial Feature Sampling Learning for Efficient Visual TrackingabstractThe tracking-by-detection tracking framework usually consists of two stages: drawing samples around the target object and classifying each sample as either the target object or background. Current popular trackers under this framework typically draw many samples from the raw image and feed them into the deep neural networks, resulting in high computational burden and low tracking speed. In this article, we propose an adversarial feature sampling learning (AFSL) method to address this problem. A convolutional neural network is designed, which takes only one cropped image around the target object as input, and samples are collected from the feature maps with spatial bilinear resampling. To enrich the appearance variations of positive samples in the feature space, which has limited spatial resolution, we fuse the high-level features and low-level features to better describe the target by using a generative adversarial network. Extensive experiments on benchmark data sets demonstrate that the proposed ASFL achieves leading tracking accuracy while significantly accelerating the speed of tracking-by-detection trackers. Yingjie Yin, De Xu, Xingang Wang 0003, Lei Zhang 0006 |
IEEE Trans Autom. Sci. Eng. | 4 |
| 2020 | Learning a Single Tucker Decomposition Network for Lossy Image Compression With Multiple Bits-per-Pixel RatesabstractLossy image compression (LIC), which aims to utilize inexact approximations to represent an image more compactly, is a classical problem in image processing. Recently, deep convolutional neural networks (CNNs) have achieved interesting results in LIC by learning an encoder-quantizer-decoder network from a large amount of data. However, existing CNN-based LIC methods generally train a network for a specific bits-perpixel (bpp). Such a "one-network-per-bpp" problem limits the generality and flexibility of CNNs to practical LIC applications. In this paper, we propose to learn a single CNN which can perform LIC at multiple bpp rates. A simple yet effective Tucker Decomposition Network (TDNet) is developed, where there is a novel tucker decomposition layer (TDL) to decompose a latent image representation into a set of projection matrices and a core tensor. By changing the rank of core tensor and its quantization, we can easily adjust the bpp rate of latent image representation within a single CNN. Furthermore, an iterative non-uniform quantization scheme is presented to optimize the quantizer, and a coarse-to-fine training strategy is introduced to reconstruct the decompressed images. Extensive experiments demonstrate the state-of-the-art compression performance of TDNet in terms of both PSNR and MS-SSIM indices. Jianrui Cai, Zisheng Cao, Lei Zhang 0006 |
IEEE Trans. Image Process. | 3 |
| 2020 | Dark and Bright Channel Prior Embedded Network for Dynamic Scene DeblurringabstractRecent years have witnessed the significant progress on convolutional neural networks (CNNs) in dynamic scene deblurring. While most of the CNN models are generally learned by the reconstruction loss defined on training data, incorporating suitable image priors as well as regularization terms into the network architecture could boost the deblurring performance. In this work, we propose a Dark and Bright Channel Priors embedded Network (DBCPeNet) to plug the channel priors into a neural network for effective dynamic scene deblurring. A novel trainable dark and bright channel priors embedded layer (DBCPeL) is developed to aggregate both channel priors and blurry image representations, and a sparse regularization is introduced to regularize the DBCPeNet model learning. Furthermore, we present an effective multi-scale network architecture, namely image full scale exploitation (IFSE), which works in both coarse-to-fine and fine-to-coarse manners for better exploiting information flow across scales. Experimental results on the GoPro and Köhler datasets show that our proposed DBCPeNet performs favorably against state-of-the-art deep image deblurring methods in terms of both quantitative metrics and visual quality. Jianrui Cai, Wangmeng Zuo, Lei Zhang 0006 |
IEEE Trans. Image Process. | 3 |
| 2020 | Deep Saliency Hashing for Fine-Grained RetrievalabstractIn recent years, hashing methods have been proved to be effective and efficient for large-scale Web media search. However, the existing general hashing methods have limited discriminative power for describing fine-grained objects that share similar overall appearance but have a subtle difference. To solve this problem, we for the first time introduce the attention mechanism to the learning of fine-grained hashing codes. Specifically, we propose a novel deep hashing model, named deep saliency hashing (DSaH), which automatically mines salient regions and learns semantic-preserving hashing codes simultaneously. DSaH is a two-step end-to-end model consisting of an attention network and a hashing network. Our loss function contains three basic components, including the semantic loss, the saliency loss, and the quantization loss. As the core of DSaH, the saliency loss guides the attention network to mine discriminative regions from pairs of images.We conduct extensive experiments on both fine-grained and general retrieval datasets for performance evaluation. Experimental results on fine-grained datasets, including Oxford Flowers, Stanford Dogs, and CUB Birds demonstrate that our DSaH performs the best for the fine-grained retrieval task and beats the strongest competitor (DTQ) by approximately 10% on both Stanford Dogs and CUB Birds. DSaH is also comparable to several state-of-the-art hashing methods on CIFAR-10 and NUS-WIDE. Sheng Jin 0002, Hongxun Yao, Xiaoshuai Sun, Shangchen Zhou, Lei Zhang 0006, Xian-Sheng Hua 0001 |
IEEE Trans. Image Process. | 5 |
| 2020 | Learning Symmetry Consistent Deep CNNs for Face CompletionabstractDeep convolutional networks (CNNs) have achieved great success in face completion to generate plausible facial structures. These methods, however, are limited in maintaining global consistency among face components and recovering fine facial details. On the other hand, reflectional symmetry is a prominent property of face images and benefits face analysis and consistency modeling, yet remaining uninvestigated in deep face completion. In this work, we leverage two kinds of symmetry-enforcing modules to form a symmetry-consistent CNN model (i.e., SymmFCNet) for effective face completion. For missing pixels on only one of the half-faces, an illumination-reweighted warping subnet is developed to guide the warping and illumination reweighting of the other half-face. As for missing pixels on both of half-faces, we present a generative reconstruction subnet together with a perceptual symmetry loss to enforce symmetry consistency of recovered structures. The SymmFCNet is constructed by stacking generative reconstruction subnet upon illumination-reweighted warping subnet, and can be learned in an end-to-end manner. Experiments show that SymmFCNet can generate globally consistent results on images with synthetic and real occlusions, and performs favorably against state-of-the-arts. Xiaoming Li 0002, Guosheng Hu, Jieru Zhu, Wangmeng Zuo, Meng Wang 0001, Lei Zhang 0006 |
IEEE Trans. Image Process. | 6 |
| 2020 | Fast Multi-Scale Structural Patch Decomposition for Multi-Exposure Image FusionabstractExposure bracketing is crucial to high dynamic range imaging, but it is prone to halos for static scenes and ghosting artifacts for dynamic scenes. The recently proposed structural patch decomposition for multi-exposure fusion (SPD-MEF) has achieved reliable performance in deghosting, but suffers from visible halo artifacts and is computationally expensive. In addition, its relationship to other MEF methods is unclear. We show that without explicitly performing structural patch decomposition, we arrive at an unnormalized version of SPD-MEF, which enjoys an order of 30× speed-up, and is closely related to pixel-level MEF methods as well as the standard two-layer decomposition method for MEF. Moreover, we develop a fast multi-scale SPD-MEF method, which can effectively reduce halo artifacts. Experimental results demonstrate the effectiveness of the proposed MEF method in terms of speed and quality. Hui Li 0029, Kede Ma, Hongwei Yong, Lei Zhang 0006 |
IEEE Trans. Image Process. | 4 |
| 2020 | Remove Cosine Window From Correlation Filter-Based Visual Trackers: When and HowabstractCorrelation filters (CFs) have been continuously advancing the state-of-the-art tracking performance and have been extensively studied in the recent few years. Nonetheless, the existing CF trackers adopt a cosine window to spatially reweight base image to alleviate boundary discontinuity. However, cosine window emphasizes more on the central regions of base image and has the risk of contaminating negative training samples during model learning. On the other hand, spatial regularization deployed in many recent CF trackers plays a similar role as cosine window by enforcing spatial penalty on CF coefficients. Therefore, we in this paper investigate the feasibility to remove cosine window from CF trackers with spatial regularization. When simply removing cosine window, CF with spatial regularization still suffers from small degree of boundary discontinuity. To tackle this issue, binary and Gaussian shaped mask functions are further introduced for eliminating boundary discontinuity while reweighting the estimation error of each training sample, and can be incorporated with multiple CF trackers with spatial regularization. In comparison to the baseline methods with cosine window, our methods are effective in handling boundary discontinuity and sample contamination, thereby benefiting tracking performance. Extensive experiments on four benchmarks show that our methods perform favorably against the state-of-the-art trackers using either handcrafted or deep CNN features. Feng Li 0031, Xiaohe Wu, Wangmeng Zuo, David Zhang 0001, Lei Zhang 0006 |
IEEE Trans. Image Process. | 5 |
| 2020 | Confidence-Based Large-Scale Dense Multi-View StereoabstractAlbeit remarkable progress has been made to improve the accuracy and completeness of multi-view stereo (MVS), existing methods still suffer from either sparse reconstructions of low-textured surfaces or heavy computational burden. In this paper, we propose a Confidence-based Large-scale Dense Multi-view Stereo (CLD-MVS) method for high resolution imagery. Firstly, we formulate MVS as a multi-view depth estimation problem, and employ a normal-aware efficient PatchMatch stereo to estimate the initial depth and normal map for each reference view. A self-supervised deep learning method is then developed to predict the spatial confidence for multi-view depth maps, which is combined with cross-view consistency to generate the ground control points. Subsequently, a confidence-driven and boundary-aware interpolation scheme using static and dynamic guidance is adopted to synthesize dense depth and normal maps. Finally, a refinement procedure which leverages synthesized depth and normal as prior is conducted to estimate cross-view consistent surface. Experiments show that the proposed CLD-MVS method achieves high geometric completeness while preserving fine-scale details. In particular, it has ranked No. 1 on the ETH3D high-resolution MVS benchmark in terms of F1-score. Zhaoxin Li, Wangmeng Zuo, Lei Zhang 0006 |
IEEE Trans. Image Process. | 4 |
| 2020 | A Unified Probabilistic Formulation of Image Aesthetic AssessmentabstractImage aesthetic assessment (IAA) has been attracting considerable attention in recent years due to the explosive growth of digital photography in Internet and social networks. The IAA problem is inherently challenging, owning to the ineffable nature of the human sense of aesthetics and beauty, and its close relationship to understanding pictorial content. Three different approaches to framing and solving the problem have been posed: binary classification, average score regression and score distribution prediction. Solutions that have been proposed have utilized different types of aesthetic labels and loss functions to train deep IAA models. However, these studies ignore the fact that the three different IAA tasks are inherently related. Here, we reveal that the use of the different types of aesthetic labels can be developed within the same statistical framework, which we use to create a unified probabilistic formulation of all the three IAA tasks. This unified formulation motivates the use of an efficient and effective loss function for training deep IAA models to conduct different tasks. We also discuss the problem of learning from a noisy raw score distribution which hinders network performance. We then show that by fitting the raw score distribution to a more stable and discriminative score distribution, we are able to train a single model which is able to obtain highly competitive performance on all three IAA tasks. Extensive qualitative analysis and experimental results on image aesthetic benchmarks validate the superior performance afforded by the proposed formulation. The source code is available at. Hui Zeng 0001, Zisheng Cao, Lei Zhang 0006, Alan C. Bovik |
IEEE Trans. Image Process. | 3 |
| 2020 | PID Controller-Based Stochastic Optimization Acceleration for Deep Neural NetworksabstractDeep neural networks (DNNs) are widely used and demonstrated their power in many applications, such as computer vision and pattern recognition. However, the training of these networks can be time consuming. Such a problem could be alleviated by using efficient optimizers. As one of the most commonly used optimizers, stochastic gradient descent-momentum (SGD-M) uses past and present gradients for parameter updates. However, in the process of network training, SGD-M may encounter some drawbacks, such as the overshoot phenomenon. This problem would slow the training convergence. To alleviate this problem and accelerate the convergence of DNN optimization, we propose a proportional-integral-derivative (PID) approach. Specifically, we investigate the intrinsic relationships between the PID-based controller and SGD-M first. We further propose a PID-based optimization algorithm to update the network parameters, where the past, current, and change of gradients are exploited. Consequently, our proposed PID-based optimization alleviates the overshoot problem suffered by SGD-M. When tested on popular DNN architectures, it also obtains up to 50% acceleration with competitive accuracy. Extensive experiments about computer vision and natural language processing demonstrate the effectiveness of our method on benchmark data sets, including CIFAR10, CIFAR100, Tiny-ImageNet, and PTB. We have released the code at https://github.com/tensorboy/PIDOptimizer. Haoqian Wang, Wangpeng An, Qingyun Sun, Jun Xu 0019, Lei Zhang 0006 |
IEEE Trans. Neural Networks Learn. Syst. | 6 |
| 2020 | Robust Multiview Subspace Learning With Nonindependently and Nonidentically Distributed Complex NoiseabstractMultiview Subspace Learning (MSL), which aims at obtaining a low-dimensional latent subspace from multiview data, has been widely used in practical applications. Most recent MSL approaches, however, only assume a simple independent identically distributed (i.i.d.) Gaussian or Laplacian noise for all views of data, which largely underestimates the noise complexity in practical multiview data. Actually, in real cases, noises among different views generally have three specific characteristics. First, in each view, the data noise always has a complex configuration beyond a simple Gaussian or Laplacian distribution. Second, the noise distributions of different views of data are generally nonidentical and with evident distinctiveness. Third, noises among all views are nonindependent but obviously correlated. Based on such understandings, we elaborately construct a new MSL model by more faithfully and comprehensively considering all these noise characteristics. First, the noise in each view is modeled as a Dirichlet process (DP) Gaussian mixture model (DPGMM), which can fit a wider range of complex noise types than conventional Gaussian or Laplacian. Second, the DPGMM parameters in each view are different from one another, which encodes the "nonidentical" noise property. Third, the DPGMMs on all views share the same high-level priors by using the technique of hierarchical DP, which encodes the "nonindependent" noise property. All the aforementioned ideas are incorporated into an integrated graphics model which can be appropriately solved by the variational Bayes algorithm. The superiority of the proposed method is verified by experiments on 3-D reconstruction simulations, multiview face modeling, and background subtraction, as compared with the current state-of-the-art MSL methods. Zongsheng Yue, Hongwei Yong, Deyu Meng, Qian Zhao 0002, Yee Leung, Lei Zhang 0006 |
IEEE Trans. Neural Networks Learn. Syst. | 6 |
| 2019 | Learning a Visual Tracker from a Single Movie without AnnotationabstractThe recent success of deep network in visual trackers learning largely relies on human labeled data, which are however expensive to annotate. Recently, some unsupervised methods have been proposed to explore the learning of visual trackers without labeled data, while their performance lags far behind the supervised methods. We identify the main bottleneck of these methods as inconsistent objectives between off-line training and online tracking stages. To address this problem, we propose a novel unsupervised learning pipeline which is based on the discriminative correlation filter network. Our method iteratively updates the tracker by alternating between target localization and network optimization. In particular, we propose to learn the network from a single movie, which could be easily obtained other than collecting thousands of video clips or millions of images. Extensive experiments demonstrate that our approach is insensitive to the employed movies, and the trained visual tracker achieves leading performance among existing unsupervised learning approaches. Even compared with the same network trained with human labeled bounding boxes, our tracker achieves similar results on many tracking benchmarks. Code is available at: https://github.com/ZjjConan/UL-Tracker-AAAI2019. Lingxiao Yang, David Zhang 0001, Lei Zhang 0006 |
AAAI | 3 |
| 2019 | Deep Plug-And-Play Super-Resolution for Arbitrary Blur KernelsabstractWhile deep neural networks (DNN) based single image super-resolution (SISR) methods are rapidly gaining popularity, they are mainly designed for the widely-used bicubic degradation, and there still remains the fundamental challenge for them to super-resolve low-resolution (LR) image with arbitrary blur kernels. In the meanwhile, plug-and-play image restoration has been recognized with high flexibility due to its modular structure for easy plug-in of denoiser priors. In this paper, we propose a principled formulation and framework by extending bicubic degradation based deep SISR with the help of plug-and-play framework to handle LR images with arbitrary blur kernels. Specifically, we design a new SISR degradation model so as to take advantage of existing blind deblurring methods for blur kernel estimation. To optimize the new degradation induced energy function, we then derive a plug-and-play algorithm via variable splitting technique, which allows us to plug any super-resolver prior rather than the denoiser prior as a modular part. Quantitative and qualitative evaluations on synthetic and real LR images demonstrate that the proposed deep plug-and-play super-resolution framework is flexible and effective to deal with blurry LR images. Kai Zhang 0008, Wangmeng Zuo, Lei Zhang 0006 |
CVPR | 3 |
| 2019 | Second-Order Attention Network for Single Image Super-ResolutionabstractRecently, deep convolutional neural networks (CNNs) have been widely explored in single image super-resolution (SISR) and obtained remarkable performance. However, most of the existing CNN-based SISR methods mainly focus on wider or deeper architecture design, neglecting to explore the feature correlations of intermediate layers, hence hindering the representational power of CNNs. To address this issue, in this paper, we propose a second-order attention network (SAN) for more powerful feature expression and feature correlation learning. Specifically, a novel train- able second-order channel attention (SOCA) module is developed to adaptively rescale the channel-wise features by using second-order feature statistics for more discriminative representations. Furthermore, we present a non-locally enhanced residual group (NLRG) structure, which not only incorporates non-local operations to capture long-distance spatial contextual information, but also contains repeated local-source residual attention groups (LSRAG) to learn increasingly abstract feature representations. Experimental results demonstrate the superiority of our SAN network over state-of-the-art SISR methods in terms of both quantitative metrics and visual quality. Tao Dai 0001, Jianrui Cai, Yongbing Zhang 0002, Shutao Xia, Lei Zhang 0006 |
CVPR | 5 |
| 2019 | Toward Convolutional Blind Denoising of Real PhotographsabstractWhile deep convolutional neural networks (CNNs) have achieved impressive success in image denoising with additive white Gaussian noise (AWGN), their performance remains limited on real-world noisy photographs. The main reason is that their learned models are easy to overfit on the simplified AWGN model which deviates severely from the complicated real-world noise model. In order to improve the generalization ability of deep CNN denoisers, we suggest training a convolutional blind denoising network (CBDNet) with more realistic noise model and real-world noisy-clean image pairs. On the one hand, both signal-dependent noise and in-camera signal processing pipeline is considered to synthesize realistic noisy images. On the other hand, real-world noisy photographs and their nearly noise-free counterparts are also included to train our CBDNet. To further provide an interactive strategy to rectify denoising result conveniently, a noise estimation subnetwork with asymmetric learning to suppress under-estimation of noise level is embedded into CBDNet. Extensive experimental results on three datasets of real-world noisy photographs clearly demonstrate the superior performance of CBDNet over state-of-the-arts in terms of quantitative met- rics and visual quality. The code has been made available at https://github.com/GuoShi28/CBDNet. Shi Guo, Zifei Yan, Kai Zhang 0008, Wangmeng Zuo, Lei Zhang 0006 |
CVPR | 5 |
| 2019 | FOCNet: A Fractional Optimal Control Network for Image DenoisingabstractDeep convolutional neural networks (DCNN) have been successfully used in many low-level vision problems such as image denoising. Recent studies on the mathematical foundation of DCNN has revealed that the forward propagation of DCNN corresponds to a dynamic system, which can be described by an ordinary differential equation (ODE) and solved by the optimal control method. However, most of these methods employ integer-order differential equation, which has local connectivity in time space and cannot describe the long-term memory of the system. Inspired by the fact that the fractional-order differential equation has long-term memory, in this paper we develop an advanced image denoising network, namely FOCNet, by solving a fractional optimal control (FOC) problem. Specifically, the network structure is designed based on the discretization of a fractional-order differential equation, which enjoys long-term memory in both forward and backward passes. Besides, multi-scale feature interactions are introduced into the FOCNet to strengthen the control of the dynamic system. Extensive experiments demonstrate the leading performance of the proposed FOCNet on image denoising. Code will be made available. Xixi Jia, Xiangchu Feng, Lei Zhang 0006 |
CVPR | 4 |
| 2019 | Reliable and Efficient Image Cropping: A Grid Anchor Based ApproachabstractImage cropping aims to improve the composition as well as aesthetic quality of an image by removing extraneous content from it. Existing image cropping databases provide only one or several human-annotated bounding boxes as the groundtruth, which cannot reflect the non-uniqueness and flexibility of image cropping in practice. The employed evaluation metrics such as intersection-over-union cannot reliably reflect the real performance of cropping models, either. This work revisits the problem of image cropping, and presents a grid anchor based formulation by considering the special properties and requirements (e.g., local redundancy, content preservation, aspect ratio) of image cropping. Our formulation reduces the searching space of candidate crops from millions to less than one hundred. Consequently, a grid anchor based cropping benchmark is constructed, where all crops of each image are annotated and more reliable evaluation metrics are defined. We also design an effective and lightweight network module, which simultaneously considers the region of interest and region of discard for more accurate image cropping. Our model can stably output visually pleasing crops for images of different scenes and run at a speed of 125 FPS. Hui Zeng 0001, Lida Li, Zisheng Cao, Lei Zhang 0006 |
CVPR | 4 |
| 2019 | Toward Real-World Single Image Super-Resolution: A New Benchmark and a New ModelabstractMost of the existing learning-based single image super-resolution (SISR) methods are trained and evaluated on simulated datasets, where the low-resolution (LR) images are generated by applying a simple and uniform degradation (i.e., bicubic downsampling) to their high-resolution (HR) counterparts. However, the degradations in real-world LR images are far more complicated. As a consequence, the SISR models trained on simulated data become less effective when applied to practical scenarios. In this paper, we build a real-world super-resolution (RealSR) dataset where paired LR-HR images on the same scene are captured by adjusting the focal length of a digital camera. An image registration algorithm is developed to progressively align the image pairs at different resolutions. Considering that the degradation kernels are naturally non-uniform in our dataset, we present a Laplacian pyramid based kernel prediction network (LP-KPN), which efficiently learns per-pixel kernels to recover the HR image. Our extensive experiments demonstrate that SISR models trained on our RealSR dataset deliver better visual quality with sharper edges and finer textures on real-world scenes than those trained on simulated datasets. Though our RealSR dataset is built by using only two cameras (Canon 5D3 and Nikon D810), the trained model generalizes well to other camera devices such as Sony a7II and mobile phones. Jianrui Cai, Hui Zeng 0001, Hongwei Yong, Zisheng Cao, Lei Zhang 0006 |
ICCV | 5 |
| 2019 | Dynamic Anchor Feature Selection for Single-Shot Object DetectionabstractThe design of anchors is critical to the performance of one-stage detectors. Recently, the anchor refinement module (ARM) has been proposed to adjust the initialization of default anchors, providing the detector a better anchor reference. However, this module brings another problem: all pixels at a feature map have the same receptive field while the anchors associated with each pixel have different positions and sizes. This discordance may lead to a less effective detector. In this paper, we present a dynamic feature selection operation to select new pixels in a feature map for each refined anchor received from the ARM. The pixels are selected based on the new anchor position and size so that the receptive filed of these pixels can fit the anchor areas well, which makes the detector, especially the regression part, much easier to optimize. Furthermore, to enhance the representation ability of selected feature pixels, we design a bidirectional feature fusion module by combining features from early and deep layers. Extensive experiments on both PASCAL VOC and COCO demonstrate the effectiveness of our dynamic anchor feature selection (DAFS) operation. For the case of high IoU threshold, our DAFS can improve the mAP by a large margin. Shuai Li 0014, Lingxiao Yang, Jianqiang Huang 0001, Xian-Sheng Hua 0001, Lei Zhang 0006 |
ICCV | 5 |
| 2019 | Homocentric Hypersphere Feature Embedding for Person Re-IdentificationabstractTriplet loss and softmax loss are two widely used loss functions in Person Re-Identification (Person ReID). However, previous works that try to apply these two loss functions have measure inconsistency during training and testing stage and among different parts of the total loss function, which would cause inferior performance of models. To address this issue, we propose a novel homocentric hypersphere embedding scheme to decouple magnitude and orientation information for both feature and weight vectors, and reformulate the triplet loss and the softmax loss to their angular versions and combine them into an angular discriminative loss. We evaluate our proposed method extensively on the widely used Person ReID benchmarks. Our method demonstrates leading performance on all datasets. Wangmeng Xiang, Jianqiang Huang 0001, Xianbiao Qi, Xian-Sheng Hua 0001, Lei Zhang 0006 |
ICIP | 5 |
| 2019 | Variational Denoising Network: Toward Blind Noise Modeling and RemovalabstractBlind image denoising is an important yet very challenging problem in computer vision due to the complicated acquisition process of real images. In this work we propose a new variational inference method, which integrates both noise estimation and image denoising into a unique Bayesian framework, for blind image denoising. Specifically, an approximate posterior, parameterized by deep neural networks, is presented by taking the intrinsic clean image and noise variances as latent variables conditioned on the input noisy image. This posterior provides explicit parametric forms for all its involved hyper-parameters, and thus can be easily implemented for blind image denoising with automatic noise estimation for the test noisy image. On one hand, as other data-driven deep learning methods, our method, namely variational denoising network (VDN), can perform denoising efficiently due to its explicit form of posterior expression. On the other hand, VDN inherits the advantages of traditional model-driven approaches, especially the good generalization capability of generative models. VDN has good interpretability and can be flexibly utilized to estimate and remove complicated non-i.i.d. noise collected in real scenarios. Comprehensive experiments are performed to substantiate the superiority of our method in blind image denoising. Zongsheng Yue, Hongwei Yong, Qian Zhao 0002, Deyu Meng, Lei Zhang 0006 |
NeurIPS | 5 |
| 2019 | Learning Support Correlation Filters for Visual TrackingabstractFor visual tracking methods based on kernel support vector machines (SVMs), data sampling is usually adopted to reduce the computational cost in training. In addition, budgeting of support vectors is required for computational efficiency. Instead of sampling and budgeting, recently the circulant matrix formed by dense sampling of translated image patches has been utilized in kernel correlation filters for fast tracking. In this paper, we derive an equivalent formulation of a SVM model with the circulant matrix expression and present an efficient alternating optimization method for visual tracking. We incorporate the discrete Fourier transform with the proposed alternating optimization process, and pose the tracking problem as an iterative learning of support correlation filters (SCFs). In the fully-supervision setting, our SCF can find the globally optimal solution with real-time performance. For a given circulant data matrix with$n^2$samples of$n \times n$pixels, the computational complexity of the proposed algorithm is$O(n^2\; \log n)$whereas that of the standard SVM-based approaches is at least$O(n^4)$. In addition, we extend the SCF-based tracking algorithm with multi-channel features, kernel functions, and scale-adaptive approaches to further improve the tracking performance. Experimental results on a large benchmark dataset show that the proposed SCF-based algorithms perform favorably against the state-of-the-art tracking methods in terms of accuracy and speed. Wangmeng Zuo, Xiaohe Wu, Liang Lin 0004, Lei Zhang 0006, Ming-Hsuan Yang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2019 | Sparse, collaborative, or nonnegative representation: Which helps pattern classification?
Jun Xu 0019, Wangpeng An, Lei Zhang 0006, David Zhang 0001 |
Pattern Recognit. | 3 |
| 2019 | Foreground Gating and Background Refining Network for Surveillance Object DetectionabstractDetecting objects in surveillance videos is an important problem due to its wide applications in traffic control and public security. Existing methods tend to face performance degradation because of false positive or misalignment problems. We propose a novel framework, namely, Foreground Gating and Background Refining Network (FG-BR Net), for surveillance object detection (SOD). To reduce false positives in background regions, which is a critical problem in SOD, we introduce a new module that first subtracts the background of a video sequence and then generates high-quality region proposals. Unlike previous background subtraction methods that may wrongly remove the static foreground objects in a frame, a feedback connection from detection results to background subtraction process is proposed in our model to distill both static and moving objects in surveillance videos. Furthermore, we introduce another module, namely, the background refining stage, to refine the detection results with more accurate localizations. Pairwise non-local operations are adopted to cope with the misalignments between the features of original and background frames. Extensive experiments on real-world traffic surveillance benchmarks demonstrate the competitive performance of the proposed FG-BR Net. In particular, FG-BR Net ranks on the top among all the methods on hard and sunny subsets of the UA-DETRAC detection dataset, without any bells and whistles. Zhihang Fu, Yaowu Chen, Hongwei Yong, Rongxin Jiang 0001, Lei Zhang 0006, Xian-Sheng Hua 0001 |
IEEE Trans. Image Process. | 5 |
| 2019 | Learning Converged Propagations With Deep Prior Ensemble for Image EnhancementabstractEnhancing visual qualities of images plays very important roles in various vision and learning applications. In the past few years, both knowledge-driven maximum a posterior (MAP) with prior modelings and fully data-dependent convolutional neural network (CNN) techniques have been investigated to address specific enhancement tasks. In this paper, by exploiting the advantages of these two types of mechanisms within a complementary propagation perspective, we propose a unified framework, named deep prior ensemble (DPE), for solving various image enhancement tasks. Specifically, we first establish the basic propagation scheme based on the fundamental image modeling cues and then introduce residual CNNs to help predicting the propagation direction at each stage. By designing prior projections to perform feedback control, we theoretically prove that even with experience-inspired CNNs, DPE is definitely converged and the output will always satisfy our fundamental task constraints. The main advantage against conventional optimization-based MAP approaches is that our descent directions are learned from collected training data, thus are much more robust to unwanted local minimums. While, compared with existing CNN type networks, which are often designed in heuristic manners without theoretical guarantees, DPE is able to gain advantages from rich task cues investigated on the bases of domain knowledges. Therefore, DPE actually provides a generic ensemble methodology to integrate both knowledge and data-based cues for different image enhancement tasks. More importantly, our theoretical investigations verify that the feedforward propagations of DPE are properly controlled toward our desired solution. Experimental results demonstrate that the proposed DPE outperforms state-of-the-arts on a variety of image enhancement tasks in terms of both quantitative measure and visual perception quality. Risheng Liu, Long Ma 0002, Yiyang Wang 0001, Lei Zhang 0006 |
IEEE Trans. Image Process. | 4 |
| 2019 | Panoramic Background Image Generation for PTZ CamerasabstractBeing able to cover a wide range of views, pan-tilt-zoom (PTZ) cameras have been widely deployed in visual surveillance systems. To achieve a global-view perception of a surveillance scene, it is necessary to generate its panoramic background image, which can be used for the subsequent applications such as road segmentation, active tracking, and so on. However, few works have been reported on this problem, partially due to the lack of benchmark dataset and the high complexity of panoramic image generation of PTZ cameras. In this paper, we build, for the first time to our best knowledge, a benchmark PTZ camera dataset with multiple views, and derive a complete set of panoramic transformation formulas for PTZ cameras. We further propose a fast multi-band blending method to address the efficiency issue in panoramic image fusion and mosaicing. Some related panoramic transformations are also developed, such as cylindrical and overlooking transformations. Our proposed approach exhibits impressive accuracy and efficiency in PTZ panorama generation as well as panoramic image mosaicing. Hongwei Yong, Jianqiang Huang 0001, Wangmeng Xiang, Xian-Sheng Hua 0001, Lei Zhang 0006 |
IEEE Trans. Image Process. | 5 |
| 2019 | A Benchmark for Edge-Preserving Image SmoothingabstractEdge-preserving image smoothing is an important step for many low-level vision problems. Though many algorithms have been proposed, there are several difficulties hindering its further development. First, most existing algorithms cannot perform well on a wide range of image contents using a single parameter setting. Second, the performance evaluation of edge-preserving image smoothing remains subjective, and there lacks a widely accepted datasets to objectively compare the different algorithms. To address these issues and further advance the state of the art, in this work we propose a benchmark for edge-preserving image smoothing. This benchmark includes an image dataset with groundtruth image smoothing results as well as baseline algorithms that can generate competitive edge-preserving smoothing results for a wide range of image contents. The established dataset contains 500 training and testing images with a number of representative visual object categories, while the baseline methods in our benchmark are built upon representative deep convolutional network architectures, on top of which we design novel loss functions well suited for edge-preserving image smoothing. The trained deep networks run faster than most state-of-the-art smoothing algorithms with leading smoothing results both qualitatively and quantitatively. The benchmark will be made publicly accessible. Feida Zhu 0002, Zhetong Liang, Xixi Jia, Lei Zhang 0006, Yizhou Yu |
IEEE Trans. Image Process. | 4 |
| 2019 | Learning Aggregated Transmission Propagation Networks for Haze Removal and BeyondabstractSingle-image dehazing is an important low-level vision task with many applications. Early studies have investigated different kinds of visual priors to address this problem. However, they may fail when their assumptions are not valid on specific images. Recent deep networks also achieve a relatively good performance in this task. But unfortunately, due to the disappreciation of rich physical rules in hazes, a large amount of data are required for their training. More importantly, they may still fail when there exist completely different haze distributions in testing images. By considering the collaborations of these two perspectives, this paper designs a novel residual architecture to aggregate both prior (i.e., domain knowledge) and data (i.e., haze distribution) information to propagate transmissions for scene radiance estimation. We further present a variational energy-based perspective to investigate the intrinsic propagation behavior of our aggregated deep model. In this way, we actually bridge the gap between prior-driven models and data-driven networks and leverage advantages but avoid limitations of previous dehazing approaches. A lightweight learning framework is proposed to train our propagation network. Finally, by introducing a task-aware image separation formulation with a flexible optimization scheme, we extend the proposed model for more challenging vision tasks, such as underwater image enhancement and single-image rain removal. Experiments on both synthetic and real-world images demonstrate the effectiveness and efficiency of the proposed framework. Risheng Liu, Xin Fan 0001, Minjun Hou, Zhiying Jiang, Zhongxuan Luo, Lei Zhang 0006 |
IEEE Trans. Neural Networks Learn. Syst. | 6 |
| 2019 | Cost-Effective Object Detection: Active Sample Mining With Switchable Selection CriteriaabstractThough quite challenging, leveraging large-scale unlabeled or partially labeled data in learning systems (e.g., model/classifier training) has attracted increasing attentions due to its fundamental importance. To address this problem, many active learning (AL) methods have been proposed that employ up-to-date detectors to retrieve representative minority samples according to predefined confidence or uncertainty thresholds. However, these AL methods cause the detectors to ignore the remaining majority samples (i.e., those with low uncertainty or high prediction confidence). In this paper, by developing a principled active sample mining (ASM) framework, we demonstrate that cost-effective mining samples from these unlabeled majority data are a key to train more powerful object detectors while minimizing user effort. Specifically, our ASM framework involves a switchable sample selection mechanism for determining whether an unlabeled sample should be manually annotated via AL or automatically pseudolabeled via a novel self-learning process. The proposed process can be compatible with mini-batch-based training (i.e., using a batch of unlabeled or partially labeled data as a one-time input) for object detection. In this process, the detector, such as a deep neural network, is first applied to the unlabeled samples (i.e., object proposals) to estimate their labels and output the corresponding prediction confidences. Then, our ASM framework is used to select a number of samples and assign pseudolabels to them. These labels are specific to each learning batch based on the confidence levels and additional constraints introduced by the AL process and will be discarded afterward. Then, these temporarily labeled samples are employed for network fine-tuning. In addition, a few samples with low-confidence predictions are selected and annotated via AL. Notably, our method is suitable for object categories that are not seen in the unlabeled data during the learning process. Extensive experiments on two public benchmarks (i.e., the PASCAL VOC 2007/2012 data sets) clearly demonstrate that our ASM framework can achieve performance comparable to that of the alternative methods but with significantly fewer annotations. Keze Wang, Liang Lin 0004, Xiaopeng Yan, Ziliang Chen 0001, Dongyu Zhang 0002, Lei Zhang 0006 |
IEEE Trans. Neural Networks Learn. Syst. | 6 |
| 2018 | A Probabilistic Hierarchical Model for Multi-View and Multi-Feature ClassificationabstractSome recent works in classification show that the data obtained from various views with different sensors for an object contributes to achieving a remarkable performance. Actually, in many real-world applications, each view often contains multiple features, which means that this type of data has a hierarchical structure, while most of existing works do not take these features with multi-layer structure into consideration simultaneously. In this paper, a probabilistic hierarchical model is proposed to address this issue and applied for classification. In our model, a latent variable is first learned to fuse the multiple features obtained from a same view, sensor or modality. Particularly, mapping matrices corresponding to a certain view are estimated to project the latent variable from a shared space to the multiple observations. Since this method is designed for the supervised purpose, we assume that the latent variables associated with different views are influenced by their ground-truth label. In order to effectively solve the proposed method, the Expectation-Maximization (EM) algorithm is applied to estimate the parameters and latent variables. Experimental results on the extensive synthetic and two real-world datasets substantiate the effectiveness and superiority of our approach as compared with state-of-the-art. Jinxing Li 0003, Hongwei Yong, Bob Zhang 0001, Mu Li 0005, Lei Zhang 0006, David Zhang 0001 |
AAAI | 5 |
| 2018 | A PID Controller Approach for Stochastic Optimization of Deep NetworksabstractDeep neural networks have demonstrated their power in many computer vision applications. State-of-the-art deep architectures such as VGG, ResNet, and DenseNet are mostly optimized by the SGD-Momentum algorithm, which updates the weights by considering their past and current gradients. Nonetheless, SGD-Momentum suffers from the overshoot problem, which hinders the convergence of network training. Inspired by the prominent success of proportional-integral-derivative (PID) controller in automatic control, we propose a PID approach for accelerating deep network optimization. We first reveal the intrinsic connections between SGD-Momentum and PID based controller, then present the optimization algorithm which exploits the past, current, and change of gradients to update the network parameters. The proposed PID method reduces much the overshoot phenomena of SGD-Momentum, and it achieves up to 50% acceleration on popular deep network architectures with competitive accuracy, as verified by our experiments on the benchmark datasets including CIFAR10, CIFAR100, and Tiny-ImageNet. Wangpeng An, Haoqian Wang, Qingyun Sun, Jun Xu 0019, Qionghai Dai, Lei Zhang 0006 |
CVPR | 6 |
| 2018 | Learning Spatial-Temporal Regularized Correlation Filters for Visual TrackingabstractDiscriminative Correlation Filters (DCF) are efficient in visual tracking but suffer from unwanted boundary effects. Spatially Regularized DCF (SRDCF) has been suggested to resolve this issue by enforcing spatial penalty on DCF coefficients, which, inevitably, improves the tracking performance at the price of increasing complexity. To tackle online updating, SRDCF formulates its model on multiple training images, further adding difficulties in improving efficiency. In this work, by introducing temporal regularization to SRDCF with single sample, we present our spatialtemporal regularized correlation filters (STRCF). The STRCF formulation can not only serve as a reasonable approximation to SRDCF with multiple training samples, but also provide a more robust appearance model than SRDCF in the case of large appearance variations. Besides, it can be efficiently solved via the alternating direction method of multipliers (ADMM). By incorporating both temporal and spatial regularization, our STRCF can handle boundary effects without much loss in efficiency and achieve superior performance over SRDCF in terms of accuracy and speed. Compared with SRDCF, STRCF with hand-crafted features provides a 5× speedup and achieves a gain of 5.4% and 3.6% AUC score on OTB-2015 and Temple-Color, respectively. Moreover, STRCF with deep features also performs favorably against state-of-the-art trackers and achieves an AUC score of 68.3% on OTB-2015. Feng Li 0031, Wangmeng Zuo, Lei Zhang 0006, Ming-Hsuan Yang 0001 |
CVPR | 4 |
| 2018 | A Hybrid l1-l0 Layer Decomposition Model for Tone MappingabstractTone mapping aims to reproduce a standard dynamic range image from a high dynamic range image with visual information preserved. State-of-the-art tone mapping algorithms mostly decompose an image into a base layer and a detail layer, and process them accordingly. These methods may have problems of halo artifacts and over-enhancement, due to the lack of proper priors imposed on the two layers. In this paper, we propose a hybrid ℓ1-ℓ0decomposition model to address these problems. Specifically, an ℓ1sparsity term is imposed on the base layer to model its piecewise smoothness property. An ℓ0sparsity term is imposed on the detail layer as a structural prior, which leads to piecewise constant effect. We further propose a multiscale tone mapping scheme based on our layer decomposition model. Experiments show that our tone mapping algorithm achieves visually compelling results with little halo artifacts, outperforming the state-of-the-art tone mapping algorithms in both subjective and objective evaluations. Zhetong Liang, Jun Xu 0019, David Zhang 0001, Zisheng Cao, Lei Zhang 0006 |
CVPR | 5 |
| 2018 | Towards Human-Machine Cooperation: Self-Supervised Sample Mining for Object DetectionabstractThough quite challenging, leveraging large-scale unlabeled or partially labeled images in a cost-effective way has increasingly attracted interests for its great importance to computer vision. To tackle this problem, many Active Learning (AL) methods have been developed. However, these methods mainly define their sample selection criteria within a single image context, leading to the suboptimal robustness and impractical solution for large-scale object detection. In this paper, aiming to remedy the drawbacks of existing AL methods, we present a principled Self-supervised Sample Mining (SSM) process accounting for the real challenges in object detection. Specifically, our SSM process concentrates on automatically discovering and pseudo-labeling reliable region proposals for enhancing the object detector via the introduced cross image validation, i.e., pasting these proposals into different labeled images to comprehensively measure their values under different image contexts. By resorting to the SSM process, we propose a new AL framework for gradually incorporating unlabeled or partially labeled data into the model learning while minimizing the annotating effort of users. Extensive experiments on two public benchmarks clearly demonstrate our proposed framework can achieve the comparable performance to the state-of-the-art methods with significantly fewer annotations. Keze Wang, Xiaopeng Yan, Dongyu Zhang 0002, Lei Zhang 0006, Liang Lin 0004 |
CVPR | 4 |
| 2018 | Learning a Single Convolutional Super-Resolution Network for Multiple DegradationsabstractRecent years have witnessed the unprecedented success of deep convolutional neural networks (CNNs) in single image super-resolution (SISR). However, existing CNN-based SISR methods mostly assume that a low-resolution (LR) image is bicubicly downsampled from a high-resolution (HR) image, thus inevitably giving rise to poor performance when the true degradation does not follow this assumption. Moreover, they lack scalability in learning a single model to nonblindly deal with multiple degradations. To address these issues, we propose a general framework with dimensionality stretching strategy that enables a single convolutional super-resolution network to take two key factors of the SISR degradation process, i.e., blur kernel and noise level, as input. Consequently, the super-resolver can handle multiple and even spatially variant degradations, which significantly improves the practicability. Extensive experimental results on synthetic and real LR images show that the proposed convolutional super-resolution network not only can produce favorable results on multiple degradations but also is computationally efficient, providing a highly effective and scalable solution to practical SISR applications. Kai Zhang 0008, Wangmeng Zuo, Lei Zhang 0006 |
CVPR | 3 |
| 2018 | Weakly-Supervised Video Summarization Using Variational Encoder-Decoder and Web Prior
Sijia Cai, Wangmeng Zuo, Larry Davis 0001, Lei Zhang 0006 |
ECCV (14) | 4 |
| 2018 | A Trilateral Weighted Sparse Coding Scheme for Real-World Image Denoising
Jun Xu 0019, Lei Zhang 0006, David Zhang 0001 |
ECCV (8) | 2 |
| 2018 | Deep Image Compression with Iterative Non-Uniform QuantizationabstractImage compression, which aims to represent an image with less storage space, is a classical problem in image processing. Recently, by training an encoder-quantizer-decoder network, deep convolutional neural networks (CNNs) have achieved promising results in image compression. As a nondifferentiable part of the compression system, quantizer is hard to be updated during the network training. Most of existing deep image compression methods adopt a uniform rounding function as the quantizer, which however restricts the capability and flexibility of CNNs in compressing complex image structures. In this paper, we present an iterative nonuniform quantization scheme for deep image compression. More specifically, we alternatively optimize the quantizer and encoder-decoder. When the encoder-decoder is fixed, a non-uniform quantizer is optimized based on the distribution of representation features. The encoder-decoder network is then updated by fixing the quantizer. Extensive experiments demonstrate the superior PSNR index of the proposed method to existing deep compressors and JPEG2000. Jianrui Cai, Lei Zhang 0006 |
ICIP | 2 |
| 2018 | Multi-Exposure Fusion with CNN FeaturesabstractMulti-exposure fusion (MEF) is a widely used approach to high dynamic range imaging. The selection of features for fusion weight calculation is important to the performance of MEF. In this paper, we investigate the effectiveness of convolutional neural network (CNN) features for MEF. Considering the fact that there are no ground-truth images in MEF to train an end-to-end CNN, we adopt the pre-trained networks in other tasks to extract the feature. Both the selection of network and the selection of convolution layer are studied. With the extracted CNN feature map, we compute the local visibility and consistency maps to determine the weight map for MEF. The proposed method works well for both static and dynamic scenes. It exhibits competitive quantitative measures, and presents perceptually pleasing MEF outputs with little halo effects. Hui Li 0029, Lei Zhang 0006 |
ICIP | 2 |
| 2018 | Blind Image Quality Assessment with a Probabilistic Quality RepresentationabstractMost existing blind image quality assessment (BIQA) methods learn a regression model to predict scalar quality scores. Such a scheme ignores the fact that an image will receive divergent subjective scores from different subjects, which cannot be adequately represented by a single scalar number. This is particularly true on complex, real-world distorted images. However, the more informative score distributions are unavailable in existing image quality assessment (IQA) databases and can be potentially noisy when limited number of opinions are collected on each image. This paper proposes a probabilistic quality representation (PQR) and employs a more robust loss function to train deep BIQA models. Using a very straightforward implementation, the proposed method is shown to not only speed up the convergence of deep model training, but also greatly improve the quality prediction accuracy relative to scalar quality score regression methods under the same setting. The source code is available at https://github.com/HuiZeng/BIQA_Toolbox. Hui Zeng 0001, Lei Zhang 0006, Alan C. Bovik |
ICIP | 2 |
| 2018 | Clustering based content and color adaptive tone mapping
Hui Li 0029, Xixi Jia, Lei Zhang 0006 |
Comput. Vis. Image Underst. | 3 |
| 2018 | On Unifying Multi-view Self-Representations for Clustering by Tensor Multi-rank Minimization
Yuan Xie 0006, Dacheng Tao, Wensheng Zhang 0002, Yan Liu 0004, Lei Zhang 0006, Yanyun Qu |
Int. J. Comput. Vis. | 5 |
| 2018 | Bayesian inference for adaptive low rank and sparse matrix estimation
Xixi Jia, Xiangchu Feng, Weiwei Wang 0005, Chen Xu 0004, Lei Zhang 0006 |
Neurocomputing | 5 |
| 2018 | An extended variational image decomposition model for color image enhancement
Xixi Jia, Xiangchu Feng, Weiwei Wang 0005, Lei Zhang 0006 |
Neurocomputing | 4 |
| 2018 | Active Self-Paced Learning for Cost-Effective and Progressive Face IdentificationabstractThis paper aims to develop a novel cost-effective framework for face identification, which progressively maintains a batch of classifiers with the increasing face images of different individuals. By naturally combining two recently rising techniques: active learning (AL) and self-paced learning (SPL), our framework is capable of automatically annotating new instances and incorporating them into training under weak expert recertification. We first initialize the classifier using a few annotated samples for each individual, and extract image features using the convolutional neural nets. Then, a number of candidates are selected from the unannotated samples for classifier updating, in which we apply the current classifiers ranking the samples by the prediction confidence. In particular, our approach utilizes the high-confidence and low-confidence samples in the self-paced and the active user-query way, respectively. The neural nets are later fine-tuned based on the updated classifiers. Such heuristic implementation is formulated as solving a concise active SPL optimization problem, which also advances the SPL development by supplementing a rational dynamic curriculum constraint. The new model finely accords with the "instructor-student-collaborative" learning mode in human education. The advantages of this proposed framework are two-folds: i) The required number of annotated samples is significantly decreased while the comparable performance is guaranteed. A dramatic reduction of user effort is also achieved over other state-of-the-art active learning techniques. ii) The mixture of SPL and AL effectively improves not only the classifier accuracy compared to existing AL/SPL methods but also the robustness against noisy data. We evaluate our framework on two challenging datasets, which include hundreds of persons under diverse conditions, and demonstrate very promising results. Please find the code of this project at: http://hcp.sysu.edu.cn/projects/aspl/. Liang Lin 0004, Keze Wang, Deyu Meng, Wangmeng Zuo, Lei Zhang 0006 |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2018 | Robust Online Matrix Factorization for Dynamic Background SubtractionabstractWe propose an effective online background subtraction method, which can be robustly applied to practical videos that have variations in both foreground and background. Different from previous methods which often model the foreground as Gaussian or Laplacian distributions, we model the foreground for each frame with a specific mixture of Gaussians (MoG) distribution, which is updated online frame by frame. Particularly, our MoG model in each frame is regularized by the learned foreground/background knowledge in previous frames. This makes our online MoG model highly robust, stable and adaptive to practical foreground and background variations. The proposed model can be formulated as a concise probabilistic MAP model, which can be readily solved by EM algorithm. We further embed an affine transformation operator into the proposed model, which can be automatically adjusted to fit a wide range of video background transformations and make the method more robust to camera movements. With using the sub-sampling technique, the proposed method can be accelerated to execute more than 250 frames per second on average, meeting the requirement of real-time background subtraction for practical video processing tasks. The superiority of the proposed method is substantiated by extensive experiments implemented on synthetic and real videos, as compared with state-of-the-art online and offline background subtraction methods. Hongwei Yong, Deyu Meng, Wangmeng Zuo, Lei Zhang 0006 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2018 | Guest Editorial Introduction to the Special Issue on Large Scale and Nonlinear Similarity Learning for Intelligent Video AnalysisabstractLearning similarity and distance measures has become increasingly important for the analysis, matching, retrieval, recognition, and categorization of video and multimedia data. With the ubiquitous use of digital imaging devices, mobile terminals and social networks, there are massive volumes of heterogeneous and homogeneous video and multimedia data from multiple sources, views, and domains, e.g., news media websites, microblog, mobile phone, social networking, etc. Similarity and distance-based constraints can also be extended and incorporated to boost classification and relationship learning. Moreover, the spatio-temporal coherence among video data can also be utilized for self-supervised learning of similarity and distance metrics. This trend has brought several challenging issues for developing similarity and metric learning methods for large scale and weakly annotated data, where outliers and incorrectly annotated data are inevitable. Recently, scalability has been investigated to cope with lightweight and large scale metric learning, while nonlinear similarity models have shown their great potentials in learning invariant representation and nonlinear measures of video and multimedia data. Wangmeng Zuo, Liang Lin 0004, Alan L. Yuille, Horst Bischof, Lei Zhang 0006, Fatih Porikli |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2018 | Learning a Deep Single Image Contrast Enhancer from Multi-Exposure ImagesabstractDue to the poor lighting condition and limited dynamic range of digital imaging devices, the recorded images are often under-/over-exposed and with low contrast. Most of previous single image contrast enhancement (SICE) methods adjust the tone curve to correct the contrast of an input image. Those methods, however, often fail in revealing image details because of the limited information in a single image. On the other hand, the SICE task can be better accomplished if we can learn extra information from appropriately collected training data. In this work, we propose to use the convolutional neural network (CNN) to train a SICE enhancer. One key issue is how to construct a training dataset of low-contrast and high-contrast image pairs for end-to-end CNN learning. To this end, we build a large-scale multi-exposure image dataset, which contains 589 elaborately selected high-resolution multi-exposure sequences with 4,413 images. Thirteen representative multi-exposure image fusion and stack-based high dynamic range imaging algorithms are employed to generate the contrast enhanced images for each sequence, and subjective experiments are conducted to screen the best quality one as the reference image of each scene. With the constructed dataset, a CNN can be easily trained as the SICE enhancer to improve the contrast of an under-/over-exposure image. Experimental results demonstrate the advantages of our method over existing SICE methods with a significant margin. Jianrui Cai, Shuhang Gu, Lei Zhang 0006 |
IEEE Trans. Image Process. | 3 |
| 2018 | Partial Deconvolution With Inaccurate Blur KernelabstractMost non-blind deconvolution methods are developed under the error-free kernel assumption, and are not robust to inaccurate blur kernel. Unfortunately, despite the great progress in blind deconvolution, estimation error remains inevitable during blur kernel estimation. Consequently, severe artifacts such as ringing effects and distortions are likely to be introduced in the non-blind deconvolution stage. In this paper, we tackle this issue by suggesting: 1) a partial map in the Fourier domain for modeling kernel estimation error, and 2) a partial deconvolution model for robust deblurring with inaccurate blur kernel. The partial map is constructed by detecting the reliable Fourier entries of estimated blur kernel. And partial deconvolution is applied to wavelet-based and learning-based models to suppress the adverse effect of kernel estimation error. Furthermore, an E-M algorithm is developed for estimating the partial map and recovering the latent sharp image alternatively. Experimental results show that our partial deconvolution model is effective in relieving artifacts caused by inaccurate blur kernel, and can achieve favorable deblurring quality on synthetic and real blurry images. Dongwei Ren, Wangmeng Zuo, David Zhang 0001, Jun Xu 0019, Lei Zhang 0006 |
IEEE Trans. Image Process. | 5 |
| 2018 | External Prior Guided Internal Prior Learning for Real-World Noisy Image DenoisingabstractMost of existing image denoising methods learn image priors from either external data or the noisy image itself to remove noise. However, priors learned from external data may not be adaptive to the image to be denoised, while priors learned from the given noisy image may not be accurate due to the interference of corrupted noise. Meanwhile, the noise in real-world noisy images is very complex, which is hard to be described by simple distributions such as Gaussian distribution, making real-world noisy image denoising a very challenging problem. We propose to exploit the information in both external data and the given noisy image, and develop an external prior guided internal prior learning method for real-world noisy image denoising. We first learn external priors from an independent set of clean natural images. With the aid of learned external priors, we then learn internal priors from the given noisy image to refine the prior model. The external and internal priors are formulated as a set of orthogonal dictionaries to efficiently reconstruct the desired image. Extensive experiments are performed on several real-world noisy image datasets. The proposed method demonstrates highly competitive denoising performance, outperforming state-of-the-art denoising methods including those designed for real-world noisy images. Jun Xu 0019, Lei Zhang 0006, David Zhang 0001 |
IEEE Trans. Image Process. | 2 |
| 2018 | FFDNet: Toward a Fast and Flexible Solution for CNN-Based Image DenoisingabstractDue to the fast inference and good performance, discriminative learning methods have been widely studied in image denoising. However, these methods mostly learn a specific model for each noise level, and require multiple models for denoising images with different noise levels. They also lack flexibility to deal with spatially variant noise, limiting their applications in practical denoising. To address these issues, we present a fast and flexible denoising convolutional neural network, namely FFDNet, with a tunable noise level map as the input. The proposed FFDNet works on downsampled subimages, achieving a good trade-off between inference speed and denoising performance. In contrast to the existing discriminative denoisers, FFDNet enjoys several desirable properties, including (i) the ability to handle a wide range of noise levels (i.e., [0, 75]) effectively with a single network, (ii) the ability to remove spatially variant noise by specifying a non-uniform noise level map, and (iii) faster speed than benchmark BM3D even on CPU without sacrificing denoising performance. Extensive experiments on synthetic and real noisy images are conducted to evaluate FFDNet in comparison with state-of-the-art denoisers. The results show that FFDNet is effective and efficient, making it highly attractive for practical denoising applications. Kai Zhang 0008, Wangmeng Zuo, Lei Zhang 0006 |
IEEE Trans. Image Process. | 3 |
| 2017 | Learning Dynamic Guidance for Depth Image EnhancementabstractThe depth images acquired by consumer depth sensors (e.g., Kinect and ToF) usually are of low resolution and insufficient quality. One natural solution is to incorporate with high resolution RGB camera for exploiting their statistical correlation. However, most existing methods are intuitive and limited in characterizing the complex and dynamic dependency between intensity and depth images. To address these limitations, we propose a weighted analysis representation model for guided depth image enhancement, which advances the conventional methods in two aspects: (i) task driven learning and (ii) dynamic guidance. First, we generalize the analysis representation model by including a guided weight function for dependency modeling. And the task-driven learning formulation is introduced to obtain the optimized guidance tailored to specific enhancement task. Second, the depth image is gradually enhanced along with the iterations, and thus the guidance should also be dynamically adjusted to account for the updating of depth image. To this end, stage-wise parameters are learned for dynamic guidance. Experiments on guided depth image upsampling and noisy depth image restoration validate the effectiveness of our method. Shuhang Gu, Wangmeng Zuo, Shi Guo, Yunjin Chen, Chongyu Chen, Lei Zhang 0006 |
CVPR | 6 |
| 2017 | G2DeNet: Global Gaussian Distribution Embedding Network and Its Application to Visual RecognitionabstractRecently, plugging trainable structural layers into deep convolutional neural networks (CNNs) as image representations has made promising progress. However, there has been little work on inserting parametric probability distributions, which can effectively model feature statistics, into deep CNNs in an end-to-end manner. This paper proposes a Global Gaussian Distribution embedding Network (G2DeNet) to take a step towards addressing this problem. The core of G2DeNet is a novel trainable layer of a global Gaussian as an image representation plugged into deep CNNs for end-to-end learning. The challenge is that the proposed layer involves Gaussian distributions whose space is not a linear space, which makes its forward and backward propagations be non-intuitive and non-trivial. To tackle this issue, we employ a Gaussian embedding strategy which respects the structures of both Riemannian manifold and smooth group of Gaussians. Based on this strategy, we construct the proposed global Gaussian embedding layer and decompose it into two sub-layers: the matrix partition sub-layer decoupling the mean vector and covariance matrix entangled in the embedding matrix, and the square-rooted, symmetric positive definite matrix sub-layer. In this way, we can derive the partial derivatives associated with the proposed structural layer and thus allow backpropagation of gradients. Experimental results on large scale region classification and fine-grained recognition tasks show that G2DeNet is superior to its counterparts, capable of achieving state-of-the-art performance. Qilong Wang 0001, Peihua Li, Lei Zhang 0006 |
CVPR | 3 |
| 2017 | Learning Deep CNN Denoiser Prior for Image RestorationabstractModel-based optimization methods and discriminative learning methods have been the two dominant strategies for solving various inverse problems in low-level vision. Typically, those two kinds of methods have their respective merits and drawbacks, e.g., model-based optimization methods are flexible for handling different inverse problems but are usually time-consuming with sophisticated priors for the purpose of good performance, in the meanwhile, discriminative learning methods have fast testing speed but their application range is greatly restricted by the specialized task. Recent works have revealed that, with the aid of variable splitting techniques, denoiser prior can be plugged in as a modular part of model-based optimization methods to solve other inverse problems (e.g., deblurring). Such an integration induces considerable advantage when the denoiser is obtained via discriminative learning. However, the study of integration with fast discriminative denoiser prior is still lacking. To this end, this paper aims to train a set of fast and effective CNN (convolutional neural network) denoisers and integrate them into model-based optimization method to solve other inverse problems. Experimental results demonstrate that the learned set of denoisers can not only achieve promising Gaussian denoising results but also can be used as prior to deliver good performance for various low-level vision applications. Kai Zhang 0008, Wangmeng Zuo, Shuhang Gu, Lei Zhang 0006 |
CVPR | 4 |
| 2017 | Higher-Order Integration of Hierarchical Convolutional Activations for Fine-Grained Visual CategorizationabstractThe success of fine-grained visual categorization (FGVC) extremely relies on the modeling of appearance and interactions of various semantic parts. This makes FGVC very challenging because: (i) part annotation and detection require expert guidance and are very expensive; (ii) parts are of different sizes; and (iii) the part interactions are complex and of higher-order. To address these issues, we propose an end-to-end framework based on higherorder integration of hierarchical convolutional activations for FGVC. By treating the convolutional activations as local descriptors, hierarchical convolutional activations can serve as a representation of local parts from different scales. A polynomial kernel based predictor is proposed to capture higher-order statistics of convolutional activations for modeling part interaction. To model inter-layer part interactions, we extend polynomial predictor to integrate hierarchical activations via kernel fusion. Our work also provides a new perspective for combining convolutional activations from multiple layers. While hypercolumns simply concatenate maps from different layers, and holistically-nested network uses weighted fusion to combine side-outputs, our approach exploits higher-order intra-layer and inter-layer relations for better integration of hierarchical convolutional features. The proposed framework yields more discriminative representation and achieves competitive results on the widely used FGVC datasets. Sijia Cai, Wangmeng Zuo, Lei Zhang 0006 |
ICCV | 3 |
| 2017 | Joint Convolutional Analysis and Synthesis Sparse Representation for Single Image Layer SeparationabstractAnalysis sparse representation (ASR) and synthesis sparse representation (SSR) are two representative approaches for sparsity-based image modeling. An image is described mainly by the non-zero coefficients in SSR, while is mainly characterized by the indices of zeros in ASR. To exploit the complementary representation mechanisms of ASR and SSR, we integrate the two models and propose a joint convolutional analysis and synthesis (JCAS) sparse representation model. The convolutional implementation is adopted to more effectively exploit the image global information. In JCAS, a single image is decomposed into two layers, one is approximated by ASR to represent image large-scale structures, and the other by SSR to represent image fine-scale textures. The synthesis dictionary is adaptively learned in JCAS to describe the texture patterns for different single image layer separation tasks. We evaluate the proposed JCAS model on a variety of applications, including rain streak removal, high dynamic range image tone mapping, etc. The results show that our JCAS method outperforms state-of-the-arts in these applications in terms of both quantitative measure and visual perception quality. Shuhang Gu, Deyu Meng, Wangmeng Zuo, Lei Zhang 0006 |
ICCV | 4 |
| 2017 | 3D Surface Detail Enhancement from a Single Normal MapabstractIn 3D reconstruction, the obtained surface details are mainly limited to the visual sensor due to sampling and quantization in the digitalization process. How to get a fine-grained 3D surface with low-cost is still a challenging obstacle in terms of experience, equipment and easyto-obtain. This work introduces a novel framework for enhancing surfaces reconstructed from normal map, where the assumptions on hardware (e.g., photometric stereo setup) and reflection model (e.g., Lambertion reflection) are not necessarily needed. We propose to use a new measure, angle profile, to infer the hidden micro-structure from existing surfaces. In addition, the inferred results are further improved in the domain of discrete geometry processing (DGP) which is able to achieve a stable surface structure under a selectable enhancement setting. Extensive simulation results show that the proposed method obtains significantly improvements over uniform sharpening method in terms of both subjective visual assessment and objective quality metric. Wuyuan Xie, Miaohui Wang, Xianbiao Qi, Lei Zhang 0006 |
ICCV | 4 |
| 2017 | Multi-channel Weighted Nuclear Norm Minimization for Real Color Image DenoisingabstractMost of the existing denoising algorithms are developed for grayscale images. It is not trivial to extend them for color image denoising since the noise statistics in R, G, and B channels can be very different for real noisy images. In this paper, we propose a multi-channel (MC) optimization model for real color image denoising under the weighted nuclear norm minimization (WNNM) framework. We concatenate the RGB patches to make use of the channel redundancy, and introduce a weight matrix to balance the data fidelity of the three channels in consideration of their different noise statistics. The proposed MC-WNNM model does not have an analytical solution. We reformulate it into a linear equality-constrained problem and solve it via alternating direction method of multipliers. Each alternative updating step has a closed-form solution and the convergence can be guaranteed. Experiments on both synthetic and real noisy image datasets demonstrate the superiority of the proposed MC-WNNM over state-of-the-art denoising methods. Jun Xu 0019, Lei Zhang 0006, David Zhang 0001, Xiangchu Feng |
ICCV | 2 |
| 2017 | Part-based convolutional neural network for visual recognitionabstractMid-level element based representations have been proven to be very effective for visual recognition. We present a method to discover discriminative elements based on deep Convolutional Neural Networks (CNNs), namely Part-based CNN (P-CNN), which acts as the role of encoding module in part-based representation. The P-CNN can be attached at arbitrary layer of a pre-trained CNN and be trained using image-level labels. The training of P-CNN essentially corresponds to the optimization and selection of discriminative mid-level visual elements. For an input image, the output of P-CNN is naturally the part-based coding and can be directly used for image recognition. By applying P-CNN to multiple layers of a pretrained CNN, more diverse visual elements can be obtained for visual recognitions. Experiments are conducted on two recognition tasks and their results demonstrate the effectiveness of the proposed method. Lingxiao Yang, Xiaohua Xie, Peihua Li, David Zhang 0001, Lei Zhang 0006 |
ICIP | 5 |
| 2017 | Learning a real-time generic tracker using convolutional neural networksabstractThis paper presents a novel frame-pair based method for visual object tracking. Instead of adopting two-stream Convolutional Neural Networks (CNNs) to represent each frame, we stack frame pairs as the input, resulting in a single-stream CNN tracker with much fewer parameters. The proposed tracker can learn generic motion patterns of objects with much less annotated videos than previous methods. Besides, it is found that trackers trained using two successive frames tend to predict the centers of searching windows as the locations of tracked targets. To alleviate this problem, we propose a novel sampling strategy for off-line training. Specifically, we construct a pair by sampling two frames with a random offset. The offset controls the moving smoothness of objects. Experiments on the challenging VOT14 and OTB datasets show that the proposed tracker performs on par with recently developed generic trackers, but with much less memory. In addition, our tracker can run in a speed of over 100 (30) fps with a GPU (CPU), much faster than most deep neural network based trackers. Linnan Zhu, Lingxiao Yang, David Zhang 0001, Lei Zhang 0006 |
ICME | 4 |
| 2017 | Deep Location-Specific TrackingabstractConvolutional Neural Network (CNN) based methods have shown significant performance gains in the problem of visual tracking in recent years. Due to many uncertain changes of objects online, such as abrupt motion, background clutter and large deformation, the visual tracking is still a challenging task. We propose a novel algorithm, namely Deep Location-Specific Tracking, which decomposes the tracking problem into a localization task and a classification task, and trains an individual network for each task. The localization network exploits the information in the current frame and provides a specific location to improve the probability of successful tracking, while the classification network finds the target among many examples generated around the target location in the previous frame, as well as the one estimated from the localization network in the current frame. CNN based trackers often have massive number of trainable parameters, and are prone to over-fitting to some particular object states, leading to less precision or tracking drift. We address this problem by learning a classification network based on 1 × 1 convolution and global average pooling. Extensive experimental results on popular benchmark datasets show that the proposed tracker achieves competitive results without using additional tracking videos for fine-tuning. The code is available at https://github.com/ZjjConan/DLST Lingxiao Yang, Risheng Liu, David Zhang 0001, Lei Zhang 0006 |
ACM Multimedia | 4 |
| 2017 | Weighted Nuclear Norm Minimization and Its Applications to Low Level Vision
Shuhang Gu, Qi Xie 0002, Deyu Meng, Wangmeng Zuo, Xiangchu Feng, Lei Zhang 0006 |
Int. J. Comput. Vis. | 6 |
| 2017 | Joint Image Denoising and Disparity Estimation via Stereo Structure PCA and Noise-Tolerant Cost
Jianbo Jiao, Qingxiong Yang, Shengfeng He, Shuhang Gu, Lei Zhang 0006, Rynson W. H. Lau |
Int. J. Comput. Vis. | 5 |
| 2017 | Perception-based adaptive quantization for transform-domain Wyner-Ziv video coding
Lei Zhang 0006, Qiang Peng, Xiao Wu 0001 |
Multim. Tools Appl. | 1 |
| 2017 | Local Log-Euclidean Multivariate Gaussian Descriptor and Its Application to Image ClassificationabstractThis paper presents a novel image descriptor to effectively characterize the local, high-order image statistics. Our work is inspired by the Diffusion Tensor Imaging and the structure tensor method (or covariance descriptor), and motivated by popular distribution-based descriptors such as SIFT and HoG. Our idea is to associate one pixel with a multivariate Gaussian distribution estimated in the neighborhood. The challenge lies in that the space of Gaussians is not a linear space but a Riemannian manifold. We show, for the first time to our knowledge, that the space of Gaussians can be equipped with a Lie group structure by defining a multiplication operation on this manifold, and that it is isomorphic to a subgroup of the upper triangular matrix group. Furthermore, we propose methods to embed this matrix group in the linear space, which enables us to handle Gaussians with Euclidean operations rather than complicated Riemannian operations. The resulting descriptor, called Local Log-Euclidean Multivariate Gaussian (L2EMG) descriptor, works well with low-dimensional and high-dimensional raw features. Moreover, our descriptor is a continuous function of features without quantization, which can model the first- and second-order statistics. Extensive experiments were conducted to evaluate thoroughly L2EMG, and the results showed that L2EMG is very competitive with state-of-the-art descriptors in image classification. Peihua Li, Qilong Wang 0001, Hui Zeng 0001, Lei Zhang 0006 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2017 | An Efficient Globally Optimal Algorithm for Asymmetric Point MatchingabstractAlthough the robust point matching algorithm has been demonstrated to be effective for non-rigid registration, there are several issues with the adopted deterministic annealing optimization technique. First, it is not globally optimal and regularization on the spatial transformation is needed for good matching results. Second, it tends to align the mass centers of two point sets. To address these issues, we propose a globally optimal algorithm for the robust point matching problem in the case that each model point has a counterpart in scene set. By eliminating the transformation variables, we show that the original matching problem is reduced to a concave quadratic assignment problem where the objective function has a low rank Hessian matrix. This facilitates the use of large scale global optimization techniques. We propose a modified normal rectangular branch-and-bound algorithm to solve the resulting problem where multiple rectangles are simultaneously subdivided to increase the chance of shrinking the rectangle containing the global optimal solution. In addition, we present an efficient lower bounding scheme which has a linear assignment formulation and can be efficiently solved. Extensive experiments on synthetic and real datasets demonstrate the proposed algorithm performs favorably against the state-of-the-art methods in terms of robustness to outliers, matching accuracy, and run-time. Wei Lian, Lei Zhang 0006, Ming-Hsuan Yang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2017 | Cross-Domain Visual Matching via Generalized Similarity Measure and Feature LearningabstractCross-domain visual data matching is one of the fundamental problems in many real-world vision tasks, e.g., matching persons across ID photos and surveillance videos. Conventional approaches to this problem usually involves two steps: i) projecting samples from different domains into a common space, and ii) computing (dis-)similarity in this space based on a certain distance. In this paper, we present a novel pairwise similarity measure that advances existing models by i) expanding traditional linear projections into affine transformations and ii) fusing affine Mahalanobis distance and Cosine similarity by a data-driven combination. Moreover, we unify our similarity measure with feature representation learning via deep convolutional neural networks. Specifically, we incorporate the similarity measure matrix into the deep architecture, enabling an end-to-end way of model optimization. We extensively evaluate our generalized similarity model in several challenging cross-domain matching tasks: person re-identification under different views and face verification over different modalities (i.e., faces from still images and videos, older and younger faces, and sketch and photo portraits). The experimental results demonstrate superior performance of our model over other state-of-the-art methods. Liang Lin 0004, Guangrun Wang, Wangmeng Zuo, Xiangchu Feng, Lei Zhang 0006 |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2017 | Evaluation of Segmentation Quality via Adaptive Composition of Reference SegmentationsabstractEvaluating image segmentation quality is a critical step for generating desirable segmented output and comparing performance of algorithms, among others. However, automatic evaluation of segmented results is inherently challenging since image segmentation is an ill-posed problem. This paper presents a framework to evaluate segmentation quality using multiple labeled segmentations which are considered as references. For a segmentation to be evaluated, we adaptively compose a reference segmentation using multiple labeled segmentations, which locally matches the input segments while preserving structural consistency. The quality of a given segmentation is then measured by its distance to the composed reference. A new dataset of 200 images, where each one has 6 to 15 labeled segmentations, is developed for performance evaluation of image segmentation. Furthermore, to quantitatively compare the proposed segmentation evaluation algorithm with the state-of-the-art methods, a benchmark segmentation evaluation dataset is proposed. Extensive experiments are carried out to validate the proposed segmentation evaluation framework. Bo Peng 0006, Lei Zhang 0006, Xuanqin Mou, Ming-Hsuan Yang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2017 | High-Order Local Pooling and Encoding Gaussians Over a Dictionary of GaussiansabstractLocal pooling (LP) in configuration (feature) space proposed by Boureau et al. explicitly restricts similar features to be aggregated, which can preserve as much discriminative information as possible. At the time it appeared, this method combined with sparse coding achieved competitive classification results with only a small dictionary. However, its performance lags far behind the state-of-the-art results as only the zero-order information is exploited. Inspired by the success of high-order statistical information in existing advanced feature coding or pooling methods, we make an attempt to address the limitation of LP. To this end, we present a novel method called high-order LP (HO-LP) to leverage the information higher than the zero-order one. Our idea is intuitively simple: we compute the first- and second-order statistics per configuration bin and model them as a Gaussian. Accordingly, we employ a collection of Gaussians as visual words to represent the universal probability distribution of features from all classes. Our problem is naturally formulated as encoding Gaussians over a dictionary of Gaussians as visual words. This problem, however, is challenging since the space of Gaussians is not a Euclidean space but forms a Riemannian manifold. We address this challenge by mapping Gaussians into the Euclidean space, which enables us to perform coding with common Euclidean operations rather than complex and often expensive Riemannian operations. Our HO-LP preserves the advantages of the original LP: pooling only similar features and using a small dictionary. Meanwhile, it achieves very promising performance on standard benchmarks, with either conventional, hand-engineered features or deep learning-based features. Peihua Li, Hui Zeng 0001, Qilong Wang 0001, Simon C. K. Shiu, Lei Zhang 0006 |
IEEE Trans. Image Process. | 5 |
| 2017 | Waterloo Exploration Database: New Challenges for Image Quality Assessment ModelsabstractThe great content diversity of real-world digital images poses a grand challenge to image quality assessment (IQA) models, which are traditionally designed and validated on a handful of commonly used IQA databases with very limited content variation. To test the generalization capability and to facilitate the wide usage of IQA techniques in real-world applications, we establish a large-scale database named the Waterloo Exploration Database, which in its current state contains 4744 pristine natural images and 94 880 distorted images created from them. Instead of collecting the mean opinion score for each image via subjective testing, which is extremely difficult if not impossible, we present three alternative test criteria to evaluate the performance of IQA models, namely, the pristine/distorted image discriminability test, the listwise ranking consistency test, and the pairwise preference consistency test (P-test). We compare 20 well-known IQA models using the proposed criteria, which not only provide a stronger test in a more challenging testing environment for existing models, but also demonstrate the additional benefits of using the proposed database. For example, in the P-test, even for the best performing no-reference IQA model, more than 6 million failure cases against the model are "discovered" automatically out of over 1 billion test pairs. Furthermore, we discuss how the new database may be exploited using innovative approaches in the future, to reveal the weaknesses of existing IQA models, to provide insights on how to improve the models, and to shed light on how the next-generation IQA models may be developed. The database and codes are made publicly available at: https://ece.uwaterloo.ca/~k29ma/exploration/. Kede Ma, Zhengfang Duanmu, Qingbo Wu 0001, Zhou Wang 0001, Hongwei Yong, Hongliang Li 0001, Lei Zhang 0006 |
IEEE Trans. Image Process. | 7 |
| 2017 | Robust Multi-Exposure Image Fusion: A Structural Patch Decomposition ApproachabstractWe propose a simple yet effective structural patch decomposition approach for multi-exposure image fusion (MEF) that is robust to ghosting effect. We decompose an image patch into three conceptually independent components: signal strength, signal structure, and mean intensity. Upon fusing these three components separately, we reconstruct a desired patch and place it back into the fused image. This novel patch decomposition approach benefits MEF in many aspects. First, as opposed to most pixel-wise MEF methods, the proposed algorithm does not require post-processing steps to improve visual quality or to reduce spatial artifacts. Second, it handles RGB color channels jointly, and thus produces fused images with more vivid color appearance. Third and most importantly, the direction of the signal structure component in the patch vector space provides ideal information for ghost removal. It allows us to reliably and efficiently reject inconsistent object motions with respect to a chosen reference image without performing computationally expensive motion estimation. We compare the proposed algorithm with 12 MEF methods on 21 static scenes and 12 deghosting schemes on 19 dynamic scenes (with camera and object motion). Extensive experimental results demonstrate that the proposed algorithm not only outperforms previous MEF algorithms on static scenes but also consistently produces high quality fused images with little ghosting artifacts for dynamic scenes. Moreover, it maintains a lower computational cost compared with the state-of-the-art deghosting schemes. Kede Ma, Hui Li 0029, Hongwei Yong, Zhou Wang 0001, Deyu Meng, Lei Zhang 0006 |
IEEE Trans. Image Process. | 6 |
| 2017 | Beyond a Gaussian Denoiser: Residual Learning of Deep CNN for Image DenoisingabstractThe discriminative model learning for image denoising has been recently attracting considerable attentions due to its favorable denoising performance. In this paper, we take one step forward by investigating the construction of feed-forward denoising convolutional neural networks (DnCNNs) to embrace the progress in very deep architecture, learning algorithm, and regularization method into image denoising. Specifically, residual learning and batch normalization are utilized to speed up the training process as well as boost the denoising performance. Different from the existing discriminative denoising models which usually train a specific model for additive white Gaussian noise at a certain noise level, our DnCNN model is able to handle Gaussian denoising with unknown noise level (i.e., blind Gaussian denoising). With the residual learning strategy, DnCNN implicitly removes the latent clean image in the hidden layers. This property motivates us to train a single DnCNN model to tackle with several general image denoising tasks, such as Gaussian denoising, single image super-resolution, and JPEG image deblocking. Our extensive experiments demonstrate that our DnCNN model can not only exhibit high effectiveness in several general image denoising tasks, but also be efficiently implemented by benefiting from GPU computing. Kai Zhang 0008, Wangmeng Zuo, Yunjin Chen, Deyu Meng, Lei Zhang 0006 |
IEEE Trans. Image Process. | 5 |
| 2017 | Distance Metric Learning via Iterated Support Vector MachinesabstractDistance metric learning aims to learn from the given training data a valid distance metric, with which the similarity between data samples can be more effectively evaluated for classification. Metric learning is often formulated as a convex or nonconvex optimization problem, while most existing methods are based on customized optimizers and become inefficient for large scale problems. In this paper, we formulate metric learning as a kernel classification problem with the positive semi-definite constraint, and solve it by iterated training of support vector machines (SVMs). The new formulation is easy to implement and efficient in training with the off-the-shelf SVM solvers. Two novel metric learning models, namely positive-semidefinite constrained metric learning (PCML) and nonnegative-coefficient constrained metric learning (NCML), are developed. Both PCML and NCML can guarantee the global optimality of their solutions. Experiments are conducted on general classification, face verification, and person re-identification to evaluate our methods. Compared with the state-of-the-art approaches, our methods can achieve comparable classification accuracy and are efficient in training. Wangmeng Zuo, David Zhang 0001, Liang Lin 0004, Yuchi Huang, Deyu Meng, Lei Zhang 0006 |
IEEE Trans. Image Process. | 7 |
| 2016 | A Probabilistic Collaborative Representation Based Approach for Pattern ClassificationabstractConventional representation based classifiers, ranging from the classical nearest neighbor classifier and nearest subspace classifier to the recently developed sparse representation based classifier (SRC) and collaborative representation based classifier (CRC), are essentially distance based classifiers. Though SRC and CRC have shown interesting classification results, their intrinsic classification mechanism remains unclear. In this paper we propose a probabilistic collaborative representation framework, where the probability that a test sample belongs to the collaborative subspace of all classes can be well defined and computed. Consequently, we present a probabilistic collaborative representation based classifier (ProCRC), which jointly maximizes the likelihood that a test sample belongs to each of the multiple classes. The final classification is performed by checking which class has the maximum likelihood. The proposed ProCRC has a clear probabilistic interpretation, and it shows superior performance to many popular classifiers, including SRC, CRC and SVM. Coupled with the CNN features, it also leads to state-of-the-art classification results on a variety of challenging visual datasets. Sijia Cai, Lei Zhang 0006, Wangmeng Zuo, Xiangchu Feng |
CVPR | 2 |
| 2016 | Group MAD Competition? A New Methodology to Compare Objective Image Quality ModelsabstractObjective image quality assessment (IQA) models aim to automatically predict human visual perception of image quality and are of fundamental importance in the field of image processing and computer vision. With an increasing number of IQA models proposed, how to fairly compare their performance becomes a major challenge due to the enormous size of image space and the limited resource for subjective testing. The standard approach in literature is to compute several correlation metrics between subjective mean opinion scores (MOSs) and objective model predictions on several well-known subject-rated databases that contain distorted images generated from a few dozens of source images, which however provide an extremely limited representation of real-world images. Moreover, most IQA models developed on these databases often involve machine learning and/or manual parameter tuning steps to boost their performance, and thus their generalization capabilities are questionable. Here we propose a novel methodology to compare IQA models. We first build a database that contains 4,744 source natural images, together with 94,880 distorted images created from them. We then propose a new mechanism, namely group MAximum Differentiation (gMAD) competition, which automatically selects subsets of image pairs from the database that provide the strongest test to let the IQA models compete with each other. Subjective testing on the selected subsets reveals the relative performance of the IQA models and provides useful insights on potential ways to improve them. We report the gMAD competition results between 16 well-known IQA models, but the framework is extendable, allowing future IQA models to be added into the competition. Kede Ma, Qingbo Wu 0001, Zhou Wang 0001, Zhengfang Duanmu, Hongwei Yong, Hongliang Li 0001, Lei Zhang 0006 |
CVPR | 7 |
| 2016 | Object Tracking via Dual Linear Structured SVM and Explicit Feature MapabstractStructured support vector machine (SSVM) based methods have demonstrated encouraging performance in recent object tracking benchmarks. However, the complex and expensive optimization limits their deployment in real-world applications. In this paper, we present a simple yet efficient dual linear SSVM (DLSSVM) algorithm to enable fast learning and execution during tracking. By analyzing the dual variables, we propose a primal classifier update formula where the learning step size is computed in closed form. This online learning method significantly improves the robustness of the proposed linear SSVM with lower computational cost. Second, we approximate the intersection kernel for feature representations with an explicit feature map to further improve tracking performance. Finally, we extend the proposed DLSSVM tracker with multi-scale estimation to address the "drift" problem. Experimental results on large benchmark datasets with 50 and 100 video sequences show that the proposed DLSSVM tracking algorithm achieves state-of-the-art performance. Jifeng Ning, Jimei Yang, Shaojie Jiang, Lei Zhang 0006, Ming-Hsuan Yang 0001 |
CVPR | 4 |
| 2016 | Dictionary Pair Classifier Driven Convolutional Neural Networks for Object DetectionabstractFeature representation and object category classification are two key components of most object detection methods. While significant improvements have been achieved for deep feature representation learning, traditional SVM/softmax classifiers remain the dominant methods for the final object category classification. However, SVM/softmax classifiers lack the capacity of explicitly exploiting the complex structure of deep features, as they are purely discriminative methods. The recently proposed discriminative dictionary pair learning (DPL) model involves a fidelity term to minimize the reconstruction loss and a discrimination term to enhance the discriminative capability of the learned dictionary pair, and thus is appropriate for balancing the representation and discrimination to boost object detection performance. In this paper, we propose a novel object detection system by unifying DPL with the convolutional feature learning. Specifically, we incorporate DPL as a Dictionary Pair Classifier Layer (DPCL) into the deep architecture, and develop an end-to-end learning algorithm for optimizing the dictionary pairs and the neural networks simultaneously. Moreover, we design a multi-task loss for guiding our model to accomplish the three correlated tasks: objectness estimation, categoryness computation, and bounding box regression. From the extensive experiments on PASCAL VOC 2007/2012 benchmarks, our approach demonstrates the effectiveness to substantially improve the performances over the popular existing object detection frameworks (e.g., R-CNN [13] and FRCN [12]), and achieves new state-of-the-arts. Keze Wang, Liang Lin 0004, Wangmeng Zuo, Shuhang Gu, Lei Zhang 0006 |
CVPR | 5 |
| 2016 | RAID-G: Robust Estimation of Approximate Infinite Dimensional Gaussian with Application to Material RecognitionabstractInfinite dimensional covariance descriptors can provide richer and more discriminative information than their low dimensional counterparts. In this paper, we propose a novel image descriptor, namely, robust approximate infinite dimensional Gaussian (RAID-G). The challenges of RAID-G mainly lie on two aspects: (1) description of infinite dimensional Gaussian is difficult due to its non-linear Riemannian geometric structure and the infinite dimensional setting, hence effective approximation is necessary, (2) traditional maximum likelihood estimation (MLE) is not robust to high (even infinite) dimensional covariance matrix in Gaussian setting. To address these challenges, explicit feature mapping (EFM) is first introduced for effective approximation of infinite dimensional Gaussian induced by additive kernel function, and then a new regularized MLE method based on von Neumann divergence is proposed for robust estimation of covariance matrix. The EFM and proposed regularized MLE allow a closed-form of RAID-G, which is very efficient and effective for high dimensional features. We extend RAID-G by using the outputs of deep convolutional neural networks as original features, and apply it to material recognition. Our approach is evaluated on five material benchmarks and one fine-grained benchmark. It achieves 84.9% accuracy on FMD and 86.3% accuracy on UIUC material database, which are much higher than state-of-the-arts. Qilong Wang 0001, Peihua Li, Wangmeng Zuo, Lei Zhang 0006 |
CVPR | 4 |
| 2016 | Joint Learning of Single-Image and Cross-Image Representations for Person Re-identificationabstractPerson re-identification has been usually solved as either the matching of single-image representation (SIR) or the classification of cross-image representation (CIR). In this work, we exploit the connection between these two categories of methods, and propose a joint learning frame-work to unify SIR and CIR using convolutional neural network (CNN). Specifically, our deep architecture contains one shared sub-network together with two sub-networks that extract the SIRs of given images and the CIRs of given image pairs, respectively. The SIR sub-network is required to be computed once for each image (in both the probe and gallery sets), and the depth of the CIR sub-network is required to be minimal to reduce computational burden. Therefore, the two types of representation can be jointly optimized for pursuing better matching accuracy with moderate computational cost. Furthermore, the representations learned with pairwise comparison and triplet comparison objectives can be combined to improve matching performance. Experiments on the CUHK03, CUHK01 and VIPeR datasets show that the proposed method can achieve favorable accuracy while compared with state-of-the-arts. Wangmeng Zuo, Liang Lin 0004, David Zhang 0001, Lei Zhang 0006 |
CVPR | 5 |
| 2016 | Multispectral Images Denoising by Intrinsic Tensor Sparsity RegularizationabstractMultispectral images (MSI) can help deliver more faithful representation for real scenes than the traditional image system, and enhance the performance of many computer vision tasks. In real cases, however, an MSI is always corrupted by various noises. In this paper, we propose a new tensor-based denoising approach by fully considering two intrinsic characteristics underlying an MSI, i.e., the global correlation along spectrum (GCS) and nonlocal self-similarity across space (NSS). In specific, we construct a new tensor sparsity measure, called intrinsic tensor sparsity (ITS) measure, which encodes both sparsity insights delivered by the most typical Tucker and CANDECOMP/ PARAFAC (CP) low-rank decomposition for a general tensor. Then we build a new MSI denoising model by applying the proposed ITS measure on tensors formed by non-local similar patches within the MSI. The intrinsic GCS and NSS knowledge can then be efficiently explored under the regularization of this tensor sparsity measure to finely rectify the recovery of a MSI from its corruption. A series of experiments on simulated and real MSI denoising problems show that our method outperforms all state-of-the-arts under comprehensive quantitative performance measures. Qi Xie 0002, Qian Zhao 0002, Deyu Meng, Zongben Xu, Shuhang Gu, Wangmeng Zuo, Lei Zhang 0006 |
CVPR | 7 |
| 2016 | Learning a lightweight deep convolutional network for joint age and gender recognitionabstractThis paper proposes a lightweight deep model to recognize age and gender from a face image. Though simple, our network architecture is able to complete the two tasks effectively and efficiently. Moreover, different from existing methods, we simultaneously perform the age and gender recognition tasks via a joint regression model. Specifically, our model employs a multi-task learning scheme to learn shared features for these two correlated tasks in an end-to-end manner. Extensive experimental results on the recent Adience benchmark demonstrate that our model achieves competitive recognition accuracy with the state-of-the-art methods but with much faster speed, i.e., about 10 times faster in the testing phase. Our model can be easily adopted and extended to other facial applications. Linnan Zhu, Keze Wang, Liang Lin 0004, Lei Zhang 0006 |
ICPR | 4 |
| 2016 | A Self-Representation Induced Classifier
Pengfei Zhu 0001, Lei Zhang 0006, Wangmeng Zuo, Xiangchu Feng, Qinghua Hu |
IJCAI | 2 |
| 2016 | A Deep Structured Model with Radius-Margin Bound for 3D Human Activity Recognition
Liang Lin 0004, Keze Wang, Wangmeng Zuo, Meng Wang 0001, Jiebo Luo 0001, Lei Zhang 0006 |
Int. J. Comput. Vis. | 6 |
| 2016 | Smart computing for large scale visual data sensing and processing
Lei Zhang 0006, Pinar Duygulu, Wangmeng Zuo, Shiguang Shan, Alex Hauptmann 0001 |
Neurocomputing | 1 |
| 2016 | Evaluation of ground distances and features in EMD-based GMM matching for texture classification
Hua Hao, Qilong Wang 0001, Peihua Li, Lei Zhang 0006 |
Pattern Recognit. | 4 |
| 2016 | Towards effective codebookless model for image classification
Qilong Wang 0001, Peihua Li, Lei Zhang 0006, Wangmeng Zuo |
Pattern Recognit. | 3 |
| 2016 | Joint Learning of Multiple Regressors for Single Image Super-ResolutionabstractUsing a global regression model for single image super-resolution (SISR) generally fails to produce visually pleasant output. The recently developed local learning methods provide a remedy by partitioning the feature space into a number of clusters and learning a simple local model for each cluster. However, in these methods the space partition is conducted separately from local model learning, which results in an abundant number of local models to achieve satisfying performance. To address this problem, we propose a mixture of experts (MoE) method to jointly learn the feature space partition and local regression models. Our MoE consists of two components: gating network learning and local regressors learning. An expectation-maximization (EM) algorithm is adopted to train MoE on a large set of LR/HR patch pairs. Experimental results demonstrate that the proposed method can use much less local models and time to achieve comparable or superior results to state-of-the-art SISR methods, providing a highly practical solution to real applications. Kai Zhang 0008, Baoquan Wang, Wangmeng Zuo, Lei Zhang 0006 |
IEEE Signal Process. Lett. | 5 |
| 2016 | A Level Set Approach to Image Segmentation With Intensity InhomogeneityabstractIt is often a difficult task to accurately segment images with intensity inhomogeneity, because most of representative algorithms are region-based that depend on intensity homogeneity of the interested object. In this paper, we present a novel level set method for image segmentation in the presence of intensity inhomogeneity. The inhomogeneous objects are modeled as Gaussian distributions of different means and variances in which a sliding window is used to map the original image into another domain, where the intensity distribution of each object is still Gaussian but better separated. The means of the Gaussian distributions in the transformed domain can be adaptively estimated by multiplying a bias field with the original signal within the window. A maximum likelihood energy functional is then defined on the whole image region, which combines the bias field, the level set function, and the piecewise constant function approximating the true image signal. The proposed level set method can be directly applied to simultaneous segmentation and bias correction for 3 and 7T magnetic resonance images. Extensive evaluation on synthetic and real-images demonstrate the superiority of the proposed method over other representative algorithms. Kaihua Zhang 0001, Lei Zhang 0006, Kin-Man Lam 0001, David Zhang 0001 |
IEEE Trans. Cybern. | 2 |
| 2016 | Detail-Preserving and Content-Aware Variational Multi-View Stereo ReconstructionabstractAccurate recovery of 3D geometrical surfaces from calibrated 2D multi-view images is a fundamental yet active research area in computer vision. Despite the steady progress in multi-view stereo (MVS) reconstruction, many existing methods are still limited in recovering fine-scale details and sharp features while suppressing noises, and may fail in reconstructing regions with less textures. To address these limitations, this paper presents a detail-preserving and content-aware variational (DCV) MVS method, which reconstructs the 3D surface by alternating between reprojection error minimization and mesh denoising. In reprojection error minimization, we propose a novel inter-image similarity measure, which is effective to preserve fine-scale details of the reconstructed surface and builds a connection between guided image filtering and image registration. In mesh denoising, we propose a content-aware ℓp-minimization algorithm by adaptively estimating the p value and regularization parameters. Compared with conventional isotropic mesh smoothing approaches, the proposed method is much more promising in suppressing noise while preserving sharp features. Experimental results on benchmark data sets demonstrate that our DCV method is capable of recovering more surface details, and obtains cleaner and more accurate reconstructions than the state-of-the-art methods. In particular, our method achieves the best results among all published methods on the Middlebury dino ring and dino sparse data sets in terms of both completeness and accuracy. Zhaoxin Li, Kuanquan Wang, Wangmeng Zuo, Deyu Meng, Lei Zhang 0006 |
IEEE Trans. Image Process. | 5 |
| 2016 | Weighted Schatten p-Norm Minimization for Image Denoising and Background SubtractionabstractLow rank matrix approximation (LRMA), which aims to recover the underlying low rank matrix from its degraded observation, has a wide range of applications in computer vision. The latest LRMA methods resort to using the nuclear norm minimization (NNM) as a convex relaxation of the nonconvex rank minimization. However, NNM tends to over-shrink the rank components and treats the different rank components equally, limiting its flexibility in practical applications. We propose a more flexible model, namely, the weighted Schatten p-norm minimization (WSNM), to generalize the NNM to the Schatten p-norm minimization with weights assigned to different singular values. The proposed WSNM not only gives better approximation to the original low-rank assumption, but also considers the importance of different rank components. We analyze the solution of WSNM and prove that, under certain weights permutation, WSNM can be equivalently transformed into independent non-convex lp-norm subproblems, whose global optimum can be efficiently solved by generalized iterated shrinkage algorithm. We apply WSNM to typical low-level vision problems, e.g., image denoising and background subtraction. Extensive experimental results show, both qualitatively and quantitatively, that the proposed WSNM can more effectively remove noise, and model the complex and dynamic scenes compared with state-of-the-art methods. Yuan Xie 0006, Shuhang Gu, Yan Liu 0004, Wangmeng Zuo, Wensheng Zhang 0002, Lei Zhang 0006 |
IEEE Trans. Image Process. | 6 |
| 2016 | Learning Iteration-wise Generalized Shrinkage-Thresholding Operators for Blind DeconvolutionabstractSalient edge selection and time-varying regularization are two crucial techniques to guarantee the success of maximum a posteriori (MAP)-based blind deconvolution. However, the existing approaches usually rely on carefully designed regularizers and handcrafted parameter tuning to obtain satisfactory estimation of the blur kernel. Many regularizers exhibit the structure-preserving smoothing capability, but fail to enhance salient edges. In this paper, under the MAP framework, we propose the iteration-wise ℓp-norm regularizers together with data-driven strategy to address these issues. First, we extend the generalized shrinkage-thresholding (GST) operator for ℓp-norm minimization with negative p value, which can sharpen salient edges while suppressing trivial details. Then, the iteration-wise GST parameters are specified to allow dynamical salient edge selection and time-varying regularization. Finally, instead of handcrafted tuning, a principled discriminative learning approach is proposed to learn the iterationwise GST operators from the training dataset. Furthermore, the multi-scale scheme is developed to improve the efficiency of the algorithm. Experimental results show that, negative p value is more effective in estimating the coarse shape of blur kernel at the early stage, and the learned GST operators can be well generalized to other dataset and real world blurry images. Compared with the state-of-the-art methods, our method achieves better deblurring results in terms of both quantitative metrics and visual quality, and it is much faster than the state-of-the-art patch-based blind deconvolution method. Wangmeng Zuo, Dongwei Ren, David Zhang 0001, Shuhang Gu, Lei Zhang 0006 |
IEEE Trans. Image Process. | 5 |
| 2015 | Discriminative learning of iteration-wise priors for blind deconvolutionabstractThe maximum a posterior (MAP)-based blind deconvolution framework generally involves two stages: blur kernel estimation and non-blind restoration. For blur kernel estimation, sharp edge prediction and carefully designed image priors are vital to the success of MAP. In this paper, we propose a blind deconvolution framework together with iteration specific priors for better blur kernel estimation. The family of hyper-Laplacian (Pr(d) ∝ e-∥d∥pp/λ) is adopted for modeling iteration-wise priors of image gra- dients, where each iteration has its own model parameters {λ(t), p(t)}. To avoid heavy parameter tuning, all iteration-wise model parameters can be learned using our principled discriminative learning model from a training set, and can be directly applied to other dataset and real blurry images. Interestingly, with the generalized shrinkage / thresholding operator, negative p value (p <;0) is allowable and we find that it contributes more in estimating the coarse shape of blur kernel. Experimental results on synthetic and real world images demonstrate that our method achieves better deblurring results than the existing gradient prior-based methods. Compared with the state-of-the-art patch prior-based method, our method is competitive in restoration results but is much more efficient. Wangmeng Zuo, Dongwei Ren, Shuhang Gu, Liang Lin 0004, Lei Zhang 0006 |
CVPR | 5 |
| 2015 | External Patch Prior Guided Internal Clustering for Image DenoisingabstractNatural image modeling plays a key role in many vision problems such as image denoising. Image priors are widely used to regularize the denoising process, which is an ill-posed inverse problem. One category of denoising methods exploit the priors (e.g., TV, sparsity) learned from external clean images to reconstruct the given noisy image, while another category of methods exploit the internal prior (e.g., self-similarity) to reconstruct the latent image. Though the internal prior based methods have achieved impressive denoising results, the improvement of visual quality will become very difficult with the increase of noise level. In this paper, we propose to exploit image external patch prior and internal self-similarity prior jointly, and develop an external patch prior guided internal clustering algorithm for image denoising. It is known that natural image patches form multiple subspaces. By utilizing Gaussian mixture models (GMMs) learning, image similar patches can be clustered and the subspaces can be learned. The learned GMMs from clean images are then used to guide the clustering of noisy-patches of the input noisy images, followed by a low-rank approximation process to estimate the latent subspace for image recovery. Numerical experiments show that the proposed method outperforms many state-of-the-art denoising algorithms such as BM3D and WNNM. Fei Chen 0012, Lei Zhang 0006 |
ICCV | 2 |
| 2015 | Convolutional Sparse Coding for Image Super-ResolutionabstractMost of the previous sparse coding (SC) based super resolution (SR) methods partition the image into overlapped patches, and process each patch separately. These methods, however, ignore the consistency of pixels in overlapped patches, which is a strong constraint for image reconstruction. In this paper, we propose a convolutional sparse coding (CSC) based SR (CSC-SR) method to address the consistency issue. Our CSC-SR involves three groups of parameters to be learned: (i) a set of filters to decompose the low resolution (LR) image into LR sparse feature maps, (ii) a mapping function to predict the high resolution (HR) feature maps from the LR ones, and (iii) a set of filters to reconstruct the HR images from the predicted HR feature maps via simple convolution operations. By working directly on the whole image, the proposed CSC-SR algorithm does not need to divide the image into overlapped patches, and can exploit the image global correlation to produce more robust reconstruction of image local structures. Experimental results clearly validate the advantages of CSC over patch based SC in SR application. Compared with state-of-the-art SR methods, the proposed CSC-SR method achieves highly competitive PSNR results, while demonstrating better edge and texture preservation performance. Shuhang Gu, Wangmeng Zuo, Qi Xie 0002, Deyu Meng, Xiangchu Feng, Lei Zhang 0006 |
ICCV | 6 |
| 2015 | Patch Group Based Nonlocal Self-Similarity Prior Learning for Image DenoisingabstractPatch based image modeling has achieved a great success in low level vision such as image denoising. In particular, the use of image nonlocal self-similarity (NSS) prior, which refers to the fact that a local patch often has many nonlocal similar patches to it across the image, has significantly enhanced the denoising performance. However, in most existing methods only the NSS of input degraded image is exploited, while how to utilize the NSS of clean natural images is still an open problem. In this paper, we propose a patch group (PG) based NSS prior learning scheme to learn explicit NSS models from natural images for high performance denoising. PGs are extracted from training images by putting nonlocal similar patches into groups, and a PG based Gaussian Mixture Model (PG-GMM) learning algorithm is developed to learn the NSS prior. We demonstrate that, owe to the learned PG-GMM, a simple weighted sparse coding model, which has a closed-form solution, can be used to perform image denoising effectively, resulting in high PSNR measure, fast speed, and particularly the best visual quality among all competing methods. Jun Xu 0019, Lei Zhang 0006, Wangmeng Zuo, David Zhang 0001, Xiangchu Feng |
ICCV | 2 |
| 2015 | Ask the dictionary: Soft-assignment location-orientation pooling for image classificationabstractThe pooling step is one of the key components of the well-known Bag-of-visual words (BoW) model widely used in image classification. In this paper, we propose a novel pooling method, which is called Soft-Assignment Location-Orientation Pooling (SALOP). Inspired by the bag of statistical sampling analysis (Bossa), SALOP also explores the effect of dictionary for pooling method, but leverages both location and orientation information between the local descriptors and the atoms of dictionary to aggregate feature codes. Moreover, different from existing pooling methods, SALOP employs a soft-assignment pooling scheme to handle ambiguity and uncertainty existing in the pooling process. The evaluation is conducted on two image benchmarks: Scene15 and PASCAL VOC 2007. The experimental results show our SALOP can achieve promising performances. Qilong Wang 0001, Xiaona Deng, Peihua Li, Lei Zhang 0006 |
ICIP | 4 |
| 2015 | High-order information for robust iris recognition under less controlled conditionsabstractIris recognition has achieved great progress in cooperative environments in the past decades. However, in less controlled conditions it is still an open and challenging problem because of severe noisy factors induced by non-cooperative subjects. For handling this challenging problem, we propose a method called ordinal measure of outer product tensor (O2PT) which leverages the high-order information of image features. O2PT consists of two components. First we compute outer product tensors of raw features (e.g. SIFT) which are vectorized and locally aggregated, characterizing the second-order statistics of raw features. And then we compute the ordinal measure of the aggregated outer product tensors to model the order relation of iris texture, which makes the representation more compact and robust to noise and illumination changes. Furthermore, we combine two modalities to improve the matching performance, namely, O2PT for iris image matching and Fisher Vector (FV), which also exploits the high-order information, for eye image matching. We have achieved competitive matching performance on the challenging UBIRIS.v2 and CASIA-Iris-Thousand databases. Guanglei Yang, Hui Zeng 0001, Peihua Li, Lei Zhang 0006 |
ICIP | 4 |
| 2015 | End-to-End Photo-Sketch Generation via Fully Convolutional Representation LearningabstractSketch-based face recognition is an interesting task in vision and multimedia research, yet it is quite challenging due to the great difference between face photos and sketches. In this paper, we propose a novel approach for photo-sketch generation, aiming to automatically transform face photos into detail-preserving personal sketches. Unlike the traditional models synthesizing sketches based on a dictionary of exemplars, we develop a fully convolutional network to learn the end-to-end photo-sketch mapping. Our approach takes whole face photos as inputs and directly generates the corresponding sketch images with efficient inference and learning, in which the architecture is stacked by only convolutional kernels of very small sizes. To well capture the person identity during the photo-sketch transformation, we define our optimization objective in the form of joint generative discriminative minimization. In particular, a discriminative regularization term is incorporated into the photo-sketch generation, enhancing the discriminability of the generated person sketches against other individuals. Extensive experiments on several standard benchmarks suggest that our approach outperforms other state-of-the-arts in both photo sketch generation and face sketch verification. Liliang Zhang, Liang Lin 0004, Xian Wu 0007, Shengyong Ding, Lei Zhang 0006 |
ICMR | 5 |
| 2015 | Robust low-rank tensor factorization by cyclic weighted median
Deyu Meng, Biao Zhang 0005, Zongben Xu, Lei Zhang 0006, Chenqiang Gao |
Sci. China Inf. Sci. | 4 |
| 2015 | Sparsely encoded local descriptor for face verification
Zhen Cui 0001, Shiguang Shan, Ruiping Wang 0001, Lei Zhang 0006, Xilin Chen 0001 |
Neurocomputing | 4 |
| 2015 | Effective texture classification by texton encoding induced statistical features
Lei Zhang 0006, Jane You, Simon C. K. Shiu |
Pattern Recognit. | 2 |
| 2015 | Unsupervised feature selection by regularized self-representation
Pengfei Zhu 0001, Wangmeng Zuo, Lei Zhang 0006, Qinghua Hu, Simon C. K. Shiu |
Pattern Recognit. | 3 |
| 2015 | Bit-Scalable Deep Hashing With Regularized Similarity Learning for Image Retrieval and Person Re-IdentificationabstractExtracting informative image features and learning effective approximate hashing functions are two crucial steps in image retrieval. Conventional methods often study these two steps separately, e.g., learning hash functions from a predefined hand-crafted feature space. Meanwhile, the bit lengths of output hashing codes are preset in the most previous methods, neglecting the significance level of different bits and restricting their practical flexibility. To address these issues, we propose a supervised learning framework to generate compact and bit-scalable hashing codes directly from raw images. We pose hashing learning as a problem of regularized similarity learning. In particular, we organize the training images into a batch of triplet samples, each sample containing two images with the same label and one with a different label. With these triplet samples, we maximize the margin between the matched pairs and the mismatched pairs in the Hamming space. In addition, a regularization term is introduced to enforce the adjacency consistency, i.e., images of similar appearances should have similar codes. The deep convolutional neural network is utilized to train the model in an end-to-end fashion, where discriminative image features and hash functions are simultaneously optimized. Furthermore, each bit of our hashing codes is unequally weighted, so that we can manipulate the code lengths by truncating the insignificant bits. Our framework outperforms state-of-the-arts on public benchmarks of similar image search and also achieves promising results in the application of person re-identification in surveillance. It is also shown that the generated bit-scalable hashing codes well preserve the discriminative powers with shorter code lengths. Ruimao Zhang, Liang Lin 0004, Wangmeng Zuo, Lei Zhang 0006 |
IEEE Trans. Image Process. | 5 |
| 2015 | A Feature-Enriched Completely Blind Image Quality EvaluatorabstractExisting blind image quality assessment (BIQA) methods are mostly opinion-aware. They learn regression models from training images with associated human subjective scores to predict the perceptual quality of test images. Such opinion-aware methods, however, require a large amount of training samples with associated human subjective scores and of a variety of distortion types. The BIQA models learned by opinion-aware methods often have weak generalization capability, hereby limiting their usability in practice. By comparison, opinion-unaware methods do not need human subjective scores for training, and thus have greater potential for good generalization capability. Unfortunately, thus far no opinion-unaware BIQA method has shown consistently better quality prediction accuracy than the opinion-aware methods. Here, we aim to develop an opinion-unaware BIQA method that can compete with, and perhaps outperform, the existing opinion-aware methods. By integrating the features of natural image statistics derived from multiple cues, we learn a multivariate Gaussian model of image patches from a collection of pristine natural images. Using the learned multivariate Gaussian model, a Bhattacharyya-like distance is used to measure the quality of each image patch, and then an overall quality score is obtained by average pooling. The proposed BIQA method does not need any distorted sample images nor subjective quality scores for training, yet extensive experiments demonstrate its superior quality-prediction performance to the state-of-the-art opinion-aware BIQA methods. The MATLAB source code of our algorithm is publicly available at www.comp.polyu.edu.hk/~cslzhang/IQA/ILNIQE/ILNIQE.htm. Lin Zhang 0014, Lei Zhang 0006, Alan C. Bovik |
IEEE Trans. Image Process. | 2 |
| 2015 | A Kernel Classification Framework for Metric LearningabstractLearning a distance metric from the given training samples plays a crucial role in many machine learning tasks, and various models and optimization algorithms have been proposed in the past decade. In this paper, we generalize several state-of-the-art metric learning methods, such as large margin nearest neighbor (LMNN) and information theoretic metric learning (ITML), into a kernel classification framework. First, doublets and triplets are constructed from the training samples, and a family of degree-2 polynomial kernel functions is proposed for pairs of doublets or triplets. Then, a kernel classification framework is established to generalize many popular metric learning methods such as LMNN and ITML. The proposed framework can also suggest new metric learning methods, which can be efficiently implemented, interestingly, using the standard support vector machine (SVM) solvers. Two novel metric learning methods, namely, doublet-SVM and triplet-SVM, are then developed under the proposed framework. Experimental results show that doublet-SVM and triplet-SVM achieve competitive classification accuracies with state-of-the-art metric learning methods but with significantly less training time. Wangmeng Zuo, Lei Zhang 0006, Deyu Meng, David Zhang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2014 | Local Generic Representation for Face Recognition with Single Sample per Person
Pengfei Zhu 0001, Meng Yang 0001, Lei Zhang 0006, Il-Yong Lee |
ACCV (3) | 3 |
| 2014 | Towards a scalable resource-driven approach for detecting repackaged Android applicationsabstractRepackaged Android applications (or simply apps) are one of the major sources of mobile malware and also an important cause of severe revenue loss to app developers. Although a number of solutions have been proposed to detect repackaged apps, the majority of them heavily rely on code analysis, thus suffering from two limitations: (1) poor scalability due to the billion opcode problem; (2) unreliability to code obfuscation/app hardening techniques. In this paper, we explore an alternative approach that exploits core resources, which have close relationships with codes, to detect repackaged apps. More precisely, we define new features for characterizing apps, investigate two kinds of algorithms for searching similar apps, and propose a two-stage methodology to speed up the detection. We realize our approach in a system named ResDroid and conduct large scale evaluation on it. The results show that ResDroid can identify repackaged apps efficiently and effectively even if they are protected by obfuscation or hardening systems. Yuru Shao, Xiapu Luo, Chenxiong Qian, Pengfei Zhu 0001, Lei Zhang 0006 |
ACSAC | 5 |
| 2014 | Weighted Nuclear Norm Minimization with Application to Image DenoisingabstractAs a convex relaxation of the low rank matrix factorization problem, the nuclear norm minimization has been attracting significant research interest in recent years. The standard nuclear norm minimization regularizes each singular value equally to pursue the convexity of the objective function. However, this greatly restricts its capability and flexibility in dealing with many practical problems (e.g., denoising), where the singular values have clear physical meanings and should be treated differently. In this paper we study the weighted nuclear norm minimization (WNNM) problem, where the singular values are assigned different weights. The solutions of the WNNM problem are analyzed under different weighting conditions. We then apply the proposed WNNM algorithm to image denoising by exploiting the image nonlocal self-similarity. Experimental results clearly show that the proposed WNNM algorithm outperforms many state-of-the-art denoising algorithms such as BM3D in terms of both quantitative measure and visual perception quality. Shuhang Gu, Lei Zhang 0006, Wangmeng Zuo, Xiangchu Feng |
CVPR | 2 |
| 2014 | Point Matching in the Presence of Outliers in Both Point Sets: A Concave Optimization ApproachabstractRecently, a concave optimization approach has been proposed to solve the robust point matching (RPM) problem. This method is globally optimal, but it requires that each model point has a counterpart in the data point set. Unfortunately, such a requirement may not be satisfied in certain applications when there are outliers in both point sets. To address this problem, we relax this condition and reduce the objective function of RPM to a function with few nonlinear terms by eliminating the transformation variables. The resulting function, however, is no longer quadratic. We prove that it is still concave over the feasible region of point correspondence. The branch-and-bound (BnB) algorithm can then be used for optimization. To further improve the efficiency of the BnB algorithm whose bottleneck lies in the costly computation of the lower bound, we propose a new lower bounding scheme which has a k-cardinality linear assignment formulation and can be efficiently solved. Experimental results show that the proposed algorithm outperforms state-of-the-arts in its robustness to disturbances and point matching accuracy. Wei Lian, Lei Zhang 0006 |
CVPR | 2 |
| 2014 | Support Vector Guided Dictionary Learning
Sijia Cai, Wangmeng Zuo, Lei Zhang 0006, Xiangchu Feng, Ping Wang 0072 |
ECCV (4) | 3 |
| 2014 | Shrinkage Expansion Adaptive Metric Learning
Qilong Wang 0001, Wangmeng Zuo, Lei Zhang 0006, Peihua Li |
ECCV (7) | 3 |
| 2014 | Fast Visual Tracking via Dense Spatio-temporal Context Learning
Kaihua Zhang 0001, Lei Zhang 0006, Qingshan Liu 0001, David Zhang 0001, Ming-Hsuan Yang 0001 |
ECCV (5) | 2 |
| 2014 | Analysis of unrestrained curve of rectangular piston ring based on energy principleabstractA rotary engine with rectangle pistons and cylinders is claimed to have higher mechanism efficiency for eliminating the lateral force between piston skirt and liner. Constrained by the present machining and assembly accuracy, surfaces of piston could not fit the internal surface of cylinder well which might lead to gas leakage from cylinder. A novel ring used to seal rectangular cylinder was invented, and the working theory of the piston ring was present. Then based on the assumption of uniform pressure distribution on the contact surface of the ring the mathematical model of unrestrained curve of the ring was established. The modeling process was present in detail. Furthermore a finite element analysis was carried out as a way to validate the mathematical model of the piston ring. Results show that deformation of piston ring achieved from calculation and simulation meets well at initial and ending segments. The maximum error appears in the middle segment, which is only 11.52%. The accuracy of mathematical model is acceptable. Lei Zhang 0006, Cunyun Pan, Tengan Zou |
ICARCV | 1 |
| 2014 | Automatic foreground extraction in videoabstractThis paper presents an automatic and efficient system for extracting dynamic objects of interest from videos. We take advantage of a saliency map and an optimization-based segmentation algorithm to extract the foreground objects automatically in some key frames. Then, the segmentation results in those key frames are propagated to other frames via an error map-based propagation scheme. Finally, a Bayesian matting-based refinement approach is employed to to handle the topology changes. Experiments show that our system is able to generate high quality results at a low computation cost. Haoqian Wang, Kai Li 0016, Yongbing Zhang 0002, Lei Zhang 0006 |
ICASSP | 5 |
| 2014 | Transductive Gaussian processes for image denoisingabstractIn this paper we are interested in exploiting self-similarity information for discriminative image denoising. Towards this goal, we propose a simple yet powerful denoising method based on transductive Gaussian processes, which introduces self-similarity in the prediction stage. Our approach allows to build a rich similarity measure by learning hyper parameters defining multi-kernel combinations. We introduce perceptual-driven kernels to capture pixel-wise, gradient-based and local-structure similarities. In addition, our algorithm can integrate several initial estimates as input features to boost performance even further. We demonstrate the effectiveness of our approach on several benchmarks. The experiments show that our proposed denoising algorithm has better performance than competing discriminative denoising methods, and achieves competitive result with respect to the state-of-the-art. Shenlong Wang, Lei Zhang 0006, Raquel Urtasun |
ICCP | 2 |
| 2014 | Robust Principal Component Analysis with Complex NoiseabstractThe research on robust principal component analysis (RPCA) has been attracting much attention recently. The original RPCA model assumes sparse noise, and use the L_1-norm to characterize the error term. In practice, however, the noise is much more complex and it is not appropriate to simply use a certain L_p-norm for noise modeling. We propose a generative RPCA model under the Bayesian framework by modeling data noise as a mixture of Gaussians (MoG). The MoG is a universal approximator to continuous distributions and thus our model is able to fit a wide range of noises such as Laplacian, Gaussian, sparse noises and any combinations of them. A variational Bayes algorithm is presented to infer the posterior of the proposed model. All involved parameters can be recursively updated in closed form. The advantage of our method is demonstrated by extensive experiments on synthetic data, face modeling and background subtraction. Qian Zhao 0002, Deyu Meng, Zongben Xu, Wangmeng Zuo, Lei Zhang 0006 |
ICML | 5 |
| 2014 | Projective dictionary pair learning for pattern classification
Shuhang Gu, Lei Zhang 0006, Wangmeng Zuo, Xiangchu Feng |
NIPS | 2 |
| 2014 | Depth map super-resolution via iterative joint-trilateral-upsamplingabstractIn this paper, we propose a new approach to solve the depth map super-resolution (SR) and denoising problems simultaneously. Inspired by joint-bilateral-upsampling (JBU), we devised the joint-trilateral-upsampling (JTU), which takes edge of the initial depth map, texture of the corresponding high-resolution color image and the values of the surrounding depth pixels, into consideration during the process of SR. To preserve the sharp edge of the up-sampled depth map and remove the noise, we introduce an iterative implementation, where current up-sampled depth map is fed into the next iteration, to refine the filter coefficients of JTU. The iterative JTU presents a high performance at many aspects such as sharping edge, denoising and none texture copying, etc. To demonstrate the superiority of the proposed method, we carry out various experiments and show an across-the-board quality improvement by both of subjective and objective evaluations compared with previous state-of-art methods. Lei Zhang 0006, Yongbing Zhang 0002, Huiming Xuan, Qionghai Dai |
VCIP | 2 |
| 2014 | Sparse Representation Based Fisher Discrimination Dictionary Learning for Image Classification
Meng Yang 0001, Lei Zhang 0006, Xiangchu Feng, David Zhang 0001 |
Int. J. Comput. Vis. | 2 |
| 2014 | Special issue on "New sensing and processing technologies for hand-based biometrics authentication"
David Zhang 0001, Lei Zhang 0006 |
Inf. Sci. | 2 |
| 2014 | Special issue on "Multi-biometrics and Mobile-biometrics: Recent Advances and Future Research"
Lei Zhang 0006, Tieniu Tan, Arun Ross, Stefanos Zafeiriou |
Image Vis. Comput. | 1 |
| 2014 | Fast Compressive TrackingabstractIt is a challenging task to develop effective and efficient appearance models for robust object tracking due to factors such as pose variation, illumination change, occlusion, and motion blur. Existing online tracking algorithms often update models with samples from observations in recent frames. Despite much success has been demonstrated, numerous issues remain to be addressed. First, while these adaptive appearance models are data-dependent, there does not exist sufficient amount of data for online algorithms to learn at the outset. Second, online tracking algorithms often encounter the drift problems. As a result of self-taught learning, misaligned samples are likely to be added and degrade the appearance models. In this paper, we propose a simple yet effective and efficient tracking algorithm with an appearance model based on features extracted from a multiscale image feature space with data-independent basis. The proposed appearance model employs non-adaptive random projections that preserve the structure of the image feature space of objects. A very sparse measurement matrix is constructed to efficiently extract the features for the appearance model. We compress sample images of the foreground target and the background using the same sparse measurement matrix. The tracking task is formulated as a binary classification via a naive Bayes classifier with online update in the compressed domain. A coarse-to-fine search strategy is adopted to further reduce the computational complexity in the detection procedure. The proposed compressive tracking algorithm runs in real-time and performs favorably against state-of-the-art methods on challenging sequences in terms of efficiency, accuracy and robustness. Kaihua Zhang 0001, Lei Zhang 0006, Ming-Hsuan Yang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2014 | Fast and robust face recognition via coding residual map learning based adaptive masking
Meng Yang 0001, Zhizhao Feng, Simon C. K. Shiu, Lei Zhang 0006 |
Pattern Recognit. | 4 |
| 2014 | Image Set-Based Collaborative Representation for Face RecognitionabstractWith the rapid development of digital imaging and communication technologies, image set-based face recognition (ISFR) is becoming increasingly important. One key issue of ISFR is how to effectively and efficiently represent the query face image set using the gallery face image sets. The set-to-set distance-based methods ignore the relationship between gallery sets, whereas representing the query set images individually over the gallery sets ignores the correlation between query set images. In this paper, we propose a novel image set-based collaborative representation and classification method for ISFR. By modeling the query set as a convex or regularized hull, we represent this hull collaboratively over all the gallery sets. With the resolved representation coefficients, the distance between the query set and each gallery set can then be calculated for classification. The proposed model naturally and effectively extends the image-based collaborative representation to an image set based one, and our extensive experiments on benchmark ISFR databases show the superiority of the proposed method to state-of-the-art ISFR methods under different set sizes in terms of both recognition rate and efficiency. Pengfei Zhu 0001, Wangmeng Zuo, Lei Zhang 0006, Simon C. K. Shiu, David Zhang 0001 |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2014 | Mixed Noise Removal by Weighted Encoding With Sparse Nonlocal RegularizationabstractMixed noise removal from natural images is a challenging task since the noise distribution usually does not have a parametric model and has a heavy tail. One typical kind of mixed noise is additive white Gaussian noise (AWGN) coupled with impulse noise (IN). Many mixed noise removal methods are detection based methods. They first detect the locations of IN pixels and then remove the mixed noise. However, such methods tend to generate many artifacts when the mixed noise is strong. In this paper, we propose a simple yet effective method, namely weighted encoding with sparse nonlocal regularization (WESNR), for mixed noise removal. In WESNR, there is not an explicit step of impulse pixel detection; instead, soft impulse pixel detection via weighted encoding is used to deal with IN and AWGN simultaneously. Meanwhile, the image sparsity prior and nonlocal self-similarity prior are integrated into a regularization term and introduced into the variational encoding framework. Experimental results show that the proposed WESNR method achieves leading mixed noise removal performance in terms of both quantitative measures and visual quality. Jielin Jiang, Lei Zhang 0006, Jian Yang 0003 |
IEEE Trans. Image Process. | 2 |
| 2014 | Blind Image Quality Assessment Using Joint Statistics of Gradient Magnitude and Laplacian FeaturesabstractBlind image quality assessment (BIQA) aims to evaluate the perceptual quality of a distorted image without information regarding its reference image. Existing BIQA models usually predict the image quality by analyzing the image statistics in some transformed domain, e.g., in the discrete cosine transform domain or wavelet domain. Though great progress has been made in recent years, BIQA is still a very challenging task due to the lack of a reference image. Considering that image local contrast features convey important structural information that is closely related to image perceptual quality, we propose a novel BIQA model that utilizes the joint statistics of two types of commonly used local contrast features: 1) the gradient magnitude (GM) map and 2) the Laplacian of Gaussian (LOG) response. We employ an adaptive procedure to jointly normalize the GM and LOG features, and show that the joint statistics of normalized GM and LOG features have desirable properties for the BIQA task. The proposed model is extensively evaluated on three large-scale benchmark databases, and shown to deliver highly competitive performance with state-of-the-art BIQA models, as well as with some well-known full reference image quality assessment models. Wufeng Xue, Xuanqin Mou, Lei Zhang 0006, Alan C. Bovik, Xiangchu Feng |
IEEE Trans. Image Process. | 3 |
| 2014 | Gradient Magnitude Similarity Deviation: A Highly Efficient Perceptual Image Quality IndexabstractIt is an important task to faithfully evaluate the perceptual quality of output images in many applications, such as image compression, image restoration, and multimedia streaming. A good image quality assessment (IQA) model should not only deliver high quality prediction accuracy, but also be computationally efficient. The efficiency of IQA metrics is becoming particularly important due to the increasing proliferation of high-volume visual data in high-speed networks. We present a new effective and efficient IQA model, called gradient magnitude similarity deviation (GMSD). The image gradients are sensitive to image distortions, while different local structures in a distorted image suffer different degrees of degradations. This motivates us to explore the use of global variation of gradient based local quality map for overall image quality prediction. We find that the pixel-wise gradient magnitude similarity (GMS) between the reference and distorted images combined with a novel pooling strategy-the standard deviation of the GMS map-can predict accurately perceptual image quality. The resulting GMSD algorithm is much faster than most state-of-the-art IQA methods, and delivers highly competitive prediction accuracy. MATLAB source code of GMSD can be downloaded at http://www4.comp.polyu.edu.hk/~cslzhang/IQA/GMSD/GMSD.htm. Wufeng Xue, Lei Zhang 0006, Xuanqin Mou, Alan C. Bovik |
IEEE Trans. Image Process. | 2 |
| 2014 | A Sparse Embedding and Least Variance Encoding Approach to HashingabstractHashing is becoming increasingly important in large-scale image retrieval for fast approximate similarity search and efficient data storage. Many popular hashing methods aim to preserve the kNN graph of high dimensional data points in the low dimensional manifold space, which is, however, difficult to achieve when the number of samples is big. In this paper, we propose an effective and efficient hashing approach by sparsely embedding a sample in the training sample space and encoding the sparse embedding vector over a learned dictionary. To this end, we partition the sample space into clusters via a linear spectral clustering method, and then represent each sample as a sparse vector of normalized probabilities that it falls into its several closest clusters. This actually embeds each sample sparsely in the sample space. The sparse embedding vector is employed as the feature of each sample for hashing. We then propose a least variance encoding model, which learns a dictionary to encode the sparse embedding feature, and consequently binarize the coding coefficients as the hash codes. The dictionary and the binarization threshold are jointly optimized in our model. Experimental results on benchmark data sets demonstrated the effectiveness of the proposed approach in comparison with state-of-the-art methods. Xiaofeng Zhu 0001, Lei Zhang 0006, Zi Huang |
IEEE Trans. Image Process. | 2 |
| 2014 | Gradient Histogram Estimation and Preservation for Texture Enhanced Image DenoisingabstractNatural image statistics plays an important role in image denoising, and various natural image priors, including gradient-based, sparse representation-based, and nonlocal self-similarity-based ones, have been widely studied and exploited for noise removal. In spite of the great success of many denoising algorithms, they tend to smooth the fine scale image textures when removing noise, degrading the image visual quality. To address this problem, in this paper, we propose a texture enhanced image denoising method by enforcing the gradient histogram of the denoised image to be close to a reference gradient histogram of the original image. Given the reference gradient histogram, a novel gradient histogram preservation (GHP) algorithm is developed to enhance the texture structures while removing noise. Two region-based variants of GHP are proposed for the denoising of images consisting of regions with different textures. An algorithm is also developed to effectively estimate the reference gradient histogram from the noisy observation of the unknown image. Our experimental results demonstrate that the proposed GHP algorithm can well preserve the texture appearance in the denoised images, making them look more natural. Wangmeng Zuo, Lei Zhang 0006, Chunwei Song, David Zhang 0001, Huijun Gao |
IEEE Trans. Image Process. | 2 |
| 2013 | A Cyclic Weighted Median Method for L1 Low-Rank Matrix Factorization with Missing EntriesabstractA challenging problem in machine learning, information retrieval and computer vision research is how to recover a low-rank representation of the given data in the presence of outliers and missing entries. The L1-norm low-rank matrix factorization (LRMF) has been a popular approach to solving this problem. However, L1-norm LRMF is difficult to achieve due to its non-convexity and non-smoothness, and existing methods are often inefficient and fail to converge to a desired solution. In this paper we propose a novel cyclic weighted median (CWM) method, which is intrinsically a coordinate decent algorithm, for L1-norm LRMF. The CWM method minimizes the objective by solving a sequence of scalar minimization sub-problems, each of which is convex and can be easily solved by the weighted median filter. The extensive experimental results validate that the CWM method outperforms state-of-the-arts in terms of both accuracy and computational efficiency. Deyu Meng, Zongben Xu, Lei Zhang 0006, Ji Zhao 0001 |
AAAI | 3 |
| 2013 | Learning without Human Scores for Blind Image Quality AssessmentabstractGeneral purpose blind image quality assessment (BIQA) has been recently attracting significant attention in the fields of image processing, vision and machine learning. State-of-the-art BIQA methods usually learn to evaluate the image quality by regression from human subjective scores of the training samples. However, these methods need a large number of human scored images for training, and lack an explicit explanation of how the image quality is affected by image local features. An interesting question is then: can we learn for effective BIQA without using human scored images? This paper makes a good effort to answer this question. We partition the distorted images into overlapped patches, and use a percentile pooling strategy to estimate the local quality of each patch. Then a quality-aware clustering (QAC) method is proposed to learn a set of centroids on each quality level. These centroids are then used as a codebook to infer the quality of each patch in a given image, and subsequently a perceptual quality score of the whole image can be obtained. The proposed QAC based BIQA method is simple yet effective. It not only has comparable accuracy to those methods using human scored images in learning, but also has merits such as high linearity to human perception of image quality, real-time implementation and availability of image local quality map. Wufeng Xue, Lei Zhang 0006, Xuanqin Mou |
CVPR | 2 |
| 2013 | Texture Enhanced Image Denoising via Gradient Histogram PreservationabstractImage denoising is a classical yet fundamental problem in low level vision, as well as an ideal test bed to evaluate various statistical image modeling methods. One of the most challenging problems in image denoising is how to preserve the fine scale texture structures while removing noise. Various natural image priors, such as gradient based prior, nonlocal self-similarity prior, and sparsity prior, have been extensively exploited for noise removal. The denoising algorithms based on these priors, however, tend to smooth the detailed image textures, degrading the image visual quality. To address this problem, in this paper we propose a texture enhanced image denoising (TEID) method by enforcing the gradient distribution of the denoised image to be close to the estimated gradient distribution of the original image. A novel gradient histogram preservation (GHP) algorithm is developed to enhance the texture structures while removing noise. Our experimental results demonstrate that the proposed GHP based TEID can well preserve the texture features of the denoised images, making them look more natural. Wangmeng Zuo, Lei Zhang 0006, Chunwei Song, David Zhang 0001 |
CVPR | 2 |
| 2013 | A Novel Earth Mover's Distance Methodology for Image Matching with Gaussian Mixture ModelsabstractThe similarity or distance measure between Gaussian mixture models (GMMs) plays a crucial role in content-based image matching. Though the Earth Mover's Distance (EMD) has shown its advantages in matching histogram features, its potentials in matching GMMs remain unclear and are not fully explored. To address this problem, we propose a novel EMD methodology for GMM matching. We first present a sparse representation based EMD called SR-EMD by exploiting the sparse property of the underlying problem. SR-EMD is more efficient and robust than the conventional EMD. Second, we present two novel ground distances between component Gaussians based on the information geometry. The perspective from the Riemannian geometry distinguishes the proposed ground distances from the classical entropy-or divergence-based ones. Furthermore, motivated by the success of distance metric learning of vector data, we make the first attempt to learn the EMD distance metrics between GMMs by using a simple yet effective supervised pair-wise based method. It can adapt the distance metrics between GMMs to specific classification tasks. The proposed method is evaluated on both simulated data and benchmark real databases and achieves very promising performance. Peihua Li, Qilong Wang 0001, Lei Zhang 0006 |
ICCV | 3 |
| 2013 | Log-Euclidean Kernels for Sparse Representation and Dictionary LearningabstractThe symmetric positive definite (SPD) matrices have been widely used in image and vision problems. Recently there are growing interests in studying sparse representation (SR) of SPD matrices, motivated by the great success of SR for vector data. Though the space of SPD matrices is well-known to form a Lie group that is a Riemannian manifold, existing work fails to take full advantage of its geometric structure. This paper attempts to tackle this problem by proposing a kernel based method for SR and dictionary learning (DL) of SPD matrices. We disclose that the space of SPD matrices, with the operations of logarithmic multiplication and scalar logarithmic multiplication defined in the Log-Euclidean framework, is a complete inner product space. We can thus develop a broad family of kernels that satisfies Mercer's condition. These kernels characterize the geodesic distance and can be computed efficiently. We also consider the geometric structure in the DL process by updating atom matrices in the Riemannian space instead of in the Euclidean space. The proposed method is evaluated with various vision problems and shows notable performance gains over state-of-the-arts. Peihua Li, Qilong Wang 0001, Wangmeng Zuo, Lei Zhang 0006 |
ICCV | 4 |
| 2013 | Perceptual Fidelity Aware Mean Squared ErrorabstractHow to measure the perceptual quality of natural images is an important problem in low level vision. It is known that the Mean Squared Error (MSE) is not an effective index to describe the perceptual fidelity of images. Numerous perceptual fidelity indices have been developed, while the representatives include the Structural SIMilarity (SSIM) index and its variants. However, most of those perceptual measures are nonlinear, and they cannot be easily dopted as an objective function to minimize in various low level vision tasks. Can MSE be perceptual fidelity aware after some minor adaptation? In this paper we propose a simple framework to enhance the perceptual fidelity awareness of MSE by introducing an l2-norm structural error term to it. Such a Structural MSE (SMSE) can lead to very competitive image quality assessment (IQA) results. More surprisingly, we show that by using certain structure extractors, SMSE can be further turned into a Gaussian smoothed MSE (i.e., the Euclidean distance between the original and distorted images after Gaussian smooth filtering), which is much simpler to calculate but achieves rather better IQA performance than SSIM. The so called Perceptual-fidelity Aware MSE (PAMSE) can have great potentials in applications such as perceptual image coding and perceptual image restoration. Wufeng Xue, Xuanqin Mou, Lei Zhang 0006, Xiangchu Feng |
ICCV | 3 |
| 2013 | Sparse Variation Dictionary Learning for Face Recognition with a Single Training Sample per PersonabstractFace recognition (FR) with a single training sample per person (STSPP) is a very challenging problem due to the lack of information to predict the variations in the query sample. Sparse representation based classification has shown interesting results in robust FR, however, its performance will deteriorate much for FR with STSPP. To address this issue, in this paper we learn a sparse variation dictionary from a generic training set to improve the query sample representation by STSPP. Instead of learning from the generic training set independently w.r.t. the gallery set, the proposed sparse variation dictionary learning (SVDL) method is adaptive to the gallery set by jointly learning a projection to connect the generic training set with the gallery set. The learnt sparse variation dictionary can be easily integrated into the framework of sparse representation based classification so that various variations in face images, including illumination, expression, occlusion, pose, etc., can be better handled. Experiments on the large-scale CMU Multi-PIE, FRGC and LFW databases demonstrate the promising performance of SVDL on FR with STSPP. Meng Yang 0001, Luc Van Gool, Lei Zhang 0006 |
ICCV | 3 |
| 2013 | From Point to Set: Extend the Learning of Distance MetricsabstractMost of the current metric learning methods are proposed for point-to-point distance (PPD) based classification. In many computer vision tasks, however, we need to measure the point-to-set distance (PSD) and even set-to-set distance (SSD) for classification. In this paper, we extend the PPD based Mahalanobis distance metric learning to PSD and SSD based ones, namely point-to-set distance metric learning (PSDML) and set-to-set distance metric learning (SSDML), and solve them under a unified optimization framework. First, we generate positive and negative sample pairs by computing the PSD and SSD between training samples. Then, we characterize each sample pair by its covariance matrix, and propose a covariance kernel based discriminative function. Finally, we tackle the PSDML and SSDML problems by using standard support vector machine solvers, making the metric learning very efficient for multiclass visual classification tasks. Experiments on gender classification, digit recognition, object categorization and face recognition show that the proposed metric learning methods can effectively enhance the performance of PSD and SSD based classification. Pengfei Zhu 0001, Lei Zhang 0006, Wangmeng Zuo, David Zhang 0001 |
ICCV | 2 |
| 2013 | A Generalized Iterated Shrinkage Algorithm for Non-convex Sparse CodingabstractIn many sparse coding based image restoration and image classification problems, using non-convex Ip-norm minimization (0 ≤ p1-norm minimization. A number of algorithms, e.g., iteratively reweighted least squares (IRLS), iteratively thresholding method (ITM-Ip), and look-up table (LUT), have been proposed for non-convex Ip-norm sparse coding, while some analytic solutions have been suggested for some specific values of p. In this paper, by extending the popular soft-thresholding operator, we propose a generalized iterated shrinkage algorithm (GISA) for Ip-norm non-convex sparse coding. Unlike the analytic solutions, the proposed GISA algorithm is easy to implement, and can be adopted for solving non-convex sparse coding problems with arbitrary p values. Compared with LUT, GISA is more general and does not need to compute and store the look-up tables. Compared with IRLS and ITM-Ip, GISA is theoretically more solid and can achieve more accurate solutions. Experiments on image restoration and sparse coding based face recognition are conducted to validate the performance of GISA. Wangmeng Zuo, Deyu Meng, Lei Zhang 0006, Xiangchu Feng, David Zhang 0001 |
ICCV | 3 |
| 2013 | An Error Resilient Depth Map Coding Scheme Using Adaptive Wyner-Ziv Frame
Xiangkai Liu, Qiang Peng, Xiao Wu 0001, Lei Zhang 0006, Ling-Yu Duan |
MMM (2) | 4 |
| 2013 | SSIM-Based End-to-End Distortion Model for Error Resilient Video Coding over Packet-Switched Networks
Lei Zhang 0006, Qiang Peng, Xiao Wu 0001 |
MMM (1) | 1 |
| 2013 | Effective stereo matching using reliable points based graph cutabstractIn this paper, we propose an effective stereo matching algorithm using reliable points and region-based graph cut. Firstly, the initial disparity maps are calculated via local windowbased method. Secondly, the unreliable points are detected according to the DSI(Disparity Space Image) and the estimated disparity values of each unreliable point are obtained by considering its surrounding points. Then, the scheme of reliable points is introduced in region-based graph cut framework to optimize the initial result. Finally, remaining errors in the disparity results are effectively handled in a multi-step refinement process. Experiment results show that the proposed algorithm achieves a significant reduction in computation cost and guarantee high matching quality. Haoqian Wang, Yongbing Zhang 0002, Lei Zhang 0006 |
VCIP | 4 |
| 2013 | A learning-based method for compressive image recovery
Weisheng Dong, Guangming Shi, Xiaolin Wu 0001, Lei Zhang 0006 |
J. Vis. Commun. Image Represent. | 4 |
| 2013 | Joint segmentation and pairing of multispectral chromosome images
Yongqiang Zhao 0001, Xiaolin Wu 0001, Seong G. Kong, Lei Zhang 0006 |
Pattern Anal. Appl. | 4 |
| 2013 | Joint discriminative dimensionality reduction and dictionary learning for face recognition
Zhizhao Feng, Meng Yang 0001, Lei Zhang 0006, Yan Liu 0004, David Zhang 0001 |
Pattern Recognit. | 3 |
| 2013 | A survey of graph theoretical approaches to image segmentation
Bo Peng 0006, Lei Zhang 0006, David Zhang 0001 |
Pattern Recognit. | 2 |
| 2013 | Gabor feature based robust representation and classification for face recognition with Gabor occlusion dictionary
Meng Yang 0001, Lei Zhang 0006, Simon C. K. Shiu, David Zhang 0001 |
Pattern Recognit. | 2 |
| 2013 | Joint Registration and Active Contour Segmentation for Object TrackingabstractThis paper presents a novel object tracking framework by joint registration and active contour segmentation (JRACS), which can robustly deal with the non-rigid shape changes of the target. The target region, which includes both foreground and background pixels, is implicitly represented by a level set. A Bhattacharyya similarity based metric is proposed to locate the region whose foreground and background distributions best match those of the tracked target. Based on this metric, a tracking framework that consists of a registration stage and a segmentation stage is then established. The registration step roughly locates the target object by modeling its motion as an affine transformation, and the segmentation step refines the registration result and computes the true contour of the target. The robust tracking performance of the proposed JRACS method is demonstrated by real video sequences where the objects have clear non-rigid shape changes. Jifeng Ning, Lei Zhang 0006, David Zhang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2013 | Robust Object Tracking Via Active Feature SelectionabstractAdaptive tracking by detection has been widely studied with promising results. The key idea of such trackers is how to train an online discriminative classifier, which can well separate an object from its local background. The classifier is incrementally updated using positive and negative samples extracted from the current frame around the detected object location. However, if the detection is less accurate, the samples are likely to be less accurately extracted, thereby leading to visual drift. Recently, the multiple instance learning (MIL) based tracker has been proposed to solve these problems to some degree. It puts samples into the positive and negative bags, and then selects some features with an online boosting method via maximizing the bag likelihood function. Finally, the selected features are combined for classification. However, in MIL tracker the features are selected by a likelihood function, which can be less informative to tell the target from complex background. Motivated by the active learning method, in this paper we propose an active feature selection approach that is able to select more informative features than the MIL tracker by using the Fisher information criterion to measure the uncertainty of the classification model. More specifically, we propose an online boosting feature selection approach via optimizing the Fisher information criterion, which can yield more robust and efficient real-time object tracking performance. Experimental evaluations on challenging sequences demonstrate the efficiency, accuracy, and robustness of the proposed tracker in comparison with state-of-the-art trackers. Kaihua Zhang 0001, Lei Zhang 0006, Ming-Hsuan Yang 0001, Qinghua Hu |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2013 | Sparse Representation Based Image Interpolation With Nonlocal Autoregressive ModelingabstractSparse representation is proven to be a promising approach to image super-resolution, where the low-resolution (LR) image is usually modeled as the down-sampled version of its high-resolution (HR) counterpart after blurring. When the blurring kernel is the Dirac delta function, i.e., the LR image is directly down-sampled from its HR counterpart without blurring, the super-resolution problem becomes an image interpolation problem. In such cases, however, the conventional sparse representation models (SRM) become less effective, because the data fidelity term fails to constrain the image local structures. In natural images, fortunately, many nonlocal similar patches to a given patch could provide nonlocal constraint to the local structure. In this paper, we incorporate the image nonlocal self-similarity into SRM for image interpolation. More specifically, a nonlocal autoregressive model (NARM) is proposed and taken as the data fidelity term in SRM. We show that the NARM-induced sampling matrix is less coherent with the representation dictionary, and consequently makes SRM more effective for image interpolation. Our extensive experimental results demonstrate that the proposed NARM-based image interpolation method can effectively reconstruct the edge structures and suppress the jaggy/ringing artifacts, achieving the best image interpolation results so far in terms of PSNR as well as perceptual quality metrics such as SSIM and FSIM. Weisheng Dong, Lei Zhang 0006, Rastislav Lukac, Guangming Shi |
IEEE Trans. Image Process. | 2 |
| 2013 | Nonlocally Centralized Sparse Representation for Image RestorationabstractSparse representation models code an image patch as a linear combination of a few atoms chosen out from an over-complete dictionary, and they have shown promising results in various image restoration applications. However, due to the degradation of the observed image (e.g., noisy, blurred, and/or down-sampled), the sparse representations by conventional models may not be accurate enough for a faithful reconstruction of the original image. To improve the performance of sparse representation-based image restoration, in this paper the concept of sparse coding noise is introduced, and the goal of image restoration turns to how to suppress the sparse coding noise. To this end, we exploit the image nonlocal self-similarity to obtain good estimates of the sparse coding coefficients of the original image, and then centralize the sparse coding coefficients of the observed image to those estimates. The so-called nonlocally centralized sparse representation (NCSR) model is as simple as the standard sparse representation model, while our extensive experiments on various types of image restoration problems, including denoising, deblurring and super-resolution, validate the generality and state-of-the-art performance of the proposed NCSR algorithm. Weisheng Dong, Lei Zhang 0006, Guangming Shi, Xin Li 0005 |
IEEE Trans. Image Process. | 2 |
| 2013 | Reconstruction Based Finger-Knuckle-Print Verification With Score Level Adaptive Binary FusionabstractRecently, a new biometrics identifier, namely finger knuckle print (FKP), has been proposed for personal authentication with very interesting results. One of the advantages of FKP verification lies in its user friendliness in data collection. However, the user flexibility in positioning fingers also leads to a certain degree of pose variations in the collected query FKP images. The widely used Gabor filtering based competitive coding scheme is sensitive to such variations, resulting in many false rejections. We propose to alleviate this problem by reconstructing the query sample with a dictionary learned from the template samples in the gallery set. The reconstructed FKP image can reduce much the enlarged matching distance caused by finger pose variations; however, both the intra-class and inter-class distances will be reduced. We then propose a score level adaptive binary fusion rule to adaptively fuse the matching distances before and after reconstruction, aiming to reduce the false rejections without increasing much the false acceptances. Experimental results on the benchmark PolyU FKP database show that the proposed method significantly improves the FKP verification accuracy. Guangwei Gao, Lei Zhang 0006, Jian Yang 0003, Lin Zhang 0014, David Zhang 0001 |
IEEE Trans. Image Process. | 2 |
| 2013 | Regularized Robust Coding for Face RecognitionabstractRecently the sparse representation based classification (SRC) has been proposed for robust face recognition (FR). In SRC, the testing image is coded as a sparse linear combination of the training samples, and the representation fidelity is measured by the l2-norm or l1 -norm of the coding residual. Such a sparse coding model assumes that the coding residual follows Gaussian or Laplacian distribution, which may not be effective enough to describe the coding residual in practical FR systems. Meanwhile, the sparsity constraint on the coding coefficients makes the computational cost of SRC very high. In this paper, we propose a new face coding model, namely regularized robust coding (RRC), which could robustly regress a given signal with regularized regression coefficients. By assuming that the coding residual and the coding coefficient are respectively independent and identically distributed, the RRC seeks for a maximum a posterior solution of the coding problem. An iteratively reweighted regularized robust coding (IR(3)C) algorithm is proposed to solve the RRC model efficiently. Extensive experiments on representative face databases demonstrate that the RRC is much more effective and efficient than state-of-the-art sparse representation based methods in dealing with face occlusion, corruption, lighting, and expression changes, etc. Meng Yang 0001, Lei Zhang 0006, Jian Yang 0003, David Zhang 0001 |
IEEE Trans. Image Process. | 2 |
| 2013 | Real-Time Object Tracking Via Online Discriminative Feature SelectionabstractMost tracking-by-detection algorithms train discriminative classifiers to separate target objects from their surrounding background. In this setting, noisy samples are likely to be included when they are not properly sampled, thereby causing visual drift. The multiple instance learning (MIL) paradigm has been recently applied to alleviate this problem. However, important prior information of instance labels and the most correct positive instance (i.e., the tracking result in the current frame) can be exploited using a novel formulation much simpler than an MIL approach. In this paper, we show that integrating such prior information into a supervised learning algorithm can handle visual drift more effectively and efficiently than the existing MIL tracker. We present an online discriminative feature selection algorithm that optimizes the objective function in the steepest ascent direction with respect to the positive samples while in the steepest descent direction with respect to the negative ones. Therefore, the trained classifier directly couples its score with the importance of samples, leading to a more robust and efficient tracker. Numerous experimental evaluations with state-of-the-art algorithms on challenging sequences demonstrate the merits of the proposed algorithm. Kaihua Zhang 0001, Lei Zhang 0006, Ming-Hsuan Yang 0001 |
IEEE Trans. Image Process. | 2 |
| 2013 | Reinitialization-Free Level Set Evolution via Reaction DiffusionabstractThis paper presents a novel reaction-diffusion (RD) method for implicit active contours that is completely free of the costly reinitialization procedure in level set evolution (LSE). A diffusion term is introduced into LSE, resulting in an RD-LSE equation, from which a piecewise constant solution can be derived. In order to obtain a stable numerical solution from the RD-based LSE, we propose a two-step splitting method to iteratively solve the RD-LSE equation, where we first iterate the LSE equation, then solve the diffusion equation. The second step regularizes the level set function obtained in the first step to ensure stability, and thus the complex and costly reinitialization procedure is completely eliminated from LSE. By successfully applying diffusion to LSE, the RD-LSE model is stable by means of the simple finite difference method, which is very easy to implement. The proposed RD method can be generalized to solve the LSE for both variational level set method and partial differential equation-based level set method. The RD-LSE method shows very good performance on boundary antileakage. The extensive and promising experimental results on synthetic and real images validate the effectiveness of the proposed RD-LSE approach. Kaihua Zhang 0001, Lei Zhang 0006, Huihui Song 0003, David Zhang 0001 |
IEEE Trans. Image Process. | 2 |
| 2013 | Sparse Representation Classifier Steered Discriminative Projection With Applications to Face RecognitionabstractA sparse representation-based classifier (SRC) is developed and shows great potential for real-world face recognition. This paper presents a dimensionality reduction method that fits SRC well. SRC adopts a class reconstruction residual-based decision rule, we use it as a criterion to steer the design of a feature extraction method. The method is thus called the SRC steered discriminative projection (SRC-DP). SRC-DP maximizes the ratio of between-class reconstruction residual to within-class reconstruction residual in the projected space and thus enables SRC to achieve better performance. SRC-DP provides low-dimensional representation of human faces to make the SRC-based face recognition system more efficient. Experiments are done on the AR, the extended Yale B, and PIE face image databases, and results demonstrate the proposed method is more effective than other feature extraction methods based on the SRC. Jian Yang 0003, Delin Chu, Lei Zhang 0006, Yong Xu 0001, Jing-Yu Yang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2013 | Robust Kernel Representation With Statistical Local Features for Face RecognitionabstractFactors such as misalignment, pose variation, and occlusion make robust face recognition a difficult problem. It is known that statistical features such as local binary pattern are effective for local feature extraction, whereas the recently proposed sparse or collaborative representation-based classification has shown interesting results in robust face recognition. In this paper, we propose a novel robust kernel representation model with statistical local features (SLF) for robust face recognition. Initially, multipartition max pooling is used to enhance the invariance of SLF to image registration error. Then, a kernel-based representation model is proposed to fully exploit the discrimination information embedded in the SLF, and robust regression is adopted to effectively handle the occlusion in face images. Extensive experiments are conducted on benchmark face databases, including extended Yale B, AR (A. Martinez and R. Benavente), multiple pose, illumination, and expression (multi-PIE), facial recognition technology (FERET), face recognition grand challenge (FRGC), and labeled faces in the wild (LFW), which have different variations of lighting, expression, pose, and occlusions, demonstrating the promising performance of the proposed method. Meng Yang 0001, Lei Zhang 0006, Simon C. K. Shiu, David Zhang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2012 | Nonlocal Spectral Prior Model for Low-Level Vision
Shenlong Wang, Lei Zhang 0006, Yan Liang 0001 |
ACCV (3) | 2 |
| 2012 | Semi-coupled dictionary learning with applications to image super-resolution and photo-sketch synthesisabstractIn various computer vision applications, often we need to convert an image in one style into another style for better visualization, interpretation and recognition; for examples, up-convert a low resolution image to a high resolution one, and convert a face sketch into a photo for matching, etc. A semi-coupled dictionary learning (SCDL) model is proposed in this paper to solve such cross-style image synthesis problems. Under SCDL, a pair of dictionaries and a mapping function will be simultaneously learned. The dictionary pair can well characterize the structural domains of the two styles of images, while the mapping function can reveal the intrinsic relationship between the two styles' domains. In SCDL, the two dictionaries will not be fully coupled, and hence much flexibility can be given to the mapping function for an accurate conversion across styles. Moreover, clustering and image nonlocal redundancy are introduced to enhance the robustness of SCDL. The proposed SCDL model is applied to image super-resolution and photo-sketch synthesis, and the experimental results validated its generality and effectiveness in cross-style image synthesis. Shenlong Wang, Lei Zhang 0006, Yan Liang 0001, Quan Pan 0001 |
CVPR | 2 |
| 2012 | Relaxed collaborative representation for pattern classificationabstractRegularized linear representation learning has led to interesting results in image classification, while how the object should be represented is a critical issue to be investigated. Considering the fact that the different features in a sample should contribute differently to the pattern representation and classification, in this paper we present a novel relaxed collaborative representation (RCR) model to effectively exploit the similarity and distinctiveness of features. In RCR, each feature vector is coded on its associated dictionary to allow flexibility of feature coding, while the variance of coding vectors is minimized to address the similarity among features. In addition, the distinctiveness of different features is exploited by weighting its distance to other features in the coding domain. The proposed RCR is simple, while our extensive experimental results on benchmark image databases (e.g., various face and flower databases) show that it is very competitive with state-of-the-art image classification methods. Meng Yang 0001, Lei Zhang 0006, David Zhang 0001, Shenlong Wang |
CVPR | 2 |
| 2012 | Content Adaptive Subsampling for Stereo Interleaving Video CodingabstractStereo interleaving video coding receives considerable attention due to its desirable property of being compatible with 2D video coding standards. The errors caused by sub sampling (causing distortion between subsampling interpolated image and the original full resolution one) and by quantization during compression lead to the final distortion in stereo interleaving video coding. In this paper, the rate and distortion analysis in stereo interleaving video coding is provided. It proves that appropriate sub sampling in stereo interleaving video coding is able to obtain good compression performance. Subsequently, a content adaptive sub sampling (CAS) is proposed. In CAS, the half resolution frames are generated by decimation, where the down sampling filter coefficients are calculated based on frame contents and the targeted interpolation coefficients. Experiment results demonstrate that the CAS is able to achieve high compression efficiency of stereo interleaving encoding scheme for stereoscopic videos. Yongbing Zhang 0002, Xiangyang Ji, Haoqian Wang, Lei Zhang 0006, Qionghai Dai |
DCC | 4 |
| 2012 | Robust Point Matching Revisited: A Concave Optimization Approach
Wei Lian, Lei Zhang 0006 |
ECCV (2) | 2 |
| 2012 | Evaluation of Image Segmentation Quality by Adaptive Ground Truth Composition
Bo Peng 0006, Lei Zhang 0006 |
ECCV (3) | 2 |
| 2012 | Efficient Misalignment-Robust Representation for Real-Time Face Recognition
Meng Yang 0001, Lei Zhang 0006, David Zhang 0001 |
ECCV (1) | 2 |
| 2012 | Real-Time Compressive Tracking
Kaihua Zhang 0001, Lei Zhang 0006, Ming-Hsuan Yang 0001 |
ECCV (3) | 2 |
| 2012 | Multi-scale Patch Based Collaborative Representation for Face Recognition with Margin Distribution Optimization
Pengfei Zhu 0001, Lei Zhang 0006, Qinghua Hu, Simon C. K. Shiu |
ECCV (1) | 2 |
| 2012 | A comprehensive evaluation of full reference image quality assessment algorithmsabstractRecent years have witnessed a growing interest in developing objective image quality assessment (IQA) algorithms that can measure the image quality consistently with subjective evaluations. For the full reference (FR) IQA problem, great progress has been made in the past decade. On the other hand, several new large scale image datasets have been released for evaluating FR IQA methods in recent years. Meanwhile, no work has been reported to evaluate and compare the performance of state-of-the-art and representative FR IQA methods on all the available datasets. In this paper, we aim to fulfill this task by reporting the performance of eleven selected FR IQA algorithms on all the seven public IQA image datasets. Our evaluation results and the associated discussions will be very helpful for relevant researchers to have a clearer understanding about the status of modern FR IQA indices. Evaluation results presented in this paper are also online available at http://sse.tongji.edu.cn/linzhang/IQA/IQA.htm. Lin Zhang 0014, Lei Zhang 0006, Xuanqin Mou, David Zhang 0001 |
ICIP | 2 |
| 2012 | Sparse Representation Classifier for microaneurysm detection and retinal blood vessel extraction
Bob Zhang 0001, Fakhri Karray, Qin Li 0001, Lei Zhang 0006 |
Inf. Sci. | 4 |
| 2012 | Color demosaicking with an image formation model and adaptive PCA
Dahua Gao, Xiaolin Wu 0001, Guangming Shi, Lei Zhang 0006 |
J. Vis. Commun. Image Represent. | 4 |
| 2012 | Arbitrary body segmentation in static images
Shifeng Li, Huchuan Lu, Lei Zhang 0006 |
Pattern Recognit. | 3 |
| 2012 | Beyond sparsity: The role of L1-optimizer in pattern classification
Jian Yang 0003, Lei Zhang 0006, Yong Xu 0001, Jing-Yu Yang 0001 |
Pattern Recognit. | 2 |
| 2012 | Phase congruency induced local features for finger-knuckle-print recognition
Lin Zhang 0014, Lei Zhang 0006, David Zhang 0001, Zhenhua Guo 0001 |
Pattern Recognit. | 2 |
| 2012 | Sparse neighbor representation for classification
Kang-hua Hui, Chun-li Li, Lei Zhang 0006 |
Pattern Recognit. Lett. | 3 |
| 2012 | Image reconstruction with locally adaptive sparsity and nonlocal robust regularization
Weisheng Dong, Guangming Shi, Xin Li 0005, Lei Zhang 0006, Xiaolin Wu 0001 |
Signal Process. Image Commun. | 4 |
| 2012 | A Novel Algorithm for Finding Reducts With Fuzzy Rough SetsabstractAttribute reduction is one of the most meaningful research topics in the existing fuzzy rough sets, and the approach of discernibility matrix is the mathematical foundation of computing reducts. When computing reducts with discernibility matrix, we find that only the minimal elements in a discernibility matrix are sufficient and necessary. This fact motivates our idea in this paper to develop a novel algorithm to find reducts that are based on the minimal elements in the discernibility matrix. Relative discernibility relations of conditional attributes are defined and minimal elements in the fuzzy discernibility matrix are characterized by the relative discernibility relations. Then, the algorithms to compute minimal elements and reducts are developed in the framework of fuzzy rough sets. Experimental comparison shows that the proposed algorithms are effective. Degang Chen 0002, Lei Zhang 0006, Suyun Zhao, Qinghua Hu, Pengfei Zhu 0001 |
IEEE Trans. Fuzzy Syst. | 2 |
| 2012 | Feature Selection for Monotonic ClassificationabstractMonotonic classification is a kind of special task in machine learning and pattern recognition. Monotonicity constraints between features and decision should be taken into account in these tasks. However, most existing techniques are not able to discover and represent the ordinal structures in monotonic datasets. Thus, they are inapplicable to monotonic classification. Feature selection has been proven effective in improving classification performance and avoiding overfitting. To the best of our knowledge, no technique has been specially designed to select features in monotonic classification until now. In this paper, we introduce a function, which is called rank mutual information, to evaluate monotonic consistency between features and decision in monotonic tasks. This function combines the advantages of dominance rough sets in reflecting ordinal structures and mutual information in terms of robustness. Then, rank mutual information is integrated with the search strategy of min-redundancy and max-relevance to compute optimal subsets of features. A collection of numerical experiments are given to show the effectiveness of the proposed technique. Qinghua Hu, Lei Zhang 0006, David Zhang 0001, Yanping Song, Maozu Guo 0001, Daren Yu |
IEEE Trans. Fuzzy Syst. | 3 |
| 2012 | On Robust Fuzzy Rough Set ModelsabstractRough sets, especially fuzzy rough sets, are supposedly a powerful mathematical tool to deal with uncertainty in data analysis. This theory has been applied to feature selection, dimensionality reduction, and rule learning. However, it is pointed out that the classical model of fuzzy rough sets is sensitive to noisy information, which is considered as a main source of uncertainty in applications. This disadvantage limits the applicability of fuzzy rough sets. In this paper, we reveal why the classical fuzzy rough set model is sensitive to noise and how noisy samples impose influence on fuzzy rough computation. Based on this discussion, we study the properties of some current fuzzy rough models in dealing with noisy data and introduce several new robust models. The properties of the proposed models are also discussed. Finally, a robust classification algorithm is designed based on fuzzy lower approximations. Some numerical experiments are given to illustrate the effectiveness of the models. The classifiers that are developed with the proposed models achieve good generalization performance. Qinghua Hu, Lei Zhang 0006, Shuang An, David Zhang 0001, Daren Yu |
IEEE Trans. Fuzzy Syst. | 2 |
| 2012 | Feature Band Selection for Online Multispectral Palmprint RecognitionabstractA palmprint is a unique and reliable biometric feature with high usability. In the past decades, many palmprint recognition systems have been successfully developed. However, most of the previous work used the white light as the illumination source, and the recognition accuracy and anti-spoof capability is limited. Recently, multispectral imaging has attracted considerable research attention as it can acquire more discriminative information in a short time. One crucial step in developing online multispectral palmprint systems is how to determine the optimal number of spectral bands and select the most representative bands to build the system. This paper presents a study on feature band selection by analyzing hyperspectral palmprint data (520-1050 nm). Our experimental results showed that three spectral bands could provide most of the discriminate information of a palmprint. This finding could be used as the guidance for designing new online multispectral palmprint systems. Zhenhua Guo 0001, David Zhang 0001, Lei Zhang 0006, Wenhuang Liu |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2012 | Monogenic Binary Coding: An Efficient Local Feature Extraction Approach to Face RecognitionabstractLocal-feature-based face recognition (FR) methods, such as Gabor features encoded by local binary pattern, could achieve state-of-the-art FR results in large-scale face databases such as FERET and FRGC. However, the time and space complexity of Gabor transformation are too high for many practical FR applications. In this paper, we propose a new and efficient local feature extraction scheme, namely monogenic binary coding (MBC), for face representation and recognition. Monogenic signal representation decomposes an original signal into three complementary components: amplitude, orientation, and phase. We encode the monogenic variation in each local region and monogenic feature in each pixel, and then calculate the statistical features (e.g., histogram) of the extracted local features. The local statistical features extracted from the complementary monogenic components (i.e., amplitude, orientation, and phase) are then fused for effective FR. It is shown that the proposed MBC scheme has significantly lower time and space complexity than the Gabor-transformation-based local feature methods. The extensive FR experiments on four large-scale databases demonstrated the effectiveness of MBC, whose performance is competitive with and even better than state-of-the-art local-feature-based FR methods. Meng Yang 0001, Lei Zhang 0006, Simon C. K. Shiu, David Zhang 0001 |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2012 | Rotation-Invariant Nonrigid Point Set Matching in Cluttered ScenesabstractThis paper addresses the problem of rotation-invariant nonrigid point set matching. The shape context (SC) feature descriptor is used because of its strong discriminative nature, whereas edges in the graphs constructed by point sets are used to determine the orientations of SCs. Similar to lengths or directions, oriented SCs constructed this way can be regarded as attributes of edges. By matching edges between two point sets, rotation invariance is achieved. Two novel ways of constructing graphs on a model point set are proposed, aiming at making the orientations of SCs as robust to disturbances as possible. The structures of these graphs facilitate the use of dynamic programming (DP) for optimization. The strong discriminative nature of SC, the special structure of the model graphs, and the global optimality of DP make our methods robust to various types of disturbances, particularly clutters. The extensive experiments on both synthetic and real data validated the robustness of the proposed methods to various types of disturbances. They can robustly detect the desired shapes in complex and highly cluttered scenes. Wei Lian, Lei Zhang 0006, David Zhang 0001 |
IEEE Trans. Image Process. | 2 |
| 2012 | l2 Restoration of l∞-Decoded Images Via Soft-Decision EstimationabstractThe l(∞)-constrained image coding is a technique to achieve substantially lower bit rate than strictly (mathematically) lossless image coding, while still imposing a tight error bound at each pixel. However, this technique becomes inferior in the l(2) distortion metric if the bit rate decreases further. In this paper, we propose a new soft decoding approach to reduce the l(2) distortion of l(∞)-decoded images and retain the advantages of both minmax and least-square approximations. The soft decoding is performed in a framework of image restoration that exploits the tight error bounds afforded by the l(∞)-constrained coding and employs a context modeler of quantization errors. Experimental results demonstrate that the l(∞)-constrained hard decoded images can be restored to gain more than 2 dB in peak signal-to-noise ratio PSNR, while still retaining tight error bounds on every single pixel. The new soft decoding technique can even outperform JPEG 2000 (a state-of-the-art encoder-optimized image codec) for bit rates higher than 1 bpp, a critical rate region for applications of near-lossless image compression. All the coding gains are made without increasing the encoder complexity as the heavy computations to gain coding efficiency are delegated to the decoder. Jiantao Zhou 0001, Xiaolin Wu 0001, Lei Zhang 0006 |
IEEE Trans. Image Process. | 3 |
| 2012 | Sample Pair Selection for Attribute Reduction with Rough SetabstractAttribute reduction is the strongest and most characteristic result in rough set theory to distinguish itself to other theories. In the framework of rough set, an approach of discernibility matrix and function is the theoretical foundation of finding reducts. In this paper, sample pair selection with rough set is proposed in order to compress the discernibility function of a decision table so that only minimal elements in the discernibility matrix are employed to find reducts. First relative discernibility relation of condition attribute is defined, indispensable and dispensable condition attributes are characterized by their relative discernibility relations and key sample pair set is defined for every condition attribute. With the key sample pair sets, all the sample pair selections can be found. Algorithms of computing one sample pair selection and finding reducts are also developed; comparisons with other methods of finding reducts are performed with several experiments which imply sample pair selection is effective as preprocessing step to find reducts. Degang Chen 0002, Suyun Zhao, Lei Zhang 0006, Yongping Yang, Xiao Zhang 0012 |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2012 | Rank Entropy-Based Decision Trees for Monotonic ClassificationabstractIn many decision making tasks, values of features and decision are ordinal. Moreover, there is a monotonic constraint that the objects with better feature values should not be assigned to a worse decision class. Such problems are called ordinal classification with monotonicity constraint. Some learning algorithms have been developed to handle this kind of tasks in recent years. However, experiments show that these algorithms are sensitive to noisy samples and do not work well in real-world applications. In this work, we introduce a new measure of feature quality, called rank mutual information (RMI), which combines the advantage of robustness of Shannon's entropy with the ability of dominance rough sets in extracting ordinal structures from monotonic data sets. Then, we design a decision tree algorithm (REMT) based on rank mutual information. The theoretic and experimental analysis shows that the proposed algorithm can get monotonically consistent decision trees, if training samples are monotonically consistent. Its performance is still good when data are contaminated with noise. Qinghua Hu, Xunjian Che, Lei Zhang 0006, David Zhang 0001, Maozu Guo 0001, Daren Yu |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2012 | Low-Dose X-ray CT Reconstruction via Dictionary LearningabstractAlthough diagnostic medical imaging provides enormous benefits in the early detection and accuracy diagnosis of various diseases, there are growing concerns on the potential side effect of radiation induced genetic, cancerous and other diseases. How to reduce radiation dose while maintaining the diagnostic performance is a major challenge in the computed tomography (CT) field. Inspired by the compressive sensing theory, the sparse constraint in terms of total variation (TV) minimization has already led to promising results for low-dose CT reconstruction. Compared to the discrete gradient transform used in the TV method, dictionary learning is proven to be an effective way for sparse representation. On the other hand, it is important to consider the statistical property of projection data in the low-dose CT case. Recently, we have developed a dictionary learning based approach for low-dose X-ray CT. In this paper, we present this method in detail and evaluate it in experiments. In our method, the sparse constraint in terms of a redundant dictionary is incorporated into an objective function in a statistical iterative reconstruction framework. The dictionary can be either predetermined before an image reconstruction task or adaptively defined during the reconstruction process. An alternating minimization scheme is developed to minimize the objective function. Our approach is evaluated with low-dose X-ray projections collected in animal and human CT studies, and the improvement associated with dictionary learning is quantified relative to filtered backprojection and TV-based reconstructions. The results show that the proposed approach might produce better images with lower noise and more detailed structural features in our selected cases. However, there is no proof that this is true for all kinds of structures. Hengyong Yu, Xuanqin Mou, Lei Zhang 0006, Jiang Hsieh, Ge Wang 0001 |
IEEE Trans. Medical Imaging | 4 |
| 2012 | Principal Line-Based Alignment Refinement for Palmprint RecognitionabstractImage alignment is an important step in various biometric authentication applications such as palmprint recognition. Most of the existing palmprint alignment methods make use of some key points between fingers or in palm boundary to establish the local coordinate system for region of interest (ROI) extraction. The ROI is consequently used for feature extraction and matching. Such alignment methods usually yield a coarse alignment of the palmprint images, while many missed and false matches are actually caused by inaccurate image alignments. To improve the palmprint verification accuracy, in this paper, we present an efficient palmprint alignment refinement method. After extracting the principal lines from the palmprint image, we apply the iterative closest point method to them to estimate the translation and rotation parameters between two images. The estimated parameters are then used to refine the alignment of palmprint feature maps for a more accurate palmprint matching. The experimental results show that the proposed method greatly improves the palmprint recognition accuracy and it works in real time. Wei Li 0016, Bob Zhang 0001, Lei Zhang 0006, Jingqi Yan |
IEEE Trans. Syst. Man Cybern. Part C | 3 |
| 2011 | Sparsity-based image denoising via dictionary learning and structural clusteringabstractWhere does the sparsity in image signals come from? Local and nonlocal image models have supplied complementary views toward the regularity in natural images - the former attempts to construct or learn a dictionary of basis functions that promotes the sparsity; while the latter connects the sparsity with the self-similarity of the image source by clustering. In this paper, we present a variational framework for unifying the above two views and propose a new denoising algorithm built upon clustering-based sparse representation (CSR). Inspired by the success of l1-optimization, we have formulated a double-header l1-optimization problem where the regularization involves both dictionary learning and structural structuring. A surrogate-function based iterative shrinkage solution has been developed to solve the double-header l1-optimization problem and a probabilistic interpretation of CSR model is also included. Our experimental results have shown convincing improvements over state-of-the-art denoising technique BM3D on the class of regular texture images. The PSNR performance of CSR denoising is at least comparable and often superior to other competing schemes including BM3D on a collection of 12 generic natural images. Weisheng Dong, Xin Li 0005, Lei Zhang 0006, Guangming Shi |
CVPR | 3 |
| 2011 | Robust sparse coding for face recognitionabstractRecently the sparse representation (or coding) based classification (SRC) has been successfully used in face recognition. In SRC, the testing image is represented as a sparse linear combination of the training samples, and the representation fidelity is measured by the l2-norm or l1-norm of coding residual. Such a sparse coding model actually assumes that the coding residual follows Gaussian or Laplacian distribution, which may not be accurate enough to describe the coding errors in practice. In this paper, we propose a new scheme, namely the robust sparse coding (RSC), by modeling the sparse coding as a sparsity-constrained robust regression problem. The RSC seeks for the MLE (maximum likelihood estimation) solution of the sparse coding problem, and it is much more robust to outliers (e.g., occlusions, corruptions, etc.) than SRC. An efficient iteratively reweighted sparse coding algorithm is proposed to solve the RSC model. Extensive experiments on representative face databases demonstrate that the RSC scheme is much more effective than state-of-the-art methods in dealing with face occlusion, corruption, lighting and expression changes, etc. Meng Yang 0001, Lei Zhang 0006, Jian Yang 0003, David Zhang 0001 |
CVPR | 2 |
| 2011 | Centralized sparse representation for image restorationabstractThis paper proposes a novel sparse representation model called centralized sparse representation (CSR) for image restoration tasks. In order for faithful image reconstruction, it is expected that the sparse coding coefficients of the degraded image should be as close as possible to those of the unknown original image with the given dictionary. However, since the available data are the degraded (noisy, blurred and/or down-sampled) versions of the original image, the sparse coding coefficients are often not accurate enough if only the local sparsity of the image is considered, as in many existing sparse representation models. To make the sparse coding more accurate, a centralized sparsity constraint is introduced by exploiting the nonlocal image statistics. The local sparsity and the nonlocal sparsity constraints are unified into a variational framework for optimization. Extensive experiments on image restoration validated that our CSR model achieves convincing improvement over previous state-of-the-art methods. Weisheng Dong, Lei Zhang 0006, Guangming Shi |
ICCV | 2 |
| 2011 | Fisher Discrimination Dictionary Learning for sparse representationabstractSparse representation based classification has led to interesting image recognition results, while the dictionary used for sparse coding plays a key role in it. This paper presents a novel dictionary learning (DL) method to improve the pattern classification performance. Based on the Fisher discrimination criterion, a structured dictionary, whose dictionary atoms have correspondence to the class labels, is learned so that the reconstruction error after sparse coding can be used for pattern classification. Meanwhile, the Fisher discrimination criterion is imposed on the coding coefficients so that they have small within-class scatter but big between-class scatter. A new classification scheme associated with the proposed Fisher discrimination DL (FDDL) method is then presented by using both the discriminative information in the reconstruction error and sparse coding coefficients. The proposed FDDL is extensively evaluated on benchmark image databases in comparison with existing sparse representation and DL based classification methods. Meng Yang 0001, Lei Zhang 0006, Xiangchu Feng, David Zhang 0001 |
ICCV | 2 |
| 2011 | Sparse representation or collaborative representation: Which helps face recognition?abstractAs a recently proposed technique, sparse representation based classification (SRC) has been widely used for face recognition (FR). SRC first codes a testing sample as a sparse linear combination of all the training samples, and then classifies the testing sample by evaluating which class leads to the minimum representation error. While the importance of sparsity is much emphasized in SRC and many related works, the use of collaborative representation (CR) in SRC is ignored by most literature. However, is it really the l1-norm sparsity that improves the FR accuracy? This paper devotes to analyze the working mechanism of SRC, and indicates that it is the CR but not the l1-norm sparsity that makes SRC powerful for face classification. Consequently, we propose a very simple yet much more efficient face classification scheme, namely CR based classification with regularized least square (CRC_RLS). The extensive experiments clearly show that CRC_RLS has very competitive classification results, while it has significantly less complexity than SRC. Lei Zhang 0006, Meng Yang 0001, Xiangchu Feng |
ICCV | 1 |
| 2011 | A linear subspace learning approach via sparse codingabstractLinear subspace learning (LSL) is a popular approach to image recognition and it aims to reveal the essential features of high dimensional data, e.g., facial images, in a lower dimensional space by linear projection. Most LSL methods compute directly the statistics of original training samples to learn the subspace. However, these methods do not effectively exploit the different contributions of different image components to image recognition. We propose a novel LSL approach by sparse coding and feature grouping. A dictionary is learned from the training dataset, and it is used to sparsely decompose the training samples. The decomposed image components are grouped into a more discriminative part (MDP) and a less discriminative part (LDP). An unsupervised criterion and a supervised criterion are then proposed to learn the desired subspace, where the MDP is preserved and the LDP is suppressed simultaneously. The experimental results on benchmark face image databases validated that the proposed methods outperform many state-of-the-art LSL schemes. Lei Zhang 0006, Pengfei Zhu 0001, Qinghua Hu, David Zhang 0001 |
ICCV | 1 |
| 2011 | Strategy of Statistics-Based Visualization for Segmented 3D Cardiac Volume Data Set
Changqing Gai, Kuanquan Wang, Lei Zhang 0006, Wangmeng Zuo |
ICIC (1) | 3 |
| 2011 | Sparsity-based image deblurring with locally adaptive and nonlocally robust regularizationabstractImportant structures in photographic images such as edges and textures are jointly characterized by local variation and nonlocal invariance (similarity). Both of them provide valuable heuristics to the regularization of image restoration process. In this pa per, we propose to explore two sets of complementary ideas: 1) locally learn PCA-based dictionaries and estimate the sparsity regularization parameters for each coefficient; and 2) nonlocally enforce the invariance constraint by introducing a patch-similarity based term into the cost functional. The minimization of this new cost functional leads to an iterative thresholding-based image deblurring algorithm and its efficient implementation is discussed. Our experimental results have shown that the proposed scheme significantly outperforms several leading deblurring techniques in the literature on both objective and visual quality assessments. Weisheng Dong, Xin Li 0005, Lei Zhang 0006, Guangming Shi |
ICIP | 3 |
| 2011 | Measuring relevance between discrete and continuous features based on neighborhood mutual information
Qinghua Hu, Lei Zhang 0006, David Zhang 0001, Shuang An, Witold Pedrycz |
Expert Syst. Appl. | 2 |
| 2011 | Online joint palmprint and palmvein verification
David Zhang 0001, Zhenhua Guo 0001, Guangming Lu 0002, Lei Zhang 0006, Wangmeng Zuo |
Expert Syst. Appl. | 4 |
| 2011 | Image retrieval based on micro-structure descriptor
Guanghai Liu 0001, Lei Zhang 0006, Yong Xu 0001 |
Pattern Recognit. | 3 |
| 2011 | Image segmentation by iterated region merging with localized graph cuts
Bo Peng 0006, Lei Zhang 0006, David Zhang 0001, Jian Yang 0003 |
Pattern Recognit. | 2 |
| 2011 | A multi-manifold discriminant analysis method for image feature extraction
Wankou Yang, Changyin Sun 0001, Lei Zhang 0006 |
Pattern Recognit. | 3 |
| 2011 | From classifiers to discriminators: A nearest neighbor rule induced discriminant analysis
Jian Yang 0003, Lei Zhang 0006, Jing-Yu Yang 0001, David Zhang 0001 |
Pattern Recognit. | 2 |
| 2011 | Ensemble of local and global information for finger-knuckle-print recognition
Lin Zhang 0014, Lei Zhang 0006, David Zhang 0001, Hailong Zhu |
Pattern Recognit. | 2 |