Radu Timofte

dblp:24/8616 · DBLP profile ↗
← Back
210ranked-venue papers
19as first author
127since 2021 · last 2026
0000-0002-1478-0402ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 168 · 17 first-author · 101 since 2021Graphics, computer vision, multimedia, augmented reality and games · 152 · 13 first-author · 87 since 2021Systems, architecture and hardware · 7 · 6 since 2021Databases, data management, data science and information retrieval · 3 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Evaluating Feedback by Iterative Repair of Multi-step Solution Documents
Tobias Lengfeld, Jakob Seitz, Radu Timofte
ICDAR (2)3
2026 P4: Place with Purpose - Pose and Prompt-Guided Human Synthesis in Real Scenes
Dadan Khan, Mohammad Zohaib, Radu Timofte, Francesca Odone
ICPR (2)3
2026 Closed-Loop LLM Discovery of Non-standard Channel Priors in Vision Models
Tolgay Atinc Uzun, Dmitry Ignatov, Radu Timofte
ICPR (12)3
2026 INRetouch: Context Aware Implicit Neural Representation for Photography Retouching
abstract
Professional photo editing remains challenging, requiring extensive knowledge of imaging pipelines and significant expertise. While recent deep learning approaches, particularly style transfer methods, have attempted to automate this process, they often struggle with output fidelity, editing control, and complex retouching capabilities. We propose a novel retouch transfer approach that learns from professional edits through before-after image pairs, enabling precise replication of complex editing operations. We develop a context-aware Implicit Neural Representation that learns to apply edits adaptively based on image content and context, and is capable of learning from a single example. Our method extracts implicit transformations from reference edits and adaptively applies them to new images. To facilitate this research direction, we introduce a comprehensive Photo Retouching Dataset comprising 100,000 high-quality images edited using over 170 professional Adobe Lightroom presets. Through extensive evaluation, we demonstrate that our approach not only surpasses existing methods in photo retouching but also enhances performance in related image reconstruction tasks like Gamut Mapping and Raw Reconstruction. By bridging the gap between professional editing capabilities and automated solutions, our work presents a significant step toward making sophisticated photo editing more accessible while maintaining high-fidelity results. The source code and the dataset are publicly available at omaralezaby.github.io/inretouch/
Omar Elezabi, Marcos V. Conde, Zongwei Wu, Radu Timofte
WACV4
2026 Illuminating Darkness: Learning to Enhance Low-light Images In-the-Wild
abstract
Single-shot low-light image enhancement (SLLIE) remains challenging due to the limited availability of diverse, realworld paired datasets. To bridge this gap, we introduce the Low-Light Smartphone Dataset (LSD), a large-scale, high-resolution (4K+) dataset collected in the wild across a wide range of challenging lighting conditions (0.1–200 lux). LSD contains 6,425 precisely aligned low and normallight image pairs, selected from over 8,000 dynamic indoor and outdoor scenes through multi-frame acquisition and expert evaluation. To evaluate generalization and aesthetic quality, we collect 2,117 unpaired low-light images from previously unseen devices. To fully exploit LSD, we propose TFFormer, a hybrid model that encodes luminance and chrominance (LC) separately to reduce color-structure entanglement. We further propose a cross-attention-driven joint decoder for context-aware fusion of LC representations, along with LC refinement and LC-guided supervision to significantly enhance perceptual fidelity and structural consistency. TFFormer achieves state-of-the-art results on LSD (+2.45 dB PSNR) and substantially improves downstream vision tasks, such as low-light object detection (+6.80 mAP on ExDark).
S. M. A. Sharif, Fayaz Ali Dharejo, Radu Timofte, Rizwan Ali Naqvi
WACV5
2026 Empowering Image Restoration: A Multi-Attention Approach
Yawei Li 0001, Chao Zhang 0094, Weiyan Hou, Luc Van Gool, Radu Timofte
Expert Syst. Appl.6
2026 DASR+: Training Domain Distance Aware Network for Unsupervised Image Super-Resolution
Xiaorui Zhao, Yunxuan Wei, Xin Deng 0002, Yawei Li 0001, Radu Timofte, Hengjie Song, Shuhang Gu
Int. J. Comput. Vis.5
2026 Spec-ViT: A Vision Transformer With Wavelet for Anti-Aliasing and Denoising in Medical Image Classification
abstract
Medical image analysis remains challenging due to inherent limitations in imaging modalities, where structural aliasing and noise artifacts persistently compromise diagnostic accuracy. While convolutional neural networks (CNNs) and vision transformers (ViTs) have achieved remarkable progress in feature extraction, their inherent sampling mechanisms and spectral biases often exacerbate these high-frequency distortions, leading to suboptimal lesion characterization. To address this critical limitation, we propose Spec-ViT, a novel wavelet-based anti-aliasing Transformer architecture that synergistically integrates adaptive spectral purification with hierarchical attentive learning. The Wavelet Antialiasing Module (WAM) first implements learnable smoothing factor in the wavelet domain to suppress highfrequency artifacts, while preserving clinically relevant lowfrequency structures and fine diagnostic details. Building upon this spectral foundation, the Lightweight Enhanced Attention (LEA) refines feature representations through a dual-path mechanism, coupling channel-spatial attention with global multi-head self-attention to enhance lesion context modeling. Finally, the Smoothed Convolutional Gate (SCG) further sharpens local discriminability through depthwise convolution and adaptive Swish gating, completing a coherent pipeline from frequency-aware purification to global-local attentive analysis. Extensive experiments on five benchmark medical image classification datasets demonstrate that Spec-ViT consistently outperforms both baseline and state-of-the-art methods, achieving up to 84.04% accuracy on the Pediatric Pneumonia Chest X-rays dataset in particular.
Yanying Rao, Yuzheng Su, Fayaz Ali Dharejo, Radu Timofte, Guo-jun Mao, Moath Alathbah
IEEE J. Biomed. Health Informatics6
2025 Centurio: On Drivers of Multilingual Ability of Large Vision-Language Model
abstract
Gregor Geigle, Florian Schneider, Carolin Holtermann, Chris Biemann, Radu Timofte, Anne Lauscher, Goran Glavaš. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Gregor Geigle, Florian Schneider 0001, Carolin Holtermann, Chris Biemann, Radu Timofte, Anne Lauscher, Goran Glavas
ACL (1)5
2025 RipVIS: Rip Currents Video Instance Segmentation Benchmark for Beach Monitoring and Safety
abstract
Rip currents are strong, localized and narrow currents of water that flow outwards into the sea, causing numerous beach-related injuries and fatalities worldwide. Accurate identification of rip currents remains challenging due to their amorphous nature and the lack of annotated data, which often requires expert knowledge. To address these issues, we present RipVIS, a large-scale video instance segmentation benchmark explicitly designed for rip current segmentation. RipVIS is an order of magnitude larger than previous datasets, featuring 184 videos (212, 328 frames), of which 150 videos (163, 528 frames) are with rip currents, collected from various sources, including drones, mobile phones, and fixed beach cameras. Our dataset encompasses diverse visual contexts, such as wave-breaking patterns, sediment flows, and water color variations, across multiple global locations, including USA, Mexico, Costa Rica, Portugal, Italy, Greece, Romania, Sri Lanka, Australia and New Zealand. Most videos are annotated at 5 FPS to ensure accuracy in dynamic scenarios, supplemented by an additional 34 videos (48, 800 frames) without rip currents. We conduct comprehensive experiments with Mask R-CNN, Cascade Mask R-CNN, SparseInst and YOLO11, fine-tuning these models for the task of rip current segmentation. Results are reported in terms of multiple metrics, with a particular focus on the F2score to prioritize recall and reduce false negatives. To enhance segmentation performance, we introduce a novel post-processing step based on Temporal Confidence Aggregation (TCA). RipVIS aims to set a new standard for rip current segmentation, contributing towards safer beach environments. We offer a benchmark website to share data, models, and results with the research community, encouraging ongoing collaboration and future contributions, at https://ripvis.ai.
Andrei Dumitriu, Florin Tatui, Florin Miron, Aakash Ralhan, Radu Tudor Ionescu, Radu Timofte
CVPR6
2025 ReCap: Better Gaussian Relighting with Cross-Environment Captures
abstract
Accurate 3D objects relighting in diverse unseen environments is crucial for realistic virtual object placement. Due to the albedo-lighting ambiguity, existing methods often fall short in producing faithful relights. Without proper constraints, observed training views can be explained by numerous combinations of lighting and material attributes, lacking physical correspondence with the actual environment maps used for relighting. In this work, we present ReCap, treating cross-environment captures as multi-task target to provide the missing supervision that cuts through the entanglement. Specifically, ReCap jointly optimizes multiple lighting representations that share a common set of material attributes. This naturally harmonizes a coherent set of lighting representations around the mutual material attributes, exploiting commonalities and differences across varied object appearances. Such coherence enables physically sound lighting reconstruction and robust material estimation — both essential for accurate relighting. Together with a streamlined shading function and effective post-processing, ReCap outperforms all leading competitors on an expanded relighting benchmark.
Zongwei Wu, Eduard Zamfir, Radu Timofte
CVPR4
2025 Complexity Experts are Task-Discriminative Learners for Any Image Restoration
abstract
Recent advancements in all-in-one image restoration models have revolutionized the ability to address diverse degradations through a unified framework. However, parameters tied to specific tasks often remain inactive for other tasks, making mixture-of-experts (MoE) architectures a natural extension. Despite this, MoEs often show inconsistent behavior, with some experts unexpectedly generalizing across tasks while others struggle within their intended scope. This hinders leveraging MoEs’ computational benefits by bypassing irrelevant experts during inference. We attribute this undesired behavior to the uniform and rigid architecture of traditional MoEs. To address this, we introduce “complexity experts” – flexible expert blocks with varying computational complexity and receptive fields. A key challenge is assigning tasks to each expert, as degradation complexity is unknown in advance. Thus, we execute tasks with a simple bias toward lower complexity. To our surprise, this preference effectively drives task-specific allocation, assigning tasks to experts with the appropriate complexity. Extensive experiments validate our approach, demonstrating the ability to bypass irrelevant experts during inference while maintaining superior performance. The proposed MoCE-IR model outperforms state-of-the-art methods, affirming its efficiency and practical applicability. The source code and models are publicly available at eduardzamfir.github.io/MoCE-IR/
Eduard Zamfir, Zongwei Wu, Nancy Mehta, Yuedong Tan, Danda Pani Paudel, Yulun Zhang 0001, Radu Timofte
CVPR7
2025 MIORe & VAR-MIORe: Benchmarks to Push the Boundaries of Restoration
George Ciubotariu, Zhuyun Zhou, Zongwei Wu, Radu Timofte
ICCV4
2025 PixTalk: Controlling Photorealistic Image Processing and Editing with Language
Marcos V. Conde, Zihao Lu, Radu Timofte
ICCV3
2025 Color Matching Using Hypernetwork-Based Kolmogorov-Arnold Networks
Artem V. Nikonorov, Georgy Perevozchikov, Andrei Korepanov, Nancy Mehta, Mahmoud Afifi, Egor Ershov, Radu Timofte
ICCV7
2025 Bokehlicious: Photorealistic Bokeh Rendering with Controllable Apertures
Tim Seizinger, Florin-Alexandru Vasluianu, Marcos V. Conde, Zongwei Wu, Radu Timofte
ICCV5
2025 What You Have is What You Track: Adaptive and Robust Multimodal Tracking
abstract
Multimodal data is known to be helpful for visual tracking by improving robustness to appearance variations. However, sensor synchronization challenges often compromise data availability, particularly in video settings where shortages can be temporal. Despite its importance, this area remains underexplored. In this paper, we present the first comprehensive study on tracker performance with temporally incomplete multimodal data. Unsurprisingly, under such a circumstance, existing trackers exhibit significant performance degradation, as their rigid architectures lack the adaptability needed to effectively handle missing modalities. To address these limitations, we propose a flexible framework for robust multimodal tracking. We venture that a tracker should dynamically activate computational units based on missing data rates. This is achieved through a novel Heterogeneous Mixture-of-Experts fusion mechanism with adaptive complexity, coupled with a video-level masking strategy that ensures both temporal consistency and spatial completeness which is critical for effective video tracking. Surprisingly, our model not only adapts to varying missing rates but also adjusts to scene complexity. Extensive experiments show that our model achieves SOTA performance across 9 benchmarks, excelling in both conventional complete and missing modality settings. The code and benchmark will be publicly available at https://github.com/supertyd/FlexTrack/tree/main.
Yuedong Tan, Jiawei Shao, Eduard Zamfir, Ruanjun Li, Zhaochong An, Chao Ma 0004, Danda Pani Paudel, Luc Van Gool, Radu Timofte, Zongwei Wu
ICCV9
2025 XTrack: Multimodal Training Boosts RGB-X Video Object Trackers
Yuedong Tan, Zongwei Wu, Yuqian Fu, Zhuyun Zhou, Guolei Sun, Eduard Zamfir, Chao Ma 0004, Danda Pani Paudel, Luc Van Gool, Radu Timofte
ICCV10
2025 After the Party: Navigating the Mapping from Color to Ambient Lighting
Florin-Alexandru Vasluianu, Tim Seizinger, Zongwei Wu, Radu Timofte
ICCV4
2025 The Return of Structural Handwritten Mathematical Expression Recognition
Jakob Seitz, Tobias Lengfeld, Radu Timofte
ICDAR (1)3
2025 LeMoRe: Learn More Details for Lightweight Semantic Segmentation
abstract
Lightweight semantic segmentation is essential for many downstream vision tasks. Unfortunately, existing methods often struggle to balance efficiency and performance due to the complexity of feature modeling. Many of these existing approaches are constrained by rigid architectures and implicit representation learning, often characterized by parameter-heavy designs and a reliance on computationally intensive Vision Transformer-based frameworks. In this work, we introduce an efficient paradigm by synergizing explicit and implicit modeling to balance computational efficiency with representational fidelity. Our method combines well-defined Cartesian directions with explicitly modeled views and implicitly inferred intermediate representations, efficiently capturing global dependencies through a nested attention mechanism. Extensive experiments on challenging datasets, including ADE20K, CityScapes, Pascal Context, and COCO-Stuff, demonstrate that LeMoRe strikes an effective balance between performance and efficiency. https://github.com/miannaeem-lab/LeMoRe
Mian Muhammad Naeem Abid, Nancy Mehta, Zongwei Wu, Radu Timofte
ICIP4
2025 Fast Iterative Enhancement for Image Signal Processing
abstract
The Image Signal Processor (ISP) is a key component in modern cameras that transforms the RAW scene radiance captured by the camera sensor into sRGB images that are suitable for the human visual system. The ISP is usually a model-based pipeline, and it comprises several stages (or blocks) such as denoising, white balance, color correction, and tone mapping. Many of these blocks are non-linear operations, or even deep neural networks. In this work, we propose the use of iterative diffusion models as an additional photo-finishing block in the imaging pipeline to produce high-quality sRGB images. This showcases the power of generative learned ISPs.
Marcos V. Conde, Radu Timofte
ICIP2
2025 Reverse Distillation Based Detection of Anomalies on a Newly Developed Fabric Dataset
abstract
Neural networks have become a main tool for anomaly detection in images. One idea is to use an auto encoder for detecting irregularities in textures, and knowledge distillation to calculate an anomaly score. The method of (Thomine and Snoussi, 2024), trained on the MVTec carpet dataset, is our starting point. We modify this method for our new Textile Stripe (TS) real-world textile dataset. TS contains images of anomalies of different shapes in a textile. It differs from the MVTec carpet class in properties such as image dimensions, illumination, or alignment of the texture to the image edges. We propose a newly developed cropping and anomaly map reconcatenation technique with overlapping and test the self-ensembling method of (Meininger and Tim-ofte, 2024). The modifications made are also cross-validated on the MVTec dataset. For generalization experiments, two smaller test datasets, called G1FW and G2SR, containing normal and anomalous images on different fabrics are provided. The approaches delivered adequate performance in terms of AUROC.
Christian Jaspert, Radu Timofte
ICIP2
2025 Learning Transformer-based World Models with Contrastive Predictive Coding
abstract
The DreamerV3 algorithm recently obtained remarkable performance across diverse environment domains by learning an accurate world model based on Recurrent Neural Networks (RNNs). Following the success of model-based reinforcement learning algorithms and the rapid adoption of the Transformer architecture for its superior training efficiency and favorable scaling properties, recent works such as STORM have proposed replacing RNN-based world models with Transformer-based world models using masked self-attention. However, despite the improved training efficiency of these methods, their impact on performance remains limited compared to the Dreamer algorithm, struggling to learn competitive Transformer-based world models. In this work, we show that the next state prediction objective adopted in previous approaches is insufficient to fully exploit the representation capabilities of Transformers. We propose to extend world model predictions to longer time horizons by introducing TWISTER (Transformer-based World model wIth contraSTivE Representations), a world model using action-conditioned Contrastive Predictive Coding to learn high-level temporal feature representations and improve the agent performance. TWISTER achieves a human-normalized mean score of 162% on the Atari 100k benchmark, setting a new record among state-of-the-art methods that do not employ look-ahead search. We release our code at https://github.com/burchim/TWISTER.
Maxime Burchi, Radu Timofte
ICLR2
2025 Accurate and Efficient World Modeling with Masked Latent Transformers
abstract
The Dreamer algorithm has recently obtained remarkable performance across diverse environment domains by training powerful agents with simulated trajectories. However, the compressed nature of its world model’s latent space can result in the loss of crucial information, negatively affecting the agent’s performance. Recent approaches, such as $\Delta$-IRIS and DIAMOND, address this limitation by training more accurate world models. However, these methods require training agents directly from pixels, which reduces training efficiency and prevents the agent from benefiting from the inner representations learned by the world model. In this work, we propose an alternative approach to world modeling that is both accurate and efficient. We introduce EMERALD (Efficient MaskEd latent tRAnsformer worLD model), a world model using a spatial latent state with MaskGIT predictions to generate accurate trajectories in latent space and improve the agent performance. On the Crafter benchmark, EMERALD achieves new state-of-the-art performance, becoming the first method to surpass human experts performance within 10M environment steps. Our method also succeeds to unlock all 22 Crafter achievements at least once during evaluation.
Maxime Burchi, Radu Timofte
ICML2
2025 Steering Prediction via a Multi-Sensor System for Autonomous Racing
abstract
Autonomous racing has rapidly gained research attention. Traditionally, racing cars rely on 2D LiDAR as their primary visual system. In this work, we explore the integration of an event camera with the existing system to provide enhanced temporal information. Our goal is to fuse the 2D LiDAR data with event data in an end-to-end learning framework for steering prediction, which is crucial for autonomous racing. To the best of our knowledge, this is the first study addressing this challenging research topic. We start by creating a multisensor dataset specifically for steering prediction. Using this dataset, we establish a benchmark by evaluating various SOTA fusion methods. Our observations reveal that existing methods often incur substantial computational costs. To address this, we apply low-rank techniques to propose a novel, efficient, and effective fusion design. We introduce a new fusion learning policy to guide the fusion process, enhancing robustness against misalignment. Our fusion architecture provides better steering prediction than LiDAR alone, significantly reducing the RMSE from 7.72 to 1.28. Compared to the second-best fusion method, our work represents only 11% of the learnable parameters while achieving better accuracy. The source code and dataset are publicly available at: https://github.com/ZZY-Zhou/F1Tenth-Steering.
Zhuyun Zhou, Zongwei Wu, Florian Bolli, Rémi Boutteau, Fan Yang 0019, Radu Timofte, Dominique Ginhac, Tobi Delbruck
ICRA6
2025 When super-resolution meets camouflaged object detection: A comparison study
Shupeng Cheng, Weiyan Hou, Luc Van Gool, Radu Timofte
Comput. Vis. Image Underst.5
2025 SF-YOLO: A Novel YOLO Framework for Small Object Detection in Aerial Scenes
abstract
ABSTRACT Object detection models are widely applied in the fields such as video surveillance and unmanned aerial vehicles to enable the identification and monitoring of various objects on a diversity of backgrounds. The general CNN‐based object detectors primarily rely on downsampling and pooling operations, often struggling with small objects that have low resolution and failing to fully leverage contextual information that can differentiate objects from complex background. To address the problems, we propose a novel YOLO framework called SF‐YOLO for small object detection. Firstly, we present a spatial information perception (SIP) module to extract contextual features for different objects through the integration of space to depth operation and large selective kernel module, which dynamically adjusts receptive field of the backbone and obtains the enhanced features for richer understanding of differentiation between objects and background. Furthermore, we design a novel multi‐scale feature weighted fusion strategy, which performs weighted fusion on feature maps by combining fast normalized fusion method and CARAFE operation, accurately assessing the importance of each feature and enhancing the representation of small objects. The extensive experiments conducted on VisDrone2019, Tiny‐Person and PESMOD datasets demonstrate that our proposed method enables comparable detection performance to state‐of‐the‐art detectors.
Wangyu Jiang, Fayaz Ali Dharejo, Guojun Mao, Radu Timofte
IET Image Process.6
2025 Parameterized Low-Rank Regularizer for High-dimensional Visual Data
Zixiang Zhao, Xiangyong Cao, Jiangjun Peng, Xi-Le Zhao, Deyu Meng, Yulun Zhang 0001, Radu Timofte, Luc Van Gool
Int. J. Comput. Vis.8
2025 DiffI2I: Efficient Diffusion Model for Image-to-Image Translation
abstract
The Diffusion Model (DM) has emerged as the SOTA approach for image synthesis. However, the existing DM cannot perform well on some image-to-image translation (I2I) tasks. Different from image synthesis, some I2I tasks, such as super-resolution, require generating results in accordance with GT images. Traditional DMs for image synthesis require extensive iterations and large denoising models to estimate entire images, which gives their strong generative ability but also leads to artifacts and inefficiency for I2I. To tackle this challenge, we propose a simple, efficient, and powerful DM framework for I2I, called DiffI2I. Specifically, DiffI2I comprises three key components: a compact I2I prior extraction network (CPEN), a dynamic I2I transformer (DI2Iformer), and a denoising network. We train DiffI2I in two stages: pretraining and DM training. For pretraining, GT and input images are fed into CPEN to capture a compact I2I prior representation (IPR) guiding DI2Iformer. In the second stage, the DM is trained to only use the input images to estimate the same IRP as CPEN. Compared to traditional DMs, the compact IPR enables DiffI2I to obtain more accurate outcomes and employ a lighter denoising network and fewer iterations. Through extensive experiments on various I2I tasks, we demonstrate that DiffI2I achieves SOTA performance while significantly reducing computational burdens.
Bin Xia 0014, Yulun Zhang 0001, Shiyin Wang, Yapeng Tian, Wenming Yang, Radu Timofte, Luc Van Gool
IEEE Trans. Pattern Anal. Mach. Intell.8
2025 Calibration-Free Raw Image Denoising via Fine-Grained Noise Estimation
abstract
Image denoising has progressed significantly due to the development of effective deep denoisers. To improve the performance in real-world scenarios, recent trends prefer to formulate superior noise models to generate realistic training data, or estimate noise levels to steer non-blind denoisers. In this paper, we bridge both strategies by presenting an innovative noise estimation and realistic noise synthesis pipeline. Specifically, we integrates a fine-grained statistical noise model and contrastive learning strategy, with a unique data augmentation to enhance learning ability. Then, we use this model to estimate noise parameters on evaluation dataset, which are subsequently used to craft camera-specific noise distribution and synthesize realistic noise. One distinguishing feature of our methodology is its adaptability: our pre-trained model can directly estimate unknown cameras, making it possible to unfamiliar sensor noise modeling using only testing images, without calibration frames or paired training data. Another highlight is our attempt in estimating parameters for fine-grained noise models, which extends the applicability to even more challenging low-light conditions. Through empirical testing, our calibration-free pipeline demonstrates effectiveness in both normal and low-light scenarios, further solidifying its utility in real-world noise synthesis and denoising tasks.
Yunhao Zou, Ying Fu 0001, Yulun Zhang 0001, Tao Zhang 0042, Chenggang Yan 0001, Radu Timofte
IEEE Trans. Pattern Anal. Mach. Intell.6
2025 High-Precision Dichotomous Image Segmentation With Frequency and Scale Awareness
abstract
Dichotomous image segmentation (DIS) with rich fine-grained details within a single image is a challenging task. Despite the plausible results achieved by deep learning-based methods, most of them fail to segment generic objects when the boundary is cluttered with the background. In fact, the gradual decrease in feature map resolution during the encoding stage and the misleading texture clue may be the main issues. To handle these issues, we devise a novel frequency- and scale-aware deep neural network (FSANet) for high-precision DIS. The core of our proposed FSANet is twofold. First, a multimodality fusion (MF) module that integrates the information in spatial and frequency domains is adopted to enhance the representation capability of image features. Second, a collaborative scale fusion module (CSFM) which deviates from the traditional serial structures is introduced to maintain high resolution during the entire feature encoding stage. In the decoder side, we introduce hierarchical context fusion (HCF) and selective feature fusion (SFF) modules to infer the segmentation results from the output features of the CSFM module. We conduct extensive experiments on several benchmark datasets and compare our proposed method with existing state-of-the-art (SOTA) methods. The experimental results demonstrate that our FSANet achieves superior performance both qualitatively and quantitatively. The code will be made available at https://github.com/chasecjg/FSANet.
Qiuping Jiang, Jinguang Cheng, Zongwei Wu, Runmin Cong, Radu Timofte
IEEE Trans. Neural Networks Learn. Syst.5
2024 NILUT: Conditional Neural Implicit 3D Lookup Tables for Image Enhancement
abstract
3D lookup tables (3D LUTs) are a key component for image enhancement. Modern image signal processors (ISPs) have dedicated support for these as part of the camera rendering pipeline. Cameras typically provide multiple options for picture styles, where each style is usually obtained by applying a unique handcrafted 3D LUT. Current approaches for learning and applying 3D LUTs are notably fast, yet not so memory-efficient, as storing multiple 3D LUTs is required. For this reason and other implementation limitations, their use on mobile devices is less popular. In this work, we propose a Neural Implicit LUT (NILUT), an implicitly defined continuous 3D color transformation parameterized by a neural network. We show that NILUTs are capable of accurately emulating real 3D LUTs. Moreover, a NILUT can be extended to incorporate multiple styles into a single network with the ability to blend styles implicitly. Our novel approach is memory-efficient, controllable and can complement previous methods, including learned ISPs. Code at https://github.com/mv-lab/nilut
Marcos V. Conde, Javier Vazquez-Corral, Michael S. Brown, Radu Timofte
AAAI4
2024 Babel-ImageNet: Massively Multilingual Evaluation of Vision-and-Language Representations
abstract
Vision-and-language (VL) models with separate encoders for each modality (e.g., CLIP) have become the go-to models for zero-shot image classification and image-text retrieval.They are, however, mostly evaluated in English as multilingual benchmarks are limited in availability.We introduce Babel-ImageNet, a massively multilingual benchmark that offers (partial) translations of ImageNet labels to 100 languages, built without machine translation or manual annotation.We instead automatically obtain reliable translations by linking them -via shared WordNet synsets -to Babel-Net, a massively multilingual lexico-semantic network.We evaluate 11 public multilingual CLIP models on zero-shot image classification (ZS-IC) on our benchmark, demonstrating a significant gap between English ImageNet performance and that of high-resource languages (e.g., German or Chinese), and an even bigger gap for low-resource languages (e.g., Sinhala or Lao).Crucially, we show that the models' ZS-IC performance highly correlates with their performance in image-text retrieval, validating the use of Babel-ImageNet to evaluate multilingual models for the vast majority of languages without gold image-text data.Finally, we show that the performance of multilingual CLIP can be drastically improved for low-resource languages with parameter-efficient languagespecific training.We make our code and data publicly available
Gregor Geigle, Radu Timofte, Goran Glavas
ACL (1)2
2024 Deep Equilibrium Diffusion Restoration with Parallel Sampling
abstract
Diffusion model-based image restoration (IR) aims to use diffusion models to recover high-quality (HQ) images from degraded images, achieving promising performance. Due to the inherent property of diffusion models, most existing methods need long serial sampling chains to restore HQ images step-by-step, resulting in expensive sampling time and high computation costs. Moreover, such long sampling chains hinder understanding the relationship between inputs and restoration results since it is hard to compute the gra-dients in the whole chains. In this work, we aim to rethink the diffusion model-based IR models through a different per-spective, i.e., a deep equilibrium (DEQ) fixed point system, called DeqIR. Specifically, we derive an analytical solution by modeling the entire sampling chain in these IR models as a joint multivariate fixed point system. Based on the analyti-cal solution, we can conduct parallel sampling and restore HQ images without training. Furthermore, we compute fast gradients via DEQ inversion and found that initialization optimization can boost image quality and control the gen-eration direction. Extensive experiments on benchmarks demonstrate the effectiveness of our method on typical IR tasks and real-world settings.
Jiezhang Cao, Kai Zhang 0008, Yulun Zhang 0001, Radu Timofte, Luc Van Gool
CVPR5
2024 Real-World Mobile Image Denoising Dataset with Efficient Baselines
abstract
The recently increased role of mobile photography has raised the standards of on-device photo processing tremendously. Despite the latest advancements in camera hardware, the mobile camera sensor area cannot be increased significantly due to physical constraints, leading to a pixel size of 0.6-2.0μm, which results in strong image noise even in moderate lighting conditions. In the era of deep learning, one can train a CNN model to perform robust image denoising. However, there is still a lack of a substantially diverse dataset for this task. To address this problem, we introduce a novel Mobile Image Denoising Dataset (MIDD) comprising over 400,000 noisy / noise-free image pairs captured under various conditions by 20 different mobile camera sensors. Additionally, we propose a new DPreview test set consisting of data from 294 different cameras for precise model evaluation. Furthermore, we present the efficient baseline model SplitterNet for the considered mobile image denoising task that achieves high numerical and visual results, while being able to process 8MP photos directly on smartphone GPUs in under one second. Thereby outperforming models with similar runtimes. This model is also compatible with recent mobile NPUs, demonstrating an even higher speed when deployed on them. The conducted experiments demonstrate high robustness of the proposed solution when applied to images from previously unseen sensors, showing its high generalizability. The datasets, code and models can be found on the official project website11https://people.ee.ethz.ch/-ihnatova/midd.html22https://github.com/rflepp/Efficient_Mobile_Denoising_Models.
Roman Flepp, Andrey Ignatov, Radu Timofte, Luc Van Gool
CVPR3
2024 Single-Model and Any-Modality for Video Object Tracking
abstract
In the realm of video object tracking, auxiliary modalities such as depth, thermal, or event data have emerged as valuable assets to complement the RGB trackers. In practice, most existing RGB trackers learn a single set of parameters to use them across datasets and applications. However, a similar single-model unification for multi-modality tracking presents several challenges. These challenges stem from the inherent heterogeneity of inputs - each with modality-specific representations, the scarcity of multi-modal datasets, and the absence of all the modalities at all times. In this work, we introduce Un-Track, a Unified Tracker of a single set of parameters for any modality. To handle any modality, our method learns their common latent space through low-rank factorization and reconstruction techniques. More importantly, we use only the RGB-X pairs to learn the common latent space. This unique shared representation seamlessly binds all modalities together, enabling effective unification and accommodating any missing modality, all within a single transformer-based architecture. Our Un-Track achieves +8.1 absolute F-score gain, on the DepthTrack dataset, by introducing only +2.14 (over 21.50) GFLOPs with +6.6M (over 93M) parameters, through a simple yet efficient prompting strategy. Extensive comparisons on five benchmark datasets with different modalities show that Un-Track surpasses both SOTA unified trackers and modality-specific counterparts, validating our effectiveness and practicality. The source code is publicly available at https://thub.com/Zongwei97/UnTrack.
Zongwei Wu, Jilai Zheng, Xiangxuan Ren, Florin-Alexandru Vasluianu, Chao Ma 0004, Danda Pani Paudel, Luc Van Gool, Radu Timofte
CVPR8
2024 Equivariant Multi-Modality Image Fusion
abstract
Multi-modality image fusion is a technique that combines information from different sensors or modalities, en-abling the fused image to retain complementary features from each modality, such as functional highlights and texture details. However, effective training of such fusion models is challenging due to the scarcity of ground truth fusion data. To tackle this issue, we propose the Equivariant Multi-Modality imAge fusion (EMMA) paradigm for end-to-end self-supervised learning. Our approach is rooted in the prior knowledge that natural imaging responses are equiv-ariant to certain transformations. Consequently, we introduce a novel training paradigm that encompasses a fusion module, a pseudo-sensing module, and an equivariant fusion module. These components enable the net training to follow the principles of the natural sensing-imaging process while satisfying the equivariant imaging prior. Extensive experiments confirm that EMMA yields high-quality fusion results for infraredvisible and medical images, concurrently facilitating downstream multi-modal segmentation and detection tasks. The code is available at https://github.com/Zhaozixiang1228/MMIF-EMMA.
Zixiang Zhao, Haowen Bai, Jiangshe Zhang 0001, Yulun Zhang 0001, Kai Zhang 0008, Radu Timofte, Luc Van Gool
CVPR8
2024 InstructIR: High-Quality Image Restoration Following Human Instructions
Marcos V. Conde, Gregor Geigle, Radu Timofte
ECCV (36)3
2024 MoVideo: Motion-Aware Video Generation with Diffusion Model
Jingyun Liang, Yuchen Fan 0001, Kai Zhang 0008, Radu Timofte, Luc Van Gool
ECCV (44)4
2024 Rawformer: Unpaired Raw-to-Raw Translation for Learnable Camera ISPs
Georgy Perevozchikov, Nancy Mehta, Mahmoud Afifi, Radu Timofte
ECCV (36)4
2024 Dataset Growth
Ziheng Qin, Zhaopan Xu, Zangwei Zheng, Zebang Cheng, Hao Tang 0005, Baigui Sun, Xiaojiang Peng, Radu Timofte, Hongxun Yao, Kai Wang 0036, Yang You 0001
ECCV (9)10
2024 Towards Image Ambient Lighting Normalization
Florin-Alexandru Vasluianu, Tim Seizinger, Zongwei Wu, Radu Timofte
ECCV (70)5
2024 African or European Swallow? Benchmarking Large Vision-Language Models for Fine-Grained Object Classification
abstract
Recent Large Vision-Language Models (LVLMs) demonstrate impressive abilities on numerous image understanding and reasoning tasks.The task of fine-grained object classification (e.g., distinction between animal species), however, has been probed insufficiently, despite its downstream importance.We fill this evaluation gap by creating FOCI (Fine-grained Object ClassIfication), a difficult multiple-choice benchmark for fine-grained object classification, from existing object classification datasets: (1) multiple-choice avoids ambiguous answers associated with casting classification as open-ended QA task; (2) we retain classification difficulty by mining negative labels with a CLIP model.FOCI complements five popular classification datasets with four domain-specific subsets from ImageNet-21k.We benchmark 12 public LVLMs on FOCI and show that it tests for a complementary skill to established image understanding and reasoning benchmarks.Crucially, CLIP models exhibit dramatically better performance than LVLMs.Since the image encoders of LVLMs come from these CLIP models, this points to inadequate alignment for fine-grained object distinction between the encoder and the LLM and warrants (pre)training data with more fine-grained annotation.We release our code at https:// github.com/gregor-ge/FOCI-Benchmark.
Gregor Geigle, Radu Timofte, Goran Glavas
EMNLP2
2024 Does Object Grounding Really Reduce Hallucination of Large Vision-Language Models?
abstract
Large vision-language models (LVLMs) have recently dramatically pushed the state of the art in image captioning and many image understanding tasks (e.g., visual question answering).LVLMs, however, often hallucinate and produce captions that mention concepts that cannot be found in the image.These hallucinations erode the trustworthiness of LVLMs and are arguably among the main obstacles to their ubiquitous adoption.Recent work suggests that addition of grounding objectives-those that explicitly align image regions or objects to text spans-reduces the amount of LVLM hallucination.Although intuitive, this claim is not empirically justified as the reduction effects have been established, we argue, with flawed evaluation protocols that (i) rely on data (i.e., MSCOCO) that has been extensively used in LVLM training and (ii) measure hallucination via question answering rather than open-ended caption generation.In this work, in contrast, we offer the first systematic analysis of the effect of fine-grained object grounding on LVLM hallucination under an evaluation protocol that more realistically captures LVLM hallucination in open generation.Our extensive experiments over three backbone LLMs reveal that grounding objectives have little to no effect on object hallucination in open caption generation."A white hound and a cat looking at the camera" "A white hound and a cat looking at the camera"
Gregor Geigle, Radu Timofte, Goran Glavas
EMNLP2
2024 Multilingual Audio-Visual Speech Recognition with Hybrid CTC/RNN-T Fast Conformer
abstract
Humans are adept at leveraging visual cues from lip movements for recognizing speech in adverse listening conditions. Audio-Visual Speech Recognition (AVSR) models follow similar approach to achieve robust speech recognition in noisy conditions. In this work, we present a multilingual AVSR model incorporating several enhancements to improve performance and audio noise robustness. Notably, we adapt the recently proposed Fast Conformer model to process both audio and visual modalities using a novel hybrid CTC/RNN-T architecture. We increase the amount of audio-visual training data for six distinct languages, generating automatic transcriptions of unlabelled multilingual datasets (VoxCeleb2 and AVSpeech). Our proposed model achieves new state-of-the-art performance on the LRS3 dataset, reaching WER of 0.8%. On the recently introduced MuAViC benchmark, our model yields an absolute average-WER reduction of 11.9% in comparison to the original baseline. Finally, we demonstrate the ability of the proposed model to perform audio-only, visual-only, and audio-visual speech recognition at test time.
Maxime Burchi, Krishna C. Puvvada, Jagadeesh Balam, Boris Ginsburg, Radu Timofte
ICASSP5
2024 Streaming Neural Images
abstract
Implicit Neural Representations (INRs) are a novel paradigm for signal representation that have attracted considerable interest for image compression. INRs offer unprecedented advantages in signal resolution and memory efficiency, enabling new possibilities for compression techniques. However, the existing limitations of INRs for image compression have not been sufficiently addressed in the literature. In this work, we explore the critical yet overlooked limiting factors of INRs, such as computational cost, unstable performance, and robustness. Through extensive experiments and empirical analysis, we provide a deeper and more nuanced understanding of implicit neural image compression methods such as Fourier Feature Networks and Siren. Our work also offers valuable insights for future research in this area.
Marcos V. Conde, Andy Bigos, Radu Timofte
ICIP3
2024 Toward Efficient Deep Blind Raw Image Restoration
abstract
Multiple low-vision tasks such as denoising, deblurring and super-resolution depart from a sRGB image and further reduce the degradations, improving the perceptual quality. However, modeling the degradations in the sRGB domain is complicated because of the Image Signal Processor (ISP) transformation. Despite of this known issue, very few methods in the literature work directly with sensor RAW images. In this work we tackle image restoration directly in the RAW domain. We design a new realistic degradation pipeline for training deep blind RAW restoration models. Our pipeline considers realistic sensor noise, motion blur, camera shake, and other common degradations. The models trained with our pipeline and data from multiple sensors, can successfully reduce noise and blur, and recover details in real RAW images captured from different cameras in-the-wild. To the best of our knowledge, this is the most exhaustive analysis on RAW image restoration.
Marcos V. Conde, Florin-Alexandru Vasluianu, Radu Timofte
ICIP3
2024 Simple Image Signal Processing using Global Context Guidance
abstract
In modern smartphone cameras, the Image Signal Processor (ISP) is the core element that converts the RAW readings from the sensor into perceptually pleasant RGB images for the end users. The ISP is typically proprietary and handcrafted and consists of several blocks such as white balance, color correction, and tone mapping. Deep learning-based ISPs aim to transform RAW images into DSLR-like RGB images using deep neural networks. However, most learned ISPs are trained using patches (small regions) due to computational limitations. Such methods lack global context, which limits their efficacy on full-resolution images and harms their ability to capture global properties such as color constancy or illumination. First, we propose a novel module that can be integrated into any neural ISP to capture the global context information from the full RAW images. Second, we propose an efficient and simple neural ISP that utilizes our proposed module. Our model achieves state-of-the-art results on different benchmarks using diverse and real smartphone images.
Omar Elezabi, Marcos V. Conde, Radu Timofte
ICIP3
2024 SFNet - A Spatial-Frequency Domain Neural Network For Image Lens Flare Removal
abstract
High-intensity light sources in the scene can cause undesired internal reflections between the multiple optical elements of lenses, resulting in loss of contrast and color change. This effect, known as a lens flare, can have artistic value, but it can also limit the performance of downstream tasks. Professional cameras and lenses have complex optical systems with an increased number of elements, designed to control reflections and refractions for optimal light convergence. However, lens flare is still a challenging problem for professional image acquisition, especially due to the limited information published by manufacturers. In this work, we propose an end-to-end deep learning solution for image lens flare removal and a novel dataset, covering popular DSLR/DSLM optical systems. Our model combines information from both the spatial and frequency domains of the image, leveraging the spatial domain local features and the global features in the frequency domain to reconstruct the flare-affected image. Our model achieves state-of-the-art results, outperforming well-established image restoration architectures for image lens flare removal.
Florin-Alexandru Vasluianu, Zongwei Wu, Radu Timofte
ICIP3
2024 Stereo Risk: A Continuous Modeling Approach to Stereo Matching
abstract
We introduce Stereo Risk, a new deep-learning approach to solve the classical stereo-matching problem in computer vision. As it is well-known that stereo matching boils down to a per-pixel disparity estimation problem, the popular state-of-the-art stereo-matching approaches widely rely on regressing the scene disparity values, yet via discretization of scene disparity values. Such discretization often fails to capture the nuanced, continuous nature of scene depth. Stereo Risk departs from the conventional discretization approach by formulating the scene disparity as an optimal solution to a continuous risk minimization problem, hence the name "stereo risk". We demonstrate that $L^1$ minimization of the proposed continuous risk function enhances stereo-matching performance for deep networks, particularly for disparities with multi-modal probability distributions. Furthermore, to enable the end-to-end network training of the non-differentiable $L^1$ risk optimization, we exploited the implicit function theorem, ensuring a fully differentiable network. A comprehensive analysis demonstrates our method's theoretical soundness and superior performance over the state-of-the-art methods across various benchmark datasets, including KITTI 2012, KITTI 2015, ETH3D, SceneFlow, and Middlebury 2014.
Ce Liu 0004, Suryansh Kumar 0001, Shuhang Gu, Radu Timofte, Yao Yao 0008, Luc Van Gool
ICML4
2024 See More Details: Efficient Image Super-Resolution by Experts Mining
abstract
Reconstructing high-resolution (HR) images from low-resolution (LR) inputs poses a significant challenge in image super-resolution (SR). While recent approaches have demonstrated the efficacy of intricate operations customized for various objectives, the straightforward stacking of these disparate operations can result in a substantial computational burden, hampering their practical utility. In response, we introduce SeemoRe, an efficient SR model employing expert mining. Our approach strategically incorporates experts at different levels, adopting a collaborative methodology. At the macro scale, our experts address rank-wise and spatial-wise informative features, providing a holistic understanding. Subsequently, the model delves into the subtleties of rank choice by leveraging a mixture of low-rank experts. By tapping into experts specialized in distinct key factors crucial for accurate SR, our model excels in uncovering intricate intra-feature details. This collaborative approach is reminiscent of the concept of “see more", allowing our model to achieve an optimal performance with minimal computational costs in efficient settings.
Eduard Zamfir, Zongwei Wu, Nancy Mehta, Yulun Zhang 0001, Radu Timofte
ICML5
2024 Event-Free Moving Object Segmentation from Moving Ego Vehicle
abstract
Moving object segmentation (MOS) in dynamic scenes is an important, challenging, but under-explored research topic for autonomous driving, especially for sequences obtained from moving ego vehicles. Most segmentation methods leverage motion cues obtained from optical flow maps. However, since these methods are often based on optical flows that are pre-computed from successive RGB frames, this neglects the temporal consideration of events occurring within the inter-frame, consequently constraining its ability to discern objects exhibiting relative staticity but genuinely in motion. To address these limitations, we propose to exploit event cameras for better video understanding, which provide rich motion cues without relying on optical flow. To foster research in this area, we first introduce a novel large-scale dataset called DSEC-MOS for moving object segmentation from moving ego vehicles, which is the first of its kind. For benchmarking, we select various mainstream methods and rigorously evaluate them on our dataset. Subsequently, we devise EmoFormer, a novel network able to exploit the event data. For this purpose, we fuse the event temporal prior with spatial semantic maps to distinguish genuinely moving objects from the static background, adding another level of dense supervision around our object of interest. Our proposed network relies only on event data for training but does not require event input during inference, making it directly comparable to frame-only methods in terms of efficiency and more widely usable in many application cases. The exhaustive comparison highlights a significant performance improvement of our method over all other methods. The source code and dataset are publicly available at: https://github.com/ZZYZhou/DSEC-MOS.
Zhuyun Zhou, Zongwei Wu, Danda Pani Paudel, Rémi Boutteau, Fan Yang 0019, Luc Van Gool, Radu Timofte, Dominique Ginhac
IROS7
2024 BSRAW: Improving Blind RAW Image Super-Resolution
abstract
In smartphones and compact cameras, the Image Signal Processor (ISP) transforms the RAW sensor image into a human-readable sRGB image. Most popular super-resolution methods depart from a sRGB image and upscale it further, improving its quality. However, modeling the degradations in the sRGB domain is complicated because of the non-linear ISP transformations. Despite this known issue, only a few methods work directly with RAW images and tackle real-world sensor degradations.We tackle blind image super-resolution in the RAW domain. We design a realistic degradation pipeline tailored specifically for training models with raw sensor data. Our approach considers sensor noise, defocus, exposure, and other common issues. Our BSRAW models trained with our pipeline can upscale real-scene RAW images and improve their quality. As part of this effort, we also present a new DSLM dataset and benchmark for this task.
Marcos V. Conde, Florin-Alexandru Vasluianu, Radu Timofte
WACV3
2024 Guest Editorial: Advanced image restoration and enhancement in the wild
abstract
Image restoration and enhancement has always been a fundamental task in computer vision and is widely used in numerous applications, such as surveillance imaging, remote sensing, and medical imaging. In recent years, remarkable progress has been witnessed with deep learning techniques. Despite the promising performance achieved on synthetic data, compelling research challenges remain to be addressed in the wild. These include: (i) degradation models for low-quality images in the real world are complicated and unknown, (ii) paired low-quality and high-quality data are difficult to acquire in the real world, and a large quantity of real data are provided in an unpaired form, (iii) it is challenging to incorporate cross-modal information provided by advanced imaging techniques (e.g. RGB-D camera) for image restoration, (iv) real-time inference on edge devices is important for image restoration and enhancement methods, and (v) it is difficult to provide the confidence or performance bounds of a learning-based method on different images/regions. This special issue invites original contributions in datasets, innovative architectures, and training methods for image restoration and enhancement to address these and other challenges. In this Special Issue, we have received 17 papers, of which 8 papers underwent the peer review process, while the rest were desk-rejected. Among these reviewed papers, 5 papers have been accepted and 3 papers have been rejected as they did not meet the criteria of IET Computer Vision. Thus, the overall submissions were of high quality, which marks the success of this Special Issue. The five eventually accepted papers can be clustered into two categories, namely video reconstruction and image super-resolution. The first category of papers aims at reconstructing high-quality videos. The papers in this category are of Zhang et al., Gu et al., and Xu et al. The second category of papers studies the task of image super-resolution. The papers in this category are of Dou et al. and Yang et al. A brief presentation of each of the paper in this special issue is as follows. Zhang et al. propose a point-image fusion network for event-based frame interpolation. Temporal information in event streams plays a critical role in this task as it provides temporal context cues complementary to images. Previous approaches commonly transform the unstructured event data to structured data formats through voxelisation and then employ advanced CNNs to extract temporal information. However, the voxelisation operation inevitably leads to information loss and introduces redundant computation. To address these limitations, the proposed method directly extracts temporal information from the events at the point level without relying on any voxelisation operation. Afterwards, a fusion module is adopted to aggregate complementary cues from both points and images for frame interpolation. Experiments on both synthetic and real-world datasets show that their method produces state-of-the-art accuracy with high efficiency. Gu et al. develop a temporal shift reconstruction network for compressive video sensing. To exploit the temporal cues between adjacent frames during the reconstruction of videos, most previous approaches commonly preform alignment between initial reconstructions. However, the estimated motions are usually too coarse to provide accurate temporal information. To remedy this, the proposed network employs stacked temporal shift reconstruction blocks to enhance the initial reconstruction progressively. Within each block, an efficient temporal shift operation is used to capture temporal structures in addition to computational overheads. Then, a bidirectional alignment module is adopted to capture the temporal dependencies in a video sequence. Different from previous methods that only extract supplementary information from the key frames, the proposed alignment module can receive temporal information from the whole video sequence via bidirectional propagations. Experiments demonstrate the superior performance of the proposed method. Qu et al. propose a lightweight video frame interpolation network with a three-scale encoding-decoding structure. Specifically, multi-scale motion information is first extracted from the input video. Then, recurrent convolutional layers are adopted to refine the resultant features. Afterwards, the resultant features are aggregated to generate high-quality interpolated frames. Experimental results on the CelebA and Helen datasets show that the proposed method outperforms state-of-the-art methods while using fewer parameters. Dou et al. introduce a decoder structure-guided CNN-Transformer network for face super-resolution. Most previous approaches follow a multi-task learning paradigm to perform landmark detection while super-resolving the low-resolution images. However, these methods require additional annotation cost, and the extracted facial prior structures are usually of low quality. To address these issues, the proposed network employs a global-local feature extraction unit to extract the global structure while capturing local texture details. In addition, a multi-state fusion module is incorporated to aggregate embeddings from different stages. Experiments show that the proposed method surpasses previous approaches by notable margins. Yang et al. study the problem of blind super-resolution and propose a method to exploit degradation information through degradation representation learning. Specifically, a generative adversarial network is employed to model the degradation process from HR images to LR images and constrain the data distribution of the synthetic LR images. Then, the learnt representation is adopted to super-resolve the input low-resolution images using a transformer-based SR network. Experiments on both synthetic and real-world datasets demonstrate the effectiveness and superiority of the proposed method. Longguang Wang received his BE and PhD degrees from Shandong University and National University of Defence Technology (NUDT) in 2015 and 2022, respectively. He is currently an assistant professor with Aviation University of Air Force. He authored more than 40 peer-reviewed journals and conference publications (including TPAMI, TIP, CVPR, ICCV, and ECCV). He has organised three workshops at CVPR 2022 and 2023. His research interests include low-level vision and 3D vision, particularly on image restoration, image enhancement, image generation, depth estimation, point cloud understanding, and network acceleration. He received the CSIG Excellent Doctoral Dissertation Nomination Award in 2022 (17 nationwide). Juncheng Li received the Ph.D. degree from the School of Computer Science and Technology, East China Normal University, in 2021. He also worked as a Postdoctoral Fellow at the Center for Mathematical Artificial Intelligence, The Chinese University of Hong Kong. He is currently an assistant professor with Shanghai University. His research interests include artificial intelligence and its applications to computer vision (e.g. image segmentation) and image processing (e.g. image super-resolution, image denoising, and image dehazing). He has published more than 25 papers in top journals and conferences, including TIP, TNNLS, TMM, ECCV, ICCV, AAAI, ACMMM, and IJCAI. He also received several premium awards, including the Shanghai Outstanding Ph.D. Graduates, CUHK Research Fellowship Scheme, and the winner of 2019 ICCV-AIM. Naoto Yokoya received the M.Eng. and Ph.D. degrees from the Department of Aeronautics and Astronautics, The University of Tokyo, Tokyo, Japan, in 2010 and 2013, respectively. From 2013 to 2017, he was an assistant professor with The University of Tokyo. From 2015 to 2017, he was an Alexander von Humboldt Fellow, working at the German Aerospace Center, Oberpfaffenhofen, Germany and at the Technical University of Munich, Munich, Germany. He is currently a lecturer with The University of Tokyo and a unit leader with the RIKEN Center for Advanced Intelligence Project, Tokyo, where he leads the Geoinformatics Unit. His research interests include image processing, data fusion, and machine learning for understanding remote sensing images with applications to disaster management. Dr. Yokoya received the First Place in the 2017 IEEE Geoscience and Remote Sensing Society (GRSS) Data Fusion Contest organised by the IEEE Image Analysis and Data Fusion Technical Committee (IADF TC). From 2019 to 2021, he was the Chair and the Co-Chair (2017–2019) of the IEEE GRSS IADF TC. Since 2018, he has been an associate editor of IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing (JSTARS). Radu Timofte received his Ph.D. degree in Electrical Engineering from the KU Leuven, Belgium, in 2013. Currently, he is a professor and holds the Chair for Computer Science IV (Computer Vision) at the University of Wurzburg, Germany. Also, he is a lecturer and a group leader at ETH Zurich, Switzerland. He is a member of the editorial board of top journals such as IEEE TPAMI, Elsevier's CVIU and NEUCOM, and SIAM's SIIMS. He regularly serves as an area chair and as a reviewer for top conferences such as CVPR, ICCV, IJCAI, and ECCV. His work received several awards. Radu Timofte is the 2022 awardee of the Alexander von Humboldt Professorship for Artificial Intelligence. He is a co-founder of Merantix and a co-organiser of NTIRE, CLIC, AIM, Mobile AI, and PIRM workshops and challenges. His current research interests include deep learning, mobile AI, visual tracking, computational photography, and image/video compression, restoration, enhancement, and manipulation. Yulan Guo received the B.E. and Ph.D. degrees from NUDT in 2008 and 2015, respectively. He has authored over 100 articles at highly referred journals and conferences. His current research interests focus on 3D vision, particularly on 3D feature learning, 3D modelling, 3D object recognition, and scene understanding. He served as an associate editor for IEEE Transactions on Image Processing, IET Computer Vision, IET Image Processing, and Computers & Graphics. He also served as an area chair for CVPR 2023/2021, ICCV 2021, and ACM Multimedia 2021. He organised several tutorials, workshops, and challenges in prestigious conferences, such as CVPR 2016, CVPR 2019, ICCV 2021, 3DV 2021, CVPR 2022, ICPR 2022, and ECCV 2022. He is a senior member of IEEE and ACM. Data sharing is not applicable to this article as no new data were created or analysed in this study. Longguang Wang received his B.E. and Ph.D. degrees from Shandong University and National University of Defense Technology in 2015 and 2022, respectively. He is currently an assistant professor with Aviation University of Air Force. He authored more than 60 peer reviewed journal and conference publications (including TPAMI, TIP, CVPR, ICCV and ECCV). He served as a reviewer for more than 10 international journals (including TPAMI and TIP) and conferences (including CVPR, ICCV and ECCV). He has organized workshops at CVPR 2022/2023/2024. His research interests include low-level vision and 3D vision, particularly on image restoration, image generation, point cloud understanding, and network acceleration. His received the CSIG Excellent Doctoral Dissertation Nomination Award in 2022 (17 nationalwide). Juncheng Li received the Ph.D. degree from the School of Computer Science and Technology, East China Normal University, in 2021. He also worked as a Postdoctoral Fellow at the Center for Mathematical Artificial Intelligence, The Chinese University of Hong Kong. He is currently an assistant professor with Shanghai University. His research interests include artificial intelligence and its applications to computer vision (e.g. image segmentation) and image processing (e.g. image super-resolution, image denoising, and image dehazing). He has published more than 25 papers in top journals and conferences, including TIP, TNNLS, TMM, ECCV, ICCV, AAAI, ACMMM and IJCAI. He also received several premium awards, including the Shanghai Outstanding Ph.D. Graduates, CUHK Research Fellowship Scheme, the winner of 2019 ICCV-AIM, etc. Meanwhile, he served as a reviewer for more than 20 international journals and conferences. Naoto Yokoya received the M.Eng. and Ph.D. degrees from the Department of Aeronautics and Astronautics, The University of Tokyo, Tokyo, Japan, in 2010 and 2013, respectively. From 2013 to 2017, he was an Assistant Professor with The University of Tokyo. From 2015 to 2017, he was an Alexander von Humboldt Fellow, working at the German Aerospace Center, Oberpfaffenhofen, Germany, and at the Technical University of Munich, Munich, Germany. He is currently a Lecturer with The University of Tokyo, and a Unit Leader with the RIKEN Center for Advanced Intelligence Project, Tokyo, where he leads the Geoinformatics Unit. His research interests include image processing, data fusion, and machine learning for understanding remote sensing images, with applications to disaster management. Dr. Yokoya received the First Place in the 2017 IEEE Geoscience and Remote Sensing Society (GRSS) Data Fusion Contest organized by the IEEE Image Analysis and Data Fusion Technical Committee (IADF TC). From 2019 to 2021, he was the Chair and the Co-Chair (2017–2019) of the IEEE GRSS IADF TC. Since 2018, he has been an Associate Editor of IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing (JSTARS). Radu Timofte received his Ph.D. degree in Electrical Engineering from the KU Leuven, Belgium, in 2013. Currently, he is a professor and holds the Chair for Computer Science IV (Computer Vision) at theUniversity of Wurzburg, Germany. He is a member of the editorial board of top journals such as IEEE TPAMI, Elsevier's CVIU and NEUCOM, and SIAM's SIIMS. He regularly serves as an area chair and as a reviewer for top conferences such as CVPR, ICCV, IJCAI and ECCV. His work received several awards. Radu Timofte is the 2022 awardee of an Alexandervon Humboldt Professorship for Artificial Intelligence. He is a co-founder of Merantix and a co-organizer of NTIRE, CLIC, AIM, Mobile AI and PIRM workshops and challenges. His current research interests include deep learning, mobile AI, visual tracking, computational photography, image/video compression, restoration, enhancement and manipulation. Yulan Guo received the B.E. and Ph.D. degrees from National University of Defense Technology (NUDT) in 2008 and 2015, respectively. He has authored over 100 articles at highly referred journals and conferences. His current research interests focus on 3D vision, particularly on 3D feature learning, 3D modeling, 3D object recognition, and scene understanding. He served as an associate editor for IEEE Transactions on Image Processing, IET Computer Vision, IET Image Processing, and Computers & Graphics. He also served as an area chair for CVPR 2023/2021, ICCV 2021, and ACM Multimedia 2021. He organized several tutorials, workshops, and challenges in prestigious conferences, such as CVPR 2016, CVPR 2019, ICCV 2021, 3DV 2021, CVPR 2022, ICPR 2022 and ECCV 2022. He is a Senior Member of IEEE and ACM.
Longguang Wang, Juncheng Li 0003, Naoto Yokoya, Radu Timofte, Yulan Guo
IET Comput. Vis.4
2024 VRT: A Video Restoration Transformer
abstract
Video restoration aims to restore high-quality frames from low-quality frames. Different from single image restoration, video restoration generally requires to utilize temporal information from multiple adjacent but usually misaligned video frames. Existing deep methods generally tackle with this by exploiting a sliding window strategy or a recurrent architecture, which are restricted by frame-by-frame restoration. In this paper, we propose a Video Restoration Transformer (VRT) with parallel frame prediction ability. More specifically, VRT is composed of multiple scales, each of which consists of two kinds of modules: temporal reciprocal self attention (TRSA) and parallel warping. TRSA divides the video into small clips, on which reciprocal attention is applied for joint motion estimation, feature alignment and feature fusion, while self attention is used for feature extraction. To enable cross-clip interactions, the video sequence is shifted for every other layer. Besides, parallel warping is used to further fuse information from neighboring frames by parallel feature warping. Experimental results on five tasks, including video super-resolution, video deblurring, video denoising, video frame interpolation and space-time video super-resolution, demonstrate that VRT outperforms the state-of-the-art methods by large margins (up to 2.16dB) on fourteen benchmark datasets. The codes are available at https://github.com/JingyunLiang/VRT.
Jingyun Liang, Jiezhang Cao, Yuchen Fan 0001, Kai Zhang 0008, Yawei Li 0001, Radu Timofte, Luc Van Gool
IEEE Trans. Image Process.7
2023 Graph Transformer GANs for Graph-Constrained House Generation
abstract
We present a novel graph Transformer generative adversarial network (GTGAN) to learn effective graph node relations in an end-to-end fashion for the challenging graph-constrained house generation task. The proposed graph-Transformer-based generator includes a novel graph Transformer encoder that combines graph convolutions and self-attentions in a Transformer to model both local and global interactions across connected and non-connected graph nodes. Specifically, the proposed connected node attention (CNA) and non-connected node attention (NNA) aim to capture the global relations across connected nodes and non-connected nodes in the input graph, respectively. The proposed graph modeling block (GMB) aims to exploit local vertex interactions based on a house layout topology. More-over, we propose a new node classification-based discriminator to preserve the high-level semantic and discriminative node features for different house components. Finally, we propose a novel graph-based cycle-consistency loss that aims at maintaining the relative spatial relationships between ground truth and predicted graphs. Experiments on two challenging graph-constrained house generation tasks (i.e., house layout and roof generation) with two public datasets demonstrate the effectiveness of GTGAN in terms of objective quantitative scores and subjective visual realism. New state-of-the-art results are established by large margins on both tasks.
Hao Tang 0005, Zhenyu Zhang 0005, Humphrey Shi, Ling Shao 0001, Nicu Sebe, Radu Timofte, Luc Van Gool
CVPR7
2023 CiaoSR: Continuous Implicit Attention-in-Attention Network for Arbitrary-Scale Image Super-Resolution
abstract
Learning continuous image representations is recently gaining popularity for image super-resolution (SR) because of its ability to reconstruct high-resolution images with arbitrary scales from low-resolution inputs. Existing methods mostly ensemble nearby features to predict the new pixel at any queried coordinate in the SR image. Such a local ensemble suffers from some limitations: i) it has no learnable parameters and it neglects the similarity of the visual features; ii) it has a limited receptive field and cannot ensemble relevant features in a large field which are important in an image. To address these issues, this paper proposes a continuous implicit attention-in-attention network, called CiaoSR. We explicitly design an implicit attention network to learn the ensemble weights for the nearby local features. Furthermore, we embed a scale-aware attention in this implicit attention network to exploit additional non-local information. Extensive experiments on benchmark datasets demonstrate CiaoSR significantly outperforms the existing single image SR methods with the same backbone. In addition, CiaoSR also achieves the state-of-the-art performance on the arbitrary-scale SR task. The effectiveness of the method is also demonstrated on the real-world SR setting. More importantly, CiaoSR can be flexibly integrated into any backbone to improve the SR performance.
Jiezhang Cao, Qin Wang 0013, Yongqin Xian, Yawei Li 0001, Bingbing Ni, Zhiming Pi, Kai Zhang 0008, Yulun Zhang 0001, Radu Timofte, Luc Van Gool
CVPR9
2023 Efficient and Explicit Modelling of Image Hierarchies for Image Restoration
abstract
The aim of this paper is to propose a mechanism to efficiently and explicitly model image hierarchies in the global, regional, and local range for image restoration. To achieve that, we start by analyzing two important properties of natural images including cross-scale similarity and anisotropic image features. Inspired by that, we propose the anchored stripe self-attention which achieves a good balance between the space and time complexity of self-attention and the modelling capacity beyond the regional range. Then we propose a new network architecture dubbed GRL to explicitly model image hierarchies in the Global, Regional, and Local range via anchored stripe self-attention, window self-attention, and channel attention enhanced convolution. Finally, the proposed network is applied to 7 image restoration types, covering both real and synthetic settings. The proposed method sets the new state-of-the-art for several of those. Code will be available at https://github.com/ofsoundof/GRL-Image-Restoration.git.
Yawei Li 0001, Yuchen Fan 0001, Xiaoyu Xiang, Denis Demandolx, Radu Timofte, Luc Van Gool
CVPR6
2023 Single Image Depth Prediction Made Better: A Multivariate Gaussian Take
abstract
Neural-network-based single image depth prediction (SIDP) is a challenging task where the goal is to predict the scene's per-pixel depth at test time. Since the problem, by definition, is illposed, the fundamental goal is to come up with an approach that can reliably model the scene depth from a set of training examples. In the pursuit of perfect depth estimation, most existing state-of-the-art learning techniques predict a single scalar depth value per-pixel. Yet, it is well-known that the trained model has accuracy limits and can predict imprecise depth. Therefore, an SIDP approach must be mindful of the expected depth variations in the model's prediction at test time. Accordingly, we introduce an approach that performs continuous modeling of per-pixel depth, where we can predict and reason about the per-pixel depth and its distribution. To this end, we model per-pixel scene depth using a multivariate Gaussian distribution. Moreover, contrary to the existing uncertainty modeling methods—in the same spirit, where per-pixel depth is assumed to be independent, we introduce per-pixel covariance modeling that encodes its depth dependency w.r.t. all the scene points. Unfortunately, per-pixel depth covariance modeling leads to a computationally expensive continuous loss function, which we solve efficiently using the learned low-rank approximation of the overall covariance matrix. Notably, when tested on benchmark datasets such as KITTI, NYU, and SUN-RGB-D, the SIDP model obtained by optimizing our loss function shows state-of-the-art results. Our method's accuracy (named MG) is among the top on the KITTI depth-prediction benchmark leaderboard11http://www.cvlibs.net/datasets/kitti/eval–depth.php?benchmark=depth–prediction.
Ce Liu 0004, Suryansh Kumar 0001, Shuhang Gu, Radu Timofte, Luc Van Gool
CVPR4
2023 CDDFuse: Correlation-Driven Dual-Branch Feature Decomposition for Multi-Modality Image Fusion
abstract
Multi-modality (MM) image fusion aims to render fused images that maintain the merits of different modalities, e.g., functional highlight and detailed textures. To tackle the challenge in modeling cross-modality features and decomposing desirable modality-specific and modality-shared features, we propose a novel Correlation-Driven feature Decomposition Fusion (CDDFuse) network. Firstly, CDDFuse uses Restormer blocks to extract cross-modality shallow features. We then introduce a dual-branch Transformer-CNN feature extractor with Lite Transformer (LT) blocks leveraging long-range attention to handle low-frequency global features and Invertible Neural Networks (INN) blocks focusing on extracting high-frequency local information. A correlation-driven loss is further proposed to make the low-frequency features correlated while the high-frequency features uncorrelated based on the embedded information. Then, the LT-based global fusion and INN-based local fusion layers output the fused image. Extensive experiments demonstrate that our CDDFuse achieves promising results in multiple fusion tasks, including infrared-visible image fusion and medical image fusion. We also show that CDDFuse can boost the performance in downstream infrared-visible semantic segmentation and object detection in a unified benchmark. The code is available at https://github.om/haozixiang1228/MMIF-CDDFuse.
Zixiang Zhao, Haowen Bai, Jiangshe Zhang 0001, Yulun Zhang 0001, Zudi Lin, Radu Timofte, Luc Van Gool
CVPR7
2023 Retinexformer: One-stage Retinex-based Transformer for Low-light Image Enhancement
abstract
When enhancing low-light images, many deep learning algorithms are based on the Retinex theory. However, the Retinex model does not consider the corruptions hidden in the dark or introduced by the light-up process. Besides, these methods usually require a tedious multi-stage training pipeline and rely on convolutional neural networks, showing limitations in capturing long-range dependencies. In this paper, we formulate a simple yet principled One-stage Retinex-based Framework (ORF). ORF first estimates the illumination information to light up the low-light image and then restores the corruption to produce the enhanced image. We design an Illumination-Guided Transformer (IGT) that utilizes illumination representations to direct the modeling of non-local interactions of regions with different lighting conditions. By plugging IGT into ORF, we obtain our algorithm, Retinexformer. Comprehensive quantitative and qualitative experiments demonstrate that our Retinexformer significantly outperforms state-of-the-art methods on thirteen benchmarks. The user study and application on low-light object detection also reveal the latent practical values of our method. Code is available at https://github.com/caiyuanhao1998/Retinexformer
Yuanhao Cai, Hao Bian, Haoqian Wang, Radu Timofte, Yulun Zhang 0001
ICCV5
2023 SQAD: Automatic Smartphone Camera Quality Assessment and Benchmarking
abstract
Smartphone photography is becoming increasingly popular, but fitting high-performing camera systems within the given space limitations remains a challenge for manufacturers. As a result, powerful mobile camera systems are in high demand. Despite recent progress in computer vision, camera system quality assessment remains a tedious and manual process. In this paper, we present the Smartphone Camera Quality Assessment Dataset (SQAD), which includes natural images captured by 29 devices. SQAD defines camera system quality based on six widely accepted criteria: resolution, color accuracy, noise level, dynamic range, Point Spread Function, and aliasing. Built on thorough examinations in a controlled laboratory environment, SQAD provides objective metrics for quality assessment, overcoming previous subjective opinion scores. Moreover, we introduce the task of automatic camera quality assessment and train deep learning-based models on the collected data to perform a precise quality prediction for arbitrary photos. The dataset, codes and pre-trained models are released at https://github.com/aiff22/SQAD.
Zilin Fang, Andrey Ignatov, Eduard Zamfir, Radu Timofte
ICCV4
2023 Alignment-free HDR Deghosting with Semantics Consistent Transformer
abstract
High dynamic range (HDR) imaging aims to retrieve information from multiple low-dynamic range inputs to generate realistic output. The essence is to leverage the contextual information, including both dynamic and static semantics, for better image generation. Existing methods often focus on the spatial misalignment across input frames caused by the foreground and/or camera motion. However, there is no research on jointly leveraging the dynamic and static context in a simultaneous manner. To delve into this problem, we propose a novel alignment-free network with a Semantics Consistent Transformer (SCTNet) with both spatial and channel attention modules in the network. The spatial attention aims to deal with the intra-image correlation to model the dynamic motion, while the channel attention enables the inter-image intertwining to enhance the semantic consistency across frames. Aside from this, we introduce a novel realistic HDR dataset with more variations in foreground objects, environmental factors, and larger motions. Extensive comparisons on both conventional datasets and ours validate the effectiveness of our method, achieving the best trade-off on the performance and the computational cost. The source code and dataset are available at https://steven-tel.github.io/sctnet/.
Steven Tel, Zongwei Wu, Yulun Zhang 0001, Barthélémy Heyrman, Cédric Demonceaux, Radu Timofte, Dominique Ginhac
ICCV6
2023 Source-free Depth for Object Pop-out
abstract
Depth cues are known to be useful for visual perception. However, direct measurement of depth is often impracticable. Fortunately, though, modern learning-based methods offer promising depth maps by inference in the wild. In this work, we adapt such depth inference models for object segmentation using the objects’ "pop-out" prior in 3D. The "pop-out" is a simple composition prior that assumes objects reside on the background surface. Such compositional prior allows us to reason about objects in the 3D space. More specifically, we adapt the inferred depth maps such that objects can be localized using only 3D information. Such separation, however, requires knowledge about contact surface which we learn using the weak supervision of the segmentation mask. Our intermediate representation of contact surface, and thereby reasoning about objects purely in 3D, allows us to better transfer the depth knowledge into semantics. The proposed adaptation method uses only the depth model without needing the source data used for training, making the learning process efficient and practical. Our experiments on eight datasets of two challenging tasks, namely salient object detection and camouflaged object detection, consistently demonstrate the benefit of our method in terms of both performance and generalizability. The source code is publicly available at https://github.com/Zongwei97/PopNet.
Zongwei Wu, Danda Pani Paudel, Deng-Ping Fan, Shuo Wang 0010, Cédric Demonceaux, Radu Timofte, Luc Van Gool
ICCV7
2023 Spherical Space Feature Decomposition for Guided Depth Map Super-Resolution
abstract
Guided depth map super-resolution (GDSR), as a hot topic in multi-modal image processing, aims to upsample low-resolution (LR) depth maps with additional information involved in high-resolution (HR) RGB images from the same scene. The critical step of this task is to effectively extract domain-shared and domain-private RGB/depth features. In addition, three detailed issues, namely blurry edges, noisy surfaces, and over-transferred RGB texture, need to be addressed. In this paper, we propose the Spherical Space feature Decomposition Network (SSDNet) to solve the above issues. To better model cross-modality features, Restormer block-based RGB/depth encoders are employed for extracting local-global features. Then, the extracted features are mapped to the spherical space to complete the separation of private features and the alignment of shared features. Shared features of RGB are fused with the depth features to complete the GDSR task. Subsequently, a spherical contrast refinement (SCR) module is proposed to further address the detail issues. Patches that are classified according to imperfect categories are input into the SCR module, where the patch features are pulled closer to the ground truth and pushed away from the corresponding imperfect samples in the spherical feature space via contrastive learning. Extensive experiments demonstrate that our method can achieve state-of-the-art results on four test datasets, as well as successfully generalize to real-world scenes. The code is available at https://github.com/Zhaozixiang1228/GDSR-SSDNet.
Zixiang Zhao, Jiangshe Zhang 0001, Chengli Tan, Yulun Zhang 0001, Radu Timofte, Luc Van Gool
ICCV7
2023 DDFM: Denoising Diffusion Model for Multi-Modality Image Fusion
abstract
Multi-modality image fusion aims to combine different modalities to produce fused images that retain the complementary features of each modality, such as functional highlights and texture details. To leverage strong generative priors and address challenges such as unstable training and lack of interpretability for GAN-based generative methods, we propose a novel fusion algorithm based on the denoising diffusion probabilistic model (DDPM). The fusion task is formulated as a conditional generation problem under the DDPM sampling framework, which is further divided into an unconditional generation subproblem and a maximum likelihood subproblem. The latter is modeled in a hierarchical Bayesian manner with latent variables and inferred by the expectation-maximization (EM) algorithm. By integrating the inference solution into the diffusion sampling iteration, our method can generate high-quality fused images with natural image generative priors and cross-modality information from source images. Note that all we required is an unconditional pre-trained generative model, and no fine-tuning is needed. Our extensive experiments indicate that our approach yields promising fusion results in infrared-visible image fusion and medical image fusion. The code is available at https://github.com/Zhaozixiang1228/MMIF-DDFM.
Zixiang Zhao, Haowen Bai, Yuanzhi Zhu 0001, Jiangshe Zhang 0001, Yulun Zhang 0001, Kai Zhang 0008, Deyu Meng, Radu Timofte, Luc Van Gool
ICCV9
2023 Edge Guided GANs with Contrastive Learning for Semantic Image Synthesis
Hao Tang 0005, Xiaojuan Qi 0001, Guolei Sun, Dan Xu 0002, Nicu Sebe, Radu Timofte, Luc Van Gool
ICLR6
2023 VA-DepthNet: A Variational Approach to Single Image Depth Prediction
Ce Liu 0004, Suryansh Kumar 0001, Shuhang Gu, Radu Timofte, Luc Van Gool
ICLR4
2023 Basic Binary Convolution Unit for Binarized Image Restoration Network
Bin Xia 0014, Yulun Zhang 0001, Yapeng Tian, Wenming Yang, Radu Timofte, Luc Van Gool
ICLR6
2023 Knowledge Distillation based Degradation Estimation for Blind Super-Resolution
Bin Xia 0014, Yulun Zhang 0001, Yapeng Tian, Wenming Yang, Radu Timofte, Luc Van Gool
ICLR6
2023 LocalViT: Analyzing Locality in Vision Transformers
abstract
The aim of this paper is to study the influence of locality mechanisms in vision transformers. Transformers originated from machine translation and are particularly good at modelling long-range dependencies within a long sequence. Although the global interaction between the token embeddings could be well modelled by the self-attention mechanism of transformers, what is lacking is a locality mechanism for infor-mation exchange within a local region. In this paper, locality mechanism is systematically investigated by carefully designed controlled experiments. We add locality to vision transformers into the feed-forward network. This seemingly simple solution is inspired by the comparison between feed-forward networks and inverted residual blocks. The importance of locality mechanisms is validated in two ways: 1) A wide range of design choices (activation function, layer placement, expansion ratio) are available for incorporating locality mechanisms and proper choices can lead to a performance gain over the baseline, and 2) The same locality mechanism is successfully applied to vision transformers with different architecture designs, which shows the generalization of the locality concept. For ImageNet2012 classification, the locality-enhanced transformers outperform the baselines Swin-T [1], DeiT-T [2] and PVT-T [3] by 1.0%, 2.6 % and 3.1 % with a negligible increase in the number of parameters and computational effort. Code is available at https://github.com/ofsoundof/LocalViT.
Yawei Li 0001, Kai Zhang 0008, Jiezhang Cao, Radu Timofte, Michele Magno, Luca Benini, Luc Van Gool
IROS4
2023 Object Segmentation by Mining Cross-Modal Semantics
abstract
Multi-sensor clues have shown promise for object segmentation, but inherent noise in each sensor, as well as the calibration error in practice, may bias the segmentation accuracy. In this paper, we propose a novel approach by mining the Cross-Modal Semantics to guide the fusion and decoding of multimodal features, with the aim of controlling the modal contribution based on relative entropy. We explore semantics among the multimodal inputs in two aspects: the modality-shared consistency and the modality-specific variation. Specifically, we propose a novel network, termed XMSNet, consisting of (1) all-round attentive fusion (AF), (2) coarse-to-fine decoder (CFD), and (3) cross-layer self-supervision. On the one hand, the AF block explicitly dissociates the shared and specific representation and learns to weight the modal contribution by adjusting the proportion, region, and pattern, depending upon the quality. On the other hand, our CFD initially decodes the shared feature and then refines the output through specificity-aware querying. Further, we enforce semantic consistency across the decoding layers to enable interaction across network hierarchies, improving feature discriminability. Exhaustive comparison on eleven datasets with depth or thermal clues, and on two challenging tasks, namely salient and camouflage object segmentation, validate our effectiveness in terms of both performance and robustness. The source code is publicly available at https://github.com/Zongwei97/XMSNet.
Zongwei Wu, Zhuyun Zhou, Zhaochong An, Qiuping Jiang, Cédric Demonceaux, Guolei Sun, Radu Timofte
ACM Multimedia8
2023 LART: Neural Correspondence Learning with Latent Regularization Transformer for 3D Motion Transfer
abstract
3D motion transfer aims at transferring the motion from a dynamic input sequence to a static 3D object and outputs an identical motion of the target with high-fidelity and realistic visual effects. In this work, we propose a novel 3D Transformer framework called LART for 3D motion transfer. With carefully-designed architectures, LART is able to implicitly learn the correspondence via a flexible geometry perception. Thus, unlike other existing methods, LART does not require any key point annotations or pre-defined correspondence between the motion source and target meshes and can also handle large-size full-detailed unseen 3D targets. Besides, we introduce a novel latent metric regularization on the Transformer for better motion generation. Our rationale lies in the observation that the decoded motions can be approximately expressed as linearly geometric distortion at the frame level. The metric preservation of motions could be translated to the formation of linear paths in the underlying latent space as a rigorous constraint to control the synthetic motions occurring in the construction of the latent space. The proposed LART shows a high learning efficiency with the need for a few samples from the AMASS dataset to generate motions with plausible visual effects. The experimental results verify the potential of our generative model in applications of motion transfer, content generation, temporal interpolation, and motion denoising. The code is made available: https://github.com/mikecheninoulu/LART.
Haoyu Chen 0001, Hao Tang 0005, Radu Timofte, Luc Van Gool, Guoying Zhao 0001
NeurIPS3
2023 Audio-Visual Efficient Conformer for Robust Speech Recognition
abstract
End-to-end Automatic Speech Recognition (ASR) systems based on neural networks have seen large improvements in recent years. The availability of large scale hand-labeled datasets and sufficient computing resources made it possible to train powerful deep neural networks, reaching very low Word Error Rate (WER) on academic benchmarks. However, despite impressive performance on clean audio samples, a drop of performance is often observed on noisy speech. In this work, we propose to improve the noise robustness of the recently proposed Efficient Conformer Connectionist Temporal Classification (CTC)-based architecture by processing both audio and visual modalities. We improve previous lip reading methods using an Efficient Conformer back-end on top of a ResNet-18 visual front-end and by adding intermediate CTC losses between blocks. We condition intermediate block features on early predictions using Inter CTC residual modules to relax the conditional independence assumption of CTC-based models. We also replace the Efficient Conformer grouped attention by a more efficient and simpler attention mechanism that we call patch attention. We experiment with publicly available Lip Reading Sentences 2 (LRS2) and Lip Reading Sentences 3 (LRS3) datasets. Our experiments show that using audio and visual modalities allows to better recognize speech in the presence of environmental noise and significantly accelerate training, reaching lower WER with 4 times less training steps. Our Audio-Visual Efficient Conformer (AVEC) model achieves state-of-the-art performance, reaching WER of 2.3% and 1.8% on LRS2 and LRS3 test sets. Code and pretrained models are available at https://github.com/burchim/AVEC.
Maxime Burchi, Radu Timofte
WACV2
2023 Perceptual Image Enhancement for Smartphone Real-Time Applications
abstract
Recent advances in camera designs and imaging pipelines allow us to capture high-quality images using smartphones. However, due to the small size and lens limitations of the smartphone cameras, we commonly find artifacts or degradation in the processed images. The most common unpleasant effects are noise artifacts, diffraction artifacts, blur, and HDR overexposure. Deep learning methods for image restoration can successfully remove these artifacts. However, most approaches are not suitable for real-time applications on mobile devices due to their heavy computation and memory requirements.In this paper, we propose LPIENet, a lightweight network for perceptual image enhancement, with the focus on deploying it on smartphones. Our experiments show that, with much fewer parameters and operations, our model can deal with the mentioned artifacts and achieve competitive performance compared with state-of-the-art methods on standard benchmarks. Moreover, to prove the efficiency and reliability of our approach, we deployed the model directly on commercial smartphones and evaluated its performance. Our model can process 2K resolution images under 1 second in mid-level commercial smartphones.
Marcos V. Conde, Florin-Alexandru Vasluianu, Javier Vazquez-Corral, Radu Timofte
WACV4
2023 Fast Online Video Super-Resolution with Deformable Attention Pyramid
abstract
Video super-resolution (VSR) has many applications that pose strict causal, real-time, and latency constraints, including video streaming and TV. We address the VSR problem under these settings, which poses additional important challenges since information from future frames is unavailable. Importantly, designing efficient, yet effective frame alignment and fusion modules remain central problems. In this work, we propose a recurrent VSR architecture based on a deformable attention pyramid (DAP). Our DAP aligns and integrates information from the recurrent state into the current frame prediction. To circumvent the computational cost of traditional attention-based methods, we only attend to a limited number of spatial locations, which are dynamically predicted by the DAP. Comprehensive experiments and analysis of the proposed key innovations show the effectiveness of our approach. We significantly reduce processing time and computational complexity in comparison to state-of-the-art methods, while maintaining a high performance. We surpass state-of-the-art method EDVR-M on two standard benchmarks with a speed-up of over 3×.
Dario Fuoli, Martin Danelljan, Radu Timofte, Luc Van Gool
WACV3
2023 An Efficient Recurrent Adversarial Framework for Unsupervised Real-Time Video Enhancement
abstract
Abstract Video enhancement is a challenging problem, more than that of stills, mainly due to high computational cost, larger data volumes and the difficulty of achieving consistency in the spatio-temporal domain. In practice, these challenges are often coupled with the lack of example pairs, which inhibits the application of supervised learning strategies. To address these challenges, we propose an efficient adversarial video enhancement framework that learns directly from unpaired video examples. In particular, our framework introduces new recurrent cells that consist of interleaved local and global modules for implicit integration of spatial and temporal information. The proposed design allows our recurrent cells to efficiently propagate spatio-temporal information across frames and reduces the need for high complexity networks. Our setting enables learning from unpaired videos in a cyclic adversarial manner, where the proposed recurrent units are employed in all architectures. Efficient training is accomplished by introducing one single discriminator that learns the joint distribution of source and target domain simultaneously. The enhancement results demonstrate clear superiority of the proposed video enhancer over the state-of-the-art methods, in all terms of visual quality, quantitative metrics, and inference speed. Notably, our video enhancer is capable of enhancing over 35 frames per second of FullHD video (1080x1920).
Dario Fuoli, Zhiwu Huang, Danda Pani Paudel, Luc Van Gool, Radu Timofte
Int. J. Comput. Vis.5
2023 PDC-Net+: Enhanced Probabilistic Dense Correspondence Network
abstract
Establishing robust and accurate correspondences between a pair of images is a long-standing computer vision problem with numerous applications. While classically dominated by sparse methods, emerging dense approaches offer a compelling alternative paradigm that avoids the keypoint detection step. However, dense flow estimation is often inaccurate in the case of large displacements, occlusions, or homogeneous regions. In order to apply dense methods to real-world applications, such as pose estimation, image manipulation, or 3D reconstruction, it is therefore crucial to estimate the confidence of the predicted matches. We propose the Enhanced Probabilistic Dense Correspondence Network, PDC-Net+, capable of estimating accurate dense correspondences along with a reliable confidence map. We develop a flexible probabilistic approach that jointly learns the flow prediction and its uncertainty. In particular, we parametrize the predictive distribution as a constrained mixture model, ensuring better modelling of both accurate flow predictions and outliers. Moreover, we develop an architecture and an enhanced training strategy tailored for robust and generalizable uncertainty prediction in the context of self-supervised training. Our approach obtains state-of-the-art results on multiple challenging geometric matching and optical flow datasets. We further validate the usefulness of our probabilistic confidence estimation for the tasks of pose estimation, 3D reconstruction, image-based localization, and image retrieval.
Prune Truong, Martin Danelljan, Radu Timofte, Luc Van Gool
IEEE Trans. Pattern Anal. Mach. Intell.3
2023 MASIC: Deep Mask Stereo Image Compression
abstract
Stereo image compression (SIC) aims to simultaneously compress a pair of left and right stereoscopic images, which can achieve higher compression efficiency than single image compression. In this paper, to benefit the SIC tasks, we collect a large real-world stereo image dataset, namely Palace, which is composed of hundreds of stereo image pairs at high-resolution. More importantly, we propose a novel mask stereo image compression network, namely MASIC, which can jointly compress the stereo images with high compression efficiency. Specifically, we first estimate the homography matrix between the stereo images through a regression model. Then, the left image is spatially transformed by the homography matrix, so that only the residual information needs to be encoded for the right image. To avoid the wrong guidance between stereo image pair, we propose a mask prediction module (MPM) to generate a multi-channel guided mask to navigate both the encoding and decoding processes. Based on the guided mask, we introduce a new mask conditional stereo entropy (MCSE) model, to fully explore the correlation between the stereo images in entropy coding. In the decoder, we develop a stereo decoding module to simultaneously decode the stereo images and enhance their compression quality. Experimental results show that our MASIC significantly advances the performance of SIC both quantitatively and qualitatively on a variety of datasets, and is robust to the change of parallax level between stereo images. The software codes are available athttps://github.com/eecoder-dyf/MASIC.
Xin Deng 0002, Yufan Deng, Radu Timofte, Mai Xu
IEEE Trans. Circuits Syst. Video Technol.5
2023 Advancing Learned Video Compression With In-Loop Frame Prediction
abstract
Recent years have witnessed an increasing interest in end-to-end learned video compression. Most previous works explore temporal redundancy by detecting and compressing a motion map to warp the reference frame towards the target frame. Yet, it failed to adequately take advantage of the historical priors in the sequential reference frames. In this paper, we propose an Advanced Learned Video Compression (ALVC) approach with the in-loop frame prediction module, which is able to effectively predict the target frame from the previously compressed frames, without consuming any bit-rate. The predicted frame can serve as a better reference than the previously compressed frame, and therefore it benefits the compression performance. The proposed in-loop prediction module is a part of the end-to-end video compression and is jointly optimized in the whole framework. We propose the recurrent and the bi-directional in-loop prediction modules for compressing P-frames and B-frames, respectively. The experiments show the state-of-the-art performance of our ALVC approach in learned video compression. We also outperform the default hierarchical B mode of x265 in terms of PSNR and beat the slowest mode of the SSIM-tuned x265 on MS-SSIM. The project page:https://github.com/RenYang-home/ALVC.
Radu Timofte, Luc Van Gool
IEEE Trans. Circuits Syst. Video Technol.2
2023 Learning Context-Based Nonlocal Entropy Modeling for Image Compression
abstract
The entropy of the codes usually serves as the rate loss in the recent learned lossy image compression methods. Precise estimation of the probabilistic distribution of the codes plays a vital role in reducing the entropy and boosting the joint rate-distortion performance. However, existing deep learning based entropy models generally assume the latent codes are statistically independent or depend on some side information or local context, which fails to take the global similarity within the context into account and thus hinders the accurate entropy estimation. To address this issue, we propose a special nonlocal operation for context modeling by employing the global similarity within the context. Specifically, due to the constraint of context, nonlocal operation is incalculable in context modeling. We exploit the relationship between the code maps produced by deep neural networks and introduce the proxy similarity functions as a workaround. Then, we combine the local and the global context via a nonlocal attention block and employ it in masked convolutional networks for entropy modeling. Taking the consideration that the width of the transforms is essential in training low distortion models, we finally produce a U-net block in the transforms to increase the width with manageable memory consumption and time complexity. Experiments on Kodak and Tecnick datasets demonstrate the priority of the proposed context-based nonlocal attention block in entropy modeling and the U-net block in low distortion situations. On the whole, our model performs favorably against the existing image compression standards and recent deep image compression models.
Mu Li 0005, Kai Zhang 0008, Jinxing Li 0003, Wangmeng Zuo, Radu Timofte, David Zhang 0001
IEEE Trans. Neural Networks Learn. Syst.5
2022 Enhanced Super-Resolution Training via Mimicked Alignment for Real-World Scenes
Omar Elezabi, Zongwei Wu, Radu Timofte
ACCV (4)3
2022 SiNeRF: Sinusoidal Neural Radiance Fields for Joint Pose Estimation and Scene Reconstruction
Yitong Xia, Hao Tang 0005, Radu Timofte, Luc Van Gool
BMVC3
2022 Mask-guided Spectral-wise Transformer for Efficient Hyperspectral Image Reconstruction
abstract
Hyperspectral image (HSI) reconstruction aims to recover the 3D spatial-spectral signal from a 2D measurement in the coded aperture snapshot spectral imaging (CASSI) system. The HSI representations are highly similar and correlated across the spectral dimension. Modeling the inter-spectra interactions is beneficial for HSI reconstruction. However, existing CNN-based methods show limitations in capturing spectral-wise similarity and long-range dependencies. Besides, the HSI information is modulated by a coded aperture (physical mask) in CASSI. Nonetheless, current algorithms have not fully explored the guidance effect of the mask for HSI restoration. In this paper, we propose a novel framework, Mask-guided Spectral-wise Transformer (MST), for HSI reconstruction. Specifically, we present a Spectral-wise Multi-head Self-Attention (S-MSA) that treats each spectral feature as a token and calculates self-attention along the spectral dimension. In addition, we customize a Mask-guided Mechanism (MM) that directs S- MSA to pay attention to spatial regions with high-fidelity spectral representations. Extensive experiments show that our MST significantly outperforms state-of-the-art (SOTA) methods on simulation and real HSI datasets while requiring dramatically cheaper computational and memory costs. https://github.com/caiyuanhao1998/MST/
Yuanhao Cai, Xiaowan Hu, Haoqian Wang, Xin Yuan 0002, Yulun Zhang 0001, Radu Timofte, Luc Van Gool
CVPR7
2022 HDNet: High-resolution Dual-domain Learning for Spectral Compressive Imaging
abstract
The rapid development of deep learning provides a better solution for the end-to-end reconstruction of hyperspectral image (HSI). However, existing learning-based methods have two major defects. Firstly, networks with self-attention usually sacrifice internal resolution to balance model performance against complexity, losing fine-grained high-resolution (HR) features. Secondly, even if the optimization focusing on spatial-spectral domain learning (SDL) converges to the ideal solution, there is still a significant visual difference between the reconstructed HSI and the truth. So we propose a high-resolution dual-domain learning network (HDNet) for HSI reconstruction. On the one hand, the proposed HR spatial-spectral attention module with its efficient feature fusion provides continuous and fine pixel-level features. On the other hand, frequency domain learning (FDL) is introduced for HSI reconstruction to narrow the frequency domain discrepancy. Dynamic FDL supervision forces the model to reconstruct fine-grained frequencies and compensate for excessive smoothing and distortion caused by pixel-level losses. The HR pixel-level attention and frequency-level refinement in our HDNet mutually promote HSI perceptual quality. Extensive quantitative and qualitative experiments show that our method achieves SOTA performance on simulated and real HSI datasets. https://github.com/Huxiaowan/HDNet
Xiaowan Hu, Yuanhao Cai, Haoqian Wang, Xin Yuan 0002, Yulun Zhang 0001, Radu Timofte, Luc Van Gool
CVPR7
2022 Revisiting Random Channel Pruning for Neural Network Compression
abstract
Channel (or 3D filter) pruning serves as an effective way to accelerate the inference of neural networks. There has been a flurry of algorithms that try to solve this practical problem, each being claimed effective in some ways. Yet, a benchmark to compare those algorithms directly is lacking, mainly due to the complexity of the algorithms and some custom settings such as the particular network configuration or training procedure. A fair benchmark is important for the further development of channel pruning. Meanwhile, recent investigations reveal that the channel configurations discovered by pruning algorithms are at least as important as the pre-trained weights. This gives channel pruning a new role, namely searching the optimal channel configuration. In this paper, we try to determine the channel configuration of the pruned models by random search. The proposed approach provides a new way to compare different methods, namely how well they behave compared with random pruning. We show that this simple strategy works quite well compared with other channel pruning methods. We also show that under this setting, there are surprisingly no clear winners among different channel importance evaluation methods, which then may tilt the research efforts into advanced channel configuration searching methods. Code will be released at https://github.com/ofsoundof/random_channel_pruning.
Yawei Li 0001, Kamil Adamczewski, Wen Li 0001, Shuhang Gu, Radu Timofte, Luc Van Gool
CVPR5
2022 RePaint: Inpainting using Denoising Diffusion Probabilistic Models
abstract
Free-form inpainting is the task of adding new content to an image in the regions specified by an arbitrary binary mask. Most existing approaches train for a certain distribution of masks, which limits their generalization capabilities to unseen mask types. Furthermore, training with pixel-wise and perceptual losses often leads to simple textural extensions towards the missing areas instead of semantically meaningful generation. In this work, we propose RePaint: A Denoising Diffusion Probabilistic Model (DDPM) based inpainting approach that is applicable to even extreme masks. We employ a pretrained unconditional DDPM as the generative prior. To condition the generation process, we only alter the reverse diffusion iterations by sampling the unmasked regions using the given image infor-mation. Since this technique does not modify or condition the original DDPM network itself, the model produces high-quality and diverse output images for any inpainting form. We validate our method for both faces and general-purpose image inpainting using standard and extreme masks. Re-Paint outperforms state-of-the-art Autoregressive, and GAN approaches for at least five out of six mask distributions. Github Repository: git.io/RePaint
Andreas Lugmayr, Martin Danelljan, Andrés Romero, Fisher Yu 0001, Radu Timofte, Luc Van Gool
CVPR5
2022 Arbitrary-Scale Image Synthesis
abstract
Positional encodings have enabled recent works to train a single adversarial network that can generate images of different scales. However, these approaches are either limited to a set of discrete scales or struggle to maintain good perceptual quality at the scales for which the model is not trained explicitly. We propose the design of scale-consistent positional encodings invariant to our generator's layers transformations. This enables the generation of arbitrary-scale images even at scales unseen during training. Moreover, we incorporate novel inter-scale augmentations into our pipeline and partial generation training to facilitate the synthesis of consistent images at arbitrary scales. Lastly, we show competitive results for a continuum of scales on various commonly used datasets for image synthesis.
Evangelos Ntavelis, Mohamad Shahbazi, Iason Kastanis, Radu Timofte, Martin Danelljan, Luc Van Gool
CVPR4
2022 Generative Flows with Invertible Attentions
abstract
Flow-based generative models have shown an excellent ability to explicitly learn the probability density function of data via a sequence of invertible transformations. Yet, learning attentions in generative flows remains understudied, while it has made breakthroughs in other domains. To fill the gap, this paper introduces two types of invertible attention mechanisms, i.e., map-based and transformer-based attentions, for both unconditional and conditional generative flows. The key idea is to exploit a masked scheme of these two attentions to learn long-range data dependencies in the context of generative flows. The masked scheme allows for invertible attention modules with tractable Jacobian determinants, enabling its seamless integration at any positions of the flow-based models. The proposed attention mechanisms lead to more efficient generative flows, due to their capability of modeling the long-term data dependencies. Evaluation on multiple image synthesis tasks shows that the proposed attention flows result in efficient models and compare favorably against the state-of-the-art unconditional and conditional generative flows.
Rhea Sanjay Sukthanker, Zhiwu Huang, Suryansh Kumar 0001, Radu Timofte, Luc Van Gool
CVPR4
2022 Predict, Prevent, and Evaluate: Disentangled Text-Driven Image Manipulation Empowered by Pre-Trained Vision-Language Model
abstract
To achieve disentangled image manipulation, previous works depend heavily on manual annotation. Meanwhile, the available manipulations are limited to a pre-defined set the models were trainedfor. We propose a novelframework, i.e., Predict, Prevent, and Evaluate (PPE), for disentangled text-driven image manipulation that requires little manual annotation while being applicable to a wide variety of ma-nipulations. Our method approaches the targets by deeply exploiting the power of the large-scale pre-trained vision-language model CLIP [32]. Concretely, we firstly Predict the possibly entangled attributes for a given text command. Then, based on the predicted attributes, we introduce an entanglement loss to Prevent entanglements during training. Finally, we propose a new evaluation metric to Evaluate the disentangled image manipulation. We verify the effectiveness of our method on the challenging face editing task. Extensive experiments show that the proposed PPE frame-work achieves much better quantitative and qualitative re-sults than the up-to-date StyleCLIP [31] baseline. Code is available at https://github.com/zipengxuc/PPE.
Zipeng Xu, Hao Tang 0005, Fu Li 0003, Dongliang He, Nicu Sebe, Radu Timofte, Luc Van Gool, Errui Ding
CVPR7
2022 Coarse-to-Fine Sparse Transformer for Hyperspectral Image Reconstruction
Yuanhao Cai, Xiaowan Hu, Haoqian Wang, Xin Yuan 0002, Yulun Zhang 0001, Radu Timofte, Luc Van Gool
ECCV (17)7
2022 Transform Your Smartphone into a DSLR Camera: Learning the ISP in the Wild
Ardhendu Shekhar Tripathi, Martin Danelljan, Samarth Shukla, Radu Timofte, Luc Van Gool
ECCV (6)4
2022 Efficient Video Enhancement Transformer
abstract
Video Enhancement is an important computer vision task aiming at the removal of the artifacts from a lossy compressed video and the improvement of the visual properties by a photo-realistic restoration of the video contents. Decades of research produced a multitude of efficient algorithms, enabling the reduction of the memory footprint of the transferred video contents in a contiguously increasing network of video streaming services. In this work, we propose VETRAN - a low latency real-time online Video Enhancement TRANsformer based on spatial and temporal attention mechanisms. We validate our method on recent Video Enhancement NTIRE and AIM challenge benchmarks, i.e. REDS/REDS4, LDV, and IntVID. We improve over the compared state-of-the-art methods both quantitatively and qualitatively, while maintaining a low inference time.
Florin-Alexandru Vasluianu, Radu Timofte
ICIP2
2022 Flow-Guided Sparse Transformer for Video Deblurring
abstract
Exploiting similar and sharper scene patches in spatio-temporal neighborhoods is critical for video deblurring. However, CNN-based methods show limitations in capturing long-range dependencies and modeling non-local self-similarity. In this paper, we propose a novel framework, Flow-Guided Sparse Transformer (FGST), for video deblurring. In FGST, we customize a self-attention module, Flow-Guided Sparse Window-based Multi-head Self-Attention (FGSW-MSA). For each $query$ element on the blurry reference frame, FGSW-MSA enjoys the guidance of the estimated optical flow to globally sample spatially sparse yet highly related $key$ elements corresponding to the same scene patch in neighboring frames. Besides, we present a Recurrent Embedding (RE) mechanism to transfer information from past frames and strengthen long-range temporal dependencies. Comprehensive experiments demonstrate that our proposed FGST outperforms state-of-the-art (SOTA) methods on both DVD and GOPRO datasets and yields visually pleasant results in real video deblurring. https://github.com/linjing7/VR-Baseline
Yuanhao Cai, Xiaowan Hu, Haoqian Wang, Youliang Yan, Xueyi Zou, Henghui Ding, Yulun Zhang 0001, Radu Timofte, Luc Van Gool
ICML9
2022 PyNet-V2 Mobile: Efficient On-Device Photo Processing With Neural Networks
abstract
The increased importance of mobile photography created a need for fast and performant RAW image processing pipelines capable of producing good visual results in spite of the mobile camera sensor limitations. While deep learning-based approaches can efficiently solve this problem, their computational requirements usually remain too large for high-resolution on-device image processing. To address this limitation, we propose a novel PyNET-V2 Mobile CNN architecture designed specifically for edge devices, being able to process RAW 12MP photos directly on mobile phones under 1.5 second and producing high perceptual photo quality. To train and to evaluate the performance of the proposed solution, we use the real-world Fujifilm UltraISP dataset consisting on thousands of RAW-RGB image pairs captured with a professional medium-format 102MP Fujifilm camera and a popular Sony mobile camera sensor. The results demonstrate that the PyNET-V2 Mobile model can substantially surpass the quality of tradition ISP pipelines, while outperforming the previously introduced neural network-based solutions designed for fast image processing. Furthermore, we show that the proposed architecture is also compatible with the latest mobile AI accelerators such as NPUs or APUs that can be used to further reduce the latency of the model to as little as 0.5 second. The dataset, code and pre-trained models used in this paper are available on the project website: https://github.com/gmalivenko/PyNET-v2
Andrey Ignatov, Grigory Malivenko, Radu Timofte, Yu Tseng, Yu-Syuan Xu, Po-Hsiang Yu, Cheng-Ming Chiang, Hsien-Kai Kuo, Min-Hung Chen, Chia-Ming Cheng, Luc Van Gool
ICPR3
2022 Perceptual Learned Video Compression with Recurrent Conditional GAN
abstract
This paper proposes a Perceptual Learned Video Compression (PLVC) approach with recurrent conditional GAN. We employ the recurrent auto-encoder-based compression network as the generator, and most importantly, we propose a recurrent conditional discriminator, which judges raw vs. compressed video conditioned on both spatial and temporal features, including the latent representation, temporal motion and hidden states in recurrent cells. This way, the adversarial training pushes the generated video to be not only spatially photo-realistic but also temporally consistent with the groundtruth and coherent among video frames. The experimental results show that the learned PLVC model compresses video with good perceptual quality at low bit-rate, and that it outperforms the official HEVC test model (HM 16.20) and the existing learned video compression approaches for several perceptual quality metrics and user studies. The project page is available at https://github.com/RenYang-home/PLVC.
Radu Timofte, Luc Van Gool
IJCAI2
2022 Degradation-Aware Unfolding Half-Shuffle Transformer for Spectral Compressive Imaging
abstract
In coded aperture snapshot spectral compressive imaging (CASSI) systems, hyperspectral image (HSI) reconstruction methods are employed to recover the spatial-spectral signal from a compressed measurement. Among these algorithms, deep unfolding methods demonstrate promising performance but suffer from two issues. Firstly, they do not estimate the degradation patterns and ill-posedness degree from CASSI to guide the iterative learning. Secondly, they are mainly CNN-based, showing limitations in capturing long-range dependencies. In this paper, we propose a principled Degradation-Aware Unfolding Framework (DAUF) that estimates parameters from the compressed image and physical mask, and then uses these parameters to control each iteration. Moreover, we customize a novel Half-Shuffle Transformer (HST) that simultaneously captures local contents and non-local dependencies. By plugging HST into DAUF, we establish the first Transformer-based deep unfolding method, Degradation-Aware Unfolding Half-Shuffle Transformer (DAUHST), for HSI reconstruction. Experiments show that DAUHST surpasses state-of-the-art methods while requiring cheaper computational and memory costs. Code and models are publicly available at https://github.com/caiyuanhao1998/MST
Yuanhao Cai, Haoqian Wang, Xin Yuan 0002, Henghui Ding, Yulun Zhang 0001, Radu Timofte, Luc Van Gool
NeurIPS7
2022 Recurrent Video Restoration Transformer with Guided Deformable Attention
abstract
Video restoration aims at restoring multiple high-quality frames from multiple low-quality frames. Existing video restoration methods generally fall into two extreme cases, i.e., they either restore all frames in parallel or restore the video frame by frame in a recurrent way, which would result in different merits and drawbacks. Typically, the former has the advantage of temporal information fusion. However, it suffers from large model size and intensive memory consumption; the latter has a relatively small model size as it shares parameters across frames; however, it lacks long-range dependency modeling ability and parallelizability. In this paper, we attempt to integrate the advantages of the two cases by proposing a recurrent video restoration transformer, namely RVRT. RVRT processes local neighboring frames in parallel within a globally recurrent framework which can achieve a good trade-off between model size, effectiveness, and efficiency. Specifically, RVRT divides the video into multiple clips and uses the previously inferred clip feature to estimate the subsequent clip feature. Within each clip, different frame features are jointly updated with implicit feature aggregation. Across different clips, the guided deformable attention is designed for clip-to-clip alignment, which predicts multiple relevant locations from the whole inferred clip and aggregates their features by the attention mechanism. Extensive experiments on video super-resolution, deblurring, and denoising show that the proposed RVRT achieves state-of-the-art performance on benchmark datasets with balanced model size, testing memory and runtime.
Jingyun Liang, Yuchen Fan 0001, Xiaoyu Xiang, Eddy Ilg, Simon Green, Jiezhang Cao, Kai Zhang 0008, Radu Timofte, Luc Van Gool
NeurIPS9
2022 Normalizing Flow as a Flexible Fidelity Objective for Photo-Realistic Super-resolution
abstract
Super-resolution is an ill-posed problem, where a ground-truth high-resolution image represents only one possibility in the space of plausible solutions. Yet, the dominant paradigm is to employ pixel-wise losses, such as L1, which drive the prediction towards a blurry average. This leads to fundamentally conflicting objectives when combined with adversarial losses, which degrades the final quality. We address this issue by revisiting the L1loss and show that it corresponds to a one-layer conditional flow. Inspired by this relation, we explore general flows as a fidelity-based alternative to the L1objective. We demonstrate that the flexibility of deeper flows leads to better visual quality and consistency when combined with adversarial losses. We conduct extensive user studies for three datasets and scale factors, where our approach is shown to outperform state-of-the-art methods for photo-realistic super-resolution. Code and trained models: git.io/AdFlow
Andreas Lugmayr, Martin Danelljan, Fisher Yu 0001, Luc Van Gool, Radu Timofte
WACV5
2022 Efficient and robust eye images iris segmentation using a lightweight U-net convolutional network
Casian Miron, Alexandru Pasarica, Vasile I. Manta, Radu Timofte
Multim. Tools Appl.4
2022 Plug-and-Play Image Restoration With Deep Denoiser Prior
abstract
Recent works on plug-and-play image restoration have shown that a denoiser can implicitly serve as the image prior for model-based methods to solve many inverse problems. Such a property induces considerable advantages for plug-and-play image restoration (e.g., integrating the flexibility of model-based method and effectiveness of learning-based methods) when the denoiser is discriminatively learned via deep convolutional neural network (CNN) with large modeling capacity. However, while deeper and larger CNN models are rapidly gaining popularity, existing plug-and-play image restoration hinders its performance due to the lack of suitable denoiser prior. In order to push the limits of plug-and-play image restoration, we set up a benchmark deep denoiser prior by training a highly flexible and effective CNN denoiser. We then plug the deep denoiser prior as a modular part into a half quadratic splitting based iterative algorithm to solve various image restoration problems. We, meanwhile, provide a thorough analysis of parameter setting, intermediate results and empirical convergence to better understand the working mechanism. Experimental results on three representative image restoration tasks, including deblurring, super-resolution and demosaicing, demonstrate that the proposed plug-and-play image restoration with deep denoiser prior not only significantly outperforms other state-of-the-art model-based methods but also achieves competitive or even superior performance against state-of-the-art learning-based methods. The source code is available at https://github.com/cszn/DPIR.
Kai Zhang 0008, Yawei Li 0001, Wangmeng Zuo, Lei Zhang 0006, Luc Van Gool, Radu Timofte
IEEE Trans. Pattern Anal. Mach. Intell.6
2021 Deep Line Encoding for Monocular 3D Object Detection and Depth Prediction
Ce Liu 0004, Shuhang Gu, Luc Van Gool, Radu Timofte
BMVC4
2021 Deep Homography for Efficient Stereo Image Compression
abstract
In this paper, we propose HESIC, an end-to-end trainable deep network for stereo image compression (SIC). To fully explore the mutual information across two stereo images, we use a deep regression model to estimate the homography matrix, i.e., H matrix. Then, the left image is spatially transformed by the H matrix, and only the residual information between the left and right images is encoded to save bitrates. A two-branch auto-encoder architecture is adopted in HESIC, corresponding to the left and right images, respectively. For entropy coding, we use two conditional stereo entropy models, i.e., Gaussian mixture model (GMM) based and context based entropy models, to fully explore the correlation between the two images to reduce the coding bit-rates. In decoding, a cross quality enhancement module is proposed to enhance the image quality based on inverse H matrix. Experimental results show that our HESIC outperforms state-of-the-art SIC methods on InStereo2K and KITTI datasets both quantitatively and qualitatively. Code is available at https://github.com/ywz978020607/HESIC.
Xin Deng 0002, Mai Xu, Enpeng Liu, Qianhan Feng, Radu Timofte
CVPR7
2021 Deep Burst Super-Resolution
abstract
While single-image super-resolution (SISR) has attracted substantial interest in recent years, the proposed approaches are limited to learning image priors in order to add high frequency details. In contrast, multi-frame super-resolution (MFSR) offers the possibility of reconstructing rich details by combining signal information from multiple shifted images. This key advantage, along with the increasing popularity of burst photography, have made MFSR an important problem for real-world applications.We propose a novel architecture for the burst super-resolution task. Our network takes multiple noisy RAW images as input, and generates a denoised, super-resolved RGB image as output. This is achieved by explicitly aligning deep embeddings of the input frames using pixel-wise optical flow. The information from all frames are then adaptively merged using an attention-based fusion module. In order to enable training and evaluation on real-world data, we additionally introduce the BurstSR dataset, consisting of smartphone bursts and high-resolution DSLR ground-truth. We perform comprehensive experimental analysis, demonstrating the effectiveness of the proposed architecture.
Goutam Bhat, Martin Danelljan, Luc Van Gool, Radu Timofte
CVPR4
2021 The Heterogeneity Hypothesis: Finding Layer-Wise Differentiated Network Architectures
abstract
In this paper, we tackle the problem of convolutional neural network design. Instead of focusing on the design of the overall architecture, we investigate a design space that is usually overlooked, i.e. adjusting the channel configurations of predefined networks. We find that this adjustment can be achieved by shrinking widened baseline networks and leads to superior performance. Based on that, we articulate the "heterogeneity hypothesis": with the same training protocol, there exists a layer-wise differentiated net-work architecture (LW-DNA) that can outperform the original network with regular channel configurations but with a lower level of model complexity.The LW-DNA models are identified without extra computational cost or training time compared with the original network. This constraint leads to controlled experiments which direct the focus to the importance of layer-wise specific channel configurations. LW-DNA models come with advantages related to overfitting, i.e. the relative relationship between model complexity and dataset size. Experiments are conducted on various networks and datasets for image classification, visual tracking and image restoration. The resultant LW-DNA models consistently outperform the baseline models. Code is available at https://github.com/ofsoundof/Heterogeneity_Hypothesis.git.
Yawei Li 0001, Wen Li 0001, Martin Danelljan, Kai Zhang 0008, Shuhang Gu, Luc Van Gool, Radu Timofte
CVPR7
2021 Flow-Based Kernel Prior With Application to Blind Super-Resolution
abstract
Kernel estimation is generally one of the key problems for blind image super-resolution (SR). Recently, Double-DIP proposes to model the kernel via a network architecture prior, while KernelGAN employs the deep linear network and several regularization losses to constrain the kernel space. However, they fail to fully exploit the general SR kernel assumption that anisotropic Gaussian kernels are sufficient for image SR. To address this issue, this paper proposes a normalizing flow-based kernel prior (FKP) for kernel modeling. By learning an invertible mapping between the anisotropic Gaussian kernel distribution and a tractable latent distribution, FKP can be easily used to replace the kernel modeling modules of Double-DIP and KernelGAN. Specifically, FKP optimizes the kernel in the latent space rather than the network parameter space, which allows it to generate reasonable kernel initialization, traverse the learned kernel manifold and improve the optimization stability. Extensive experiments on synthetic and real-world images demonstrate that the proposed FKP can significantly improve the kernel estimation accuracy with less parameters, runtime and memory usage, leading to state-of-the-art blind SR results.
Jingyun Liang, Kai Zhang 0008, Shuhang Gu, Luc Van Gool, Radu Timofte
CVPR5
2021 Learning Accurate Dense Correspondences and When To Trust Them
abstract
Establishing dense correspondences between a pair of images is an important and general problem. However, dense flow estimation is often inaccurate in the case of large displacements or homogeneous regions. For most applications and down-stream tasks, such as pose estimation, image manipulation, or 3D reconstruction, it is crucial to know when and where to trust the estimated matches.In this work, we aim to estimate a dense flow field relating two images, coupled with a robust pixel-wise confidence map indicating the reliability and accuracy of the prediction. We develop a flexible probabilistic approach that jointly learns the flow prediction and its uncertainty. In particular, we parametrize the predictive distribution as a constrained mixture model, ensuring better modelling of both accurate flow predictions and outliers. Moreover, we develop an architecture and training strategy tailored for robust and generalizable uncertainty prediction in the context of self-supervised training. Our approach obtains state- of-the-art results on multiple challenging geometric matching and optical flow datasets. We further validate the usefulness of our probabilistic confidence estimation for the task of pose estimation. Code and models are available at https://github.com/PruneTruong/PDCNet.
Prune Truong, Martin Danelljan, Luc Van Gool, Radu Timofte
CVPR4
2021 Unsupervised Real-World Image Super Resolution via Domain-Distance Aware Training
abstract
These days, unsupervised super-resolution (SR) is soaring due to its practical and promising potential in real scenarios. The philosophy of off-the-shelf approaches lies in the augmentation of unpaired data, i.e. first generating synthetic low-resolution (LR) images ${\mathcal{Y}^g}$ corresponding to real-world high-resolution (HR) images ${\mathcal{X}^r}$ in the real-world LR domain ${\mathcal{Y}^r}$, and then utilizing the pseudo pairs $\left\{ {{\mathcal{Y}^g},{\mathcal{X}^r}} \right\}$ for training in a supervised manner. Unfortunately, since image translation itself is an extremely challenging task, the SR performance of these approaches is severely limited by the domain gap between generated synthetic LR images and real LR images. In this paper, we propose a novel domain-distance aware super-resolution (DASR) approach for unsupervised real-world image SR. The domain gap between training data (e.g. ${\mathcal{Y}^g}$) and testing data (e.g. ${\mathcal{Y}^r}$) is addressed with our domain-gap aware training and domain-distance weighted supervision strategies. Domain-gap aware training takes additional benefit from real data in the target domain while domain-distance weighted supervision brings forward the more rational use of labeled source domain data. The proposed method is validated on synthetic and real datasets and the experimental results show that DASR consistently outperforms state-of-the-art unsupervised SR approaches in generating SR outputs with more realistic and natural textures. Codes are available at https://github.com/ShuhangGu/DASR.
Yunxuan Wei, Shuhang Gu, Yawei Li 0001, Radu Timofte, Longcun Jin, Hengjie Song
CVPR4
2021 DeFlow: Learning Complex Image Degradations From Unpaired Data With Conditional Flows
abstract
The difficulty of obtaining paired data remains a major bottleneck for learning image restoration and enhancement models for real-world applications. Current strategies aim to synthesize realistic training data by modeling noise and degradations that appear in real-world settings. We propose DeFlow, a method for learning stochastic image degradations from unpaired data. Our approach is based on a novel unpaired learning formulation for conditional normalizing flows. We model the degradation process in the latent space of a shared flow encoder-decoder network. This allows us to learn the conditional distribution of a noisy image given the clean input by solely minimizing the negative log-likelihood of the marginal distributions. We validate our DeFlow formulation on the task of joint image restoration and super-resolution. The models trained with the synthetic data generated by DeFlow outperform previous learnable approaches on three recent datasets. Code and trained models will be made available at: https://github.com/volflow/DeFlow
Valentin Wolf, Andreas Lugmayr, Martin Danelljan, Luc Van Gool, Radu Timofte
CVPR5
2021 Designing a Practical Degradation Model for Deep Blind Image Super-Resolution
abstract
It is widely acknowledged that single image super-resolution (SISR) methods would not perform well if the assumed degradation model deviates from those in real images. Although several degradation models take additional factors into consideration, such as blur, they are still not effective enough to cover the diverse degradations of real images. To address this issue, this paper proposes to design a more complex but practical degradation model that consists of randomly shuffled blur, downsampling and noise degradations. Specifically, the blur is approximated by two convolutions with isotropic and anisotropic Gaussian kernels; the downsampling is randomly chosen from nearest, bilinear and bicubic interpolations; the noise is synthesized by adding Gaussian noise with different noise levels, adopting JPEG compression with different quality factors, and generating processed camera sensor noise via reverse-forward camera image signal processing (ISP) pipeline model and RAW image noise model. To verify the effectiveness of the new degradation model, we have trained a deep blind ES-RGAN super-resolver and then applied it to super-resolve both synthetic and real images with diverse degradations. The experimental results demonstrate that the new degradation model can help to significantly improve the practicability of deep super-resolvers, thus providing a powerful alternative solution for real SISR applications.
Kai Zhang 0008, Jingyun Liang, Luc Van Gool, Radu Timofte
ICCV4
2021 Deep Reparametrization of Multi-Frame Super-Resolution and Denoising
abstract
We propose a deep reparametrization of the maximum a posteriori formulation commonly employed in multi-frame image restoration tasks. Our approach is derived by introducing a learned error metric and a latent representation of the target image, which transforms the MAP objective to a deep feature space. The deep reparametrization allows us to directly model the image formation process in the latent space, and to integrate learned image priors into the prediction. Our approach thereby leverages the advantages of deep learning, while also benefiting from the principled multi-frame fusion provided by the classical MAP formulation. We validate our approach through comprehensive experiments on burst denoising and burst super-resolution datasets. Our approach sets a new state-of-the-art for both tasks, demonstrating the generality and effectiveness of the proposed formulation.
Goutam Bhat, Martin Danelljan, Fisher Yu 0001, Luc Van Gool, Radu Timofte
ICCV5
2021 Fourier Space Losses for Efficient Perceptual Image Super-Resolution
abstract
Many super-resolution (SR) models are optimized for high performance only and therefore lack efficiency due to large model complexity. As large models are often not practical in real-world applications, we investigate and propose novel loss functions, to enable SR with high perceptual quality from much more efficient models. The representative power for a given low-complexity generator network can only be fully leveraged by strong guidance towards the optimal set of parameters. We show that it is possible to improve the performance of a recently introduced efficient generator architecture solely with the application of our proposed loss functions. In particular, we use a Fourier space supervision loss for improved restoration of missing high-frequency (HF) content from the ground truth image and design a discriminator architecture working directly in the Fourier domain to better match the target HF distribution. We show that our losses’ direct emphasis on the frequencies in Fourier-space significantly boosts the perceptual image quality, while at the same time retaining high restoration quality in comparison to previously proposed loss functions for this task. The performance is further improved by utilizing a combination of spatial and frequency domain losses, as both representations provide complementary information during training. On top of that, the trained generator achieves comparable results with and is 2.4× and 48× faster than state-of-the-art perceptual SR methods RankSRGAN and SRFlow respectively.
Dario Fuoli, Luc Van Gool, Radu Timofte
ICCV3
2021 Towards Flexible Blind JPEG Artifacts Removal
abstract
Training a single deep blind model to handle different quality factors for JPEG image artifacts removal has been attracting considerable attention due to its convenience for practical usage. However, existing deep blind methods usually directly reconstruct the image without predicting the quality factor, thus lacking the flexibility to control the output as the non-blind methods. To remedy this problem, in this paper, we propose a flexible blind convolutional neural network, namely FBCNN, that can predict the adjustable quality factor to control the trade-off between artifacts removal and details preservation. Specifically, FBCNN decouples the quality factor from the JPEG image via a decoupler module and then embeds the predicted quality factor into the subsequent reconstructor module through a quality factor attention block for flexible control. Besides, We find existing methods are prone to fail on non-aligned double JPEG images even with only one pixel shift, and we thus propose a double JPEG degradation model to augment the training data. Extensive experiments on single JPEG images, more general double JPEG images and real-world JPEG images demonstrate that our proposed FBCNN achieves favorable performance against state-of-the-art methods in terms of both quantitative metrics and visual quality.
Jiaxi Jiang, Kai Zhang 0008, Radu Timofte
ICCV3
2021 Towards Efficient Graph Convolutional Networks for Point Cloud Handling
abstract
We aim at improving the computational efficiency of graph convolutional networks (GCNs) for learning on point clouds. The basic graph convolution that is composed of a K-nearest neighbor (KNN) search and a multilayer perceptron (MLP) is examined. By mathematically analyzing the operations there, two findings to improve the efficiency of GCNs are obtained. (1) The local geometric structure information of 3D representations propagates smoothly across the GCN that relies on KNN search to gather neighborhood features. This motivates the simplification of multiple KNN searches in GCNs. (2) Shuffling the order of graph feature gathering and an MLP leads to equivalent or similar composite operations. Based on those findings, we optimize the computational procedure in GCNs. A series of experiments show that the optimized networks have reduced computational complexity, decreased memory consumption, and accelerated inference speed while maintaining comparable accuracy for learning on point clouds.
Yawei Li 0001, Zhaopeng Cui, Radu Timofte, Marc Pollefeys, Gregory S. Chirikjian, Luc Van Gool
ICCV4
2021 Hierarchical Conditional Flow: A Unified Framework for Image Super-Resolution and Image Rescaling
abstract
Normalizing flows have recently demonstrated promising results for low-level vision tasks. For image super-resolution (SR), it learns to predict diverse photo-realistic high-resolution (HR) images from the low-resolution (LR) image rather than learning a deterministic mapping. For image rescaling, it achieves high accuracy by jointly modelling the downscaling and upscaling processes. While existing approaches employ specialized techniques for these two tasks, we set out to unify them in a single formulation. In this paper, we propose the hierarchical conditional flow (HCFlow) as a unified framework for image SR and image rescaling. More specifically, HCFlow learns a bijective mapping between HR and LR image pairs by modelling the distribution of the LR image and the rest high-frequency component simultaneously. In particular, the high-frequency component is conditional on the LR image in a hierarchical manner. To further enhance the performance, other losses such as perceptual loss and GAN loss are combined with the commonly used negative log-likelihood loss in training. Extensive experiments on general image SR, face image SR and image rescaling have demonstrated that the proposed HCFlow achieves state-of-the-art performance in terms of both quantitative metrics and visual quality.
Jingyun Liang, Andreas Lugmayr, Kai Zhang 0008, Martin Danelljan, Luc Van Gool, Radu Timofte
ICCV6
2021 Mutual Affine Network for Spatially Variant Kernel Estimation in Blind Image Super-Resolution
abstract
Existing blind image super-resolution (SR) methods mostly assume blur kernels are spatially invariant across the whole image. However, such an assumption is rarely applicable for real images whose blur kernels are usually spatially variant due to factors such as object motion and out-of-focus. Hence, existing blind SR methods would inevitably give rise to poor performance in real applications. To address this issue, this paper proposes a mutual affine network (MANet) for spatially variant kernel estimation. Specifically, MANet has two distinctive features. First, it has a moderate receptive field so as to keep the locality of degradation. Second, it involves a new mutual affine convolution (MAConv) layer that enhances feature expressiveness without increasing receptive field, model size and computation burden. This is made possible through exploiting channel interdependence, which applies each channel split with an affine transformation module whose input are the rest channel splits. Extensive experiments on synthetic and real images show that the proposed MANet not only performs favorably for both spatially variant and invariant kernel estimation, but also leads to state-of-the-art blind SR performance when combined with non-blind SR methods.
Jingyun Liang, Guolei Sun, Kai Zhang 0008, Luc Van Gool, Radu Timofte
ICCV5
2021 Generating Masks from Boxes by Mining Spatio-Temporal Consistencies in Videos
abstract
Segmenting objects in videos is a fundamental computer vision task. The current deep learning based paradigm offers a powerful, but data-hungry solution. However, current datasets are limited by the cost and human effort of annotating object masks in videos. This effectively limits the performance and generalization capabilities of existing video segmentation methods. To address this issue, we explore weaker form of bounding box annotations.We introduce a method for generating segmentation masks from per-frame bounding box annotations in videos. To this end, we propose a spatio-temporal aggregation module that effectively mines consistencies in the object and background appearance across multiple frames. We use our predicted accurate masks to train video object segmentation (VOS) networks for the tracking domain, where only manual bounding box annotations are available. The additional data provides substantially better generalization performance, leading to state-of-the-art results on standard tracking benchmarks. The code and models are available at https://github.com/visionml/pytracking.
Goutam Bhat, Martin Danelljan, Luc Van Gool, Radu Timofte
ICCV5
2021 MS-RANAS: Multi-Scale Resource-Aware Neural Architecture Search
abstract
Neural Architecture Search (NAS) has proved effective in offering outperforming alternatives to handcrafted neural networks. In this paper we analyse the benefits of NAS for image classification tasks under strict computational constraints. Our aim is to automate the design of highly efficient deep neural networks, capable of offering fast and accurate predictions and that could be deployed on a low-memory, low-power system-on-chip. The task thus becomes a three-party trade-off between accuracy, computational complexity, and memory requirements. To address this concern, we propose Multi-Scale Resource-Aware Neural Architecture Search (MS-RANAS). We employ a one-shot architecture search approach in order to obtain a reduced search cost and we focus on an anytime prediction setting. Through the usage of multiple-scaled features and early classifiers, we achieved state-of-the-art results in terms of accuracy-speed trade-off.
Cristian Cioflan, Radu Timofte
ICRA2
2021 Fast Few-Shot Classification by Few-Iteration Meta-Learning
abstract
Autonomous agents interacting with the real world need to learn new concepts efficiently and reliably. This requires learning in a low-data regime, which is a highly challenging problem. We address this task by introducing a fast optimization-based meta-learning method for few-shot classification. It consists of an embedding network, providing a general representation of the image, and a base learner module. The latter learns a linear classifier during the inference through an unrolled optimization procedure. We design an inner learning objective composed of (i) a robust classification loss on the support set and (ii) an entropy loss, allowing transductive learning from unlabeled query samples. By employing an efficient initialization module and a Steepest Descent based optimization algorithm, our base learner predicts a powerful classifier within only a few iterations. Further, our strategy enables important aspects of the base learner objective to be learned during meta-training. To the best of our knowledge, this work is the first to integrate both induction and transduction into the base learner in an optimization-based meta-learning framework. We perform a comprehensive experimental analysis, demonstrating the speed and effectiveness of our approach on four few-shot classification datasets. The Code is available at https://github.com/4rdhendu/FIML.
Ardhendu Shekhar Tripathi, Martin Danelljan, Luc Van Gool, Radu Timofte
ICRA4
2021 Local Memory Attention for Fast Video Semantic Segmentation
abstract
We propose a novel neural network module that transforms an existing single-frame semantic segmentation model into a video semantic segmentation pipeline. In contrast to prior works, we strive towards a simple, fast, and general module that can be integrated into virtually any single-frame architecture. Our approach aggregates a rich representation of the semantic information in past frames into a memory module. Information stored in the memory is then accessed through an attention mechanism. In contrast to previous memory-based approaches, we propose a fast local attention layer, providing temporal appearance cues in the local region of prior frames. We further fuse these cues with an encoding of the current frame through a second attention-based module. The segmentation decoder processes the fused representation to predict the final semantic segmentation. We integrate our approach into two popular semantic segmentation networks: ERFNet and PSPNet. We observe an improvement in segmentation performance on Cityscapes by 1.7% and 2.1% in mIoU respectively, while increasing inference time of ERFNet by only 1.5ms. Source code is available at https://github.com/mattpfr/lmanet.
Matthieu Paul, Martin Danelljan, Luc Van Gool, Radu Timofte
IROS4
2021 Deep Learning for Visual Data Compression
abstract
In this paper, we will introduce the recent progress in deep learning based visual data compression, including image compression, video compression and point cloud compression. In the past few years, deep learning techniques have been successfully applied to various computer vision and image processing applications. However, for the data compression task, the traditional approaches (i.e., block based motion estimation and motion compensation, etc.) are still widely employed in the mainstream codecs. Considering the powerful representation capability of neural networks, it is feasible to improve the data compression performance by employing the advanced deep learning technologies. To this end, the deep leaning based compression approaches have recently received increasing attention from both academia and industry in the field of computer vision and signal processing.
Guo Lu, Shenlong Wang, Shan Liu 0001, Radu Timofte
ACM Multimedia5
2021 Unsupervised Multimodal Video-to-Video Translation via Self-Supervised Learning
abstract
Existing unsupervised video-to-video translation methods fail to produce translated videos which are frame-wise realistic, semantic information preserving and video-level consistent. In this work, we propose UVIT, a novel unsupervised video-to-video translation model. Our model decomposes the style and the content, uses the specialized encoder-decoder structure and propagates the inter-frame information through bidirectional recurrent neural network (RNN) units. The style-content decomposition mechanism enables us to achieve style consistent video translation results as well as provides us with a good interface for modality flexible translation. In addition, by changing the input frames and style codes incorporated in our translation, we propose a video interpolation loss, which captures temporal information within the sequence to train our building blocks in a self-supervised manner. Our model can produce photo-realistic, spatio-temporal consistent translated videos in a multimodal way. Subjective and objective experimental results validate the superiority of our model over existing methods.
Kangning Liu, Shuhang Gu, Andrés Romero, Radu Timofte
WACV4
2021 Zero-Pair Image to Image Translation using Domain Conditional Normalization
abstract
In this paper, we propose an approach based on domain conditional normalization (DCN) for zero-pair image-to-image translation, i.e., translating between two domains which have no paired training data available but each have paired training data with a third domain. We employ a single generator which has an encoder-decoder structure and analyze different implementations of domain conditional normalization to obtain the desired target domain output. The validation benchmark uses RGB-depth pairs and RGB-semantic pairs for training and compares performance for the depth-semantic translation task. The proposed approaches improve in qualitative and quantitative terms over the compared methods, while using much fewer parameters.
Samarth Shukla, Andrés Romero, Luc Van Gool, Radu Timofte
WACV4
2021 Adversarial feature distribution alignment for semi-supervised learning
Christoph Mayer 0007, Matthieu Paul, Radu Timofte
Comput. Vis. Image Underst.3
2021 Towards closing the gap in weakly supervised semantic segmentation with DCNNs: Combining local and global models
Christoph Mayer 0007, Radu Timofte, Grégory Paul
Comput. Vis. Image Underst.2
2021 Single-Image super-resolution - When model adaptation matters
Yudong Liang, Radu Timofte, Jinjun Wang, Sanping Zhou, Yihong Gong, Nanning Zheng 0001
Pattern Recognit.2
2020 Self-Supervised 2D Image to 3D Shape Translation with Disentangled Representations
abstract
We present a framework to translate between 2D image views and 3D object shapes. Recent progress in deep learning enabled us to learn structure-aware representations from a scene. However, the existing literature assumes that pairs of images and 3D shapes are available for training in full supervision. In this paper, we propose SIST, a Self-supervised Image to Shape Translation framework that fulfills three tasks: (i) reconstructing the 3D shape from a single image; (ii) learning disentangled representations for shape, appearance and viewpoint; and (iii) generating a realistic RGB image from these independent factors. In contrast to the existing approaches, our method does not require image-shape pairs for training. Instead, it uses unpaired image and shape datasets from the same object class and jointly trains image generator and shape reconstruction networks. Our translation method achieves promising results, comparable in quantitative and qualitative terms to the state-of-the-art achieved by fully-supervised methods1.
Berk Kaya, Radu Timofte
3DV2
2020 DeepSEE: Deep Disentangled Semantic Explorative Extreme Super-Resolution
Marcel C. Bühler, Andrés Romero, Radu Timofte
ACCV (4)3
2020 How to Train Your Energy-Based Model for Regression
Fredrik Gustafsson, Martin Danelljan, Radu Timofte, Thomas B. Schön
BMVC3
2020 Probabilistic Regression for Visual Tracking
abstract
Visual tracking is fundamentally the problem of regressing the state of the target in each video frame. While significant progress has been achieved, trackers are still prone to failures and inaccuracies. It is therefore crucial to represent the uncertainty in the target estimation. Although current prominent paradigms rely on estimating a state-dependent confidence score, this value lacks a clear probabilistic interpretation, complicating its use. In this work, we therefore propose a probabilistic regression formulation and apply it to tracking. Our network predicts the conditional probability density of the target state given an input image. Crucially, our formulation is capable of modeling label noise stemming from inaccurate annotations and ambiguities in the task. The regression network is trained by minimizing the Kullback-Leibler divergence. When applied for tracking, our formulation not only allows a probabilistic representation of the output, but also substantially improves the performance. Our tracker sets a new state-of-the-art on six datasets, achieving 59.8% AUC on LaSOT and 75.8% Success on TrackingNet. The code and models are available at https://github.com/visionml/pytracking.
Martin Danelljan, Luc Van Gool, Radu Timofte
CVPR3
2020 Group Sparsity: The Hinge Between Filter Pruning and Decomposition for Network Compression
abstract
In this paper, we analyze two popular network compression techniques, i.e. filter pruning and low-rank decomposition, in a unified sense. By simply changing the way the sparsity regularization is enforced, filter pruning and low-rank decomposition can be derived accordingly. This provides another flexible choice for network compression because the techniques complement each other. For example, in popular network architectures with shortcut connections (e.g. ResNet), filter pruning cannot deal with the last convolutional layer in a ResBlock while the low-rank decomposition methods can. In addition, we propose to compress the whole network jointly instead of in a layer-wise manner. Our approach proves its potential as it compares favorably to the state-of-the-art on several benchmarks. Code is available at https://github.com/ofsoundof/group_sparsity.
Yawei Li 0001, Shuhang Gu, Christoph Mayer 0007, Luc Van Gool, Radu Timofte
CVPR5
2020 GLU-Net: Global-Local Universal Network for Dense Flow and Correspondences
abstract
Establishing dense correspondences between a pair of images is an important and general problem, covering geometric matching, optical flow and semantic correspondences. While these applications share fundamental challenges, such as large displacements, pixel-accuracy, and appearance changes, they are currently addressed with specialized network architectures, designed for only one particular task. This severely limits the generalization capabilities of such networks to new scenarios, where e.g. robustness to larger displacements or higher accuracy is required. In this work, we propose a universal network architecture that is directly applicable to all the aforementioned dense correspondence problems. We achieve both high accuracy and robustness to large displacements by investigating the combined use of global and local correlation layers. We further propose an adaptive resolution strategy, allowing our network to operate on virtually any input image resolution. The proposed GLU-Net achieves state-of-the-art performance for geometric and semantic matching as well as optical flow, when using the same network and weights. Code and trained models are available at https://github.com/PruneTruong/GLU-Net.
Prune Truong, Martin Danelljan, Radu Timofte
CVPR3
2020 Learning for Video Compression With Hierarchical Quality and Recurrent Enhancement
abstract
In this paper, we propose a Hierarchical Learned Video Compression (HLVC) method with three hierarchical quality layers and a recurrent enhancement network. The frames in the first layer are compressed by an image compression method with the highest quality. Using these frames as references, we propose the Bi-Directional Deep Compression (BDDC) network to compress the second layer with relatively high quality. Then, the third layer frames are compressed with the lowest quality, by the proposed Single Motion Deep Compression (SMDC) network, which adopts a single motion map to estimate the motions of multiple frames, thus saving bits for motion information. In our deep decoder, we develop the Weighted Recurrent Quality Enhancement (WRQE) network, which takes both compressed frames and the bit stream as inputs. In the recurrent cell of WRQE, the memory and update signal are weighted by quality features to reasonably leverage multi-frame information for enhancement. In our HLVC approach, the hierarchical quality benefits the coding efficiency, since the high quality information facilitates the compression and enhancement of low quality frames at encoder and decoder sides, respectively. Finally, the experiments validate that our HLVC approach advances the state-of-the-art of deep video compression methods, and outperforms the "Low-Delay P (LDP) very fast" mode of x265 in terms of both PSNR and MS-SSIM. The project page is at https://github.com/RenYang-home/HLVC.
Fabian Mentzer, Luc Van Gool, Radu Timofte
CVPR4
2020 Deep Unfolding Network for Image Super-Resolution
abstract
Learning-based single image super-resolution (SISR) methods are continuously showing superior effectiveness and efficiency over traditional model-based methods, largely due to the end-to-end training. However, different from model-based methods that can handle the SISR problem with different scale factors, blur kernels and noise levels under a unified MAP (maximum a posteriori) framework, learning-based methods generally lack such flexibility. To address this issue, this paper proposes an end-to-end trainable unfolding network which leverages both learningbased methods and model-based methods. Specifically, by unfolding the MAP inference via a half-quadratic splitting algorithm, a fixed number of iterations consisting of alternately solving a data subproblem and a prior subproblem can be obtained. The two subproblems then can be solved with neural modules, resulting in an end-to-end trainable, iterative network. As a result, the proposed network inherits the flexibility of model-based methods to super-resolve blurry, noisy images for different scale factors via a single model, while maintaining the advantages of learning-based methods. Extensive experiments demonstrate the superiority of the proposed deep unfolding network in terms of flexibility, effectiveness and also generalizability.
Kai Zhang 0008, Luc Van Gool, Radu Timofte
CVPR3
2020 Know Your Surroundings: Exploiting Scene Information for Object Tracking
Goutam Bhat, Martin Danelljan, Luc Van Gool, Radu Timofte
ECCV (23)4
2020 Learning What to Learn for Video Object Segmentation
Goutam Bhat, Felix Järemo Lawin, Martin Danelljan, Andreas Robinson, Michael Felsberg, Luc Van Gool, Radu Timofte
ECCV (2)7
2020 DHP: Differentiable Meta Pruning via HyperNetworks
Yawei Li 0001, Shuhang Gu, Kai Zhang 0008, Luc Van Gool, Radu Timofte
ECCV (8)5
2020 SRFlow: Learning the Super-Resolution Space with Normalizing Flow
Andreas Lugmayr, Martin Danelljan, Luc Van Gool, Radu Timofte
ECCV (5)4
2020 SESAME: Semantic Editing of Scenes by Adding, Manipulating or Erasing Objects
Evangelos Ntavelis, Andrés Romero, Iason Kastanis, Luc Van Gool, Radu Timofte
ECCV (22)5
2020 DeepSVG: A Hierarchical Generative Network for Vector Graphics Animation
abstract
Scalable Vector Graphics (SVG) are ubiquitous in modern 2D interfaces due to their ability to scale to different resolutions. However, despite the success of deep learning-based models applied to rasterized images, the problem of vector graphics representation learning and generation remains largely unexplored. In this work, we propose a novel hierarchical generative network, called DeepSVG, for complex SVG icons generation and interpolation. Our architecture effectively disentangles high-level shapes from the low-level commands that encode the shape itself. The network directly predicts a set of shapes in a non-autoregressive fashion. We introduce the task of complex SVG icons generation by releasing a new large-scale dataset along with an open-source library for SVG manipulation. We demonstrate that our network learns to accurately reconstruct diverse vector graphics, and can serve as a powerful animation tool by performing interpolations and other latent space operations. Our code is available at https://github.com/alexandre01/deepsvg.
Alexandre Carlier, Martin Danelljan, Alexandre Alahi, Radu Timofte
NeurIPS4
2020 GOCor: Bringing Globally Optimized Correspondence Volumes into Your Neural Network
abstract
The feature correlation layer serves as a key neural network module in numerous computer vision problems that involve dense correspondences between image pairs. It predicts a correspondence volume by evaluating dense scalar products between feature vectors extracted from pairs of locations in two images. However, this point-to-point feature comparison is insufficient when disambiguating multiple similar regions in an image, severely affecting the performance of the end task. We propose GOCor, a fully differentiable dense matching module, acting as a direct replacement to the feature correlation layer. The correspondence volume generated by our module is the result of an internal optimization procedure that explicitly accounts for similar regions in the scene. Moreover, our approach is capable of effectively learning spatial matching priors to resolve further matching ambiguities. We analyze our GOCor module in extensive ablative experiments. When integrated into state-of-the-art networks, our approach significantly outperforms the feature correlation layer for the tasks of geometric matching, optical flow, and dense semantic matching. The code and trained models will be made available at github.com/PruneTruong/GOCor.
Prune Truong, Martin Danelljan, Luc Van Gool, Radu Timofte
NeurIPS4
2020 Learned image and video compression with deep neural networks
abstract
This tutorial aims at reviewing the recent progress in the deep learning based data compression, including image compression and video compression. In the past years, deep learning techniques have been successfully applied to a large number of computer vision and image processing tasks. However, for the data compression task, the traditional approaches (i.e., block based motion estimation and motion compensation, etc.) are still widely employed in the mainstream codecs. Considering the powerful representation capability, it is possible to improve the data compression performance by employing the advanced deep learning technologies. To this end, deep leaning based compression approaches have recently received significant attention from both academia and industry in the field of computer vision and image/video compression. In this tutorial, we will introduce the related deep learning techniques for image compression and video compression. Specifically, in this tutorial, we will first introduce the basic pipeline for the traditional codecs, such as JPEG, H.264 and HEVC. Then, we will discuss the common network architectures for visual data compression and analyse different learning based entropy models. Based on these techniques, we will describe several widely used end-to-end optimized frameworks for visual data compression. In summary, our tutorial will cover both the traditional data coding techniques and the popular learning based visual data compression algorithms, which will help the audiences with different backgrounds learn the recent progresses in this emerging research area.
Dong Xu 0001, Guo Lu, Radu Timofte
VCIP4
2020 Adversarial Sampling for Active Learning
abstract
This paper proposes ASAL, a new GAN based active learning method that generates high entropy samples. Instead of directly annotating the synthetic samples, ASAL searches similar samples from the pool and includes them for training. Hence, the quality of new samples is high and annotations are reliable. To the best of our knowledge, ASAL is the first GAN based AL method applicable to multi-class problems that outperforms random sample selection. Another benefit of ASAL is its small run-time complexity (sub-linear) compared to traditional uncertainty sampling (linear). We present a comprehensive set of experiments on multiple traditional data sets and show that ASAL outperforms similar methods and clearly exceeds the established baseline (random sampling). In the discussion section we analyze in which situations ASAL performs best and why it is sometimes hard to outperform random sample selection.
Christoph Mayer 0007, Radu Timofte
WACV2
2020 Efficient Video Semantic Segmentation with Labels Propagation and Refinement
abstract
This paper tackles the problem of real-time semantic segmentation of high definition videos using a hybrid GPU-CPU approach. We propose an Efficient Video Segmentation (EVS) pipeline that combines:(i)On the CPU, a very fast optical flow method, that is used to exploit the temporal aspect of the video and propagate semantic information from one frame to the next. It runs in parallel with the GPU.(ii)On the GPU, two Convolutional Neural Networks: A main segmentation network that is used to predict dense semantic labels from scratch, and a Refiner that is designed to improve predictions from previous frames with the help of a fast Inconsistencies Attention Module (IAM). The latter can identify regions that cannot be propagated accurately.We suggest several operating points depending on the desired frame rate and accuracy. Our pipeline achieves accuracy levels competitive to the existing real-time methods for semantic image segmentation (mIoU above 60%), while achieving much higher frame rates. On the popular Cityscapes dataset with high resolution frames (2048 × 1024), the proposed operating points range from 80 to 1000 Hz on a single GPU and CPU.
Matthieu Paul, Christoph Mayer 0007, Luc Van Gool, Radu Timofte
WACV4
2020 Learned Dynamic Guidance for Depth Image Reconstruction
abstract
The depth images acquired by consumer depth sensors (e.g., Kinect and ToF) usually are of low resolution and insufficient quality. One natural solution is to incorporate a high resolution RGB camera and exploit the statistical correlation of its data and depth. In recent years, both optimization-based and learning-based approaches have been proposed to deal with the guided depth reconstruction problems. In this paper, we introduce a weighted analysis sparse representation (WASR) model for guided depth image enhancement, which can be considered a generalized formulation of a wide range of previous optimization-based models. We unfold the optimization by the WASR model and conduct guided depth reconstruction with dynamically changed stage-wise operations. Such a guidance strategy enables us to dynamically adjust the stage-wise operations that update the depth image, thus improving the reconstruction quality and speed. To learn the stage-wise operations in a task-driven manner, we propose two parameterizations and their corresponding methods: dynamic guidance with Gaussian RBF nonlinearity parameterization (DG-RBF) and dynamic guidance with CNN nonlinearity parameterization (DG-CNN). The network structures of the proposed DG-RBF and DG-CNN methods are designed with the the objective function of our WASR model in mind and the optimal network parameters are learned from paired training data. Such optimization-inspired network architectures enable our models to leverage the previous expertise as well as take benefit from training data. The effectiveness is validated for guided depth image super-resolution and for realistic depth image reconstruction tasks using standard benchmarks. Our DG-RBF and DG-CNN methods achieve the best quantitative results (RMSE) and better visual quality than the state-of-the-art approaches at the time of writing. The code is available at https://github.com/ShuhangGu/GuidedDepthSR.
Shuhang Gu, Shi Guo, Wangmeng Zuo, Yunjin Chen, Radu Timofte, Luc Van Gool, Lei Zhang 0006
IEEE Trans. Pattern Anal. Mach. Intell.5
2019 Tracking the Known and the Unknown by Leveraging Semantic Information
Ardhendu Shekhar Tripathi, Martin Danelljan, Luc Van Gool, Radu Timofte
BMVC4
2019 3D Appearance Super-Resolution With Deep Learning
abstract
We tackle the problem of retrieving high-resolution (HR) texture maps of objects that are captured from multiple view points. In the multi-view case, model-based super-resolution (SR) methods have been recently proved to recover high quality texture maps. On the other hand, the advent of deep learning-based methods has already a significant impact on the problem of video and image SR. Yet, a deep learning-based approach to super-resolve the appearance of 3D objects is still missing. The main limitation of exploiting the power of deep learning techniques in the multi-view case is the lack of data. We introduce a 3D appearance SR (3DASR) dataset based on the existing ETH3D [42], SyB3R [31], MiddleBury, and our Collection of 3D scenes from TUM [21], Fountain [51] and Relief [53]. We provide the high- and low-resolution texture maps, the 3D geometric model, images and projection matrices. We exploit the power of 2D learning-based SR methods and design networks suitable for the 3D multi-view case. We incorporate the geometric information by introducing normal maps and further improve the learning process. Experimental results demonstrate that our proposed networks successfully incorporate the 3D geometric information and super-resolve the texture maps.
Yawei Li 0001, Vagia Tsiminaki, Radu Timofte, Marc Pollefeys, Luc Van Gool
CVPR3
2019 Practical Full Resolution Learned Lossless Image Compression
abstract
We propose the first practical learned lossless image compression system, L3C, and show that it outperforms the popular engineered codecs, PNG, WebP and JPEG 2000. At the core of our method is a fully parallelizable hierarchical probabilistic model for adaptive entropy coding which is optimized end-to-end for the compression task. In contrast to recent autoregressive discrete probabilistic models such as PixelCNN, our method i) models the image distribution jointly with learned auxiliary representations instead of exclusively modeling the image distribution in RGB space, and ii) only requires three forward-passes to predict all pixel probabilities instead of one for each pixel. As a result, L3C obtains over two orders of magnitude speedups when sampling compared to the fastest PixelCNN variant (Multiscale-PixelCNN). Furthermore, we find that learning the auxiliary representation is crucial and outperforms predefined auxiliary representations such as an RGB pyramid significantly.
Fabian Mentzer, Eirikur Agustsson, Michael Tschannen, Radu Timofte, Luc Van Gool
CVPR4
2019 Generative Adversarial Networks for Extreme Learned Image Compression
abstract
We present a learned image compression system based on GANs, operating at extremely low bitrates. Our proposed framework combines an encoder, decoder/generator and a multi-scale discriminator, which we train jointly for a generative learned compression objective. The model synthesizes details it cannot afford to store, obtaining visually pleasing results at bitrates where previous methods fail and show strong artifacts. Furthermore, if a semantic label map of the original image is available, our method can fully synthesize unimportant regions in the decoded image such as streets and trees from the label map, proportionally reducing the storage cost. A user study confirms that for low bitrates, our approach is preferred to state-of-the-art methods, even when they use more than double the bits.
Eirikur Agustsson, Michael Tschannen, Fabian Mentzer, Radu Timofte, Luc Van Gool
ICCV4
2019 Learning Discriminative Model Prediction for Tracking
abstract
The current strive towards end-to-end trainable computer vision systems imposes major challenges for the task of visual tracking. In contrast to most other vision problems, tracking requires the learning of a robust target-specific appearance model online, during the inference stage. To be end-to-end trainable, the online learning of the target model thus needs to be embedded in the tracking architecture itself. Due to the imposed challenges, the popular Siamese paradigm simply predicts a target feature template, while ignoring the background appearance information during inference. Consequently, the predicted model possesses limited target-background discriminability. We develop an end-to-end tracking architecture, capable of fully exploiting both target and background appearance information for target model prediction. Our architecture is derived from a discriminative learning loss by designing a dedicated optimization process that is capable of predicting a powerful model in only a few iterations. Furthermore, our approach is able to learn key aspects of the discriminative loss itself. The proposed tracker sets a new state-of-the-art on 6 tracking benchmarks, achieving an EAO score of 0.440 on VOT2018, while running at over 40 FPS. The code and models are available at https://github.com/visionml/pytracking.
Goutam Bhat, Martin Danelljan, Luc Van Gool, Radu Timofte
ICCV4
2019 Self-Guided Network for Fast Image Denoising
abstract
During the past years, tremendous advances in image restoration tasks have been achieved using highly complex neural networks. Despite their good restoration performance, the heavy computational burden hinders the deployment of these networks on constrained devices, \eg smart phones and consumer electronic products. To tackle this problem, we propose a self-guided network (SGN), which adopts a top-down self-guidance architecture to better exploit image multi-scale information. SGN directly generates multi-resolution inputs with the shuffling operation. Large-scale contextual information extracted at low resolution is gradually propagated into the higher resolution sub-networks to guide the feature extraction processes at these scales. Such a self-guidance strategy enables SGN to efficiently incorporate multi-scale information and extract good local features to recover noisy images. We validate the effectiveness of SGN through extensive experiments. The experimental results demonstrate that SGN greatly improves the memory and runtime efficiency over state-of-the-art efficient methods, without trading off PSNR accuracy.
Shuhang Gu, Yawei Li 0001, Luc Van Gool, Radu Timofte
ICCV4
2019 Fast Image Restoration With Multi-Bin Trainable Linear Units
abstract
Tremendous advances in image restoration tasks such as denoising and super-resolution have been achieved using neural networks. Such approaches generally employ very deep architectures, large number of parameters, large receptive fields and high nonlinear modeling capacity. In order to obtain efficient and fast image restoration networks one should improve upon the above mentioned requirements. In this paper we propose a novel activation function, the multi-bin trainable linear unit (MTLU), for increasing the nonlinear modeling capacity together with lighter and shallower networks. We validate the proposed fast image restoration networks for image denoising (FDnet) and super-resolution (FSRnet) on standard benchmarks. We achieve large improvements in both memory and runtime over current state-of-the-art for comparable or better PSNR accuracies.
Shuhang Gu, Wen Li 0001, Luc Van Gool, Radu Timofte
ICCV4
2019 Learning Filter Basis for Convolutional Neural Network Compression
abstract
Convolutional neural networks (CNNs) based solutions have achieved state-of-the-art performances for many computer vision tasks, including classification and super-resolution of images. Usually the success of these methods comes with a cost of millions of parameters due to stacking deep convolutional layers. Moreover, quite a large number of filters are also used for a single convolutional layer, which exaggerates the parameter burden of current methods. Thus, in this paper, we try to reduce the number of parameters of CNNs by learning a basis of the filters in convolutional layers. For the forward pass, the learned basis is used to approximate the original filters and then used as parameters for the convolutional layers. We validate our proposed solution for multiple CNN architectures on image classification and image super-resolution benchmarks and compare favorably to the existing state-of-the-art in terms of reduction of parameters and preservation of accuracy. Code is available at https://github.com/ofsoundof/learning_filter_basis.
Yawei Li 0001, Shuhang Gu, Luc Van Gool, Radu Timofte
ICCV4
2019 Dense-Haze: A Benchmark for Image Dehazing with Dense-Haze and Haze-Free Images
abstract
Single image dehazing is an ill-posed problem that has recently drawn important attention. Despite the significant increase in interest shown for dehazing over the past few years, the validation of the dehazing methods remains largely unsatisfactory, due to the lack of pairs of real hazy and corresponding haze-free reference images. To address this limitation, we introduce Dense-Haze - a novel dehazing dataset. Characterized by dense and homogeneous hazy scenes, Dense-Haze contains 33 pairs of real hazy and corresponding haze-free images of various outdoor scenes. The hazy scenes have been recorded by introducing real haze, generated by professional haze machines. The hazy and haze-free corresponding scenes contain the same visual content captured under the same illumination parameters. Dense-Haze dataset aims to push significantly the state-of-the-art in single-image dehazing by promoting robust methods for real and various hazy scenes. We also provide a comprehensive qualitative and quantitative evaluation of state-of-the-art single image dehazing techniques based on the Dense-Haze dataset. Not surprisingly, our study reveals that the existing dehazing techniques perform poorly for dense homogeneous hazy scenes and that there is still much room for improvement.
Codruta O. Ancuti, Cosmin Ancuti, Mateu Sbert, Radu Timofte
ICIP4
2019 Optimal Transport Maps For Distribution Preserving Operations on Latent Spaces of Generative Models
Eirikur Agustsson, Alexander Sage, Radu Timofte, Luc Van Gool
ICLR (Poster)3
2019 Night-to-Day Image Translation for Retrieval-based Localization
abstract
Visual localization is a key step in many robotics pipelines, allowing the robot to (approximately) determine its position and orientation in the world. An efficient and scalable approach to visual localization is to use image retrieval techniques. These approaches identify the image most similar to a query photo in a database of geo-tagged images and approximate the query's pose via the pose of the retrieved database image. However, image retrieval across drastically different illumination conditions, e.g. day and night, is still a problem with unsatisfactory results, even in this age of powerful neural models. This is due to a lack of a suitably diverse dataset with true correspondences to perform end-to-end learning. A recent class of neural models allows for realistic translation of images among visual domains with relatively little training data and, most importantly, without ground-truth pairings.In this paper, we explore the task of accurately localizing images captured from two traversals of the same area in both day and night. We propose ToDayGAN - a modified image-translation model to alter nighttime driving images to a more useful daytime representation. We then compare the daytime and translated night images to obtain a pose estimate for the night image using the known 6-DOF position of the closest day image. Our approach improves localization performance by over 250% compared the current state-of-the-art, in the context of standard metrics in multiple categories.
Asha Anoosheh, Torsten Sattler, Radu Timofte, Marc Pollefeys, Luc Van Gool
ICRA3
2018 I-HAZE: A Dehazing Benchmark with Real Hazy and Haze-Free Indoor Images
Cosmin Ancuti, Codruta O. Ancuti, Radu Timofte, Christophe De Vleeschouwer
ACIVS3
2018 Conditional Probability Models for Deep Image Compression
abstract
Deep Neural Networks trained as image auto-encoders have recently emerged as a promising direction for advancing the state-of-the-art in image compression. The key challenge in learning such networks is twofold: To deal with quantization, and to control the trade-off between reconstruction error (distortion) and entropy (rate) of the latent image representation. In this paper, we focus on the latter challenge and propose a new technique to navigate the rate-distortion trade-off for an image compression auto-encoder. The main idea is to directly model the entropy of the latent representation by using a context model: A 3D-CNN which learns a conditional probability model of the latent distribution of the auto-encoder. During training, the auto-encoder makes use of the context model to estimate the entropy of its representation, and the context model is concurrently updated to learn the dependencies between the symbols in the latent representation. Our experiments show that this approach, when measured in MS-SSIM, yields a state-of-the-art image compression system based on a simple convolutional auto-encoder.
Fabian Mentzer, Eirikur Agustsson, Michael Tschannen, Radu Timofte, Luc Van Gool
CVPR4
2018 Logo Synthesis and Manipulation With Clustered Generative Adversarial Networks
abstract
Designing a logo for a new brand is a lengthy and tedious back-and-forth process between a designer and a client. In this paper we explore to what extent machine learning can solve the creative task of the designer. For this, we build a dataset - LLD - of 600k+ logos crawled from the world wide web. Training Generative Adversarial Networks (GANs) for logo synthesis on such multi-modal data is not straightforward and results in mode collapse for some state-of-the-art methods. We propose the use of synthetic labels obtained through clustering to disentangle and stabilize GAN training, and validate this approach on CIFAR-10 and ImageNet-small to demonstrate its generality. We are able to generate a high diversity of plausible logos and demonstrate latent space exploration techniques to ease the logo design task in an interactive manner. GANs can cope with multi-modal data by means of synthetic labels achieved through clustering, and our results show the creative potential of such techniques for logo synthesis and manipulation. Our dataset and models are publicly available at https://data.vision.ee.ethz.ch/sagea/lld/.
Alexander Sage, Eirikur Agustsson, Radu Timofte, Luc Van Gool
CVPR3
2018 Towards Image Understanding from Deep Compression Without Decoding
Robert Torfason, Fabian Mentzer, Eirikur Agustsson, Michael Tschannen, Radu Timofte, Luc Van Gool
ICLR (Poster)5
2018 Integrating Local and Non-local Denoiser Priors for Image Restoration
abstract
Image local structural prior and non-local self-similarity (NSS) prior are two categories of priors which have been commonly used for solving the ill-posed image restoration problem. As they exploit different properties of the natural images, it is interesting to investigate whether the two categories of priors can be combined to achieve better restoration performance. Inspired by recently proposed Regularization by denoising [1] idea, we propose LNIR which incorporates a Local CNN denoiser prior and a NSS-based denoiser prior implicitly for Image Restoration. Our experimental results on the image deblurring and super-resolution tasks demonstrate the effectiveness of the proposed method. The proposed LNIR algorithm can not only flexibly adapt to different restoration tasks, but also delivers state-of-the-art restoration results.
Shuhang Gu, Radu Timofte, Luc Van Gool
ICPR2
2018 Guest Editorial: Vision and Computational Photography and Graphics
Radu Timofte, Luc Van Gool, Ming-Hsuan Yang 0001, Shai Avidan, Yasuyuki Matsushita, Qingxiong Yang
Comput. Vis. Image Underst.1
2018 Deep Expectation of Real and Apparent Age from a Single Image Without Facial Landmarks
Rasmus Rothe, Radu Timofte, Luc Van Gool
Int. J. Comput. Vis.2
2017 Apparent and Real Age Estimation in Still Images with Deep Residual Regressors on Appa-Real Database
abstract
After decades of research, the real (biological) age estimation from a single face image reached maturity thanks to the availability of large public face databases and impressive accuracies achieved by recently proposed methods. The estimation of “apparent age” is a related task concerning the age perceived by human observers. Significant advances have been also made in this new research direction with the recent Looking At People challenges. In this paper we make several contributions to age estimation research. (i) We introduce APPA-REAL, a large face image database with both real and apparent age annotations. (ii)We study the relationship between real and apparent age. (iii) We develop a residual age regression method to further improve the performance. (iv) We show that real age estimation can be successfully tackled as an apparent age estimation followed by an apparent to real age residual regression. (v) We graphically reveal the facial regions on which the CNN focuses in order to perform apparent and real age estimation tasks.
Eirikur Agustsson, Radu Timofte, Sergio Escalera, Xavier Baró, Isabelle Guyon, Rasmus Rothe
FG2
2017 Anchored Regression Networks Applied to Age Estimation and Super Resolution
abstract
We propose the Anchored Regression Network (ARN), a nonlinear regression network which can be seamlessly integrated into various networks or can be used stand-alone when the features have already been fixed. Our ARN is a smoothed relaxation of a piecewise linear regressor through the combination of multiple linear regressors over soft assignments to anchor points. When the anchor points are fixed the optimal ARN regressors can be obtained with a closed form global solution, otherwise ARN admits end-to-end learning with standard gradient based methods. We demonstrate the power of the ARN by applying it to two very diverse and challenging tasks: age prediction from face images and image super-resolution. In both cases, ARNs yield strong results.
Eirikur Agustsson, Radu Timofte, Luc Van Gool
ICCV2
2017 DSLR-Quality Photos on Mobile Devices with Deep Convolutional Networks
abstract
Despite a rapid rise in the quality of built-in smartphone cameras, their physical limitations - small sensor size, compact lenses and the lack of specific hardware, - impede them to achieve the quality results of DSLR cameras. In this work we present an end-to-end deep learning approach that bridges this gap by translating ordinary photos into DSLR-quality images. We propose learning the translation function using a residual convolutional neural network that improves both color rendition and image sharpness. Since the standard mean squared loss is not well suited for measuring perceptual image quality, we introduce a composite perceptual error function that combines content, color and texture losses. The first two losses are defined analytically, while the texture loss is learned in an adversarial fashion. We also present DPED, a large-scale dataset that consists of real photos captured from three different phones and one high-end reflex camera. Our quantitative and qualitative assessments reveal that the enhanced image quality is comparable to that of DSLR-taken photos, while the methodology is generalized to any type of digital camera.
Andrey Ignatov, Nikolay Kobyshev, Radu Timofte, Kenneth Vanhoey, Luc Van Gool
ICCV3
2017 Soft-to-Hard Vector Quantization for End-to-End Learning Compressible Representations
abstract
We present a new approach to learn compressible representations in deep architectures with an end-to-end training strategy. Our method is based on a soft (continuous) relaxation of quantization and entropy, which we anneal to their discrete counterparts throughout training. We showcase this method for two challenging applications: Image compression and neural network compression. While these tasks have typically been approached with different methods, our soft-to-hard quantization approach gives results competitive with the state-of-the-art for both.
Eirikur Agustsson, Fabian Mentzer, Michael Tschannen, Lukas Cavigelli, Radu Timofte, Luca Benini, Luc Van Gool
NIPS5
2017 k^2 k 2 -means for Fast and Accurate Large Scale Clustering
Eirikur Agustsson, Radu Timofte, Luc Van Gool
ECML/PKDD (2)2
2017 Leveraging observation uncertainty for robust visual tracking
Junseok Kwon, Radu Timofte, Luc Van Gool
Comput. Vis. Image Underst.2
2016 From Face Images and Attributes to Attributes
Robert Torfason, Eirikur Agustsson, Rasmus Rothe, Radu Timofte
ACCV (3)4
2016 Some Like It Hot - Visual Guidance for Preference Prediction
abstract
For people first impressions of someone are of determining importance. They are hard to alter through further information. This begs the question if a computer can reach the same judgement. Earlier research has already pointed out that age, gender, and average attractiveness can be estimated with reasonable precision. We improve the state of-the-art, but also predict - based on someone's known preferences - how much that particular person is attracted to a novel face. Our computational pipeline comprises a face detector, convolutional neural networks for the extraction of deep features, standard support vector regression for gender, age and facial beauty, and - as the main novelties - visual regularized collaborative filtering to infer interperson preferences as well as a novel regression technique for handling visual queries without rating history. We validate the method using a very large dataset from a dating site as well as images from celebrities. Our experiments yield convincing results, i.e. we predict 76% of the ratings correctly solely based on an image, and reveal some sociologically relevant conclusions. We also validate our collaborative filtering solution on the standard MovieLens rating dataset, augmented with movie posters, to predict an individuals movie rating. We demonstrate our algorithms on howhot.io which went viral around the Internet with more than 50 million pictures evaluated in the first month.
Rasmus Rothe, Radu Timofte, Luc Van Gool
CVPR2
2016 Seven Ways to Improve Example-Based Single Image Super Resolution
abstract
In this paper we present seven techniques that everybody should know to improve example-based single image super resolution (SR): 1) augmentation of data, 2) use of large dictionaries with efficient search structures, 3) cascading, 4) image self-similarities, 5) back projection refinement, 6) enhanced prediction by consistency check, and 7) context reasoning. We validate our seven techniques on standard SR benchmarks (i.e. Set5, Set14, B100) and methods (i.e. A+, SRCNN, ANR, Zeyde, Yang) and achieve substantial improvements. The techniques are widely applicable and require no changes or only minor adjustments of the SR methods. Moreover, our Improved A+ (IA) method sets new stateof-the-art results outperforming A+ by up to 0.9dB on average PSNR whilst maintaining a low time complexity.
Radu Timofte, Rasmus Rothe, Luc Van Gool
CVPR1
2016 Fast Optical Flow Using Dense Inverse Search
Till Kroeger, Radu Timofte, Dengxin Dai, Luc Van Gool
ECCV (4)2
2016 Regressor Basis Learning for anchored super-resolution
abstract
A+ aka Adjusted Anchored Neighborhood Regression - is a state-of-the-art method for exemplar-based single image super-resolution with low time complexity at both train and test time. By robustly training a clustered regression model over a low-resolution dictionary, its performance keeps improving with the dictionary size - even when using tens of thousands of regressors. However, this can pose a memory issue where the model size can grow to more than a gigabyte, limiting applicability in memory constrained scenarios. To address this, we propose Regressor Basis Learning (RB), a novel variant of A+ where we restrict the regressor set to a learned low-dimensional subspace, such that each regressor is coded as a linear combination of few basis regressors. We learn the regressor basis by alternating between closed form solutions of the optimal coding of the regressor set (given the basis) and the optimal regressor basis (given the coding). We validate RB on several standard benchmarks and achieve comparable performance to A+ but by using orders of magnitude fewer basis regressors, ie. 32 basis regressors instead of 1024 regressors. This makes our RB method ideal for memory constrained applications.
Eirikur Agustsson, Radu Timofte, Luc Van Gool
ICPR2
2016 Anchored fusion for image restoration
abstract
Recently, a series of advances were made for image restoration tasks such as image denoising and single image super-resolution. It is particularly remarkable that methods employing different formulations and assumptions achieve comparable top performances. Moreover, the top methods operate at their best on some particular image contents and poorer on other. No method is the best on all the image contents. The methods are complementary in both formulation and performance. We propose a locally adaptive fusion of results of such methods towards an improved restoration result. We work patch-wise and partition the patch space such that per each partition to train anchored regressors from the fused methods' output patches to the fusion target result. At test our anchored fusion method applies efficiently the anchored regressors corresponding to the input patches to be fused. Whilst having a low time complexity, we achieve significant improvements over the fused state-of-the-art methods on standard test images for both image denoising and super-resolution tasks (e.g. 0.1 - 0.5dB PSNR).
Radu Timofte
ICPR1
2016 Leveraging single for multi-target tracking using a novel trajectory overlap affinity measure
abstract
Multi-target tracking (MTT) is the task of localizing objects of interest in a video and associating them through time. Accurate affinity measures between object detections is crucial for MTT. Previous methods use simple affinity measures, based on heuristics, that are unable to track through occlusions and missing detections. To address this problem, this paper proposes a novel affinity measure by leveraging the power of single-target visual tracking (VT), which has proven reliable to locally track objects of interest given a bounding-box initialization. In particular, given two detections at different frames, we perform VT starting from each of them and towards the frame of the other. We then learn a metric with features extracted from the behaviours (e.g. overlaps and distances) of the two tracking trajectories. By plugging our learned affinity into the standard MTT framework, we are able to cope with occlusions and large amounts of missing or inaccurate detections. We evaluate our method on public datasets, including the popular MOT benchmark, and show improvements over previously published methods.
Santiago Manen, Radu Timofte, Dengxin Dai, Luc Van Gool
WACV2
2016 PICASO: PIxel correspondences and SOft match selection for real-time tracking
Radu Timofte, Junseok Kwon, Luc Van Gool
Comput. Vis. Image Underst.1
2016 Semantic super-resolution: When and where is it useful?
Radu Timofte, Vincent De Smet, Luc Van Gool
Comput. Vis. Image Underst.1
2016 Demosaicing Based on Directional Difference Regression and Efficient Regression Priors
abstract
Color demosaicing is a key image processing step aiming to reconstruct the missing pixels from a recorded raw image. On the one hand, numerous interpolation methods focusing on spatial-spectral correlations have been proved very efficient, whereas they yield a poor image quality and strong visible artifacts. On the other hand, optimization strategies, such as learned simultaneous sparse coding and sparsity and adaptive principal component analysis-based algorithms, were shown to greatly improve image quality compared with that delivered by interpolation methods, but unfortunately are computationally heavy. In this paper, we propose efficient regression priors as a novel, fast post-processing algorithm that learns the regression priors offline from training data. We also propose an independent efficient demosaicing algorithm based on directional difference regression, and introduce its enhanced version based on fused regression. We achieve an image quality comparable to that of the state-of-the-art methods for three benchmarks, while being order(s) of magnitude faster.
Jiqing Wu, Radu Timofte, Luc Van Gool
IEEE Trans. Image Process.2
2015 Metric imitation by manifold transfer for efficient vision applications
abstract
Metric learning has proved very successful. However, human annotations are necessary. In this paper, we propose an unsupervised method, dubbed Metric Imitation (MI), where metrics over cheap features (target features, TFs) are learned by imitating the standard metrics over more sophisticated, off-the-shelf features (source features, SFs) by transferring view-independent property manifold structures. In particular, MI consists of: 1) quantifying the properties of source metrics as manifold geometry, 2) transferring the manifold from source domain to target domain, and 3) learning a mapping of TFs so that the manifold is approximated as well as possible in the mapped feature domain. MI is useful in at least two scenarios where: 1) TFs are more efficient computationally and in terms of memory than SFs; and 2) SFs contain privileged information, but are not available during testing. For the former, MI is evaluated on image clustering, category-based image retrieval, and instance-based object retrieval, with three SFs and three TFs. For the latter, MI is tested on the task of example-based image super-resolution, where high-resolution patches are taken as SFs and low-resolution patches as TFs. Experiments show that MI is able to provide good metrics while avoiding expensive data labeling efforts and that it achieves state-of-the-art performance for image super-resolution. In addition, manifold transfer is an interesting direction of transfer learning.
Dengxin Dai, Till Kroeger, Radu Timofte, Luc Van Gool
CVPR3
2015 Efficient regression priors for reducing image compression artifacts
abstract
Lossy image compression allows for large storage savings but at the cost of reduced fidelity of the compressed images. There is a fair amount of literature aiming at restoration by suppressing the compression artifacts. Very recently a learned semi-local Gaussian Processes-based solution (SLGP) has been proposed with impressive results. However, when applied to top compression schemes such as JPEG 2000, the improvement is less significant. In our paper we propose an efficient novel artifact reduction algorithm based on the adjusted anchored neighborhood regression (A+), a method from image super-resolution literature. We double the relative gains in PSNR when compared with the state-of-the-art methods such as SLGP, while being order(s) of magnitude faster.
Rasmus Rothe, Radu Timofte, Luc Van Gool
ICIP2
2015 Efficient regression priors for post-processing demosaiced images
abstract
Color demosaicing is a process of reconstructing lost pixels in an incomplete color image. By extracting spatial-spectral correlations of RGB channels various interpolation methods have been proposed with low computational complexity. Meanwhile, optimization strategies such as sparsity and adaptive PCA based algorithm (SAPCA) were developed. SAPCA outperforms many interpolation techniques by impressive margins at the cost of dramatically increasing the computational time. In this paper we propose an efficient novel post-processing algorithm based on the adjusted anchored neighborhood regression (A+) method from image super-resolution literature. We greatly improve the results of the demosaicing methods, and achieve image quality as competitive as SAPCA but orders of magnitude faster.
Jiqing Wu, Radu Timofte, Luc Van Gool
ICIP2
2015 Sparse Flow: Sparse Matching for Small to Large Displacement Optical Flow
abstract
Despite recent advances, the extraction of optical flow with large displacements is still challenging for state-of the-art methods. The approaches that are the most successful at handling large displacements blend sparse correspondences from a matching algorithm with an optimization that refines the optical flow. We follow the scheme of Deep-Flow [33]. We first extract sparse pixel correspondences by means of a matching procedure and then apply a variational approach to obtain a refined optical flow. In our approach, coined 'Sparse Flow', the novelty lies in the matching. This uses an efficient sparse decomposition of a pixel's surrounding patch as a linear sum of those found around candidate corresponding pixels. As matching pixel the one dominating the decomposition is chosen. The pixel pairs matching in both directions, i.e. in a forward-backward fashion, are used as guiding points in the variational approach. Sparse-Flow is competitive on standard optical flow benchmarks with large displacements, while showing excellent performance for small and medium displacements. Moreover, it is fast in comparison to methods with a similar performance.
Radu Timofte, Luc Van Gool
WACV1
2015 Learned Collaborative Representations for Image Classification
abstract
The collaborative representation-based classifier (CRC) is proposed as an alternative to the sparse representation based classifier (SRC) for image face recognition. CRC solves an l2-regularized least squares formulation, with algebraic solution, while SRC optimizes over an I1-regularized least squares problem. As an extension of CRC, the weighted collaborative representation-based classifier (WCRC) is further proposed. The weights in WCRC are picked intuitively, it remains unclear why such choice of weights works and how we optimize those weights. In this paper, we propose a learned collaborative representation based classifier (LCRC) and attempt to answer the above questions. Our learning technique is based on the fixed point theorem and we use a weights formulation similar to WCRC as the starting point. Through extensive experiments on face datasets we show that the learning procedure is stable and convergent, and that LCRC is able to improve in performance over CRC and WCRC, while keeping the same computational efficiency at test.
Jiqing Wu, Radu Timofte, Luc Van Gool
WACV2
2015 Jointly Optimized Regressors for Image Super-resolution
abstract
Abstract Learning regressors from low‐resolution patches to high‐resolution patches has shown promising results for image super‐resolution. We observe that some regressors are better at dealing with certain cases, and others with different cases. In this paper, we jointly learn a collection of regressors, which collectively yield the smallest super‐resolving error for all training data. After training, each training sample is associated with a label to indicate its ‘best’ regressor, the one yielding the smallest error. During testing, our method bases on the concept of ‘adaptive selection’ to select the most appropriate regressor for each input patch. We assume that similar patches can be super‐resolved by the same regressor and use a fast, approximate kNN approach to transfer the labels of training patches to test patches. The method is conceptually simple and computationally efficient, yet very effective. Experiments on four datasets show that our method outperforms competing methods.
Dengxin Dai, Radu Timofte, Luc Van Gool
Comput. Graph. Forum2
2015 An Elastic Deformation Field Model for Object Detection and Tracking
Marco Pedersoli, Radu Timofte, Tinne Tuytelaars, Luc Van Gool
Int. J. Comput. Vis.2
2015 Iterative Nearest Neighbors
Radu Timofte, Luc Van Gool
Pattern Recognit.1
2014 A+: Adjusted Anchored Neighborhood Regression for Fast Super-Resolution
Radu Timofte, Vincent De Smet, Luc Van Gool
ACCV (4)1
2014 Using a Deformation Field Model for Localizing Faces and Facial Points under Weak Supervision
abstract
Face detection and facial points localization are interconnected tasks. Recently it has been shown that solving these two tasks jointly with a mixture of trees of parts (MTP) leads to state-of-the-art results. However, MTP, as most other methods for facial point localization proposed so far, requires a complete annotation of the training data at facial point level. This is used to predefine the structure of the trees and to place the parts correctly. In this work we extend the mixtures from trees to more general loopy graphs. In this way we can learn in a weakly supervised manner (using only the face location and orientation) a powerful deformable detector that implicitly aligns its parts to the detected face in the image. By attaching some reference points to the correct parts of our detector we can then localize the facial points. In terms of detection our method clearly outperforms the state-of-the-art, even if competing with methods that use facial point annotations during training. Additionally, without any facial point annotation at the level of individual training images, our method can localize facial points with an accuracy similar to fully supervised approaches.
Marco Pedersoli, Radu Timofte, Tinne Tuytelaars, Luc Van Gool
CVPR2
2014 Scale-invariant line descriptors for wide baseline matching
abstract
In this paper we propose a method to add scale-invariance to line descriptors for wide baseline matching purposes. While finding point correspondences among different views is a well-studied problem, there still remain difficult cases where it performs poorly, such as textureless scenes, ambiguities and extreme transformations. For these cases using line segment correspondences is a valuable addition for finding sufficient matches. Our general method for adding scale-invariance to line segment descriptors consist of 5 basic rules. We apply these rules to enhance both the line descriptor described by Bay et al. [1] and the mean-standard deviation line descriptor (MSLD) proposed by Wang et al. [14]. Moreover, we examine the effect of the line descriptors when combined with the topological filtering method proposed by Bay et al. and the recent proposed graph matching strategy from K-VLD [6]. We validate the method using standard point correspondence benchmarks and more challenging new ones. Adding scale-invariance increases the accuracy when confronted with big scale changes and increases the number of inliers in the general case, both resulting in smaller calibration errors by means of RANSAC-like techniques and epipolar estimations.
Bart Verhagen, Radu Timofte, Luc Van Gool
WACV2
2014 Multi-view traffic sign detection, recognition, and 3D localisation
Radu Timofte, Karel Zimmermann, Luc Van Gool
Mach. Vis. Appl.1
2014 Adaptive and Weighted Collaborative Representations for image classification
Radu Timofte, Luc Van Gool
Pattern Recognit. Lett.1
2013 Robust Scene Stitching in Large Scale Mobile Mapping
abstract
We provide a solution for the loop closure problem in an image-based mobile mapping context.A van drives through a city while taking images in multiple directions.Local feature matching in two stages detects when a particular site is revisited, in order to enforce correspondences between such images, that may have been taken with large time lapses in between.Our system relies on GPS but does not use odometric information.We extend the original image-to-image matching approach to a pose-to-pose matching approach, combining several images and achieving robust scene matching results.Parameter optimization is followed by extensive experiments.Our pipeline, which facilitates parallel execution, reaches matching rates higher than those reported for typical state-of-the-art algorithms.We also demonstrate robustness to odometric inconsistencies resulting from poor prior model build-up.
Filip Schouwenaars, Radu Timofte, Luc Van Gool
BMVC2
2013 Handling Occlusions with Franken-Classifiers
abstract
Detecting partially occluded pedestrians is challenging. A common practice to maximize detection quality is to train a set of occlusion-specific classifiers, each for a certain amount and type of occlusion. Since training classifiers is expensive, only a handful are typically trained. We show that by using many occlusion-specific classifiers, we outperform previous approaches on three pedestrian datasets, INRIA, ETH, and Caltech USA. We present a new approach to train such classifiers. By reusing computations among different training stages, 16 occlusion-specific classifiers can be trained at only one tenth the cost of one full training. We show that also test time cost grows sub-linearly.
Markus Mathias, Rodrigo Benenson, Radu Timofte, Luc Van Gool
ICCV3
2013 Anchored Neighborhood Regression for Fast Example-Based Super-Resolution
abstract
Recently there have been significant advances in image up scaling or image super-resolution based on a dictionary of low and high resolution exemplars. The running time of the methods is often ignored despite the fact that it is a critical factor for real applications. This paper proposes fast super-resolution methods while making no compromise on quality. First, we support the use of sparse learned dictionaries in combination with neighbor embedding methods. In this case, the nearest neighbors are computed using the correlation with the dictionary atoms rather than the Euclidean distance. Moreover, we show that most of the current approaches reach top performance for the right parameters. Second, we show that using global collaborative coding has considerable speed advantages, reducing the super-resolution mapping to a precomputed projective matrix. Third, we propose the anchored neighborhood regression. That is to anchor the neighborhood embedding of a low resolution patch to the nearest atom in the dictionary and to precompute the corresponding embedding matrix. These proposals are contrasted with current state-of-the-art methods on standard images. We obtain similar or improved quality and one or two orders of magnitude speed improvements.
Radu Timofte, Vincent De Smet, Luc Van Gool
ICCV1
2013 Traffic sign recognition - How far are we from the solution?
abstract
Traffic sign recognition has been a recurring application domain for visual objects detection. The public datasets have only recently reached large enough size and variety to enable proper empirical studies. We revisit the topic by showing how modern methods perform on two large detection and classification datasets (thousand of images, tens of categories) captured in Belgium and Germany. We show that, without any application specific modification, existing methods for pedestrian detection, and for digit and face classification; can reach performances in the range of 95% ~ 99% of the perfect solution. We show detailed experiments and discuss the trade-off of different options. Our top performing methods use modern variants of HOG features for detection, and sparse representations for classification.
Markus Mathias, Radu Timofte, Rodrigo Benenson, Luc Van Gool
IJCNN2
2012 Automatic Stave Discovery for Musical Facsimiles
Radu Timofte, Luc Van Gool
ACCV (4)1
2012 Naive Bayes Image Classification: Beyond Nearest Neighbors
Radu Timofte, Tinne Tuytelaars, Luc Van Gool
ACCV (1)1
2012 A Training-free Classification Framework for Textures, Writers, and Materials
abstract
We advocate the idea of a training-free texture classification scheme. This we demonstrate not only for traditional texture benchmarks, but also for the identification of materials and of the writers of musical scores. State-of-the-art methods operate using local descriptors, their intermediate representation over trained dictionaries, and classifiers. For the first two steps, we work with pooled local Gaussian derivative filters and a small dictionary not obtained through training, resp. Moreover, we build a multi-level representation similar to a spatial pyramid which captures region-level information. An extra step robustifies the final representation by means of comparative reasoning. As to the classification step, we achieve robust results using nearest neighbor classification, and state-of-the-art results with a collaborative strategy. Also these classifiers need no training. To the best of our knowledge, the proposed system yields top results on five standard benchmarks: 99.4% for CUReT, 97.3% for Brodatz, 99.5% for UMD, 99.4% for KTHTIPS, and 99% for UIUC. We significantly improve the state-of-the-art for three other benchmarks: KTHTIPS2b - 66.3% (from 58.1%), CVC-MUSCIMA - 99.8% (from 77.0%), and FMD - 55.8% (from 54%).
Radu Timofte, Luc Van Gool
BMVC1
2012 Pedestrian detection at 100 frames per second
abstract
We present a new pedestrian detector that improves both in speed and quality over state-of-the-art. By efficiently handling different scales and transferring computation from test time to training time, detection speed is improved. When processing monocular images, our system provides high quality detections at 50 fps. We also propose a new method for exploiting geometric context extracted from stereo images. On a single CPU+GPU desktop machine, we reach 135 fps, when processing street scenes, from rectified input to detections output.
Rodrigo Benenson, Markus Mathias, Radu Timofte, Luc Van Gool
CVPR3
2012 Iterative Nearest Neighbors for classification and dimensionality reduction
abstract
Representative data in terms of a set of selected samples is of interest for various machine learning applications, e.g. dimensionality reduction and classification. The best-known techniques probably still are k-Nearest Neighbors (kNN) and its variants. Recently, richer representations have become popular. Examples are methods based on l1-regularized least squares (Sparse Representation (SR)) or l2-regularized least squares (Collaborative Representation (CR)), or on l1-constrained least squares (Local Linear Embedding (LLE)). We propose Iterative Nearest Neighbors (INN). This is a novel sparse representation that combines the power of SR and LLE with the computational simplicity of kNN. We test our method in terms of dimensionality reduction and classification, using standard benchmarks such as faces (AR), traffic signs (GTSRB), and PASCAL VOC 2007. INN performs better than NN and comparable with CR and SR, while being orders of magnitude faster than the latter.
Radu Timofte, Luc Van Gool
CVPR1
2012 Stixels Motion Estimation without Optical Flow Computation
Bertan Günyel, Rodrigo Benenson, Radu Timofte, Luc Van Gool
ECCV (6)3
2012 Weighted collaborative representation and classification of images
Radu Timofte, Luc Van Gool
ICPR1
2012 Non-parametric motion-priors for flow understanding
abstract
We present a novel method for extracting the dominant dynamic properties of crowded scenes from a single, static, uncalibrated camera using a codebook of tracklets. Our approach relies only on tracklets of fixed length which are generated based on sparse optical flow. A grid of points is placed on the image plane and local meanshift clustering is employed to extract dominant directions of tracklets in the neighborhood. A Gaussian Process (GP) is fitted to each tracklet resulting in a codebook, with each codeword representing a local motion model. At test time, a mixture of weighted local GP experts is applied, providing multimodal density estimates for next object location and simulation of full object trajectories. Our scenarios come from challenging crowded scenes, from which we extract dominant local motion-patterns and use the model to simulate full object trajectories. In addition, we apply the learnt model to multiple object tracking. Random trajectories are sampled from the model that match the learnt scene dynamics. Minimum Description Length (MDL) is employed to pick the best trajectories in order to associate sparse detections over short time windows. Also, we modify a state-of-the-art multiple object tracking algorithm leading to significant improvement. Our results compare favorably to a state-of-the-art algorithm and we introduce a new challenging dataset for multiple object tracking.
Vasilis Lasdas, Radu Timofte, Luc Van Gool
WACV2
2011 Sparse Representation Based Projections
abstract
In dimensionality reduction most methods aim at preserving one or a few properties of the original space in the resulting embedding.As our results show, preserving the sparse representation of the signals from the original space in the (lower) dimensional projected space is beneficial for several benchmarks (faces, traffic signs, and handwritten digits).The intuition behind is that taking a sparse representation for the different samples as point of departure highlights the important correlations among the samples that one then wants to exploit to arrive at the final, effective low-dimensional embedding.We explicitly adapt the LPP and LLE techniques to work with the sparse representation criterion and compare to the original methods on the referenced databases, and this for both unsupervised and supervised cases.The improved results corroborate the usefulness of the proposed sparse representation based linear and non-linear projections.
Radu Timofte, Luc Van Gool
BMVC1
2010 Four Color Theorem for Fast Early Vision
Radu Timofte, Luc Van Gool
ACCV (1)1
2010 Hough Transform and 3D SURF for Robust Three Dimensional Classification
Jan Knopp, Mukta Prasad, Geert Willems, Radu Timofte, Luc Van Gool
ECCV (6)4
2010 Integrating Object Detection with 3D Tracking Towards a Better Driver Assistance System
abstract
Driver assistance helps save lives. Accurate 3D pose is required to establish if a traffic sign is relevant to the driver. We propose a real-time system that integrates single view detection with region-based 3D tracking of road signs. The optimal set of candidate detections is found, followed by AdaBoost cascades and SVMs. The 2D detections are then employed in simultaneous 2D segmentation and 3D pose tracking, using the known 3D model of the recognised traffic sign. We demonstrate the abilities of our system by tracking multiple road signs in real world scenarios.
Victor Adrian Prisacariu, Radu Timofte, Karel Zimmermann, Ian D. Reid 0001, Luc Van Gool
ICPR2
2009 Multi-view traffic sign detection, recognition, and 3D localisation
abstract
Several applications require information about street furniture. Part of the task is to survey all traffic signs. This has to be done for millions of km of road, and the exercise needs to be repeated every so often. A van with 8 roof-mounted cameras drove through the streets and took images every meter. The paper proposes a pipeline for the efficient detection and recognition of traffic signs. The task is challenging, as illumination conditions change regularly, occlusions are frequent, 3D positions and orientations vary substantially, and the actual signs are far less similar among equal types than one might expect. We combine 2D and 3D techniques to improve results beyond the state-of-the-art, which is still very much preoccupied with single view analysis.
Radu Timofte, Karel Zimmermann, Luc Van Gool
WACV1