EDBT 2026 Demo / reviewers in the wild / expert
Luis Herranz
dblp:64/224
· DBLP profile ↗
75ranked-venue papers
15as first author
34since 2021 · last 2026
0000-0002-7022-3395ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 57 · 14 first-author · 22 since 2021Artificial intelligence and machine learning · 31 · 2 first-author · 16 since 2021Databases, data management, data science and information retrieval · 2Applied, interdisciplinary, general and emerging computing · 2 · 1 first-authorSystems, architecture and hardware · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Adaptive Blind All-in-One Image RestorationabstractBlind all-in-one image restoration models aim to recover a high-quality image from an input degraded with unknown distortions. However, these models require all the possible degradation types to be defined during the training stage while showing limited generalization to unseen degradations, which limits their practical application in complex cases. In this paper, we introduce ABAIR, a simple yet effective adaptive blind all-in-one restoration model that not only handles multiple degradations and generalizes well to unseen distortions but also efficiently integrates new degradations by training only a small subset of parameters. We first train our baseline model on a large dataset of natural images with multiple synthetic degradations. To enhance its ability to recognize distortions, we incorporate a segmentation head that estimates per-pixel degradation types. Second, we adapt our initial model to varying image restoration tasks using independent low-rank adapters. Third, we learn to adaptively combine adapters to versatile images via a flexible and lightweight degradation estimator. This specialize-then-merge approach is both powerful in addressing specific distortions and flexible in adapting to complex tasks. Moreover, our model not only surpasses state-of-the-art performance on five- and three-task IR setups but also demonstrates superior generalization to unseen degradations and composite distortions. David Serrano-Lozano, Luis Herranz, Shaolin Su, Javier Vazquez-Corral |
Comput. Vis. Image Underst. | 2 |
| 2026 | Text and Non-Text Latent Feature Disentanglement for Screen Content Image CompressionabstractWith the growing prevalence of screen content images in multimedia communication, efficient compression has become increasingly crucial. Unlike natural scene images, screen content typically contains rich text regions that exhibit unique characteristics and low correlation with surrounding non-text elements. The intricate mixture of text and non-text within images poses significant challenges for existing learned compression networks, as the text and non-text features are severely entangled in the latent domain along the channel dimension, leading to compromised reconstruction quality and suboptimal entropy estimation. In this paper, we propose a novel Disentangled Image Compression Architecture (DICA) that enhances the analysis module and the entropy model of existing compression architectures to address these limitations. First, we introduce a Disentangled Analysis Module (DAM) by augmenting original analysis modules with an additional text approximation branch and a disentangling network. They work in concert to disentangle latent features into text and non-text classes along the channel dimension, resulting in a more structured feature distribution that better aligns with compression requirements. Second, we propose a Disentangled Channel-Conditional Entropy Model (DCEM) that efficiently leverages the feature distribution bias introduced by DAM, thereby further improving compression performance. Experimental results demonstrate that the proposed DICA, along with DAM and DCEM can be integrated into various channel-conditional compression backbones, significantly improving their performance in screen content compression—particularly in hard-to-compress text regions. When integrated with an advanced WACNN backbone, our method achieves a 13% overall BD-Rate gain and a 16% BD-Rate gain in text regions on the SIQAD dataset. Hao Wang 0184, Junyan Huo, Fei Yang 0004, Shuai Wan, Gaoxing Chen, Luis Herranz, Fuzheng Yang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 7 |
| 2026 | Camera Motion-Conditioned Motion Estimation for Neural Video Coding for Cloud GamingabstractThe rising popularity of cloud gaming highlights the importance of effective video compression for this domain. Despite the potential of neural video codecs to surpass traditional codecs, their application to cloud gaming videos generally achieves limited performance primarily due to two reasons: (1) most advanced neural video codecs are designed for natural videos with only a few optimized for cloud gaming content, and (2) the distinctive characteristics of cloud gaming videos including abrupt and large camera movements, coupled with repetitive textures, make the direct application of existing codecs suboptimal. To this end, this paper introduces an enhanced neural video codec for cloud gaming with camera motion-conditioned motion estimation. By leveraging the unique feature of cloud gaming, we effectively utilize the camera motion information to condition motion estimation with an attention mechanism, overcoming the pronounced challenges of motion estimation and facilitating the accurate learning of optical flow. Furthermore, we optimize the loss function of the motion estimation network by employing a comprehensive loss across multiple feature levels, ensuring the learned optical flow is multi-faceted and well-suited to subsequent motion compensation. Extensive experiments demonstrate the effectiveness of the proposed codec, providing improved motion estimation capabilities and superior rate-distortion performance. Moreover, the robustness of our codec to rapid camera movements is validated, which makes it highly suitable for cloud gaming scenarios. Fei Yang 0004, Luka Murn, Juil Sock, Marc Gorriz, Shuai Wan, Wei Zhang 0072, Fuzheng Yang 0001, Luis Herranz |
IEEE Trans. Multim. | 9 |
| 2025 | Revisiting Image Fusion for Multi-Illuminant White-Balance CorrectionabstractWhite balance (WB) correction in scenes with multiple illuminants remains a persistent challenge in computer vision. Recent methods explored fusion-based approaches, where a neural network linearly blends multiple sRGB versions of an input image, each processed with predefined WB presets. However, we demonstrate that these methods are suboptimal for common multi-illuminant scenarios. Additionally, existing fusion-based methods rely on sRGB WB datasets lacking dedicated multi-illuminant images, limiting both training and evaluation. To address these challenges, we introduce two key contributions. First, we propose an efficient transformer-based model that effectively captures spatial dependencies across sRGB WB presets, substantially improving upon linear fusion techniques. Second, we introduce a large-scale multi-illuminant dataset comprising over 16,000 sRGB images rendered with five different WB settings, along with WB-corrected images. Our method achieves up to 100\% improvement over existing techniques on our new multi-illuminant image fusion dataset. David Serrano-Lozano, Aditya Arora, Luis Herranz, Konstantinos G. Derpanis, Michael S. Brown, Javier Vazquez-Corral |
ICCV | 3 |
| 2025 | Enhanced neural video compression for cloud gaming videos with aligned frame generationabstractThe burgeoning popularity of cloud gaming makes it critical for efficient video compression to relieve the growing bandwidth pressure. While existing neural video coding approaches have demonstrated strong compression potential on natural videos, there is an absence of efficient neural codecs dedicated to gaming videos. To bridge this gap, in this paper, we propose an end-to-end neural video compression method designed specifically for cloud gaming videos. By effectively utilizing the unique camera motion information inherent to cloud gaming, the previous reconstructed frame is maximally aligned to the current frame through a learningbased module with multiple losses, which then replaces the previous reconstructed frame for optical flow estimation. By significantly reducing the displacement between two consecutive frames caused by camera motion, the motion estimation accuracy is enhanced, effectively handling the large and abrupt motion scenarios frequently present in gaming videos. Furthermore, the aligned tensor obtained in the previous step is used to enhance the latent prior of the entropy model, providing a superior temporal prior for coding. Extensive experimental results demonstrate the superior performance of our proposed method compared to one of the previous state-of-the-art approaches, DCVC-HEM, providing significant progress in end-to-end neural compression in cloud gaming videos Fei Yang 0004, Luka Murn, Juil Sock, Marc Gorriz, Shuai Wan, Wei Zhang 0072, Fuzheng Yang 0001, Luis Herranz |
Expert Syst. Appl. | 9 |
| 2025 | Curriculum learning-based slimmable cross-component prediction for video coding
Chengyi Zou, Shuai Wan, Marc Gorriz, Luka Murn, Juil Sock, Fei Yang 0004, Luis Herranz |
Neurocomputing | 7 |
| 2025 | Lightweight Deep Exemplar Colorization via Semantic Attention-Guided Laplacian PyramidabstractExemplar-based colorization aims to generate plausible colors for a grayscale image with the guidance of a color reference image. The main challenging problem is finding the correct semantic correspondence between the target image and the reference image. However, the colors of the object and background are often confused in the existing methods. Besides, these methods usually use simple encoder-decoder architectures or pyramid structures to extract features and lack appropriate fusion mechanisms, which results in the loss of high-frequency information or high complexity. To address these problems, this article proposes a lightweight semantic attention-guided Laplacian pyramid network (SAGLP-Net) for deep exemplar-based colorization, exploiting the inherent multi-scale properties of color representations. They are exploited through a Laplacian pyramid, and semantic information is introduced as high-level guidance to align the object and background information. Specially, a semantic guided non-local attention fusion module is designed to exploit the long-range dependency and fuse the local and global features. Moreover, a Laplacian pyramid fusion module based on criss-cross attention is proposed to fuse high frequency components in the large-scale domain. An unsupervised multi-scale multi-loss training strategy is further introduced for network training, which combines pixel loss, color histogram loss, total variance regularisation, and adversarial loss. Experimental results demonstrate that our colorization method achieves better subjective and objective performance with lower complexity than the state-of-the-art methods. Chengyi Zou, Shuai Wan, Marc Gorriz, Luka Murn, Marta Mrak, Juil Sock, Fei Yang 0004, Luis Herranz |
IEEE Trans. Vis. Comput. Graph. | 8 |
| 2024 | NamedCurves: Learned Image Enhancement via Color Naming
David Serrano-Lozano, Luis Herranz, Michael S. Brown, Javier Vazquez-Corral |
ECCV (71) | 2 |
| 2024 | Rate Control for Slimmable Video Codec using Multilayer PerceptronabstractAs a flexible design for practical end-to-end video compression, slimmable video coding achieves variable rate coding with an adaptable complexity. In this paper, rate control using a multilayer perceptron is proposed to achieve accurate bitrate control for slimmable video coding. An average bitrate error of less than 0.3% is achieved on tested sequences at a slight decrease in the rate-distortion performance. Furthermore, the results show that there is no overflow or underflow in the buffer occupancy during the rate control process. Defa Wang, Shuai Wan, Fei Yang 0004, Luis Herranz |
ISCAS | 5 |
| 2024 | MineGAN++: Mining Generative Models for Efficient Knowledge Transfer to Limited Data Domains
Yaxing Wang, Abel Gonzalez-Garcia, Chenshen Wu, Luis Herranz, Fahad Shahbaz Khan, Shangling Jui, Jian Yang 0003, Joost van de Weijer 0001 |
Int. J. Comput. Vis. | 4 |
| 2024 | A slimmable framework for practical neural video compressionabstractDeep learning is being increasingly applied to image and video compression in a new paradigm known as neural video compression. While achieving impressive rate–distortion (RD) performance, neural video codecs (NVC) require heavy neural networks, which in turn have large memory and computational costs and often lack important functionalities such as variable rate. These are significant limitations to their practical application. Addressing these problems, recent slimmable image codecs can dynamically adjust their model capacity to elegantly reduce the memory and computation requirements, without harming RD performance. However, the extension to video is not straightforward due to the non-trivial interplay with complex motion estimation and compensation modules in most NVC architectures. In this paper we propose the slimmable video codec framework (SlimVC) that integrates an slimmable autoencoder and a motion-free conditional entropy model. We show that the slimming mechanism is also applicable to the more complex case of video architectures, providing SlimVC with simultaneous control of the computational cost, memory and rate, which are all important requirements in practice. We further provide detailed experimental analysis, and describe application scenarios that can benefit from slimmable video codecs. Zhaocheng Liu, Fei Yang 0004, Defa Wang, Marc Gorriz, Luka Murn, Shuai Wan, Saiping Zhang, Marta Mrak, Luis Herranz |
Neurocomputing | 9 |
| 2024 | Main product detection with graph networks for fashionabstractAbstract Computer vision has established a foothold in the online fashion retail industry. Main product detection is a crucial step of vision-based fashion product feed parsing pipelines, focused on identifying the bounding boxes that contain the product being sold in the gallery of images of the product page. The current state-of-the-art approach does not leverage the relations between regions in the image, and treats images of the same product independently, therefore not fully exploiting visual and product contextual information. In this paper, we propose a model that incorporates Graph Convolutional Networks (GCN) that jointly represent all detected bounding boxes in the gallery as nodes. We show that the proposed method is better than the state-of-the-art, especially, when we consider the scenario where title-input is missing at inference time and for cross-dataset evaluation, our method outperforms previous approaches by a large margin. Vacit Oguz Yazici, Arnau Ramisa, Luis Herranz, Joost van de Weijer 0001 |
Multim. Tools Appl. | 4 |
| 2024 | Palette-Based Color Harmonization via Color NamingabstractColor harmony refers to combinations of colors that look pleasing together. We present a novel strategy to harmonize an image's colors using color-palette manipulation and color naming. Palette-based color manipulation is a method that extracts a few colors to represent the image. Modifying the palette colors modifies the color appearance of the image. A color-naming model is a mechanism to categorize colors into a fixed number of basic color terms. Working from a color-naming model, we derive a set ofprototype colorsand demonstrate that mapping an image's extracted color palette to the nearest prototype colors effectively harmonizes the image's colors. This straightforward approach yields visually compelling, outperforming more complex color harmony methods. Danna Xue, Javier Vazquez-Corral, Luis Herranz, Yanning Zhang 0001, Michael S. Brown |
IEEE Signal Process. Lett. | 3 |
| 2024 | Task-Switchable Pre-Processor for Image Compression for Multiple Machine Vision TasksabstractVisual content is increasingly being processed by machines for various automated content analysis tasks instead of being consumed by humans. Despite the existence of several compression methods tailored for machine tasks, few consider real-world scenarios with multiple tasks. In this paper, we aim to address this gap by proposing a task-switchable pre-processor that optimizes input images specifically for machine consumption prior to encoding by an off-the-shelf codec designed for human consumption. The proposed task-switchable pre-processor adeptly maintains relevant semantic information based on the specific characteristics of different downstream tasks, while effectively suppressing irrelevant information to reduce bitrate. To enhance the processing of semantic information for diverse tasks, we leverage pre-extracted semantic features to modulate the pixel-to-pixel mapping within the pre-processor. By switching between different modulations, multiple tasks can be seamlessly incorporated into the system. Extensive experiments demonstrate the practicality and simplicity of our approach. It significantly reduces the number of parameters required for handling multiple tasks while still delivering impressive performance. Our method showcases the potential to achieve efficient and effective compression for machine vision tasks, supporting the evolving demands of real-world applications. Mingyi Yang, Fei Yang 0004, Luka Murn, Marc Gorriz, Juil Sock, Shuai Wan, Fuzheng Yang 0001, Luis Herranz |
IEEE Trans. Circuits Syst. Video Technol. | 8 |
| 2024 | Chroma Intra Prediction With Lightweight Attention-Based Neural NetworksabstractNeural networks can be successfully used for cross-component prediction in video coding. In particular, attention-based architectures are suitable for chroma intra prediction using luma information because of their capability to model relations between difierent channels. However, the complexity of such methods is still very high and should be further reduced, especially for decoding. In this paper, a cost-effective attention-based neural network is designed for chroma intra prediction. Moreover, with the goal of further improving coding performance, a novel approach is introduced to utilize more boundary information effectively. In addition to improving prediction, a simplification methodology is also proposed to reduce inference complexity by simplifying convolutions. The proposed schemes are integrated into H.266/Versatile Video Coding (VVC) pipeline, and only one additional binary block-level syntax flag is introduced to indicate whether a given block makes use of the proposed method. Experimental results demonstrate that the proposed scheme achieves up to −0.46%/−2.29%/−2.17% BD-rate reduction on Y/Cb/Cr components, respectively, compared with H.266/VVC anchor. Reductions in the encoding and decoding complexity of up to 22% and 61%, respectively, are achieved by the proposed scheme with respect to the previous attention-based chroma intra prediction method while maintaining coding performance. Chengyi Zou, Shuai Wan, Tiannan Ji, Marc Gorriz, Marta Mrak, Luis Herranz |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2023 | Efficient Super-Resolution for Compression Of Gaming VideosabstractDue to the increasing demand for game-streaming services, efficient compression of computer-generated video is more critical than ever, especially when the available bandwidth is low. This paper proposes a super-resolution framework that improves the coding efficiency of computer-generated gaming videos at low bitrates. Most state-of-the-art super-resolution networks generalize over a variety of RGB inputs and use a unified network architecture for frames of different levels of degradation, leading to high complexity and redundancy. Since games usually consist of a limited number of fixed scenarios, we specialize one model for each scenario and assign appropriate network capacities for different QPs to perform super-resolution under the guidance of reconstructed high-quality luma components. Experimental results show that our framework achieves a superior quality-complexity trade-off compared to the ESRnet baseline, saving at most 93.59% parameters while maintaining comparable performance. The compression efficiency compared to HEVC is also improved by more than 17% BD-rate gain. Luka Murn, Luis Herranz, Fei Yang 0004, Marta Mrak, Wei Zhang 0072, Shuai Wan, Marc Gorriz |
ICASSP | 3 |
| 2023 | Burst Perception-Distortion Tradeoff: Analysis and EvaluationabstractBurst image restoration attempts to effectively utilize the complementary cues appearing in sequential images to produce a high-quality image. Most current methods use all the available images to obtain the reconstructed image. However, using more images for burst restoration is not always the best option regarding reconstruction quality and efficiency, as the images acquired by handheld imaging devices suffer from degradation and misalignment caused by the camera noise and shake. In this paper, we extend the perception-distortion tradeoff theory by introducing multiple-frame information. We propose the area of the unattainable region as a new metric for perception-distortion tradeoff evaluation and comparison. Based on this metric, we analyse the performance of burst restoration from the perspective of the perception-distortion tradeoff under both aligned bursts and misaligned bursts situations. Our analysis reveals the importance of inter-frame alignment for burst restoration and shows that the optimal burst length for the restoration model depends both on the degree of degradation and misalignment. Danna Xue, Luis Herranz, Javier Vazquez-Corral, Yanning Zhang 0001 |
ICASSP | 2 |
| 2023 | Semantic Preprocessor for Image Compression for MachinesabstractVisual content is being increasingly transmitted and consumed by machines rather than humans to perform automated content analysis tasks. In this paper, we propose an image preprocessor that optimizes the input image for machine consumption prior to encoding by an off-the-shelf codec designed for human consumption. To achieve a better trade-off between the accuracy of the machine analysis task and bitrate, we propose leveraging pre-extracted semantic information to improve the preprocessor’s ability to accurately identify and filter out task-irrelevant information. Furthermore, we propose a two-part loss function to optimize the preprocessor, consisted of a rate-task performance loss and a semantic distillation loss, which helps the reconstructed image obtain more information that contributes to the accuracy of the task. Experiments show that the proposed preprocessor can save up to 48.83% bitrate compared with the method without the preprocessor, and save up to 36.24% bitrate compared to existing preprocessors for machine vision. Mingyi Yang, Luis Herranz, Fei Yang 0004, Luka Murn, Marc Gorriz, Shuai Wan, Fuzheng Yang 0001, Marta Mrak |
ICASSP | 2 |
| 2023 | Integrating High-Level Features for Consistent Palette-based Multi-image RecoloringabstractAbstract Achieving visually consistent colors across multiple images is important when images are used in photo albums, websites, and brochures. Unfortunately, only a handful of methods address multi‐image color consistency compared to one‐to‐one color transfer techniques. Furthermore, existing methods do not incorporate high‐level features that can assist graphic designers in their work. To address these limitations, we introduce a framework that builds upon a previous palette‐based color consistency method and incorporates three high‐level features: white balance, saliency, and color naming. We show how these features overcome the limitations of the prior multi‐consistency workflow and showcase the user‐friendly nature of our framework. D. Xue, Javier Vazquez-Corral, Luis Herranz, Michael S. Brown |
Comput. Graph. Forum | 3 |
| 2023 | Casting a BAIT for offline and online source-free domain adaptation
Shiqi Yang 0002, Yaxing Wang, Luis Herranz, Shangling Jui, Joost van de Weijer 0001 |
Comput. Vis. Image Underst. | 3 |
| 2023 | Trust Your Good Friends: Source-Free Domain Adaptation by Reciprocal Neighborhood ClusteringabstractDomain adaptation (DA) aims to alleviate the domain shift between source domain and target domain. Most DA methods require access to the source data, but often that is not possible (e.g., due to data privacy or intellectual property). In this paper, we address the challenging source-free domain adaptation (SFDA) problem, where the source pretrained model is adapted to the target domain in the absence of source data. Our method is based on the observation that target data, which might not align with the source domain classifier, still forms clear clusters. We capture this intrinsic structure by defining local affinity of the target data, and encourage label consistency among data with high local affinity. We observe that higher affinity should be assigned to reciprocal neighbors. To aggregate information with more context, we consider expanded neighborhoods with small affinity values. Furthermore, we consider the density around each target sample, which can alleviate the negative impact of potential outliers. In the experimental results we verify that the inherent structure of the target features is an important source of information for domain adaptation. We demonstrate that this local structure can be efficiently captured by considering the local neighbors, the reciprocal neighbors, and the expanded neighborhood. Finally, we achieve state-of-the-art performance on several 2D image and 3D point cloud recognition datasets. Shiqi Yang 0002, Yaxing Wang, Joost van de Weijer 0001, Luis Herranz, Shangling Jui, Jian Yang 0003 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2022 | DCNGAN: A Deformable Convolution-Based GAN with QP Adaptation for Perceptual Quality Enhancement of Compressed VideoabstractIn this paper, we propose a deformable convolution-based generative adversarial network (DCNGAN) for perceptual quality enhancement of compressed videos. DCNGAN is also adaptive to the quantization parameters (QPs). Compared with optical flows, deformable convolutions are more effective and efficient to align frames. Deformable convolutions can operate on multiple frames, thus leveraging more temporal information, which is beneficial for enhancing the perceptual quality of compressed videos. Instead of aligning frames in a pairwise manner, the deformable convolution can process multiple frames simultaneously, which leads to lower computational complexity. Experimental results demonstrate that the proposed DCNGAN outperforms other state-of-the-art compressed video quality enhancement algorithms. Saiping Zhang, Luis Herranz, Marta Mrak, Marc Gorriz, Shuai Wan, Fuzheng Yang 0001 |
ICASSP | 2 |
| 2022 | Towards Lightweight Neural Network-based Chroma Intra Prediction for Video CodingabstractIn video compression the luma channel can be useful for predicting chroma channels (Cb, Cr), as has been demonstrated with the Cross-Component Linear Model (CCLM) used in Versatile Video Coding (VVC) standard. More recently, it has been shown that neural networks can even better capture the relationship among different channels. In this paper, a new attention-based neural network is proposed for cross-component intra prediction. With the goal to simplify neural network design, the new framework consists of four branches: boundary branch and luma branch for extracting features from reference samples, attention branch for fusing the first two branches, and prediction branch for computing the predicted chroma samples. The proposed scheme is integrated into VVC test model together with one additional binary block-level syntax flag which indicates whether a given block makes use of the proposed method. Experimental results demonstrate 0.31%/2.36%/2.00% BD-rate reductions on Y/Cb/Cr components, respectively, on top of the VVC Test Model (VTM) 7.0 which uses CCLM. Chengyi Zou, Shuai Wan, Marta Mrak, Marc Gorriz, Luis Herranz, Tiannan Ji |
ICIP | 5 |
| 2022 | SlimSeg: Slimmable Semantic Segmentation with Boundary SupervisionabstractAccurate semantic segmentation models typically require significant computational resources, inhibiting their use in practical applications. Recent works rely on well-crafted lightweight models to achieve fast inference. However, these models cannot flexibly adapt to varying accuracy and efficiency requirements. In this paper, we propose a simple but effective slimmable semantic segmentation (SlimSeg) method, which can be executed at different capacities during inference depending on the desired accuracy-efficiency tradeoff. More specifically, we employ parametrized channel slimming by stepwise downward knowledge distillation during training. Motivated by the observation that the differences between segmentation results of each submodel are mainly near the semantic borders, we introduce an additional boundary guided semantic segmentation loss to further improve the performance of each submodel. We show that our proposed SlimSeg with various mainstream networks can produce flexible models that provide dynamic adjustment of computational cost and better performance than independent models. Extensive experiments on semantic segmentation benchmarks, Cityscapes and CamVid, demonstrate the generalization ability of our framework. Danna Xue, Fei Yang 0004, Luis Herranz, Jinqiu Sun, Yu Zhu 0004, Yanning Zhang 0001 |
ACM Multimedia | 4 |
| 2022 | A novel framework for image-to-image translation and image compression
Fei Yang 0004, Yaxing Wang, Luis Herranz, Yongmei Cheng, Mikhail G. Mozerov |
Neurocomputing | 3 |
| 2021 | HCV: Hierarchy-Consistency Verification for Incremental Implicitly-Refined Classification
Kai Wang 0060, Xialei Liu, Luis Herranz, Joost van de Weijer 0001 |
BMVC | 3 |
| 2021 | Slimmable Compressive Autoencoders for Practical Neural Image CompressionabstractNeural image compression leverages deep neural networks to outperform traditional image codecs in rate-distortion performance. However, the resulting models are also heavy, computationally demanding and generally optimized for a single rate, limiting their practical use. Focusing on practical image compression, we propose slimmable compressive autoencoders (SlimCAEs), where rate (R) and distortion (D) are jointly optimized for different capacities. Once trained, encoders and decoders can be executed at different capacities, leading to different rates and complexities. We show that a successful implementation of Slim-CAEs requires suitable capacity-specific RD tradeoffs. Our experiments show that SlimCAEs are highly flexible models that provide excellent rate-distortion performance, variable rate, and dynamic adjustment of memory, computational cost and latency, thus addressing the main requirements of practical image compression. Fei Yang 0004, Luis Herranz, Yongmei Cheng, Mikhail G. Mozerov |
CVPR | 2 |
| 2021 | Generalized Source-free Domain AdaptationabstractDomain adaptation (DA) aims to transfer the knowledge learned from a source domain to an unlabeled target domain. Some recent works tackle source-free domain adaptation (SFDA) where only a source pre-trained model is available for adaptation to the target domain. However, those methods do not consider keeping source performance which is of high practical value in real world applications. In this paper, we propose a new domain adaptation paradigm called Generalized Source-free Domain Adaptation (G-SFDA), where the learned model needs to perform well on both the target and source domains, with only access to current unlabeled target data during adaptation. First, we propose local structure clustering (LSC), aiming to cluster the target features with its semantically similar neighbors, which successfully adapts the model to the target domain in the absence of source data. Second, we propose sparse domain attention (SDA), it produces a binary domain specific attention to activate different feature channels for different domains, meanwhile the domain attention will be utilized to regularize the gradient during adaptation to keep source information. In the experiments, for target performance our method is on par with or better than existing DA and SFDA methods, specifically it achieves state-of-the-art performance (85.4%) on VisDA, and our method works well for all domains after adapting to single or multiple target domains. Code is available in https://github.com/Albert0147/G-SFDA. Shiqi Yang 0002, Yaxing Wang, Joost van de Weijer 0001, Luis Herranz, Shangling Jui |
ICCV | 4 |
| 2021 | Exploiting the Intrinsic Neighborhood Structure for Source-free Domain AdaptationabstractDomain adaptation (DA) aims to alleviate the domain shift between source domain and target domain. Most DA methods require access to the source data, but often that is not possible (e.g. due to data privacy or intellectual property). In this paper, we address the challenging source-free domain adaptation (SFDA) problem, where the source pretrained model is adapted to the target domain in the absence of source data. Our method is based on the observation that target data, which might no longer align with the source domain classifier, still forms clear clusters. We capture this intrinsic structure by defining local affinity of the target data, and encourage label consistency among data with high local affinity. We observe that higher affinity should be assigned to reciprocal neighbors, and propose a self regularization loss to decrease the negative impact of noisy neighbors. Furthermore, to aggregate information with more context, we consider expanded neighborhoods with small affinity values. In the experimental results we verify that the inherent structure of the target features is an important source of information for domain adaptation. We demonstrate that this local structure can be efficiently captured by considering the local neighbors, the reciprocal neighbors, and the expanded neighborhood. Finally, we achieve state-of-the-art performance on several 2D image and 3D point cloud recognition datasets. Code is available in https://github.com/Albert0147/SFDA_neighbors. Shiqi Yang 0002, Yaxing Wang, Joost van de Weijer 0001, Luis Herranz, Shangling Jui |
NeurIPS | 4 |
| 2021 | DVC-P: Deep Video Compression with Perceptual OptimizationsabstractRecent years have witnessed the significant development of learning-based video compression methods, which aim at optimizing objective or perceptual quality and bit rates. In this paper, we introduce deep video compression with perceptual op-timizations (DVC-P), which aims at increasing perceptual quality of decoded videos. Our proposed DVC-P is based on Deep Video Compression (DVC) network, but improves it with perceptual optimizations. Specifically, a discriminator network and a mixed loss are employed to help our network trade off among distortion, perception and rate. Furthermore, nearest-neighbor interpolation is used to eliminate checkerboard artifacts which can appear in sequences encoded with DVC frameworks. Thanks to these two improvements, the perceptual quality of decoded sequences is improved. Experimental results demonstrate that, compared with the baseline DVC, our proposed method can generate videos with higher perceptual quality achieving 12.27% reduction in a perceptual BD- rate equivalent, on average. Saiping Zhang, Marta Mrak, Luis Herranz, Marc Gorriz, Shuai Wan, Fuzheng Yang 0001 |
VCIP | 3 |
| 2021 | Controlling biases and diversity in diverse image-to-image translation
Yaxing Wang, Abel Gonzalez-Garcia, Luis Herranz, Joost van de Weijer 0001 |
Comput. Vis. Image Underst. | 3 |
| 2021 | ACAE-REMIND for online continual learning with compressed feature replay
Kai Wang 0060, Joost van de Weijer 0001, Luis Herranz |
Pattern Recognit. Lett. | 3 |
| 2021 | On Implicit Attribute Localization for Generalized Zero-Shot LearningabstractZero-shot learning (ZSL) aims to discriminate images from unseen classes by exploiting relations to seen classes via their attribute-based descriptions. Since attributes are often related to specific parts of objects, many recent works focus on discovering discriminative regions. However, these methods usually require additional complex part detection modules or attention mechanisms. In this paper, 1) we show that common ZSL backbones (without explicit attention nor part detection) can implicitly localize attributes, yet this property is not exploited. 2) Exploiting it, we then propose SELAR, a simple method that further encourages attribute localization, surprisingly achieving very competitive generalized ZSL (GZSL) performance when compared with more complex state-of-the-art methods. Our findings provide useful insight for designing future GZSL methods, and SELAR provides an easy to implement yet strong baseline. Shiqi Yang 0002, Kai Wang 0060, Luis Herranz, Joost van de Weijer 0001 |
IEEE Signal Process. Lett. | 3 |
| 2021 | Distributed Learning and Inference With Compressed ImagesabstractModern computer vision requires processing large amounts of data, both while training the model and/or during inference, once the model is deployed. Scenarios where images are captured and processed in physically separated locations are increasingly common (e.g. autonomous vehicles, cloud computing, smartphones). In addition, many devices suffer from limited resources to store or transmit data (e.g. storage space, channel capacity). In these scenarios, lossy image compression plays a crucial role to effectively increase the number of images collected under such constraints. However, lossy compression entails some undesired degradation of the data that may harm the performance of the downstream analysis task at hand, since important semantic information may be lost in the process. Moreover, we may only have compressed images at training time but are able to use original images at inference time (i.e. test), or vice versa, and in such a case, the downstream model suffers from covariate shift. In this paper, we analyze this phenomenon, with a special focus on vision-based perception for autonomous driving as a paradigmatic scenario. We see that loss of semantic information and covariate shift do indeed exist, resulting in a drop in performance that depends on the compression rate. In order to address the problem, we propose dataset restoration, based on image restoration with generative adversarial networks (GANs). Our method is agnostic to both the particular image compression method and the downstream task; and has the advantage of not adding additional cost to the deployed models, which is particularly important in resource-limited devices. The presented experiments focus on semantic segmentation as a challenging use case, cover a broad range of compression rates and diverse datasets, and show how our method is able to significantly alleviate the negative effects of compression on the downstream visual task. Sudeep Katakol, Basem Elbarashy, Luis Herranz, Joost van de Weijer 0001, Antonio M. López 0001 |
IEEE Trans. Image Process. | 3 |
| 2020 | Semantic Drift Compensation for Class-Incremental LearningabstractClass-incremental learning of deep networks sequentially increases the number of classes to be classified. During training, the network has only access to data of one task at a time, where each task contains several classes. In this setting, networks suffer from catastrophic forgetting which refers to the drastic drop in performance on previous tasks. The vast majority of methods have studied this scenario for classification networks, where for each new task the classification layer of the network must be augmented with additional weights to make room for the newly added classes. Embedding networks have the advantage that new classes can be naturally included into the network without adding new weights. Therefore, we study incremental learning for embedding networks. In addition, we propose a new method to estimate the drift, called semantic drift, of features and compensate for it without the need of any exemplars. We approximate the drift of previous tasks based on the drift that is experienced by current task data. We perform experiments on fine-grained datasets, CIFAR100 and ImageNet-Subset. We demonstrate that embedding networks suffer significantly less from catastrophic forgetting. We outperform existing methods which do not require exemplars and obtain competitive results compared to methods which store exemplars. Furthermore, we show that our proposed SDC when combined with existing methods to prevent forgetting consistently improves results. Lu Yu 0004, Bartlomiej Twardowski, Xialei Liu, Luis Herranz, Kai Wang 0060, Yongmei Cheng, Shangling Jui, Joost van de Weijer 0001 |
CVPR | 4 |
| 2020 | MineGAN: Effective Knowledge Transfer From GANs to Target Domains With Few ImagesabstractOne of the attractive characteristics of deep neural networks is their ability to transfer knowledge obtained in one domain to other related domains. As a result, high-quality networks can be trained in domains with relatively little training data. This property has been extensively studied for discriminative networks but has received significantly less attention for generative models. Given the often enormous effort required to train GANs, both computationally as well as in the dataset collection, the re-use of pretrained GANs is a desirable objective. We propose a novel knowledge transfer method for generative models based on mining the knowledge that is most beneficial to a specific target domain, either from a single or multiple pretrained GANs. This is done using a miner network that identifies which part of the generative distribution of each pretrained GAN outputs samples closest to the target domain. Mining effectively steers GAN sampling towards suitable regions of the latent space, which facilitates the posterior finetuning and avoids pathologies of other methods such as mode collapse and lack of flexibility. We perform experiments on several complex datasets using various GAN architectures (BigGAN, Progressive GAN) and show that the proposed method, called MineGAN, effectively transfers knowledge to domains with few target images, outperforming existing methods. In addition, MineGAN can successfully transfer knowledge from multiple pretrained GANs. Our code is available at: \url{https://github.com/yaxingwang/MineGAN}. Yaxing Wang, Abel Gonzalez-Garcia, David Berga, Luis Herranz, Fahad Shahbaz Khan, Joost van de Weijer 0001 |
CVPR | 4 |
| 2020 | Generalized Zero-shot Learning with Multi-source Semantic Embeddings for Scene RecognitionabstractRecognizing visual categories from semantic descriptions is a promising way to extend the capability of a visual classifier beyond the concepts represented in the training data (i.e. seen categories). This problem is addressed by (generalized) zero-shot learning methods (GZSL), which leverage semantic descriptions that connect them to seen categories (e.g. label embedding, attributes). Conventional GZSL are designed mostly for object recognition. In this paper we focus on zero-shot scene recognition, a more challenging setting with hundreds of categories where their differences can be subtle and often localized in certain objects or regions. Conventional GZSL representations are not rich enough to capture these local discriminative differences. Addressing these limitations, we propose a feature generation framework with two novel components: 1) multiple sources of semantic information (i.e. attributes, word embeddings and descriptions), 2) region descriptions that can enhance scene discrimination. To generate synthetic visual features we propose a two-step generative approach, where local descriptions are sampled and used as conditions to generate visual features. The generated features are then aggregated and used together with real features to train a joint classifier. In order to evaluate the proposed method, we introduce a new dataset for zero-shot scene recognition with multi-semantic annotations. Experimental results on the proposed dataset and SUN Attribute dataset illustrate the effectiveness of the proposed method. Xinhang Song, Haitao Zeng, Sixian Zhang, Luis Herranz, Shuqiang Jiang |
ACM Multimedia | 4 |
| 2020 | Mix and Match Networks: Cross-Modal Alignment for Zero-Pair Image-to-Image Translation
Yaxing Wang, Luis Herranz, Joost van de Weijer 0001 |
Int. J. Comput. Vis. | 2 |
| 2020 | Variable Rate Deep Image Compression With Modulated AutoencoderabstractVariable rate is a requirement for flexible and adaptable image and video compression. However, deep image compression methods (DIC) are optimized for a single fixed rate-distortion (R-D) tradeoff. While this can be addressed by training multiple models for different tradeoffs, the memory requirements increase proportionally to the number of models. Scaling the bottleneck representation of a shared autoencoder can provide variable rate compression with a single shared autoencoder. However, the R-D performance using this simple mechanism degrades in low bitrates, and also shrinks the effective range of bitrates. To address these limitations, we formulate the problem of variable R-D optimization for DIC, and propose modulated autoencoders (MAEs), where the representations of a shared autoencoder are adapted to the specific R-D tradeoff via a modulation network. Jointly training this modulated autoencoder and the modulation network provides an effective way to navigate the R-D operational curve. Our experiments show that the proposed method can achieve almost the same R-D performance of independent models with significantly fewer parameters. Fei Yang 0004, Luis Herranz, Joost van de Weijer 0001, José Antonio Iglesias Guitián, Antonio M. López 0001, Mikhail G. Mozerov |
IEEE Signal Process. Lett. | 2 |
| 2019 | SDIT: Scalable and Diverse Cross-domain Image TranslationabstractRecently, image-to-image translation research has witnessed remarkable progress. Although current approaches successfully generate diverse outputs or perform scalable image transfer, these properties have not been combined into a single method. To address this limitation, we propose SDIT: Scalable and Diverse image-to-image translation. These properties are combined into a single generator. The diversity is determined by a latent variable which is randomly sampled from a normal distribution. The scalability is obtained by conditioning the network on the domain attributes. Additionally, we also exploit an attention mechanism that permits the generator to focus on the domain-specific attribute. We empirically demonstrate the performance of the proposed method on face mapping and other datasets beyond faces. Yaxing Wang, Abel Gonzalez-Garcia, Joost van de Weijer 0001, Luis Herranz |
ACM Multimedia | 4 |
| 2019 | Learning Effective RGB-D Representations for Scene RecognitionabstractDeep convolutional networks (CNN) can achieve impressive results on RGB scene recognition thanks to large datasets such as Places. In contrast, RGB-D scene recognition is still underdeveloped in comparison, due to two limitations of RGB-D data we address in this paper. The first limitation is the lack of depth data for training deep learning models. Rather than fine tuning or transferring RGB-specific features, we address this limitation by proposing an architecture and a twostep training approach that directly learns effective depth-specific features using weak supervision via patches. The resulting RGBD model also benefits from more complementary multimodal features. Another limitation is the short range of depth sensors (typically 0.5m to 5.5m), resulting in depth images not capturing distant objects in the scenes that RGB images can. We show that this limitation can be addressed by using RGB-D videos, where more comprehensive depth information is accumulated as the camera travels across the scenes. Focusing on this scenario, we introduce the ISIA RGB-D video dataset to evaluate RGB-D scene recognition with videos. Our video recognition architecture combines convolutional and recurrent neural networks (RNNs) that are trained in three steps with increasingly complex data to learn effective features (i.e. patches, frames and sequences). Our approach obtains state-of-the-art performances on RGB-D image (NYUD2 and SUN RGB-D) and video (ISIA RGB-D) scene recognition. Xinhang Song, Shuqiang Jiang, Luis Herranz, Chengpeng Chen |
IEEE Trans. Image Process. | 3 |
| 2018 | Mix and Match Networks: Encoder-Decoder Alignment for Zero-Pair Image TranslationabstractWe address the problem of image translation between domains or modalities for which no direct paired data is available (i.e. zero-pair translation). We propose mix and match networks, based on multiple encoders and decoders aligned in such a way that other encoder-decoder pairs can be composed at test time to perform unseen image translation tasks between domains or modalities for which explicit paired samples were not seen during training. We study the impact of autoencoders, side information and losses in improving the alignment and transferability of trained pairwise translation models to unseen translations. We show our approach is scalable and can perform colorization and style transfer between unseen combinations of domains. We evaluate our system in a challenging cross-modal setting where semantic segmentation is estimated from depth images, without explicit access to any depth-semantic segmentation training pairs. Our model outperforms baselines based on pix2pix and CycleGAN models. Yaxing Wang, Joost van de Weijer 0001, Luis Herranz |
CVPR | 3 |
| 2018 | Transferring GANs: Generating Images from Limited Data
Yaxing Wang, Chenshen Wu, Luis Herranz, Joost van de Weijer 0001, Abel Gonzalez-Garcia, Bogdan Raducanu |
ECCV (6) | 3 |
| 2018 | Rotate your Networks: Better Weight Consolidation and Less Catastrophic ForgettingabstractIn this paper we propose an approach to avoiding catastrophic forgetting in sequential task learning scenarios. Our technique is based on a network reparameterization that approximately diagonalizes the Fisher Information Matrix of the network parameters. This reparameterization takes the form of a factorized rotation of parameter space which, when used in conjunction with Elastic Weight Consolidation (which assumes a diagonal Fisher Information Matrix), leads to significantly better performance on lifelong learning of sequential tasks. Experimental results on the MNIST, CIFAR-100, CUB-200 and Stanford-40 datasets demonstrate that we significantly improve the results of standard elastic weight consolidation, and that we obtain competitive results when compared to the state-of-the-art in lifelong learning without forgetting. Xialei Liu, Marc Masana, Luis Herranz, Joost van de Weijer 0001, Antonio M. López 0001, Andrew D. Bagdanov |
ICPR | 3 |
| 2018 | Memory Replay GANs: Learning to Generate New Categories without ForgettingabstractPrevious works on sequential learning address the problem of forgetting in discriminative models. In this paper we consider the case of generative models. In particular, we investigate generative adversarial networks (GANs) in the task of learning new categories in a sequential fashion. We first show that sequential fine tuning renders the network unable to properly generate images from previous categories (i.e. forgetting). Addressing this problem, we propose Memory Replay GANs (MeRGANs), a conditional GAN framework that integrates a memory replay generator. We study two methods to prevent forgetting by leveraging these replays, namely joint training with replay and replay alignment. Qualitative and quantitative experimental results in MNIST, SVHN and LSUN datasets show that our memory replay approach can generate competitive images while significantly mitigating the forgetting of previous categories. Chenshen Wu, Luis Herranz, Xialei Liu, Yaxing Wang, Joost van de Weijer 0001, Bogdan Raducanu |
NeurIPS | 2 |
| 2017 | Depth CNNs for RGB-D Scene Recognition: Learning from Scratch Better than Transferring from RGB-CNNsabstractScene recognition with RGB images has been extensively studied and has reached very remarkable recognition levels, thanks to convolutional neural networks (CNN) and large scene datasets. In contrast, current RGB-D scene data is much more limited, so often leverages RGB large datasets, by transferring pretrained RGB CNN models and fine-tuning with the target RGB-D dataset. However, we show that this approach has the limitation of hardly reaching bottom layers, which is key to learn modality-specific features. In contrast, we focus on the bottom layers, and propose an alternative strategy to learn depth features combining local weakly supervised training from patches followed by global fine tuning with images. This strategy is capable of learning very discriminative depth-specific features with limited depth images, without resorting to Places-CNN. In addition we propose a modified CNN architecture to further match the complexity of the model and the amount of data available. For RGB-D scene recognition, depth and RGB features are combined by projecting them in a common space and further leaning a multilayer classifier, which is jointly optimized in an end-to-end network. Our framework achieves state-of-the-art accuracy on NYU2 and SUN RGB-D in both depth only and combined RGB-D data. Xinhang Song, Luis Herranz, Shuqiang Jiang |
AAAI | 2 |
| 2017 | Domain-Adaptive Deep Network Compression
Marc Masana, Joost van de Weijer 0001, Luis Herranz, Andrew D. Bagdanov, José M. Álvarez 0004 |
ICCV | 3 |
| 2017 | Combining Models from Multiple Sources for RGB-D Scene RecognitionabstractDepth can complement RGB with useful cues about object volumes and scene layout. However, RGB-D image datasets are still too small for directly training deep convolutional neural networks (CNNs), in contrast to the massive monomodal RGB datasets. Previous works in RGB-D recognition typically combine two separate networks for RGB and depth data, pretrained with a large RGB dataset and then fine tuned to the respective target RGB and depth datasets. These approaches have several limitations: 1) only use low-level filters learned from RGB data, thus not being able to exploit properly depth-specific patterns, and 2) RGB and depth features are only combined at high-levels but rarely at lower-levels. In this paper, we propose a framework that leverages both knowledge acquired from large RGB datasets together with depth-specific cues learned from the limited depth data, obtaining more effective multi-source and multi-modal representations. We propose a multi-modal combination method that selects discriminative combinations of layers from the different source models and target modalities, capturing both high-level properties of the task and intrinsic low-level properties of both modalities. Xinhang Song, Shuqiang Jiang, Luis Herranz |
IJCAI | 3 |
| 2017 | A survey on context-aware mobile visual recognition
Weiqing Min, Shuqiang Jiang, Shuhui Wang, Ruihan Xu 0001, Yushan Cao, Luis Herranz, Zhiqiang He 0002 |
Multim. Syst. | 6 |
| 2017 | Multi-Scale Multi-Feature Context Modeling for Scene Recognition in the Semantic ManifoldabstractBefore the big data era, scene recognition was often approached with two-step inference using localized intermediate representations (objects, topics, and so on). One of such approaches is the semantic manifold (SM), in which patches and images are modeled as points in a semantic probability simplex. Patch models are learned resorting to weak supervision via image labels, which leads to the problem of scene categories co-occurring in this semantic space. Fortunately, each category has its own co-occurrence patterns that are consistent across the images in that category. Thus, discovering and modeling these patterns are critical to improve the recognition performance in this representation. Since the emergence of large data sets, such as ImageNet and Places, these approaches have been relegated in favor of the much more powerful convolutional neural networks (CNNs), which can automatically learn multi-layered representations from the data. In this paper, we address many limitations of the original SM approach and related works. We propose discriminative patch representations using neural networks and further propose a hybrid architecture in which the semantic manifold is built on top of multiscale CNNs. Both representations can be computed significantly faster than the Gaussian mixture models of the original SM. To combine multiple scales, spatial relations, and multiple features, we formulate rich context models using Markov random fields. To solve the optimization problem, we analyze global and local approaches, where a top-down hierarchical algorithm has the best performance. Experimental results show that exploiting different types of contextual relations jointly consistently improves the recognition accuracy. Xinhang Song, Shuqiang Jiang, Luis Herranz |
IEEE Trans. Image Process. | 3 |
| 2017 | Modeling Restaurant Context for Food RecognitionabstractFood photos are widely used in food logs for diet monitoring and in social networks to share social and gastronomic experiences. A large number of these images are taken in restaurants. Dish recognition in general is very challenging, due to different cuisines, cooking styles, and the intrinsic difficulty of modeling food from its visual appearance. However, contextual knowledge can be crucial to improve recognition in such scenario. In particular, geocontext has been widely exploited for outdoor landmark recognition. Similarly, we exploit knowledge about menus and location of restaurants and test images. We first adapt a framework based on discarding unlikely categories located far from the test image. Then, we reformulate the problem using a probabilistic model connecting dishes, restaurants, and locations. We apply that model in three different tasks: dish recognition, restaurant recognition, and location refinement. Experiments on six datasets show that by integrating multiple evidences (visual, location, and external knowledge) our system can boost the performance in all tasks. Luis Herranz, Shuqiang Jiang, Ruihan Xu 0001 |
IEEE Trans. Multim. | 1 |
| 2017 | Being a Supercook: Joint Food Attributes and Multimodal Content Modeling for Recipe Retrieval and ExplorationabstractThis paper considers the problem of recipe-oriented image-ingredient correlation learning with multi-attributes for recipe retrieval and exploration. Existing methods mainly focus on food visual information for recognition while we model visual information, textual content (e.g., ingredients), and attributes (e.g., cuisine and course) together to solve extended recipe-oriented problems, such as multimodal cuisine classification and attribute-enhanced food image retrieval. As a solution, we propose a multimodal multitask deep belief network ($\mathrm{M}^{3}$TDBN) to learn joint image-ingredient representation regularized by different attributes. By grouping ingredients into visible ingredients (which are visible in the food image, e.g., “chicken” and “mushroom”) and nonvisible ingredients (e.g., “salt” and “oil”),$\mathrm{M}^{3}$TDBN is capable of learning both midlevel visual representation between images and visible ingredients and nonvisual representation. Furthermore, in order to utilize different attributes to improve the intermodality correlation,$\mathrm{M}^{3}$TDBN incorporates multitask learning to make different attributes collaborate each other. Based on the proposed$\mathrm{M}^{3}$TDBN, we exploit the derived deep features and the discovered correlations for three extended novel applications: 1) multimodal cuisine classification; 2) attribute-augmented cross-modal recipe image retrieval; and 3) ingredient and attribute inference from food images. The proposed approach is evaluated on the constructed Yummly dataset and the evaluation results have validated the effectiveness of the proposed approach. Weiqing Min, Shuqiang Jiang, Huayang Wang, Xinda Liu, Luis Herranz |
IEEE Trans. Multim. | 6 |
| 2016 | Scene Recognition with CNNs: Objects, Scales and Dataset BiasabstractSince scenes are composed in part of objects, accurate recognition of scenes requires knowledge about both scenes and objects. In this paper we address two related problems: 1) scale induced dataset bias in multi-scale convolutional neural network (CNN) architectures, and 2) how to combine effectively scene-centric and object-centric knowledge (i.e. Places and ImageNet) in CNNs. An earlier attempt, Hybrid-CNN[23], showed that incorporating ImageNet did not help much. Here we propose an alternative method taking the scale into account, resulting in significant recognition gains. By analyzing the response of ImageNet-CNNs and Places-CNNs at different scales we find that both operate in different scale ranges, so using the same network for all the scales induces dataset bias resulting in limited performance. Thus, adapting the feature extractor to each particular scale (i.e. scale-specific CNNs) is crucial to improve recognition, since the objects in the scenes have their specific range of scales. Experimental results show that the recognition accuracy highly depends on the scale, and that simple yet carefully chosen multi-scale combinations of ImageNet-CNNs and Places-CNNs, can push the stateof-the-art recognition accuracy in SUN397 up to 66.26% (and even 70.17% with deeper architectures, comparable to human performance). Luis Herranz, Shuqiang Jiang, Xiangyang Li 0002 |
CVPR | 1 |
| 2016 | Image Captioning with both Object and Scene InformationabstractRecently, automatic generation of image captions has attracted great interest not only because of its extensive applications but also because it connects computer vision and natural language processing. By combining convolutional neural networks (CNNs), which learn visual representations from images, and recurrent neural networks (RNNs), which translate the learned features into text sequences, the content of a image can be transformed into linguistic sequences. Existing approaches typically focus on visual features extracted form an object-oriented CNN (train on ImageNet) and then decode them into natural language. In this paper, we propose a novel model using not only object-related, but also scene-related information extracted from the images. To make full use of both object and scene information, we first combine object information and scene information (extracted from a scene-oriented CNN), and then using as inputs to RNNs. Both types of information provide complementary aspects that help in generating a more complete description of the image. Qualitative and quantitative evaluation results validate the effectiveness of our method. Xiangyang Li 0002, Xinhang Song, Luis Herranz, Shuqiang Jiang |
ACM Multimedia | 3 |
| 2016 | Guest Editorial: Image Analysis and Processing Leveraging Additional Information
Luis Herranz, Jian Cheng 0001, Yue Gao 0002, Shuqiang Jiang |
Multim. Tools Appl. | 1 |
| 2016 | Scalable storyboards in handheld devices: applications and evaluation metrics
Luis Herranz, Shuqiang Jiang |
Multim. Tools Appl. | 1 |
| 2016 | Category co-occurrence modeling for large scale scene recognition
Xinhang Song, Shuqiang Jiang, Luis Herranz, Yan Kong, Kai Zheng 0001 |
Pattern Recognit. | 3 |
| 2015 | Joint multi-feature spatial context for scene recognition in the semantic manifoldabstractIn the semantic multinomial framework patches and images are modeled as points in a semantic probability simplex. Patch theme models are learned resorting to weak supervision via image labels, which leads the problem of scene categories co-occurring in this semantic space. Fortunately, each category has its own co-occurrence patterns that are consistent across the images in that category. Thus, discovering and modeling these patterns is critical to improve the recognition performance in this representation. In this paper, we observe that not only global co-occurrences at the image-level are important, but also different regions have different category co-occurrence patterns. We exploit local contextual relations to address the problem of discovering consistent co-occurrence patterns and removing noisy ones. Our hypothesis is that a less noisy semantic representation, would greatly help the classifier to model consistent co-occurrences and discriminate better between scene categories. An important advantage of modeling features in a semantic space is that this space is feature independent. Thus, we can combine multiple features and spatial neighbors in the same common space, and formulate the problem as minimizing a context-dependent energy. Experimental results show that exploiting different types of contextual relations consistently improves the recognition accuracy. In particular, larger datasets benefit more from the proposed method, leading to very competitive performance. Xinhang Song, Shuqiang Jiang, Luis Herranz |
CVPR | 3 |
| 2015 | A probabilistic model for food image recognition in restaurantsabstractA large amount of food photos are taken in restaurants for diverse reasons. This dish recognition problem is very challenging, due to different cuisines, cooking styles and the intrinsic difficulty of modeling food from its visual appearance. Contextual knowledge is crucial to improve recognition in such scenario. In particular, geocontext has been widely exploited for outdoor landmark recognition. Similarly, we exploit knowledge about menus and geolocation of restaurants and test images. We first adapt a framework based on discarding unlikely categories located far from the test image. Then we reformulate the problem using a probabilistic model connecting dishes, restaurants and geolocations. We apply that model in three different tasks: dish recognition, restaurant recognition and geolocation refinement. Experiments on a dataset including 187 restaurants and 701 dishes show that combining multiple evidences (visual, geolocation, and external knowledge) can boost the performance in all tasks. Luis Herranz, Ruihan Xu 0001, Shuqiang Jiang |
ICME | 1 |
| 2015 | Hand-Object Sense: A Hand-held Object Recognition System Based on RGB-D InformationabstractHand-held objects play an important role in human-human and human-machine interaction. It can be used as a reference for understanding user intentions or user requirements. In this technical demonstration, we introduce an object recognition system called Hand-Object Sense that can automatically recognize the object held by user. This system first detects and segments the hand-held object by exploiting skeleton information combined with depth information. Second, in the object recognition stage, this system exploits features computed in different ways and fuses them to improve the recognition accuracy. Our system can recognize objects in real-time and have a good tolerance to angle and scale transformation. Furthermore, it has a good generalization capability for unknown objects. Xiong Lv, Shuqiang Jiang, Luis Herranz |
ACM Multimedia | 3 |
| 2015 | RGB-D Hand-Held Object Recognition Based on Heterogeneous Feature Fusion
Xiong Lv, Shuqiang Jiang, Luis Herranz |
J. Comput. Sci. Technol. | 3 |
| 2015 | Geolocalized Modeling for Dish RecognitionabstractFood-related photos have become increasingly popular , due to social networks, food recommendations, and dietary assessment systems. Reliable annotation is essential in those systems, but unconstrained automatic food recognition is still not accurate enough. Most works focus on exploiting only the visual content while ignoring the context. To address this limitation, in this paper we explore leveraging geolocation and external information about restaurants to simplify the classification problem. We propose a framework incorporating discriminative classification in geolocalized settings and introduce the concept of geolocalized models, which, in our scenario, are trained locally at each restaurant location. In particular, we propose two strategies to implement this framework: geolocalized voting and combinations of bundled classifiers. Both models show promising performance, and the later is particularly efficient and scalable. We collected a restaurant-oriented food dataset with food images, dish tags, and restaurant-level information, such as the menu and geolocation. Experiments on this dataset show that exploiting geolocation improves around 30% the recognition performance, and geolocalized models contribute with an additional 3-8% absolute gain, while they can be trained up to five times faster. Ruihan Xu 0001, Luis Herranz, Shuqiang Jiang, Xinhang Song, Ramesh Jain 0001 |
IEEE Trans. Multim. | 2 |
| 2014 | Accuracy and Specificity Trade-off in k -nearest Neighbors Classification
Luis Herranz, Shuqiang Jiang |
ACCV (2) | 1 |
| 2013 | Flexible navigation in smartphones and tablets using scalable storyboardsabstractIn this demo paper we present a multiscale browsing interface for handheld devices, in which the user can interactively change the scale of the storyboard to easily adjust the amount of information desired. Conventional and hierarchical storyboards provide one or very few possible lengths. In contrast, scalable storyboards allow the number of images and the storyboard itself to be adapted to the device constraints (e.g. aspect ratio, resolution) and navigation state with much finer granularity. Several levels and modes, including segment of interest, are provided for more intuitive and convenient navigation. Shuai Zheng 0004, Luis Herranz, Shuqiang Jiang |
ICMR | 2 |
| 2013 | Combining MPEG Tools to Generate Video Summaries Adapted to the Terminal and NetworkabstractMoving picture experts group (MPEG) standards provide tools for a broad range of purposes, covering from coding to metadata description tools. In this paper, the combined use of tools from different MPEG standards is described in the context of a video summarization application. The main objective of the framework is the efficient generation of summaries, integrated with their adaptation to the user's terminal and network. The MPEG-4 Scalable Video Coding specification is used for fast adaptation and summary bitstream generation. MPEG-21 digital item adaptation tools are used to describe metadata related to the user's terminal and network. MPEG-7 tools are used to describe the summary. Finally, the framework is compared with alternative approaches (variations and transcoding), in terms of efficiency, rate-distortion performance and other aspects. Luis Herranz, José María Martínez Sanchez |
Comput. J. | 1 |
| 2012 | Scalable Comic-Like Video Summaries and Layout DisturbanceabstractThis paper describes an efficient system for scalable video summarization that exploits comic-like summaries and multi-scale representations to facilitate interactivity and balance between content coverage and compactness. Due to the layout disturbance induced by the transitions between scales, a new heuristic algorithm is proposed to restrict changes to bounded summary segments. Conducted user evaluations show that the proposed methodology improves usability while keeping the summaries compact and informative. Luis Herranz, Janko Calic, José María Martínez Sanchez, Marta Mrak |
IEEE Trans. Multim. | 1 |
| 2010 | On the Advantages of the Use of Bitstream Extraction for Video Summary Generation
Luis Herranz, José María Martínez Sanchez |
MMM | 1 |
| 2010 | A Framework for Scalable Summarization of VideoabstractVideo summaries provide compact representations of video sequences, with the length of the summary playing an important role, trading off the amount of information conveyed and how fast it can be visualized. This letter proposes scalable summarization as a method to easily adapt the summary to a suitable length, according to the requirements in each case, along with a suitable framework. The analysis algorithm uses a novel iterative ranking procedure in which each summary is the result of the extension of the previous one, balancing information coverage and visual pleasantness. The result of the algorithm is a ranked list, a scalable representation of the sequence useful for summarization. The summary is then efficiently generated from the bitstream of the sequence using bitstream extraction. Luis Herranz, José María Martínez Sanchez |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2009 | An efficient summarization algorithm based on clustering and bitstream extractionabstractVisualizing video is a very time consuming task. Video summaries are compact representations and very useful in video management systems. However, in some applications, in which the summary must be generated on demand, it usually implies a high delay for the user. Video summarization is a very demanding task in terms of processing resources, and it is very difficult to obtain a summary of a large video with low delay. In this paper a flexible summarization framework is presented which can create storyboards and video skims using bitstream extraction to generate the summary after an appropriate analysis algorithm. This technique is very efficient and achieves very low delay while keeping a reasonable quality of video summaries. Luis Herranz, José María Martínez Sanchez |
ICME | 1 |
| 2009 | An integrated approach to summarization and adaptation using H.264/MPEG-4 SVC
Luis Herranz, José María Martínez Sanchez |
Signal Process. Image Commun. | 1 |
| 2009 | On the use of hierarchical prediction structures for efficient summary generation of H.264/AVC bitstreams
Luis Herranz, José María Martínez Sanchez |
Signal Process. Image Commun. | 1 |
| 2008 | Generation of scalable summaries based on iterative GoP rankingabstractVideo skims and image storyboards are two widely used abstractions for representing the essence of a video sequence, crucial for effective browsing and retrieval applications. In this paper we propose a flexible approach to find a reasonable balance between the semantic coverage and naturalness of the generated summaries, targeting a wide range of summarization ratios. The result of the algorithm is a scalable representation of the information required for summarization (a ranked list of GoPs) with a number of advantages in terms of efficient generation and potential applications. Luis Herranz, José María Martínez Sanchez |
ICIP | 1 |
| 2007 | Integrating semantic analysis and scalable video coding for efficient content-based adaptation
Luis Herranz |
Multim. Syst. | 1 |
| 2007 | Content-driven adaptation of on-line video
Jesús Bescós, José María Martínez Sanchez, Luis Herranz, Fabricio Tiburzi |
Signal Process. Image Commun. | 3 |
| 2006 | Video summaries generation and access via personalized delivery of multimedia presentations adapted to service and terminalabstractThis article is centered on describing the provision of universal multimedia access services for video summaries via multimedia presentations that allow the integration of multimedia messaging service (MMS)-enabled terminals in the framework of the deferred time environment (DTE) of the DYMAS system. The system uses the framework of MPEG-7 and MPEG-21 to provide description metadata of the multimedia content and the usage context (including terminal, network capabilities, and user preferences), respectively. These descriptions are the base for the main functionalities of the complete system that provides personalized access to content (filtering by user preferences or via querying) that is first adapted to a multimedia presentation (generating a video summary that is represented via keyframes with synchronized audio clips) by the Presentation Builder and afterward adapted to the current service and terminal by the Adaptation Engine. © 2006 Wiley Periodicals, Inc. Int J Int Syst 21: 785–800, 2006. Marta Padilla, José María Martínez Sanchez, Luis Herranz |
Int. J. Intell. Syst. | 3 |