EDBT 2026 Demo / reviewers in the wild / expert
Fei Yang 0004
dblp:19/2504-4
· DBLP profile ↗
30ranked-venue papers
3as first author
28since 2021 · last 2026
0000-0003-4099-6511ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 19 · 2 first-author · 17 since 2021Artificial intelligence and machine learning · 16 · 2 first-author · 16 since 2021Systems, architecture and hardware · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Bridging Inter-Task Gap of Continual Self-Supervised Learning With External DataabstractRecent research on Self-Supervised Learning (SSL) has demonstrated its ability to extract high-quality representations from unlabeled samples. However, in continual learning scenarios where training data arrives sequentially, SSL’s performance tends to deteriorate. This study focuses on Continual Contrastive Self-Supervised Learning (CCSSL) and highlights that the absence of inter-task contrastive learning, due to the unavailability of historical samples, leads to a significant drop in performance. To tackle this issue, we introduce a simple and effective method called BGE, which Bridges the inter-task Gap of CCSSL using External data from publicly available datasets. BGE enables the contrastive learning of each task data with external data, allowing relationships between them to be passed along the tasks, thereby facilitatingimplicitinter-task data comparisons. To overcome the limitation of the external data selection and maintain its effectiveness, we further propose the One-Propose-One algorithm to collect more relevant and diverse high-quality samples from external sources while filtering out distractions from the out-of-distribution data. Experiments show that BGE can generate better discriminative representation in CCSSL, especially for inter-task data, and improve classification results with various external data compositions. Additionally, BGE can be seamlessly integrated into existing continual learning methods, yielding significant performance improvement. Haori Lu, Linlan Huang, Enguang Wang, Fei Yang 0004, Xialei Liu |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2026 | Text and Non-Text Latent Feature Disentanglement for Screen Content Image CompressionabstractWith the growing prevalence of screen content images in multimedia communication, efficient compression has become increasingly crucial. Unlike natural scene images, screen content typically contains rich text regions that exhibit unique characteristics and low correlation with surrounding non-text elements. The intricate mixture of text and non-text within images poses significant challenges for existing learned compression networks, as the text and non-text features are severely entangled in the latent domain along the channel dimension, leading to compromised reconstruction quality and suboptimal entropy estimation. In this paper, we propose a novel Disentangled Image Compression Architecture (DICA) that enhances the analysis module and the entropy model of existing compression architectures to address these limitations. First, we introduce a Disentangled Analysis Module (DAM) by augmenting original analysis modules with an additional text approximation branch and a disentangling network. They work in concert to disentangle latent features into text and non-text classes along the channel dimension, resulting in a more structured feature distribution that better aligns with compression requirements. Second, we propose a Disentangled Channel-Conditional Entropy Model (DCEM) that efficiently leverages the feature distribution bias introduced by DAM, thereby further improving compression performance. Experimental results demonstrate that the proposed DICA, along with DAM and DCEM can be integrated into various channel-conditional compression backbones, significantly improving their performance in screen content compression—particularly in hard-to-compress text regions. When integrated with an advanced WACNN backbone, our method achieves a 13% overall BD-Rate gain and a 16% BD-Rate gain in text regions on the SIQAD dataset. Hao Wang 0184, Junyan Huo, Fei Yang 0004, Shuai Wan, Gaoxing Chen, Luis Herranz, Fuzheng Yang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2026 | Sharpness-Aware Dynamic Anchor Selection for Generalized Category DiscoveryabstractGeneralized category discovery (GCD) is an important and challenging task in open-world learning. Specifically, given some labeled data of known classes, GCD aims to cluster unlabeled data that contain both known and unknown classes. Current GCD methods based on parametric classification adopt the DINO-like pseudo-labeling strategy, where the sharpened probability output of one view is used as supervision information for the other view. However, large pre-trained models have a preference for some specific visual patterns, resulting in encoding spurious correlation for unlabeled data and generating noisy pseudo-labels. To address this issue, we propose a novel method, which contains two modules: Loss Sharpness Penalty (LSP) and Dynamic Anchor Selection (DAS). LSP enhances the robustness of model parameters to small perturbations by minimizing the worst-case loss sharpness of the model, which suppressing the encoding of trivial features, thereby reducing overfitting of noise samples and improving the quality of pseudo-labels. Meanwhile, DAS selects representative samples for the unknown classes based on KNN density and class probability during the model training and assigns hard pseudo-labels to them, which not only alleviates the confidence difference between known and unknown classes but also enables the model to quickly learn more accurate feature distribution for the unknown classes, thus further improving the clustering accuracy. Extensive experiments demonstrate that the proposed method can effectively mitigate the noise of pseudo-labels, and achieve state-of-the-art results on multiple GCD benchmarks. Zhimao Peng, Enguang Wang, Fei Yang 0004, Xialei Liu, Ming-Ming Cheng |
IEEE Trans. Multim. | 3 |
| 2026 | Camera Motion-Conditioned Motion Estimation for Neural Video Coding for Cloud GamingabstractThe rising popularity of cloud gaming highlights the importance of effective video compression for this domain. Despite the potential of neural video codecs to surpass traditional codecs, their application to cloud gaming videos generally achieves limited performance primarily due to two reasons: (1) most advanced neural video codecs are designed for natural videos with only a few optimized for cloud gaming content, and (2) the distinctive characteristics of cloud gaming videos including abrupt and large camera movements, coupled with repetitive textures, make the direct application of existing codecs suboptimal. To this end, this paper introduces an enhanced neural video codec for cloud gaming with camera motion-conditioned motion estimation. By leveraging the unique feature of cloud gaming, we effectively utilize the camera motion information to condition motion estimation with an attention mechanism, overcoming the pronounced challenges of motion estimation and facilitating the accurate learning of optical flow. Furthermore, we optimize the loss function of the motion estimation network by employing a comprehensive loss across multiple feature levels, ensuring the learned optical flow is multi-faceted and well-suited to subsequent motion compensation. Extensive experiments demonstrate the effectiveness of the proposed codec, providing improved motion estimation capabilities and superior rate-distortion performance. Moreover, the robustness of our codec to rapid camera movements is validated, which makes it highly suitable for cloud gaming scenarios. Fei Yang 0004, Luka Murn, Juil Sock, Marc Gorriz, Shuai Wan, Wei Zhang 0072, Fuzheng Yang 0001, Luis Herranz |
IEEE Trans. Multim. | 2 |
| 2025 | KAC: Kolmogorov-Arnold Classifier for Continual LearningabstractContinual learning requires models to train continuously across consecutive tasks without forgetting. Most existing methods utilize linear classifiers, which struggle to maintain a stable classification space while learning new tasks. Inspired by the success of Kolmogorov-Arnold Networks (KAN) in preserving learning stability during simple continual regression tasks, we set out to explore their potential in more complex continual learning scenarios. In this paper, we introduce the Kolmogorov-Arnold Classifier (KAC), a novel classifier developed for continual learning based on the KAN structure. We delve into the impact of KAN’s spline functions and introduce Radial Basis Functions (RBF) for improved compatibility with continual learning. We replace linear classifiers with KAC in several recent approaches and conduct experiments across various continual learning benchmarks, all of which demonstrate performance improvements, highlighting the effectiveness and robustness of KAC in continual learning. The code is available at https://github.com/Ethanhuhuhu/KAC. Yusong Hu, Zichen Liang, Fei Yang 0004, Qibin Hou, Xialei Liu, Ming-Ming Cheng |
CVPR | 3 |
| 2025 | GET: Unlocking the Multi-modal Potential of CLIP for Generalized Category DiscoveryabstractGiven unlabelled datasets containing both old and new categories, generalized category discovery (GCD) aims to accurately discover new classes while correctly classifying old classes. Current GCD methods only use a single visual modality of information, resulting in a poor classification of visually similar classes. As a different modality, text information can provide complementary discriminative information, which motivates us to introduce it into the GCD task. However, the lack of class names for unlabelled data makes it impractical to utilize text information. To tackle this challenging problem, in this paper, we propose a Text Embedding Synthesizer (TES) to generate pseudo text embeddings for unlabelled samples. Specifically, our TES leverages the property that CLIP can generate aligned vision-language features, converting visual embeddings into tokens of the CLIP’s text encoder to generate pseudo text embeddings. Besides, we employ a dual-branch framework, through the joint learning and instance consistency of different modality branches, visual and semantic information mutually enhance each other, promoting the interaction and fusion of visual and text knowledge. Our method unlocks the multi-modal potentials of CLIP and outperforms the baseline methods by a large margin on all GCD benchmarks, achieving new state-of-the-art. Our code is available at: https://github.com/enguangW/GET. Enguang Wang, Zhimao Peng, Zhengyuan Xie, Fei Yang 0004, Xialei Liu, Ming-Ming Cheng |
CVPR | 4 |
| 2025 | Mind the Gap: Preserving and Compensating for the Modality Gap in CLIP-Based Continual Learning
Linlan Huang, Haori Lu, Yifan Meng, Fei Yang 0004, Xialei Liu |
ICCV | 5 |
| 2025 | Improving Continual Learning Performance and Efficiency with Auxiliary ClassifiersabstractContinual learning is crucial for applying machine learning in challenging, dynamic, and often resource-constrained environments. However, catastrophic forgetting — overwriting previously learned knowledge when new information is acquired — remains a major challenge. In this work, we examine the intermediate representations in neural network layers during continual learning and find that such representations are less prone to forgetting, highlighting their potential to accelerate computation. Motivated by these findings, we propose to use auxiliary classifiers (ACs) to enhance performance and demonstrate that integrating ACs into various continual learning methods consistently improves accuracy across diverse evaluation settings, yielding an average 10% relative gain. We also leverage the ACs to reduce the average cost of the inference by 10-60% without compromising accuracy, enabling the model to return the predictions before computing all the layers. Our approach provides a scalable and efficient solution for continual learning. Filip Szatkowski, Yaoyue Zheng, Fei Yang 0004, Tomasz Trzcinski, Bartlomiej Twardowski, Joost van de Weijer 0001 |
ICML | 3 |
| 2025 | Knowledge Graph Enhanced Generative Multi-modal Models for Class-Incremental LearningabstractContinual learning in computer vision faces the critical challenge of catastrophic forgetting, where models struggle to retain prior knowledge while adapting to new tasks.
Although recent studies have attempted to leverage the generalization capabilities of pre-trained models to mitigate overfitting on current tasks, models still tend to forget details of previously learned categories as tasks progress, leading to misclassification. To address these limitations, we introduce a novel Knowledge Graph Enhanced Generative Multi-modal model (KG-GMM) that builds an evolving knowledge graph throughout the learning process. Our approach utilizes relationships within the knowledge graph to augment the class labels and assigns different relations to similar categories to enhance model differentiation. During testing, we propose a Knowledge Graph Augmented Inference method that locates specific categories by analyzing relationships within the generated text, thereby reducing the loss of detailed information about old classes when learning new knowledge and alleviating forgetting. Experiments demonstrate that our method effectively leverages relational information to help the model correct mispredictions, achieving state-of-the-art results in both conventional CIL and few-shot CIL settings, confirming the efficacy of knowledge graphs at preserving knowledge in the continual learning scenarios. Haori Lu, Linlan Huang, Fei Yang 0004, Xialei Liu, Ming-Ming Cheng |
NeurIPS | 4 |
| 2025 | Multi-Class Textual-Inversion Secretly Yields a Semantic-Agnostic ClassifierabstractWith the advent of large pre-trained vision-language models such as CLIP, prompt learning methods aim to enhance the transferability of the CLIP model. They learn the prompt given few samples from the downstream task given the specific class names as prior knowledge, which we term as semantic-aware classification. However, in many realistic scenarios, we only have access to few samples and no knowledge of the class names (e.g., when considering instances of classes). This challenging scenario represents the semantic-agnostic discriminative case. Text-to-Image (T2I) personalization methods aim to adapt T2I models to unseen concepts by learning new tokens and endowing these tokens with the capability of generating the learned concepts. These methods do not require knowledge of class names as a semantic-aware prior. Therefore, in this paper, we first explore Textual Inversion and reveal that the new concept tokens possess both generation and classification capabilities by regarding each category as a single concept. However, learning classifiers from single-concept textual inversion is limited since the learned tokens are sub-optimal for the discriminative tasks. To mitigate this issue, we propose Multi-Class textual inversion, which includes a discriminative regularization term for the token updating process. Using this technique, our method MC-TI achieves stronger Semantic-Agnostic Classification while preserving the generation capability of these modifier tokens given only few samples per category. In the experiments, we extensively evaluate MC-TI on 12 datasets covering various scenarios, which demonstrates that MC-TI achieves superior results in terms of both classification and generation outcomes. Kai Wang 0060, Fei Yang 0004, Bogdan Raducanu, Joost van de Weijer 0001 |
WACV | 2 |
| 2025 | Enhanced neural video compression for cloud gaming videos with aligned frame generationabstractThe burgeoning popularity of cloud gaming makes it critical for efficient video compression to relieve the growing bandwidth pressure. While existing neural video coding approaches have demonstrated strong compression potential on natural videos, there is an absence of efficient neural codecs dedicated to gaming videos. To bridge this gap, in this paper, we propose an end-to-end neural video compression method designed specifically for cloud gaming videos. By effectively utilizing the unique camera motion information inherent to cloud gaming, the previous reconstructed frame is maximally aligned to the current frame through a learningbased module with multiple losses, which then replaces the previous reconstructed frame for optical flow estimation. By significantly reducing the displacement between two consecutive frames caused by camera motion, the motion estimation accuracy is enhanced, effectively handling the large and abrupt motion scenarios frequently present in gaming videos. Furthermore, the aligned tensor obtained in the previous step is used to enhance the latent prior of the entropy model, providing a superior temporal prior for coding. Extensive experimental results demonstrate the superior performance of our proposed method compared to one of the previous state-of-the-art approaches, DCVC-HEM, providing significant progress in end-to-end neural compression in cloud gaming videos Fei Yang 0004, Luka Murn, Juil Sock, Marc Gorriz, Shuai Wan, Wei Zhang 0072, Fuzheng Yang 0001, Luis Herranz |
Expert Syst. Appl. | 2 |
| 2025 | Exemplar-Free Continual Learning of Vision Transformers via Gated Class-Attention and Cascaded Feature Drift CompensationabstractAbstract Vision transformers (ViTs) have achieved remarkable successes across a broad range of computer vision applications. As a consequence, there has been increasing interest in extending continual learning theory and techniques to ViT architectures. We propose a new method for exemplar-free class incremental training of ViTs. The main challenge of exemplar-free continual learning is maintaining plasticity of the learner without causing catastrophic forgetting of previously learned tasks. This is often achieved via exemplar replay which can help recalibrate previous task classifiers to the feature drift which occurs when learning new tasks. Exemplar replay, however, comes at the cost of retaining samples from previous tasks which for many applications may not be possible. To address the problem of continual ViT training, we first propose gated class-attention to minimize the drift in the final ViT transformer block. This mask-based gating is applied to class-attention mechanism of the last transformer block and strongly regulates the weights crucial for previous tasks. Importantly, gated class-attention does not require the task-ID during inference, which distinguishes it from other parameter isolation methods. Secondly, we propose a new method of feature drift compensation that accommodates feature drift in the backbone when learning new tasks. The combination of gated class-attention and cascaded feature drift compensation allows for plasticity towards new tasks while limiting forgetting of previous ones. Extensive experiments performed on CIFAR-100, Tiny-ImageNet and ImageNet100 demonstrate that our exemplar-free method obtains competitive results when compared to rehearsal based ViT methods.(Code: https://github.com/OcraM17/GCAB-CFDC ) Marco Cotogni, Fei Yang 0004, Claudio Cusano, Andrew D. Bagdanov, Joost van de Weijer 0001 |
Int. J. Comput. Vis. | 2 |
| 2025 | Curriculum learning-based slimmable cross-component prediction for video coding
Chengyi Zou, Shuai Wan, Marc Gorriz, Luka Murn, Juil Sock, Fei Yang 0004, Luis Herranz |
Neurocomputing | 6 |
| 2025 | Implicit neural representation model for camera relocalization in multiple scenes
Yongmei Cheng, Fei Yang 0004, Mikhail G. Mozerov |
Pattern Recognit. | 3 |
| 2025 | Enhancing Continual Semantic Segmentation via Uncertainty and Class Balance Re-WeightingabstractContinual Semantic Segmentation (CSS) primarily aims to continually learn new semantic segmentation categories while avoiding catastrophic forgetting. In semantic segmentation tasks, images can comprise both familiar old categories and novel unseen categories and they are treated as background in the incremental stage. Therefore, it is necessary to utilize the old model to generate pseudo-labels. However, the quality of these pseudo-labels significantly influences the model's forgetting of the old categories. Erroneous pseudo-labels can introduce harmful gradients, thus exacerbating model forgetting. In addition, the issue of class imbalance poses a significant challenge within the realm of CSS. Although traditional methods frequently diminish the emphasis placed on new classes to address this imbalance, we discover that the imbalance extends beyond the distinction between old and new classes. In this paper, we specifically address two previously overlooked problems in CSS: the impact of erroneous pseudo-labels on model forgetting and the confusion induced by class imbalance. We propose an Uncertainty and Class Balance Re-weighting approach (UCB) that assigns higher weights to pixels with pseudo-labels exhibiting lower uncertainty and to categories with smaller proportions during the training process. Our proposed approach enhances the impact of essential pixels during the continual learning process, thereby reducing model forgetting and dynamically balancing category weights based on the dataset. Our method is simple yet effective and can be applied to any method that uses pseudo-labels. Extensive experiments on the Pascal-VOC and ADE20K datasets demonstrate the efficacy of our approach in improving model performance across three state-of-the-art methods. The code will be available at https://github.com/JACK-Chen-2019/UCB. Zichen Liang, Yusong Hu, Fei Yang 0004, Xialei Liu |
IEEE Trans. Image Process. | 3 |
| 2025 | Lightweight Deep Exemplar Colorization via Semantic Attention-Guided Laplacian PyramidabstractExemplar-based colorization aims to generate plausible colors for a grayscale image with the guidance of a color reference image. The main challenging problem is finding the correct semantic correspondence between the target image and the reference image. However, the colors of the object and background are often confused in the existing methods. Besides, these methods usually use simple encoder-decoder architectures or pyramid structures to extract features and lack appropriate fusion mechanisms, which results in the loss of high-frequency information or high complexity. To address these problems, this article proposes a lightweight semantic attention-guided Laplacian pyramid network (SAGLP-Net) for deep exemplar-based colorization, exploiting the inherent multi-scale properties of color representations. They are exploited through a Laplacian pyramid, and semantic information is introduced as high-level guidance to align the object and background information. Specially, a semantic guided non-local attention fusion module is designed to exploit the long-range dependency and fuse the local and global features. Moreover, a Laplacian pyramid fusion module based on criss-cross attention is proposed to fuse high frequency components in the large-scale domain. An unsupervised multi-scale multi-loss training strategy is further introduced for network training, which combines pixel loss, color histogram loss, total variance regularisation, and adversarial loss. Experimental results demonstrate that our colorization method achieves better subjective and objective performance with lower complexity than the state-of-the-art methods. Chengyi Zou, Shuai Wan, Marc Gorriz, Luka Murn, Marta Mrak, Juil Sock, Fei Yang 0004, Luis Herranz |
IEEE Trans. Vis. Comput. Graph. | 7 |
| 2024 | GRIF-DM: Generation of Rich Impression Fonts Using Diffusion ModelsabstractFonts are integral to creative endeavors, design processes, and artistic productions. The appropriate selection of a font can significantly enhance artwork and endow advertisements with a higher level of expressivity. Despite the availability of numerous diverse font designs online, traditional retrieval-based methods for font selection are increasingly being supplanted by generation-based approaches. These newer methods offer enhanced flexibility, catering to specific user preferences and capturing unique stylistic impressions. However, current impression font techniques based on Generative Adversarial Networks (GANs) necessitate the utilization of multiple auxiliary losses to provide guidance during generation. Furthermore, these methods commonly employ weighted summation for the fusion of impression-related keywords. This leads to generic vectors with the addition of more impression keywords, ultimately lacking in detail generation capacity. In this paper, we introduce a diffusion-based method, termed GRIF-DM, to generate fonts that vividly embody specific impressions, utilizing an input consisting of a single letter and a set of descriptive impression keywords. The core innovation of GRIF-DM lies in the development of dual cross-attention modules, which process the characteristics of the letters and impression keywords independently but synergistically, ensuring effective integration of both types of information. Our experimental results, conducted on the MyFonts dataset, affirm that this method is capable of producing realistic, vibrant, and high-fidelity fonts that are closely aligned with user specifications. This confirms the potential of our approach to revolutionize font generation by accommodating a broad spectrum of user-driven design requirements. Our code is publicly available at https://github.com/leitro/GRIF-DM. Lei Kang 0002, Fei Yang 0004, Kai Wang 0060, Mohamed Ali Souibgui, Lluís Gómez i Bigorda, Alicia Fornés, Ernest Valveny, Dimosthenis Karatzas |
ECAI | 2 |
| 2024 | Machine Unlearning for Document Classification
Lei Kang 0002, Mohamed Ali Souibgui, Fei Yang 0004, Lluís Gómez i Bigorda, Ernest Valveny, Dimosthenis Karatzas |
ICDAR (4) | 3 |
| 2024 | Rate Control for Slimmable Video Codec using Multilayer PerceptronabstractAs a flexible design for practical end-to-end video compression, slimmable video coding achieves variable rate coding with an adaptable complexity. In this paper, rate control using a multilayer perceptron is proposed to achieve accurate bitrate control for slimmable video coding. An average bitrate error of less than 0.3% is achieved on tested sequences at a slight decrease in the rate-distortion performance. Furthermore, the results show that there is no overflow or underflow in the buffer occupancy during the rate control process. Defa Wang, Shuai Wan, Fei Yang 0004, Luis Herranz |
ISCAS | 4 |
| 2024 | A slimmable framework for practical neural video compressionabstractDeep learning is being increasingly applied to image and video compression in a new paradigm known as neural video compression. While achieving impressive rate–distortion (RD) performance, neural video codecs (NVC) require heavy neural networks, which in turn have large memory and computational costs and often lack important functionalities such as variable rate. These are significant limitations to their practical application. Addressing these problems, recent slimmable image codecs can dynamically adjust their model capacity to elegantly reduce the memory and computation requirements, without harming RD performance. However, the extension to video is not straightforward due to the non-trivial interplay with complex motion estimation and compensation modules in most NVC architectures. In this paper we propose the slimmable video codec framework (SlimVC) that integrates an slimmable autoencoder and a motion-free conditional entropy model. We show that the slimming mechanism is also applicable to the more complex case of video architectures, providing SlimVC with simultaneous control of the computational cost, memory and rate, which are all important requirements in practice. We further provide detailed experimental analysis, and describe application scenarios that can benefit from slimmable video codecs. Zhaocheng Liu, Fei Yang 0004, Defa Wang, Marc Gorriz, Luka Murn, Shuai Wan, Saiping Zhang, Marta Mrak, Luis Herranz |
Neurocomputing | 2 |
| 2024 | Task-Switchable Pre-Processor for Image Compression for Multiple Machine Vision TasksabstractVisual content is increasingly being processed by machines for various automated content analysis tasks instead of being consumed by humans. Despite the existence of several compression methods tailored for machine tasks, few consider real-world scenarios with multiple tasks. In this paper, we aim to address this gap by proposing a task-switchable pre-processor that optimizes input images specifically for machine consumption prior to encoding by an off-the-shelf codec designed for human consumption. The proposed task-switchable pre-processor adeptly maintains relevant semantic information based on the specific characteristics of different downstream tasks, while effectively suppressing irrelevant information to reduce bitrate. To enhance the processing of semantic information for diverse tasks, we leverage pre-extracted semantic features to modulate the pixel-to-pixel mapping within the pre-processor. By switching between different modulations, multiple tasks can be seamlessly incorporated into the system. Extensive experiments demonstrate the practicality and simplicity of our approach. It significantly reduces the number of parameters required for handling multiple tasks while still delivering impressive performance. Our method showcases the potential to achieve efficient and effective compression for machine vision tasks, supporting the evolving demands of real-world applications. Mingyi Yang, Fei Yang 0004, Luka Murn, Marc Gorriz, Juil Sock, Shuai Wan, Fuzheng Yang 0001, Luis Herranz |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2023 | Efficient Super-Resolution for Compression Of Gaming VideosabstractDue to the increasing demand for game-streaming services, efficient compression of computer-generated video is more critical than ever, especially when the available bandwidth is low. This paper proposes a super-resolution framework that improves the coding efficiency of computer-generated gaming videos at low bitrates. Most state-of-the-art super-resolution networks generalize over a variety of RGB inputs and use a unified network architecture for frames of different levels of degradation, leading to high complexity and redundancy. Since games usually consist of a limited number of fixed scenarios, we specialize one model for each scenario and assign appropriate network capacities for different QPs to perform super-resolution under the guidance of reconstructed high-quality luma components. Experimental results show that our framework achieves a superior quality-complexity trade-off compared to the ESRnet baseline, saving at most 93.59% parameters while maintaining comparable performance. The compression efficiency compared to HEVC is also improved by more than 17% BD-rate gain. Luka Murn, Luis Herranz, Fei Yang 0004, Marta Mrak, Wei Zhang 0072, Shuai Wan, Marc Gorriz |
ICASSP | 4 |
| 2023 | Semantic Preprocessor for Image Compression for MachinesabstractVisual content is being increasingly transmitted and consumed by machines rather than humans to perform automated content analysis tasks. In this paper, we propose an image preprocessor that optimizes the input image for machine consumption prior to encoding by an off-the-shelf codec designed for human consumption. To achieve a better trade-off between the accuracy of the machine analysis task and bitrate, we propose leveraging pre-extracted semantic information to improve the preprocessor’s ability to accurately identify and filter out task-irrelevant information. Furthermore, we propose a two-part loss function to optimize the preprocessor, consisted of a rate-task performance loss and a semantic distillation loss, which helps the reconstructed image obtain more information that contributes to the accuracy of the task. Experiments show that the proposed preprocessor can save up to 48.83% bitrate compared with the method without the preprocessor, and save up to 36.24% bitrate compared to existing preprocessors for machine vision. Mingyi Yang, Luis Herranz, Fei Yang 0004, Luka Murn, Marc Gorriz, Shuai Wan, Fuzheng Yang 0001, Marta Mrak |
ICASSP | 3 |
| 2023 | Dynamic Prompt Learning: Addressing Cross-Attention Leakage for Text-Based Image EditingabstractLarge-scale text-to-image generative models have been a ground-breaking development in generative AI, with diffusion models showing their astounding ability to synthesize convincing images following an input text prompt. The goal of image editing research is to give users control over the generated images by modifying the text prompt. Current image editing techniques are susceptible to unintended modifications of regions outside the targeted area, such as on the background or on distractor objects which have some semantic or visual relationship with the targeted object. According to our experimental findings, inaccurate cross-attention maps are at the root of this problem. Based on this observation, we propose $\textit{Dynamic Prompt Learning}$ ($DPL$) to force cross-attention maps to focus on correct $\textit{noun}$ words in the text prompt. By updating the dynamic tokens for nouns in the textual input with the proposed leakage repairment losses, we achieve fine-grained image editing over particular objects while preventing undesired changes to other image regions. Our method $DPL$, based on the publicly available $\textit{Stable Diffusion}$, is extensively evaluated on a wide range of images, and consistently obtains superior results both quantitatively (CLIP score, Structure-Dist) and qualitatively (on user-evaluation). We show improved prompt editing results for Word-Swap, Prompt Refinement, and Attention Re-weighting, especially for complex multi-object scenes. Kai Wang 0060, Fei Yang 0004, Shiqi Yang 0002, Muhammad Atif Butt, Joost van de Weijer 0001 |
NeurIPS | 2 |
| 2022 | Attention Distillation: self-supervised vision transformer students need more guidance
Kai Wang 0060, Fei Yang 0004, Joost van de Weijer 0001 |
BMVC | 2 |
| 2022 | SlimSeg: Slimmable Semantic Segmentation with Boundary SupervisionabstractAccurate semantic segmentation models typically require significant computational resources, inhibiting their use in practical applications. Recent works rely on well-crafted lightweight models to achieve fast inference. However, these models cannot flexibly adapt to varying accuracy and efficiency requirements. In this paper, we propose a simple but effective slimmable semantic segmentation (SlimSeg) method, which can be executed at different capacities during inference depending on the desired accuracy-efficiency tradeoff. More specifically, we employ parametrized channel slimming by stepwise downward knowledge distillation during training. Motivated by the observation that the differences between segmentation results of each submodel are mainly near the semantic borders, we introduce an additional boundary guided semantic segmentation loss to further improve the performance of each submodel. We show that our proposed SlimSeg with various mainstream networks can produce flexible models that provide dynamic adjustment of computational cost and better performance than independent models. Extensive experiments on semantic segmentation benchmarks, Cityscapes and CamVid, demonstrate the generalization ability of our framework. Danna Xue, Fei Yang 0004, Luis Herranz, Jinqiu Sun, Yu Zhu 0004, Yanning Zhang 0001 |
ACM Multimedia | 2 |
| 2022 | A novel framework for image-to-image translation and image compression
Fei Yang 0004, Yaxing Wang, Luis Herranz, Yongmei Cheng, Mikhail G. Mozerov |
Neurocomputing | 1 |
| 2021 | Slimmable Compressive Autoencoders for Practical Neural Image CompressionabstractNeural image compression leverages deep neural networks to outperform traditional image codecs in rate-distortion performance. However, the resulting models are also heavy, computationally demanding and generally optimized for a single rate, limiting their practical use. Focusing on practical image compression, we propose slimmable compressive autoencoders (SlimCAEs), where rate (R) and distortion (D) are jointly optimized for different capacities. Once trained, encoders and decoders can be executed at different capacities, leading to different rates and complexities. We show that a successful implementation of Slim-CAEs requires suitable capacity-specific RD tradeoffs. Our experiments show that SlimCAEs are highly flexible models that provide excellent rate-distortion performance, variable rate, and dynamic adjustment of memory, computational cost and latency, thus addressing the main requirements of practical image compression. Fei Yang 0004, Luis Herranz, Yongmei Cheng, Mikhail G. Mozerov |
CVPR | 1 |
| 2020 | Variable Rate Deep Image Compression With Modulated AutoencoderabstractVariable rate is a requirement for flexible and adaptable image and video compression. However, deep image compression methods (DIC) are optimized for a single fixed rate-distortion (R-D) tradeoff. While this can be addressed by training multiple models for different tradeoffs, the memory requirements increase proportionally to the number of models. Scaling the bottleneck representation of a shared autoencoder can provide variable rate compression with a single shared autoencoder. However, the R-D performance using this simple mechanism degrades in low bitrates, and also shrinks the effective range of bitrates. To address these limitations, we formulate the problem of variable R-D optimization for DIC, and propose modulated autoencoders (MAEs), where the representations of a shared autoencoder are adapted to the specific R-D tradeoff via a modulation network. Jointly training this modulated autoencoder and the modulation network provides an effective way to navigate the R-D operational curve. Our experiments show that the proposed method can achieve almost the same R-D performance of independent models with significantly fewer parameters. Fei Yang 0004, Luis Herranz, Joost van de Weijer 0001, José Antonio Iglesias Guitián, Antonio M. López 0001, Mikhail G. Mozerov |
IEEE Signal Process. Lett. | 1 |
| 2019 | Sparse Data Interpolation Using the Geodesic Distance Affinity SpaceabstractIn this letter, we adapt the geodesic distance-based recursive filter to the sparse data interpolation problem. The proposed technique is general and can be easily applied to any kind of sparse data. We demonstrate its superiority over other interpolation techniques in three experiments for qualitative and quantitative evaluation. In addition, we compare our method with the popular interpolation algorithm presented in the paper on EpicFlow optical flow, which is intuitively motivated by a similar geodesic distance principle. The comparison shows that our algorithm is more accurate and considerably faster than the EpicFlow interpolation technique. Mikhail G. Mozerov, Fei Yang 0004, Joost van de Weijer 0001 |
IEEE Signal Process. Lett. | 2 |