EDBT 2026 Demo / reviewers in the wild / expert
Li Yu 0004
dblp:70/5913-4
· DBLP profile ↗
24ranked-venue papers
15as first author
21since 2021 · last 2026
0000-0002-8855-8089ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 20 · 12 first-author · 17 since 2021Artificial intelligence and machine learning · 6 · 3 first-author · 6 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | FANeRV: frequency separation and augmentation based neural representation for video
Li Yu 0004, Jimin Xiao, Moncef Gabbouj |
Expert Syst. Appl. | 1 |
| 2026 | Deep learning-based point cloud upsampling: A survey of methodologies, performance comparisons, and noise robustness analysis
Yihang Yin, Li Yu 0004, Wei Zhou 0021, Moncef Gabbouj |
Neurocomputing | 2 |
| 2026 | VATP360: viewport adaptive 360-degree video streaming based on tile priority
Li Yu 0004, Zhiyu Pang, Yao Zhao 0001 |
Neural Comput. Appl. | 1 |
| 2026 | DVLTA-VQA: Decoupled Vision-Language Modeling With Text-Guided Adaptation for Blind Video Quality AssessmentabstractInspired by the dual-stream (dorsal and ventral streams) theory of the human visual system (HVS), recent Video Quality Assessment (VQA) methods have integrated Contrastive Language-Image Pretraining (CLIP) to enhance semantic understanding. However, as CLIP is originally designed for images, it lacks the ability to adequately capture the temporal dynamics and motion perception (dorsal stream) inherent in videos. To address this limitation, we propose DVLTA-VQA (Decoupled Vision-Language Modeling with Text-Guided Adaptation), which decouples CLIP’s visual and textual components to better align with the NR-VQA pipeline. Specifically, we introduce a Video-Based Temporal CLIP module and a Temporal Context Module to explicitly model motion dynamics, effectively enhancing the dorsal stream representation. Complementing this, a Basic Visual Feature Extraction Module is employed to strengthen spatial detail analysis in the ventral stream. Furthermore, we propose a text-guided adaptive fusion strategy that leverages textual semantics to dynamically weight visual features, facilitating effective spatiotemporal integration. Extensive experiments on multiple public datasets demonstrate that the proposed method achieves state-of-the-art performance, significantly improving prediction accuracy and generalization capability. Li Yu 0004, Situo Wang, Wei Zhou 0021, Moncef Gabbouj |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2026 | Perception-Inspired Network for Stereo Image Quality AssessmentabstractExisting stereo image quality assessment (SIQA) methods generally have limitations in binocular fusion and fine-grained perception modeling. To address these issues, we propose a Perception-Inspired Network for SIQA that simulates binocular difference-guided fusion, high-frequency sensitivity, and hierarchical perception mechanisms of the human visual system (HVS). First, a difference-guided binocular fusion (DGBF) module is designed to mimic the binocular difference sensitivity mechanism, which exploits difference information at both the feature-level and image-level to optimize binocular fusion. Furthermore, the image distortion primarily affects the high-frequency components, which are critical for perceptual quality. To reflect this, we propose a high-frequency enhancement module (HFEM) to simulate the human eye's sensitivity to edge and texture distortions. Finally, to better achieve fine-grained perception modeling, we propose a hierarchical quality regression strategy that simulates the human perceptual process, from perceiving local details to forming a global quality judgment, thereby achieving a quality prediction more aligned with human subjective evaluation. Experimental results demonstrate that the proposed method outperforms mainstream approaches, achieving a PLCC of 0.9734 on the LIVE I database, and a PLCC of 0.9632 on the LIVE II database. Yongli Chang, Guanghui Yue 0001, Li Yu 0004, Yakun Ju, Hadi Amirpour, Moncef Gabbouj, Wei Zhou 0021 |
IEEE Trans. Image Process. | 4 |
| 2025 | Unsupervised Domain Adaptation on Point Cloud Classification via Imposing Structural Manifolds into Representation Space
Hongchao Zhong, Li Yu 0004, Longkun Zou, Ke Chen 0004 |
CVM (3) | 2 |
| 2025 | Hierarchical Transformer for Panoramic Image Inpainting with Comprehensive Attention Module
Li Yu 0004, Yanjun Gao, Yihang Yin, Farhad Pakdaman, Moncef Gabbouj |
ICIG (2) | 1 |
| 2025 | Spatial-Spectral Aware Learning with Deformable Affinity for Weakly Supervised Semantic SegmentationabstractWeakly supervised semantic segmentation (WSSS) leverages image-level labels for semantic segmentation, reducing the reliance on pixel-level annotations. While recent studies have explored CLIP for this task, challenges remain in capturing sufficient local information and achieving complete object representations. To address these limitations, we propose a spatial-spectral awareness learning strategy that leverages spectral features to extract intricate structural information, compensating for the lack of pixel-level supervision. By effectively integrating spectral and spatial tokens, our method enhances feature representation and preserves fine-grained details. Furthermore, we introduce deformable offsets to incorporate positional constraints, refine the affinity matrix, reduce noise interference, and accurately delineate object boundaries. Our framework significantly improves CLIP’s performance for WSSS. Experiments on PASCAL VOC 2012 and MS COCO 2014 show that our single-stage WSSS approach outperforms state-of-the-art, and even surpasses some multi-stage methods in segmentation results. Code is available at: https://github.com/I2-Multimedia-Lab/CLIP-SSD Yuzhen Zhou, Pan Gao 0001, Li Yu 0004 |
ICME | 3 |
| 2025 | MCHM25: Multimedia Computing for Health and MedicineabstractRecent years have witnessed an unprecedented growth of multimodal data in healthcare, ranging from distributed sensors and medical imaging devices (MRI, CT, X-rays) to digital health platforms that integrate audio, video, 3D geometry, and clinical text. The increasing availability of such data presents significant opportunities for computer-aided diagnosis and intelligent healthcare solutions, yet also poses substantial challenges in multimodal integration, large-scale analysis, and real-world deployment. The 2nd International Workshop on Multimedia Computing for Health and Medicine (MCHM'25), held in conjunction with ACM Multimedia 2025, focuses on advanced multimedia computing techniques, including mobile and hardware solutions, for tackling real-world problems in healthcare. The workshop brings together researchers and practitioners in multimedia computing, artificial intelligence, and medicine to explore emerging methods, applications, and systems that have a direct impact on human health. Wei Zhou 0021, Hadi Amirpour, Li Yu 0004, Jungong Han, Richang Hong, Paul L. Rosin |
ACM Multimedia | 3 |
| 2025 | High-Frequency Enhanced Hybrid Neural Representation for video compression
Li Yu 0004, Jimin Xiao, Moncef Gabbouj |
Expert Syst. Appl. | 1 |
| 2025 | Bridging Domain Gap of Point Cloud Representations via Self-Supervised Geometric AugmentationabstractRecent progress of semantic point clouds analysis is largely driven by synthetic data (e.g., the ModelNet and the ShapeNet), which are typically complete, well-aligned and noisy-free. Therefore, representations of those ideal synthetic point clouds have limited variations in the geometric perspective and can gain good performance on a number of 3D vision tasks such as point cloud classification. In the context of unsupervised domain adaptation (UDA), representation learning designed for synthetic point clouds can hardly capture domain invariant geometric patterns from incomplete and noisy point clouds. To address such a problem, we introduce a novel scheme for induced geometric invariance of point cloud representations across domains, via regularizing representation learning with two self-supervised geometric augmentation tasks. On one hand, a novel pretext task of predicting translation distances of augmented samples is proposed to alleviate centroid shift of point clouds due to occlusion and noises. On the other hand, we pioneer an integration of the self-supervised relational learning on geometrically-augmented point clouds in a cascade manner, utilizing the intrinsic relationship of augmented variants and other samples as extra constraints of cross-domain geometric features. Experiments on the PointDA-10 dataset demonstrate the effectiveness of the proposed method, achieving the state-of-the-art performance. Li Yu 0004, Hongchao Zhong, Longkun Zou, Ke Chen 0004, Pan Gao 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2024 | Panoramic Image Inpainting with Gated Convolution and Contextual Reconstruction LossabstractDeep learning-based methods have demonstrated encouraging results in tackling the task of panoramic image inpainting. However, it is challenging for existing methods to distinguish valid pixels from invalid pixels and find suitable references for corrupted areas, thus leading to artifacts in the inpainted results. In response to these challenges, we propose a panoramic image inpainting framework that consists of a Face Generator, a Cube Generator, a side branch, and two discriminators. We use the Cubemap Projection (CMP) format as network input. The generator employs gated convolutions to distinguish valid pixels from invalid ones, while a side branch is designed utilizing contextual reconstruction (CR) loss to guide the generators to find the most suitable reference patch for inpainting the missing region. The proposed method is compared with state-of-the-art (SOTA) methods on SUN360 Street View dataset in terms of PSNR and SSIM. Experimental results and ablation study demonstrate that the proposed method outperforms SOTA both quantitatively and qualitatively. Li Yu 0004, Yanjun Gao, Farhad Pakdaman, Moncef Gabbouj |
ICASSP | 1 |
| 2024 | Joint End-to-End Image Compression and Denoising: Leveraging Contrastive Learning and Multi-Scale Self-OnnsabstractNoisy images are a challenge to image compression algorithms due to the inherent difficulty of compressing noise. As noise cannot easily be discerned from image details, such as high-frequency signals, its presence leads to extra bits needed for compression Since the emerging learned image compression paradigm enables end-to-end optimization of codecs, recent efforts were made to integrate denoising into the compression model, relying on clean image features to guide denoising. However, these methods exhibit suboptimal performance under high noise levels, lacking the capability to generalize across diverse noise types. In this paper, we propose a novel method integrating a multi-scale denoiser comprising of Self Organizing Operational Neural Networks, for joint image compression and denoising. We employ contrastive learning to boost the network ability to differentiate noise from high frequency signal components, by emphasizing the correlation between noisy and clean counterparts. Experimental results demonstrate the effectiveness of the proposed method both in rate-distortion performance, and codec speed, outperforming the current state-of-the-art. Li Yu 0004, Farhad Pakdaman, Moncef Gabbouj |
ICIP | 2 |
| 2024 | Pointsoup: High-Performance and Extremely Low-Decoding-Latency Learned Geometry Codec for Large-Scale Point Cloud Scenes
Kang You, Li Yu 0004, Pan Gao 0001, Dandan Ding |
IJCAI | 3 |
| 2024 | MAESR360: Masked autoencoder-based 360-degree video streaming via multi-scale feature fusionabstract360-degree video streaming is becoming increasingly popular for its immersive experience. Traditional adaptive tile-based streaming methods allocate the bitrates according to view-port prediction, which effectively reduces required transmission bandwidth, but it will cause serious quality degradation when the viewport prediction is inaccurate. Thus, some researchers propose visual reconstruction and enhancement-based 360-degree video streaming framework, which can reconstructs the whole frame at very low bitrates. However, existing frameworks are built upon image-based visual reconstruction methods, which do not fully consider the characteristics of videos. In this paper, we propose a masked autoencoder-based, multi-scale optimized framework for 360-degree video streaming (MAESR360), which fully considers the temporal relevance of the video. We utilize spatio-temporal downsampling and high-ratio tube masking strategies to effectively reduce the amount of transmitted data. Additionally, we design a lightweight visual reconstruction model based on multi-scale feature fusion to recover the visual quality of video frames. The effectiveness of our proposed method is demonstrated through extensive experiments. Li Yu 0004, Zhiyu Pang, Moncef Gabbouj |
VCIP | 1 |
| 2024 | Multi-Swin Transformer Based Spatio-Temporal Information Exploration for Compressed Video Quality EnhancementabstractSpatio-temporal information plays an important role in compressed video quality enhancement. Most advanced studies use deformable convolution or Swin transformer to explore spatio-temporal information. However, deformable convolution based methods may incur inaccurate motion compensation due to the compression artifacts and limited receptive fields. The Swin transformer based approaches are unable to fully explore the spatio-temporal information, limited by its rigid window-based mechanism. To solve the above problems, we propose a novel multi-Swin transformer-based network for compressed video quality enhancement to better explore spatio-temporal information. The whole workflow consists of the Local Alignment (LA) Module, the Global Refinement Fusion (GRF) Module, and the Quality Enhancement (QE) Module. The LA module roughly perceives the local motion through the deformable fusion. Subsequently, the GRF module employs the proposed multi-Swin transformer to enhance the spatio-temporal perception. Finally, the QE module effectively restores the texture details across various scales. Extensive experimental results prove the effectiveness of the proposed method. Li Yu 0004, Shiyu Wu, Moncef Gabbouj |
IEEE Signal Process. Lett. | 1 |
| 2023 | MAMIQA: No-Reference Image Quality Assessment Based on Multiscale Attention Mechanism With Natural Scene StatisticsabstractNo-Reference Image Quality Assessment aims to evaluate the perceptual quality of an image, according to human perception. Many recent studies use Transformers to assign different self-attention mechanisms to distinguish regions of an image, simulating the perception of the human visual system (HVS). However, the quadratic computational complexity caused by the self-attention mechanism is time-consuming and expensive. Meanwhile, the image resizing in the feature extraction stage loses the full-size image quality. To address these issues, we propose a lightweight attention mechanism using decomposed large-kernel convolutions to extract multiscale features, and a novel feature enhancement module to simulate HVS. We also propose to compensate the information loss caused by image resizing, with supplementary features from natural scene statistics. Experimental results on five standard datasets show that the proposed method surpasses the SOTA, while significantly reducing the computational costs. Li Yu 0004, Farhad Pakdaman, Miaogen Ling, Moncef Gabbouj |
IEEE Signal Process. Lett. | 1 |
| 2022 | High-frequency guided CNN for video compression artifacts reductionabstractIn this paper, we propose a high-frequency guided CNN for video compression artifacts reduction. In the proposed method, high frequency component in Y channel is extracted and used to guide the quality enhancement of all Y, U, V channels. As high frequency component contains the edge and contour information of the objects in the image, which is of vital importance to both subjective and objective quality. In general, the proposed method consists of two modules: the high frequency guidance module and the quality enhancement module. The high-frequency guidance module uses multiple octave convolutions to extract the high-frequency component in Y channel and then fuse it into the features of Y, U, and V channels. While in the quality enhancement module, multiple CNN residual blocks are used for the quality enhancement of Y, U, and V channels. The proposed method was integrated into both HM-16.22 and VTM-16.0. The results on the JVET test sequence under All Intra configuration shows the effectiveness of the proposed method. Compared with HEVC, the proposed method achieves the average BD-rate reductions of -12.3%, -22.7% and -23.5% for Y, U and V channels respectively. Compared with VVC, the average BD-rate reductions are -6.7%, -12.3% and -13.2% correspondingly. Li Yu 0004, Wenshuai Chang, Moncef Gabbouj |
VCIP | 1 |
| 2022 | Neural texture transfer assisted video coding with adaptive up-sampling
Li Yu 0004, Wenshuai Chang, Weize Quan, Jimin Xiao, Dong-Ming Yan 0001, Moncef Gabbouj |
Signal Process. Image Commun. | 1 |
| 2021 | Discrete-Continuous Action Space Policy Gradient-Based Attention for Image-Text MatchingabstractImage-text matching is an important multi-modal task with massive applications. It tries to match the image and the text with similar semantic information. Existing approaches do not explicitly transform the different modalities into a common space. Meanwhile, the attention mechanism which is widely used in image-text matching models does not have supervision. We propose a novel attention scheme which projects the image and text embedding into a common space and optimises the attention weights directly towards the evaluation metrics. The proposed attention scheme can be considered as a kind of supervised attention and requiring no additional annotations. It is trained via a novel Discrete-continuous action space policy gradient algorithm, which is more effective in modelling complex action space than previous continuous action space policy gradient. We evaluate the proposed methods on two widely-used benchmark datasets: Flickr30k and MS-COCO, outperforming the previous approaches by a large margin. Shiyang Yan, Li Yu 0004, Yuan Xie 0006 |
CVPR | 2 |
| 2021 | SVM based approach for complexity control of HEVC intra coding
Farhad Pakdaman, Li Yu 0004, Mahmoud Reza Hashemi, Mohammed Ghanbari 0001, Moncef Gabbouj |
Signal Process. Image Commun. | 2 |
| 2018 | Convolutional Neural Network for Intermediate View Enhancement in Multiview StreamingabstractMultiview video streaming continues to gain popularity due to the great viewing experience it offers, as well as its availability that has been enabled by increased network throughput and other recent technical developments. User demand for interactive multiview video streaming that provides seamless view switching upon request is also increasing. However, it is a highly challenging task to stream stable and high quality videos that allow real-time scene navigation within the bandwidth constraint. In this paper, a convolutional neural network (ConvNet)-assisted seamless multiview video streaming system is proposed to tackle the challenge. The proposed method solves the problem from two perspectives. First, a ConvNet-assisted multiview representation method is proposed, which provides flexible interactivity without compromising on multiview video compression efficiency. Second, a bit allocation mechanism guided by a navigation model is developed to provide seamless navigation and adapt to network bandwidth fluctuations at the same time. These two blocks work closely to provide an optimized viewing experience to users. They can be integrated into any existing multiview video streaming framework to enhance overall performance. Experimental results demonstrate the effectiveness of the proposed method for seamless multiview streaming. Li Yu 0004, Tammam Tillo, Jimin Xiao, Marco Grangetto |
IEEE Trans. Multim. | 1 |
| 2015 | Statistical approach for motion estimation skipping (SAMEK)abstractHigh Efficient Video Coding (HEVC) standard has achieved significant rate-distortion improvement over the previous standard H.264/AVC. However, the complexity that comes from its flexible data structure representation is an obstacle for its wide application. To reduce the overall complexity and encoding time, this paper proposes a statistical approach for motion estimation skipping (SAMEK). The SAMEK method avoids some unnecessary motion estimations in units with less probability of being referenced. These units can be recognized by the two rules of SAMEK method, namely ZeroCase and DecreaseCase. These two rules are summarized by statistically analyzing the relationships between each PU and its references. The experimental results demonstrate that our method can save up to 9.5% encoding time (averagely 6.87%) with negligible rate-distortion losses (averagely 0.006 dB) when compared with HEVC encoder with Test Zone search (TZ search) enabled. Li Yu 0004, Jimin Xiao, Tammam Tillo, Ce Zhu |
ICIP | 1 |
| 2014 | Dynamic redundancy allocation for video streaming using Sub-GOP based FEC codeabstractReed-Solomon erasure code is one of the most studied protection methods for video streaming over unreliable networks. As a block-based error correcting code, large block size and increased number of parity packets will enhance its protection performance. However, for video applications this enhancement is sacrificed by the error propagation and the increased bitrate. So, to tackle this paradox, we propose a rate-distortion optimized redundancy allocation scheme, which takes into consideration the distortion caused by losing each slice and the propagated error. Different from other approaches, the amount of introduced redundancy and the way it is introduced are automatically selected without human interventions based on the network condition and video characteristics. The redundancy allocation problem is formulated as a constraint optimization problem, which allows to have more flexibility in setting the block-wise redundancy. The proposed scheme is implemented in JM14.0 for H.264, and it achieves an average gain of 1dB over the state-of-the-art approach. Li Yu 0004, Jimin Xiao, Tammam Tillo |
VCIP | 1 |