EDBT 2026 Demo / reviewers in the wild / expert
Tiesong Zhao
dblp:25/5740
· DBLP profile ↗
118ranked-venue papers
13as first author
94since 2021 · last 2026
0000-0002-7497-8883ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 98 · 13 first-author · 75 since 2021Artificial intelligence and machine learning · 11 · 11 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 7 since 2021Computer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | KASS: Efficient video artifact removal via Kernel-Adaptive Spatiotemporal Synchronization
Liqun Lin, Fawei Tang, Yipeng Liao, Tiesong Zhao |
Comput. Vis. Image Underst. | 5 |
| 2026 | Structural Similarity in Deep Features: Unified Image Quality Assessment Robust to Geometrically Disparate ReferenceabstractImage Quality Assessment (IQA) with references plays an important role in optimizing and evaluating computer vision tasks. Traditional methods assume that all pixels of the reference and test images are fully aligned. Such Aligned-Reference IQA (AR-IQA) approaches fail to address many real-world problems with various geometric deformations between the two images. Although significant effort has been made to attack Geometrically-Disparate-Reference IQA (GDR-IQA) problem, it has been addressed in a task-dependent fashion, for example, by dedicated designs for image super-resolution and retargeting, or by assuming the geometric distortions to be small that can be countered by translation-robust filters or by explicit image registrations. Here we rethink this problem and propose a unified, non-training-based Deep Structural Similarity (DeepSSIM) approach to address the above problems in a single framework, which assesses structural similarity of deep features in a simple but efficient way and uses an attention calibration strategy to alleviate attention deviation. The proposed method, without application-specific design, achieves state-of-the-art performance on AR-IQA datasets and meanwhile shows strong robustness to various GDR-IQA test cases. Interestingly, our test also shows the effectiveness of DeepSSIM as an optimization tool for training image super-resolution, enhancement and restoration, implying an even wider generalizability. Tiesong Zhao, Zhou Wang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2026 | Leveraging CV to haptic processing: Cross tactile-visual mapping based on shared information
Qian Liu 0001, Tiesong Zhao |
Pattern Recognit. | 5 |
| 2026 | Deep learning-based underwater species recognition: Advances, benchmarks and challenges
Tiesong Zhao |
Signal Process. Image Commun. | 3 |
| 2026 | Pixel-Level Just Noticeable Difference in Sonar Images: Modeling and ApplicationsabstractSonar images are vital in ocean explorations but face transmission challenges due to limited bandwidth and unstable channels. The Just Noticeable Difference (JND) represents the minimum distortion detectable by human observers. By eliminating perceptual redundancy, JND offers a solution for efficient compression and accurate Image Quality Assessment (IQA) to enable reliable transmission. However, existing JND models prove inadequate for sonar images due to their unique redundancy distributions and the absence of pixel-level annotated data. To bridge these gaps, we propose the first sonar-specific, picture-level JND dataset and a weakly supervised JND model that infers pixel-level JND from picture-level annotations. Our approach starts with pretraining a perceptually lossy/lossless predictor, which collaborates with sonar image properties to drive an unsupervised generator producing Critically Distorted Images (CDIs). These CDIs maximize pixel differences while preserving perceptual fidelity, enabling precise JND map derivation. Furthermore, we systematically investigate JND-guided optimization for sonar image compression and IQA algorithms, demonstrating favorable performance enhancements. Weiming Lin, Qianxue Feng, Rongxin Zhang, Tiesong Zhao |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2026 | FTGID: Fine-Grained Text-Driven Framework for Universal Generative Image DetectionabstractThe rapid progress of generative models has made detecting realistic forgeries a critical challenge for security and trust. Existing image and frequency-based methods depend on dataset-specific artifacts with poor generalization, while Vision-Language Model (VLM)-based methods remain limited by coarse prompts and underused cross-modal alignment. To address these issues, we propose a Fine-grained Text-driven Generative Image Detection (FTGID) framework, which enables comprehensive detection through multi-modal cues. First, we design a Layer-wise Adaptive Global Extractor (LAGE) that stabilizes multi-level global representations through adaptive CLS token fusion with lightweight calibration and parameter-efficient tuning. Second, we propose a Fine-grained Text-guided Local Enhancer (FTLE) that performs patch-level text-visual interaction to enhance the localization of forgery-relevant regions. Third, we introduce a High-frequency Artifact Feature Extractor (HAFE) that adaptively captures discriminative high-frequency cues, enabling more reliable detection of subtle generative artifacts. Extensive experiments demonstrate that FTGID consistently outperforms state-of-the-art GID methods across diverse generative models and unseen datasets, achieving superior performance, thereby enhancing both robustness and interpretability in open-world generative image detection. Our codes will be made publicly available after the peer review process. Liqun Lin, Tiesong Zhao |
IEEE Trans. Image Process. | 5 |
| 2026 | Multimodal Image Representation Learning With Limited Visual-Tactile DataabstractPrevious multimodal visual-tactile image representation learning (VTL) methods have achieved significant success in object understanding through large-scale training data. However, obtaining sufficient training data is often infeasible, and the above methods struggle to effectively focus on discriminative visual and tactile features with limited data, resulting in degraded performance. To solve the above issue, we introduce a new task called visual-tactile image representation learning with limited data (VTL-L), which better facilitates real-world applications. To address the challenges of limited data and modality discrepancy in the VTL-L task, we propose a novel multi-order feature enhancement-based, alignment-free fusion network (MOA-Net). First, we introduce a multi-order feature enhancement (MFE) module to hierarchically strengthen the detailed and structural representation by aggregating the low- and high-order topological information. This approach can effectively reduce the attention noise and obtain discriminative features with limited data. Then, we propose the alignment-free visual-tactile fusion (AVTF) module to achieve representative spatial and channel features and perform the cross-modality fusion without alignment, which efficiently mitigates the modality discrepancy. Finally, we develop a dual counterfactual intervention (DCI) loss to jointly optimize fused visual-tactile feature and probability distributions, thereby improving the performance of the MOA-Net in the VTL-L task. Extensive experiments demonstrate the superiority of the proposed method across three types of tasks on four datasets under diverse limited-data settings (source code available at: https://github.com/liuxiangqiu007/MOA-Net). Liuxiang Qiu, Hui Da, Wenxi Liu, Yuzhen Niu, Hanli Wang, Tiesong Zhao |
IEEE Trans. Image Process. | 6 |
| 2026 | D2C-SID: A Divergence to Convergence Strategy for Sonar Image DenoisingabstractSonar imaging serves as a crucial sensing modality in underwater environments, yet its outputs are inherently de graded by substantial speckle noise arising from background in terference, biological activity, and device self-noise, posing unique challenges for effective denoising. Existing sonar image denoising techniques take a deterministic approach and produce only one outcome. However, this may neglect the complexities present in real-world situations, thereby limiting the effectiveness of denoising. To address this challenge, we first create a sonar noise dataset with corresponding reference images obtained through well esigned pairwise comparisons. Subsequently, we introduce the Divergence-to-Convergence Sonar Image Denoising (D2C SID) network. This network treats denoising as an uncertainty problem and generates multiple plausible denoised samples. It accomplishes this by leveraging a latent space constructed via a conditional variational autoencode. By this process, our method reduces the reliance on any single estimation. Experimental results on both forward-looking and side-scan sonar images confirm the efficacy of the D2C-SID approach. Fengquan Lan, Boqin Cai, Dongjie Chen, Tiesong Zhao |
IEEE Trans. Multim. | 6 |
| 2026 | A Lightweight and Unified Framework for Underwater Image Utility EvaluationabstractTransmitting images through narrow-bandwidth underwater acoustic channels presents a significant challenge. However, for many applications, transmitting the entire image is not necessary. Instead, conveying only the critical information is typically sufficient to ensure task success. To address this, we have introduced a preprocessing strategy at the image acquisition stage, which involves down-sampling or feature extraction to distill the data into its essential components. This strategy is guided by utility considerations to ensure that the transmitted data is both concise and informative. However, existing quality evaluation methods are inadequate for providing accurate utility assessments and fail to simultaneously evaluate the utility of both images and features within a unified framework. To bridge this gap in underwater object detection (one of the typical underwater task), we have developed an underwater image utility dataset and introduced the Lightweight and Unified Utility for Underwater Image Evaluation (LU3IE). LU3IE leverages features derived from multiple detection models to enhance generalization. It then refines these features through dimensionality reduction and commonality extraction, and simplifies complexity through the use of knowledge distillation. Experimental results validate the effectiveness and scalability of LU3IE in accurately assessing the utility of underwater images. Lishi Xu, Honggang Liao, Nanfeng Jiang, Tiesong Zhao |
IEEE Trans. Multim. | 5 |
| 2026 | Illumination-Guided Grouped Attention and Masked Progressive Denoising for Low-Light Image EnhancementabstractExisting Transformer- or Mamba-based low-light image enhancement (LLIE) methods can capture long-range dependencies in images and achieve global degradation restoration, but suffer from insufficient detail recovery. Besides, these methods do not consider the different illumination levels in various regions or enhance the regional illumination accordingly. Furthermore, existing denoising methods often overfit to single noise distribution and type, showing poor generalization to the noises in low-light images. To address these issues, we propose an Illumination-guided Grouped Attention (IGA) module and a Masked Progressive Denoising (MPD) module for low-light image enhancement. The IGA module first improves the local feature representation through detail recovery operations, and then iteratively groups the image representation based on the illumination levels and conducts grouped attention to achieve both fine detail recovery and adaptive regional illumination optimization. The MPD module explores correlations within and among regions to enhance the texture and structure representations from a local to global perspective, thus achieving both local and global denoising and well denoising generalization capability. Extensive quantitative and qualitative experiments on eight datasets demonstrate that our proposed method outperforms the existing state-of-the-art low-light image enhancement methods. Furthermore, plugging our proposed IGA and MPD modules into existing Transformer- or Mamba-based LLIE methods can significantly improve their performance, further demonstrating the modules' ability to address the common issues in such methods. Hui Da, Yuzhen Niu, Liuxiang Qiu, Tiesong Zhao, Yuzhong Chen 0001 |
IEEE Trans. Multim. | 5 |
| 2026 | YACT-Net: Asymmetric YUV Color Transfer for Reference-Based ColorizationabstractMulti-camera systems have become extensively integrated into smartphones. One prominent implementation employs a color-monochrome dual-camera system, which synergistically combines data from both sensors to produce high-quality color images. The dual-camera system can capture asymmetric image pairs, consisting of relatively high-resolution monochrome images and low-resolution color images. Existing methods do not effectively address the issues of the inconsistency in spectral and spatial resolutions between color-monochrome image pairs in asymmetric color transfer. In this paper, we propose a novel framework based on YUV color space to address these issues. First, we conduct some analyses and experiments on several color spaces, demonstrating the effectiveness of asymmetric color transfer in YUV color space. Second, we design a YUV-color-space-based Asymmetric Color Transfer Network (YACT-Net) for reference-based colorization. In this network, we use a restoration module and a chrominance reconstruction module that utilize YUV information from image pairs to generate high-quality color images. Third, we create a Stereo Color Transfer (SCT) dataset, which includes both synthetic and real-world datasets, to evaluate the performance of YACT-Net. Experimental results demonstrate that YACT-Net achieves state-of-the-art performance in both quantitative and qualitative results. Our code and dataset are available athttps://github.com/fengshenghong/YACT-Net. Shenghong Feng, Junhong Lin 0001, Shufan Pei, Tiesong Zhao |
IEEE Trans. Multim. | 5 |
| 2026 | Compressed Video Quality Assessment With Fine-Grained Artifact Perception and EvaluationabstractVideos are generally compressed to save storage and transmission bandwidth. Popular lossy video compression inevitably leads to Perceivable Encoding Artifacts (PEAs) that affect user's visual experience. Thus, Compression Artifact Removal (CAR) methods have emerged to eliminate perceivable encoding artifacts after video coding. However, there still lacks of an efficient artifact discrimination and evaluation method to guide the optimization of CAR methods. To solve this problem, we make the first attempt to propose an Artifact Perception and Evaluation Network (APE-Net) that can accurately locate artifacts and evaluate their impacts on user experience. First, we propose an Artifact Perception Module (APM) that captures various types and long-tailed-distributed PEAs with attention learning and data re-weighting, thus greatly improving the perception capability for video compression artifacts. Second, we design an Artifact Evaluation Module (AEM) to fuse all recognized PEAs with visual saliency and random forest regression, which assists the artifact perception model to be in line with human visual characteristics in video quality assessment tasks. Experimental results demonstrate that our proposed APE-Net is superior to the state-of-the-art algorithms on compressed video quality assessment. Our codes will be made publicly available after the peer review process. Liqun Lin, Tiesong Zhao |
IEEE Trans. Multim. | 5 |
| 2026 | Multi-Scale Feature Compression via Multi-Receptive-Field Convolutional Neural Network for Machine VisionabstractMulti-scale feature compression is essential in machine vision tasks for reducing storage and transmission costs while maintaining task performance. However, existing multi-scale feature compression methods fail to effectively extract and aggregate the local and global correlations of multi-scale features, resulting in incomplete elimination of feature redundancies. Moreover, these multi-scale feature compression methods mainly rely on mean square error-based loss functions to optimize signal fidelity, but they fail to adequately preserve semantic information critical to machine vision tasks. To address these issues, a Multi-receptive-field Convolutional Neural Network (MCNN)-based multi-scale feature compression method is proposed in this paper, which not only achieves compact fusion of multi-scale features but also enhances the semantic fidelity of reconstructed features. To effectively eliminate feature redundancies, a multi-receptive-field-based feature fusion module is designed for capturing both local and global correlations in multi-scale features. To enhance the quality of the reconstructed features, a cosine similarity-based multi-fidelity loss function is developed by considering both signal and semantic fidelity. Extensive experiments on the object detection and instance segmentation tasks show that the proposed MCNN outperforms the state-of-the-art multi-scale feature compression methods in terms of compression efficiency. The source code of our proposed MCNN is publicly available athttps://github.com/NUIST-Videocoding/MCNN.git. Zhaoqing Pan, Haihang Wang, Tiesong Zhao, Haoran Xie 0001, Sam Kwong |
IEEE Trans. Multim. | 3 |
| 2026 | Deep Video Coding With Bit-Depth ScalabilityabstractDeep video coding techniques have achieved significant advancements, leading to enhanced compression performance. However, existing approaches are primarily optimized for 8-bit content, thereby limiting their effectiveness in scenarios with different bit-depths. In this paper, we propose a deep bit-depth scalable video codec (DB-SVC) that supports two-layer scalability for different bit-depths. First, we design a base layer (BL) for low bit-depth (LBD) videos, incorporating a dual-stage multi-scale feature extraction module (DFEM) to enhance compression efficiency while providing reference features for subsequent coding. Second, we introduce an inter-layer bit-depth enhancement module (IBEM) that refines the bit-depth of BL reconstructed frames by leveraging interlayer information, thus enhancing the reference quality without increasing coding overhead. Third, we design an enhancement layer (EL) tailored for high bit-depth (HBD) videos, employing a bit-depth residual compression (BRC) method to achieve a more accurate reconstruction of HBD videos. DB-SVC supports progressive decoding of LBD and HBD videos, accommodating diverse display requirements. Experimental results demonstrate that DB-SVC outperforms state-of-the-art codecs in LBD and HBD scenarios. At the same PSNR/MS-SSIM levels, DB-SVC achieves average bit-rate savings of 11.94%/53.35% for 8-bit videos and 55.98%/73.16% for 10-bit videos while comparing with VTM13.2, showcasing its superior compression performance. Zhaoqing Pan, Tiesong Zhao, Xiaoming Tao 0001 |
IEEE Trans. Multim. | 4 |
| 2026 | SSVD: Efficient Video Deinterlacing With Spatiotemporal Synchronization and RefinementabstractVideo deinterlacing remains a significant challenge due to structural artifacts and information loss. When displayed on modern digital devices, early interlaced videos often suffer from complex interlacing and compression artifacts, which severely degrade visual quality. Existing deinterlacing methods typically struggle to handle such diverse artifacts while preserving fine-grained details. To address these issues, we propose a novel Spatiotemporal Synchronization for Video Deinterlacing (SSVD). First, we design a Multi-Directional Shuffling Module (MDSM) to enhance the model's ability to capture spatial dependencies, thereby guiding the prediction of missing fields. Second, a Dynamic Cross-frame Interaction Module (DCIM) is incorporated to implicitly model inter-frame correspondences, effectively leveraging cross-frame information to alleviate the blurring and artifacts. Third, we develop a Gated Refinement Module (GRM) to achieve fine-grained reconstruction. Experimental results demonstrate SSVD is superior to the state-of-the-art algorithms in video deinterlacing tasks, achieving superior visual quality and detail reconstruction. Source code will be made public after the review is completed. Liqun Lin, Ruipeng Gang, Tiesong Zhao, Sam Kwong |
IEEE Trans. Multim. | 4 |
| 2026 | VikitaFusion: Object Recognition Based on Heterogeneous Visual-Kinesthetic-Tactile InformationabstractKinesthetic and tactile information can represent the physical states of objects, encompassing roughness, stiffness, motion, force, and other attributes. The introduction of these can enhance imprecise recognition that relies solely on visual information in cases of light disturbance, occlusion and camouflage. Nevertheless, this task is still challenging due to the heterogeneity among visual, tactile and kinesthetic data. To address this issue, this paper delves into the alignment of heterogeneous data dimensions, the fusion of heterogeneous data features, and the optimization of learning rates for multi-source heterogeneous sensor learning models. Consequently, an effective Visual-Kinesthetic-Tactile Information Fusion (VikitaFusion) network is proposed, which comprises: 1) heterogeneous data extractors that align visual images with tactile and kinesthetic data through image-to-sequence projection; 2) a visual-kinesthetic-tactile Transformer-based domain fusion that mimics human multi-sensory fusion perception through a feature-level fusion block and dynamic fusion blocks; 3) a Periodic Triangulation Learning Rate (PTLR) method aimed at optimizing the learning rate for performance enhancement in multi-source heterogeneous sensor learning models. Extensive experiments demonstrate that VikitaFusion outperforms current state-of-the-art methods with higher recognition accuracy and a lower parameter size. Source code will be made publicly available after the peer review process. Rouqi Zhang, Jinde Zhu, Tiesong Zhao |
IEEE Trans. Multim. | 5 |
| 2025 | Keep the Balance: A Parameter-Efficient Symmetrical Framework for RGB+X Semantic SegmentationabstractMultimodal semantic segmentation is a critical challenge in computer vision, with early methods suffering from high computational costs and limited transferability due to full fine-tuning of RGB-based pre-trained parameters. Recent studies, while leveraging additional modalities as supplementary prompts to RGB, still predominantly rely on RGB, which restricts the full potential of other modalities. To address these issues, we propose a novel symmetric parameter-efficient fine-tuning framework for multimodal segmentation, featuring with a modality-aware prompting and adaptation scheme, to simultaneously adapt the capabilities of a powerful pre-trained model to both RGB and X modalities. Furthermore, prevalent approaches use the global cross-modality correlations of attention mechanism for modality fusion, which inadvertently introduces noise across modalities. To mitigate this noise, we propose a dynamic sparse cross-modality fusion module to facilitate effective and efficient cross-modality fusion. To further strengthen the above two modules, we propose a training strategy that leverages accurately predicted dual-modality results to self-teach the single-modality outcomes. In comprehensive experiments, we demonstrate that our method outperforms previous state-of-the-art approaches across six multimodal segmentation scenarios with minimal computation cost. Jingze Su, Qi Li 0038, Wenjie Yang 0005, Tiesong Zhao, Shengfeng He, Wenxi Liu |
CVPR | 6 |
| 2025 | Efficient Semantic Codec for Real-time Vibrotactile TransmissionabstractNowadays, haptic data has gained a fast-growing volume with enormous interaction points during human-computer interaction and embodied AI. In the near future, the massive haptic signals -encompassing both kinesthetic and vibrotactile signals- will place significant demands on both communication and computing resources. To address this challenge, we propose the first task-oriented sematic codec of low-delay vibrotactile transmission, namely, vibrotactile semantic codec (VTSC). Specifically, we design a perception-based vibrotactile semantic extraction mechanism (PSEM) that considers the high and low thresholds of vibrotactile perception in effective semantic coding while adhering to the low delay constraint. Inspired by this principle, we then propose a vibrotactile semantic encoder (VSE) with local and global semantic extractors, which can efficiently extract and preserve semantic features within the short frame context. Besides, we present a semantic distribution loss function to enhance the learning of meaningful representations. Comprehensive experiments demonstrate the superiority of our VTSC, achieving significantly higher task accuracy than the state-of-the-art vibrotactile codecs at the same compression ratio (CR), e.g. 60% improvement when CR=256. When compared to transferred audio-visual sematic codecs, our VTSC also shows promising improvements, validating the effectiveness our approach. Runjie Wang, Kemi Chen, Shuijie Li, Mingkai Chen 0001, Tiesong Zhao |
ACM Multimedia | 5 |
| 2025 | RobustVisH: Robust Visual-Haptic Cross-Modal Recognition under Transmission InterferenceabstractEmbodied AI calls for a reliable, cross-modal object recognition that deeply mines High-Quality (HQ) object appearance (i.e., visual information) and touch details (i.e., haptic information). While in real-world scenarios, cross-modal data is usually degraded due to data acquisition and delivery in complex environments. In this paper, we propose a Robust Visual-Haptic recognition (RobustVisH) model that identifies Low-Quality (LQ) visual-haptic data with transmission distortion for the first time. First, we introduce the WIreless Transmission Interference-based Multi-modal benchmark (WITIM) as a visual-haptic dataset under transmission interference. In particular, the dataset consists of WITIM/AU and WITIM/PHAC-2, in which the original signals are obtained from AU and PHAC-2, respectively. Second, we design a trainable weighted fusion and a Transformer encoder based on the bi-directional self-attention mechanism, enabling RobustVisH to form and learn fused visual-haptic features after modality-specific one-dimensional feature encoding. Third, we employ a covariate shift paradigm, transferring knowledge of RobustVisH from HQ data to LQ data, thereby increasing its robustness against transmission-interference inputs. Experimental results demonstrate that the proposed RobustVisH improves the accuracy of the state-of-the-art method by 2.06% and 9.28% on WITIM/AU and WITIM/PHAC-2, respectively. Source code is available at: https://github.com/lylibylily/RobustVisH. Rouqi Zhang, Chengdi Lu, Hancheng Lu, Yang Cao 0010, Tiesong Zhao |
ACM Multimedia | 5 |
| 2025 | MSMA'2025: The 1st International Workshop on Multi-Sensorial Media and ApplicationsabstractToday, truly immersive multimedia systems demand the integration of emerging multi-sensorial media, which go beyond traditional audiovisual signals to include haptics, olfaction, motion capture, electroencephalograms, and other novel media forms. To effectively incorporate these modalities into cutting-edge multimedia systems, advances are needed across the entire pipeline, from processing and encoding to seamless integration. In addition, human-centric factors such as ergonomics and user experience must be considered to ensure practical implementation. Our workshop, the International Workshop on Multi-Sensorial Media and Applications (MSMA'2025), seeks to attract contributions related to multi-sensorial media systems, including system design, evaluation, coding, delivery, media analysis, multi-modal interaction, human factors, ergonomics, and related areas. By fostering collaboration among researchers, MSMA aims to bridge existing work in the field, spark innovation, and push the boundaries of multimedia technology. Tiesong Zhao, Qian Liu 0001, Zhisheng Yan |
ACM Multimedia | 1 |
| 2025 | Exploring underwater image quality: A review of current methodologies and emerging trends
Xiaoyi Xu, Rongxin Zhang, Tiesong Zhao |
Image Vis. Comput. | 6 |
| 2025 | REHair: Efficient hairstyle transfer robust to face misalignment
Liping Ling, Qingxu Lin, Tiesong Zhao |
Pattern Recognit. | 5 |
| 2025 | GAN-based multi-view video coding with spatio-temporal EPI reconstruction
Chengdong Lan, Tiesong Zhao |
Signal Process. Image Commun. | 4 |
| 2025 | SJND: A Spherical Just Noticeable Difference Modelling for 360° video coding
Liqun Lin, Hongan Wei, Tiesong Zhao |
Signal Process. Image Commun. | 7 |
| 2025 | Dynamic-Clustering-Based Color Quantization for Electrophoretic DisplayabstractElectrophoretic Display (EPD) is a reflective technology that closely mimics traditional paper, making it a popular choice in E-readers, IoT devices, and wearables. However, color quantization, which is a critical step to display natural images on EPD with reduced color scales, usually leads to grayscale distortion and edge loss. In this paper, we propose a Dynamic-Clustering-based E-paper Color Quantization (DCECQ) method to address the above issue. First, it employs a dynamically adjustable Particle Swarm Optimization (PSO) clustering, facilitating adaptive threshold optimization for diverse image content. Second, it introduces a Human Visual System (HVS) based model to quantify visual errors and compensates for grayscale ghosting, effectively reducing artifacts such as edge blurring and color distortion. Third, it implements a validation platform for EPD to assess performance under real-world conditions. Experimental results demonstrate that our approach outperforms existing methods across multiple metrics, which attests to its effectiveness and practical applicability. Source code will be made publicly available after the peer review process. Tingyu Cheng, Gongning Yang, Tiesong Zhao |
IEEE Signal Process. Lett. | 5 |
| 2025 | SIAVC: Semi-Supervised Framework for Industrial Accident Video ClassificationabstractSemi-supervised learning suffers from the imbalance of labeled and unlabeled training data in the video surveillance scenario. In this paper, we propose a new semi-supervised learning method called SIAVC for industrial accident video classification. Specifically, we design a video augmentation module called the Super Augmentation Block (SAB). SAB adds Gaussian noise and randomly masks video frames according to historical loss on the unlabeled data for model optimization. Then, we propose a Video Cross-set Augmentation Module (VCAM) to generate diverse pseudo-label samples from the high-confidence unlabeled samples, which alleviates the mismatch of sampling experience and provides high-quality training data. Additionally, we construct a new industrial accident surveillance video dataset with frame-level annotation, namely ECA9, to evaluate our proposed method. Compared with the state-of-the-art semi-supervised learning based methods, SIAVC demonstrates outstanding video classification performance, achieving 88.76% and 89.13% accuracy on ECA9 and Fire Detection datasets, respectively. The source code and the constructed dataset ECA9 will be released inhttps://github.com/AlchemyEmperor/SIAVC. Qinghua Lin, Haoyi Fan, Tiesong Zhao, David Zhang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2025 | Unified No-Reference Quality Assessment for Sonar Imaging and ProcessingabstractSonar technology has been widely used in underwater surface mapping and remote object detection for its light-independent characteristics. Recently, the booming of artificial intelligence further surges sonar image (SI) processing and understanding techniques. However, the intricate marine environments and diverse nonlinear postprocessing operations may degrade the quality of SIs, impeding accurate interpretation of underwater information. Efficient image quality assessment (IQA) methods are crucial for quality monitoring in sonar imaging and processing. Existing IQA methods overlook the unique characteristics of SIs or focus solely on typical distortions in specific scenarios, which limits their generalization capability. In this article, we propose a unified sonar IQA method, which overcomes the challenges posed by diverse distortions. Though degradation conditions are changeable, ideal SIs consistently require certain properties that must be task-centered and exhibit attribute consistency. We derive a comprehensive set of quality attributes from both the task background and visual content of SIs. These attribute features are represented in just ten dimensions and ultimately mapped to the quality score. To validate the effectiveness of our method, we construct the first comprehensive SI dataset. Experimental results demonstrate the superior performance and robustness of the proposed method. Boqin Cai, Jianghe Zhang, Naveed Ur Rehman Junejo, Tiesong Zhao |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2025 | Sonar Image Super-Resolution Based on Structure-Texture Dual PreservationabstractSonar imaging system plays a crucial role in ocean exploration since it can overcome the limitation of light conditions. However, the challenge of low resolution remains in Sonar Images (SIs) due to sonar imaging characteristics and varying compression for low-bandwidth transmission. Most existing image Super-Resolution (SR) methods treated both the structure and texture in the same way, thus failing to simultaneously capture the rich global-local information. Nevertheless, both structure and texture are essential for the visual quality and applications of SIs. In this study, we propose a Structure-Texture Dual-Preserving Network (STDPNet) tailored to capture both local texture details and global structure in a parallel manner for Sonar Image Super-Resolution (SISR). To further explore the internal correlation between structure and texture features, a feature interaction strategy is introduced. Moreover, conventional loss functions for SR often yield smooth results. We propose a hybrid loss function with spectral and local gradient-aware components to preserve frequency content and enhance texture detail. Experimental results validate the superior performance of the proposed STDPNet. Fengquan Lan, Naveed Ur Rehman Junejo, Tiesong Zhao |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2025 | Low-Light Aerial Imaging With Color and Monochrome CamerasabstractAerial imaging aims to produce well-exposed images with rich details. However, aerial photography may encounter low-light conditions during dusk or dawn, as well as on cloudy or foggy days. In such low-light scenarios, aerial images often suffer from issues such as underexposure, noise, and color distortion. Most existing low-light imaging methods struggle with achieving realistic exposure and retaining rich details. To address these issues, we propose an Aerial Low-light Imaging with Color-monochrome Engagement (ALICE), which employs a coarse-to-fine strategy to correct low-light aerial degradation. First, we introduce wavelet transform to design a perturbation corrector for coarse exposure recovery while preserving details. Second, inspired by the binocular low-light imaging mechanism of the human visual system, we introduce uniformly well-exposed monochrome images to guide a refinement restorer, processing luminance and chrominance branches separately for further improved reconstruction. Within this framework, we design a Reference-based Illumination Fusion Module (RIFM) and an Illumination Detail Transformation Module (IDTM) for targeted exposure and detail restoration. Third, we develop a Dual-camera Low-light Aerial Imaging (DuLAI) dataset to evaluate our proposed ALICE. Extensive qualitative and quantitative experiments demonstrate the effectiveness of our ALICE, achieving a PSNR improvement of at least 19.52% over 12 state-of-the-art methods on the DuLAI Syn-R1440 dataset, while providing more balanced exposure and richer details. Our codes and datasets will be made publicly available after the peer review process. Pengwu Yuan, Liqun Lin, Junhong Lin 0001, Yipeng Liao, Tiesong Zhao |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2025 | ColorAssist: Perception-Based Recoloring for Color Vision Deficiency CompensationabstractImage enhancement methods have been widely studied to improve the visual quality of diverse images, implicitly assuming that all human observers have normal vision. However, a large population around the world suffers from Color Vision Deficiency (CVD). Enhancing images to compensate for their perceptions remains a challenging issue. Existing CVD compensation methods have two drawbacks: first, the available datasets and validations have not been rigorously tested by CVD individuals; second, these methods struggle to strike an optimal balance between contrast enhancement and naturalness preservation, which often results in suboptimal outcomes for individuals with CVD. To address these issues, we develop the first large-scale, CVD-individual-labeled dataset called FZU-CVDSet and a CVD-friendly recoloring algorithm called ColorAssist. In particular, we design a perception-guided feature extraction module and a perception-guided diffusion transformer module that jointly achieve efficient image recoloring for individuals with CVD. Comprehensive experiments on both FZU-CVDSet and subjective tests in hospitals demonstrate that the proposed ColorAssist closely aligns with the visual perceptions of individuals with CVD, achieving superior performance compared with the state-of-the-arts. The source code is available at https://github.com/xsx-fzu/ColorAssist. Liqun Lin, Shangxi Xie, Xiahai Zhuang, Tiesong Zhao |
IEEE Trans. Image Process. | 7 |
| 2025 | Facing Differences of Similarity: Intra- and Inter-Correlation Unsupervised Learning for Chest X-Ray Anomaly DetectionabstractAnomaly detection can significantly aid doctors in interpreting chest X-rays. The commonly used strategy involves utilizing the pre-trained network to extract features from normal data to establish feature representations. However, when a pre-trained network is applied to more detailed X-rays, differences of similarity can limit the robustness of these feature representations. Therefore, we propose an intra- and inter-correlation learning framework for chest X-ray anomaly detection. Firstly, to better leverage the similar anatomical structure information in chest X-rays, we introduce the Anatomical-Feature Pyramid Fusion Module for feature fusion. This module aims to obtain fusion features with both local details and global contextual information. These fusion features are initialized by a trainable feature mapper and stored in a feature bank to serve as centers for learning. Furthermore, to Facing Differences of Similarity (FDS) introduced by the pre-trained network, we propose an intra- and inter-correlation learning strategy: 1) We use intra-correlation learning to establish intra-correlation between mapped features of individual images and semantic centers, thereby initially discovering lesions; 2) We employ inter-correlation learning to establish inter-correlation between mapped features of different images, further mitigating the differences of similarity introduced by the pre-trained network, and achieving effective detection results even in diverse chest disease environments. Finally, a comparison with 18 state-of-the-art methods on three datasets demonstrates the superiority and effectiveness of the proposed method across various scenarios. Wei Li 0227, Tiesong Zhao, Bob Zhang 0001 |
IEEE Trans. Medical Imaging | 4 |
| 2025 | Prototype Alignment With Dedicated Experts for Test-Agnostic Long-Tailed RecognitionabstractUnlike vanilla long-tailed recognition trains on imbalanced data but assumes a uniform test class distribution, test-agnostic long-tailed recognition aims to handle arbitrary test class distributions. Existing methods require prior knowledge of test sets for post-adjustment through multi-stage training, resulting in static decisions at the dataset-level. This pipeline overlooks instance diversity and is impractical in real situations. In this work, we introduce Prototype Alignment with Dedicated Experts (PADE), a one-stage framework for test-agnostic long-tailed recognition. PADE tackles unknown test distributions at the instance-level, without depending on test priors. It reformulates the task as a domain detection problem, dynamically adjusting the model for each instance. PADE comprises three main strategies: 1) parameter customization strategy for multi-experts skilled at different categories; 2) normalized target knowledge distillation for mutual guidance among experts while maintaining diversity; 3) re-balanced compactness learning with momentum prototypes, promoting instance alignment with the corresponding class centroid. We evaluate PADE on various long-tailed recognition benchmarks with diverse test distributions. The results verify its effectiveness in both vanilla and test-agnostic long-tailed recognition. Aiping Huang, Tiesong Zhao |
IEEE Trans. Multim. | 4 |
| 2025 | Cross-Modal Haptic Compression Inspired by Embodied AI for Haptic CommunicationsabstractHaptic data compression has gradually become a key issue for emerging real-time haptic communications in Tactile Internet (TI). However, it is challenging to achieve a trade-off between high perceptual quality and compression ratio in haptic data compression scheme. Inspired by the perspective of embodied AI, we propose a cross-modal haptic compression scheme for haptic communications to improve the perception quality on TI devices in this paper. Since multimodal fusion is routinely employed to improve the ability of system in cognition, we assume that haptic codec is guided by visual semantics to optimize parameter settings in the coding process. We first design a multi-dimensional tactile feature fusion network (MTFFN) relying on multi-head attention mechanism. The MTFFN extracts the multi-dimensional features from the material surface and maps them to infer the coding parameters. Secondly, we provide second-order difference and linear interpolation to establish an criterion for the determination of optimal codec parameters, which are customized by the material categories so as to give high robustness. Finally, the simulation results reveal that our compression scheme can efficiently make a personalized codec procedure for different materials, obtaining more than 17% improvement in terms of compression ratio with high perceptual quality at the same time. Xinmeng Tan, Mingkai Chen 0001, Zhe Zhang 0010, Xin Wei 0001, Tiesong Zhao |
IEEE Trans. Multim. | 8 |
| 2025 | Geometry-Aware Self-Supervised Indoor 360$^{\circ }$ Depth Estimation via Asymmetric Dual-Domain Collaborative LearningabstractBeing able to estimate monocular depth for spherical panoramas is of fundamental importance in 3D scene perception. However, spherical distortion severely limits the effectiveness of vanilla convolutions. To push the envelope of accuracy, recent approaches attempt to utilize Tangent projection (TP) to estimate the depth of$360 ^{\circ }$images. Yet, these methods still suffer from discrepancies and inconsistencies among patch-wise tangent images, as well as the lack of accurate ground truth depth maps under a supervised fashion. In this paper, we propose a geometry-aware self-supervised$360 ^{\circ }$image depth estimation methodology that explores the complementary advantages of TP and Equirectangular projection (ERP) by an asymmetric dual-domain collaborative learning strategy. Especially, we first develop a lightweight asymmetric dual-domain depth estimation network, which enables to aggregate depth-related features from a single TP domain, and then produce depth distributions of the TP and ERP domains via collaborative learning. This effectively mitigates stitching artifacts and preserves fine details in depth inference without overspending model parameters. In addition, a frequent-spatial feature concentration module is devised to simultaneously capture non-local Fourier features and local spatial features, such that facilitating the efficient exploration of monocular depth cues. Moreover, we introduce a geometric structural alignment module to further improve geometric structural consistency among tangent images. Extensive experiments illustrate that our designed approach outperforms existing self-supervised$360 ^{\circ }$depth estimation methods on three publicly available benchmark datasets. Xu Wang 0006, Ziyan He, Qiudan Zhang, You Yang 0002, Tiesong Zhao, Jianmin Jiang |
IEEE Trans. Multim. | 5 |
| 2025 | RDVC: Efficient Deep Video Compression With Regulable Rate and Complexity OptimizationabstractDeep video coding has paved a way to break through the performance bottleneck of reigning hybrid video coding. However, unlike hybrid video codecs, existing deep video codecs cannot offer both flexible rates and regulable complexities within one single codec, which limits their applications. In this article, we propose a Regulable Deep Video Codec (RDVC) to address the above issue. First, we propose an Adaptive Feature Compression (AFC) network that generates variable rates while ensuring Rate-Distortion (RD) performance. The network introduces a two-stage coarse-to-fine rate adjustment that can be controlled by a user-specified rate level. Second, we propose a Spatio-Temporal Feature Propagation (STFP) mechanism to provide high-quality reference information for AFC process. Third, we also utilize slimmable convolutional components in our framework to adjust decoding complexity constrained by user configuration. Experimental results demonstrate that RDVC can adjust the codec structure flexibly according to different user configurations while maintaining advanced performance. On average, it reduces the bit-per-pixel (bpp) by 9.35%$/$58.12% while maintaining the same PSNR/MS-SSIM as the reference software VTM-13.2. Xiaojie Wei 0001, Jielian Lin, Wei Gao 0003, Tiesong Zhao |
IEEE Trans. Multim. | 5 |
| 2025 | HUGS-Net: A Lightweight and Unified Network for Adverse Weather Image DenoisingabstractImage denoising under adverse weather conditions aims to eliminate multiple weather-related noises and restore bright and clear images. Until now, most methods are task-specific while all-in-one algorithms often require a large number of parameters, limiting their model efficiency. Our theoretical analysis and statistical experiments reveal that adverse weather images in Hue channel contain rich contextual information for further processing. With this observation, we propose a novel lightweight HUe-Guided Synergistic Network (HUGS-Net) with multi-scale detail refinement. First, we design a Fourier interaction and evolution module to capture global information from Hue channel without introducing excessive network parameters. Second, we develop a lightweight residue group convolution block to model local texture features, incorporating them with global information to guide noises removal. Third, we introduce a multi-scale fusion module to enhance high-frequency details at a small feature resolution in RGB color space. With the above design, HUGS-Net further supervises and supplements refined background information. Comprehensive experiments showcase the superiority of HUGS-Net across various adverse weather datasets (e.g., image deraining, desnowing, dehazing) with the least parameter size and fast running speed.The source code will be made public after peer review process. Runjie Wang, Yuzhen Niu, Tiesong Zhao |
IEEE Trans. Multim. | 5 |
| 2024 | ViewPCGC: View-Guided Learned Point Cloud Geometry CompressionabstractWith the rise of immersive media applications such as digital museums, virtual reality, and interactive exhibitions, point clouds, as a three-dimensional data storage format, have gained increasingly widespread attention. The massive data volume of point clouds imposes extremely high requirements on transmission bandwidth in the above applications, gradually becoming a bottleneck for immersive media applications. Although existing learning-based point cloud compression methods have achieved specific successes in compression efficiency by mining the spatial redundancy of their local structural features, these methods often overlook the intrinsic connections between point cloud data and other modality data (such as image modality), thereby limiting further improvements in compression efficiency. To address the limitation, we innovatively propose a view-guided learned point cloud geometry compression scheme, namely ViewPCGC. We adopt a novel self-attention mechanism and cross-modality attention mechanism based on sparse convolution to align the modality features of the point cloud and the view image, removing view redundancy through Modality Redundancy Removal Module (MRRM). Simultaneously, side information of the view image is introduced into the Conditional Checkboard Entropy Model (CCEM), significantly enhancing the accuracy of the probability density function estimation for point cloud geometry. In addition, we design a View-Guided Quality Enhancement Module (VG-QEM) in the decoder, utilizing the contour information of the point cloud in the view image to supplement reconstruction details. The superior experimental performance demonstrates the effectiveness of our method. Compared to the state-of-the-art point cloud geometry compression methods, ViewPCGC exhibits an average performance gain exceeding 10% on D1-PSNR metric. Huiming Zheng, Wei Gao 0003, Zhuozhen Yu, Tiesong Zhao, Ge Li 0002 |
ACM Multimedia | 4 |
| 2024 | Underwater image quality optimization: Researches, challenges, and future trends
Hongan Wei, Tiesong Zhao |
Image Vis. Comput. | 5 |
| 2024 | Blind consumer video quality assessment with spatial-temporal perception and fusion
Yuzhen Niu, Yuming Zheng, Zhenlong Wang, Mengzhen Zhong, Tiesong Zhao |
Multim. Tools Appl. | 5 |
| 2024 | Underwater image restoration based on progressive guidance
Jianghe Zhang, Zuxin Lin, Hongan Wei, Tiesong Zhao |
Signal Process. | 5 |
| 2024 | A comprehensive review of quality of experience for emerging video services
Fengquan Lan, Hongan Wei, Tiesong Zhao, Wei Liu 0112 |
Signal Process. Image Commun. | 4 |
| 2024 | Distillation-Based Utility Assessment for Compacted Underwater InformationabstractThe limited bandwidth of underwater acoustic channels poses a challenge to the efficiency of multimedia information transmission. To improve efficiency, the system aims to transmit less data while maintaining image utility at the receiving end. Although assessing utility within compressed information is essential, the current methods exhibit limitations in addressing utility-driven quality assessment. Therefore, this paper introduces a Distillation-based Compacted Information Quality assessment metric (DCIQ) for utility-oriented quality evaluation in the context of underwater machine vision. This method is conducted in the Utility-oriented compacted Image Quality Dataset (UIQD) that contains utility qualities of reference images and their corresponding compressed information at different levels. The utility score is derived from the average confidence of various object detection models. In DCIQ, utility features of compacted information are acquired through transfer learning and mapped using a Transformer. Besides, we propose a utility-oriented cross-model feature fusion mechanism to address different detection algorithm preferences. After that, a utility-oriented feature quality measure assesses compacted feature utility. Finally, we utilize distillation to compress the model by reducing its parameters by 55%. Experiment results effectively demonstrate that our proposed DCIQ can predict utility-oriented quality within compressed underwater information Honggang Liao, Nanfeng Jiang, Hongan Wei, Tiesong Zhao |
IEEE Signal Process. Lett. | 5 |
| 2024 | Efficient Vibrotactile Codec Based on Nbeats NetworkabstractWithin the domain of multimodal communication, the compression of audio, image, and video information is well-established, but compressing haptic signals, including vibrotactile signals, remains challenging. Particularly with the enhancement of haptic signal sampling rate and degrees of freedom, there is a substantial increase in data volume. While existing algorithms have made progress in vibrotactile codecs, there remains significant room for improvement in compression ratios. We propose an innovative Nbeats Network-based Vibrotactile Codec (NNVC) that leverages the statistical characteristics of vibrotactile data. This advanced codec integrates the Nbeats network for precise vibrotactile prediction, residual quantization, efficient Run-Length Encoding, and Huffman coding. The algorithm not only captures the intricate details of vibrotactile signals but also ensures high-efficiency data compression. It exhibits robust overall performance in terms of Signal-to-Noise Ratio (SNR) and Peak Signal-to-Noise Ratio (PSNR), significantly surpassing the state-of-the-art. Dongfang Chen, Tiesong Zhao |
IEEE Signal Process. Lett. | 5 |
| 2024 | Perception-Based Prediction for Efficient Kinesthetic CodingabstractIntegrating haptic feedback with audio and video not only expands the perceptual dimensions of multimedia applications but also enhances user engagement and experience. However, higher signal sampling rates and multi-degree-freedom in haptic interaction increase data significantly. For low-latency and reliable transmission of haptic signal (i.e. tactile and kinesthetic signals), efficient haptic coding is crucial. Existing algorithms overlook haptic signal characteristics, leaving room for improvement. We analyze the statistical characteristics of kinesthetic signals in-depth. Based on the local linear characteristics of position and velocity signals, and the sparse distribution of force signal, we propose an improved kinesthetic coding algorithm by combining dead-zone coding with segmented linear prediction. Extensive experiments on the standard datasets of the IEEE P1918.1.1 Haptic Codecs Task Group demonstrate the superior performance compared to state-of-the-art methods, achieving a more than halved reduction in data transmission rates with high signal-to-noise ratios and structural similarity. Qingfeng Huang, Quanfei Zheng, Tiesong Zhao |
IEEE Signal Process. Lett. | 5 |
| 2024 | Multi-View Graph Embedding Learning for Image Co-Segmentation and Co-LocalizationabstractImage co-segmentation and co-localization exploit inter-image information to identify and extract foreground objects with a batch mode. However, they remain challenging when confronted with large object variations or complex backgrounds. This paper proposes a multi-view graph embedding (MV-Gem) learning scheme which integrates diversity, robustness and discernibility of object features to alleviate this phenomenon. To encourage the diversity, the deep co-information containing both low-layer general representations and high-layer semantic information is generated to form a multi-view feature pool for comprehensive co-object description. To enhance the robustness, a multi-view adaptive weighted learning is formulated to fuse the deep co-information for feature complementation. To ensure the discernibility, the graph embedding and sparse constraint are embedded into the fusion formulation for feature selection. The former aims to inherit important structures from multiple views, and the latter further selects important features to restrain irrelevant backgrounds. With these techniques, MV-Gem gradually recovers all co-objects through optimization iterations. Extensive experimental results on real-world datasets demonstrate that MV-Gem is capable of locating and delineating co-objects in an image group. Aiping Huang, Lijian Li 0004, Le Zhang 0001, Yuzhen Niu, Tiesong Zhao, Chia-Wen Lin |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2024 | LightViD: Efficient Video Deblurring With Spatial-Temporal Feature FusionabstractNatural video capturing suffers from visual blurriness due to high-motion of cameras or objects. Until now, the video blurriness removal task has been extensively explored for both human vision and machine processing. However, its computational cost is still a critical issue and has not yet been fully addressed. In this paper, we propose a novel Lightweight Video Deblurring (LightViD) method that achieves the top-tier performance with an extremely low parameter size. The proposed LightViD consists of a blur detector and a deblurring network. In particular, the blur detector effectively separate blurriness regions, thus avoid both unnecessary computation and over-enhancement on non-blurriness regions. The deblurring network is designed as a lightweight model. It employs a Spatial Feature Fusion Block (SFFB) to extract hierarchical spatial features, which are further fused by ConvLSTM for effective spatial-temporal feature representation. Comprehensive experiments with quantitative and qualitative comparisons demonstrate the effectiveness of our LightViD method, which achieves competitive performances on GoPro and DVD datasets, with reduced computational costs of 1.63M parameters and 96.8 GMACs. Trained model available: https://github.com/wgp/LightVid. Liqun Lin, Guangpeng Wei, Kanglin Liu, Wanjian Feng, Tiesong Zhao |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2024 | Enlarged Motion-Aware and Frequency-Aware Network for Compressed Video Artifact ReductionabstractMaking full use of spatial-temporal information is the key factor for removing compressed video artifacts. Recently, many deep learning-based compression artifact reduction methods have emerged. Among them, a series of methods based on deformable convolution have shown excellent capabilities in spatio-temporal feature extraction. However, local deformable offset prediction and pixel-wise inter-frame feature alignment in the unidirectional form limit the full utilization of temporal features in the existing method. Additionally, compressed video shows inconsistent degrees of distortion on different frequency components, and their restoration difficulty is also nonuniform. For the above problems presented by existing methods, we propose anenlarged motion-aware and frequency-aware network(EMAFA) to further extract spatio-temporal information and enhance information of different frequency components. To perceive different degrees of motion artifacts between compressed frames as accurately as possible, we design a bidirectional dense propagation pattern withpixel-wise and patch-wise deformable convolution(PIPA) module in the feature domain. In addition, we propose amulti-scale atrous deformable alignment(MSADA) module to enrich spatio-temporal features in image domain. Moreover, we design amulti-direction frequency enhancement(MDFE) module with multiple direction convolution to enhance the features of different frequency components. The experimental results show that the proposed method performs better than the state-of-the-art methods in both objective evaluation and visual perception experience. Supplementary experiments for Internet Streamed Video with hybrid-distortion demonstrate that our method also exhibits considerable generalizability for quality enhancement. Wang Liu 0001, Wei Gao 0003, Ge Li 0002, Siwei Ma 0001, Tiesong Zhao, Hui Yuan 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2024 | Target-Aware Camera Placement for Large-Scale Video SurveillanceabstractIn large-scale surveillance of urban or rural areas, an effective placement of cameras is critical in maximizing surveillance coverage or minimizing economic cost of cameras. Existing Surveillance Camera Placement (SCP) methods generally focus on physical coverage of surveillance by implicitly assuming uniform distribution of interested targets or objects across all blocks, which is, however, uncommon in real-world scenarios. In this paper, we are the first to propose a target-aware SCP (tSCP) model, which prioritizes optimizing the task based on uneven target densities, allowing cameras to preferentially cover blocks with more interested targets. First, we define target density as the likelihood of interested targets occurring in a block, which is positively correlated with the importance of the block. Second, we combine aerial imagery with a lightweight object detection network to identify target density. Third, we formulate tSCP as an optimization problem to maximize target coverage in surveillance area, and solve this problem with a target-guided genetic algorithm. Our method optimizes the rational and economical utilization of cameras in large-scale video survillance. Compared with the state-of-the-art methods, our tSCP achieves the highest target coverage with a fixed number of cameras (8.31%-14.81% more than its peers), or utilizes the minimum number of cameras to achieve a preset target coverage. Codes are available athttps://github.com/wu-hongxin/tSCP_main. Hongxin Wu, Qinghou Zeng, Tiesong Zhao, Chang Wen Chen |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2024 | Perception-Driven Similarity-Clarity Tradeoff for Image Super-Resolution Quality AssessmentabstractSuper-Resolution (SR) algorithms aim to enhance the resolutions of images. Massive deep-learning-based SR techniques have emerged in recent years. In such case, a visually appealing output may contain additional details compared with its reference image. Accordingly, fully referenced Image Quality Assessment (IQA) cannot work well; however, reference information remains essential for evaluating the qualities of SR images. This poses a challenge to SR-IQA: How to balance the referenced and no-reference scores for user perception? In this paper, we propose a Perception-driven Similarity-Clarity Tradeoff (PSCT) model for SR-IQA. Specifically, we investigate this problem from both referenced and no-reference perspectives, and design two deep-learning-based modules to obtain referenced and no-reference scores. We present a theoretical analysis based on Human Visual System (HVS) properties on their tradeoff and also calculate adaptive weights for them. Experimental results indicate that our PSCT model is superior to the state-of-the-arts on SR-IQA. In addition, the proposed PSCT model is also capable of evaluating quality scores in other image enhancement scenarios, such as deraining, dehazing and underwater image enhancement. The source code is available at https://github.com/kekezhang112/PSCT. Tiesong Zhao, Yuzhen Niu, Jinsong Hu 0001, Weisi Lin |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2024 | FUVC: A Flexible Codec for Underwater Video TransmissionabstractSmart oceanic exploration has greatly benefitted from AI-driven underwater image and video processing. However, the volume of underwater video content is subject to narrow-band and time-varying underwater acoustic channels. How to support high-utility video transmission at such a limited capacity is still an open issue. In this paper, we propose a Flexible Underwater Video Codec (FUVC) with separate designs for targets-of-interest regions and backgrounds. The encoder locates all targets-of-interests, compresses their corresponding regions with x.265 and if bandwidth allows, compresses the background with a lower bitrate. The decoder reconstructs both streams, identifies clean targets-of-interest and fuses them with the background via a Mask Detection and Background Recovery (MDBR) network. When the background stream is unavailable, the decoder adapts all targets-of-interest to a virtual background via Poisson blending. Experimental results show that FUVC outperforms other codecs with a lower bitrate at the same quality. It also supports a flexible codec for underwater acoustic channels. The database and source code are available at https://github.com/z21110008/FUVC. Yannan Zheng, Tiesong Zhao |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2024 | Perception-and-Cognition-Inspired Quality Assessment for Sonar Image Super-ResolutionabstractDue to the light-independent imaging characteristics, sonar images play a crucial role in fields such as underwater detection and rescue. However, the resolution of sonar images is negatively correlated with the imaging distance. To overcome this limitation, Super-Resolution (SR) techniques have been introduced into sonar image processing. Nevertheless, it is not always guaranteed that SR maintains the utility of the image. Therefore, quantifying the utility of SR reconstructed Sonar Images (SRSIs) can facilitate their optimization and usage. Existing Image Quality Assessment (IQA) methods are inadequate for evaluating SRSIs as they fail to consider both the unique characteristics of sonar images and reconstruction artifacts while meeting task requirements. In this paper, we propose a Perception-and-Cognition-inspired quality Assessment method for Sonar image Super-resolution (PCASS). Our approach incorporates a hierarchical feature fusion-based framework inspired by the cognitive process in the human brain to comprehensively evaluate SRSIs' quality under object recognition tasks. Additionally, we select features at each level considering visual perception characteristics introduced by SR reconstruction artifacts such as texture abundance, contour details, and semantic information to measure image quality accurately. Importantly, our method does not require training data and is suitable for scenarios with limited available images. Experimental results validate its superior performance. Boqin Cai, Sumei Zheng, Tiesong Zhao, Ke Gu 0001 |
IEEE Trans. Multim. | 4 |
| 2024 | Video Compression Artifacts Removal With Spatial-Temporal Attention-Guided EnhancementabstractRecently, many compression algorithms are applied to decrease the cost of video storage and transmission. This will introduce undesirable artifacts, which severely degrade visual quality. Therefore, Video Compression Artifacts Removal (VCAR) aims at reconstructing a high-quality video from its corrupted version of compression. Generally, this task is considered as a vision-related instead of media-related problem. In vision-related research, the visual quality has been significantly improved while the computational complexity and bitrate issues are less considered. In this work, we review the performance constraints of video coding and transfer to evaluate the VCAR outputs. Based on the analyses, we propose a Spatial-Temporal Attention-Guided Enhancement Network (STAGE-Net). First, we employ dynamic filter processing, instead of conventional optical flow method, to reduce the computational cost of VCAR. Second, we introduce self-attention mechanism to design Sequential Residual Attention Blocks (SRABs) to improve visual quality of enhanced video frames with bitrate constraints. Both quantitative and qualitative experimental results have demonstrated the superiority of our proposed method, which achieves high visual qualities and low computational costs. Nanfeng Jiang, Jielian Lin, Tiesong Zhao, Chia-Wen Lin |
IEEE Trans. Multim. | 4 |
| 2024 | Perceptual Decoupling With Heterogeneous Auxiliary Tasks for Joint Low-Light Image Enhancement and DeblurringabstractCapturing images at night are susceptible to inadequate illumination conditions and motion blurring. Given the typical coupling of these two forms of degradation, a pioneer work takes a compact approach of brightening followed by deblurring. However, this sequential approach may compromise informative features and elevate the likelihood of generating unintended artifacts. In this paper, we observe that the co-existing low light and blurs intuitively impair multiple perceptions, making it difficult to produce visually appealing results. To meet these challenges, we propose perceptual decoupling with heterogeneous auxiliary tasks (PDHAT) for joint low-light image enhancement and deblurring. Based on the crucial perceptual properties of the two degradations, we construct two individual auxiliary tasks: coarse preview prediction (CPP) and high-frequency reconstruction (HFR), so that the perception of color, brightness, edges, and details are decoupled into heterogeneous auxiliary tasks to obtain task-specific representations for parallel assisting the main task: joint low-light enhancement and deblurring (LLE-Deblur). Furthermore, we develop dedicated modules to build the network blocks in each branch based on the exclusive properties of each task. Comprehensive experiments are conducted on LOL-Blur and Real-LOL-Blur datasets, showing that our method outperforms existing methods on quantitative metrics and qualitative results. Yuezhou Li, Rui Xu 0028, Yuzhen Niu, Wenzhong Guo, Tiesong Zhao |
IEEE Trans. Multim. | 5 |
| 2024 | Toward Efficient Video Compression Artifact Detection and Removal: A Benchmark DatasetabstractVideo compression leads to compression artifacts, among which Perceivable Encoding Artifacts (PEAs) degrade user perception. Most of existing state-of-the-art Video Compression Artifact Removal (VCAR) methods indiscriminately process all artifacts, thus leading to over-enhancement in non-PEA regions. Therefore, accurate detection and location of PEAs is crucial. In this paper, we propose the largest-ever Fine-grained PEA database (FPEA). First, we employ the popular video codecs, VVC and AVS3, as well as their common test settings, to generate four types of spatial PEAs (blurring, blocking, ringing and color bleeding) and two types of temporal PEAs (flickering and floating). Second, we design a labeling platform and recruit sufficient subjects to manually locate all the above types of PEAs. Third, we propose a voting mechanism and feature matching to synthesize all subjective labels to obtain the final PEA labels with fine-grained locations. Besides, we also provide Mean Opinion Score (MOS) values of all compressed video sequences. Experimental results show the effectiveness of FPEA database on both VCAR and compressed Video Quality Assessment (VQA). We envision that FPEA database will benefit the future development of VCAR, VQA and perception-aware video encoders. The FPEA database has been made publicly available. Liqun Lin, Tiesong Zhao |
IEEE Trans. Multim. | 5 |
| 2024 | Distortion-Aware Self-Supervised Indoor 360$^{\circ }$ Depth Estimation via Hybrid Projection Fusion and Structural RegularitiesabstractOwing to the rapid development of emerging 360$^{\circ }$panoramic imaging techniques, indoor 360$^{\circ }$depth estimation has aroused extensive attention in the community. Due to the lack of available ground truth depth data, it is extremely urgent to model indoor 360$^{\circ }$depth estimation in self-supervised mode. However, self-supervised 360$^{\circ }$depth estimation suffers from two major limitations. One is the distortion and network training problems caused by Equirectangular projection (ERP), and the other is that texture-less regions are quite difficult to back-propagate in self-supervised mode. Hence, to address the above issues, we introduce spherical view synthesis for learning self-supervised 360$^{\circ }$depth estimation. Specifically, to alleviate the ERP-related problems, we first propose a dual-branch distortion-aware network to produce the coarse depth map, including a distortion-aware module and a hybrid projection fusion module. Subsequently, the coarse depth map is utilized for spherical view synthesis, in which a spherically weighted loss function for view reconstruction and depth smoothing is investigated to optimize the projection distribution problem of 360$^{\circ }$images. In addition, two structural regularities of indoor 360$^{\circ }$scenes are devised as two additional supervisory signals to efficiently optimize our self-supervised 360$^{\circ }$depth estimation model, containing the principal-direction normal constraint and the co-planar depth constraint. The principal-direction normal constraint is designed to align the normal of the 360$^{\circ }$image with the direction of the vanishing points. Meanwhile, we employ the co-planar depth constraint to fit the estimated depth of each pixel through its 3D plane. Finally, a depth map is obtained for the 360$^{\circ }$image. Experimental results illustrate that our proposed method achieves superior performance than the current advanced depth estimation methods on four publicly available datasets. Xu Wang 0006, Weifeng Kong, Qiudan Zhang, You Yang 0002, Tiesong Zhao, Jianmin Jiang |
IEEE Trans. Multim. | 5 |
| 2024 | Bilateral Interaction for Local-Global Collaborative Perception in Low-Light Image EnhancementabstractLow-light image enhancement is a challenging task due to the limited visibility in dark environments. While recent advances have shown progress in integrating CNNs and Transformers, the inadequate local-global perceptual interactions still impedes their application in complex degradation scenarios. To tackle this issue, we propose BiFormer, a lightweight framework that facilitates local-global collaborative perception via bilateral interaction. Specifically, our framework introduces a core CNN-Transformer collaborative perception block (CPB) that combines local-aware convolutional attention (LCA) and global-aware recursive Transformer (GRT) to simultaneously preserve local details and ensure global consistency. To promote perceptual interaction, we adopt bilateral interaction strategy for both local and global perception, which involves local-to-global second-order interaction (SoI) in the dual-domain, as well as a mixed-channel fusion (MCF) module for global-to-local interaction. The MCF is also a highly efficient feature fusion module tailored for degraded features. Extensive experiments conducted on low-level and high-level tasks demonstrate that BiFormer achieves state-of-the-art performance. Furthermore, it exhibits a significant reduction in model parameters and computational cost compared to existing Transformer-based low-light image enhancement methods. Rui Xu 0028, Yuezhou Li, Yuzhen Niu, Huangbiao Xu, Yuzhong Chen 0001, Tiesong Zhao |
IEEE Trans. Multim. | 6 |
| 2023 | DeepSVC: Deep Scalable Video Coding for Both Machine and Human VisionabstractNowadays, end-to-end video coding for both machine and human vision has become an emerging research topic. In complicated systems such as large-scale internet of video things (IoVT), feature streams and video streams can be separately encoded and delivered for machine judgement and human viewing. In this paper, we propose a deep scalable video codec (DeepSVC) to support three-layer scalability from machine to human vision. First, we design a semantic layer that encodes semantic features extracted from the captured video for machine analysis. This layer employs a conditional semantic compression (CSC) method to remove redundancies between semantic features. Second, we design a structure layer that can be combined with semantic layer to predict the captured video at a low quality. This layer effectively estimates video frames based on semantic layer with an interlayer frame prediction (IFP) network. Third, we design a texture layer that can be combined with the above two layers to reconstruct high-quality video signals. This layer also takes advantage of the IFP network to improve its coding efficiency. In large-scale IoVT systems, DeepSVC can deliver semantic layer for regular use and transmit the other layers on demand. Experimental results indicate that the proposed DeepSVC outperforms popular codecs for machine and human vision. Compared with scalable extension of H.265/HEVC (SHVC), the proposed DeepSVC reduces average bit-per-pixel (bpp) by 25.51%/27.63%/59.87% at the same mAP/PSNR/MS-SSIM. Sourcecode is available at: https://github.com/LHB116/DeepSVC. Zhichen Zhang, Jielian Lin, Xu Wang 0006, Tiesong Zhao |
ACM Multimedia | 6 |
| 2023 | LDRM: Degradation Rectify Model for Low-light Imaging via Color-Monochrome CamerasabstractLow-light imaging task aims to approximate low-light scenes as perceived by human eyes. Existing methods usually pursue higher brightness, resulting in unrealistic exposure. Inspired by Human Vision System (HVS), where rods perceive more lights while cones perceive more colors, we propose a Low-light Degradation Rectify Model (LDRM) with color-monochrome cameras to solve this problem. First, we propose to use a low-ISO color camera and a high-ISO monochrome camera for low-light imaging under short-exposure of less than 0.1s. Short-exposure could avoid motion blurriness, while monochrome camera captures more photons than color camera. By mimicing HVS, this capture system could benefit low-light imaging. Second, we propose an LDRM model to fuse the color-monochrome image pair into a high-quality image. In this model, we separately restore UV and Y channels through chrominance and luminance branches and use monochrome image to guide the restoration of luminance. We also propose a latent code embedding method to improve the restorations of both branches. Third, we create a Low-light Color-Monochrome benchmark (LCM), including both synthetic and real-world datasets, to examine low-light imaging quality of LDRM and the state-of-the-art methods. Experimental results demonstrate the superior performance of LDRM with visually pleasing results. Codes and datasets are available at https://github.com/StephenLinn/LDRM. Junhong Lin 0001, Shufan Pei, Nanfeng Jiang, Wei Gao 0003, Tiesong Zhao |
ACM Multimedia | 6 |
| 2023 | ELFIC: A Learning-based Flexible Image Codec with Rate-Distortion-Complexity OptimizationabstractLearning-based image coding has attracted increasing attentions for its higher compression efficiency than reigning image codecs. However, most existing learning-based codecs do not support variable rates with a single encoder; their decoders are also of fixed, high computational complexity. In this paper, we propose an End-to-end, Learning-based and Flexible Image Codec (ELFIC) that supports variable rate and flexible decoding complexity. First, we propose a general image codec with Nonlinear Feature Fusion Transform (NFFT) as nonlinear transforms to improve its Rate-Distortion (RD) performance. Second, we propose an Instance-aware Decoding Complexity Allocation (IDCA) approach, which exploits image contents for a tradeoff between reconstruction quality and computational complexity in the decoding process. Third, we propose an RD-Complexity (RDC) optimization algorithm, which maximizes the image quality under given rate and complexity constraints for the whole framework. Experimental results show that ELFIC achie-ves variable rate, flexible decoding complexity with the state-of-the-art RD performance. It also supports a more efficient decoding process by focusing on image contents. Source codes are available at https://github.com/Zhichen-Zhang/ELFIC-Image-Compression. Zhichen Zhang, Jielian Lin, Xu Wang 0006, Tiesong Zhao |
ACM Multimedia | 6 |
| 2023 | HCSD-Net: Single Image Desnowing with Color Space TransformationabstractSingle-image desnowing aims at depressing snowflake noises while preserving a clean background. Existing methods usually mask the locations of noises and remove them in RGB color space. In this paper, we rethink this problem by investigating the impacts of color space selection. Theoretical analysis and experiments reveal that the feature of snowflake noises exhibit different distributions in different color spaces. In particular, these noises are barely seen in Hue channel, which inspires us to recover global structure and texture information of the clean background from Hue channel. More low-frequency information is also found in the Hue channel. With these observations, we propose a novel Hybrid-Color-Space-based Desnowing Network (HCSD-Net). The proposed HCSD-Net extracts low-frequency and high-frequency features in Hue channel and RGB color space, respectively. After that, it utilizes a multi-scale fusion module to enhance high-frequency details at a small feature resolution. These details are further used to supervise and supplement the background information. Extensive experiments demonstrate that our proposed HCSD-Net outperforms state-of-the-art methods on various synthetic and real-world desnowing datasets. Codes are available at https://github.com/ttz-rainbow/HCSD-Net. Nanfeng Jiang, Hongxin Wu, Yuzhen Niu, Tiesong Zhao |
ACM Multimedia | 6 |
| 2023 | Pixel-Level Sonar Image JND Based on Inexact Supervised Learning
Qianxue Feng, Tiesong Zhao |
PRCV (11) | 4 |
| 2023 | Deep hybrid model for single image dehazing and detail refinement
Nanfeng Jiang, Kejian Hu, Tiesong Zhao |
Pattern Recognit. | 6 |
| 2023 | Lightweight Semi-supervised Network for Single Image Rain RemovalabstractDeep learning technologies have shown their advantages in Single Image Rain Removal (SIRR) tasks. However, the derained results of most methods are limited to some challenges. First, due to the lack of real-world rainy/clean image pairs, many methods seriously rely on the labeled synthetic training images and will not effectively remove complex rain streaks in real-world scenarios. Second, most existing SIRR models require high computing power, which considerably limits their real-world applications. To address these issues, we propose a Lightweight Semi-supervised Network (LSNet) for SIRR. Our LSNet utilizes a compact semi-supervised framework to improve generalization ability in real-world rainy images removal. Meanwhile, in our semi-supervised framework, we also design a cascaded sub-network, which progressively removes complex rain streaks via a multi-stage manner. Specially, the multi-stage manner is based on a series of cascaded blocks, where we conduct recursive learning strategy to reduce model parameters. Extensive experimental results demonstrate that our method achieves comparable performance to the state-of-the-arts while has fewer parameters. Nanfeng Jiang, Junhong Lin 0001, Tiesong Zhao |
Pattern Recognit. | 5 |
| 2023 | Saliency-Aware Spatio-Temporal Artifact Detection for Compressed Video Quality AssessmentabstractCompressed videos often exhibit visually annoying artifacts, known as Perceivable Encoding Artifacts (PEAs), which dramatically degrade video visual quality. Subjective and objective measures capable of identifying and quantifying various types of PEAs are critical in improving visual quality. In this letter, we investigate the influence of four spatial PEAs (i.e.blurring, blocking, bleeding, and ringing) and two temporal PEAs (i.e.flickering and floating) on video quality. For spatial artifacts, we propose a visual saliency model with a low computational cost and higher consistency with human visual perception. In terms of temporal artifacts, self-attention based TimeSFormer is improved to detect temporal artifacts. Based on the six types of PEAs, a quality metric called Saliency-Aware Spatio-Temporal Artifacts Measurement (SSTAM) is proposed. Experimental results demonstrate that the proposed method outperforms state-of-the-art metrics. We believe that SSTAM will be beneficial for optimizing video coding techniques. Liqun Lin, Chengdong Lan, Tiesong Zhao |
IEEE Signal Process. Lett. | 5 |
| 2023 | Joint Shared-and-Specific Information for Deep Multi-View ClusteringabstractMulti-view data describes an image sample with different modalities of features, thus provides a more comprehensive description of data. Its three basic characteristics, i.e., consensus, complementary and redundancy, determine its performances in computer vision tasks. In this paper, we effectively exploit the above three characteristics to propose a deep learning scheme with joint shared-and-specific information (JSSI) for multi-view clustering. Aiming at facilitating the consensus, JSSI extracts shared information of multi-view data via an adversarial similarity constraint, which is realized by classification and discrimination interactions. Aiming at reducing the redundancy, JSSI separate out view-specific features and prevent them from interfering with the shared features via a difference constraint. Aiming at ensuring the complementary, JSSI aligns the shared features and then concatenates them with the specific features. We examine the effectiveness of JSSI with multi-view clustering on real-world datasets, such as faces and indoor scenes. Extensive experiments and comparisons show that JSSI outperforms other state-of-the-art methods in most of these datasets. Aiping Huang, Wei Gao 0003, Yuzhen Niu, Tiesong Zhao |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2023 | Low-Light Image Enhancement via Stage-Transformer-Guided NetworkabstractImages collected in low-light environments usually suffer from multiple, non-uniform distributed distortions, including local dark, dim light, backlit and so on. In this paper, we propose a Stage-Transformer-Guided Network (STGNet) that effectively handles region-specific distributions and enhance diverse low-light images. Specifically, our STGNet adopts a multi-stage way to progressively learn hierarchical features that benefit the robustness of our model. At each stage, we design an efficient transformer with horizontal and vertical attentions that jointly capture degradation distributions with different magnitudes and orientations. We also introduce learnable degradation queries to adaptively select task-specific features of degradations for enhancement. In addition, we design a histogram loss for enhancement and combine it with other loss functions, in order to exploit both global contrast and local details during network training. Benefiting from the above contributions, our STGNet achieves the state-of-the-art performances on both synthetic and real-world datasets. Nanfeng Jiang, Junhong Lin 0001, Haifeng Zheng, Tiesong Zhao |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2023 | λ-Domain VVC Rate Control Based on Nash EquilibriumabstractWith a significant Rate-Distortion (RD) improvement than H.265/HEVC, Versatile Video Coding (VVC) has set a new milestone in lossy video compression. It also incorporates the emerging$\lambda $-domain rate control technique, aiming at a higher visual quality under a fixed bit constraint. However, the challenge remains how to efficiently allocate bits to all frames and Coding Tree Units (CTUs). In this paper, we propose an effective solution by formulating the above task as a Nash equilibrium problem, where all CTUs are treated as players that bargains with each other. By introducing$\lambda $-domain RD models, a constrained optimization is derived with no closed-form solution. We then propose a two-step strategy to address this issue: a Newton method to iteratively calculate an intermediate variable, and a final solution of Nash equilibrium to obtain an approximately optimal$\lambda $. Finally, we utilize the derived$\lambda $to perform an effective CTU-level bit allocation, which is the very first attempt to introduce Nash equilibrium in$\lambda $-domain rate control. Experimental results with Common Test Conditions (CTC) demonstrate the effectiveness and superiority of our method, which outperforms the state-of-the-art CTU-level rate allocation algorithms for VVC. Jielian Lin, Aiping Huang, Tiesong Zhao, Xu Wang 0006, Sam Kwong |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2023 | LMQFormer: A Laplace-Prior-Guided Mask Query Transformer for Lightweight Snow RemovalabstractSnow removal aims to locate snow areas and recover clean images without repairing traces. Unlike the regularity and semitransparency of rain, snow with various patterns and degradations seriously occludes the background. As a result, the state-of-the-art snow removal methods usually retains a large parameter size. In this paper, we propose a lightweight but high-efficient snow removal network called Laplace Mask Query Transformer (LMQFormer). Firstly, we present a Laplace-VQVAE to generate a coarse mask as prior knowledge of snow. Instead of using the mask in dataset, we aim at reducing both the information entropy of snow and the computational cost of recovery. Secondly, we design a Mask Query Transformer (MQFormer) to remove snow with the coarse mask, where we use two parallel encoders and a hybrid decoder to learn extensive snow features under lightweight requirements. Thirdly, we develop a Duplicated Mask Query Attention (DMQA) that converts the coarse mask into a specific number of queries, which constraint the attention areas of MQFormer with reduced parameters. Experimental results in popular datasets have demonstrated the efficiency of our proposed model, which achieves the state-of-the-art snow removal quality with significantly reduced parameters and the lowest running time. Codes and models are available athttps://github.com/StephenLinn/LMQFormer. Junhong Lin 0001, Nanfeng Jiang, Tiesong Zhao |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2023 | Deep Quality Assessment of Compressed Videos: A Subjective and Objective StudyabstractVideo quality assessment is critical in optimizing video coding techniques. However, the state-of-the-art methods have limited performance, which is largely due to the lack of large-scale subjective databases for training. In this work, a semi-automatic labeling method is adopted to build a large-scale compressed video quality database, which allows us to label a large number of compressed videos with manageable human workload. The resulting Compressed Video quality database with Semi-Automatic Ratings (CVSAR), so far the largest of compressed video quality database. We train a no-reference compressed video quality assessment model with a 3D CNN for SpatioTemporal Feature Extraction and Evaluation (STFEE). Experimental results demonstrate that the proposed method outperforms state-of-the-art metrics and achieves promising generalization performance in cross-database tests. The CVSAR database has been made publicly available. It can be accessed athttps://github.com/Rocknroll194/CVSAR. Liqun Lin, Zheng Wang 0007, Jiachen He, Tiesong Zhao |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2023 | Decentralized Federated Learning With Markov Chain Based Consensus for Industrial IoT NetworksabstractFederated learning (FL) provides a novel framework to collaboratively train a shared model in a distribution fashion by virtue of a central server. However, FL is inappropriate for a serverless scenario and also suffers from some major drawbacks in Industrial Internet of Things (IIoT) networks, such as unresilience to network failures and communication bottleneck effect. In this article, we propose a novel decentralized federated learning (DFL) approach for IIoT devices to achieve model consensus by exchanging model parameters only with their neighbors rather than a central server. We firstly formulate the problem of model consensus in DFL as a fastest mixing Markov chain problem and then optimize the consensus matrix to improve the convergence rate. Meanwhile, a practical medium access control protocol with time slotted channel hopping is taken into account to implement the proposed approach. Furthermore, we also propose an accumulated update compression method to alleviate communication cost. Finally, extensive simulation results demonstrate that the proposed approach improves accuracy and reduces communication cost especially under the nonindependent identically distribution data distribution. Mengxuan Du, Haifeng Zheng, Xinxin Feng, Youjia Chen, Tiesong Zhao |
IEEE Trans. Ind. Informatics | 5 |
| 2023 | High Efficiency Vibrotactile Codec Based on Gate Recurrent NetworkabstractThe multimedia has achieved dominant positions in both local storage and internet bandwidth, which inevitably promotes the compression of audio, image and video information. Nowadays, the emerging haptic technology, which enhances the immersion in virtual reality and remote control, has also brought new challenges in its codec design. It is thus imperative to develop haptic codecs, including kinesthetic and vibrotactile codecs, with high efficiency and low delay. In this paper, we exploit statistical features of vibrotactile data to develop a Recurrent-Network-based Vibrotactile Codec (RNVC) with high compression efficiency and low coding delay. The proposed encoder consists of vibrotactile estimation by Gate Recurrent Unit (GRU), non-uniform quantization/compensation of residuals and an entropy encoder. In particular, the GRU-based recurrent network is utilized for its high efficiency to predict signals and low complexity to converge. The decoder consists of all counterparts of encoder. Experimental results show the proposed RNVC significantly reduces of original bitrates with negligible encoding delay, which achieves the state-of-the-art coding performance of vibrotactile signal. Tiesong Zhao, Qian Liu 0001, Yuzhen Niu |
IEEE Trans. Multim. | 1 |
| 2023 | Efficient VVC Intra Prediction Based on Deep Feature Fusion and Probability EstimationabstractThe ever-growing multimedia traffic has underscored the importance of effective multimedia codecs. Among them, the up-to-date lossy video coding standard, Versatile Video Coding (VVC), has been attracting attentions of video coding community. However, the gain of VVC is achieved at the cost of significant encoding complexity, which brings the need to realize fast encoder with comparable Rate Distortion (RD) performance. In this paper, we propose to optimize the VVC complexity at intra-frame prediction, with a two-stage framework of deep feature fusion and probability estimation. At the first stage, we employ the deep convolutional network to extract the spatial-temporal neighboring coding features. Then we fuse all reference features obtained by different convolutional kernels to determine an optimal intra coding depth. At the second stage, we employ a probability-based model and the spatial-temporal coherence to select the candidate partition modes within the optimal coding depth. Finally, these selected depths and partitions are executed whilst unnecessary computations are excluded. Experimental results on standard database demonstrate the superiority of proposed method, especially for High Definition (HD) and Ultra-HD (UHD) video sequences. Tiesong Zhao, Weize Feng, Sam Kwong |
IEEE Trans. Multim. | 1 |
| 2022 | Learning-Based Video Coding with Joint Deep Compression and EnhancementabstractEnd-to-end learning-based video coding has attracted substantial attentions by compressing video signals as stacked visual features. This paper proposes an end-to-end deep video codec with jointly optimized compression and enhancement modules (JCEVC). First, we propose a dual-path generative adversarial network (DPEG) to reconstruct video details after compression. An α-path and a β-path concurrently reconstruct the structure information and local textures. Second, we reuse the DPEG network in both motion compensation and quality enhancement modules, which are further combined with other necessary modules to formulate our JCEVC framework. Third, we employ a joint training of deep video compression and enhancement that further improves the rate-distortion (RD) performance of compression. Compared with x265 LDP very fast mode, our JCEVC reduces the average bit-per-pixel (bpp) by 39.39%/54.92% at the same PSNR/MS-SSIM, which outperforms the state-of-the-art deep video codecs by a considerable margin. Sourcecode is available at: https://github.com/fwz1021/JCEVC. Tiesong Zhao, Weize Feng, Hongji Zeng, Yuzhen Niu, Jiaying Liu 0001 |
ACM Multimedia | 1 |
| 2022 | HD-Net: Hierarchical Distillation Network for High-Efficiency Single Image DerainingabstractRain streaks usually result in severe image visual degradation and foreground occlusion, affecting the quality of computer tasks in outdoor scenes. Currently, the mainstream methods in single-image deraining are based on data-driven. However, the deep learning network could be imperfect, with limited power for learning the global information from rain streaks all over the map. In order to solve this problem, we proposed a novel Hierarchical Distillation Network (HD-Net). In this network, Hierarchical Feature Extraction Block (HFEB) can fully utilize the Transformer's learning ability in high-level features, integrate local detail extraction and global structure representation, and compensate for the weakness of the Convolutional Neural Network (CNN), which is overattentive to the underlying image features. Furthermore, the Distillation-Calibration Block (DCB) are adopted to avoid feature redundancy during model training and calibrate the channel and spatial information through the feature transmission, which could significantly improve the learning efficiency. Finally, the experiment results show that our model performs better than traditional CNN models and state-of-the-art methods. Kejian Hu, Zhichen Zhang, Xiang Chen 0015, Nanfeng Jiang, Yu Zhou 0048, Tiesong Zhao |
MMSP | 7 |
| 2022 | Learning-Based Multi-Stage Intra Partition for Versatile Video CodingabstractThe latest standard, Versatile Video Coding (VVC), doubles the coding efficiency over the previous generation standard. However, better performance is at the cost of a sharp increase in coding complexity. In order to reduce the complexity of VVC intra coding, this paper proposes a multi-stage block partition decision framework based on deep learning. First, we propose a three-stage redundant modes removal framework that decreases the number of modes checked in the brute-force process. Then, we build a lightweight CNN to complete the classification task of each stage. To reduce the burden of CNN and adapt to different Coding Unit (CU) sizes, we pre-process the luminance component of CU and use the results as input of the network. Finally, the multi-threshold adjusting scheme is proposed for trading off complexity reduction with the bit-rate increase. The experimental results shows our method can reduce the encoding time ranging from 16.93% to 69.40% with the bit-rate increase ranging from 0.31% to 3.59%. Such results demonstrate that our method has superior performance with a wide range of adjustments compared with other state-of-the-art methods. Hongji Zeng, Tiesong Zhao, Weize Feng, Jielian Lin, Xu Wang 0006 |
MMSP | 2 |
| 2022 | Effective VVC Intra Prediction Based on Ensemble LearningabstractThis paper proposes a fast VVC coding unit partition algorithm based on ensemble convolutional neural network (CNN) by investigating and bagging spatial-temporal adjacent coding features. First, we propose an ensemble CNN framework to aggregate the reference features to predict the depths of uncoded CUs. The proposed model consists of three lightweight CNNs, which can compromise prediction accuracy with overhead. Then a majority voting mechanism is used to unify the predicted depth. By extracting the majority prediction of base learners, the outputs of three CNNs are integrated to obtain the final prediction. To avoid Rate Distortion (RD) loss caused by a small probability of prediction failure, we introduce the optimal depth strategy. During the encoding process, the optimal depth is used for the decision-making of coding unit partition, thus avoiding redundant rate distortion optimization process. Compared with the original encoder, the proposed algorithm saves 21.56% encoding time on average, with a BDBR loss of 0.39%. The performance is even superior in High-Definition (HD) and Ultra HD (UHD) sequences, up to 59.52%. This approach has a great efficiency of time reduction compared with state-of-the-arts with negligible RD performance loss. Hongji Zeng, Tiesong Zhao, Ludi Wu, Weize Feng |
PCS | 3 |
| 2022 | Self-supervised Indoor 360-Degree Depth Estimation via Structural Regularization
Weifeng Kong, Qiudan Zhang, You Yang 0002, Tiesong Zhao, Wenhui Wu 0001, Xu Wang 0006 |
PRICAI (3) | 4 |
| 2022 | DesnowFormer: an effective transformer-based image desnowing networkabstractSingle image desnowing is an important and challenge task for lots of computer vision applications, such as visual tracking and video surveillance. Although existing deep learning-based methods have achieved promising results, most of them rely on the local deep features and neglect global relationship information between the local regions. Therefore, inevitably leading to over-smooth or detail loss results. To solve this issue, we design a UNet-based end-to-end architecture for image desnowing. Specially, to better characterize global information and preserve image detail, we combine Window-based Self-Attention (WSA) transformer block with Residue Spatial Attention (RSA) to build basic unit of our network. Besides, to protect the structure of the image effectively, we also introduce a Residue Channel (RC) loss to guide high-quality image restoration. Extensive experimental results on both synthetic and real-world datasets demonstrate that the proposed model achieves new state-of-the-art results. Nanfeng Jiang, Junhong Lin 0001, Jielian Lin, Tiesong Zhao |
VCIP | 5 |
| 2022 | Joint Front-Edge-Cloud IoVT Analytics: Resource-Effective Design and SchedulingabstractA tremendous amount of visual data are bing collected by the Internet of Video Things (IoVT) systems in which ubiquitous cameras deployed in cities enable new applications in the domains of smart transportation and public security. However, the limited resources in terms of communication, computing, and caching (3C) in the conventional cellular network make it challenging to adopt centralized artificial intelligence (AI) to conduct real-time video-based data analytics. In this work, based on the 5G network architecture with edge servers, a three-phase resource-effective solution is proposed to perform surveillance operations in a large-scale wireless IoVT. The proposed strategy integrates front-end cameras with simple on-chip neural networks performing real-time object-of-interest segmentation, edge servers, and cloud servers with AI functionality carrying out image-based target recognition and video-based target analytics tasks. More importantly, we design the optimal 3C strategy to achieve the best video analytics performance constrained by computing offload ratio, network resource allocation and video-related parameters. Extensive simulations with deep neural networks implemented both at the front-end cameras and in the cloud server have validated the effectiveness of the proposed solution. Youjia Chen, Tiesong Zhao, Peng Cheng 0002, Ming Ding 0001, Chang Wen Chen |
IEEE Internet Things J. | 2 |
| 2022 | Contrastive semantic similarity learning for image captioning evaluation
Chao Zeng 0005, Sam Kwong, Tiesong Zhao, Hanli Wang |
Inf. Sci. | 3 |
| 2022 | SSIM-Variation-Based Complexity Optimization for Versatile Video CodingabstractHitherto, Versatile Video Coding (VVC) has a more magnificent overall performance than High Efficiency Video Coding (HEVC). The Quadtree with Nested Multi-Type Tree (QTMT) coding block structure can substantially enhance video coding quality in VVC. However, the coding gain also leads to a greater coding complexity. Therefore, this letter proposes a Fast Decision Scheme Based on Structural Similarity Index Metric Variation (FDS-SSIMV) to solve this problem. Firstly, the Structural Similarity Index Metric Variation (SSIMV) characteristic among the sub coding units of the spit mode is illustrated. Next, to evaluate the SSIMV value, SSIMV measure strategies are designed for different split modes in this letter. Then, the desired split modes are selected by the SSIMV values. Experimental results show that the proposed method achieves an average encoding Time Saving (TS) and Bjøntegaard Delta Bit Rate (BDBR) with 64.74% and 2.79%, respectively, outperforming the benchmarks. Jielian Lin, Zhichen Zhang, Tiesong Zhao |
IEEE Signal Process. Lett. | 5 |
| 2022 | A Display-Independent Quality Assessment for HDR ImagesabstractHigh Dynamic Range (HDR) images improve visual quality by representing a wide range of luminance in real world. HDR Image Quality Assessment (IQA) is essential in HDR processing technologies. This work focus on the IQA for the HDR applications represented by HDR compression. Towards designing IQA for these tasks, user-end display information should not be considered, given that these tasks only focus on the media intrinsic features. However, most existing HDR metrics are associated with display luminance, hence they can not handle IQA during HDR processing. In this work, we find that it is possible to evaluate HDR image quality without prior knowledge of luminance for display. Based on this, we proposed a low complexity Display-Independent Gradient Magnitude Similarity (DIGMS) metric. Experimental results demonstrate that the proposed DIGMS outperforms the state-of-the-arts in HDR IQA. This work opens up a new perspective for designing IQA algorithms for HDR images. Tiesong Zhao |
IEEE Signal Process. Lett. | 5 |
| 2022 | Weak Supervision Learning for Object Co-SegmentationabstractThe booming of multimedia technologies has promoted the diversity of visual big data. To learn common features across heterogeneous image data, the image co-processing has exhibited its advantages over the separate one. Recently, an active topic of image co-processing is the object co-segmentation, which aims at simultaneously extracting and segmenting shared objects from relevant images. In this paper, we address this problem with a weak-supervision-based probabilistic model. We introduce the weakly supervised priors to alleviate the confusion between common foreground and background, thereby facilitating performance improvement. To ensure the validity of potential background prior knowledge, the nodes on four sides of image are respectively leveraged as the labelled queries. After that, we develop quantitative probabilistic metrics for precisely measuring internal consistencies within single image and correlations between multiple images. Combining the intra-image consistencies with the inter-image correlations, we propose an optimized energy function coupled with binary labeling and graph connectivity to carry out the object co-segmentation. Extensively experimental results on real-world datasets demonstrate that the proposed method achieves superior co-segmentation performance to the state-of-the-arts, with a significantly reduced time consumption. Aiping Huang, Tiesong Zhao |
IEEE Trans. Big Data | 2 |
| 2022 | UIF: An Objective Quality Assessment for Underwater Image EnhancementabstractDue to complex and volatile lighting environment, underwater imaging can be readily impaired by light scattering, warping, and noises. To improve the visual quality, Underwater Image Enhancement (UIE) techniques have been widely studied. Recent efforts have also been contributed to evaluate and compare the UIE performances with subjective and objective methods. However, the subjective evaluation is time-consuming and uneconomic for all images, while existing objective methods have limited capabilities for the newly-developed UIE approaches based on deep learning. To fill this gap, we propose an Underwater Image Fidelity (UIF) metric for objective evaluation of enhanced underwater images. By exploiting the statistical features of these images in CIELab space, we present the naturalness, sharpness, and structure indexes. Among them, the naturalness and sharpness indexes represent the visual improvements of enhanced images; the structure index indicates the structural similarity between the underwater images before and after UIE. We combine all indexes with a saliency-based spatial pooling and thus obtain the final UIF metric. To evaluate the proposed metric, we also establish a first-of-its-kind large-scale UIE database with subjective scores, namely Underwater Image Enhancement Database (UIED). Experimental results confirm that the proposed UIF metric outperforms a variety of underwater and general-purpose image quality metrics. The database and source code are available at https://github.com/z21110008/UIF. Yannan Zheng, Rongfu Lin, Tiesong Zhao, Patrick Le Callet |
IEEE Trans. Image Process. | 4 |
| 2022 | Underwater Image Enhancement With Lightweight Cascaded NetworkabstractDue to light scatter and absorption in waterbody, underwater imaging can be easily impaired with low contrast and visual distortion. The resulting images are often unable to meet the quality requirements of human perception and computer processing. Therefore, Underwater Image Enhancement (UIE) has been attracting extensive research efforts. Although deep learning has demonstrated its great success in many vision tasks, its huge amounts of parameters and computations are not conducive to UIE in resource-limited scenarios. In this paper, we address this issue by proposing a Lightweight Cascaded Network (LCNet) based on Laplacian image pyramids. At each pyramid level, we implement cascaded blocks upon a residual network. Specifically, high quality residuals can be progressively predicted with significantly reduced complexity in a coarse-to-fine fashion. Furthermore, these sub-networks are recursively nested to build our LCNet, thereby reducing the overall computational complexity with reused parameters. Extensive experiments demonstrate that the proposed method performs favorably against the state-of-the-arts in terms of visual quality, model parameters and complexity. Nanfeng Jiang, Yuting Lin 0006, Tiesong Zhao, Chia-Wen Lin |
IEEE Trans. Multim. | 4 |
| 2022 | Learning-Based Quality Assessment for Image Super-ResolutionabstractImage Super-Resolution (SR) techniques improve visual quality by enhancing the spatial resolution of images. Quality evaluation metrics play a critical role in comparing and optimizing SR algorithms, but current metrics achieve only limited success, largely due to the lack of large-scale quality databases, which are essential for learning accurate and robust SR quality metrics. In this work, we first build a large-scale SR image database using a novel semi-automatic labeling approach, which allows us to label a large number of images with manageable human workload. The resulting SR Image quality database with Semi-Automatic Ratings (SISAR), so far the largest of SR-IQA database, contains 12 600 images of 100 natural scenes. We train an end-to-end Deep Image SR Quality (DISQ) model by employing two-stream Deep Neural Networks (DNNs) for feature extraction, followed by a feature fusion network for quality prediction. Experimental results demonstrate that the proposed method outperforms state-of-the-art metrics and achieves promising generalization performance in cross-database tests. The SISAR database and DISQ model will be made publicly available to facilitate reproducible research. Tiesong Zhao, Yuting Lin 0006, Zhou Wang 0001 |
IEEE Trans. Multim. | 1 |
| 2021 | Video Playback Quality Evaluation Based on User Expectation and Memory
Shuyi Ji, Liqun Lin, Tiesong Zhao |
ICIG (3) | 6 |
| 2021 | Game Theory-driven Rate Control for 360-Degree Video CodingabstractThe 360-degree video (omnidirectional video) has become popular recently due to its capability of providing immersive experience, which is generally achieved via spherical moving pictures with freedom of viewpoint changing. Nevertheless, the support of full-view visual contents has inevitably reshaped its perceptual quality metric and dramatically increased its bitrate output after video coding. Therefore in 360-degree video coding, the Rate Control (RC) problem, which aims to maximize the resulted perceptual quality under bitrate constraint, has become a challenging task yet to be addressed. In this paper, we observe a latitude-based bitrate discrepancy in equirectangular-projected 360-degree video coding and further utilize this feature in bitrate allocation under panoramic vision. We introduce game theory to find optimal inter/intra-frame bit allocations that maximize the overall RC performance in terms of utility function. Finally, an overall framework is proposed that is capable of providing both an improved bitrate accuracy and an enhanced perceptual quality. Experimental results demonstrate the efficiency of proposed method, with promising RC performances for 4K and 8K 360-degree videos. Tiesong Zhao, Jielian Lin, Xu Wang 0006, Yuzhen Niu |
ACM Multimedia | 1 |
| 2021 | Single image rain removal via multi-module deep grid network
Nanfeng Jiang, Liqun Lin, Tiesong Zhao |
Comput. Vis. Image Underst. | 4 |
| 2021 | Image Retargeting Quality Assessment Based on Registration Confidence Measure and Noticeability-Based PoolingabstractNowadays, image retargeting approaches have been widely applied to adapt images of various resolutions to heterogenous display devices. To assess the quality of the retargeted images, image retargeting quality assessment (IRQA) has emerged as a critical problem in image quality assessment. In this paper, we address the IRQA problem with a newly proposed framework based on registration confidence measurement (RCM) and noticeability-based pooling (NBP). First, we define the RCM to evaluate the accuracy of image registration, which aligns scenes between the original and retargeted images. We then integrate the proposed RCM with the computed local fidelity of each image block to alleviate the negative influence of inaccurate registration on fidelity measurements. Meanwhile, we present a visual attention fusion (VAF) framework to enhance faces and lines in the saliency map, which are observed to be highly sensitive in the human visual system (HVS). Finally, we propose the NBP strategy, which aggregates the local fidelity of each image block into the overall quality of the retargeted image. Specifically, the NBP strategy sets larger quality ranges for the regions where the visual distortions are more accessible to HVS to reflect the easy noticeability of these regions. Experimental results on the MIT RetargetMe and CUHK datasets demonstrate that the proposed IRQA metric based on RCM and NBP outperforms the state-of-the-art IRQA metrics. Yuzhen Niu, Zhishan Wu, Tiesong Zhao |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2021 | Flexible Complexity Optimization in Multiview Video CodingabstractFlexible optimization on video coding computational complexity is imperative to adapt to diverse terminals in video streaming, especially with high definition videos and increasingly comprehensive encoders. Until now, few efforts have been contributed to the encoder of multiview videos, which complicates its encoding structure with duplex frame display. To address this issue, we propose a flexible complexity optimization framework in this paper. In the first place, the proposed algorithm flexibly reduces the encoder complexity at different external-defined complexity constraints; furthermore, it is also committed to compression efficiency optimization at different complexity levels. The framework is achieved by a hybrid approach:complexity allocation, which allocates the external-defined complexity constraint to all views, hierarchical layers and frames based on linear programming; andcomplexity regulation, which dynamically adjusts local candidate partitions to fulfil the targeted local complexity constraint, with a probability-driven Alternate Partition Cost (APC) minimization. The overall algorithm is implemented on the popular multiview encoder, Multiview High Efficiency Video Coding (MV-HEVC), with promising Rate-Distortion (RD) performances at different complexity constraints, which are also superior to two recent complexity optimization algorithms. Kaiying Xing, Tiesong Zhao, Sam Kwong |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2021 | Joint Learning of Latent Similarity and Local Embedding for Multi-View ClusteringabstractSpectral clustering has been an attractive topic in the field of computer vision due to the extensive growth of applications, such as image segmentation, clustering and representation. In this problem, the construction of the similarity matrix is a vital element affecting clustering performance. In this paper, we propose a multi-view joint learning (MVJL) framework to achieve both a reliable similarity matrix and a latent low-dimensional embedding. Specifically, the similarity matrix to be learned is represented as a convex hull of similarity matrices from different views, where the nuclear norm is imposed to capture the principal information of multiple views and improve robustness against noise/outliers. Moreover, an effective low-dimensional representation is obtained by applying local embedding on the similarity matrix, which preserves the local intrinsic structure of data through dimensionality reduction. With these techniques, we formulate the MVJL as a joint optimization problem and derive its mathematical solution with the alternating direction method of multipliers strategy and the proximal gradient descent method. The solution, which consists of a similarity matrix and a low-dimensional representation, is ultimately integrated with spectral clustering or K-means for multi-view clustering. Extensive experimental results on real-world datasets demonstrate that MVJL achieves superior clustering performance over other state-of-the-art methods. Aiping Huang, Tiesong Zhao, Chang Wen Chen |
IEEE Trans. Image Process. | 3 |
| 2021 | Embedding Regularizer Learning for Multi-View Semi-Supervised ClassificationabstractClassification remains challenging when confronted with the existence of multi-view data with limited labels. In this paper, we propose an embedding regularizer learning scheme for multi-view semi-supervised classification (ERL-MVSC). The proposed framework integrates diversity, sparsity and consensus to dexterously manipulate multi-view data with limited labels. To encourage diversity, ERL-MVSC recasts a linear regression model to derive view-specific embedding regularizers and automatically determines their weights. This is able to tactfully incorporate complementary information of different views. To ensure sparsity, ERL-MVSC imposes$\ell _{2,1}$-norm on a fused embedding regularizer to exploit the sparse local structure of samples, thereby conveying valuable classification information and enhancing the robustness against noise/outliers. To enhance consensus, ERL-MVSC learns a shared predicted label matrix, which serves as the comment target of multi-view classification. With these techniques, we formulate ERL-MVSC as a joint optimization problem of an embedding regularizer and a predicted label matrix, which can be solved by a coordinate descent method. Extensive experimental results on real-world datasets demonstrate the effectiveness and superiority of the proposed algorithm. Aiping Huang, Zheng Wang 0007, Yannan Zheng, Tiesong Zhao, Chia-Wen Lin |
IEEE Trans. Image Process. | 4 |
| 2021 | Semi-Reference Sonar Image Quality Assessment Based on Task and Visual PerceptionabstractIn submarine and underwater detection tasks, conventional optical imaging and analysis methods are not universally applicable due to the limited penetration depth of visible light. Instead, sonar imaging has become a preferred alternative. However, the capture and transmission conditions in complicated and dynamic underwater environments inevitably lead to visual quality degradation of sonar images, which might also impede further recognition, analysis and understanding. To measure this quality decrease and provide a solid quality indicator for sonar image enhancement, we propose a task- and perception-oriented sonar image quality assessment (TPSIQA) method, in which a semi-reference (SR) approach is applied to adapt to the limited bandwidth of underwater communication channels. In particular, we exploit reduced visual features that are critical for both human perception of and object recognition in sonar images. The final quality indicator is obtained through ensemble learning, which aggregates an optimal subset of multiple base learners to achieve both high accuracy and a high generalization ability. In this way, we are able to develop a compact but generalized quality metric using a small database of sonar images. Experimental results demonstrate competitive performance, high efficiency, and strong robustness of our method compared to the latest available image quality metrics. Ke Gu 0001, Tiesong Zhao, Gangyi Jiang, Patrick Le Callet |
IEEE Trans. Multim. | 3 |
| 2020 | Perception-Lossless Codec of Haptic Data with Low DelayabstractIn multimedia services, the introduction of haptic signals provides a more immersive user experience besides of conventional audio-visual perceptions. To support synchronous streaming and display of these information, it is imperative to efficiently compress and store the haptic signals, which promotes the development and optimization of haptic codecs. In this paper, we propose an end-to-end haptic codec for high-efficiency, low-delay and perception-lossless compression of kinesthetic signal, one of two major components of haptic signals. The proposed encoder consists of amplifier, DCT, quantizer, run-length encoder and entropy encoder, while the decoder includes all counterpart modules of the encoder. In particular, all parameters of these modules are deliberately calibrated aimed at a high compression efficiency of kinesthetic information. We allow a maximal DCT length of 8 samples, in order to guarantee a maximal encoding delay of 7ms for a popular haptic simulator of 1000Hz. Incorporating the model of perception deadband, the proposed codec is capable of realizing perception-lossless kinesthetic bitsteam. Finally, we examine the proposed codec on the standard database of IEEE P1918.1.1 Haptic Codecs Task Group. Comprehensive experiments reveal that our codec outperforms its rivals with 50% bit rate reduction, improved perception quality and a negligible encoder delay. Chaoyang Zeng, Tiesong Zhao, Qian Liu 0001 |
ACM Multimedia | 2 |
| 2020 | Single image reflection removal based on structure-texture layering
Nanfeng Jiang, Yuzhen Niu, Liqun Lin, Nadir Mustafa, Tiesong Zhao |
Signal Process. Image Commun. | 6 |
| 2020 | PEA265: Perceptual Assessment of Video Compression ArtifactsabstractThe most widely used video encoders share a common hybrid coding framework that includes block-based motion estimation/compensation and block-based transform coding. Despite their high coding efficiency, the encoded videos often exhibit visually annoying artifacts, denoted as Perceivable Encoding Artifacts (PEAs), which significantly degrade the visual Quality-of-Experience (QoE) of end users. To monitor and improve visual QoE, it is crucial to develop subjective and objective measures that can identify and quantify various types of PEAs. In this work, we make the first attempt to build a large-scale subject-labeled database composed of H.265/HEVC compressed videos containing various PEAs. The database, namely the PEA265, includes 4 types of spatial PEAs (i.e. blurring, blocking, ringing and color bleeding) and 2 types of temporal PEAs (i.e. flickering and floating). Each containing at least 60,000 image or video patches with positive and negative labels. Based on the PEA265 database, we develop and optimize Convolutional Neural Networks (CNNs) to objectively recognize different types of PEAs. Experiments show that our architecture is capable of identifying the 6 types of PEAs with an accuracy over 86%. To further demonstrate its application, we explore the relationship between collected PEA intensities and subjective quality scores of compressed videos. A quality metric is consequently proposed with superior performance in terms of correlation to Mean Opinion Score (MOS) values. We believe that the PEA265 database and our findings will benefit the future development of video quality assessment methods and perceptually motivated video encoders. Liqun Lin, Shiqi Yu 0004, Tiesong Zhao, Zhou Wang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2020 | Matting-Based Residual Optimization for Structurally Consistent Image Color CorrectionabstractImage color correction aims to eliminate color differences between images, especially for color consistency in panoramic or stereoscopic images. Nowadays, the global color correction methods cannot correct local color differences, while the local color correction methods usually lead to structural inconsistency between local regions and image clarity reduction. To address these problems, we propose a matting-based residual optimization for structurally consistent image color correction (MROC). Inspired by the residual image, we formulate the image color correction problem as the optimization of a residual image between the input target image and the resulting image. The residual image is initialized and improved by the soft matting method with a closed-form solution. Besides, a data term is introduced to identify those pixels with higher color and structural consistencies and preserve these pixels during optimization. The whole computational infrastructure operates at the pixel level to correct local color differences while maintaining image clarity. Experimental results demonstrate that the performance of the proposed MROC method is superior to the state-of-the-art image color correction methods. Furthermore, the proposed matting-based residual optimization can also be incorporated in a variety of color correction methods, with enhanced outcomes justified by a group of image quality assessment metrics. Yuzhen Niu, Tiesong Zhao |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2020 | Visually Consistent Color Correction for Stereoscopic Images and VideosabstractIn stereoscopic 3D (S3D) color correction, visual inconsistency is a common problem that leads to perceptual quality degradations. In this paper, we propose an S3D image/video color correction strategy that resolves global, local, and temporal color discrepancies simultaneously. We achieve the image-based S3D color correction by three steps: a coarse-grain color correction for global color matching, a fine-grain color correction to further improve both global and local color consistencies, and a guided filtering process to guarantee the structural consistency before and after color correction. In addition, we extend the above strategy to S3D and multiview video color correction. To achieve temporal consistency between successive video frames, we develop an improved histogram matching within a sliding window on time axis. In our method, the mapping functions for each color channel change gradually following the video stream to avoid abrupt temporal changes in colors. The experimental results demonstrate that the proposed strategy outperforms the state-of-the-art color correction algorithms for images and videos. Yuzhen Niu, Xiaohua Zheng, Tiesong Zhao |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2020 | Multi-View Data Fusion Oriented Clustering via Nuclear Norm MinimizationabstractImage clustering remains challenging when handling image data from heterogeneous sources. Fusing the independent and complementary information existing in heterogeneous sources together facilitates to improve the image clustering performance. To this end, we propose a joint learning framework of multi-view image data fusion and clustering based on nuclear norm minimization. Specifically, we first formulate the problem as matrix factorization to a shared clustering indicator matrix and a representative coefficient matrix. The former is constrained with orthogonality and nonnegativity, which ensures the validation of clustering assignments. The latter is imposed with nuclear norm minimization to achieve compression of principal components for performance improvement. Then, an alternating minimization strategy is employed to efficiently decompose the multi-variable optimization problem into several small solvable sub-problems with closed-form solutions. Extensive experimental results on real-world image and video datasets demonstrate the superiority of proposed method over other state-of-the-art methods. Aiping Huang, Tiesong Zhao, Chia-Wen Lin |
IEEE Trans. Image Process. | 2 |
| 2019 | Complexity Control in the HEVC Intracoding for Industrial Video ApplicationsabstractA large number of industrial video applications are expected to work in real-time and power-constrained scenarios. For such applications, the coding complexity has a great impact on their performance. In this paper, we propose a complexity control method in the high-efficiency video coding intracoding to facilitate these video applications. The proposed method is performed on the coding tree unit (CTU) level, which consists of three steps, namely complexity estimation, complexity allocation, and prediction unit (PU) adaption. In the first step, a complexity estimation model is proposed to estimate the coding complexity of each CTU based on the sum of absolute transformed difference. Then, the complexity budget is allocated to each CTU proportionally to its estimated coding complexity. In PU adaption, only a subset of PU sizes are selected for the CTU according to its allocated complexity and prediction performance. A feedback-based error elimination scheme removes the complexity error during the encoding process. Experimental results show that the proposed method is able to adjust the complexity ratio from 100% to 20%. Meanwhile, the rate-distortion performance and complexity control accuracy of the proposed method are superior to those of the state-of-the-art methods. Jia Zhang 0002, Sam Kwong, Tiesong Zhao, Horace Ho-Shing Ip |
IEEE Trans. Ind. Informatics | 3 |
| 2018 | Complexity Control for HEVC Inter Coding Based on Two-Level Complexity Allocation and Mode SortingabstractHigh coding complexity is an obstruction when promoting the HEVC standard. In this paper, we propose a complexity control scheme for HEVC, which improves our previous work by employing a finer-grained complexity allocation scheme and mode sorting. In CTU-level complexity allocation, the time budget is allocated to each CTU proportionally to its estimated complexity. Then, the time budget for a CTU is further allocated to each CTU quadtree depth according to its probability of being the dominant depth. At last, a subset of modes are selected to reach the target complexity based on mode sorting. Compared with our previous work of which the least supported complexity ratio is 40%, the proposed scheme further reduces the lower bound of the ratio range to 20%. Jia Zhang 0002, Sam Kwong, Tiesong Zhao, Xu Wang 0006, Shiqi Wang 0001 |
ICIP | 3 |
| 2018 | A Hybrid Quality Metric for Non-Integer Image InterpolationabstractA great need of High-Resolution (HR) images has boosted the development of interpolation techniques. However, it is still a challenging task to objectively evaluate the perceptual quality of interpolated images, especially when the interpolation factor is a non-integer. To address this issue, we propose a hybrid quality metric for non-integer image interpolation that combines both reduced-reference and no-reference philosophies. To validate the proposed metric, we construct a non-integer interpolated image database and conduct a subjective user study to collect subjective opinions for each image. Experiments on the new database show that the proposed metric outperforms previous methods by a large margin. Jinling Chen, Kede Ma, Huiwen Huang, Tiesong Zhao |
QoMEX | 5 |
| 2018 | Eye-tracking-Based Quality Assessment for Image InterpolationabstractIn an Image Quality Assessment (IQA) scenario, the Human Vision System (HVS) always acts as the ultimate receiver and valuator of generated images. As an important feature of HVS, the visual attention data has been demonstrated to be able to effectively improve the performance of existing objective quality metrics. However, this feature has not yet been well explored in the IQA of image interpolation. In this paper, we conduct an eye-tracking test on an interpolated image database and investigate the impact of visual attention on IQA of image interpolation. Two visual attention models, saliency map and Region Of Interest (ROI), are then obtained from the eye-tracking data. We further incorporate these models into non-integer interpolated IQA metric and examine their performances. Experiments show that the introduction of eye-tracking features obviously improves the conventional IQA metric for non-integer image interpolation. Jinling Chen, Ludi Wu, Yisang Liu, Tiesong Zhao |
VCIP | 5 |
| 2018 | Probability-Based Intra Encoder Optimization in High Efficiency Video CodingabstractThe High Efficiency Video Coding (HEVC) adopts increasing number of intra prediction modes and Coding Unit (CU) partitions, in order to diminish its cost at required bitrates. However, its attendant computational complexity is unfavorable in lots of real-world applications. In this paper, we propose a probability-based strategy that benefits a lowcomplexity HEVC intra-encoder and also provides guidance to design low-complexity framework of next generation encoder. The proposed strategy comprises three steps: a candidate mode list initialization based on the estimated winning probabilities of all intra modes, an early termination of intra prediction based on the estimated probability of obtaining the best intra mode, and a pre-decision of Coding Unit (CU) split based on the estimated distributions of the Rate-Distortion (RD) costs. Comprehensive experiments have validated the effectiveness of the proposed algorithm, with a promising simulation performance under Common Test Conditions (CTC). Hongan Wei, Minghai Wang, Yisang Liu, Tiesong Zhao |
VCIP | 5 |
| 2018 | CTU-Level Complexity Control for High Efficiency Video CodingabstractAmong the existing video-related applications, a large proportion have requirements for the scalability of the video coding complexity, such as live video chatting and video coding on power-limited mobile devices. Hence, the complexity control algorithms, which aim to make an effective and flexible tradeoff between coding complexity and rate-distortion (RD) performance, have a great practical value. In this paper, a novel complexity control scheme for high efficiency video coding (HEVC) is proposed by dynamically adjusting the depth range for each coding tree unit (CTU). To control the complexity accurately, a statistical model is proposed to estimate the coding complexity of each CTU. Then the complexity budget is allocated to each CTU proportionally to its estimated complexity. At last, the depth range is optimized for each CTU based on the allocated complexity and the probability that contains the actual maximum depth. Our method works well even if the ratio of target complexity to full complexity drops to 40%. The experimental results show that our proposed method outperforms other four state-of-the-art methods in terms of the RD performance, and has superior complexity control accuracy and complexity control stability compared with other one-pass complexity control strategies. Jia Zhang 0002, Sam Kwong, Tiesong Zhao, Zhaoqing Pan |
IEEE Trans. Multim. | 3 |
| 2016 | Adaptive Quantization Parameter Cascading in HEVC Hierarchical CodingabstractThe state-of-the-art High Efficiency Video Coding (HEVC) standard adopts a hierarchical coding structure to improve its coding efficiency. This allows for the quantization parameter cascading (QPC) scheme that assigns quantization parameters (Qps) to different hierarchical layers in order to further improve the rate-distortion (RD) performance. However, only static QPC schemes have been suggested in HEVC test model, which are unable to fully explore the potentials of QPC. In this paper, we propose an adaptive QPC scheme for an HEVC hierarchical structure to code natural video sequences characterized by diversified textures, motions, and encoder configurations. We formulate the adaptive QPC scheme as a non-linear programming problem and solve it in a scientifically sound way with a manageable low computational overhead. The proposed model addresses a generic Qp assignment problem of video coding. Therefore, it also applies to group-of-picture-level, frame-level and coding unit-level Qp assignments. Comprehensive experiments have demonstrated that the proposed QPC scheme is able to adapt quickly to different video contents and coding configurations while achieving noticeable RD performance enhancement over all static and adaptive QPC schemes under comparison as well as HEVC default frame-level rate control. We have also made valuable observations on the distributions of adaptive QPC sets in the videos of different types of contents, which provide useful insights on how to further improve static QPC schemes. Tiesong Zhao, Zhou Wang 0001, Chang Wen Chen |
IEEE Trans. Image Process. | 1 |
| 2015 | Objective Quality Assessment for Color-to-Gray Image ConversionabstractColor-to-gray (C2G) image conversion is the process of transforming a color image into a grayscale one. Despite its wide usage in real-world applications, little work has been dedicated to compare the performance of C2G conversion algorithms. Subjective evaluation is reliable but is also inconvenient and time consuming. Here, we make one of the first attempts to develop an objective quality model that automatically predicts the perceived quality of C2G converted images. Inspired by the philosophy of the structural similarity index, we propose a C2G structural similarity (C2G-SSIM) index, which evaluates the luminance, contrast, and structure similarities between the reference color image and the C2G converted image. The three components are then combined depending on image type to yield an overall quality measure. Experimental results show that the proposed C2G-SSIM index has close agreement with subjective rankings and significantly outperforms existing objective quality metrics for C2G conversion. To explore the potentials of C2G-SSIM, we further demonstrate its use in two applications: 1) automatic parameter tuning for C2G conversion algorithms and 2) adaptive fusion of C2G converted images. Kede Ma, Tiesong Zhao, Kai Zeng 0003, Zhou Wang 0001 |
IEEE Trans. Image Process. | 2 |
| 2013 | Fast Mode Decision Based on Optimal Stopping Theory for Multiview Video Coding
Hanli Wang, Yue Heng, Tiesong Zhao, Bo Xiao 0005 |
MMM (2) | 3 |
| 2013 | Multiview Coding Mode Decision With Hybrid Optimal Stopping ModelabstractIn a generic decision process, optimal stopping theory aims to achieve a good tradeoff between decision performance and time consumed, with the advantages of theoretical decision-making and predictable decision performance. In this paper, optimal stopping theory is employed to develop an effective hybrid model for the mode decision problem, which aims to theoretically achieve a good tradeoff between the two interrelated measurements in mode decision, as computational complexity reduction and rate-distortion degradation. The proposed hybrid model is implemented and examined with a multiview encoder. To support the model and further promote coding performance, the multiview coding mode characteristics, including predicted mode probability and estimated coding time, are jointly investigated with inter-view correlations. Exhaustive experimental results with a wide range of video resolutions reveal the efficiency and robustness of our method, with high decision accuracy, negligible computational overhead, and almost intact rate-distortion performance compared to the original encoder. Tiesong Zhao, Sam Kwong, Hanli Wang, Zhou Wang 0001, Zhaoqing Pan, C.-C. Jay Kuo |
IEEE Trans. Image Process. | 1 |
| 2012 | H.264/SVC Mode Decision Based on Optimal Stopping TheoryabstractFast mode decision algorithms have been widely used in the video encoder implementation to reduce encoding complexity yet without much sacrifice in the coding performance. Optimal stopping theory, which addresses early termination for a generic class of decision problems, is adopted in this paper to achieve fast mode decision for the H.264/Scalable Video Coding standard. A constrained model is developed with optimal stopping, and the solutions to this model are employed to initialize the candidate mode list and predict the early termination. Comprehensive simulation results are conducted to demonstrate that the proposed method strikes a good balance between low encoding complexity and high coding efficiency. Tiesong Zhao, Sam Kwong, Hanli Wang, C.-C. Jay Kuo |
IEEE Trans. Image Process. | 1 |
| 2011 | Priority pyramid based bit allocation for multiview video codingabstractIn Multivew Video Coding (MVC), Hierarchial B Pictures (HBP) structure is adopted to remove the redundancy of multiview videos. It makes the rate control of MVC more difficult with respect to accurate bit control and good compression efficiency. The existing rate control algorithms of MVC are based on those of monoview video coding standards, which didn't exploit the correlations of multiview video frames fully. In this paper, a Group of pictures from multiple views (called GoGOP) are rearranged in the form of a pyramid. The top layer of the pyramid containing anchor frames receives the highest priority in bits consuming. And the next layer is inferior to it but superior to others in bits consuming. Secondly, a matrix of weighting factors is further introduced to perform the bit allocation for MVC. Thirdly, the parameter updating is handled in a vector operation, which is somewhat robust and has a low level of computational complexity. The experimental results demonstrate that our proposed bit allocation cooperating with the conventional rate-distortion (R-D) model in H.264/AVC is efficient in rate control of MVC. The coding performance of the proposed algorithm is comparable to that of hierarchical quantization scheme (HQS) of MVC. Meanwhile, a small bit control error is obtained by our algorithm. Long Xu 0001, Sam Kwong, Tiesong Zhao, Yu Zhou 0015 |
VCIP | 3 |
| 2011 | Hierarchical B-picture mode decision in H.264/SVC
Tiesong Zhao, Hanli Wang, Sam Kwong, C.-C. Jay Kuo, Wolfgang A. Halang |
J. Vis. Commun. Image Represent. | 1 |
| 2011 | Rate Control Optimization for Temporal-Layer Scalable Video CodingabstractA novel frame-level rate control (RC) algorithm is presented in this paper for temporal scalability of scalable video coding. First, by introducing a linear quality dependency model, the quality dependency between a coding frame and its references is investigated for the hierarchical B-picture prediction structure. Second, linear rate-quantization (R-Q) and distortion-quantization (D-Q) models are introduced based on different characteristics of temporal layers. Third, according to the proposed quality dependency model and R-Q and D-Q models for each temporal layer, adaptive weighting factors are derived to allocate bits efficiently among temporal layers. Experimental results on not only traditional quarter common intermediate format/common intermediate format but also standard definition and high definition sequences demonstrate that the proposed algorithm achieves excellent coding efficiency as compared to other benchmark RC schemes. Sudeng Hu, Hanli Wang, Sam Kwong, Tiesong Zhao, C.-C. Jay Kuo |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2010 | Fast inter-layer mode decision in Scalable Video CodingabstractThe recently developed video encoding standard Scalable Video Coding (SVC) is designed as the extension of H.264/AVC. By employing new techniques such as hierarchal layer encoding, SVC could produce bit streams with different frame rates, bit rates and spatial resolutions; but meanwhile, the computational complexity of video encoder is also increased, which makes optimization of SVC encoding a necessity. In this paper, a novel algorithm is proposed, which utilizes the encoding mode of Base Layer (BL) macroblocks (MBs) to initialize the potential optimal candidate encoding mode list for Enhancement Layer (EL) MBs, by mapping encoding modes onto a Two-Dimensional (2D) map. Experimental results demonstrate that, the proposed algorithm could achieve significantly better and more robust time saving than two recent works while maintaining the video performances. Tiesong Zhao, Hanli Wang, Sam Kwong |
ICIP | 1 |
| 2010 | Probability-based coding mode prediction for H.264/AVCabstractFast mode decision (FMD) algorithm aims to reduce the computational complexity of video encoders, especially for H.264/AVC, and meanwhile keep the coding efficiency. It is necessary for coding high resolution sequences, such as standard definition (SD) or high definition (HD) sequences. In this paper, an efficient algorithm is proposed to speed up inter-mode decision for H.264/AVC. Firstly, for a macroblock (MB) being coded, with modes of neighboring blocks, a probability model is utilized to build up an optimal encoding mode list; after that, rate-distortion (RD) cost-based early termination is employed to skip unnecessary inter-modes. Simulation results demonstrate that, the proposed algorithm could achieve significant encoding time saving for QCIF/CIF and SD/HD sequences, meanwhile with almost the same RD performance as the original encoder. Tiesong Zhao, Hanli Wang, Sam Kwong, Sudeng Hu |
ICIP | 1 |
| 2010 | Frame level rate control for H.264/AVC with novel Rate-Quantization modelabstractIn this paper, a frame level rate control algorithm is proposed with a novel Rate-Quantization (R-Q) model for H.264/AVC. Firstly, a two-stage rate control scheme is adopted to decouple the inter-dependency between Rate Distortion Optimization (RDO) and rate control. Secondly, in order to predict the frame complexity accurately, instead of the Mean Absolute Difference (MAD) of the residual signal, bits information in the RDO-based mode decision process is employed to predict the frame complexity. Thirdly, a self-adaptive exponential R-Q model is proposed for rate control. Experimental results reveal that the proposed R-Q model can estimate the actual output bits very well, and the novel rate control scheme has excellent performance both in bit rate accuracy and coding efficiency as compared to JVT-W043 and the FixedQp tool in the Joint Scalable Video Model reference software. Sudeng Hu, Hanli Wang, Sam Kwong, Tiesong Zhao |
ICME | 4 |
| 2010 | Fast Mode Decision Based on Mode AdaptationabstractThis paper proposes an efficient algorithm for fast mode decision in H.264/advanced video coding by adaptively predicting the optimal mode for each macroblock (MB) to be coded. Firstly, encoding modes are projected as points onto a 2-D map, and an optimal 2-D point of the MB to be coded is predicted based on the encoding information of spatial-temporal neighboring blocks. Then, a priority-based mode candidate list with a descending order to be the best mode is constructed based on the optimal 2-D point. Finally, mode decision is performed according to the priority-based mode candidate list in the checking order, from the most important mode to the least one, with early termination conditions. Extensive experimental results demonstrate that the proposed algorithm is superior to three recent fast mode decision algorithms, with the entire encoding time being reduced by about 60% for quarter common intermediate format/common intermediate format/standard-definition sequences on average and the rate distortion performance being kept almost intact. Tiesong Zhao, Hanli Wang, Sam Kwong, C.-C. Jay Kuo |
IEEE Trans. Circuits Syst. Video Technol. | 1 |