EDBT 2026 Demo / reviewers in the wild / expert
Yucheng Zhu
dblp:180/2618
· DBLP profile ↗
48ranked-venue papers
12as first author
34since 2021 · last 2026
0000-0002-3069-060XORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 37 · 10 first-author · 27 since 2021Artificial intelligence and machine learning · 5 · 4 since 2021Systems, architecture and hardware · 4 · 2 first-authorComputer networks · 4 · 2 first-author · 4 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 first-authorApplied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Market-Bench: Benchmarking Large Language Models on Economic and Trade CompetitionabstractYushuo Zheng, Huiyu Duan, Zicheng Zhang, Yucheng Zhu, Xiongkuo Min, Guangtao Zhai. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Yushuo Zheng, Huiyu Duan, Yucheng Zhu, Xiongkuo Min, Guangtao Zhai |
ACL (1) | 4 |
| 2026 | D-DSQN: A Coordinated Botnet Suppression Mechanism for Social IoT Networks
Shigen Shen, Yucheng Zhu, Xinmin Cheng, Zhaoxi Fang, Tian Wang 0001, Xiao Zhi Gao 0001 |
IEEE Internet Things J. | 2 |
| 2026 | DHQA-4D: A large-scale dataset and LMM-based metric for dynamic 4D digital human quality assessment
Sijing Wu, Yucheng Zhu, Huiyu Duan, Wei Sun 0029, Xiongkuo Min, Guangtao Zhai |
Pattern Recognit. | 3 |
| 2026 | AGHI-QA: A Subjective-Aligned Dataset and Metric for AI-Generated Human ImagesabstractThe rapid development of text-to-image (T2I) generation approaches has attracted extensive interest in evaluating the quality of generated images, leading to the development of various quality assessment methods for general-purpose T2I outputs. However, existing image quality assessment (IQA) methods are limited to providing global quality scores, failing to deliver fine-grained perceptual evaluations for structurally complex subjects like humans, which is a critical challenge considering the frequent anatomical and textural distortions in AI-generated human images (AGHIs). To address this gap, we introduce AGHI-QA, a large-scale benchmark specifically designed for quality assessment of AGHIs. The dataset comprises 4, 000 images generated from 400 carefully crafted text prompts using 10 state-of-the-art T2I models. We conduct a systematic subjective study to collect multidimensional annotations, including perceptual quality scores, text-image correspondence scores, visible and distorted body part labels. Based on AGHI-QA, we evaluate the strengths and weaknesses of current T2I methods in generating human images from multiple dimensions. Furthermore, we propose AGHI-Assessor, a novel quality metric that integrates the large multimodal model (LMM) with domain-specific human features for precise quality prediction and identification of visible and distorted body parts in AGHIs. Extensive experimental results demonstrate that AGHI-Assessor showcases state-of-the-art performance, significantly outperforming existing IQA methods in multidimensional quality assessment and surpassing leading LMMs in detecting structural distortions in AGHIs. Sijing Wu, Wei Sun 0029, Yucheng Zhu, Huiyu Duan, Xiongkuo Min, Guangtao Zhai |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2026 | Future Fixation Sequence Prediction for Audio-Visual 360° VideosabstractFuture fixation sequence prediction plays a crucial role in various aspects of virtual reality content production, transmission, rendering, and display. Accurate prediction of future fixation sequence can significantly enhance the quality of user experience, particularly in resource-constrained scenarios. In this paper, we present a novel framework for predicting future fixation sequence and achieves state-of-the-art performance. Specifically, the anti-projection-distortion FoV patch extraction algorithm is proposed to mitigate projection distortions. A comprehensive contextual representation is then constructed by integrating multiple data sources, including visual and audio information, historical fixation sequence, user identity, timestamp, and positional embeddings. The transformer-based predictor is proposed to perform the future fixation sequence prediction based on the integrated contextual representations. Additionally, we propose a framework that effectively utilizes saliency information as supervision and conduct saliency contrastive distillation during the training phase, eliminating the need for saliency data during inference. Overall, by integrating anti-projection-distortion and multimodal representations, along with key embeddings, a dedicated predictor, and contrastive distillation, our approach is designed to accurately predict future fixation sequences. Extensive experiments validate the effectiveness of our framework, demonstrating its superior performance in fixation prediction tasks. Yucheng Zhu, Guangtao Zhai, Xiongkuo Min, Huiyu Duan, Xiaokang Yang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2026 | SingingHead: A Large-Scale 4D Dataset for Singing Head AnimationabstractSinging, as a common facial movement second only to talking, can be regarded as a universal language across ethnicities and cultures, plays an important role in emotional communication, art, and entertainment. However, it is often overlooked in the field of audio-driven 3D facial animation due to the lack of singing head datasets and the domain gap between singing and talking in rhythm and amplitude. To this end, we collect a large-scale high-quality multi-modal singing head dataset,SingingHead, which consists of more than 27 hours of synchronized singing video, 3D facial motion, singing audio, and background music from 76 individuals and 8 types of music. Along with the SingingHead dataset, we benchmark existing audio-driven 3D facial animation methods and 2D talking head methods on the singing task. Existing 3D facial animation methods and 2D talking head methods fail to produce satisfactory singing results. Focusing on the 3D singing head animation, we first utilize the proposed singing-specific dataset to retrain the 3D facial animation methods, resulting in substantial performance improvements. Besides, considering the absence of background music and the slow generation speed of existing methods, we propose a simple but efficient non-autoregressive VAE-based framework with background music as an input signal to generate diverse and accurate 3D singing facial motions in real time. Extensive experiments demonstrate the significance of the SingingHead dataset in promoting the development of singing head animation. The dataset is released for research purposes at:https://wsj-sjtu.github.io/SingingHead/. Sijing Wu, Weitian Zhang, Jun Jia, Yucheng Zhu, Yichao Yan, Guangtao Zhai, Xiaokang Yang 0001 |
IEEE Trans. Multim. | 5 |
| 2025 | ESVQA: Perceptual Quality Assessment of Egocentric Spatial VideosabstractWith the rapid development of eXtended Reality (XR), egocentric spatial shooting and display technologies have further enhanced immersion and engagement for users, delivering more captivating and interactive experiences. Assessing the quality of experience (QoE) of egocentric spatial videos is crucial to ensure a high-quality viewing experience. However, the corresponding research is still lacking. In this paper, we use the concept of embodied experience to highlight this more immersive experience and study the new problem, i.e., embodied perceptual quality assessment for egocentric spatial videos. Specifically, we introduce the first Egocentric Spatial Video Quality Assessment Database (ESVQAD), which comprises 600 egocentric spatial videos captured using the Apple Vision Pro and their corresponding mean opinion scores (MOSs). Furthermore, we propose a novel multi-dimensional binocular feature fusion model, termed ESVQAnet, which integrates binocular spatial, motion, and semantic features to predict the overall perceptual quality. Experimental results demonstrate the ESVQAnet significantly outperforms 16 state-of-the-art VQA models on the embodied perceptual quality assessment task, and exhibits strong generalization capability on traditional VQA tasks. The database and code are available at https://github.com/IntMeGroup/ESVQA. Xilei Zhu, Huiyu Duan, Yucheng Zhu, Xiongkuo Min, Guangtao Zhai, Patrick Le Callet |
ICME | 4 |
| 2025 | HVEval: Towards Unified Evaluation of Human-Centric Video Generation and UnderstandingabstractHuman-centric videos play a significant role in the pervasive video content of modern life. However, the capabilities of text-to-video (T2V) generation models and video-to-text (V2T) understanding models for human-centric videos remain largely unexplored. To this end, we present HVEval, the first comprehensive evaluation dataset focusing on human-centric videos, which consists of 20,000 videos, 60k MOS annotations across 3 dimensions (i.e., spatial quality, temporal quality, and text-video correspondence), and 20k category-specific Q&A pairs. Based on the HVEval dataset, this paper aims to answer three questions: (1) can today's T2V models effectively generate human-centric videos following the given prompts? (2) how effective are today's V2T LMMs in understanding and evaluating human-centric videos? (3) are current VQA metrics good enough for evaluating human-centric videos? Comprehensive evaluations of 24 T2V models, 20 LMMs, and 18 VQA metrics reveal their limitations in fine-grained text-controlled generation and human-aligned perception and understanding, highlighting the significant potential of our dataset and benchmarks to advance research in human-centric video generation and understanding. Sijing Wu, Huiyu Duan, Yanwei Jiang, Yucheng Zhu, Guangtao Zhai |
ACM Multimedia | 5 |
| 2025 | Omni2: Unifying Omnidirectional Image Generation and Editing in an Omni Modelabstract360° omnidirectional images (ODIs) have gained considerable attention recently, and are widely used in various virtual reality (VR) and augmented reality (AR) applications. However, capturing such images is expensive and requires specialized equipment, making ODI synthesis increasingly important. While common 2D image generation and editing methods are rapidly advancing, these models struggle to deliver satisfactory results when generating or editing ODIs due to the unique format and broad 360° Field-of-View (FoV) of ODIs. To bridge this gap, we construct Any2Omni , the first comprehensive ODI generation-editing dataset comprises 60,000+ training data covering diverse input conditions and up to 9 ODI generation and editing tasks. Built upon Any2Omni, we propose an Omni model for Omni-directional image generation and editing ( Omni 2), with the capability of handling various ODI generation and editing tasks under diverse input conditions using one model. Extensive experiments demonstrate the superiority and effectiveness of the proposed Omni2 model for both the ODI generation and editing tasks. Both the Any2Omni dataset and the Omni2 model are publicly available at: https://github.com/IntMeGroup/Omni2. Huiyu Duan, Yucheng Zhu, Xiaohong Liu 0001, Lu Liu 0005, Zitong Xu, Guangji Ma, Xiongkuo Min, Guangtao Zhai, Patrick Le Callet |
ACM Multimedia | 3 |
| 2025 | Time-Smooth Wireless Transmission of Probabilistic Slicing VR 360 Video in MISO-OFDM SystemsabstractThe multiple-input and single-output (MISO)-orthogonal frequency-division multiplexing (OFDM) systems afford low latency and high reliability for virtual reality (VR) 360 video in multi-user scenarios. Motivated by the goal of maintaining time-smoothness while holding acceptably low complexity, a crucial factor in VR video transmission, we conduct a comprehensive study that integrates the characteristics of VR video with the strategies for subcarrier assignment and power allocation. By analyzing the pre-transmitted tile-segments, the missing tile-segments, and the video frame structure, we propose two probabilistic slicing schemes (PSPs) to minimize the size of required tile-segments of VR video scenes. In time-smoothness maximization, the desired discrete encoding rate set, discrete subcarrier assignment, continuous power allocation, and fixed total power constraint make it a challenging mixed-integer nonlinear programming (MINLP) problem. Unlike the straightforward relaxation-recovery method, we firstly prove that a near-optimal recovered encoding rate is the discrete value closest to the optimal relaxed-continuous encoding rate. We then propose a Two-step Encoding Rate Maximization (TERM) method, including the relaxed-continuous sum-rate maximization and the discrete encoding rate recovery, to achieve the near-optimal subcarrier assignment and the power allocation with low complexity. Simulation results on real-world VR video dataset validate that the two PSPs can effectively minimize the number of transmitted tile-segments. The proposed TERM with PSPs can maintain time-smoothness of VR 360 video with an acceptably low level of complexity in MISO-OFDM systems. Guangtao Zhai, Yongpeng Wu 0001, Xiongkuo Min, Biqian Feng, Yucheng Zhu, Wenjun Zhang 0001 |
IEEE Trans. Commun. | 6 |
| 2025 | Subjective and Objective Audio-Visual Quality Assessment for Omnidirectional VideosabstractVirtual Reality (VR) has attracted widespread attention in recent years due to its capability to create immersive experiences by presenting multi-modal information to users. Omnidirectional videos (ODVs), as a prominent component of VR content, are essential across diverse applications. This necessitates service providers to monitor and optimize the quality of ODVs throughout the filming, encoding, decoding, and transmission stages to ensure a high-quality viewing experience. However, most existing Quality of Experience (QoE) studies for ODVs only focus on the visual quality, while overlooking the impact of the audio modality on perceptual quality. This paper presents a comprehensive study of omnidirectional audio-visual quality assessment (OD-AVQA) from both subjective and objective perspectives. Specifically, we first establish a large-scale audio-visual quality assessment database for ODVs named OAVQAD+, which includes 625 distorted omnidirectional audio-visual sequences derived from 25 pristine ODVs, and the corresponding collected mean opinion scores (MOSs) for the QoE of these ODVs. This contributes to the largest database for assessing the audio-visual quality of ODVs. To advance the fields of objective OD-AVQA, we construct a benchmark that includes three types of benchmark models. Type I and Type II models integrate well-known video quality assessment (VQA) and audio quality assessment (AQA) methods using support vector regression (SVR) and multi-layer perceptron (MLP), respectively, while Type III consists of AVQA models specifically designed for traditional 2D audio-visual sequences. We also propose a novel Omnidirectional Audio-Visual quality assessment Network (OmniAVNet) that integrates quality-aware audio, visual, and motion features to predict overall audio-visual quality for ODVs effectively, which supports both full-reference (FR) and no-reference (NR) assessment. Extensive experimental results demonstrate that OmniAVNet outperforms the aforementioned benchmark OD-AVQA models on two OD-AVQA databases, and shows great performance on one omnidirectional VQA database. The database and code are available at https://github.com/IntMeGroup/OmniAVNet. Xilei Zhu, Huiyu Duan, Yuqin Cao, Yucheng Zhu, Jing Liu 0002, Xiongkuo Min, Guangtao Zhai, Patrick Le Callet |
IEEE Trans. Image Process. | 4 |
| 2025 | How Does Audio Influence Visual Attention in Omnidirectional Videos? Database and ModelabstractUnderstanding and predicting viewer attention in omnidirectional videos (ODVs) is crucial for enhancing user engagement in virtual and augmented reality applications. Although both audio and visual modalities are essential for saliency prediction in ODVs, the joint exploitation of these two modalities has been limited, primarily due to the absence of large-scale audio-visual saliency databases and comprehensive analyses. This paper comprehensively investigates audio-visual attention in ODVs from both subjective and objective perspectives. Specifically, we first introduce a new audio-visual saliency database for omnidirectional videos, termed AVS-ODV database, containing 162 ODVs and corresponding eye movement data collected from 60 subjects under three audio modes including mute, mono, and ambisonics. Based on the constructed AVS-ODV database, we perform an in-depth analysis of how audio influences visual attention in ODVs. To advance the research on audio-visual saliency prediction for ODVs, we further establish a new benchmark based on the AVS-ODV database by testing numerous state-of-the-art saliency models, including visual-only models and audio-visual models. In addition, given the limitations of current models, we propose an innovative omnidirectional audio-visual saliency prediction network (OmniAVS), which is built based on the U-Net architecture, and hierarchically fuses audio and visual features from the multimodal aligned embedding space. Extensive experimental results demonstrate that the proposed OmniAVS model outperforms other state-of-the-art models on both ODV AVS prediction and traditional AVS prediction tasks. The AVS-ODV database and the OmniAVS model are available at: https://github.com/IntMeGroup/AVS-ODV. Huiyu Duan, Kaiwei Zhang, Yucheng Zhu, Xilei Zhu, Xiongkuo Min, Guangtao Zhai |
IEEE Trans. Image Process. | 4 |
| 2025 | Evaluating Point Cloud From Moving Camera Videos: A No-Reference MetricabstractPoint cloud is one of the most widely used digital representation formats for three-dimensional (3D) contents, the visual quality of which may suffer from noise and geometric shift distortions during the production procedure as well as compression and downsampling distortions during the transmission process. To tackle the challenge of point cloud quality assessment (PCQA), many PCQA methods have been proposed to evaluate the visual quality levels of point clouds by assessing the rendered static 2D projections. Although such projectionbased PCQA methods achieve competitive performance with the assistance of mature image quality assessment (IQA) methods, they neglect that the 3D model is also perceived in a dynamic viewing manner, where the viewpoint is continually changed according to the feedback of the rendering device. Therefore, in this paper, we evaluate the point clouds from moving camera videos and explore the way of dealing with PCQA tasks via using video quality assessment (VQA) methods. First, we generate the captured videos by rotating the camera around the point clouds through several circular pathways. Then we extract both spatial and temporal quality-aware features from the selected key frames and the video clips through using trainable 2D-CNN and pretrained 3D-CNN models respectively. Finally, the visual quality of point clouds is represented by the video quality values. The experimental results reveal that the proposed method is effective for predicting the visual quality levels of the point clouds and even competitive with full-reference (FR) PCQA methods. The ablation studies further verify the rationality of the proposed framework and confirm the contributions made by the qualityaware features extracted via the dynamic viewing manner. The code is available athttps://github.com/zzc-1998/VQA_PC. Wei Sun 0029, Yucheng Zhu, Xiongkuo Min, Wei Wu 0002, Ying Chen 0011, Guangtao Zhai |
IEEE Trans. Multim. | 3 |
| 2024 | SG-JND: Semantic-Guided Just Noticeable Distortion Predictor for Image CompressionabstractJust noticeable distortion (JND), representing the threshold of distortion in an image that is minimally perceptible to the human visual system (HVS), is crucial for image compression algorithms to achieve a trade-off between transmission bit rate and image quality. However, traditional JND prediction methods only rely on pixel-level or sub-band level features, lacking the ability to capture the impact of image content on JND. To bridge this gap, we propose a Semantic-Guided JND (SG-JND) network to leverage semantic information for JND prediction. In particular, SG-JND consists of three essential modules: the image preprocessing module extracts semantic-level patches from images, the feature extraction module extracts multi-layer features by utilizing the cross-scale attention layers, and the JND prediction module regresses the extracted features into the final JND value. Experimental results show that SG-JND achieves the state-of-the-art performance on two publicly available JND datasets, which demonstrates the effectiveness of SG-JND and highlight the significance of incorporating semantic information in JND assessment. Linhan Cao, Wei Sun 0029, Xiongkuo Min, Jun Jia, Zijian Chen 0001, Yucheng Zhu, Lizhou Liu, Qiubo Chen, Guangtao Zhai |
ICIP | 7 |
| 2024 | AIGCOIQA2024: Perceptual Quality Assessment of AI Generated Omnidirectional Imagesabstract[?]In recent years, the rapid advancement of Artificial Intelligence Generated Content (AIGC) has attracted widespread attention. Among the AIGC, AI generated omnidirectional images hold significant potential for Virtual Reality (VR) and Augmented Reality (AR) applications, hence omnidirectional AIGC techniques have also been widely studied. AI-generated omnidirectional images exhibit unique distortions compared to natural omnidirectional images, however, there is no dedicated Image Quality Assessment (IQA) criteria for assessing them. This study addresses this gap by establishing a large-scale AI generated omnidirectional image IQA database named AIGCOIQA2024 and constructing a comprehensive benchmark. We first generate 300 omnidirectional images based on 5 AIGC models utilizing 25 text prompts. A subjective IQA experiment is conducted subsequently to assess human visual preferences from three perspectives including quality, comfortability, and correspondence. Finally, we conduct a benchmark experiment to evaluate the performance of state-of-the-art IQA models on our database. The AIGCOIQA2024 database is released to facilitate future research on https://github.com/IntMeGroup/AIGCOIQA. Huiyu Duan, Yucheng Zhu, Xiaohong Liu 0001, Menghan Hu, Xiongkuo Min, Guangtao Zhai, Patrick Le Callet |
ICIP | 4 |
| 2024 | MVBind: Self-Supervised Music Recommendation for Videos via Embedding Space BindingabstractRecent years have witnessed the rapid development of short videos, which usually contain both visual and audio modalities. Background music is important to the short videos, which can significantly influence the emotions of the viewers. However, at present, the background music of short videos is generally chosen by the video producer, and there is a lack of automatic music recommendation methods for short videos. This paper introduces MVBind, an innovative Music-Video embedding space Binding model for cross-modal retrieval. MVBind operates as a self-supervised approach, acquiring inherent knowledge of intermodal relationships directly from data, without the need of manual annotations. Additionally, to compensate the lack of a corresponding musical-visual pair dataset for short videos, we construct a dataset, SVM-10K (Short Video with Music-10K), which mainly consists of meticulously selected short videos. On this dataset, MVBind manifests significantly improved performance compared to other baseline methods. The database and code are available at: https://github.com/IntMeGroup/MVBind. Jiajie Teng, Huiyu Duan, Yucheng Zhu, Sijing Wu, Guangtao Zhai |
VCIP | 3 |
| 2024 | Perceptual video quality assessment: a surveyabstractAbstract Perceptual video quality assessment plays a vital role in the field of video processing due to the existence of quality degradations introduced in various stages of video signal acquisition, compression, transmission and display. With the advancement of Internet communication and cloud service technology, video content and traffic are growing exponentially, which further emphasizes the requirement for accurate and rapid assessment of video quality. Therefore, numerous subjective and objective video quality assessment studies have been conducted over the past two decades for both generic videos and specific videos such as streaming, user-generated content, 3D, virtual and augmented reality, high dynamic range, high frame rate, audio-visual, etc. This survey provides an up-to-date and comprehensive review of these video quality assessment studies. Specifically, we first review the subjective video quality assessment methodologies and databases, which are necessary for validating the performance of video quality metrics. Second, the objective video quality assessment measures for general purposes are categorized and surveyed according to the methodologies utilized in the quality measures. Third, we overview the objective video quality assessment measures for specific applications and emerging topics. Finally, the performance of the state-of-the-art video quality assessment measures is compared and analyzed. This survey provides a systematic overview of both classical works and recent progress in the realm of video quality assessment, which can help other researchers quickly access the field and conduct relevant research. Xiongkuo Min, Huiyu Duan, Wei Sun 0029, Yucheng Zhu, Guangtao Zhai |
Sci. China Inf. Sci. | 4 |
| 2024 | Blind Image Quality Assessment: A Fuzzy Neural Network for Opinion Score Distribution PredictionabstractImage quality assessment (IQA) has always been a popular research topic. There have been many methods proposed for predicting image quality, also known as the mean opinion score (MOS). However, it is worth noting that different people may assign different opinion scores to the same image. Image quality described by all subjective opinion scores can express rich subjective information about the image, such as diversity and uncertainty, which cannot be accurately described by a single MOS. Therefore, this paper proposes a fuzzy neural network to predict the opinion score distribution (OSD) of image quality. The fuzzy neural network includes three sub-networks: a feature extraction network, a feature fuzzification network, and a fuzzy learning network. First, a novel network is designed to extract image features. The extracted features are then fuzzified by fuzzy theory to model the epistemic uncertainty in the feature extraction process. Finally, the OSD of image quality is predicted using the fuzzy learning network by learning the mapping from fuzzy features to fuzzy uncertainty when rating image quality. In addition, to train the proposed fuzzy neural network, we employ a new loss function based on the quantile and the cumulative density function. We experimentally validate the feasibility and superiority of the proposed method in two aspects. On the one hand, we demonstrate the performance of the proposed method in predicting the OSD of image quality on the SJTU IQSD and KonIQ-10K databases. On the other hand, we also prove the feasibility of the proposed method in predicting the MOS of image quality on several popular IQA databases, including CSIQ, TID2013, LIVE MD, and LIVE Challenge. Xiongkuo Min, Yucheng Zhu, Xiao-Ping Zhang 0002, Guangtao Zhai |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2023 | Audio-Visual Saliency for Omnidirectional Videos
Xilei Zhu, Huiyu Duan, Kaiwei Zhang, Yucheng Zhu, Li Chen 0021, Xiongkuo Min, Guangtao Zhai |
ICIG (5) | 6 |
| 2023 | EEP-3DQA: Efficient and Effective Projection-Based 3D Model Quality AssessmentabstractCurrently, great numbers of efforts have been put into improving the effectiveness of 3D model quality assessment (3DQA) methods. However, little attention has been paid to the computational costs and inference time, which is also important for practical applications. Unlike 2D media, 3D models are represented by more complicated and irregular digital formats, such as point cloud and mesh. Thus it is normally difficult to perform an efficient module to extract quality-aware features of 3D models. In this paper, we address this problem from the aspect of projection-based 3DQA and develop a no-reference (NR) Efficient and Effective Projection-based 3D Model Quality Assessment (EEP-3DQA) method. The input projection images of EEP-3DQA are randomly sampled from the six perpendicular viewpoints of the 3D model and are further spatially downsampled by the grid-mini patch sampling strategy. Further, the lightweight Swin-Transformer tiny is utilized as the backbone to extract the quality-aware features. Finally, the proposed EEP-3DQA and EEP-3DQA-t (tiny version) achieve the best performance than the existing state-of-the-art NR-3DQA methods and even outperforms most full-reference (FR) 3DQA methods on the point cloud and mesh quality assessment databases while consuming less inference time than the compared 3DQA methods. Wei Sun 0029, Yingjie Zhou 0003, Wei Lu 0021, Yucheng Zhu, Xiongkuo Min, Guangtao Zhai |
ICME | 5 |
| 2023 | Blind Image Quality Assessment via Cross-View ConsistencyabstractImage quality assessment (IQA) is very important for both end-users and service-providers since a high-quality image can significantly improve the user's quality of experience (QoE). Most existing blind image quality assessment (BIQA) models were developed for synthetically distorted images, however, they perform poorly on in-the-wild images, which are widely existed in various practical applications. In this paper, a BIQA model is proposed that consists of a desirable self-supervised feature learning approach to mitigate the data shortage problem and learn comprehensive feature representations, and a self-attention-based feature fusion module to introduce self-attention mechanism. We develop the image quality assessment model under the framework of contrastive learning with multi views. Since human visual system perceives signals through multiple channels, the most important visual information should exist among all views of the channels. So we design the cross-view consistent information mining (CVC-IM) module to extract compact mutual information between different views. Color information and pseudo-reference image (PRI) of different distortion types are employed to formulate rich feature embeddings and preserve the quality-aware fidelity of learned representations. We employ the Transformer as the self-attention-based architecture to integrate feature embeddings. Extensive experiments show that our model achieves remarkable image quality assessment results on in-the-wild IQA datasets. Yucheng Zhu, Wei Sun 0029, Xiongkuo Min, Guangtao Zhai, Xiaokang Yang 0001 |
IEEE Trans. Multim. | 1 |
| 2023 | Toward Visual Behavior and Attention Understanding for Augmented 360 Degree VideosabstractAugmented reality (AR) overlays digital content onto reality. In an AR system, correct and precise estimations of user visual fixations and head movements can enhance the quality of experience by allocating more computational resources for analyzing, rendering, and 3D registration on the areas of interest. However, there is inadequate research to help in understanding the visual explorations of the users when using an AR system or modeling AR visual attention. To bridge the gap between the saliency prediction on real-world scenes and on scenes augmented by virtual information, we construct the ARVR saliency dataset. The virtual reality (VR) technique is employed to simulate the real-world. Annotations of object recognition and tracking as augmented contents are blended into omnidirectional videos. The saliency annotations of head and eye movements for both original and augmented videos are collected and together constitute the ARVR dataset. We also design a model that is capable of solving the saliency prediction problem in AR. Local block images are extracted to simulate the viewport and offset the projection distortion. Conspicuous visual cues in the local block images are extracted to constitute the spatial features. The optical flow information is estimated as an important temporal feature. We also consider the interplay between virtual information and reality. The composition of the augmentation information is distinguished, and the joint effects of adversarial augmentation and complementary augmentation are estimated. The Markov chain is constructed with block images as graph nodes. In the determination of the edge weights, both the characteristics of the viewing behaviors and the visual saliency mechanisms are considered. The order of importance for block images is estimated through the state of equilibrium of the Markov chain. Extensive experiments are conducted to demonstrate the effectiveness of the proposed method. Yucheng Zhu, Xiongkuo Min, Dandan Zhu 0001, Guangtao Zhai, Xiaokang Yang 0001, Wenjun Zhang 0001, Ke Gu 0001, Jiantao Zhou 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2022 | How Sound Affects Visual Attention in Omnidirectional VideosabstractIn this paper, we propose a new audio-visual attention dataset that records eye movement for omnidirectional videos with and without sound. We classify the videos into three types according to the number of salient objects and sound sources and analyze the impact of sound on visual attention distribution and inter-observer consistency of viewing area in different types of videos. From the quantitative and qualitative analysis, we find that visual attention will be drawn to and concentrated on the sound source with the presence of sound, especially when there are several visually salient objects and only one sound source. Also, the sound will enhance the consistency of observation areas among viewers to some extent. For more investigations on the impact of sound on visual attention and prospective audio-visual saliency model, we still need further study. Guangtao Zhai, Yucheng Zhu, Jun Zhou 0007, Xiao-Ping Zhang 0002 |
ICIP | 3 |
| 2022 | Image Quality Assessment: From Mean Opinion Score to Opinion Score DistributionabstractRecently, many methods have been proposed to predict the image quality which is generally described by the mean opinion score (MOS) of all subjective ratings given to an image. However, few efforts focus on predicting the opinion score distribution of the image quality ratings. In fact, the opinion score distribution reflecting subjective diversity, uncertainty, etc., can provide more subjective information about the image quality than a single MOS, which is worthy of in-depth study. In this paper, we propose a convolutional neural network based on fuzzy theory to predict the opinion score distribution of image quality. The proposed method consists of three main steps: feature extraction, feature fuzzification and fuzzy transfer. Specifically, we first use the pre-trained VGG16 without fully-connected layers to extract image features. Then, the extracted features are fuzzified by fuzzy theory, which is used to model epistemic uncertainty in the process of feature extraction. Finally, a fuzzy transfer network is used to predict the opinion score distribution of image quality by learning the mapping from epistemic uncertainty to the uncertainty existing in the image quality ratings. In addition, a new loss function is designed based on the subjective uncertainty of the opinion score distribution. Extensive experimental results prove the superior prediction performance of our proposed method. Xiongkuo Min, Yucheng Zhu, Jing Li 0026, Xiao-Ping Zhang 0002, Guangtao Zhai |
ACM Multimedia | 3 |
| 2022 | Skeleton2Humanoid: Animating Simulated Characters for Physically-plausible Motion In-betweeningabstractHuman motion synthesis is a long-standing problem with various applications in digital twins and the Metaverse. However, modern deep learning based motion synthesis approaches barely consider the physical plausibility of synthesized motions and consequently they usually produce unrealistic human motions. In order to solve this problem, we propose a system "Skeleton2Humanoid" which performs physics-oriented motion correction at test time by regularizing synthesized skeleton motions in a physics simulator. Concretely, our system consists of three sequential stages: (I) test time motion synthesis network adaptation, (II) skeleton to humanoid matching and (III) motion imitation based on reinforcement learning (RL). Stage I introduces a test time adaptation strategy, which improves the physical plausibility of synthesized human skeleton motions by optimizing skeleton joint locations. Stage II performs an analytical inverse kinematics strategy, which converts the optimized human skeleton motions to humanoid robot motions in a physics simulator, then the converted humanoid robot motions can be served as reference motions for the RL policy to imitate. Stage III introduces a curriculum residual force control policy, which drives the humanoid robot to mimic complex converted reference motions in accordance with the physical law. We verify our system on a typical human motion synthesis task, motion-in-betweening. Experiments on the challenging LaFAN1 dataset show our system can outperform prior methods significantly in terms of both physical plausibility and accuracy. Code will be released for research purposes at: https://github.com/michaelliyunhao/Skeleton2Humanoid. Zhenbo Yu, Yucheng Zhu, Bingbing Ni, Guangtao Zhai, Wei Shen 0002 |
ACM Multimedia | 3 |
| 2022 | Viewing Behavior Supported Visual Saliency Predictor for 360 Degree VideosabstractIn virtual reality (VR), correct and precise estimations of user’s visual fixations and head movements can enhance the quality of experience by allocating more computation resources for analysing and rendering on the areas of interest. However, there is insufficient research about understanding the visual exploration of users when modeling VR visual attention. To bridge the gap between the saliency prediction for traditional 2D content and omnidirectional content, we construct the visual attention dataset and propose the visual saliency prediction framework for panoramic videos. Around the instantaneous viewing behavior, we propose a traditional method to adapt 2D saliency models and design a CNN-based model to better predict visual saliency. In the proposed traditional model, mechanism of visual attention and viewing behaviors are considered in the computation of edge weights on graphs which are interpreted as Markov chains. The fraction of the visual attention that is diverted to each high-clarity vision (HCV) area is estimated through equilibrium distribution of this chain. We also propose the Graph-Based CNN model. The RGB channel and optical flow form the spatial-temporal units of HCVs, from which node feature vectors are extracted. Graph convolution is used to learn the mutual information between node feature vectors of HCVs and retain geometric information. Then feature vectors are aligned according to geometry structure of equirectangular format, and the feature decoder maps the aligned feature maps to the data distribution. We also construct the dynamic omnidirectional monocular (DOM) saliency dataset with 64 diverse videos evaluated by 28 people. The subjective results show that the instantaneous viewing behavior is important in the VR experience. Extensive experiments are conducted on the dataset and the results demonstrate the effectiveness of the proposed framework. The dataset will be released to facilitate the future studies related to visual saliency prediction for 360-degree contents. Yucheng Zhu, Guangtao Zhai, Yiwei Yang 0007, Huiyu Duan, Xiongkuo Min, Xiaokang Yang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2022 | Confusing Image Quality Assessment: Toward Better Augmented Reality ExperienceabstractWith the development of multimedia technology, Augmented Reality (AR) has become a promising next-generation mobile platform. The primary value of AR is to promote the fusion of digital contents and real-world environments, however, studies on how this fusion will influence the Quality of Experience (QoE) of these two components are lacking. To achieve better QoE of AR, whose two layers are influenced by each other, it is important to evaluate its perceptual quality first. In this paper, we consider AR technology as the superimposition of virtual scenes and real scenes, and introduce visual confusion as its basic theory. A more general problem is first proposed, which is evaluating the perceptual quality of superimposed images, i.e., confusing image quality assessment. A ConFusing Image Quality Assessment (CFIQA) database is established, which includes 600 reference images and 300 distorted images generated by mixing reference images in pairs. Then a subjective quality perception experiment is conducted towards attaining a better understanding of how humans perceive the confusing images. Based on the CFIQA database, several benchmark models and a specifically designed CFIQA model are proposed for solving this problem. Experimental results show that the proposed CFIQA model achieves state-of-the-art performance compared to other benchmark models. Moreover, an extended ARIQA study is further conducted based on the CFIQA study. We establish an ARIQA database to better simulate the real AR application scenarios, which contains 20 AR reference images, 20 background (BG) reference images, and 560 distorted images generated from AR and BG references, as well as the correspondingly collected subjective quality ratings. Three types of full-reference (FR) IQA benchmark variants are designed to study whether we should consider the visual confusion when designing corresponding IQA algorithms. An ARIQA metric is finally proposed for better evaluating the perceptual quality of AR images. Experimental results demonstrate the good generalization ability of the CFIQA model and the state-of-the-art performance of the ARIQA model. The databases, benchmark models, and proposed metrics are available at: https://github.com/DuanHuiyu/ARIQA. Huiyu Duan, Xiongkuo Min, Yucheng Zhu, Guangtao Zhai, Xiaokang Yang 0001, Patrick Le Callet |
IEEE Trans. Image Process. | 3 |
| 2022 | HazDesNet: An End-to-End Network for Haze Density PredictionabstractVision-based intelligent systems such as driver assistance systems and transportation systems should take into account weather conditions. The presence of haze in images can be a critical threat to driving scenarios. Haze density measures the visibility and usability of hazy images captured in real-world conditions. The prediction of haze density can be valuable in various vision-based intelligent systems, especially in those systems deployed in outdoor environments. Haze density prediction is a challenging task since the haze and many scene contents have a lot in common in appearance. Existing methods generally utilize different priors and design complex handcrafted features to predict the visibility or haze density of the image. In this article, we propose a novel end-to-end convolutional neural network (CNN) based method to predict haze density, named as HazDesNet. Our HazDesNet takes a hazy image as input and predicts a pixel-level haze density map. The density map is then refined and smoothed, and the average of the refined map is calculated as the global haze density of the image. To verify the performance of HazDesNet, a subjective human study is performed to build a Human Perceptual Haze Density (HPHD) database, which includes 500 real-world hazy images and 100 synthetic hazy images, and the corresponding human-rated perceptual haze density scores. Experimental results show that our method achieves the best haze density prediction performance on our built HPHD database and existing databases. Besides the global quantitative results, our HazDesNet is capable of predicting a continuous, stable, fine, and high-resolution haze density map. We will make the database and code publicly available athttps://github.com/JiaheZhang/HazDesNet. Xiongkuo Min, Yucheng Zhu, Guangtao Zhai, Jiantao Zhou 0001, Xiaokang Yang 0001, Wenjun Zhang 0001 |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2021 | Looking here or there? Gaze Following in 360-Degree ImagesabstractGaze following, i.e., detecting the gaze target of a human subject, in 2D images has become an active topic in computer vision. However, it usually suffers from the out of frame issue due to the limited field-of-view (FoV) of 2D images. In this paper, we introduce a novel task, gaze following in 360-degree images which provide an omnidirectional FoV and can alleviate the out of frame issue. We collect the first dataset, "GazeFollow360"1, for this task, containing around 10,000 360-degree images with complex gaze behaviors under various scenes. Existing 2D gaze following methods suffer from performance degradation in 360degree images since they may use the assumption that a gaze target is in the 2D gaze sight line. However, this assumption is no longer true for long-distance gaze behaviors in 360-degree images, due to the distortion brought by sphere-to-plane projection. To address this challenge, we propose a 3D sight line guided dual-pathway framework, to detect the gaze target within a local region (here) and from a distant region (there), parallelly. Specifically, the local region is obtained as a 2D cone-shaped field along the 2D projection of the sight line starting at the human subject’s head position, and the distant region is obtained by searching along the sight line in 3D sphere space. Finally, the location of the gaze target is determined by fusing the estimations from both the local region and the distant region. Experimental results show that our method achieves significant improvements over previous 2D gaze following methods on our GazeFollow360 dataset. Wei Shen 0002, Zhongpai Gao, Yucheng Zhu, Guangtao Zhai, Guodong Guo |
ICCV | 4 |
| 2021 | Muiqa: Image Quality Assessment Database And Algorithm For Medical Ultrasound ImagesabstractIn the process of medical image acquisition, medical images may be blurred or ghosted due to machine noise, electromagnetic interference, man-made disturbance, etc. This can result in poor image quality and severely affect the diagnosis accuracy and confidence of doctors. IntraVascular UltraSound (IVUS) is an important supplementary method for the diagnosis of coronary angiography. IVUS images can be distorted for many reasons and some severe distortions can affect diagnosis confidence. However, existing manual medical image quality control method is extremely time-consuming and requires a lot of manpower. To solve this problem, we first construct an Medical UltraSound Image Quality Assessment (MUIQA) database, which consists of 10766 IVUS images with quality labels given by professional doctors. Then we propose a deep-learning network to automatically distinguish the low, medium and high level images from each other. We achieve good classification accuracy of 96.34% on the testing set. Xiongkuo Min, Huiyu Duan, Yucheng Zhu, Guangtao Zhai |
ICIP | 4 |
| 2021 | SalGFCN: Graph Based Fully Convolutional Network for Panoramic Saliency PredictionabstractThe saliency prediction of panoramic images is dramatically affected by the distortion caused by non-Euclidean geometry characteristic. Traditional CNN based saliency pre-diction algorithms for 2D images are no longer suitable for 360-degree images. Intuitively, we propose a graph based fully convolutional network for saliency prediction of 360-degree images, which can reasonably map panoramic pixels to spherical graph data structures for representation. The saliency prediction network is based on residual U-Net architecture, with dilated graph convolutions and attention mechanism in the bottleneck. Furthermore, we design a fully convolutional layer for graph pooling and unpooling operations in spherical graph space to retain node-to-node features. Experimental results show that our proposed method outperforms other state-of-the-art saliency models on the large-scale dataset. Yiwei Yang 0007, Yucheng Zhu, Zhongpai Gao, Guangtao Zhai |
VCIP | 2 |
| 2021 | RANSP: Ranking attention network for saliency prediction on omnidirectional images
Dandan Zhu 0001, Yongqing Chen, Xiongkuo Min, Yucheng Zhu, Qiangqiang Zhou, Guangtao Zhai, Xiaokang Yang 0001 |
Neurocomputing | 4 |
| 2021 | Comparative Perceptual Assessment of Visual Signals Using Free Energy FeaturesabstractIn this paper, we put forward the concept of comparative perceptual quality assessment (C-PQA), which refers to the judgment of relative qualities of two visual signals of the same content, but subject to different types and levels of distortions. While it is straightforward for human observers to fulfill the CPQA task in daily lives, it remains a difficult challenge for the current research of perceptual quality assessment (PQA). Among the existing PQA algorithms, the full-reference (FR) and reducedreference (RR) methods both need prior knowledge of the original images while the no-reference (NR) algorithms usually work with a single input image. C-PQA is inherently different from those existing methods in that it takes an image pair as input and predicts their relative quality without using any knowledge about the original image. In this paper, we propose a brain theory inspired approach to C-PQA that emulates the process of comparing the relative quality of two visual stimuli as performed by the human visual system (HVS) within the framework of free energy minimization. The brain's internal generative models initialized on the inputs are then used to explain both images. During the internal generative modeling, a group of features are extracted and then integrated to determine the relative quality of two images. We designed a dedicated image database to test the proposed C-PQA algorithm. Experimental results show that the proposed method achieves up to 98% prediction accuracy in line with the subjective ratings, outperforming many state of the art PQA algorithms. Guangtao Zhai, Yucheng Zhu, Xiongkuo Min |
IEEE Trans. Multim. | 2 |
| 2021 | Learning a Deep Agent to Predict Head Movement in 360-Degree ImagesabstractVirtual reality adequately stimulates senses to trick users into accepting the virtual environment. To create a sense of immersion, high-resolution images are required to satisfy human visual system, and low latency is essential for smooth operations, which put great demands on data processing and transmission. Actually, when exploring in the virtual environment, viewers only perceive the content in the current field of view. Therefore, if we can predict the head movements that are important behaviors of viewers, more processing resources can be allocated to the active field of view. In this article, we propose a model to predict the trajectory of head movement. Deep reinforcement learning is employed to mimic the decision making. In our framework, to characterize each state, features for viewport images are extracted by convolutional neural networks. In addition, the spherical coordinate maps and visited maps are generated for each viewport image, which facilitate the multiple dimensions of the state information by considering the impact of historical head movement and position information. To ensure the accurate simulation of visual behaviors during the watching of panoramas, we stipulate that the model imitates the behaviors of human demonstrators. To allow the model to generalize to more conditions, the intrinsic motivation is employed to guide the agent’s action toward reducing uncertainty, which can enhance robustness during the exploration. The experimental results demonstrate the effectiveness of the proposed stepwise head movement predictor. Yucheng Zhu, Guangtao Zhai, Xiongkuo Min, Jiantao Zhou 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2020 | Automatic Region Selection For Objective Sharpness Assessment Of Mobile Device PhotosabstractMobile devices are the source of a vast majority of digital photos today. Photos taken by mobile devices generally have fairly good visual quality. When evaluating high-quality mobile device photos, people have to manually zoom in to local regions to discern the subtle difference. Understandably, a global objective quality assessment method cannot perform well on such task. Therefore, local region selection is widely recognized as a prerequisite for the following quality evaluation. Clearly, subjective regions selection suffers from the drawbacks in terms of productivity, reproducibility and optimality. In this paper, we propose an automatic local region selection algorithm for sharpness measurement of mobile device photos. Specifically, local texture statistics, depth, saliency, as well as inter-pictures difference, are used as main features to select an optimal local region, in which the sharpness is then measured. For validation, we have built a largescale database for sharpness evaluation of mobile device photos, with 100 different scenes shot by several flagship mobile phones. The experimental results show that the performance of classic sharpness evaluation algorithms can be substantially improved with the region selected by the proposed algorithm. Guangtao Zhai, Wenhan Zhu, Yucheng Zhu, Xiongkuo Min, Xiao-Ping Zhang 0002, Hua Yang 0001 |
ICIP | 4 |
| 2020 | Ransp: Ranking Attention Network For Saliency Prediction On Omnidirectional ImagesabstractVarious convolutional neural network (CNN)-based methods have shown the ability to boost the performance of saliency prediction on omnidirectional images (ODIs). However, these methods are limited by sub-optimal accuracy, because not all the features extracted by the CNN model are not useful for the final fine-grained saliency prediction. Features are redundant and have negative impact on the final fine-grained saliency prediction. To tackle this problem, we propose a novel Ranking Attention Network for saliency prediction (RANSP) of head fixations on ODIs. Specifically, the part-guided attention (PA) module and channel-wise feature (CF) extraction module are integrated in a unified framework and are trained in an end-to-end manner for fine-grained saliency prediction. To better utilize the channel-wise feature map, we further propose a new Ranking Attention Module (RAM), which automatically ranks and selects these maps based on scores for fine-grained saliency prediction. Extensive experiments are conducted to show the effectiveness of our method for saliency prediction of ODIs. Dandan Zhu 0001, Yongqing Chen, Tian Han 0001, Defang Zhao, Yucheng Zhu, Qiangqiang Zhou, Guangtao Zhai, Xiaokang Yang 0001 |
ICME | 5 |
| 2020 | Saliency Prediction on Omnidirectional Images with Brain-Like Shallow Neural NetworkabstractDeep feedforward convolutional neural networks (CNNs) perform well in the saliency prediction of omnidirectional images (ODIs), and have become the leading class of candidate models of the visual processing mechanism in the primate ventral stream. These CNNs have evolved from shallow network architecture to extremely deep and branching architecture to achieve superb performance in various vision tasks, yet it is unclear how brain-like they are. In particular, these deep feedforward CNNs are difficult to mapping to ventral stream structure of the brain visual system due to their vast number of layers and missing biologically-important connections, such as recurrence. To tackle this issue, some brain-like shallow neural networks are introduced. In this paper, we propose a novel brain-like network model for saliency prediction of head fixations on ODIs. Specifically, our proposed model consists of three modules: a CORnet-S module, a template feature extraction module and a ranking attention module (RAM). The CORnet-S module is a lightweight artificial neural network (ANN) with four anatomically mapped areas (V1, V2, V4 and IT) and it can simulate the visual processing mechanism of ventral visual stream in the human brain. The template features extraction module is introduced to extract attention maps of ODIs and provide guidance for the feature ranking in the following RAM module. The RAM module is used to rank and select features that are important for fine-grained saliency prediction. Extensive experiments have validated the effectiveness of the proposed model in predicting saliency maps of ODIs, and the proposed model outperforms other state-of-the-art methods with similar scale. Dandan Zhu 0001, Yongqing Chen, Xiongkuo Min, Defang Zhao, Yucheng Zhu, Qiangqiang Zhou, Xiaokang Yang 0001, Tian Han 0001 |
ICPR | 5 |
| 2020 | An Improved Algorithm for Real-Time Dual-View DisplayabstractDual-view display based on spatial psychovisual modulation (SPVM) aims to present two different views on a single screen. Users with special glasses see a personal view. Users without special glasses see a shared view which can be completely irrelevant to the personal view. This dual-view display technology can be used for information security, QR hiding, education, etc. In this paper, we propose an improved algorithm called pseudo-inv algorithm, using the pseudo-inverse method to solve the optimization problem. Moreover, a new Gaussian integration window and a new method to adjust the luminance range, are implemented for a better personal view. The experiments prove that the new algorithm results in a better image quality of both the shared view and the personal view. The time cost of the new algorithm is much less than that of the previous algorithms. Zhongpai Gao, Guangtao Zhai, Zhaodi Wang 0002, Yucheng Zhu |
ISCAS | 6 |
| 2020 | The Prediction of Saliency Map for Head and Eye Movements in 360 Degree ImagesabstractBy recording the whole scene around the capturer, virtual reality (VR) techniques can provide viewers the sense of presence. To provide a satisfactory quality of experience, there should be at least 60 pixels per degree, so the resolution of panoramas should reach 21600 × 10800. The huge amount of data will put great demands on data processing and transmission. However, when exploring in the virtual environment, viewers only perceive the content in the current field of view (FOV). Therefore if we can predict the head and eye movements which are important behaviors of viewer, more processing resources can be allocated to the active FOV. But conventional saliency prediction methods are not fully adequate for panoramic images. In this paper, a new panorama-oriented model, to predict head and eye movements, is proposed. Due to the superiority of computation in the spherical domain, the spherical harmonics are employed to extract features at different frequency bands and orientations. Related low- and high-level features including the rare components in the frequency domain and color domain, the difference between center vision and peripheral vision, visual equilibrium, person and car detection, and equator bias are extracted to estimate the saliency. To predict head movements, visual mechanisms including visual uncertainty and equilibrium are incorporated, and the graphical model and functional representation for the switch of head orientation are established. Extensive experimental results on the publicly available database demonstrate the effectiveness of our methods. Yucheng Zhu, Guangtao Zhai, Xiongkuo Min, Jiantao Zhou 0001 |
IEEE Trans. Multim. | 1 |
| 2019 | Quality Evaluation of Image Dehazing Methods Using Synthetic Hazy ImagesabstractTo enhance the visibility and usability of images captured in hazy conditions, many image dehazing algorithms (DHAs) have been proposed. With so many image DHAs, there is a need to evaluate and compare these DHAs. Due to the lack of the reference haze-free images, DHAs are generally evaluated qualitatively using real hazy images. But it is possible to perform quantitative evaluation using synthetic hazy images since the reference haze-free images are available and full-reference (FR) image quality assessment (IQA) measures can be utilized. In this paper, we follow this strategy and study DHA evaluation using synthetic hazy images systematically. We first build a synthetic haze removing quality (SHRQ) database. It consists of two subsets: regular and aerial image subsets, which include 360 and 240 dehazed images created from 45 and 30 synthetic hazy images using 8 DHAs, respectively. Since aerial imaging is an important application area of dehazing, we create an aerial image subset specifically. We then carry out subjective quality evaluation study on these two subsets. We observe that taking DHA evaluation as an exact FR IQA process is questionable, and the state-of-the-art FR IQA measures are not effective for DHA evaluation. Thus, we propose a DHA quality evaluation method by integrating some dehazing-relevant features, including image structure recovering, color rendition, and over-enhancement of low-contrast areas. The proposed method works for both types of images, but we further improve it for aerial images by incorporating its specific characteristics. Experimental results on two subsets of the SHRQ database validate the effectiveness of the proposed measures. Xiongkuo Min, Guangtao Zhai, Ke Gu 0001, Yucheng Zhu, Jiantao Zhou 0001, Guodong Guo, Xiaokang Yang 0001, Xin-Ping Guan, Wenjun Zhang 0001 |
IEEE Trans. Multim. | 4 |
| 2018 | Perceptual Quality Assessment of Omnidirectional ImagesabstractOmnidirectional images and videos can provide immersive experience of real-world scenes in Virtual Reality (VR) environment. We present a perceptual omnidirectional image quality assessment (IQA) study in this paper since it is extremely important to provide a good quality of experience under the VR environment. We first establish an omnidirectional IQA (OIQA) database, which includes 16 source images and 320 distorted images degraded by 4 commonly encountered distortion types, namely JPEG compression, JPEG2000 compression, Gaussian blur and Gaussian noise. Then a subjective quality evaluation study is conducted on the OIQA database in the VR environment. Considering that humans can only see a part of the scene at one movement in the VR environment, visual attention becomes extremely important. Thus we also track head and eye movement data during the quality rating experiments. The original and distorted omnidirectional images, subjective quality ratings, and the head and eye movement data together constitute the OIQA database. State-of-the-art full-reference (FR) IQA measures are tested on the OIQA database, and some new observations different from traditional IQA are made. The OIQA database will be released to facilitate further research. Huiyu Duan, Guangtao Zhai, Xiongkuo Min, Yucheng Zhu, Yi Fang 0009, Xiaokang Yang 0001 |
ISCAS | 4 |
| 2018 | An Image Augmentation Method for Quality Assessment DatabaseabstractImage databases for quality assessment are helpful to evaluate the performance of objective assessment methods. Recommendations in regard to the constitution of databases and experimental methods of the subjective assessment have been proposed to ensure the database a good ground truth for the validation of objective quality assessment methods. However, these restrictions make databases scale-limited by covering small number of scenes distorted by few levels. To enrich IQA databases and increase the generalization capability of IQA models, we devise an effective image augmentation method. The two-stages scheme consists of the image-label pairs generation by minimizing the free energy between the pristine image and its augmentation as well as the distortion level interpolation which is based on the monotonicity of the perceptual quality with the severity of distortion. The experimental results show the ability of the augmented database to improve the prediction accuracy of learning-based no-reference image quality assessment metrics which in turn demonstrates the effectiveness of our method. Yucheng Zhu, Guangtao Zhai, Wenhan Zhu, Jiantao Zhou 0001 |
ISCAS | 1 |
| 2018 | The prediction of head and eye movement for 360 degree images
Yucheng Zhu, Guangtao Zhai, Xiongkuo Min |
Signal Process. Image Commun. | 1 |
| 2017 | No-reference quality assessment for JPEG compressed imagesabstractJPEG is a most commonly used standard of compression for digital images. Quality factor (Qfactor) for JPEG compressed image is actually a suitable indicator to the perceptual quality. However, the information of the compressor might be unknown due to various reasons. To evaluate the Qfactor, we recompress the formerly compressed image and measure the consistency between them. Then we define the fixed points (the points on the Qfactor-axis where the content of recompressed images are almost the same with that of directly compressed images) by following the Qfactor based specifications and form the image set. The quality of JPEG compressed images are measured by combining the estimated Qfactor with the features extracted from the image set. The experimental results confirm that the proposed image quality assessment technique, which is no-reference, is able to faithfully predict the visual quality of JPEG compressed images. Yucheng Zhu, Guangtao Zhai, Ke Gu 0001, Wenhan Zhu |
QoMEX | 1 |
| 2016 | No-reference image quality assessment for photographic images of consumer deviceabstractIn this paper we study common, camera-specific kinds of distortions and propose a no-reference image quality assessment algorithm for photographic images produced by consumer devices. Those real consumer-type images, being different from simulated-distortion images, are with realistic artifacts and quality ranges. We find that the state-of-the-art no-reference image quality assessment approaches do not perform well on those photographic images, and propose an approach that achieves high prediction performance on a dataset of consumer-centric images. The proposed method, with no need for the original image, is able to reveal camera-specific problems and differentiate consumer cameras. Yucheng Zhu, Guangtao Zhai, Ke Gu 0001, Zhaohui Che |
ICASSP | 1 |
| 2016 | Blindly evaluating stereoscopic image quality with free-energy principleabstractThree-dimensional (3D) imaging technology has been growingly prevalent in today's world. But objective quality assessment of 3D images is a challenging task. In this paper, we propose a blind metric to predict the perceptual quality of stereopairs within the concept of free energy. On the basis of a psychological measure, the free energy is a principle telling where supervises more and attracts human attention. We believe that the “surprise” can account for the binocular rivalry and thus be used to predict the quality of stereopairs. We first evaluate the quality of the monoscopic image, then introduce the computation process of binocular rivalry's results for deciding the relative importance of the left and right views, and finally infer the overall quality score. Our algorithm is tested on the symmetric LIVE3D-I and asymmetric LIVE3D-II databases. Experimental results confirm that the proposed blind 3D IQA technique, without distortion identification, is able to faithfully predict the visual quality of stereopairs. Yucheng Zhu, Guangtao Zhai, Ke Gu 0001, Min Liu 0003 |
ISCAS | 1 |
| 2016 | Closing the gap: Visual quality assessment considering viewing conditionsabstractMost of existing visual quality assessment algorithms are tested on standard databases that are created in controlled viewing conditions (e.g. display device, viewing distance and lighting). This implies that all the recoded subjective scores are only valid for the specific settings used in the database. However, with the prevalence of mobile devices, the practical viewing environments can significantly vary from moment to moment. It is our daily experience that the same image can look drastically different on dissimilar devices under changed viewing distance and/or lighting conditions. In other words, a gap exists between the eyes and the visual contents behind the screen in current research of quality assessment. Therefore, in this work, we perform subjective quality evaluation with varied actual viewing conditions. To make the research reproducible, we build a prototype system to record what the eyes really see from the screen and construct the viewing environment-changed image database. The database will be made available to the public. Meanwhile we design a dedicated effective environment-assessing algorithm. We believe that this work will benefit the research of visual quality assessment towards more practical applications. Yucheng Zhu, Guangtao Zhai, Ke Gu 0001, Zhaohui Che |
QoMEX | 1 |
| 2016 | Quality assessment for dual-view display systemabstractSpatial psychovisual modulation (SPVM) is a new information display technology, which aims to generate multiple visual percepts for different viewers on a single display simultaneously. After the proposal of SPVM, lots of efforts have been made and several applications (i.e., dual-view display system) have been implemented based on this technology. The dual-view display (DVD) system is considered as an effective digital image hiding system based on SPVM theory, but little work has been dedicated to the perceptual quality assessment of DVD system. Up to now, there is no clear and standard method to evaluate the performance of the dual-view display system. It is important for the viewers to see a clear and non-aliasing image when they are front of the screen. Therefore, in this paper, we will build a DVD database and carry out a subjective experiment to evaluate the performance of the DVD system, and then we investigate and analyze the performance of prevailing no-reference (NR) image quality metrics on the particular DVD system. We have a sufficient belief that this paper can supply the guideline for the performance on the DVD system and serve as a good testing bed for future research of SPVM technology. Yuanchun Chen, Guangtao Zhai, Ke Gu 0001, Jia Wang 0004, Zhongpai Gao, Yucheng Zhu |
VCIP | 7 |