EDBT 2026 Demo / reviewers in the wild / expert
Guangtao Zhai
dblp:19/3230
· DBLP profile ↗
563ranked-venue papers
30as first author
375since 2021 · last 2026
0000-0001-8165-9322ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 428 · 25 first-author · 283 since 2021Artificial intelligence and machine learning · 116 · 101 since 2021Systems, architecture and hardware · 42 · 4 first-author · 14 since 2021Computer networks · 30 · 1 first-author · 27 since 2021Applied, interdisciplinary, general and emerging computing · 17 · 1 first-author · 14 since 2021Human-computer interaction and ubiquitous computing · 9 · 7 since 2021Databases, data management, data science and information retrieval · 5 · 3 since 2021Security and privacy · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | VQAThinker: Exploring Generalizable and Explainable Video Quality Assessment via Reinforcement LearningabstractVideo quality assessment (VQA) aims to objectively quantify perceptual quality degradation in alignment with human visual perception. Despite recent advances, existing VQA models still suffer from two critical limitations: poor generalization to out-of-distribution (OOD) videos and limited explainability, which restrict their applicability in real-world scenarios. To address these challenges, we propose VQAThinker, a reasoning-based VQA framework that leverages large multimodal models (LMMs) with reinforcement learning to jointly model video quality understanding and scoring, emulating human perceptual decision-making. Specifically, we adopt group relative policy optimization (GRPO), a rule-guided reinforcement learning algorithm that enables reasoning over video quality under score-level supervision, and introduce three VQA-specific rewards: (1) a bell-shaped regression reward that increases rapidly as the prediction error decreases and becomes progressively less sensitive near the ground truth; (2) a pairwise ranking reward that guides the model to correctly determine the relative quality between video pairs; and (3) a temporal consistency reward that encourages the model to prefer temporally coherent videos over their perturbed counterparts. Extensive experiments demonstrate that VQAThinker achieves state-of-the-art performance on both in-domain and OOD VQA benchmarks, showing strong generalization for video quality scoring. Furthermore, evaluations on video quality understanding tasks validate its superiority in distortion attribution and quality description compared to existing explainable VQA models and LMMs. These findings demonstrate that reinforcement learning offers an effective pathway toward building generalizable and explainable VQA models solely with score-level supervision. Linhan Cao, Wei Sun 0029, Weixia Zhang, Jun Jia, Kaiwei Zhang, Dandan Zhu 0001, Guangtao Zhai, Xiongkuo Min |
AAAI | 8 |
| 2026 | Scaling-up Perceptual Video Quality AssessmentabstractThe data scaling law has significantly enhanced large multi-modal models (LMMs) performance across various downstream tasks. However, in the domain of perceptual video quality assessment (VQA), the potential of data scaling remains unprecedented due to the scarcity of labeled resources and the insufficient scale of datasets. To address this, we propose OmniVQA, a framework designed to efficiently build high-quality, machine-dominated synthetic multi-modal instruction databases (MIDBs) for VQA. We then scale up to create OmniVQA-Chat-400K, the largest dataset in the VQA field concurrently. Our focus is on the technical and aesthetic quality dimensions, with abundant in-context instruction data to provide fine-grained VQA knowledge. Additionally, we build the OmniVQA-MOS-20K dataset to enhance the model's quantitative quality rating capabilities. We then introduce a complementary training strategy that effectively leverages the knowledge from datasets for different tasks. Furthermore, we propose the OmniVQA-FG (fine-grain)-Benchmark to evaluate the fine-grained performance of models. Our results demonstrate that our models achieve state-of-the-art performance in both tasks. Ziheng Jia, Xiaorong Zhu, Chunyi Li 0001, Jinliang Han, Xiaohong Liu 0001, Guangtao Zhai, Xiongkuo Min |
AAAI | 7 |
| 2026 | GeoX-Bench: Benchmarking Cross-View Geo-Localization and Pose Estimation Capabilities of Large Multimodal ModelsabstractLarge multimodal models (LMMs) have demonstrated remarkable capabilities across a wide range of tasks, however their knowledge and abilities in the cross-view geo-localization and pose estimation domains remain unexplored, despite potential benefits for navigation, autonomous driving, outdoor robotics, etc. To bridge this gap, we introduce GeoX-Bench, a comprehensive Benchmark designed to explore and evaluate the capabilities of LMMs in cross-view Geo-localization and pose estimation. Specifically, GeoX-Bench contains 10,859 panoramic-satellite image pairs spanning 128 cities in 49 countries, along with corresponding 755,976 question-answering (QA) pairs. Among these, 42,900 QA pairs are designated for benchmarking, while the remaining are intended to enhance the capabilities of LMMs. Based on GeoX-Bench, we evaluate the capabilities of 25 state-of-the-art LMMs on cross-view geo-localization and pose estimation tasks, and further explore the empowered capabilities of instruction-tuning. Our benchmark demonstrate that while current LMMs achieve impressive performance in geo-localization tasks, their effectiveness declines significantly on the more complex pose estimation tasks, highlighting a critical area for future improvement, and instruction-tuning LMMs on the training data of GeoX-Bench can significantly improve the cross-view geo-sense abilities. Yushuo Zheng, Jiangyong Ying, Huiyu Duan, Chunyi Li 0001, Jing Liu 0002, Xiaohong Liu 0001, Guangtao Zhai |
AAAI | 8 |
| 2026 | One Battle After Another: Probing LLMs' Limits on Multi-Turn Instruction Following with a Benchmark Evolving FrameworkabstractQi Jia, Ye Shen, Xiujie Song, Kaiwei Zhang, Shibo Wang, Dun Pei, Xiangyang Zhu, Guangtao Zhai. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Ye Shen, Xiujie Song, Kaiwei Zhang, Dun Pei, Guangtao Zhai |
ACL (1) | 8 |
| 2026 | Market-Bench: Benchmarking Large Language Models on Economic and Trade CompetitionabstractYushuo Zheng, Huiyu Duan, Zicheng Zhang, Yucheng Zhu, Xiongkuo Min, Guangtao Zhai. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Yushuo Zheng, Huiyu Duan, Yucheng Zhu, Xiongkuo Min, Guangtao Zhai |
ACL (1) | 6 |
| 2026 | ReGAF: A Relational Graph Attention Fusion Method for Chinese-Kazakh CLIR
Changle Yin, Junyao Shen, Ainur Zhumadillayeva, Guangtao Zhai |
ICIC (24) | 5 |
| 2026 | DocIQ: A Benchmark Dataset and Feature Fusion Network for Document Image Quality AssessmentabstractDocument image quality assessment (DIQA) is an important component for various applications, including optical character recognition (OCR), document restoration, and the evaluation of document image processing systems. In this paper, we introduce a subjective DIQA dataset DIQA-5000. The DIQA-5000 dataset comprises 5,000 document images, generated by applying multiple document enhancement techniques to 500 real-world images with diverse distortions. Each enhanced image was rated by 15 subjects across three rating dimensions: overall quality, sharpness, and color fidelity. Furthermore, we propose a specialized no-reference DIQA model that exploits document layout features to maintain quality perception at reduced resolutions to lower computational cost. Recognizing that image quality is influenced by both low-level and high-level visual features, we designed a feature fusion module to extract and integrate multi-level features from document images. To generate multi-dimensional scores, our model employs independent quality heads for each dimension to predict score distributions, allowing it to learn distinct aspects of document image quality. Experimental results demonstrate that our method outperforms current state-of-the-art general-purpose IQA models on both DIQA-5000 and an additional document image dataset focused on OCR accuracy. Fengjun Guo, Guangtao Zhai, Xiongkuo Min |
ISCAS | 5 |
| 2026 | Q-Agent: An MLLM-Driven Framework for Universal Visual Quality Assessment
Peihang Chen, Huiyu Duan, Zitong Xu, Yuqin Cao, Sijing Wu, Wei Sun 0029, Xiongkuo Min, Guangtao Zhai |
QoMEX | 14 |
| 2026 | Assessing Personality Consistency in Large Language Models: A Psychometric Framework for Human-Centric Quality of Experience
Yitian Kou, Dandan Zhu 0001, Wei Sun 0029, Kaiwei Zhang, Xiongkuo Min, Guangtao Zhai |
QoMEX | 6 |
| 2026 | LEIQ-Assessor: Multi-Dimensional Quality Assessment of Low-Light Enhanced Images via Multi-Task Learning
Wei Sun 0029, Yanwei Jiang, Dandan Zhu 0001, Jinqiu Sang, Jikai Xu, Weixia Zhang, Guangtao Zhai |
QoMEX | 7 |
| 2026 | M3DGCQA: A Quality Assessment Dataset for Multi-Object 3D Generated Contents
Farong Wen, Yuanhao Xue, Xiahui Ren, Ziying Wang, Yingjie Zhou 0003, Jun Jia, Jiezhang Cao, Xiaohong Liu 0001, Guangtao Zhai |
QoMEX | 11 |
| 2026 | QoMEX 2026 Grand Challenge on Video Quality Assessment for Asymmetric Encoded Videos: Methods and Results
Yixu Chen, Hai Wei, Pierre R. Lebreton, Patrick Le Callet, Alexander Kopte, Amritha Premkumar, Anna Meyer, Baojun Li, Changsheng Gao, Christian Herglotz, Christian Timmerer, Dandan Zhu 0001, Diwakara Reddy, Dong Liu 0002, Dounia Hammou, Guangtao Zhai, Hadi Amirpour, Hao Cheng 0015, Hichem Faraoun, Jonas Janzen, Krishna Srikar Durbha, Li Li 0040, Marc Windsheimer, MohammadAli Hamidi, Mykyta Skipenko, Paul Wawerek-Lopez, Pragyadipta Adhya, Prajit T. Rajendran, Rafal Mantiuk, Shien Ke, Sid Ahmed Fezza, Simon Deniffel, Wei Sun 0029, Weixia Zhang, Xiangguang Chen, Zuowei Cao, Minhao Tang, Xiaoyan Sun 0001, Xingwei Liu, Yeganeh Chatri, Yenan Xu |
QoMEX | 17 |
| 2026 | Towards versatile multimedia quality assessment for visual communications
Ziheng Jia, Chunyi Li 0001, Yingjie Zhou 0003, Xiaohong Liu 0001, Xiongkuo Min, Guangtao Zhai |
Sci. China Inf. Sci. | 7 |
| 2026 | Enhancing blind video quality assessment with rich quality-aware features
Wei Sun 0029, Linhan Cao, Jun Jia, Xiongkuo Min, Guangtao Zhai |
Expert Syst. Appl. | 7 |
| 2026 | A trajectory-based framework for diagnosing and calibrating social order in large language models
Yan Zhao 0012, Peitong Han, Guangtao Zhai |
Expert Syst. Appl. | 4 |
| 2026 | Beyond catastrophic forgetting: A continual learning-driven multi-modal fusion model for saliency prediction in dynamic scenes
Jiaqi Wang 0003, Dandan Zhu 0001, Kun Zhu 0024, Xiongkuo Min, Guangtao Zhai |
Expert Syst. Appl. | 6 |
| 2026 | InvJND: Just Noticeable Difference Estimation via Deep Invertible Network
Qiuping Jiang, Zhihua Wang 0002, Shiqi Wang 0001, Feng Shao 0001, Guangtao Zhai, Weisi Lin |
Int. J. Comput. Vis. | 7 |
| 2026 | Lumina-mGPT: Flexible Photorealistic Autoregressive Text-to-Image Generation
Yi Xin 0003, Shitian Zhao, Le Zhuo, Weifeng Lin, Xinyue Li 0001, Guangtao Zhai, Xiaohong Liu 0001, Hongsheng Li 0001, Yu Qiao 0001, Peng Gao 0007 |
Int. J. Comput. Vis. | 8 |
| 2026 | Parameter-Efficient Fine-Tuning for Pre-Trained Vision Models: A Survey and Benchmark
Yi Xin 0003, Jianjiang Yang, Yuntao Du 0001, Haoxing Chen, Kangrui Cen, Yangfan He, Yuewen Cao, Junjun He, Xiaokang Yang 0001, Guangtao Zhai, Ming-Hsuan Yang 0001, Xiaohong Liu 0001 |
Int. J. Comput. Vis. | 13 |
| 2026 | ResAD++: Towards Class Agnostic Anomaly Detection via Residual Feature Learning
Xincheng Yao, Muming Zhao, Guangtao Zhai |
Int. J. Comput. Vis. | 4 |
| 2026 | Light-VQA+: A Video Quality Assessment Model for Exposure Correction with Vision-Language Guidance
Xunchu Zhou, Xiaohong Liu 0001, Yudong Zhang 0001, Tengchuan Kou, Chunyi Li 0001, Haoning Wu 0001, Guangtao Zhai |
Int. J. Comput. Vis. | 9 |
| 2026 | MI3S: A multimodal large language model assisted quality assessment framework for AI-generated talking heads
Yingjie Zhou 0003, Sijing Wu, Jun Jia, Yanwei Jiang, Wei Sun 0029, Xiaohong Liu 0001, Xiongkuo Min, Guangtao Zhai |
Inf. Process. Manag. | 9 |
| 2026 | An All-in-One Quality Assessment Agent for 4D digital human: Bridging talking heads and animated human
Yingjie Zhou 0003, Farong Wen, Li Xu 0008, Yu Zhou 0016, Jiezhang Cao, Xiaohong Liu 0001, Xiongkuo Min, Yu Wang 0002, Guangtao Zhai |
Inf. Process. Manag. | 12 |
| 2026 | Preference-guided debiasing for no-reference enhancement image quality assessment
Shiqi Gao, Zitong Xu, Huiyu Duan, Xiongkuo Min, Jia Wang 0004, Guangtao Zhai |
Image Vis. Comput. | 7 |
| 2026 | Linking-free online spatio temporal action detection
Ningyu Sun, Yichao Yan, Guangtao Zhai, Xiaokang Yang 0001 |
Image Vis. Comput. | 4 |
| 2026 | GOBench: Benchmarking and instruction-tuned assessment of geometric optics in multimodal LLMs
Xiaorong Zhu, Ziheng Jia, Guangtao Zhai |
J. Vis. Commun. Image Represent. | 4 |
| 2026 | Hierarchical Mesh Representation Learning With Spectral Dictionary EmbeddingabstractLearning mesh representation is important for many 3D tasks. Conventional convolution for regular data (i.e., images) cannot directly be applied to meshes since each vertex's neighbors are unordered. Previous methods use isotropic filters or predefined local coordinate systems or learning weighting matrices for each template vertex to overcome the irregularity. Learning weighting matrices to resample the vertex's neighbors into an implicit canonical order is the most effective way to capture the local structure of each vertex. However, learning weighting matrices for each vertex increases the model size linearly with the vertex number. Thus, large parameters are required for high-resolution 3D shapes, which is not favorable for many applications. In this paper, we learn spectral dictionary (i.e., bases) for the weighting matrices such that the model size is independent of the resolution of 3D shapes. The coefficients of the weighting matrix bases are learned from the spectral features of the template and its hierarchical levels in a weight-sharing manner. Furthermore, we introduce an adaptive sampling method that learns the hierarchical mapping matrices directly to improve the performance without increasing the model size at the inference stage. Comprehensive experiments demonstrate that our model produces state-of-the-art results with a much smaller model size. Zhongpai Gao, Junchi Yan, Tianyu Luan, Guangtao Zhai, Xiaokang Yang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2026 | Variational Bayesian Personalized RankingabstractPairwise learning underpins implicit collaborative filtering, yet its effectiveness is often hindered by sparse supervision, noisy interactions, and popularity-driven exposure bias. In this paper, we propose Variational Bayesian Personalized Ranking (VarBPR), a tractable variational framework for implicit-feedback pairwise learning that offers principled exposure controllability and theoretical interpretability. VarBPR reformulates pairwise learning as variational inference over discrete latent indexing variables, explicitly modeling noise and indexing uncertainty, and divides training into two stages: variational inference, which solve variational posteriors, and variational learning, which updates model parameters based on these posteriors. In the variational inference stage, we develop a variational formulation that integrates preference alignment, denoising, and popularity debiasing under a unified ELBO/regularization objective, deriving closed-form posteriors with clear control semantics: the prior encodes a target exposure pattern, while temperature/regularization strength controls posterior-prior adherence. As a result, exposure controllability becomes an endogenous and interpretable outcome of variational inference. In the variational learning stage, we propose a posterior-compression objective that reduces the ideal ELBO's computational complexity from polynomial to linear, with the approximation justified by an explicit Jensen-gap upper bound. Theoretically, we provide interpretable generalization guarantees by identifying a structural error component and revealing the opportunity cost of prioritizing certain exposure patterns (e.g., long-tail), offering a concrete analytical lens for designing controllable recommender systems. Empirically, We validate VarBPR across popular backbones; it demonstrates consistent gains in ranking accuracy, enables controlled long-tail exposure, and preserves the linear-time complexity of BPR. Bin Liu 0076, Xiaohong Liu 0001, Ziqiao Shang, Jielei Chu, Fei Teng 0001, Guangtao Zhai, Tianrui Li 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 9 |
| 2026 | SMC++: Masked Learning of Unsupervised Video Semantic CompressionabstractMost video compression methods focus on human visual perception, neglecting semantic preservation. This leads to severe semantic loss during the compression, hampering downstream video analysis tasks. In this paper, we propose a Masked Video Modeling (MVM)-powered compression framework that particularly preserves video semantics, by jointly mining and compressing the semantics in a self-supervised manner. While MVM is proficient at learning generalizable semantics through the masked patch prediction task, it may also encode non-semantic information like trivial textural details, wasting bitcost and bringing semantic noises. To suppress this, we explicitly regularize the non-semantic entropy of the compressed video in the MVM token space. The proposed framework is instantiated as a simple Semantic-Mining-then-Compression (SMC) model. Furthermore, we extend SMC as an advanced SMC++ model from several aspects. First, we equip it with a masked motion prediction objective, leading to better temporal semantic learning ability. Second, we introduce a Transformer-based compression module, to improve the semantic compression efficacy. Considering that directly mining the complex redundancy among heterogeneous features in different coding stages is non-trivial, we introduce a compact blueprint semantic representation to align these features into a similar form, fully unleashing the power of the Transformer-based compression module. Extensive results demonstrate the proposed SMC and SMC++ models show remarkable superiority over previous traditional, learnable, and perceptual quality-oriented video codecs, on three video analysis tasks and seven datasets. Yuan Tian 0017, Xiaoyue Ling, Cong Geng, Qiang Hu 0003, Guo Lu, Guangtao Zhai |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2026 | Developing Evolving Adaptability in Biological Intelligence: A Novel Biologically-Inspired Continual Learning Model for Video Saliency PredictionabstractIn the era of deep learning, video saliency prediction task still remains major challenge due to the issue of catastrophic forgetting during feature learning. Most prior works commonly employ generative replay strategies to generate pseudo-samples from previous tasks, enabling them to recall the data distribution. However, scaling up generative replay to accommodate class-incremental and task-incremental settings poses challenges, as generated data with low quality can severely deteriorate performance. Additionally, existing advances mainly focus on preserving memory stability to alleviate catastrophic forgetting, but they remain difficult to flexibly adapt to incremental changes in dynamic scenes. To achieve a better balance between memory stability and learning plasticity, we propose a novel biologically-inspired continual learning (BICL) model tailored to effectively predict human attention in dynamic scenes while mitigate catastrophic forgetting. In particular, inspired by the function of the hippocampus in the human neural system, we elaborately design a visual saliency memory bank module to explicitly store and retrieve representative features from previous tasks. Furthermore, drawing inspiration from the Drosophila $\gamma$γMB system, we propose an active forgetting strategy equipped with multiple parallel adaptive learner modules, which can appropriately attenuate old memories in parameter distribution to enhance learning plasticity to adapt to new tasks, and accordingly to ensure compatibility among multiple learners. Notably, without compromising the performance of old tasks, our proposed model can achieve a better trade-off between memory stability and learning plasticity. Through extensive experiments on several benchmark datasets, our model not only enhances performance in task-incremental settings, but also potentially provides deep insights into neurological adaptive mechanisms. Dandan Zhu 0001, Kaiwei Zhang, Kun Zhu 0024, Xiongkuo Min, Guangtao Zhai, Xiaokang Yang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2026 | DHQA-4D: A large-scale dataset and LMM-based metric for dynamic 4D digital human quality assessment
Sijing Wu, Yucheng Zhu, Huiyu Duan, Wei Sun 0029, Xiongkuo Min, Guangtao Zhai |
Pattern Recognit. | 8 |
| 2026 | HMSR: Hypercomplex-guided mamba for fine-texture coupling in single image super-resolution
Fengqian Sun, Qiqi Kou, Deqiang Cheng 0001, Guangtao Zhai, Wenjun Zhang 0001 |
Pattern Recognit. | 7 |
| 2026 | Phase Transition Hypothesis of Perception and Cognition in the Visually ImpairedabstractPerception and cognition are core processes that transform external sensory signals into internal representations for knowledge construction and understanding, and in visually impaired individuals, this transformation is reorganized through auditory and tactile feedback. To explain how perceptual information evolves into stable cognitive representations under limited sensory bandwidth, this study proposes Phase Transition Hypothesis of Perception and Cognition. The proposed hypothesis models the perceptual–cognitive process as a dynamic phase transition, in which sensory information evolves from fragmented perception into organized cognition. To counteract perceptual bias induced by information collapse, the Perceptual Dynamic Optimization Mechanism adaptively regulates sensory deviations to stabilize the perceptual–cognitive transition, whereas the Cognitive Potential Model, derived from the Free-Energy Principle, elucidates how stable and self-organizing cognition emerges from this dynamic process. A cognitive simulation system and a blind writing navigation experiment are conducted to validate the hypothesis. Experiments demonstrate the proposed adaptive correction of perceptual bias and the phase transition mechanism from perception to cognition. Ji-Feng Luo, Zhengqiang Jiang, Jian Zhang 0060, Guangtao Zhai, Menghan Hu |
IEEE Signal Process. Lett. | 7 |
| 2026 | UNQA: Unified No-Reference Quality Assessment for Audio, Image, Video, and Audio-Visual ContentabstractAs multimedia data flourishes on the Internet, quality assessment (QA) of multimedia data becomes paramount for digital media applications. Since multimedia data includes multiple modalities including audio, image, video, and audio-visual (A/V) content, researchers have developed a range of QA methods to evaluate the quality of different modality data. While they exclusively focus on addressing the single modality QA issues, a unified QA model that can handle diverse media across multiple modalities is still missing, whereas the latter can better resemble human perception behaviour and also have a wider range of applications. In this paper, we propose the Unified No-reference Quality Assessment model (UNQA) for audio, image, video, and A/V content, which tries to train a single QA model across different media modalities. To tackle the issue of inconsistent quality scales among different QA databases, we develop a multi-modality strategy to jointly train UNQA on multiple QA databases. Based on the input modality, UNQA selectively extracts the spatial features, motion features, and audio features, and calculates a final quality score via the four corresponding modality regression modules. Compared with existing QA methods, UNQA has two advantages: 1) the multi-modality training strategy makes the QA model learn more general and robust quality-aware feature representation as evidenced by the superior performance of UNQA compared to state-of-the-art QA methods. 2) UNQA reduces the number of models required to assess multimedia data across different modalities. and is friendly to deploy to practical applications. Code are available at https://github.com/charlotte9524/UNQA. Yuqin Cao, Xiongkuo Min, Wei Sun 0029, Long Ye, Weisi Lin, Guangtao Zhai |
IEEE Trans. Circuits Syst. Video Technol. | 7 |
| 2026 | A Lightweight Deep and Wide Network for Image-Based Detection of Industrial Waste GasabstractDue to inadequate monitoring, key pollutants (e.g., PM2.5, VOCs, etc) very possibly leak into atmosphere, thus to endanger the long-term and short-term life safety of people that work and live in the environment. Therefore, it is imperative to effectively and efficiently detect the leakage of industrial waste gas, for the purpose of timely lowering the risk of pollution and explosions. To solve such a problem, we in this paper propose a new lightweight deep and wide network (LdwNet) for detecting the leakage of industrial waste gas from an image, which brings about the two main merits: 1) Compensating for the deficiencies of sensor-based detection methods, which can accurately detect the leakage of waste gas and even measure its concentrations but require to seek leakage sources beforehand; 2) Overcoming the shortcomings of image-based detection methods, which leverage DNN-based recognition technologies and usually suffer from low efficacy, low efficiency and high energy consumption during the model training and inference. To specify, the proposed LdwNet is developed by simulating human perception, motivated by the method which detects the leakage of industrial waste gas from surveillance images with the human observation and judgement. First, based on the inspiration that the human eyes are highly sensitive to horizontal and vertical stimuli, we construct a novel lightweight parallel-series-stripe (PS2) module to validly extract features with very few parameters. Second, to fully exploit deep and shallow features for fusing the global and local information, we extend the PS2 module as a backbone along both the deep and wide directions to build the multi-channel network. Third, to achieve effective, efficient and low-carbon detection in model running, we constraint the extended PS2 modules with parameter sharing to prodigiously reduce the model parameters and thus to make the proposed model ultra-lightweight. Experiments on the datasets of carbon particulate matters and ethylene leakage prove that our LdwNet with ten thousand parameters outperforms the state-of-the-art models with millions of parameters in detection accuracy and implementation cost, and this renders our proposed LdwNet more suitable for real industrial applications. Ke Gu 0001, Hongyan Liu 0004, Jingchao Cao, Lai-Kuan Wong, Junfei Qiao 0001, Guangtao Zhai, Wenjun Zhang 0001, Weisi Lin, Sam Kwong |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2026 | Subjective and Objective Quality Assessment of Display Content VideosabstractDisplay quality assessment plays a crucial role in evaluating the performance of display devices. However, existing video quality assessment methods primarily target compression-related distortions, failing to capture display-specific degradations including definition loss, color distortions, and motion artifacts that critically affect user subjective experiences during video playback. To address these limitations, we develop a specialized video dataset, namely Video Displaying Quality Assessment Dataset (VDQA), constructed using a DSLR camera with standardized parameter optimization of exposure settings (aperture, ISO sensitivity, and shutter speed). VDQA comprises 250 high-resolution video clips covering diverse content categories, providing a robust foundation for evaluating display devices across multiple quality dimensions. Additionally, we propose a deep learning-based model specifically designed for display quality assessment that employs three complementary pathways to independently evaluate definition, color fidelity, and motion quality. The model integrates Canny edge detection for explicit sharpness measurement, a color attention mechanism to enhance sensitivity to display color reproduction characteristics, and temporal modeling for motion artifact assessment. Experimental results demonstrate that the proposed model achieves superior performance in reflecting user subjective experiences for display content videos compared to state-of-the-art methods, with significant improvements in both color fidelity assessment and definition evaluation. Fangfang Lu, Huiqun Yu, Kaiwei Zhang, Wei Sun 0029, Xiongkuo Min, Guangtao Zhai |
IEEE Trans. Circuits Syst. Video Technol. | 7 |
| 2026 | Surveillance Facial Image Quality Assessment: A Multi-Dimensional Dataset and Lightweight ModelabstractSurveillance facial images are often captured under unconstrained conditions, resulting in severe quality degradation due to factors such as low resolution, motion blur, occlusion, and poor lighting. Although recent face restoration techniques applied to surveillance cameras can significantly enhance visual quality, they often compromise fidelity (i.e., identity-preserving features), which directly conflicts with the primary objective of surveillance images -- reliable identity verification. Existing facial image quality assessment (FIQA) predominantly focus on either visual quality or recognition-oriented evaluation, thereby failing to jointly address visual quality and fidelity, which are critical for surveillance applications. To bridge this gap, we propose the first comprehensive study on surveillance facial image quality assessment (SFIQA), targeting the unique challenges inherent to surveillance scenarios. Specifically, we first construct SFIQA-Bench, a multi-dimensional quality assessment benchmark for surveillance facial images, which consists of 5,004 surveillance facial images captured by three widely deployed surveillance cameras in real-world scenarios. A subjective experiment is conducted to collect six dimensional quality ratings, including noise, sharpness, colorfulness, contrast, fidelity and overall quality, covering the key aspects of SFIQA. Furthermore, we propose SFIQA-Assessor, a lightweight multi-task FIQA model that jointly exploits complementary facial views through cross-view feature interaction, and employs learnable task tokens to guide the unified regression of multiple quality dimensions. The experiment results on the proposed dataset show that our method achieves the best performance compared with the state-of-the-art general image quality assessment (IQA) and FIQA methods, validating its effectiveness for real-world surveillance applications. Yanwei Jiang, Wei Sun 0029, Yingjie Zhou 0003, Yuqin Cao, Jun Jia, Sijing Wu, Dandan Zhu 0001, Xiongkuo Min, Guangtao Zhai |
IEEE Trans. Circuits Syst. Video Technol. | 11 |
| 2026 | AGHI-QA: A Subjective-Aligned Dataset and Metric for AI-Generated Human ImagesabstractThe rapid development of text-to-image (T2I) generation approaches has attracted extensive interest in evaluating the quality of generated images, leading to the development of various quality assessment methods for general-purpose T2I outputs. However, existing image quality assessment (IQA) methods are limited to providing global quality scores, failing to deliver fine-grained perceptual evaluations for structurally complex subjects like humans, which is a critical challenge considering the frequent anatomical and textural distortions in AI-generated human images (AGHIs). To address this gap, we introduce AGHI-QA, a large-scale benchmark specifically designed for quality assessment of AGHIs. The dataset comprises 4, 000 images generated from 400 carefully crafted text prompts using 10 state-of-the-art T2I models. We conduct a systematic subjective study to collect multidimensional annotations, including perceptual quality scores, text-image correspondence scores, visible and distorted body part labels. Based on AGHI-QA, we evaluate the strengths and weaknesses of current T2I methods in generating human images from multiple dimensions. Furthermore, we propose AGHI-Assessor, a novel quality metric that integrates the large multimodal model (LMM) with domain-specific human features for precise quality prediction and identification of visible and distorted body parts in AGHIs. Extensive experimental results demonstrate that AGHI-Assessor showcases state-of-the-art performance, significantly outperforming existing IQA methods in multidimensional quality assessment and surpassing leading LMMs in detecting structural distortions in AGHIs. Sijing Wu, Wei Sun 0029, Yucheng Zhu, Huiyu Duan, Xiongkuo Min, Guangtao Zhai |
IEEE Trans. Circuits Syst. Video Technol. | 9 |
| 2026 | Video Respiratory Rate Measurement in Walking Scenarios Using Multi-Strategy Adaptive Denoising
Gan Pei, Junhao Ning, Chenrui Niu, Siqiong Yao, Menghan Hu, Guangtao Zhai |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2026 | Mitigating Low-Level Visual Hallucinations Requires Self-Awareness: Database, Model, and Training StrategyabstractThe rapid development of multimodal large language models has resulted in remarkable advancements in visual perception and understanding, consolidating several tasks into a single visual question-answering framework. However, these models are prone to hallucinations, which limit their reliability as artificial intelligence systems. While this issue is extensively researched in natural language processing and image captioning, there remains a lack of investigation of hallucinations in Low-level Visual Perception and Understanding (HLPU), especially in the context of image quality assessment tasks. We consider that these hallucinations arise from an absence of clear self-awareness within the models. To address this issue, we first introduce the HLPU instruction database, the first instruction database specifically focused on hallucinations in low-level vision tasks. This database contains approximately 200K question-answer pairs and comprises four subsets, each covering different types of instructions. Subsequently, we propose the Self-Awareness Failure Elimination (SAFEQA) model, which utilizes image features, salient region features and quality features to improve the perception and comprehension abilities of the model in low-level vision tasks. Furthermore, we propose the Enhancing Self-Awareness Preference Optimization (ESA-PO) framework to increase the model’s awareness of knowledge boundaries, thereby mitigating the incidence of hallucination. Finally, we conduct comprehensive experiments on low-level vision tasks, with the results demonstrating that our proposed method significantly enhances self-awareness of the model in these tasks and reduces hallucinations. Notably, our proposed method improves both accuracy and self-awareness of the proposed model and outperforms close-source models in terms of various evaluation metrics. This research contributes to the advancement of self-awareness capabilities in multimodal large language models, particularly for low-level visual perception and understanding tasks. Xiongkuo Min, Yuqin Cao, Guangtao Zhai |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2026 | Quality Assessment and Distortion-Aware Saliency Prediction for AI-Generated Omnidirectional ImagesabstractWith the rapid advancement of Artificial Intelligence Generated Content (AIGC) techniques, AI generated images (AIGIs) have attracted widespread attention, among which AI generated omnidirectional images (AIGODIs) hold significant potential for Virtual Reality (VR) and Augmented Reality (AR) applications. AI generated omnidirectional images exhibit unique quality issues, however, research on the quality assessment and optimization of AI-generated omnidirectional images is still lacking. To this end, this work first studies the quality assessment and distortion-aware saliency prediction problems for AIGODIs, and further presents a corresponding optimization process. Specifically, we first establish a comprehensive database to reflecthumanfeedback for AI-generatedomnidirectionals, termed OHF2024, which includes both subjective quality ratings evaluated from three perspectives and distortion-aware salient regions. Based on the constructed OHF2024 database, we propose two models with shared encoders based on the BLIP-2 model to evaluate the human visual experience and predict distortion-aware saliency for AI-generated omnidirectional images, which are named as BLIP2OIQA and BLIP2OISal, respectively. Finally, based on the proposed models, we present an automatic optimization process that utilizes the predicted visual experience scores and distortion regions to further enhance the visual quality of an AI-generated omnidirectional image. Extensive experiments show that our BLIP2OIQA model and BLIP2OISal model achieve state-of-the-art (SOTA) results in the human visual experience evaluation task and the distortion-aware saliency prediction task for AI generated omnidirectional images, and can be effectively used in the optimization process. The database and codes will be released on https://github.com/IntMeGroup/AIGCOIQA to facilitate future research. Huiyu Duan, Jing Liu 0002, Menghan Hu, Xiongkuo Min, Guangtao Zhai, Patrick Le Callet |
IEEE Trans. Circuits Syst. Video Technol. | 7 |
| 2026 | LMVQ: Label-Free Metric-Learning for General AI-Generated Video Quality AssessmentabstractThe recent rapid development of video generation technology has led to a significant demand for quality assessment of the latest AI-generated videos. However, current supervised approaches depend on expensive and quickly outdated human scores, and label-free methods overlook the general distortions of AI-generated videos. To address these limitations, we introduce LMVQ, a Label-free Metric-learning framework for general AI-generated Video Quality assessment of three dimensions, spatial, temporal, and alignment. The LMVQ is the first to introduce sample degradations specially designed for AIGC-specific distortions, and constructs a comprehensive training set through two complementary sample generation strategies. It then employs two synergistic modules, the Intra-Quality Token Transformer (IQ-Trans), which explicitly refines dimension-specific quality representations, and the Inter-Quality Mixture of Experts (IQ-MoE), which fuses interactions across multiple quality dimensions. Finally, a Multi-Proxy Metric-Learning (MPML) strategy aligns the learned representations with multi-dimensional quality scores and constrains the model to learn discriminative quality-aware representations. Extensive experiments on four public AIGC-VQA benchmarks show that MPML outperforms previous label-free methods by over 20%, and greatly narrows the gap with supervised methods. This provides a scalable, adaptive foundation for evaluating the ever-evolving quality of AI-generated videos. Xinyue Li 0001, Wei Sun 0029, Xiaohong Liu 0001, Xiongkuo Min, Guangtao Zhai |
IEEE Trans. Circuits Syst. Video Technol. | 7 |
| 2026 | Unveiling the Modular Design of Text-to-Video Quality AssessmentabstractArtificial intelligence generated content (AIGC) are reshaping digital media creation, with text-to-video (T2V) generation emerging as one of its most powerful and widely-used techniques. Despite rapid progress, videos generated by T2V models still suffer from issues such as unrealistic spatial details, temporal inconsistencies, and content misalignment with the input textual prompts. It is of high importance to develop computational video quality assessment (VQA) models for T2V videos to ensure a favorable quality-of-experience (QoE) for end users. Towards comprehensive quality evaluation of AI-generated videos, modern T2V quality assessment (T2V QA) models typically integrate multiple modules that excel in capturing different and complementary quality-aware features. In this paper, we categorize the constituent modules of modern T2V QA models into four types: base quality evaluators, spatial perception modules, temporal perception modules, and text-video alignment modules. Within this framework, we systematically evaluate and compare the relative strengths and weaknesses of candidate models. Through experiments on multiple T2V datasets, we verify that the top-performing models from this competition demonstrate very competitive performance against existing VQA methods. The resulting framework is agnostic to specific architectural designs and can be continuously refined by integrating advancements from each of its constituent modules, making it well-suited for adapting T2V QA methods to the fast-evolving T2V generation techniques. Our source code is available at https://github.com/CH053N0N3/mineBVQA. Bingkun Zheng, Weixia Zhang, Guangtao Zhai, Xiaokang Yang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2026 | Future Fixation Sequence Prediction for Audio-Visual 360° VideosabstractFuture fixation sequence prediction plays a crucial role in various aspects of virtual reality content production, transmission, rendering, and display. Accurate prediction of future fixation sequence can significantly enhance the quality of user experience, particularly in resource-constrained scenarios. In this paper, we present a novel framework for predicting future fixation sequence and achieves state-of-the-art performance. Specifically, the anti-projection-distortion FoV patch extraction algorithm is proposed to mitigate projection distortions. A comprehensive contextual representation is then constructed by integrating multiple data sources, including visual and audio information, historical fixation sequence, user identity, timestamp, and positional embeddings. The transformer-based predictor is proposed to perform the future fixation sequence prediction based on the integrated contextual representations. Additionally, we propose a framework that effectively utilizes saliency information as supervision and conduct saliency contrastive distillation during the training phase, eliminating the need for saliency data during inference. Overall, by integrating anti-projection-distortion and multimodal representations, along with key embeddings, a dedicated predictor, and contrastive distillation, our approach is designed to accurately predict future fixation sequences. Extensive experiments validate the effectiveness of our framework, demonstrating its superior performance in fixation prediction tasks. Yucheng Zhu, Guangtao Zhai, Xiongkuo Min, Huiyu Duan, Xiaokang Yang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2026 | Multi-Dimensional Quality Assessment for Single-Image-to-3D Contents: Dataset and ModelabstractThe rapid advancement of AI generation technologies has led to the widespread use of AI-generated multimedia content, including images, videos, and 3D contents, across various applications. While significant progress has been made in quality evaluation for 2D content, evaluating the quality of 3D content synthesized from single image remains an underexplored problem. To bridge this gap, we introduce the first comprehensive subjective evaluation database tailored for assessing the quality of 3D content generated from single image. Our database, named AIGC-SI23DCQA, includes three distinct categories of input images, i.e., realistic images, AI-generated images, and computer graphic (CG) images, with 100 images in each category. Using five representative single-image-to-3D algorithms, we produce 1,500 3D contents and collect 94,500 annotations across three quality dimensions, including texture fidelity, shape accuracy, and overall quality. Based on the constructed database, we first benchmark and evaluate the performance of existing quality assessment methods revealing their limitations in addressing this novel task. Thus, we further propose a novel objective quality assessment method, termed I3DQA, for effective single-image-to-3D content quality assessment. Specifically, I3DQA first extracts the reference features from the source image, and the multi-modal features from the generated 3D content, including the projected video, patches, and large-multimodal model (LMM) features. These features are integrated through symmetric transformer blocks, enabling effective quality-related feature fusion and score prediction. Extensive experiments demonstrate the superior performance of our method and validate the effectiveness of its components. This work provides a foundational resource and a robust framework for advancing research in this emerging field, and our database and model are released at https://github.com/ZedFu/SI23DCQA. Huiyu Duan, Jing Liu 0002, Yun Liu 0009, Xiaohong Liu 0001, Jia Wang 0004, Xiongkuo Min, Patrick Le Callet, Guangtao Zhai |
IEEE Trans. Image Process. | 10 |
| 2026 | Self-Supervised Unfolding Network With Shared Reflectance Learning for Low-Light Image EnhancementabstractRecently, incorporating Retinex theory with unfolding networks has attracted increasing attention in the low-light image enhancement field. However, existing methods have two limitations, i.e., ignoring the modeling of the physical prior of Retinex theory and relying on a large amount of paired data. To advance this field, we propose a novel self-supervised unfolding network, named S2UNet, for the LIE task. Specifically, we formulate a novel optimization model based on the principle that content-consistent images under different illumination should share the same reflectance. The model simultaneously decomposes two illumination-different images into a shared reflectance component and two independent illumination components. Due to the absence of the normal-light image, we process the low-light image with gamma correction to create the illumination-different image pair. Then, we translate this model into a multi-stage unfolding network, in which each stage alternately optimizes the shared reflectance component and the respective illumination components of the two images. During progressive multi-stage optimization, the network inherently encodes the reflectance consistency prior by jointly estimating an optimal reflectance across varying illumination conditions. Finally, considering the presence of noise in low-light images and to suppress noise amplification, we propose a self-supervised denoising mechanism. Extensive experiments on nine benchmark datasets demonstrate that our proposed S2UNet outperforms state-of-the-art unsupervised methods in terms of both quantitative metrics and visual quality, while achieving competitive performance compared to supervised methods. The source code will be available at https://github.com/J-Liu-DL/S2UNet. Jia Liu 0025, Yu Luo 0004, Guanghui Yue 0001, Jie Ling 0002, Chia-Wen Lin, Guangtao Zhai, Wei Zhou 0021 |
IEEE Trans. Image Process. | 7 |
| 2026 | Infrared Image Quality Estimation With Node-to-Graph RegressionabstractBy comparison with the commonly seen visible light images that can be effectively characterized within a Euclidean space, infrared images have non-Euclidean characteristics since their pixels contain rich thermal radiation information, such as heat distribution, surface temperature and thermal radiation. Considering the advantages of Graph Convolutional Networks (GCNs) in processing non-Euclidean data, this study proposes to introduce the GCNs to estimate the quality of infrared images by developing the Node-to-Graph Regression (NGR) model. To specify, the proposed NGR model is composed of two main steps, namely network establishment and network training. In the first step, following the classical researches of image quality estimation that include local distortion measurement followed by pooling for inferring the image quality score, this study captures the local distortion of the input infrared images by stacking up a set of Vision Graph (VSG) blocks to generate one node map, and then conducts the weighted pooling method on the node map to yield the graph output as the estimated quality score. In the second step, for enhancing the model's performance and generalization ability in the network training process, this study implements the node regression with the big data pre-training method to raise the local distortion extraction ability in a broad range of image scenarios and distortion intensities, and then performs the graph regression by using the knowledge distillation method to reduce the over-fitting risk. Using the largest-size infrared image quality evaluation database (I2QED), this study compared the proposed NGR model with three dozen mainstream and state-of-the-art competitors, and results showed that our proposed NGR model achieved the optimal performance. Ke Gu 0001, Hongyan Liu 0004, Yubin Gao, Chen Wang 0019, Lai-Kuan Wong, Weisi Lin, Guangtao Zhai, Wenjun Zhang 0001, Daniel Thalmann |
IEEE Trans. Multim. | 7 |
| 2026 | TFFN: Three-Branch Feature Fusion Network for Stereoscopic Omnidirectional Image Quality AssessmentabstractStereoscopic omnidirectional image (SOI) has both omnidirectional and stereoscopic perception features. Many previous models have proved the viewport characteristics and stereoscopic visual features are crucial for quality perception of SOI. However, effective monocular and binocular visual features extraction and fusion are difficult due to the size of SOI and inaccuracy of feature representation. In this paper, we proposed a three-branch feature fusion network (TFFN) by fusing two-stream binocular visual features and the important monocular features based on the viewport perspective. The hierarchical fusion module is first designed to fuse effective binocular visual features from different semantic scales, and the pseudo-difference information extraction module is built to obtain the accuracy monocular visual features to complement the binocular visual features. Finally, the above monocular and binocular visual features are fused together to measure the quality of SOI. The comparison experiments are conducted on three public datasets and the analysis of the results demonstrate the effectiveness of the proposed method. Yun Liu 0009, Daoxin Fan, Huiyu Duan, Peiguang Jing, Guanghui Yue 0001, Guangtao Zhai |
IEEE Trans. Multim. | 7 |
| 2026 | SingingHead: A Large-Scale 4D Dataset for Singing Head AnimationabstractSinging, as a common facial movement second only to talking, can be regarded as a universal language across ethnicities and cultures, plays an important role in emotional communication, art, and entertainment. However, it is often overlooked in the field of audio-driven 3D facial animation due to the lack of singing head datasets and the domain gap between singing and talking in rhythm and amplitude. To this end, we collect a large-scale high-quality multi-modal singing head dataset,SingingHead, which consists of more than 27 hours of synchronized singing video, 3D facial motion, singing audio, and background music from 76 individuals and 8 types of music. Along with the SingingHead dataset, we benchmark existing audio-driven 3D facial animation methods and 2D talking head methods on the singing task. Existing 3D facial animation methods and 2D talking head methods fail to produce satisfactory singing results. Focusing on the 3D singing head animation, we first utilize the proposed singing-specific dataset to retrain the 3D facial animation methods, resulting in substantial performance improvements. Besides, considering the absence of background music and the slow generation speed of existing methods, we propose a simple but efficient non-autoregressive VAE-based framework with background music as an input signal to generate diverse and accurate 3D singing facial motions in real time. Extensive experiments demonstrate the significance of the SingingHead dataset in promoting the development of singing head animation. The dataset is released for research purposes at:https://wsj-sjtu.github.io/SingingHead/. Sijing Wu, Weitian Zhang, Jun Jia, Yucheng Zhu, Yichao Yan, Guangtao Zhai, Xiaokang Yang 0001 |
IEEE Trans. Multim. | 7 |
| 2026 | FreeQR: Free Lunch for Aesthetic QR Codes Emerging From the Latent Space in Diffusion ModelsabstractIn the modern digital age, Quick Response (QR) codes serve as a critical interface for bridging the physical and virtual worlds, widely utilized in multimedia applications. However, traditional binary QR codes often lack the visual appeal desired in contexts. Aesthetic QR codes address this limitation by enabling the customization of QR code patterns to enhance visual attractiveness while retaining compatibility with standard QR decoders. Previous works have explored the use of diffusion models for generating such codes but often require extensive training of ControlNets and face challenges in maintaining scannability. To address these issues, we present FreeQR, a streamlined and effective approach that enables the stable generation of QR code images with diffusion models. Our methodology involves the strategic fusion between the specific channel in the latent space of the denoising process with the noised latent representations of the QR blueprint image at corresponding timesteps. This ensures that the generated images adhere to the brightness distribution required for effective scanning while achieving a balance between aesthetics and functionality. Additionally, we introduce gradient guidance based on scanning errors directly in the latent space, enabling the generation of scannable QR codes in seconds without additional model parameters. Experimental results demonstrate that FreeQR significantly enhances the aesthetics and scannability of QR codes compared to existing methods, making it a lightweight and efficient solution for multimedia applications. Yiwei Yang 0007, Jun Jia, Zheyuan Liu 0011, Zhongpai Gao, Wei Sun 0029, Guangtao Zhai |
IEEE Trans. Multim. | 7 |
| 2026 | MEScan360: A Memory-Enhanced Scanpath Prediction Model for Omnidirectional Images
Dandan Zhu 0001, Kaiwei Zhang, Kun Zhu 0024, Xiongkuo Min, Guangtao Zhai |
ACM Trans. Multim. Comput. Commun. Appl. | 7 |
| 2025 | Medical Manifestation-Aware De-IdentificationabstractFace de-identification (DeID) has been widely studied for common scenes, but remains under-researched for medical scenes, mostly due to the lack of large-scale patient face datasets. In this paper, we release MeMa, consisting of over 40,000 photo-realistic patient faces. MeMa is re-generated from massive real patient photos. By carefully modulating the generation and data-filtering procedures, MeMa avoids breaching real patient privacy, while ensuring rich and plausible medical manifestations. We recruit expert clinicians to annotate MeMa with both coarse- and fine-grained labels, building the first medical-scene DeID benchmark. Additionally, we propose a baseline approach for this new medical-aware DeID task, by integrating data-driven medical semantic priors into the DeID procedure. Despite its conciseness and simplicity, our approach substantially outperforms previous ones. Yuan Tian 0017, Guangtao Zhai |
AAAI | 3 |
| 2025 | VRVVC: Variable-Rate NeRF-Based Volumetric Video CompressionabstractNeural Radiance Field (NeRF)-based volumetric video has revolutionized visual media by delivering photorealistic Free-Viewpoint Video (FVV) experiences that provide audiences with unprecedented immersion and interactivity. However, the substantial data volumes pose significant challenges for storage and transmission. Existing solutions typically optimize NeRF representation and compression independently or focus on a single fixed rate-distortion (RD) tradeoff. In this paper, we propose VRVVC, a novel end-to-end joint optimization variable-rate framework for volumetric video compression that achieves variable bitrates using a single model while maintaining superior RD performance. Specifically, VRVVC introduces a compact tri-plane implicit residual representation for inter-frame modeling of long-duration dynamic scenes, effectively reducing temporal redundancy. We further propose a variable-rate residual representation compression scheme that leverages a learnable quantization and a tiny MLP-based entropy model. This approach enables variable bitrates through the utilization of predefined Lagrange multipliers to manage the quantization error of all latent representations. Finally, we present an end-to-end progressive training strategy combined with a multi-rate-distortion loss function to optimize the entire framework. Extensive experiments demonstrate that VRVVC achieves a wide range of variable bitrates within a single model and surpasses the RD performance of existing methods across various datasets. Qiang Hu 0003, Houqiang Zhong, Zihan Zheng, Xiaoyun Zhang 0001, Zhengxue Cheng, Li Song 0001, Guangtao Zhai, Yanfeng Wang 0001 |
AAAI | 7 |
| 2025 | Textured Mesh Saliency: Bridging Geometry and Texture for Human Perception in 3D GraphicsabstractTextured meshes significantly enhance the realism and detail of objects by mapping intricate texture details onto the geometric structure of 3D models. This advancement is valuable across various applications, including entertainment, education, and industry. While traditional mesh saliency studies focus on non-textured meshes, our work explores the complexities introduced by detailed texture patterns. We present a new dataset for textured mesh saliency, created through an innovative eye-tracking experiment in a six degrees of freedom (6-DOF) VR environment. This dataset addresses the limitations of previous studies by providing comprehensive eye-tracking data from multiple viewpoints, thereby advancing our understanding of human visual behavior and supporting more accurate and effective 3D content creation. Our proposed model predicts saliency maps for textured mesh surfaces by treating each triangular face as an individual unit and assigning a saliency density value to reflect the importance of each local surface region. The model incorporates a texture alignment module and a geometric extraction module, combined with an aggregation module to integrate texture and geometry for precise saliency prediction. We believe this approach will enhance the visual fidelity of geometric processing while ensuring computational efficiency, essential for real-time rendering and high-detail applications such as VR and gaming. Kaiwei Zhang, Dandan Zhu 0001, Xiongkuo Min, Guangtao Zhai |
AAAI | 4 |
| 2025 | Low-Light Image Enhancement via Generative Perceptual PriorsabstractAlthough significant progress has been made in enhancing visibility, retrieving texture details, and mitigating noise in Low-Light (LL) images, the challenge persists in applying current Low-Light Image Enhancement (LLIE) methods to real-world scenarios, primarily due to the diverse illumination conditions encountered. Furthermore, the quest for generating enhancements that are visually realistic and attractive remains an underexplored realm. In response to these challenges, we present a novel LLIE framework with the guidance of Generative Perceptual Priors (GPP-LLIE) derived from vision-language models (VLMs). Specifically, we first propose a pipeline that guides VLMs to assess multiple visual attributes of the LL image and quantify the assessment to output the global and local perceptual priors. Subsequently, to incorporate these generative perceptual priors to benefit LLIE, we introduce a transformer-based backbone in the diffusion process, and develop a new layer normalization (GPP-LN) and an attention mechanism (LPP-Attn) guided by global and local perceptual priors. Extensive experiments demonstrate that our model outperforms current SOTA methods on paired LL datasets and exhibits superior generalization on real-world data. Han Zhou 0003, Wei Dong 0011, Xiaohong Liu 0010, Yulun Zhang 0001, Guangtao Zhai, Jun Chen 0005 |
AAAI | 5 |
| 2025 | Redundancy Principles for MLLMs BenchmarksabstractZicheng Zhang, Xiangyu Zhao, Xinyu Fang, Chunyi Li, Xiaohong Liu, Xiongkuo Min, Haodong Duan, Kai Chen, Guangtao Zhai. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Xinyu Fang, Chunyi Li 0001, Xiaohong Liu 0001, Xiongkuo Min, Haodong Duan, Kai Chen 0026, Guangtao Zhai |
ACL (1) | 9 |
| 2025 | OmniAlign-V: Towards Enhanced Alignment of MLLMs with Human PreferenceabstractXiangyu Zhao, Shengyuan Ding, Zicheng Zhang, Haian Huang, Maosongcao Maosongcao, Jiaqi Wang, Weiyun Wang, Xinyu Fang, Wenhai Wang, Guangtao Zhai, Hua Yang, Haodong Duan, Kai Chen. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Shengyuan Ding, Haian Huang, Maosongcao, Jiaqi Wang 0003, Weiyun Wang, Xinyu Fang, Wenhai Wang, Guangtao Zhai, Hua Yang 0001, Haodong Duan, Kai Chen 0026 |
ACL (1) | 10 |
| 2025 | TF-Fusion: Time-Frequency Feature Fusion for High-Quality Single-Angle Ultrasound ImagingabstractSingle-angle plane wave (SAPW) ultrasound imaging has gained significant attention in ultrafast and wearable ultrasound systems due to its high frame rate and hardware efficiency. However, the lack of transmit diversity leads to severe degradation in image quality, limiting its clinical utility. In contrast, Coherent Plane Wave Compounding (CPWC) improves spatial resolution and contrast by aggregating data from multiple transmission angles, but at the expense of a substantially reduced frame rate. To overcome this trade-off, we propose TF-Fusion, a novel Time-Frequency Domain Feature Fusion framework that enhances SAPW imaging by jointly exploiting both time- and frequency-domain features extracted from raw in-phase/quadrature (IQ) data. Specifically, our model utilizes a dual-branch encoder to learn compact and complementary representations from each domain, followed by a learnable fusion module that adaptively integrates the multi-domain information. This design facilitates effective noise and artifact suppression while preserving fine anatomical structures. We evaluate TF-Fusion on two public benchmarks—PICMUS and CUBDL—and demonstrate that our method significantly improves image resolution and contrast, achieving performance comparable to multi-angle CPWC methods. Notably, TF-Fusion maintains the high temporal resolution of SAPW, making it well-suited for real-time and resource-constrained ultrasound imaging scenarios. Yankun Cao, Baolin Sun, Guangtao Zhai, Li-Zhen Cui 0001, Zhi Liu 0004 |
BIBM | 7 |
| 2025 | FineVQ: Fine-Grained User Generated Content Video Quality AssessmentabstractThe rapid growth of user-generated content (UGC) videos has produced an urgent need for effective video quality assessment (VQA) algorithms to monitor video quality and guide optimization and recommendation procedures. However, current VQA models generally only give an overall rating for a UGC video, which lacks fine-grained labels for serving video processing and recommendation applications. To address the challenges and promote the development of UGC videos, we establish the first large-scale Fine-grained Video quality assessment Database, termed FineVD, which comprises 6104 UGC videos with fine-grained quality scores and descriptions across multiple dimensions. Based on this database, we propose a Fine-grained Video Quality assessment (FineVQ) model to learn the fine-grained quality of UGC videos, with the capabilities of quality rating, quality scoring, and quality attribution. Extensive experimental results demonstrate that our proposed FineVQ can produce fine-grained video-quality results and achieve state-of-the-art performance on FineVD and other commonly used UGC-VQA datasets. Both FineVD and FineVQ are publicly available at: https://github.com/IntMeGroup/FineVQ. Huiyu Duan, Qiang Hu 0003, Zitong Xu, Lu Liu 0005, Xiongkuo Min, Chunlei Cai, Tianxiao Ye, Xiaoyun Zhang 0001, Guangtao Zhai |
CVPR | 11 |
| 2025 | 4DGC: Rate-Aware 4D Gaussian Compression for Efficient Streamable Free-Viewpoint Videoabstract3D Gaussian Splatting (3DGS) has substantial potential for enabling photorealistic Free-Viewpoint Video (FVV) experiences. However, the vast number of Gaussians and their associated attributes poses significant challenges for storage and transmission. Existing methods typically handle dynamic 3DGS representation and compression separately, neglecting motion information and the rate-distortion (RD) trade-off during training, leading to performance degradation and increased model redundancy. To address this gap, we propose 4DGC, a novel rate-aware 4D Gaussian compression framework that significantly reduces storage size while maintaining superior RD performance for FVV. Specifically, 4DGC introduces a motion-aware dynamic Gaussian representation that utilizes a compact motion grid combined with sparse compensated Gaussians to exploit inter-frame similarities. This representation effectively handles large motions, preserving quality and reducing temporal redundancy. Furthermore, we present an end-to-end compression scheme that employs differentiable quantization and a tiny implicit entropy model to compress the motion grid and compensated Gaussians efficiently. The entire framework is jointly optimized using a rate-distortion trade-off. Extensive experiments demonstrate that 4DGC supports variable bitrates and consistently outperforms existing methods in RD performance across multiple datasets. Qiang Hu 0003, Zihan Zheng, Houqiang Zhong, Sihua Fu, Li Song 0001, Xiaoyun Zhang 0001, Guangtao Zhai, Yanfeng Wang 0001 |
CVPR | 7 |
| 2025 | Image Quality Assessment: From Human to Machine PreferenceabstractImage Quality Assessment (IQA) based on human subjective preferences has undergone extensive research in the past decades. However, with the development of communication protocols, the visual data consumption volume of machines has gradually surpassed that of humans. For machines, the preference depends on downstream tasks such as segmentation and detection, rather than visual appeal. Considering the huge gap between human and machine visual systems, this paper proposes the topic: Image Quality Assessment for Machine Vision for the first time. Specifically, we (1) defined the subjective preferences of machines, including downstream tasks, test models, and evaluation metrics; (2) established the Machine Preference Database (MPD), which contains 2.25M fine-grained annotations and 30k reference/distorted image pair instances; (3) verified the performance of mainstream IQA algorithms on MPD. Experiments show that current IQA metrics are human-centric and cannot accurately characterize machine preferences. We sincerely hope that MPD can promote the evolution of IQA from human to machine preferences. Project page is on: https://github.com/lcysyzxdxc/MPD. Chunyi Li 0001, Yuan Tian 0017, Xiaoyue Ling, Haodong Duan, Haoning Wu 0001, Ziheng Jia, Xiaohong Liu 0001, Xiongkuo Min, Guo Lu, Weisi Lin, Guangtao Zhai |
CVPR | 12 |
| 2025 | Towards All-in-One Medical Image Re-IdentificationabstractMedical image re-identification (MedReID) is underexplored so far, despite its critical applications in personalized healthcare and privacy protection. In this paper, we introduce a thorough benchmark and a unified model for this problem. First, to handle various medical modalities, we propose a novel Continuous Modality-based Parameter Adapter (ComPA). ComPA condenses medical content into a continuous modality representation and dynamically adjusts the modality-agnostic model with modalityspecific parameters at runtime. This allows a single model to adaptively learn and process diverse modality data. Furthermore, we integrate medical priors into our model by aligning it with a bag of pre-trained medical foundation models, in terms of the differential features. Compared to single-image feature, modeling the inter-image difference better fits the re-identification problem, which involves discriminating multiple images. We evaluate the proposed model against 25 foundation models and 8 large multimodal language models across 11 image datasets, demonstrating consistently superior performance. Additionally, we deploy the proposed MedReID technique to two realworld applications, i.e., history-augmented personalized diagnosis and medical privacy protection. Codes and model is available at https://github.com/tianyuan168326/All-inOne-MedReID-Pytorch. Yuan Tian 0017, Kaiyuan Ji, Rongzhao Zhang, Yankai Jiang 0003, Chunyi Li 0001, Xiaosong Wang 0001, Guangtao Zhai |
CVPR | 7 |
| 2025 | AIGV-Assessor: Benchmarking and Evaluating the Perceptual Quality of Text-to-Video Generation with LMM
Huiyu Duan, Guangtao Zhai, Juntong Wang, Xiongkuo Min |
CVPR | 3 |
| 2025 | Learning Hazing to Dehazing: Towards Realistic Haze Generation for Real-World Image DehazingabstractExisting real-world image dehazing methods primarily attempt to fine-tune pre-trained models or adapt their inference procedures, thus heavily relying on the pre-trained models and associated training data. Moreover, restoring heavily distorted information under dense haze requires generative diffusion models, whose potential in de-hazing remains underutilized partly due to their lengthy sampling processes. To address these limitations, we introduce a novel hazing-dehazing pipeline consisting of a Realistic Hazy Image Generation framework (HazeGen) and a Diffusion-based Dehazing framework (DiffDehaze). Specifically, HazeGen harnesses robust generative diffusion priors of real-world hazy images embedded in a pre-trained text-to-image diffusion model. By employing specialized hybrid training and blended sampling strategies, HazeGen produces realistic and diverse hazy images as high-quality training data for DiffDehaze. To alleviate the inefficiency and fidelity concerns associated with diffusion-based methods, DiffDehaze adopts an Accelerated Fidelity-Preserving Sampling process (AccSamp). The core of AccSamp is the Tiled Statistical Alignment Operation (AlignOp), which can provide a clean and faithful dehazing estimate within a small fraction of sampling steps to reduce complexity and enable effective fidelity guidance. Extensive experiments demonstrate the superior dehazing performance and visual quality of our approach over existing methods. The code is available at https://github.com/ruiyi-w/Learning-Hazing-to-Dehazing. Ruiyi Wang, Yushuo Zheng, Chunyi Li 0001, Shuaicheng Liu, Guangtao Zhai, Xiaohong Liu 0001 |
CVPR | 6 |
| 2025 | Mesh Mamba: A Unified State Space Model for Saliency Prediction in Non-Textured and Textured MeshesabstractMesh saliency enhances the adaptability of 3D vision by identifying and emphasizing regions that naturally attract visual attention. To investigate the interaction between geometric structure and texture in shaping visual attention, we establish a comprehensive mesh saliency dataset, which is the first to systematically capture the differences in saliency distribution under both textured and non-textured visual conditions. Furthermore, we introduce mesh Mamba, a unified saliency prediction model based on a state space model (SSM), designed to adapt across various mesh types. Mesh Mamba effectively analyzes the geometric structure of the mesh while seamlessly incorporating texture features into the topological framework, ensuring coherence throughout appearance-enhanced modeling. More importantly, by sub-graph embedding and a bidirectional SSM, the model enables global context modeling for both local geometry and texture, preserving the topological structure and improving the understanding of visual details and structural complexity. Through extensive theoretical and empirical validation, our model not only improves performance across various mesh types but also demonstrates high scalability and versatility, particularly through cross validations of various visual features. Kaiwei Zhang, Dandan Zhu 0001, Xiongkuo Min, Guangtao Zhai |
CVPR | 4 |
| 2025 | Q-Bench-Video: Benchmark the Video Quality Understanding of LMMsabstractWith the rising interest in research on Large Multi-modal Models (LMMs) for video understanding, many studies have emphasized general video comprehension capabilities, neglecting the systematic exploration into video quality understanding. To address this oversight, we introduce Q-Bench-Video in this paper, a new benchmark specifically designed to evaluate LMMs' proficiency in discerning video quality. a) To ensure video source diversity, Q-Bench-Video encompasses videos from natural scenes, AI-generated content (AIGC), and computer graphics (CG). b) Building on the traditional multiple-choice questions format with the Yes-or-No and What-How categories, we include Open-ended questions to better evaluate complex scenarios. Additionally, we incorporate the video pair quality comparison question to enhance comprehensiveness. c) Beyond the traditional Technical, Aesthetic, and Temporal distortions, we have expanded our evaluation aspects to include the dimension of AIGC distortions, which addresses the increasing demand for video generation. Finally, we collect a total of 2,378 question-answer pairs and test them on 12 open-source & 5 proprietary LMMs. Our findings indicate that while LMMs have a foundational understanding of perceptual video quality, their performance remains incomplete and imprecise, with a notable discrepancy compared to the performance of human beings. Through Q-Bench-Video, we seek to catalyze community interest, stimulate further research, and unlock the untapped potential of LMMs to close the gap in video quality understanding. Ziheng Jia, Haoning Wu 0001, Chunyi Li 0001, Zijian Chen 0001, Yingjie Zhou 0003, Wei Sun 0029, Xiaohong Liu 0001, Xiongkuo Min, Weisi Lin, Guangtao Zhai |
CVPR | 11 |
| 2025 | Q-Eval-100K: Evaluating Visual Quality and Alignment Level for Text-to-Vision ContentabstractEvaluating text-to-vision content hinges on two crucial aspects: visual quality and alignment. While significant progress has been made in developing objective models to assess these dimensions, the performance of such models heavily relies on the scale and quality of human annotations. According to Scaling Law, increasing the number of human-labeled instances follows a predictable pattern that enhances the performance of evaluation models. Therefore, we introduce a comprehensive dataset designed to Evaluate Visual quality and Alignment Level for text-to-vision content (Q-EVAL-100K), featuring the largest collection of human-labeled Mean Opinion Scores (MOS) for the mentioned two aspects. The Q-EVAL-100K dataset encompasses both text-to-image and text-to-video models, with 960K human annotations specifically focused on visual quality and alignment for 100K instances (60K images and 40K videos). Leveraging this dataset with context prompt, we propose Q-Eval-Score, a unified model capable of evaluating both visual quality and alignment with special improvements for handling long-text prompt alignment. Experimental results indicate that the proposed Q-Eval-Score achieves superior performance on both visual quality and alignment, with strong generalization capabilities across other benchmarks. These findings highlight the significant value of the Q-EVAL-100K dataset. Data and codes will be available at https://github.com/zzc-1998/Q-Eval. Tengchuan Kou, Shushi Wang, Chunyi Li 0001, Wei Sun 0029, Wei Wang 0213, Zongyu Wang, Xuezhi Cao, Xiongkuo Min, Xiaohong Liu 0001, Guangtao Zhai |
CVPR | 12 |
| 2025 | Shadow Generation Using Diffusion Model with Geometry PriorabstractImage composition involves integrating foreground object into background image to obtain a composite image. One of the key challenges is to produce realistic shadow for the inserted foreground object. Recently, diffusion-based methods have shown superior performance compared to GAN-based methods in shadow generation. However, they are still struggling to generate shadows with plausible geometry in complex cases. In this paper, we focus on promoting diffusion-based methods by leveraging geometry priors. Specifically, we first predict the rotated bounding box and matched shadow shapes for the foreground shadow. Then, the geometry information of rotated bounding box and matched shadow shapes is injected into ControlNet to facilitate shadow generation. Extensive experiments on both DESOBAv2 dataset and real composite images validate the effectiveness of our proposed method. The code and model are released at https://github.com/bcmi/GPSDiffusion-Object-Shadow-Generation. Qingyang Liu 0002, Xinhao Tao, Li Niu 0002, Guangtao Zhai |
CVPR | 5 |
| 2025 | Explore the Hallucination on Low-level Perception for MLLMsabstractThe rapid development of Multi-modality Large Language Models (MLLMs) has significantly influenced various aspects of industry and daily life, showcasing impressive capabilities in visual perception and understanding. However, these models also exhibit hallucinations, which limit their reliability as AI systems, especially in tasks involving low-level visual perception and understanding. We believe that hallucinations stem from a lack of explicit self-awareness in these models, which directly impacts their overall performance. In this paper, we aim to define and evaluate the self-awareness of MLLMs in low-level visual perception and understanding tasks. To this end, we present QL-Bench, a benchmark settings to simulate human responses to low-level vision, investigating self-awareness in low-level visual perception through visual question answering related to low-level attributes such as clarity and lighting. Specifically, we construct the LLSAVisionQA dataset, comprising 2,990 single images and 1,999 image pairs, each accompanied by an open-ended question about its low-level features. Through the evaluation of 15 MLLMs, we demonstrate that while some models exhibit robust low-level visual capabilities, their self-awareness remains relatively underdeveloped. Notably, for the same model, simpler questions are often answered more accurately than complex ones. However, self-awareness appears to improve when addressing more challenging questions. We hope that our benchmark will motivate further research, particularly focused on enhancing the self-awareness of MLLMs in tasks involving low-level visual perception and understanding. Haoning Wu 0001, Xiaohong Liu 0001, Weisi Lin, Guangtao Zhai, Xiongkuo Min |
ICASSP | 6 |
| 2025 | HazeCLIP: Towards Language Guided Real-World Image DehazingabstractExisting methods have achieved remarkable performance in image dehazing, particularly on synthetic datasets. However, they often struggle with real-world hazy images due to domain shift, limiting their practical applicability. This paper introduces HazeCLIP, a language-guided adaptation framework designed to enhance the real-world performance of pre-trained dehazing networks. Inspired by the Contrastive Language-Image Pre-training (CLIP) model’s ability to distinguish between hazy and clean images, we leverage it to evaluate dehazing results. Combined with a region-specific dehazing technique and tailored prompt sets, the CLIP model accurately identifies hazy areas, providing a high-quality, human-like prior that guides the fine-tuning process of pre-trained networks. Extensive experiments demonstrate that HazeCLIP achieves state-of-the-art performance in real-word image dehazing, evaluated through both visual quality and image quality assessment metrics. Codes are available at https://github.com/Troivyn/HazeCLIP. Ruiyi Wang, Wenhao Li 0018, Xiaohong Liu 0001, Chunyi Li 0001, Xiongkuo Min, Guangtao Zhai |
ICASSP | 7 |
| 2025 | Contrastive Learning via Randomly Generated Deep SupervisionabstractUnsupervised visual representation learning has gained significant attention in the computer vision community, driven by recent advancements in contrastive learning. Most existing contrastive learning frameworks rely on instance discrimination as a pretext task, treating each instance as a distinct category. However, this often leads to intra-class collision in a large latent space, compromising the quality of learned representations. To address this issue, we propose a novel contrastive learning method that utilizes randomly generated supervision signals. Our framework incorporates two projection heads: one handles conventional classification tasks, while the other employs a random algorithm to generate fixed-length vectors representing different classes. The second head executes a supervised contrastive learning task based on these vectors, effectively clustering instances of the same class and increasing the separation between different classes. Our method, Contrastive Learning via Randomly Generated Supervision(CLRGS), significantly improves the quality of feature representations across various datasets and achieves state-of-the-art performance in contrastive learning tasks. Zili Ma, Ka-Hou Chan, Yue Liu 0001, Tong Tong 0001, Qinquan Gao, Guangtao Zhai, Xiaohong Liu 0001, Tao Tan 0002 |
ICASSP | 7 |
| 2025 | 3DGCQA: A Quality Assessment Database for 3D AI-Generated ContentsabstractAlthough 3D generated content (3DGC) offers advantages in reducing production costs and accelerating design timelines, its quality often falls short when compared to 3D professionally generated content. Common quality issues frequently affect 3DGC, highlighting the importance of timely and effective quality assessment. Such evaluations not only ensure a higher standard of 3DGCs for end-users but also provide critical insights for advancing generative technologies. To address existing gaps in this domain, this paper introduces a novel 3DGC quality assessment dataset, 3DGCQA, built using 7 representative Text-to-3D generation methods. During the dataset’s construction, 50 fixed prompts are utilized to generate contents across all methods, resulting in the creation of 313 textured meshes that constitute the 3DGCQA dataset. The visualization intuitively reveals the presence of 6 common distortion categories in the generated 3DGCs. To further explore the quality of the 3DGCs, subjective quality assessment is conducted by evaluators, whose ratings reveal significant variation in quality across different generation methods. Additionally, several objective quality assessment algorithms are tested on the 3DGCQA dataset. The results expose limitations in the performance of existing algorithms and underscore the need for developing more specialized quality assessment methods. To provide a valuable resource for future research and development in 3D content generation and quality assessment, the dataset has been open-sourced in https://github.com/zyj-2000/3DGCQA. Yingjie Zhou 0003, Farong Wen, Jun Jia, Yanwei Jiang, Xiaohong Liu 0001, Xiongkuo Min, Guangtao Zhai |
ICASSP | 8 |
| 2025 | Information Density Principle for MLLM BenchmarksabstractWith the emergence of Multimodal Large Language Models (MLLMs), hundreds of benchmarks have been developed to ensure the reliability of MLLMs in downstream tasks. However, the evaluation mechanism itself may not be reliable. For developers of MLLMs, questions remain about which benchmark to use and whether the test results meet their requirements. Therefore, we propose a critical principle of Information Density, which examines how much insight a benchmark can provide for the development of MLLMs. We characterize it from four key dimensions: (1) Fallacy, (2) Difficulty, (3) Redundancy, (4) Diversity. Through a comprehensive analysis of more than 10,000 samples, we measured the information density of 19 MLLM benchmarks. Experiments show that using the latest benchmarks in testing can provide more insight compared to previous ones, but there is still room for improvement in their information density. We hope this principle can promote the development and application of future MLLM benchmarks. Project page: https://github.com/lcysyzxdxc/bench4bench Chunyi Li 0001, Xiaozhe Li, Yuan Tian 0017, Ziheng Jia, Xiaohong Liu 0001, Xiongkuo Min, Jia Wang 0004, Haodong Duan, Kai Chen 0026, Guangtao Zhai |
ICCV | 11 |
| 2025 | FPEM: Face Prior Enhanced Facial Attractiveness Prediction for Live Videos with Face Retouching
Hongjiu Yu, Ying Chen 0011, Kai Li 0012, Xiongkuo Min, Huiyu Duan, Guangtao Zhai, Xu Liu 0006 |
ICCV | 9 |
| 2025 | F-Bench: Rethinking Human Preference Evaluation Metrics for Benchmarking Face Generation, Customization, and RestorationabstractArtificial intelligence generative models exhibit remarkable capabilities in content creation, particularly in face image generation, customization, and restoration. However, current AI-generated faces (AIGFs) often fall short of human preferences due to unique distortions, unrealistic details, and unexpected identity shifts, underscoring the need for a comprehensive quality evaluation framework for AIGFs. To address this need, we introduce FaceQ, a large-scale, comprehensive database of AI-generated Face images with fine-grained Quality annotations reflecting human preferences. The FaceQ database comprises 12,255 images generated by 29 models across three tasks: (1) face generation, (2) face customization, and (3) face restoration. It includes 32,742 mean opinion scores (MOSs) from 180 annotators, assessed across multiple dimensions: quality, authenticity, identity (ID) fidelity, and text-image correspondence. Using the FaceQ database, we establish F-Bench, a benchmark for comparing and evaluating face generation, customization, and restoration models, highlighting strengths and weaknesses across various prompts and evaluation dimensions. Additionally, we assess the performance of existing image quality assessment (IQA), face quality assessment (FQA), AI-generated content image quality assessment (AIGCIQA), and preference evaluation metrics, manifesting that these standard metrics are relatively ineffective in evaluating authenticity, ID fidelity, and text-image correspondence. The FaceQ database will be publicly available upon publication. Lu Liu 0005, Huiyu Duan, Qiang Hu 0003, Chunlei Cai, Tianxiao Ye, Huayu Liu, Xiaoyun Zhang 0001, Guangtao Zhai |
ICCV | 9 |
| 2025 | TR-PTS: Task-Relevant Parameter and Token Selection for Efficient Tuning
Yi Xin 0003, Mingyang Yi, Guangyang Wu, Guangtao Zhai, Xiaohong Liu 0001 |
ICCV | 6 |
| 2025 | Semantics Versus Identity: A Divide-and-Conquer Approach Towards Adjustable Medical Image De-Identification
Yuan Tian 0017, Rongzhao Zhang, Zijian Chen 0001, Yankai Jiang 0003, Chunyi Li 0001, Fang Yan 0002, Qiang Hu 0003, Xiaosong Wang 0001, Guangtao Zhai |
ICCV | 11 |
| 2025 | LMM4LMM: Benchmarking and Evaluating Large-Multimodal Image Generation With LMMsabstractRecent breakthroughs in large multimodal models (LMMs) have significantly advanced both text-to-image (T2I) generation and image-to-text (I2T) interpretation. However, many generated images still suffer from issues related to perceptual quality and text-image alignment. Given the high cost and inefficiency of manual evaluation, an automatic metric that aligns with human preferences is desirable. To this end, we present EvalMi-50K, a comprehensive dataset and benchmark for evaluating large-multimodal image generation, which features (i) comprehensive tasks, encompassing 2,100 extensive prompts across 20 fine-grained task dimensions, and (ii) large-scale human-preference annotations, including 100K mean-opinion scores (MOSs) and 50K question-answering (QA) pairs annotated on 50,400 images generated from 24 T2I models. Based on EvalMi-50K, we propose LMM4LMM, an LMM-based metric for evaluating large multimodal T2I generation from multiple dimensions including perception, text-image correspondence, and task-specific accuracy. Extensive experimental results show that LMM4LMM achieves state-of-the-art performance on EvalMi-50K, and exhibits strong generalization ability on other AI-generated image evaluation benchmark datasets, manifesting the generality of both the EvalMi-50K dataset and LMM4LMM metric. Both EvalMi-50K and LMM4LMM will be released at https://github.com/IntMeGroup/LMM4LMM. Huiyu Duan, Juntong Wang, Guangtao Zhai, Xiongkuo Min |
ICCV | 5 |
| 2025 | Who is a Better Talker: Subjective and Objective Quality Assessment for AI-Generated Talking HeadsabstractSpeech-driven methods for portraits are figuratively known as "Talkers" because of their capability to synthesize speaking mouth shapes and facial movements. Especially with the rapid development of the Text-to-Image (T2I) models, AI-Generated Talking Heads (AGTHs) have gradually become an emerging digital human media. However, challenges persist regarding the quality of these talkers and AGTHs they generate, and comprehensive studies addressing these issues remain limited. To address this gap, this paper presents the largest AGTH quality assessment dataset THQA-10K to date, which selects 12 prominent T2I models and 14 advanced talkers to generate AGTHs for 14 prompts. After excluding instances where AGTH generation is unsuccessful, the THQA-10K dataset contains 10,457 AGTHs. Then, volunteers are recruited to subjectively rate the AGTHs and give the corresponding distortion categories. In our analysis for subjective experimental results, we evaluate the performance of talkers in terms of generalizability and quality, and also expose the distortions of existing AGTHs. Finally, an objective quality assessment method based on the first frame, Y-T slice and tone-lip consistency is proposed. Experimental results show that this method can achieve state-of-the-art (SOTA) performance in AGTH quality assessment. The work is released at https://github.com/zyj-2000/Talker. Yingjie Zhou 0003, Jiezhang Cao, Farong Wen, Yanwei Jiang, Jun Jia, Xiaohong Liu 0001, Xiongkuo Min, Guangtao Zhai |
ICCV | 9 |
| 2025 | CDHQA: A Quality Assessment Database for Conversational Digital Human
Yingjie Zhou 0003, Yinghan Xia, Zhixiang Lu, Farong Wen, Yu Wang 0002, Yu Zhou 0016, Xiaohong Liu 0001, Xiongkuo Min, Jiezhang Cao, Guangtao Zhai |
ICIG (3) | 13 |
| 2025 | Exploring The Potential of Vision-Language Models for Pure-Image and Text-Guided-Image Saliency PredictionabstractWe introduce VLSal, a saliency prediction framework that leverages Vision-Language Models (VLMs) to unify pure-image and text-guided-image saliency prediction tasks and achieve high performance in both. We extract visual features from the visual encoder and retrieve the corresponding visual token features from the language decoder, which serves as a natural feature fusion mechanism. These features are then processed through a U-Net-based saliency decoder to generate accurate saliency maps. To efficiently adapt the large-scale pretrained model, we apply Low-Rank Adaptation (LoRA) finetuning, reducing computational costs while preserving performance. Extensive experiments on benchmark datasets, including SALICON, MIT1003, and TIS, demonstrate that VLSal outperforms existing methods in both pure-image and text-guided-image saliency prediction. Huiyu Duan, Xiongkuo Min, Guangtao Zhai, Patrick Le Callet |
ICIP | 4 |
| 2025 | OBI-Bench: Can LMMs Aid in Study of Ancient Script on Oracle Bones?abstractWe introduce OBI-Bench, a holistic benchmark crafted to systematically evaluate large multi-modal models (LMMs) on whole-process oracle bone inscriptions (OBI) processing tasks demanding expert-level domain knowledge and deliberate cognition. OBI-Bench includes 5,523 meticulously collected diverse-sourced images, covering five key domain problems: recognition, rejoining, classification, retrieval, and deciphering. These images span centuries of archaeological findings and years of research by front-line scholars, comprising multi-stage font appearances from excavation to synthesis, such as original oracle bone, inked rubbings, oracle bone fragments, cropped single characters, and handprinted characters. Unlike existing benchmarks, OBI-Bench focuses on advanced visual perception and reasoning with OBI-specific knowledge, challenging LMMs to perform tasks akin to those faced by experts. The evaluation of 6 proprietary LMMs as well as 17 open-source LMMs highlights the substantial challenges and demands posed by OBI-Bench. Even the latest versions of GPT-4o, Gemini 1.5 Pro, and Qwen-VL-Max are still far from public-level humans in some fine-grained perception tasks. However, they perform at a level comparable to untrained humans in deciphering tasks, indicating remarkable capabilities in offering new interpretative perspectives and generating creative guesses. We hope OBI-Bench can facilitate the community to develop domain-specific multi-modal foundation models towards ancient language research and delve deeper to discover and enhance these untapped potentials of LMMs. Zijian Chen 0001, Tingzhu Chen, Wenjun Zhang 0001, Guangtao Zhai |
ICLR | 4 |
| 2025 | A-Bench: Are LMMs Masters at Evaluating AI-generated Images?abstractHow to accurately and efficiently assess AI-generated images (AIGIs) remains a critical challenge for generative models. Given the high costs and extensive time commitments required for user studies, many researchers have turned towards employing large multi-modal models (LMMs) as AIGI evaluators, the precision and validity of which are still questionable. Furthermore, traditional benchmarks often utilize mostly natural-captured content rather than AIGIs to test the abilities of LMMs, leading to a noticeable gap for AIGIs. Therefore, we introduce **A-Bench** in this paper, a benchmark designed to diagnose *whether LMMs are masters at evaluating AIGIs*. Specifically, **A-Bench** is organized under two key principles: 1) Emphasizing both high-level semantic understanding and low-level visual quality perception to address the intricate demands of AIGIs. 2) Various generative models are utilized for AIGI creation, and various LMMs are employed for evaluation, which ensures a comprehensive validation scope. Ultimately, 2,864 AIGIs from 16 text-to-image models are sampled, each paired with question-answers annotated by human experts. We hope that **A-Bench** will significantly enhance the evaluation process and promote the generation quality for AIGIs. Haoning Wu 0001, Chunyi Li 0001, Yingjie Zhou 0003, Wei Sun 0029, Xiongkuo Min, Zijian Chen 0001, Xiaohong Liu 0001, Weisi Lin, Guangtao Zhai |
ICLR | 10 |
| 2025 | SI23DCQA: Perceptual Quality Assessment of Single Image-to-3D ContentabstractIn recent years, significant efforts have been dedicated to advancing 3D content generation. However, existing quality assessment research predominantly focuses on evaluating Text-to-3D Content (T23DC) while ignoring Single Image-to-3D Content (SI23DC). In this paper, we establish the first Single Image-to-3D Content Quality Assessment (SI23DCQA) database to comprehensively study the perceptual quality of SI23DCs. The database contains 1500 SI23DCs, which are generated by 5 common SI23DC algorithms from 300 images including realistic images, AI generated images, and model rendered images. Afterward, we carry out a well-designed subjective experiment to collect subjective quality ratings for SI23DCs from three perspectives including overall, color, and shape. Additionally, a benchmark experiment is conducted with the state-of-the-art no reference image quality assessment (NR-IQA), no reference video quality assessment (NR-VQA), and no reference 3D quality assessment (NR-3DQA) and the experimental results show that current quality assessment methods are limited in evaluating the perceptual loss of SI23DCs. The database is released on https://github.com/ZedFu/SI23DCQA. Huiyu Duan, Xiaohong Liu 0001, Xiongkuo Min, Jia Wang 0004, Guangtao Zhai |
ICME | 7 |
| 2025 | VIP-PCQA: A Multi-Modal Framework for No-reference Point Cloud Quality AssessmentabstractPoint clouds often suffer from geometric and color noise, as well as compression artifacts, during their production, storage, and transmission. Therefore, accurately and automatically evaluating the quality of point clouds is crucial for optimizing storage and compression strategies. This paper introduces the VIP-PCQA, a novel framework that combines Video, Image, and Point cloud modalities for no-reference Point Cloud Quality Assessment. The framework begins by rendering projection videos and normal images from point clouds, followed by sampling patches and computing statistical features related to color and geometry. Subsequently, a video encoder, two image encoders, and a point cloud encoder are employed to extract modality-specific features. Finally, these features are fused to regress the quality score. Experimental results on three publicly available benchmark databases demonstrate that VIP-PCQA achieves outstanding performance with excellent generalization capabilities. An ablation study further highlights the indispensable contribution of each modality to the framework’s success. The code is released on https://github.com/ZedFu/VIP-PCQA. Huiyu Duan, Xiaohong Liu 0001, Xiongkuo Min, Guangtao Zhai |
ICME | 7 |
| 2025 | Perceiving Smoothness: Temporal Consistency Learning for Multi-Frame-Rate Video Quality AssessmentabstractThe video industry is continuously evolving, with videos featuring a wide range of frame rates. Video quality assessment (VQA) aims to automatically monitor and select the optimal frame rate for video communication systems. Existing VQA research has achieved impressive results in high frame rate (HFR) VQA tasks, but often lacks specific designs to address various frame rate distortions, such as smoothness distortions caused by frame rate variations, artifacts from the coupling of frame rate and compression, and confusion between low frame rate and slow motion. To address these challenges, we propose a VQA framework to perceive smoothness (PSVQA), which includes a novel frame-rate-driven feature processing module and a new feature fusion strategy. The module aggregates smoothness features from multi-scale temporal embeddings and incorporates frame rate guidance to resolve the discrepancies between temporal features and real perceptual experience. Furthermore, we combine spatial video features with temporal consistency features for quality modeling, optimizing the feature fusion module to enhance multi-frame-rate perception. Through extensive experiments on HFR and variable-frame-rate datasets, we validate the effectiveness of PSVQA. Jinliang Han, Xiongkuo Min, Wei Sun 0029, Guangtao Zhai |
ICME | 4 |
| 2025 | A Novel Framework for Realistic 3D Scene Regeneration with Graph of ThoughtsabstractIn embodied intelligence applications, highly realistic 3D scenes lay the foundation for perception and decision-making, while 3D scene regeneration creates more coherent and personalized virtual spaces, facilitating more efficient task adaptation and agent training. To address this, we propose a reasoning framework based on the Graph of Thoughts (GoT), which enhances the prompting capabilities of large language models (LLM) and integrates a synergistic mechanism of retrospective memory and feedback loops into the regeneration process. During the initial generation phase, we retain the Holodeck paradigm, combining LLM-driven scene design inferences with the spatial layout of 3D assets from Objaverse. In the regeneration phase, dynamic feedback loops trigger backtracking of reasoning memory to adjust relevant elements according to evolving requirements, while maintaining stability and consistency in unrelated elements, ensuring the scene’s overall coherence. We conduct both subjective and objective experiments to validate the effectiveness of this framework, demonstrating significant improvements in 3D scene generation. Yitian Kou, Kaiwei Zhang, Dandan Zhu 0001, Xiongkuo Min, Guangtao Zhai |
ICME | 5 |
| 2025 | DiffDeid: High-Quality Face De-identification and Recovery via Diffusion InversionabstractNowadays, personal privacy protection is extremely emphasised. Face de-identification is considered as an effective way to protect the visual privacy through disguising or replacing identity attributes. Existing methods compromise either high fidelity or reversibility. To address these issues, this paper proposes DiffDeid, the first diffusion-based face de-identification and recovery method. Leveraging recent diffusion inversion and control technicques, DiffDeid achieves both high quality imperceptible de-identification and exact recovery with passwords. DiffDeid has three attractions: (1) It can generate de-identified faces with the superior fidelity while maintaining other non-identity attributes for visual tasks. (2) The correct password is powerful enough to restore facial images with extreme details. Meanwhile, incorrect passwords can lead to vastly different decryption results. (3) DiffDeid demands minimal computing resources and instant training time compared to others. We conducted experiments on various face datasets to showcase the superiority of our proposed method. Additional experiments show that DiffDeid is powerful with diverse text prompts and control instructions even beyond human faces. Codes are available at project page. Zheyuan Liu 0011, Jun Jia, Hongyi Miao, Yiwei Yang 0007, Yanwei Jiang, Yingjie Zhou 0003, Zhi Liu 0004, Guangtao Zhai |
ICME | 8 |
| 2025 | HarmonyIQA: Pioneering Benchmark and Model for Image Harmonization Quality AssessmentabstractImage composition involves extracting a foreground object from one image and pasting it into another image through Image harmonization algorithms (IHAs), which aim to adjust the appearance of the foreground object to better match the background. Existing image quality assessment (IQA) methods may fail to align with human visual preference on image harmonization due to the insensitivity to minor color or light inconsistency. To address the issue and facilitate the advancement of IHAs, we introduce the first Image Quality Assessment Database for image Harmony evaluation (HarmonyIQAD), which consists of 1,350 harmonized images generated by 9 different IHAs, and the corresponding human visual preference scores. Based on this database, we propose a Harmony Image Quality Assessment (HarmonyIQA), to predict human visual preference for harmonized images. Extensive experiments show that HarmonyIQA achieves state-of-the-art performance on human visual preference evaluation for harmonized images, and also achieves competing results on traditional IQA tasks. Furthermore, cross-dataset evaluation also shows that HarmonyIQA exhibits better generalization ability than self-supervised learning-based IQA methods. The dataset and code are available at https://github.com/IntMeGroup/HarmonyIQA. Zitong Xu, Huiyu Duan, Guangji Ma, Qingbo Wu 0001, Xiongkuo Min, Guangtao Zhai, Patrick Le Callet |
ICME | 8 |
| 2025 | LD2Scan: A Lightweight Dual-Temporal Constrained Scanpath Prediction Model for Omnidirectional ImagesabstractPredicting scanpaths in omnidirectional images (ODIs) is essential for simulating human gaze behaviors. However, current methods often struggle with long-term dependencies and exhibit high complexity, which limits their efficiency and scalability. To tackle these challenges, we propose LD2Scan, a lightweight diffusion-based model specifically designed for scanpath prediction in ODIs. It employs Efficient Equivariant (E4) convolution to enhance feature extraction from distorted ODIs while improving computational performance, thereby reducing resource demands. LD2Scan utilizes a dual-graph convolutional network (GCN) to enforce internal time constraints between fixations, integrating semantic-level GCN for sequential fixation modeling and image-level GCN to capture relationships across different images, enriching contextual information. We formulate the scanpath prediction issue as a conditional generation task, refining noisy scanpaths using features encoded by the dual-GCN and robust E4-processed features. Experimental results on several benchmark datasets demonstrate that LD2Scan outperforms existing methods in terms of both accuracy and efficiency. Dandan Zhu 0001, Kun Zhu 0024, Xiongkuo Min, Guangtao Zhai |
ICME | 6 |
| 2025 | IllusionBench: A Large-scale and Comprehensive Benchmark for Visual Illusion Understanding in Vision-Language ModelsabstractCurrent Visual Language Models (VLMs) show impressive image understanding but struggle with visual illusions, especially in real-world scenarios. Existing benchmarks focus on classical cognitive illusions, which have been learned by state-of-the-art (SOTA) VLMs, revealing issues such as hallucinations and limited perceptual abilities. To address this gap, we introduce IllusionBench, a comprehensive visual illusion dataset that encompasses not only classic cognitive illusions but also real-world scene illusions. This dataset features 1,051 images, 5,548 question-answer pairs, and 1,051 golden text descriptions that address the presence, causes, and content of the illusions. We evaluate ten SOTA VLMs on this dataset using true-or-false, multiple-choice, and open-ended tasks. In addition to real-world illusions, we design trap illusions that resemble classical patterns but differ in reality, highlighting hallucination issues in SOTA models. The top-performing model, GPT-4o, achieves 80.59% accuracy on true-or-false tasks and 76.75% on multiple-choice questions, but still lags behind human performance. In the semantic description task, GPT-4o’s hallucinations on classical illusions result in low scores for trap illusions, even falling behind some open-source models. IllusionBench is, to the best of our knowledge, the largest and most comprehensive benchmark for visual illusions in VLMs to date. Xinyi Wei, Xiaohong Liu 0001, Guangtao Zhai, Xiongkuo Min |
ICME | 5 |
| 2025 | MEScan360: A Memory-Enhanced Scanpath Prediction Model for Omnidirectional ImagesabstractScanpath prediction for omnidirectional images (ODIs) aims to capture the dynamic human visual attention. However, the complicated gaze behavior and inevitable projection distortion make scanpath prediction in ODIs extremely challenging. Most existing models neither capture the long-term dependencies across visual states nor fully incorporate historical memory information, leading to limited performance. To this end, we propose MEScan360, a memory-enhanced scanpath prediction model for ODIs. We introduce two key innovations: long-term memory storage unit and memory interaction module. These two components establish a more explicit link between past visual information and current visual inputs, thereby significantly enhancing the performance of scanpath prediction. Furthermore, a robust feature extraction module is designed to extract semantic feature precisely from distorted ODIs with a more lightweight structure. Extensive experiments on several benchmark datasets demonstrate that our proposed model achieves competitive performance in both accuracy and efficiency. Dandan Zhu 0001, Kaiwei Zhang, Fei Jiang 0006, Guangtao Zhai |
ICME | 5 |
| 2025 | CAP: An Advanced No-Reference Quality Assessment Method for AI-Generated 3D MeshesabstractThe advent of generative AI has revolutionized 3D content design, significantly enhancing modelers’ efficiency. However, the quality of generated 3D content, particularly Generated Meshes (GMs), remains a critical concern. GMs pose unique challenges for quality assessment due to their complex geometry, detailed texture mapping, and distortions that differ from traditional meshes. Existing methods fail to address these GM-specific issues. To tackle this gap, we introduce a novel no-reference quality assessment method, CAP, which integrates CT-Slice, prompt Alignment, and Projections. CAP employs a six-face projection to capture external features and a CT-like slicing approach to extract internal quality features. Additionally, it leverages Contrastive Language-Image Pre-Training (CLIP) to measure the alignment between projection embeddings and prompts as a key quality indicator. Experimental results demonstrate that CAP effectively evaluates GM quality by combining internal, external, and alignment features. The code for this work has been open-sourced in https://github.com/zyj-2000/CAP. Yingjie Zhou 0003, Farong Wen, Yanwei Jiang, Jun Jia, Xiaohong Liu 0001, Xiongkuo Min, Guangtao Zhai |
ICME | 8 |
| 2025 | ESVQA: Perceptual Quality Assessment of Egocentric Spatial VideosabstractWith the rapid development of eXtended Reality (XR), egocentric spatial shooting and display technologies have further enhanced immersion and engagement for users, delivering more captivating and interactive experiences. Assessing the quality of experience (QoE) of egocentric spatial videos is crucial to ensure a high-quality viewing experience. However, the corresponding research is still lacking. In this paper, we use the concept of embodied experience to highlight this more immersive experience and study the new problem, i.e., embodied perceptual quality assessment for egocentric spatial videos. Specifically, we introduce the first Egocentric Spatial Video Quality Assessment Database (ESVQAD), which comprises 600 egocentric spatial videos captured using the Apple Vision Pro and their corresponding mean opinion scores (MOSs). Furthermore, we propose a novel multi-dimensional binocular feature fusion model, termed ESVQAnet, which integrates binocular spatial, motion, and semantic features to predict the overall perceptual quality. Experimental results demonstrate the ESVQAnet significantly outperforms 16 state-of-the-art VQA models on the embodied perceptual quality assessment task, and exhibits strong generalization capability on traditional VQA tasks. The database and code are available at https://github.com/IntMeGroup/ESVQA. Xilei Zhu, Huiyu Duan, Yucheng Zhu, Xiongkuo Min, Guangtao Zhai, Patrick Le Callet |
ICME | 6 |
| 2025 | AGAV-Rater: Adapting Large Multimodal Model for AI-Generated Audio-Visual Quality AssessmentabstractMany video-to-audio (VTA) methods have been proposed for dubbing silent AI-generated videos. An efficient quality assessment method for AI-generated audio-visual content (AGAV) is crucial for ensuring audio-visual quality. Existing audio-visual quality assessment methods struggle with unique distortions in AGAVs, such as unrealistic and inconsistent elements. To address this, we introduce AGAVQA-3k, the first large-scale AGAV quality assessment dataset, comprising $3,382$ AGAVs from $16$ VTA methods. AGAVQA-3k includes two subsets: AGAVQA-MOS, which provides multi-dimensional scores for audio quality, content consistency, and overall quality, and AGAVQA-Pair, designed for optimal AGAV pair selection. We further propose AGAV-Rater, a LMM-based model that can score AGAVs, as well as audio and music generated from text, across multiple dimensions, and selects the best AGAV generated by VTA methods to present to the user. AGAV-Rater achieves state-of-the-art performance on AGAVQA-3k, Text-to-Audio, and Text-to-Music datasets. Subjective tests also confirm that AGAV-Rater enhances VTA performance and user experience. The dataset and code is available at https://github.com/charlotte9524/AGAV-Rater. Yuqin Cao, Xiongkuo Min, Wei Sun 0029, Guangtao Zhai |
ICML | 5 |
| 2025 | A Dataset and Method for Assessing the Quality of Display DevicesabstractThis paper proposes a novel objective quality assessment method for display devices, aiming to comprehensively evaluate their overall quality, with a particular focus on the user’s subjective experience. By constructing a dedicated video dataset and integrating a no-reference video quality assessment (NR VQA) model, we directly learn spatiotemporal features related to quality from video frames and introduce a color attention mechanism to enhance the model’s sensitivity to color distortions. The model consists of feature extraction, quality regression, and quality pooling modules, capable of extracting quality-aware features from spatial, spatiotemporal, and color domains, and determining the final quality score using a temporal averaging pooling strategy. Experimental results demonstrate that the proposed model performs excellently in evaluating the overall quality of display devices, showing high practical applicability. The main contributions of this paper include constructing a video dataset, proposing a new model, and introducing a color attention mechanism, which effectively enhances the comprehensive evaluation of the subjective quality of display devices, filling the gap in display device quality assessment. Haoyang Ni, Kaiwei Zhang, Fangfang Lu, Xiongkuo Min, Guangtao Zhai |
ISCAS | 6 |
| 2025 | Machine Vision Quality Assessment for Image RestorationabstractIn recent years, substantial progress has been made in the realm of No-Reference Image Quality Assessment (NR-IQA) for image restoration, where performance has been predominantly evaluated using metrics such as BRISQUE [1] and Hyper-IQA [2]. However, these NR-IQA metrics assess the perceptual quality of images without considering their utility in specific machine vision tasks, such as object detection and semantic segmentation. In this paper, we propose a Machine Vision Quality Assessment (MVQA) framework for image restoration. Specifically, we introduce the weighted Alternative Free-response Operating Characteristic (wAFROC) [3] as a metric to assess the machine vision quality of three image restoration sub-tasks: dehazing, denoising, and superresolution. By accounting for both detection sensitivity and spatial localization, wAFROC provides a more comprehensive evaluation of image quality in machine vision contexts. Its effectiveness is validated through downstream tasks, specifically object detection and semantic segmentation. We construct an IQA dataset for image restoration to explore the impact of various image restoration algorithms on the accuracy of object detection and semantic segmentation algorithms. Extensive experimental results demonstrate that the MVQA framework, leveraging wAFROC, effectively predicts the influence of image quality on machine vision tasks, bridging the gap between perceptual IQA and task-specific quality requirements in machine vision applications. Yiming Shi, Xiongkuo Min, Guangtao Zhai |
ISCAS | 4 |
| 2025 | Visual Saliency Prediction for Augmented Reality VideosabstractAugmented Reality (AR) is an emerging technology that allows users to perceive both virtual-world contents and real-world scenes simultaneously. It has numerous applications in industrial manufacturing, entertainment, gaming, education, etc. In AR environments, the visual confusion phenomenon caused by the overlay of augmented content and real backgrounds is evident, yet the understanding and research of visual saliency under the AR visual confusion condition remains limited. This paper primarily analyzes the interaction between real-world scenes and AR content, explores human visual saliency when using AR devices. First, we conduct a large-scale eye-tracking experiment based on a head-mounted AR device, and construct an AR saliency dataset containing 2160 videos, with corresponding collected eye movement data. Through qualitative analysis of the visual attention heat maps, we conclude that visual confusion significantly influences visual attention in AR video. Additionally, we quantitatively evaluate the performance of a series of classical saliency models and deep neural network saliency models on the dataset constructed in this project. For better predicting saliency in AR, we propose a general saliency prediction model, InternSal, which achieves state-of-the-art performance compared to other methods. The database and codes will be released to facilitate future research. Zongyi Xie, Huiyu Duan, Xiongkuo Min, Guangtao Zhai |
ISCAS | 6 |
| 2025 | LPerceptual Quality Assessment of AI Generated Content Videos: a Dataset and BenchmarkabstractIn recent years, artificial intelligence (AI) driven video generation has garnered significant attention due to advancements in large language model techniques. Thus, there is a great demand to explore the effectiveness of video quality assessment (VQA) models in evaluating the perceptual quality of AI-generated content (AIGC) videos and in optimizing video generation techniques. Therefore, in this paper, we try to systemically investigate the AIGC-VQA problem from both subjective and objective quality assessment perspectives. For the subjective perspective, we construct a Large-scale Generated Video Quality assessment (LGVQ) dataset, consisting of 2,808 AIGC videos generated by 6 video generation models using 468 carefully selected text prompts. We evaluate the perceptual quality of AIGC videos from three dimensions: spatial quality, temporal quality, and text-to-video alignment, which hold the utmost importance for current video generation techniques. For the objective perspective, we establish a benchmark for evaluating existing quality assessment metrics on the LGVQ dataset, which fully demonstrates the performance of current mainstream VQA methods in evaluating AIGV quality. We hope that this work can contribute to the advancement of AIGC video generation technology as well as the evaluation techniques for AIGC videos. The LGVQ dataset will release publicly. Wei Sun 0029, Xinyue Li 0001, Jun Jia, Xiongkuo Min, Chunyi Li 0001, Zhongpeng Ji, Fengyu Sun, Shangling Jui, Guangtao Zhai |
ISCAS | 12 |
| 2025 | Mitigating Long-tail Distribution in Oracle Bone Inscriptions: Dataset, Model, and BenchmarkabstractThe oracle bone inscription (OBI) recognition plays a significant role in understanding the history and culture of ancient China. However, the existing OBI datasets suffer from a long-tail distribution problem, leading to biased performance of OBI recognition models across majority and minority classes. With recent advancements in generative models, OBI synthesis-based data augmentation has become a promising avenue to expand the sample size of minority classes. Unfortunately, current OBI datasets lack large-scale structure-aligned image pairs for generative model training. To address these problems, we first present the Oracle-P15K, a structure-aligned OBI dataset for OBI generation and denoising, consisting of 14,542 images infused with domain knowledge from OBI experts. Second, we propose a diffusion model-based pseudo OBI generator, called OBIDiff, to achieve realistic and controllable OBI generation. Given a clean glyph image and a target rubbing-style image, it can effectively transfer the noise style of the original rubbing to the glyph image. Extensive experiments on OBI downstream tasks and user preference studies show the effectiveness of the proposed Oracle-P15K dataset and demonstrate that OBIDiff can accurately preserve inherent glyph structures while transferring authentic rubbing styles effectively. The dataset, code, and pre-trained models are available at https://github.com/LJHolyGround/Oracle-P15K. Jinhao Li 0001, Zijian Chen 0001, Runze Jiang, Tingzhu Chen, Changbo Wang, Guangtao Zhai |
ACM Multimedia | 6 |
| 2025 | EEmo-Bench: A Benchmark for Multi-modal Large Language Models on Image Evoked Emotion AssessmentabstractThe furnishing of multi-modal large language models (MLLMs) has led to the emergence of numerous benchmark studies, particularly those evaluating their perception and understanding capabilities. Among these, understanding image-evoked emotions aims to enhance MLLMs' empathy, with significant applications such as human-machine interaction and advertising recommendations. However, current evaluations of this MLLM capability remain coarse-grained, and a systematic and comprehensive assessment is still lacking. To this end, we introduce EEmo-Bench, a novel benchmark dedicated to the analysis of the evoked emotions in images across diverse content categories. Our core contributions include: 1) Regarding the diversity of the evoked emotions, we adopt an emotion ranking strategy and employ the Valence-Arousal-Dominance (VAD) as emotional attributes for emotional assessment. In line with this methodology, 1,960 images are collected and manually annotated. 2) We design four tasks to evaluate MLLMs' ability to capture the evoked emotions by single images and their associated attributes: Perception, Ranking, Description, and Assessment. Additionally, image-pairwise analysis is introduced to investigate the model's proficiency in performing joint and comparative analysis. In total, we collect 6,773 question-answer pairs and perform a thorough assessment on 19 commonly-used MLLMs. The results indicate that while some proprietary and large-scale open-source MLLMs achieve promising overall performance, the analytical capabilities in certain evaluation dimensions remain suboptimal. Our EEmo-Bench paves the path for further research aimed at enhancing the comprehensive perceiving and understanding capabilities of MLLMs concerning image-evoked emotions, which is crucial for machine-centric emotion perception and understanding. Our code and benchmark datasets are available at https://github.com/workerred/EEmo-Bench. Lancheng Gao, Ziheng Jia, Yunhao Zeng, Wei Sun 0029, Wei Zhou 0021, Guangtao Zhai, Xiongkuo Min |
ACM Multimedia | 7 |
| 2025 | Multi-Dimensional Text-to-Face Image Quality Assessment Using LLM: Database and MethodabstractWith the rise of Text-to-Image (T2I) models, generating face images from text prompts has emerged as a prominent research area. However, evaluating the quality of these generated face images, particularly with respect to fine-grained facial attributes, remains a significant challenge. To address this, we introduce the Fine -grained Text-to-Face Image Quality Assessment (FineTFIQA) database, which is designed to evaluate the ability of T2I models to generate fine-grained face images. To the best of our knowledge, this database is the largest of its kind, containing 7,218 face images generated from 1,000 text prompts that cover 111 distinct facial attributes. A large group of subjects was invited to assess the quality of text-to-face images on four evaluation dimensions: perceptual quality, human likeness, attractiveness, and consistency. Additionally, we develop the Multi-Dimensional Text-to-Face Image Quality Assessment (MDTFIQA) method based on the Large Language Model (LLM), which combines both face image features and text features to evaluate generated images on all evaluation dimensions. Extensive experimental results demonstrate that traditional face image assessment methods and general image quality assessment methods are inadequate for accurately evaluating generated text-to-face images. Our method significantly outperforms these existing methods on all evaluation dimensions, proving to be an effective method for assessing the quality of generated text-to-face images. Xiongkuo Min, Jinliang Han, Yuqin Cao, Sijing Wu, Yunze Dou, Guangtao Zhai |
ACM Multimedia | 7 |
| 2025 | VQA2: Visual Question Answering for Video Quality AssessmentabstractThe advent and proliferation of large multi-modal models (LMMs) have introduced new paradigms to computer vision, transforming various tasks into a unified visual question answering framework. Video Quality Assessment (VQA), a classic field in low-level visual perception, focused initially on quantitative video quality scoring. However, driven by advances in LMMs, it is now progressing toward more holistic visual quality understanding tasks. Recent studies in the image domain have demonstrated that Visual Question Answering (VQA) can markedly enhance low-level visual quality evaluation. Nevertheless, related work has not been explored in the video domain, leaving substantial room for improvement. To address this gap, we introduce the VQA² Instruction Dataset-the first visual question answering instruction dataset that focuses on video quality assessment. This dataset consists of 3 subsets and covers various video types, containing 157,755 instruction question-answer pairs. Then, leveraging this foundation, we present the VQA² series models. The VQA² series models interleave visual and motion tokens to enhance the perception of spatial-temporal quality details in videos. We conduct extensive experiments on video quality scoring and understanding tasks, and results demonstrate that the VQA² series models achieve excellent performance in both tasks. Notably, our final model, the VQA²-Assistant, exceeds the renowned GPT-4o in visual quality understanding tasks while maintaining strong competitiveness in quality scoring tasks. Our work provides a foundation and feasible approach for integrating low-level video quality assessment and understanding with LMMs. Ziheng Jia, Jiaying Qian, Haoning Wu 0001, Wei Sun 0029, Chunyi Li 0001, Xiaohong Liu 0001, Weisi Lin, Guangtao Zhai, Xiongkuo Min |
ACM Multimedia | 9 |
| 2025 | RGC-VQA: An Exploration Database for Robotic-Generated Video Quality AssessmentabstractAs camera-equipped robotic platforms become increasingly integrated into daily life, robotic-generated videos have begun to appear on streaming media platforms, enabling us to envision a future where humans and robots coexist. We innovatively propose the concept of Robotic-Generated Content (RGC) to term these videos generated from egocentric perspective of robots. The perceptual quality of RGC videos is critical in human-robot interaction scenarios, and RGC videos exhibit unique distortions and visual requirements that differ markedly from those of professionally-generated content (PGC) videos and user-generated content (UGC) videos. However, dedicated research on quality assessment of RGC videos is still lacking. To address this gap and to support broader robotic applications, we establish the first Robotic-Generated Content Database (RGCD), which contains a total of 2,100 videos drawn from three robot categories and sourced from diverse platforms. A subjective VQA experiment is conducted subsequently to assess human visual perception of robotic-generated videos. Finally, we conduct a benchmark experiment to evaluate the performance of 11 state-of-the-art VQA models on our database. Experimental results reveal significant limitations in existing VQA models when applied to complex, robotic-generated content, highlighting a critical need for RGC-specific VQA models. Our RGCD is publicly available at: https://github.com/IntMeGroup/RGC-VQA. Jianing Jin, Jiangyong Ying, Huiyu Duan, Sijing Wu, Yushuo Zheng, Xiongkuo Min, Guangtao Zhai |
ACM Multimedia | 9 |
| 2025 | Towards a New Paradigm of Visual Signal CompressionabstractUltra-low bitrate image compression is a challenging and demand- ing topic. With the development of Large Multimodal Models (LMMs), a Cross Modality Compression (CMC) paradigm of Image-Text- Image has emerged. Compared with traditional codecs, this semantic- level compression can reduce image data size to 0.1% or even lower, which has strong potential applications. However, CMC has cer- tain defects in consistency with the original image and perceptual quality. To inspire insights into such a problem, we introduce CMC- Bench, a benchmark of the cooperative performance of Image-to- Text (I2T) and Text-to-Image (T2I) models for image compression. This benchmark covers 18,000 and 40,000 images respectively to verify 6 mainstream I2T and 12 T2I models, including 160,000 sub- jective preference scores annotated by human experts. At ultra-low bitrates, it proves that the combination of some I2T and T2I models has surpassed the most advanced visual signal codecs; meanwhile, it highlights where LMMs can be further optimized toward the compression task. We encourage LMM developers to participate in this test to promote the evolution of visual signal codec protocols. Chunyi Li 0001, Xiele Wu, Haoning Wu 0001, Donghui Feng 0003, Guo Lu, Xiongkuo Min, Xiaohong Liu 0001, Guangtao Zhai, Weisi Lin |
ACM Multimedia | 9 |
| 2025 | Towards Explainable Partial-AIGC Image Quality Assessment
Jiaying Qian, Ziheng Jia, Guangtao Zhai, Xiongkuo Min |
ACM Multimedia | 5 |
| 2025 | DFBench: Benchmarking Deepfake Image Detection Capability of Large Multimodal ModelsabstractWith the rapid advancement of generative models, the realism of AI-generated images has significantly improved, posing critical challenges for verifying digital content authenticity. Current deepfake detection methods often depend on datasets with limited generation models and content diversity that fail to keep pace with the evolving complexity and increasing realism of the AI-generated content. Large multimodal models (LMMs), widely adopted in various vision tasks, have demonstrated strong zero-shot capabilities, yet their potential in deepfake detection remains largely unexplored. To bridge this gap, we present DFBench, a large-scale DeepFake Benchmark featuring (i) broad diversity, including 540,000 images across real, AI-edited, and AI-generated content, (ii) latest model, the fake images are generated by 12 state-of-the-art generation models, and (iii) bidirectional benchmarking and evaluating for both the detection accuracy of deepfake detectors and the evasion capability of generative models. Based on DFBench, we propose MoA-DF, Mixture of Agents for DeepFake detection, leveraging a combined probability strategy from multiple LMMs. MoA-DF achieves state-of-the-art performance, further proving the effectiveness of leveraging LMMs for deepfake detection. Database and codes are publicly available at https://github.com/IntMeGroup/DFBench. Huiyu Duan, Juntong Wang, Ziheng Jia, Woo Yi Yang, Xiaorong Zhu, Jiaying Qian, Yuke Xing, Guangtao Zhai, Xiongkuo Min |
ACM Multimedia | 10 |
| 2025 | AU-IQA: A Benchmark Dataset for Perceptual Quality Assessment of AI-Enhanced User-Generated ContentabstractAI-based image enhancement techniques have been widely adopted in various visual applications, significantly improving the perceptual quality of user-generated content (UGC). However, the lack of specialized quality assessment models has become a significant limiting factor in this field, limiting user experience and hindering the advancement of enhancement methods. While perceptual quality assessment methods have shown strong performance on UGC and AIGC individually, their effectiveness on AI-enhanced UGC (AI-UGC) which blends features from both-remains largely unexplored. To address this gap, we construct AU-IQA, a benchmark dataset comprising 4,800 AI-UGC images produced by three representative enhancement types which include super-resolution, low-light enhancement, and denoising. On this dataset, we further evaluate a range of existing quality assessment models, including traditional IQA methods and large multimodal models. Finally, we provide a comprehensive analysis of how well current approaches perform in assessing the perceptual quality of AI-UGC. The access link to the AU-IQA is https://github.com/WNNGGU/AU-IQA-Dataset. Shushi Wang, Chunyi Li 0001, Han Zhou 0003, Wei Dong 0011, Jun Chen 0005, Guangtao Zhai, Xiaohong Liu 0001 |
ACM Multimedia | 7 |
| 2025 | Evaluating Perceptual Color Preferences in Smartphone Photography: Dataset and ChallengesabstractInternational audience Zhihua Wang 0002, Weixia Zhang, Wei Zhou 0021, Xiaohong Liu 0001, Guangtao Zhai, Patrick Le Callet |
ACM Multimedia | 5 |
| 2025 | HVEval: Towards Unified Evaluation of Human-Centric Video Generation and UnderstandingabstractHuman-centric videos play a significant role in the pervasive video content of modern life. However, the capabilities of text-to-video (T2V) generation models and video-to-text (V2T) understanding models for human-centric videos remain largely unexplored. To this end, we present HVEval, the first comprehensive evaluation dataset focusing on human-centric videos, which consists of 20,000 videos, 60k MOS annotations across 3 dimensions (i.e., spatial quality, temporal quality, and text-video correspondence), and 20k category-specific Q&A pairs. Based on the HVEval dataset, this paper aims to answer three questions: (1) can today's T2V models effectively generate human-centric videos following the given prompts? (2) how effective are today's V2T LMMs in understanding and evaluating human-centric videos? (3) are current VQA metrics good enough for evaluating human-centric videos? Comprehensive evaluations of 24 T2V models, 20 LMMs, and 18 VQA metrics reveal their limitations in fine-grained text-controlled generation and human-aligned perception and understanding, highlighting the significant potential of our dataset and benchmarks to advance research in human-centric video generation and understanding. Sijing Wu, Huiyu Duan, Yanwei Jiang, Yucheng Zhu, Guangtao Zhai |
ACM Multimedia | 6 |
| 2025 | FVQ: A Large-Scale Dataset and an LMM-based Method for Face Video Quality AssessmentabstractFace video quality assessment (FVQA) deserves to be explored in addition to general video quality assessment (VQA), as face videos are the primary content on social media platforms and human visual system (HVS) is particularly sensitive to human faces. However, FVQA is rarely explored due to the lack of large-scale FVQA datasets. To fill this gap, we present the first large-scale in-the-wild FVQA dataset, FVQ-20K, which contains 20,000 in-the-wild face videos together with corresponding mean opinion score (MOS) annotations. Along with the FVQ-20K dataset, we further propose a specialized FVQA method named FVQ-Rater to achieve human-like rating and scoring for face video, which is the first attempt to explore the potential of large multimodal models (LMMs) for the FVQA task. Concretely, we elaborately extract multi-dimensional features including spatial features, temporal features, and face-specific features (i.e., portrait features and face embeddings) to provide comprehensive visual information, and take advantage of the LoRA-based instruction tuning technique to achieve quality-specific fine-tuning, which shows superior performance on both FVQ-20K and CFVQA datasets. Extensive experiments and comprehensive analysis demonstrate the significant potential of the FVQ-20K dataset and FVQ-Rater method in promoting the development of FVQA. The code and dataset will be released at: https://github.com/wsj-sjtu/FVQ. Sijing Wu, Ziwen Xu, Huiyu Duan, Wei Sun 0029, Guangtao Zhai |
ACM Multimedia | 7 |
| 2025 | 3DGS-IEval-15K: A Large-scale Image Quality Evaluation Database for 3D Gaussian-Splatting
Yuke Xing, Peizhi Niu, Guangtao Zhai, Yiling Xu |
ACM Multimedia | 5 |
| 2025 | LMM4Edit: Benchmarking and Evaluating Multimodal Image Editing with LMMsabstractThe rapid advancement of Text-guided Image Editing (TIE) enables image modifications through text prompts. However, current TIE models still struggle to balance image quality, editing alignment, and consistency with the original image, limiting their practical applications. Existing TIE evaluation benchmarks and metrics have limitations on scale or alignment with human perception. To this end, we introduce EBench-18K, the first large-scale image Editing Benchmark including 18K edited images with fine-grained human preference annotations for evaluating TIE. Specifically, EBench-18K includes 1,080 source images with corresponding editing prompts across 21 tasks, 18K+ edited images produced by 17 state-of-the-art TIE models, 55K+ mean opinion scores (MOSs) assessed from three evaluation dimensions, and 18K+ question-answering (QA) pairs. Based on EBench-18K, we employ outstanding LMMs to assess edited images, while the evaluation results, in turn, provide insights into assessing the alignment between the LMMs' understanding ability and human preferences. Then, we propose LMM4Edit, a LMM-based metric for evaluating image Editing models from perceptual quality, editing alignment, attribute preservation, and task-specific QA accuracy in an all-in-one manner. Extensive experiments show that LMM4Edit achieves outstanding performance and aligns well with human preference. Zero-shot validation on the other datasets also shows the generalization ability of our model. The dataset and code are available at https://github.com/IntMeGroup/LMM4Edit. Zitong Xu, Huiyu Duan, Bingnan Liu, Guangji Ma, Shiqi Gao, Jia Wang 0004, Xiongkuo Min, Guangtao Zhai, Weisi Lin |
ACM Multimedia | 11 |
| 2025 | Omni2: Unifying Omnidirectional Image Generation and Editing in an Omni Modelabstract360° omnidirectional images (ODIs) have gained considerable attention recently, and are widely used in various virtual reality (VR) and augmented reality (AR) applications. However, capturing such images is expensive and requires specialized equipment, making ODI synthesis increasingly important. While common 2D image generation and editing methods are rapidly advancing, these models struggle to deliver satisfactory results when generating or editing ODIs due to the unique format and broad 360° Field-of-View (FoV) of ODIs. To bridge this gap, we construct Any2Omni , the first comprehensive ODI generation-editing dataset comprises 60,000+ training data covering diverse input conditions and up to 9 ODI generation and editing tasks. Built upon Any2Omni, we propose an Omni model for Omni-directional image generation and editing ( Omni 2), with the capability of handling various ODI generation and editing tasks under diverse input conditions using one model. Extensive experiments demonstrate the superiority and effectiveness of the proposed Omni2 model for both the ODI generation and editing tasks. Both the Any2Omni dataset and the Omni2 model are publicly available at: https://github.com/IntMeGroup/Omni2. Huiyu Duan, Yucheng Zhu, Xiaohong Liu 0001, Lu Liu 0005, Zitong Xu, Guangji Ma, Xiongkuo Min, Guangtao Zhai, Patrick Le Callet |
ACM Multimedia | 9 |
| 2025 | LMME3DHF: Benchmarking and Evaluating Multimodal 3D Human Face Generation with LMMsabstractThe rapid advancement in generative artificial intelligence have enabled the creation of 3D human faces (HFs) for applications including media production, virtual reality, security, healthcare, and game development, etc. However, assessing the quality and realism of these AI-generated 3D human faces remains a significant challenge due to the subjective nature of human perception and innate perceptual sensitivity to facial features. To this end, we conduct a comprehensive study on the quality assessment of AI-generated 3D human faces. We first introduce Gen3DHF, a large-scale benchmark comprising 2,000 videos of AI-Generated 3D Human Faces along with 4,000 Mean Opinion Scores (MOS) collected across two dimensions, i.e., quality and authenticity, 2,000 distortion-aware saliency maps and distortion descriptions. Based on Gen3DHF, we propose LMME3DHF, a Large Multimodal Model (LMM)-based metric for Evaluating 3DHF capable of quality and authenticity score prediction, distortion-aware visual question answering, and distortion-aware saliency prediction. Experimental results show that LMME3DHF achieves state-of-the-art performance, surpassing existing methods in both accurately predicting quality scores for AI-generated 3D human faces and effectively identifying distortion-aware salient regions and distortion types, while maintaining strong alignment with human perceptual judgments. Both the Gen3DHF database and the LMME3DHF will be released upon the publication. Woo Yi Yang, Sijing Wu, Huiyu Duan, Guangtao Zhai, Xiongkuo Min |
ACM Multimedia | 8 |
| 2025 | Human-Activity AGV Quality Assessment: A Benchmark Dataset and an Objective Evaluation MetricabstractAI-driven video generation techniques have made significant progress in recent years. However, AI-generated videos (AGVs) involving human activities often exhibit substantial visual and semantic distortions, hindering the practical application of video generation technologies in real-world scenarios. To address this challenge, we conduct a pioneering study on human activity AGV quality assessment, focusing on visual quality evaluation and the identification of semantic distortions. First, we construct the AI-Generated Human activity Video Quality Assessment (Human-AGVQA) dataset, consisting of 6,000 AGVs derived from 15 popular text-to-video (T2V) models using 400 text prompts that describe diverse human activities. We conduct a subjective study to evaluate the human appearance quality, action continuity quality, and overall video quality of AGVs, and identify semantic issues of human body parts. Based on Human-AGVQA, we benchmark the performance of T2V models and analyze their strengths and weaknesses in generating different categories of human activities. Second, we develop an objective evaluation metric, named AI-Generated Human activity Video Quality metric (GHVQ), to automatically analyze the quality of human activity AGVs. GHVQ systematically extracts human-focused quality features, AI-generated content-aware quality features, and temporal continuity features, making it a comprehensive and explainable quality metric for human activity AGVs. The extensive experimental results show that GHVQ outperforms existing quality metrics on the Human-AGVQA dataset by a large margin, demonstrating its efficacy in assessing the quality of human activity AGVs. The Human-AGVQA dataset and GHVQ metric will be released at https://github.com/zczhang-sjtu/GHVQ.git. Wei Sun 0029, Xinyue Li 0001, Qihang Ge, Jun Jia, Zhongpeng Ji, Fengyu Sun, Shangling Jui, Xiongkuo Min, Guangtao Zhai |
ACM Multimedia | 12 |
| 2025 | GOBench: Benchmarking Geometric Optics Generation and Understanding of MLLMsabstractThe rapid evolution of Multi-modality Large Language Models (MLLMs) is driving significant advancements in visual understanding and generation. Nevertheless, a comprehensive assessment of their capabilities, concerning the fine-grained physical principles especially in geometric optics, remains underexplored. To address this gap, we introduce GOBench, the first benchmark to systematically evaluate MLLMs' ability across two tasks: 1) Generating Optically Authentic Imagery and 2) Understanding Underlying Optical Phenomena. We curate high-quality prompts of geometric optical scenarios and use MLLMs to construct the GOBench-Gen-1k dataset. We then organize subjective experiments to assess the generated imagery based on Optical Authenticity, Aesthetic Quality, and Instruction Fidelity, revealing MLLMs' generation flaws that violate optical principles. For the understanding task, we apply crafted evaluation instructions to test the optical understanding ability of eleven prominent MLLMs. The experimental results demonstrate that current models face significant challenges in both optical generation and understanding. The top-performing generative model, GPT-4o-Image, cannot perfectly complete all generation tasks, and the best-performing MLLM model, Gemini-2.5Pro, attains a mere 37.35% accuracy in optical understanding. Database and codes are publicly available at: https://github.com/aiben-ch/GOBench. Xiaorong Zhu, Ziheng Jia, Haodong Duan, Xiongkuo Min, Jia Wang 0004, Guangtao Zhai |
ACM Multimedia | 9 |
| 2025 | CompBench: Benchmarking and Comparing Image Generation with Large Multimodal ModelsabstractRecent advancements in large multimodal models (LMMs) have significantly enhanced both text-to-image (T2I) generation and image-to-text (I2T) interpretation. However, critical challenges in perceptual quality and text-image correspondence remain hindering the practicality of AI-generated images (AIGIs). Therefore, a reliable benchmark and automatic model for AIGI evaluation is desirable, which heavily relies on the scale and quality of human annotations. To this end, we present CompBench, the largest dataset for benchmarking and comparing image generation, which features: (i) the largest AIGI pair comparison dataset, comprising 616,346 carefully curated image pairs generated by 24 state-of-the-art AIGI models annotated with 1.6M+ human annotations, enabling robust relative quality assessment through pairwise comparison, (ii) multi-dimensional pairwise comparison from perceptual and text-image correspondence perspectives across three difficulty levels, and (iii) bidirectional benchmarking and evaluating for both T2I generation models and AIGI comparison models. Based on CompBench, we propose LMM4Comp, a LMM-based evaluation metric that learns nuanced quality distinctions from multiple dimensions for pairwise comparison at both instance level and model level. Experiments demonstrate that LMM4Comp achieves state-of-the-art performance, highly aligning to human preference. Both of the CompBench dataset and LMM4Comp metric will be released at https://github.com/IntMeGroup/CompBench. Huiyu Duan, Yuke Xing, Yiling Xu, Guangtao Zhai, Xiongkuo Min |
MMSP | 5 |
| 2025 | AnimateQR: Bridging Aesthetics and Functionality in Dynamic QR Code GenerationabstractAnimated QR codes present an exciting frontier for dynamic content delivery and digital interaction. However, despite their potential, there has been no prior work focusing on the generation of animated QR codes that are both visually appealing and universally scannable. In this paper, we introduce AnimateQR, **the first generative framework** for creating **animated QR codes** that balance aesthetic flexibility with scannability. Unlike previous methods that focus on static QR codes, AnimateQR leverages **hierarchical luminance guidance** and **progressive spatiotemporal control** to produce high-quality dynamic QR codes. Our first innovation is a multi-scale hierarchical control signal that adjusts luminance across different spatial scales, ensuring that the QR code remains decodable while allowing for artistic expression. The second innovation is a progressive control mechanism that dynamically adjusts spatiotemporal guidance throughout the diffusion denoising steps, enabling fine-grained balance between visual quality and scannability. Extensive experimental results demonstrate that AnimateQR achieves state-of-the-art performance in both decoding success rates (96\% vs. 56\% baseline) and visual quality (user preference: 7.2 vs. 2.3 on a 10-point scale). Codes are availble at https://github.com/mulns/AnimateQR. Guangyang Wu, Huayu Zheng, Guangtao Zhai, Xiaohong Liu 0001 |
NeurIPS | 4 |
| 2025 | Envisioning Beyond the Pixels: Benchmarking Reasoning-Informed Visual EditingabstractLarge Multi-modality Models (LMMs) have made significant progress in visual understanding and generation, but they still face challenges in General Visual Editing, particularly in following complex instructions, preserving appearance consistency, and supporting flexible input formats. To study this gap, we introduce RISEBench, the first benchmark for evaluating Reasoning-Informed viSual Editing (RISE). RISEBench focuses on four key reasoning categories: Temporal, Causal, Spatial, and Logical Reasoning. We curate high-quality test cases for each category and propose an robust evaluation framework that assesses Instruction Reasoning, Appearance Consistency, and Visual Plausibility with both human judges and the LMM-as-a-judge approach. We conducted experiments evaluating nine prominent visual editing models, comprising both open-source and proprietary models. The evaluation results demonstrate that current models face significant challenges in reasoning-based editing tasks. Even the most powerful model evaluated, GPT-image-1, achieves an accuracy of merely 28.8%. RISEBench effectively highlights the limitations of contemporary editing models, provides valuable insights, and indicates potential future directions for the field of reasoning-aware visual editing. Our code and data have been released at https://github.com/PhoenixZ810/RISEBench. Peiyuan Zhang, Kexian Tang, Xiaorong Zhu, Hao Li 0069, Wenhao Chai, Renqiu Xia, Guangtao Zhai, Junchi Yan, Hua Yang 0001, Xue Yang 0005, Haodong Duan |
NeurIPS | 9 |
| 2025 | Mixed-Reference Quality Assessment for Novel View Synthesis Scenes
Huiyu Duan, Dong Zhang 0006, Yongming Han, Guangtao Zhai |
PRCV (8) | 7 |
| 2025 | A Light-Aware Quality Assessment Method for Relighted Human Heads Based on Multi-task Learning
Farong Wen, Yingjie Zhou 0003, Xiaohong Liu 0001, Jia Wang 0004, Jiezhang Cao, Yu Wang 0002, Guangtao Zhai |
PRCV (12) | 8 |
| 2025 | Learning from Model Rankings Improves Blind Super-Resolution Image Quality AssessmentabstractImage super-resolution (SR) aims to generate a high-resolution (HR) image from a low-resolution (LR) input. Traditionally, full-reference image quality assessment (FR-IQA) models have been widely used to evaluate the perceptual quality of super-resolved images, relying on pristine reference images as the gold standard. However, in real-world SR applications, such reference images are often unavailable, posing challenges for the use of FR-IQA. While blind image quality assessment (BIQA) models can assess the perceptual quality of super-resolved images without requiring a reference, there remains a lack of comprehensive studies evaluating the effectiveness of existing BIQA models for real-world SR tasks. This dilemma can largely be attributed to the high cost of subjective testing required to collect sufficient human quality annotations, which in turn hinders the development of effective SR-IQA models. In this study, we tackle this challenge with a data-efficient approach. We first generate super-resolved images from LR inputs using state-of-the-art real-world SR methods. Then, we use the maximum differentiation competition (MAD) to select a diverse set of images for subjective testing, allowing us to efficiently gather human preferences and assess the alignment between BIQA model predictions and human judgments. The resulting global ranking of SR methods not only indicates the relative performance of recent real-world SR models, but also gives us an opportunity to develop a new BIQA model tailored for real-world SR-IQA. By utilizing the global rankings of SR algorithms as prior knowledge, we can refine pretrained BIQA models using vast amounts of super-resolved images without any supervisory signal. Experimental results show that our approach substantially enhances IQA performance for real-world SR while preserving robust predictive accuracy across various distortion scenarios. The dataset and the code are available at https://github.com/cschenjunlin/SR-IQA-SMC25. Junlin Chen, Peibei Cao, Guangtao Zhai, Xiaokang Yang 0001, Weixia Zhang |
SMC | 3 |
| 2025 | MM-MQA: Multi-Modal Learning for No-reference 3D Colored Mesh Models Quality AssessmentabstractThe proliferation of 3D colored mesh models lead to growing scholarly interest in assessing their visual quality. Existing 3D mesh quality assessment methods primarily rely on single-modal features derived from either 2D projections or the three-dimensional structure of the model. 2D projections contain rich semantic and texture information, but they cannot show enough quality loss caused by structured distortion. However, the 3D structure of the model can accurately reflect geometric distortions but does not fully utilize color information, making it difficult to comprehensively represent the perception of human visual system of complex distortions. Therefore, we propose a multi-modal learning for no-reference 3D colored mesh models quality assessment method (MM-MQA), First, we obtain projection images of the 3D colored mesh model from different viewpoints and split the model into equally sized patches. We then employ MeshNet and ResNet to encode the structural features of the model patches and the texture features of the projection images. Finally, double cross-attention is employed to achieve multi-modal fusion to perceive the overall quality of the mesh model. Experiments demonstrate that our method outperforms existing approaches in both CMDM and TMQ datasets, validating the effectiveness of multi-modal learning in mesh quality assessment tasks. Guoquan Zheng, Dong Zhang 0006, Guangtao Zhai |
SMC | 6 |
| 2025 | Ges-QA: A Multidimensional Quality Assessment Dataset for Audio-to-3D Gesture GenerationabstractThe Audio-to-3D-Gesture (A2G) task exhibits significant potential across domains including virtual reality, computer graphics, and 3D animation production. However, current evaluation metrics, such as Fréchet Gesture Distance or Beat Constancy, fail at reflecting the human preference of the generated 3D gestures. To cope with this problem, exploring human preference and an objective quality assessment metric for AI-generated 3D human gestures is becoming increasingly significant. In this paper, we introduce the Ges-QA dataset, which includes 1,400 samples with multidimensional scores for gesture quality and audio-gesture consistency. Moreover, we collect binary classification labels to determine whether the generated gestures match the emotions of the audio. Equipped with our Ges-QA dataset, we propose a multi-modal transformer-based neural network with 3 branches for video, audio and 3D skeleton modalities, which can score A2G contents in multiple dimensions. Comparative experimental results and ablation studies demonstrate that Ges-QAer yields state-of-the-art performance on our dataset. Zhilin Gao, Sijing Wu, Yuqin Cao, Huiyu Duan, Guangtao Zhai |
VCIP | 6 |
| 2025 | CVBench: Benchmarking and Comparing Video Generation with Large Multimodal ModelsabstractLarge multimodal models (LMMs) have revolutionized both text-to-video (T2V) generation and video-to-text (V2T) interpretation. However, despite these advancements, issues such as imperfect perceptual quality and inconsistent text-video alignment continue to limit the practical deployment of AI-generated videos (AIGVs). Consequently, there is a pressing need for a reliable benchmark and automatic evaluation framework tailored for AIGVs. To this end, we propose CVBench, the largest and most comprehensive dataset for Comparative Video Benchmarking, including 60K video pairs generated by 30 state-of-the-art T2V models and 600K pairwise comparisons annotated with over 1.7 million human judgments from perspectives of both perceptual quality and text-video correspondence. This dataset enables bidirectional benchmarking and evaluation of both T2V generation models and V2T interpretation models. Based on CVBench, we propose VComp, a novel LMM-based evaluation metric that captures fine-grained quality differences from multiple perspectives for pairwise comparison at both the instance level and model level. Extensive experiments show that VComp achieves state-of-the-art alignment with human preferences. Both the CVBench dataset and VComp metric will be available at https://github.com/IntMeGroup/CVBench. Huiyu Duan, Yuke Xing, Wei Zhou 0021, Guangtao Zhai, Xiongkuo Min |
VCIP | 5 |
| 2025 | MGM-DPO: Multi-Generative-Model Guided Preference Optimization for Advanced Text-to-Image ModelsabstractText-to-image generation has rapidly progressed with autoregressive, diffusion, and distillation-based models, enabling high-fidelity and semantically relevant image synthesis from natural language prompts. Despite their success, these models still suffer from text–image misalignment, limited detail, and outputs that may violate human commonsense or aesthetic preferences. Existing human-feedback-based fine-tuning approaches partially alleviate these issues but primarily focus on earlier diffusion models (e.g., Stable Diffusion v1.4, v2.0, SDXL) and often exhibit over-optimization or under-optimization, leaving their effectiveness on state-of-the-art models unclear. In this work, we advance preference alignment for modern text-to-image models. We utilize EvalMi-50K, a large-scale dataset with human preference scores, to train our image quality assessment (IQA) model and use the predicted scores from the IQA model to reward the generation model. Building upon the dataset and IQA model, we propose Multi-Generative-Model Guided Diffusion Preference Optimization(MGM-DPO), which leverages diverse model outputs and a stable optimization strategy to align advanced diffusion models. Applied to Stable Diffusion 3.5 Large, MGM-DPO significantly improves fidelity, prompt alignment, and human preference alignment across multiple benchmarks, achieving state-of-the-art results among human-feedback-tuned models. Code is available at https://github.com/IntMeGroup/MGM-DPO. Huiyu Duan, Guangtao Zhai, Xiongkuo Min |
VCIP | 4 |
| 2025 | Subjective and Objective Quality-of-Experience Evaluation Study for Live Video StreamingabstractIn recent years, live video streaming has gained widespread popularity across various social media platforms. Quality of experience (QoE), which reflects end-users’ satisfaction and overall experience, plays a critical role for media service providers to optimize large-scale live compression and transmission strategies to achieve perceptually optimal rate-distortion trade-off. Although many QoE metrics for video-on-demand (VoD) have been proposed, there remain significant challenges in developing QoE metrics for live video streaming. To bridge this gap, we conduct a comprehensive study of subjective and objective QoE evaluations for live video streaming. For the subjective QoE study, we introduce the first live video streaming QoE dataset, TaoLive QoE, which consists of 42 source videos collected from real live broadcasts and 1, 155 corresponding distorted ones degraded due to a variety of streaming distortions, including conventional streaming distortions such as compression, stalling, as well as live streaming-specific distortions like frame skipping, variable frame rate, etc. Subsequently, a human study was conducted to derive subjective QoE scores of videos in the TaoLive QoE dataset. For the objective QoE study, we benchmark existing QoE models on the TaoLive QoE dataset as well as publicly available QoE datasets for VoD scenarios, highlighting that current models struggle to accurately assess video QoE, particularly for live content. Hence, we propose an end-to-end QoE evaluation model, Tao-QoE, which integrates multi-scale semantic features and optical flow-based motion features to predicting a retrospective QoE score, eliminating reliance on statistical quality of service (QoS) features. Extensive experiments demonstrate that Tao-QoE outperforms other models on the TaoLive QoE dataset and five publicly available QoE datasets, showcasing the effectiveness and feasibility of Tao-QoE. Zehao Zhu, Wei Sun 0029, Jun Jia, Jia Wang 0004, Guangtao Zhai |
VCIP | 5 |
| 2025 | Situation-adaptive neural network for fast pre-computing image enhancement
Xinyue Li 0001, Huiyu Duan, Jia Wang 0004, Xiaohong Liu 0001, Guangtao Zhai |
Sci. China Inf. Sci. | 6 |
| 2025 | Large multimodal models evaluation: a survey
Farong Wen, Yijin Guo, Xinyu Fang, Shengyuan Ding, Ziheng Jia, Jiahao Xiao, Ye Shen, Yushuo Zheng, Xiaorong Zhu, Yalun Wu, Ziheng Jiao, Wei Sun 0029, Zijian Chen 0001, Kaiwei Zhang, Yuqin Cao, Yue Zhou 0005, Xuemei Zhou, Juntai Cao, Wei Zhou 0021, Jinyu Cao, Ronghui Li, Yuan Tian 0017, Chunyi Li 0001, Haoning Wu 0001, Xiaohong Liu 0001, Junjun He, Yu Zhou 0016, Zesheng Wang 0004, Huiyu Duan, Yingjie Zhou 0003, Xiongkuo Min, Dongzhan Zhou, Jiezhang Cao, Xue Yang 0005, Junzhi Yu 0001, Songyang Zhang 0001, Haodong Duan, Guangtao Zhai |
Sci. China Inf. Sci. | 49 |
| 2025 | A study on the user viewing experience of implanted advertisement videos based on visual saliency
Fangfang Lu, Yingjie Lian, Wei Sun 0029, Xiongkuo Min, Guangtao Zhai |
Expert Syst. Appl. | 5 |
| 2025 | VARFVV: View-Adaptive Real-Time Interactive Free-View Video Streaming With Edge ComputingabstractFree-view video (FVV) allows users to explore immersive video content from multiple views. However, delivering FVV poses significant challenges due to the uncertainty in view switching, combined with the substantial bandwidth and computational resources required to transmit and decode multiple video streams, which may result in frequent playback interruptions. Existing approaches, either client-based or cloud-based, struggle to meet high Quality of Experience (QoE) requirements under limited bandwidth and computational resources. To address these issues, we propose VARFVV, a bandwidth- and computationally-efficient system that enables real-time interactive FVV streaming with high QoE and low switching delay. Specifically, VARFVV introduces a low-complexity FVV generation scheme that reassembles multiview video frames at the edge server based on user-selected view tracks, eliminating the need for transcoding and significantly reducing computational overhead. This design makes it well-suited for large-scale, mobile-based UHD FVV experiences. Furthermore, we present a popularity-adaptive bit allocation method, leveraging a graph neural network, that predicts view popularity and dynamically adjusts bit allocation to maximize QoE within bandwidth constraints. We also construct an FVV dataset comprising 330 videos from 10 scenes, including basketball, opera, etc. Extensive experiments show that VARFVV surpasses existing methods in video quality, switching latency, computational efficiency, and bandwidth usage, supporting over 500 users on a single edge server with a switching delay of 71.5ms. Our code and dataset are available at https://github.com/qianghu-huber/VARFVV. Qiang Hu 0003, Qihan He, Houqiang Zhong, Guo Lu, Xiaoyun Zhang 0001, Guangtao Zhai, Yanfeng Wang 0001 |
IEEE J. Sel. Areas Commun. | 6 |
| 2025 | Content adaptive JND profile by leveraging HVS inspired channel modeling and perception oriented energy allocation optimization
Haibing Yin, Xia Wang 0006, Guangtao Zhai, Xiaofei Zhou 0003, Chenggang Yan 0001 |
Signal Process. | 3 |
| 2025 | CT-PCQA: A Convolutional Neural Network and Transformer combined Method for Point Cloud Quality Assessment
Yingjie Zhou 0003, Wei Sun 0029, Xiongkuo Min, Guangtao Zhai |
Signal Process. Image Commun. | 5 |
| 2025 | CoughSlowFast: Cough Recognition With Audio and Video Signal Fusion
Mingke Feng, Guangtao Zhai, Xiao-Ping Zhang 0002, Menghan Hu |
IEEE Signal Process. Lett. | 2 |
| 2025 | FsPN: Blind Image Quality Assessment Based on Feature-Selected Pyramid NetworkabstractBlind image quality assessment (BIQA) is crucial for user satisfaction and the performance of various image processing applications. Most BIQA methods directly use the pre-trained model to extract features and then perform feature fusion. However, the features extracted by pre-trained models may contain irrelevant information to BIQA. Although some methodspre-train the feature extraction network from scratch, these approaches raise computational costs and resource demands. In this letter, a Feature-selected Pyramid Network(FsPN) is proposed to address this issue from a different perspective. First, a spatial selection module selects useful information from the features extracted by the pre-trained model. Additionally, a pyramid network based on skip connections is utilized to fuse the selected multi-scale features. The proposed method is verified in six public datasets, where it consistently outperformed existing state-of-the-art methods, affirming its effectiveness and adaptability. Yongming Han, Guangtao Zhai |
IEEE Signal Process. Lett. | 4 |
| 2025 | Real-Time Respiration Monitoring via Motion Artifact Suppression and Quality-Guided Peak DetectionabstractReal-time respiration monitoring faces several challenges including network latency in remote settings, limited computational resources, and increased motion artifacts. Although many existing non-contact respiration algorithms are designed for offline processing and thus overlook these limitations, real-time applications demand greater robustness, efficiency, and adaptability to dynamic conditions. In this study, a lightweight framework called Quality-Guided Respiration Monitoring (QGRM) is proposed. This framework integrates a two-stage motion artifact suppression module and a quality-guided peak detection (QGPD) module. The former enhances signal stability through FIR filtering and amplitude limiting, while the latter improves the estimation of the respiration rate by filtering false peaks based on amplitude and zero-crossing constraints. The experimental results obtained with both the public OVRM dataset and a self-constructed simulated dataset demonstrate that QGRM achieves superior accuracy and robustness compared to state-of-the-art methods. The dataset and code are available athttps://github.com/zxx5058/QGRM. Chenrui Niu, Zhanzhan Cheng, Nengfeng Qian, Changyin Wu, Guangtao Zhai, Menghan Hu |
IEEE Signal Process. Lett. | 6 |
| 2025 | Energy-Efficient VR 360 Video Streaming in the IRS-Aided Rate-Splitting Multiple Access NetworkabstractMaximizing energy efficiency in VR 360 video transmission is essential for advancing VR applications. Motivated by this goal, we conduct a comprehensive study that integrates the characteristics of VR 360 video with beamforming strategies and intelligent reflecting surface (IRS) shifting techniques in the rate-splitting multiple access (RSMA) network. In the IRS-aided RSMA network, we propose a stable energy-efficient transmission (SEET) scheme aimed at minimizing the number of transmitted VR video chunks. The SEET scheme constructs a stable pre-transmission and playback flow, ensuring seamless and continuous display of the upcoming content without latency. We also propose a mixed-format-based chunk (MFC) method that simultaneously pre-transmits both 2D and 3D chunk frames to each user, further enhancing energy efficiency. We utilize an alternating optimization method to divide the original energy-efficient problem into three subproblems. To tackle the non-convex and NP-hard beamforming subproblem, we utilize the first-order Taylor expansion and then obtain the approximate transmission rates of common messages and private messages regarding the quadratic form of beamforming vectors. We then utilize quadratically constrained programming, fractional programming, and linear programming to obtain the near-optimal solutions for beamforming vectors, IRS phase shifts, and RSMA parameters, respectively. The final numerical results affirm that the proposed SEET scheme can notably minimize the beamforming power of the base station. Through the SEET scheme, the MFC method with the approximation method exhibits superior energy efficiency, outperforming existing transmission methods in terms of both energy utility and consumption by HMDs. Qingqing Wu 0001, Huiyu Duan, Xiongkuo Min, Guangtao Zhai |
IEEE Trans. Commun. | 5 |
| 2025 | Time-Smooth Wireless Transmission of Probabilistic Slicing VR 360 Video in MISO-OFDM SystemsabstractThe multiple-input and single-output (MISO)-orthogonal frequency-division multiplexing (OFDM) systems afford low latency and high reliability for virtual reality (VR) 360 video in multi-user scenarios. Motivated by the goal of maintaining time-smoothness while holding acceptably low complexity, a crucial factor in VR video transmission, we conduct a comprehensive study that integrates the characteristics of VR video with the strategies for subcarrier assignment and power allocation. By analyzing the pre-transmitted tile-segments, the missing tile-segments, and the video frame structure, we propose two probabilistic slicing schemes (PSPs) to minimize the size of required tile-segments of VR video scenes. In time-smoothness maximization, the desired discrete encoding rate set, discrete subcarrier assignment, continuous power allocation, and fixed total power constraint make it a challenging mixed-integer nonlinear programming (MINLP) problem. Unlike the straightforward relaxation-recovery method, we firstly prove that a near-optimal recovered encoding rate is the discrete value closest to the optimal relaxed-continuous encoding rate. We then propose a Two-step Encoding Rate Maximization (TERM) method, including the relaxed-continuous sum-rate maximization and the discrete encoding rate recovery, to achieve the near-optimal subcarrier assignment and the power allocation with low complexity. Simulation results on real-world VR video dataset validate that the two PSPs can effectively minimize the number of transmitted tile-segments. The proposed TERM with PSPs can maintain time-smoothness of VR 360 video with an acceptably low level of complexity in MISO-OFDM systems. Guangtao Zhai, Yongpeng Wu 0001, Xiongkuo Min, Biqian Feng, Yucheng Zhu, Wenjun Zhang 0001 |
IEEE Trans. Commun. | 2 |
| 2025 | Multi-Scale Local and Global Feature Fusion for Blind Quality Assessment of Enhanced ImagesabstractImage enhancement plays a crucial role in computer vision by improving visual quality while minimizing distortion. Traditional methods enhance images through pixel value transformations, yet they often introduce new distortions. Recent advancements in deep learning-based techniques promise better results but challenge the preservation of image fidelity. Therefore, it is essential to evaluate the visual quality of enhanced images. However, existing quality assessment methods frequently encounter difficulties due to the unique distortions introduced by these enhancements, thereby restricting their effectiveness. To address these challenges, this paper proposes a novel blind image quality assessment (BIQA) method for enhanced natural images, termed multi-scale local feature fusion and global feature representation-based quality assessment (MLGQA). This model integrates three key components: a multi-scale Feature Attention Mechanism (FAM) for local feature extraction, a Local Feature Fusion (LFF) module for cross-scale feature synthesis, and a Global Feature Representation (GFR) module using Vision Transformers to capture global perceptual attributes. This synergistic framework effectively captures both fine-grained local distortions and broader global features that collectively define the visual quality of enhanced images. Furthermore, in the absence of a dedicated benchmark for enhanced natural images, we design the Natural Image Enhancement Database (NIED), a large-scale dataset consisting of 8,581 original images and 102,972 enhanced natural images generated through a wide array of traditional and deep learning-based enhancement techniques. Extensive experiments on NIED demonstrate that the proposed MLGQA model significantly outperforms current state-of-the-art BIQA methods in terms of both prediction accuracy and robustness. Jingchao Cao, Yutao Liu 0002, Feng Gao 0005, Ke Gu 0001, Guangtao Zhai, Junyu Dong, Sam Kwong |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2025 | Joint Luminance-Chrominance Learning for Image DebandingabstractBanding is a visually annoying artifact that frequently occurs along the chain of video acquisition, production, distribution, and display, showing a significant need for improvement in many fields. Thus far, efforts on banding removal are mainly knowledge-driven or merely learning on RGB space, which is either limited by domain knowledge or lacks the consideration for banding in chrominance channels. In this work, we propose a unified deep neural network that explicitly disentangles the luminance and chrominance channels, and simultaneously recovers intensity gradients and color discontinuity from detection-free measurement in an end-to-end manner. Our debanding model is comprised of a luminance restoration network (LR-Net) and a chrominance restoration network (CR-Net). Each of them follows an encoder-decoder architecture, where a cascade of residual blocks is employed to exploit hierarchical non-local features in spatial dimensions for more powerful feature representation. Moreover, we investigate the characteristics of banding artifacts and apply specific loss functions to guide the debanding in different channels, thus boosting the restoration performance. Both qualitative and quantitative experiments show that our model significantly surpasses the existing method in terms of all 7 metrics. Ultimately, our network trained on simulated data exhibits good adaptiveness under various compression scenarios, which further demonstrates the effectiveness of the proposed model. Zijian Chen 0001, Wei Sun 0029, Jun Jia, Ru Huang 0002, Fangfang Lu, Ying Chen 0011, Xiongkuo Min, Guangtao Zhai, Wenjun Zhang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 8 |
| 2025 | Study of Subjective and Objective Naturalness Assessment of AI-Generated ImagesabstractThe proliferation of Artificial Intelligence-Generated Images (AIGIs) has greatly expanded the Image Naturalness Assessment (INA) problem. Different from early definitions that mainly focus on tone-mapped images with limited distortions (e.g., exposure, contrast, and color reproduction), INA on AI-generated images is especially challenging as it owns more diverse contents and could be affected by factors from multiple perspectives, including low-level technical distortions and high-level rationality distortions. In this paper, we take the first step to benchmark and assess the visual naturalness of AI-generated images. First, we construct the AI-Generated Image Naturalness (AGIN) dataset by conducting a large-scale subjective study to collect human opinions on the overall naturalness as well as perceptions from the technical quality and rationality perspectives. AGIN verifies several insights for the first time that naturalness is universally and disparately affected by both technical and rational distortions, while its manifestations vary with different generation tasks. Second, to automatically assess the naturalness of AIGIs that align with human opinions, we propose the Joint Objective Image Naturalness evaluaTor (JOINT). Specifically, JOINT imitates human reasoning in naturalness evaluation by jointly learning technical and rationality features with several specific designs to guide model behavior from respective perspectives. Experiments demonstrate that JOINT significantly outperforms existing methods for providing more subjectively consistent results on naturalness assessment. The dataset can be accessed athttps://github.com/zijianchen98/AGIN. Zijian Chen 0001, Wei Sun 0029, Haoning Wu 0001, Jun Jia, Ru Huang 0002, Xiongkuo Min, Guangtao Zhai, Wenjun Zhang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 8 |
| 2025 | No-Reference Image Quality Assessment: Obtain MOS From Image Quality Score DistributionabstractRecent image quality assessment (IQA) methods typically focus on predicting the mean opinion score (MOS) of image quality, ignoring the image quality score distribution. This distribution provides valuable information beyond the MOS, including the standard deviation of opinion scores (SOS) and opinion scores at different quality levels. This paper introduces a novel no-reference IQA method that predicts the image quality score distribution to estimate the MOS. The proposed method consists of three modules: a visual feature extraction module, a graph convolutional module, and a MOS prediction module. In the visual feature extraction module, a convolutional neural network is designed to extract both first- and second-order visual features of images. The graph convolutional module employs a graph convolutional network (GCN)-based mapper to map these visual features to the image quality score distribution by exploring correlations between quality labels. The MOS is then derived from the predicted image quality score distribution in the MOS prediction module. We are the first to jointly train the method using both the MOS and the image quality score distribution, enabling it to learn richer subjective information and improve prediction performance. To address the lack of the ground-truth image quality score distribution in some IQA databases, we propose to use a SOS assumption to generate a Gaussian-based image quality score distribution that better reflects subjective perception. Additionally, we design appropriate loss functions for training. Experimental results demonstrate that our method effectively predicts both the image quality score distribution and the MOS, outperforming most state-of-the-art IQA methods. Xiongkuo Min, Yuqin Cao, Xiaohong Liu 0001, Guangtao Zhai |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2025 | LMM-VQA: Advancing Video Quality Assessment With Large Multimodal ModelsabstractThe explosive growth of videos on streaming media platforms has underscored the urgent need for effective video quality assessment (VQA) algorithms to monitor and perceptually optimize the quality of streaming videos. However, VQA remains an extremely challenging task due to the diverse video content and the complex spatial and temporal distortions, thus necessitating more advanced methods to address these issues. Nowadays, large multimodal models (LMMs), such as GPT-4V, have exhibited strong capabilities for various visual understanding tasks, motivating us to leverage the powerful multimodal representation ability of LMMs to solve the VQA task. Therefore, we propose anLargeMulti-Modal basedVideoQuality Assessment (LMM-VQA) model, which introduces a novel spatiotemporal visual modeling strategy for quality-aware feature extraction. Specifically, we reformulate the quality regression problem into a question and answering (Q&A) task and construct Q&A prompts for VQA instruction tuning. Then, we design a spatiotemporal vision encoder to extract spatial and temporal features to represent the quality characteristics of videos, which are subsequently mapped into the language space by the spatiotemporal projector for modality alignment. Finally, the aligned visual tokens and the quality-inquired text tokens are aggregated as inputs for the large language model (LLM) to generate the quality score as well as the quality level. Extensive experiments demonstrate thatLMM-VQAachieves state-of-the-art performance across five VQA benchmarks, exhibiting an average improvement of 5% in generalization ability over existing methods. Furthermore, due to the advanced design of the spatiotemporal encoder and projector, LMM-VQA also performs exceptionally well on general video understanding tasks, further validating its effectiveness. Our code will be released at https://github.com/Sueqk/LMM-VQA. Qihang Ge, Wei Sun 0029, Yu Zhang 0133, Zhongpeng Ji, Fengyu Sun, Shangling Jui, Xiongkuo Min, Guangtao Zhai |
IEEE Trans. Circuits Syst. Video Technol. | 9 |
| 2025 | Perceptual Information Fidelity for Quality Estimation of Industrial ImagesabstractDepending on high quality images, industrial vision technologies can basically oversee all the industrial production processes, such as workpiece processing and assembly automation, which play a highly significant role in promoting detection automation and production capacity in assembly lines. Unlike the natural scene images which consist of richer colors and natural lines, industrial images that cover complex industrial goods and equipment are made up of fewer colors, more regular shapes, massive graphic elements, etc., causing existing image processing methods for quality estimation, enhancement and monitoring to fail. Human beings usually play the part of the final receiver of an industrial image, so in the researches of image quality estimation, it is necessary to take the perception process of human eyes and brain to the input images into consideration. On this basis, we in this paper propose a novel perceptual information fidelity based image quality estimation model, abbreviated as PIF. Particularly, we first introduce a visual-cell low-pass filter and an optical-nerve noise model, which are separately inspired by the two processes: one is that an image in the form of optical signals arrives at the retina through the eye’s optical system to form the stimuli; the other is that the aforesaid stimuli in the form of electrical signals transfer to the human brain through the optical nerve. Second, we construct a novel image content-aware adjustor to optimize the above visual-cell low-pass filter and optical-nerve noise model. Third, we compare the two quantities of the information that is present in the clean image and how much of the information can be extracted from the lossy image to generate the overall quality score. Experiments on the two large-size industrial image quality databases demonstrate the excellent performance achieved by our proposed PIF model, with a remarkable performance gain over the existing state-of-the-art competitors. Ke Gu 0001, Hongyan Liu 0004, Junfei Qiao 0001, Guangtao Zhai, Wenjun Zhang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2025 | Full-Reference and No-Reference Quality Assessment for Video Frame InterpolationabstractVideo frame interpolation (VFI) synthesizes new frames from original video frames to produce high frame-rate videos and enhance their visual appeal. The quality of these interpolated frames significantly affects the perceptual experience of the synthesized video. Recent research in VFI has increasingly focused on perceptual quality of the interpolated frames and the overall video. However, most existing quality metrics do not align well with human perceptual experiences and often suffer from unnatural artifacts in the interpolated frames. Consequently, there is an urgent need for VFI video quality assessment (VFIVQA) methods to assess the quality of the synthesized videos. In this paper, we propose both a full-reference (FR) method and a no-reference (NR) method for VFIVQA. The FR method employs two feature extraction blocks to measure continuous frame changes, extracting flow features with short temporal spans and motion features with long temporal spans. By calculating multilevel similarities in the temporal dimension of 3D convolutional neural networks and fusing these similarity features, the quality score of the VFI video is obtained from the quality regression network. Since the flow feature extraction block does not utilize the reference VFI video, the proposed NR method consists solely of this feature block. Extensive validation on several VFIVQA datasets demonstrates that the proposed methods outperform state-of-the-art FR and NR methods. Jinliang Han, Xiongkuo Min, Jun Jia, Xiaohong Liu 0001, Guangtao Zhai |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2025 | Multi-Task Guided No-Reference Omnidirectional Image Quality Assessment With Feature InteractionabstractOmnidirectional image quality assessment (OIQA) has become an increasingly vital problem in recent years. Most previous no-reference OIQA methods only extract local features from the distorted viewports, or extract global features from the entire distorted image, lacking the interaction and fusion between local and global features. Moreover, the lack of reference information also limits their performance. Thus, we propose a no-reference OIQA model which consists of three novel modules, including a bidirectional pseudo-reference module, a Mamba-based global feature extraction module, and a multi-scale local-global feature aggregation module. Specifically, by considering the image distortion degradation process, a bidirectional pseudo-reference module capturing the error maps on viewports is first constructed to refine the multi-scale local visual features, which can supply rich quality degradation reference information without the reference image. To well complement the local features, the VMamba module is adopted to extract the representative multi-scale global visual features. Inspired by human hierarchical visual perception characteristics, a novel multi-scale aggregation module is built to strengthen the feature interaction and effective fusion which can extract deep semantic information. Finally, motivated by the multi-task managing mechanism of human brain, a multi-task learning module is introduced to assist the main quality assessment task by digging the hidden information in compression type and distortion degree. Extensive experimental results demonstrate that our proposed method achieves the state-of-the-art performance on the no-reference OIQA task compared to other models. Yun Liu 0009, Huiyu Duan, Yu Zhou 0009, Daoxin Fan, Guangtao Zhai |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2025 | Exploring Rich Subjective Quality Information for Image Quality Assessment in the WildabstractTraditional in the wild image quality assessment (IQA) models are generally trained with the quality labels of mean opinion score (MOS), while missing the rich subjective quality information contained in the quality ratings, for example, the standard deviation of opinion scores (SOS) or even distribution of opinion scores (DOS). In this paper, we propose a novel IQA method namedRichIQAto explore the rich subjective rating information beyond MOS to predict image quality in the wild. RichIQA is characterized by two key novel designs: 1) a three-stage image quality prediction network which exploits the powerful feature representation capability of the Convolutional vision Transformer (CvT) and mimics the short-term and long-term memory mechanisms of human brain; 2) a multi-label training strategy in which rich subjective quality information like MOS, SOS and DOS are concurrently used to train the quality prediction network. Powered by these two novel designs, RichIQA is able to predict the image quality in terms of a distribution, from which the mean image quality can be subsequently obtained. Extensive experimental results verify that the three-stage network is tailored to predict rich quality information, while the multi-label training strategy can fully exploit the potentials within subjective quality rating and enhance the prediction performance and generalizability of the network. RichIQA outperforms state-of-the-art competitors on multiple large-scale in the wild IQA databases with rich subjective rating labels. The code of RichIQA will be made publicly available on GitHub. Xiongkuo Min, Yuqin Cao, Guangtao Zhai, Wenjun Zhang 0001, Huifang Sun, Chang Wen Chen |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2025 | Who Is a Better Imitator: Subjective and Objective Quality Assessment of Animated HumansabstractAnimated human (AH) have gained popularity due to their vivid appearance and smooth, natural movements. Various animation methods based on artificial intelligence (AI) have been introduced, which are viewed as “Imitators,” offering new solutions for designing AHs. However, the effectiveness of these AI-generated AHs varies significantly across different categories and within the same category, leading to visual distortions that adversely affect the viewer’s experience. Consequently, it is essential to evaluate the quality of AHs to provide reliable and objective indicators for their further development and to ensure the delivery of higher-quality AH videos to users. In this paper, the first Animated Human Quality Assessment (AHQA) dataset is constructed by selecting 6 advanced and popular imitators and 10 common actions to animate 20 AI-generated characters. The constructed dataset integrates different genders and age groups of character images, and two types of poses, standing and sitting, are selected, highlighting the comprehensiveness and diversity of the AHQA dataset. Subjective experiments reveal significant differences in the quality of AHs produced by different imitators. Finally, we propose a quality assessment method, VIP-QA, incorporating Video quality, Identity consistency, and Posture similarity for the AHQA dataset. Experimental results show that VIP-QA significantly outperforms existing assessment methods on multiple datasets by about 5%, more closely approximates human visual perception, and provides a valid objective metric for assessing imitators. All the work in this paper has been released at https://github.com/zyj-2000/Imitator. Yingjie Zhou 0003, Jun Jia, Yanwei Jiang, Xiaohong Liu 0001, Xiongkuo Min, Guangtao Zhai |
IEEE Trans. Circuits Syst. Video Technol. | 7 |
| 2025 | ScanDTM: A Novel Dual-Temporal Modulation Scanpath Prediction Model for Omnidirectional ImagesabstractScanpath prediction for omnidirectional images aims to effectively simulate the human visual perception mechanism to generate dynamic realistic fixation trajectories. However, the majority of scanpath prediction methods for omnidirectional images are still in their infancy as they fail to accurately capture the time-dependency of viewing behavior and suffer from sub-optimal performance along with limited generalization capability. A desirable solution should achieve a better trade-off between prediction performance and generalization ability. To this end, we propose a novel dual-temporal modulation scanpath prediction (ScanDTM) model for omnidirectional images. Such a model is designed to effectively capture long-range time-dependencies between various fixation regions across both internal and external time dimensions, thereby generating more realistic scanpaths. In particular, we design a Dual Graph Convolutional Network (Dual-GCN) module comprising a semantic-level GCN and an image-level GCN. This module servers as a robust visual encoder that captures spatial relationships among various object regions within an image and fully utilizes similar images as complementary information to capture similarity relations across relevant images. Notably, the proposed Dual-GCN focuses on modeling temporal correlations from both local and global perspectives within the internal time dimension. Furthermore, drawing inspiration from the promising generalization capabilities of diffusion models across various generative tasks, we introduce a novel diffusion-guided saliency module. This module formulates the prediction issue as a conditional generative process for the saliency map, utilizing extracted semantic-level and image-level visual features as conditions. With the well-designed diffusion-guided saliency module, our proposed ScanDTM model acting as an external temporal modulator, we can progressively refine the generated scanpath from the noisy map. We conduct extensive experiments on several benchmark datasets, and the results demonstrate that our ScanDTM model significantly outperforms other competitors. Meanwhile, when applied to tasks such as saliency prediction and image quality assessment, our ScanDTM model consistently achieves superior generalization performance. Dandan Zhu 0001, Kaiwei Zhang, Xiongkuo Min, Guangtao Zhai, Xiaokang Yang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2025 | Blind Image Quality Assessment by Gaussian Mixture DistributionabstractIn the field of image quality assessment (IQA), researchers have been studying the mean opinion score (MOS) of image quality for decades. They focus on developing IQA methods with the help of MOS without using the potential of the distribution of opinion scores (DOS). We find that the Gaussian mixture distribution (GMD) can more accurately describe the DOS of image quality on SJTU IQSD and KonIQ-10K databases compared to some traditional distributions. Therefore, this paper proposes a blind IQA method that predicts the MOS of image quality by learning the GMD-based image quality. The proposed method consists of a visual feature learning module and a GMD learning module. The visual feature learning module uses a multi-stage Swin Transformer model and a CLIP feature extractor to extract visual features from an image. The GMD learning module then maps the extracted visual features to the GMD-based image quality using a mixture density network, where the mean of the GMD represents the MOS of image quality. We not only use the MOS of image quality to train the proposed method, but also employ the DOS of image quality for auxiliary training to improve the prediction performance of the proposed method. To address the lack of DOS in some existing IQA databases, we introduce a pseudo DOS generation strategy to generate the DOS of image quality for training, which significantly improves the applicability of the proposed method. Numerous analyses show that the proposed method is superior to most state-of-the-art IQA methods in predicting both the MOS and the DOS, thus facilitating a deeper investigation into the DOS of image quality in IQA. Xiongkuo Min, Yuqin Cao, Weisi Lin, Bu-Sung Lee, Guangtao Zhai |
IEEE Trans. Image Process. | 6 |
| 2025 | MISC: Ultra-Low Bitrate Image Semantic Compression Driven by Large Multimodal ModelabstractWith the evolution of storage and communication protocols, ultra-low bitrate image compression has become a highly demanding topic. However, all existing compression algorithms must sacrifice either consistency with the ground truth or perceptual quality at ultra-low bitrate. During recent years, the rapid development of the Large Multimodal Model (LMM) has made it possible to balance these two goals. To solve this problem, this paper proposes a method called Multimodal Image Semantic Compression (MISC), which consists of an LMM encoder for extracting the semantic information of the image, a map encoder to locate the region corresponding to the semantic, an image encoder generates an extremely compressed bitstream, and a decoder reconstructs the image based on the above information. Experimental results show that our proposed MISC is suitable for compressing both traditional Natural Sense Images (NSIs) and emerging AI-Generated Images (AIGIs) content. It can achieve optimal consistency and perception results while saving 50% bitrate, which has strong potential applications in the next generation of storage and communication. The code will be released on https://github.com/lcysyzxdxc/MISC. Chunyi Li 0001, Guo Lu, Donghui Feng 0003, Haoning Wu 0001, Xiaohong Liu 0001, Guangtao Zhai, Weisi Lin, Wenjun Zhang 0001 |
IEEE Trans. Image Process. | 7 |
| 2025 | Advancing Zero-Shot Digital Human Quality Assessment Through Text-Prompted EvaluationabstractDigital humans have witnessed extensive applications in various domains, necessitating related quality assessment studies. However, there is a lack of comprehensive digital human quality assessment (DHQA) databases. To address this gap, we propose SJTU-H3D, a subjective quality assessment database specifically designed for full-body digital humans. It comprises 40 high-quality reference digital humans and 1,120 labeled distorted counterparts generated with seven types of distortions. The SJTU-H3D database can serve as a benchmark for DHQA research, allowing evaluation and refinement of processing algorithms. Further, we propose a zero-shot DHQA approach that focuses on no-reference (NR) scenarios to ensure generalization capabilities while mitigating database bias. Our method leverages semantic and distortion features extracted from projections, as well as geometry features derived from the mesh structure of digital humans. Specifically, we employ the Contrastive Language-Image Pre-training (CLIP) model to measure semantic affinity and incorporate the Naturalness Image Quality Evaluator (NIQE) model to capture low-level distortion information. Additionally, we utilize dihedral angles as geometry descriptors to extract mesh features. By aggregating these measures, we introduce the Digital Human Quality Index (DHQI), which demonstrates significant improvements in zero-shot performance. The DHQI can also serve as a robust baseline for DHQA tasks, facilitating advancements in the field. The database and the code are available at https://github.com/zzc-1998/SJTU-H3D. Wei Sun 0029, Yingjie Zhou 0003, Haoning Wu 0001, Chunyi Li 0001, Xiongkuo Min, Xiaohong Liu 0001, Guangtao Zhai, Weisi Lin |
IEEE Trans. Image Process. | 8 |
| 2025 | Subjective and Objective Audio-Visual Quality Assessment for Omnidirectional VideosabstractVirtual Reality (VR) has attracted widespread attention in recent years due to its capability to create immersive experiences by presenting multi-modal information to users. Omnidirectional videos (ODVs), as a prominent component of VR content, are essential across diverse applications. This necessitates service providers to monitor and optimize the quality of ODVs throughout the filming, encoding, decoding, and transmission stages to ensure a high-quality viewing experience. However, most existing Quality of Experience (QoE) studies for ODVs only focus on the visual quality, while overlooking the impact of the audio modality on perceptual quality. This paper presents a comprehensive study of omnidirectional audio-visual quality assessment (OD-AVQA) from both subjective and objective perspectives. Specifically, we first establish a large-scale audio-visual quality assessment database for ODVs named OAVQAD+, which includes 625 distorted omnidirectional audio-visual sequences derived from 25 pristine ODVs, and the corresponding collected mean opinion scores (MOSs) for the QoE of these ODVs. This contributes to the largest database for assessing the audio-visual quality of ODVs. To advance the fields of objective OD-AVQA, we construct a benchmark that includes three types of benchmark models. Type I and Type II models integrate well-known video quality assessment (VQA) and audio quality assessment (AQA) methods using support vector regression (SVR) and multi-layer perceptron (MLP), respectively, while Type III consists of AVQA models specifically designed for traditional 2D audio-visual sequences. We also propose a novel Omnidirectional Audio-Visual quality assessment Network (OmniAVNet) that integrates quality-aware audio, visual, and motion features to predict overall audio-visual quality for ODVs effectively, which supports both full-reference (FR) and no-reference (NR) assessment. Extensive experimental results demonstrate that OmniAVNet outperforms the aforementioned benchmark OD-AVQA models on two OD-AVQA databases, and shows great performance on one omnidirectional VQA database. The database and code are available at https://github.com/IntMeGroup/OmniAVNet. Xilei Zhu, Huiyu Duan, Yuqin Cao, Yucheng Zhu, Jing Liu 0002, Xiongkuo Min, Guangtao Zhai, Patrick Le Callet |
IEEE Trans. Image Process. | 8 |
| 2025 | How Does Audio Influence Visual Attention in Omnidirectional Videos? Database and ModelabstractUnderstanding and predicting viewer attention in omnidirectional videos (ODVs) is crucial for enhancing user engagement in virtual and augmented reality applications. Although both audio and visual modalities are essential for saliency prediction in ODVs, the joint exploitation of these two modalities has been limited, primarily due to the absence of large-scale audio-visual saliency databases and comprehensive analyses. This paper comprehensively investigates audio-visual attention in ODVs from both subjective and objective perspectives. Specifically, we first introduce a new audio-visual saliency database for omnidirectional videos, termed AVS-ODV database, containing 162 ODVs and corresponding eye movement data collected from 60 subjects under three audio modes including mute, mono, and ambisonics. Based on the constructed AVS-ODV database, we perform an in-depth analysis of how audio influences visual attention in ODVs. To advance the research on audio-visual saliency prediction for ODVs, we further establish a new benchmark based on the AVS-ODV database by testing numerous state-of-the-art saliency models, including visual-only models and audio-visual models. In addition, given the limitations of current models, we propose an innovative omnidirectional audio-visual saliency prediction network (OmniAVS), which is built based on the U-Net architecture, and hierarchically fuses audio and visual features from the multimodal aligned embedding space. Extensive experimental results demonstrate that the proposed OmniAVS model outperforms other state-of-the-art models on both ODV AVS prediction and traditional AVS prediction tasks. The AVS-ODV database and the OmniAVS model are available at: https://github.com/IntMeGroup/AVS-ODV. Huiyu Duan, Kaiwei Zhang, Yucheng Zhu, Xilei Zhu, Xiongkuo Min, Guangtao Zhai |
IEEE Trans. Image Process. | 8 |
| 2025 | From Haziness to Clarity: A Novel Iterative Memory-Retrospective Emergence Model for Omnidirectional Image Saliency PredictionabstractTo achieve saliency prediction in omnidirectional images (ODIs), the majority of prior works typically adopt the convolutional neural networks (CNNs)-based saliency models to extract semantic features to predict prominent regions in ODIs. Albeit achieving substantially performance gains, these works all employed purely visual computing paradigms and ignore to explore the nature of human visual attention mechanisms. In other words, existing saliency prediction works for ODIs are insufficient to capture the biological characteristics of the visual attention mechanism in the human brain. To establish a more explicit link between saliency prediction performance and brain-like visual attention mechanism, we simulate the mechanism of human retrospective memory in neuropsychology and propose IMRE model, a novel iterative memory-retrospective emergence model can predict and infer the salient features by recalling previously learned information. In IMRE model, we introduce four key modules to simulate the visual attention mechanism for predicting human fixations in the human brain. Firstly, the visual stimulus response module is designed to effectively extract semantic features and capture the intricate relationship between these features, acting as the human visual cortex. Secondly, the retrospective integration module serves to distill valuable information from a fuzzy memory ensemble, resembling the role of the basal ganglia in the neural system. Thirdly, the memory bank module explicitly records and stores subconscious response information and learned knowledge, acting like the hippocampus in neural system. Lastly, the prospective inference module accurately infers saliency maps from the refined useful information, resembling the role of the prefrontal cortex. During prediction, we utilize the introduced memory bank to retrieve and recall previously learned information, which simulates the process of memory emergence from haziness to clarity. Such a process aligns with the retrospective memory mechanism of the human brain. To validate the superiority of the proposed model in ODIs saliency prediction tasks, we conduct extensive experiments on two benchmark datasets. Experiments show impressive performances that IMRE model outperforms other state-of-the-art methods across all benchmark datasets. Importantly, experiments also highlight the IMRE model's ability to trace back to specific instances during prediction, thereby reducing model inference costs and enhancing interpretability. Dandan Zhu 0001, Kaiwei Zhang, Xiongkuo Min, Guangtao Zhai, Xiaokang Yang 0001 |
IEEE Trans. Image Process. | 4 |
| 2025 | Attention and Mamba-Driven Quality Assessment for Underwater ImagesabstractUnderwater imaging is essential in a variety of fields, including resource exploration, marine observation, and scientific research. However, the quality of underwater images is often compromised by environmental factors such as light scattering, absorption, and the presence of fog, leading to distortions such as color shifts, low contrast, and blurriness. To address these challenges, we propose a novel underwater image quality assessment (UIQA) method, the Attention and Mamba-driven Quality Index (AMQI). The AMQI model employs a multi-stage architecture designed to capture both local and global image features critical for underwater quality evaluation. First, a Shallow Feature Extractor (SFE) captures essential spatial details. Next, the Local Information Representation Network (LIR-Net), equipped with Channel Attention (CA) and Large Kernel-guided Spatial (LKS) mechanisms, enhances fine details and captures long-range dependencies to address underwater-specific distortions. The Global Information Representation Network (GIR-Net) further processes the features using a combination of the Visual State-Space Model (VSSM) and ResNet-50 to capture high-level semantic and contextual information. Finally, the Feature-Quality Mapping Network (FQM) converts the learned features into a quality score, ensuring precise predictions of image quality. Extensive experiments on the Underwater Image Quality Database (UIQD) demonstrate that AMQI outperforms current state-of-the-art IQA and UIQA models in terms of accuracy and correlation with human subjective evaluations. The model's robustness and generalization capabilities are further validated through detailed ablation studies and cross-database evaluations, showcasing its strong performance across diverse underwater environments. The source code is available athttps://github.com/ibaochao/AMQI. Jingchao Cao, Baochao Zhang, Yutao Liu 0002, Runze Hu, Ke Gu 0001, Guangtao Zhai, Junyu Dong |
IEEE Trans. Multim. | 6 |
| 2025 | No-Reference Point Cloud Quality Assessment via Graph Convolutional NetworkabstractThree-dimensional (3D) point cloud, as an emerging visual media format, is increasingly favored by consumers as it can provide more realistic visual information than two-dimensional (2D) data. Similar to 2D plane images and videos, point clouds inevitably suffer from quality degradation and information loss through multimedia communication systems. Therefore, automatic point cloud quality assessment (PCQA) is of critical importance. In this work, we propose a novel no-reference PCQA method by using a graph convolutional network (GCN) to characterize the mutual dependencies of multi-view 2D projected image contents. The proposed GCN-based PCQA (GC-PCQA) method contains three modules, i.e., multi-view projection, graph construction, and GCN-based quality prediction. First, multi-view projection is performed on the test point cloud to obtain a set of horizontally and vertically projected images. Then, a perception-consistent graph is constructed based on the spatial relations among different projected images. Finally, reasoning on the constructed graph is performed by GCN to characterize the mutual dependencies and interactions between different projected images, and aggregate feature information of multi-view projected images for final quality prediction. Experimental results on two publicly available benchmark databases show that our proposed GC-PCQA can achieve superior performance than state-of-the-art quality assessment metrics. Qiuping Jiang, Wei Zhou 0021, Feng Shao 0001, Guangtao Zhai, Weisi Lin |
IEEE Trans. Multim. | 5 |
| 2025 | Multi-Dimensional Quality Assessment for Text-to-3D Assets: Dataset and ModelabstractRecent advancements in text-to-image (T2I) generation have spurred the development of text-to-3D asset (T23DA) generation, leveraging pretrained 2D text-to-image diffusion models for text-to-3D asset synthesis. Despite the growing popularity of text-to-3D asset generation, its evaluation has not been well considered and studied. However, given the significant quality discrepancies among various text-to-3D assets, there is a pressing need for quality assessment models aligned with human subjective judgments. To tackle this challenge, we conduct a comprehensive study to explore the T23DA quality assessment (T23DAQA) problem in this work from both subjective and objective perspectives. Given the absence of corresponding databases, we first establish the largest text-to-3D asset quality assessment database to date, termed the AIGC-T23DAQA database. This database encompasses 969 validated 3D assets generated from 170 prompts via 6 popular text-to-3D asset generation models, and corresponding subjective quality ratings for these assets from the perspectives of quality, authenticity, and text-asset correspondence, respectively. Subsequently, we establish a comprehensive benchmark based on the AIGC-T23DAQA database, and devise an effective T23DAQA model to evaluate the generated 3D assets from the aforementioned three perspectives, respectively. Specifically, the proposed method utilizes the projection videos of text-to-3D assets to extract 3D shape, texture and text-asset correspondence features, then fuses them to calculate the final three preference scores respectively. Extensive experimental results demonstrate the effectiveness of the proposed T23DAQA method in evaluating the quality of AI generated 3D asset, which is more consistent with human perception. To the best of our knowledge, this is the first work that studies the problem of text-guided 3D generation quality assessment, and our database and codes will be released to facilitate future research. Huiyu Duan, Xiaohong Liu 0001, Xiongkuo Min, Jia Wang 0004, Guangtao Zhai |
IEEE Trans. Multim. | 7 |
| 2025 | Quality-Guided Skin Tone Enhancement for Portrait PhotographyabstractIn recent years, learning-based color and tone enhancement methods for photos have become increasingly popular. However, most learning-based image enhancement methods just learn a mapping from one distribution to another based on one dataset, lacking the ability to adjust images continuously and controllably. It is important to enable the learning-based enhancement models to adjust an image continuously, since in many cases we may want to get a slighter or stronger enhancement effect rather than one fixed adjusted result. In this paper, we propose a quality-guided image enhancement paradigm that enables image enhancement models to learn the distribution of images with various quality ratings. By learning this distribution, image enhancement models can associate image features with their corresponding perceptual qualities, which can be used to adjust images continuously according to different quality scores. To validate the effectiveness of our proposed method, a subjective quality assessment experiment is first conducted, focusing on skin tone adjustment in portrait photography. Guided by the subjective quality ratings obtained from this experiment, our method can adjust the skin tone corresponding to different quality requirements. Furthermore, an experiment conducted on 10 natural raw images corroborates the effectiveness of our model in situations with fewer subjects and fewer shots, and also demonstrates its general applicability to natural images. Shiqi Gao, Huiyu Duan, Xinyue Li 0001, Yicong Peng, Qihang Xu, Yuanyuan Chang, Jia Wang 0004, Xiongkuo Min, Guangtao Zhai |
IEEE Trans. Multim. | 10 |
| 2025 | Dataset and Metric for Quality Assessment of HDR Tone Mapping: Detail Visibility, Color Naturalness, and Overall QualityabstractTone-Mapping Operators (TMOs) aim at converting high dynamic range (HDR) images into standard dynamic range (SDR) ones that are suitable for being displayed on standard screens. As the visual quality of tone-mapped image (TMI) is paramount, conducting quality assessment of TMIs becomes crucial. Despite the growing body of research on TMI quality assessment, the existing metrics are often limited to a narrow selection of hand-picked examples generated by a restricted range of TMOs. Consequently, their ability of generalizing to the wide array of TMIs encountered in practical scenarios remains unclear. Moreover, the quality degradation in practical TMIs can be intricate, diverse, and complex. To overcome these limitations, we construct so far the largest subjective-annotated TMI quality assessment dataset which comprises a total number of 14,000 TMIs generated by applying 20 representative TMOs to 700 HDR images. The dataset is accompanied by subjective scores that encompass multiple quality dimensions, i.e., TMI quality dataset in terms of Detail visibility, Color naturalness, and overall Quality (TDCQ). In addition, we also design a multi-branch deep neural network tailored to characterize the multi-dimensional quality perception of TMIs, i.e., Color naturalness-, Detail visibility-aware TMI Quality (CDTIQ) metric, allowing for a comprehensive and multifaceted quality assessment of TMIs. Through extensive experiments, we demonstrate the superiority of our proposed metric, showcasing a higher correlation with subjective rating results compared to other relevant no-reference image quality metrics. Qiuping Jiang, Xiwen Li, Zhihua Wang 0002, Guangtao Zhai |
IEEE Trans. Multim. | 5 |
| 2025 | Learning to Generate Realistic Images for Bit-Depth Enhancement via Camera Imaging ProcessingabstractWith the prevalence of advanced displays devices, many attempts have been successfully made in bit-depth enhancement (BDE) to restore the low bit-depth (LBD) images to visually pleasant high bit-depth (HBD) images. However, most methods are still far from satisfactory when addressing real-world LBD images owing to their heavy dependence on LBD-HBD data pairs through direct pixel quantization. Therefore, in this paper, we propose a novel network dubbed RealGAN to generate real-world LBD images by simulating the complex quantization procedure in camera imaging process. Particularly, we design a two-mode differentiable quantization block embedded in the synthesis network facilitating adaptively simulation of the complicated quantization distortions. Furthermore, a simple residual group network is proposed in order to learn the distribution of degradation and non-linear processing in the Image Signal Processing (ISP) pipeline. In the absence of paired HBD and LBD data, the synthesis model is trained end-to-end within the generative adversarial framework using non-paired LBD and HBD images. Finally, we demonstrate that a series of BDE models can benefit from the proposed synthetic dataset and exhibit improved visual quality with sharper edges and finer textures on real-world scenes compared with the original versions trained on directly quantized LBD-HBD pairs. Jing Liu 0002, Huiyu Duan, Yuting Su 0001, Guangtao Zhai |
IEEE Trans. Multim. | 6 |
| 2025 | Deep No-Reference Quality Assessment for Underwater Enhanced ImagesabstractThe goal of underwater image enhancement (UIE) is to boost the acquired underwater image quality, which increases the value of the underwater image significantly. However, without effective underwater enhanced image quality assessment (UEIQA) measures that benchmark the UIE, the process of UIE becomes driftless and the enhanced results of different UIE algorithms cannot be fairly compared. Toward this end, we in this work construct a dedicated UEIQA scheme on the basis of deep investigation of the underwater enhanced image characteristics. Specifically, in our proposed method, we respectively design deep neural networks to represent the unique attributes of the underwater enhanced image, such as color cast, local distortions, naturalness degree, sharpness, contrast, fog density, etc., that are highly correlated with the image quality. Then we introduce the Vision Transformer (ViT) to capture the dependencies among different image attributes and infer the image quality level. Extensive experiments conducted on three typical UEIQA databases, i.e., SOTA, UID2021 and SAUD, show that the proposed UEIQA model yields noteworthy higher prediction accuracy than the representative IQA and UEIQA metrics, e.g., achieving SRCC values of 0.891 ( vs. 0.749 in SAUD) and 0.933 ( vs. 0.798 in UID2021). The proposed UEIQA model will be released athttps://github.com/YT2015?tab=repositories. Yutao Liu 0002, Baochao Zhang, Runze Hu, Ke Gu 0001, Guangtao Zhai, Junyu Dong |
IEEE Trans. Multim. | 5 |
| 2025 | Aggregate and Discriminate: Pseudo Clips-Guided Boundary Perception for Video Moment RetrievalabstractVideo moment retrieval (VMR) aims to localize a video segment in an untrimmed video that is semantically relevant to a language query. The challenge of this task lies in effectively aligning the intricate and information-dense video modality with the succinctly summarized textual modality, and further localizing the starting and ending timestamps of the target moments. Previous works have attempted to achieve multi-granularity alignment of video and query in a coarse-to-fine manner, yet these efforts still fall short in addressing the inherent disparities in representation and information density between videos and queries, leading to modal misalignments. In this paper, we propose a progressive video moment retrieval framework, initially retrieving the most relevant and irrelevant video clips to the query as semantic guidance, thereby bridging the semantic gap between video modality and language modality. Futhermore, we introduce a pseudo clips guided aggregation module to aggregate densely relevant moment clips closer together and propose a discriminative boundary-enhanced decoder with the guidance of pseudo clips to push the semantically confusing proposals away. Extensive experiments on the Charades-STA, ActivityNet Captions and TACoS datasets demonstrate that our method outperforms existing methods. Jing Liu 0002, Zongbing Zhang, Yuting Su 0001, Bing Yang 0003, Xiongkuo Min, Guangtao Zhai |
IEEE Trans. Multim. | 6 |
| 2025 | Human-Centered Financial Signal Analysis Based on Visual Patterns in Stock ChartsabstractThe study adopted a human-centered perspective to research the financial markets, focusing on identifying variations in eye movement patterns between professional and non-professional traders as they analyze a series of stock charts. Eye movement data was selected as the analysis target based on the hypothesis that it represents a behavioral phenotype indicative of stock analysts' cognitive processes during market analysis. Disparities were identified by conducting variance analysis and the Wilcoxon signed-rank test on statistical metrics derived from eye fixations and saccades. Psychological and behavioral economic interpretations were provided to understand the underlying reasons for these observed patterns. To showcase the practical application potential of the human-centered perspective, eye movement data and human visual characteristics were used to construct visual saliency prediction models of professional stock analysts. Leveraging this human-centered model, we developed two practical application demonstrations specifically designed to support and instruct novice traders. Based on the above demonstrations, a training program was designed that demonstrates how, with ongoing training, the non-professional traders' ability to observe stock charts improves progressively. Ji-Feng Luo, Kaixun Zhang, Xudong An, Menghan Hu, Guangtao Zhai, Xiao-Ping Zhang 0002 |
IEEE Trans. Multim. | 6 |
| 2025 | Weakly Supervised Referring Video Object Segmentation With Object-Centric Pseudo-GuidanceabstractReferring video object segmentation (RVOS) is an emerging task for multimodal video comprehension while the expensive annotating process of object masks restricts the scalability and diversity of RVOS datasets. To relax the dependency on expensive mask annotations and take advantage from large-scale partially annotated data, in this paper, we explore a novel extended RVOS task, namely weakly supervised referring video object segmentation (WRVOS), which employs multiple weak supervision sources, including object points and bounding boxes. Correspondingly, we propose a unified WRVOS framework. Specifically, an object-centric pseudo mask generation method is introduced to provide effective shape priors for the pseudo guidance of spatial object location. Then, a pseudo-guided optimization strategy is proposed to effectively optimize the object outlines in terms of spatial location and projection density with a multi-stage online learning strategy. Furthermore, a multimodal cross-frame level set evolution method is proposed to iteratively refine the object boundaries considering both temporal consistency and cross-modal interactions. Extensive experiments are conducted on four publicly available RVOS datasets, including A2D Sentences, J-HMDB Sentences, Ref-DAVIS, and Ref-YoutubeVOS. Performance comparison shows that the proposed method achieves state-of-the-art performance in both point-supervised and box-supervised settings. Weikang Wang 0002, Yuting Su 0001, Jing Liu 0002, Wei Sun 0029, Guangtao Zhai |
IEEE Trans. Multim. | 5 |
| 2025 | Evaluating Point Cloud From Moving Camera Videos: A No-Reference MetricabstractPoint cloud is one of the most widely used digital representation formats for three-dimensional (3D) contents, the visual quality of which may suffer from noise and geometric shift distortions during the production procedure as well as compression and downsampling distortions during the transmission process. To tackle the challenge of point cloud quality assessment (PCQA), many PCQA methods have been proposed to evaluate the visual quality levels of point clouds by assessing the rendered static 2D projections. Although such projectionbased PCQA methods achieve competitive performance with the assistance of mature image quality assessment (IQA) methods, they neglect that the 3D model is also perceived in a dynamic viewing manner, where the viewpoint is continually changed according to the feedback of the rendering device. Therefore, in this paper, we evaluate the point clouds from moving camera videos and explore the way of dealing with PCQA tasks via using video quality assessment (VQA) methods. First, we generate the captured videos by rotating the camera around the point clouds through several circular pathways. Then we extract both spatial and temporal quality-aware features from the selected key frames and the video clips through using trainable 2D-CNN and pretrained 3D-CNN models respectively. Finally, the visual quality of point clouds is represented by the video quality values. The experimental results reveal that the proposed method is effective for predicting the visual quality levels of the point clouds and even competitive with full-reference (FR) PCQA methods. The ablation studies further verify the rationality of the proposed framework and confirm the contributions made by the qualityaware features extracted via the dynamic viewing manner. The code is available athttps://github.com/zzc-1998/VQA_PC. Wei Sun 0029, Yucheng Zhu, Xiongkuo Min, Wei Wu 0002, Ying Chen 0011, Guangtao Zhai |
IEEE Trans. Multim. | 7 |
| 2025 | Optimizing Video-Based Respiration Monitoring: Motion Artifact Reduction and Adaptive ROI SelectionabstractIn non-contact respiratory monitoring, reducing motion artifact and selecting the appropriate Region of Interest (ROI) pose significant challenges. Most motion artifact removal methods rely on signal periodicity assumptions, while respiratory signals usually are non-periodic in real-world scenarios. Existing automated ROI selection approaches are mostly primarily impacted by the texture of clothing, absence of chest landmarks, and obstruction of face. To improve the quality of respiratory signals, in this study, we propose a framework for automatic respiratory ROI selection based on video, namely, Optimizing Video-based Respiration Monitoring (OVRM), which consists of peak-trough adaptive motion artifact removal and characteristic-driven adaptive ROI selection. This motion artifact removal strategy removes motion artifacts by using a dynamic ratio-based judgment mechanism, and reconstructs signals using sinusoidal interpolation. The adaptive ROI method scores signals based on periodicity, similarity, smoothness, and energy, selecting the highest-scoring blocks as the ROIs to match respiratory signals efficiently. Experimental results, validated across four datasets, demonstrate that OVRM effectively reduces signal noise caused by subject movement and outperforms state-of-the-art non-contact respiratory monitoring algorithms. The dataset and code are publicly available at:https://github.com/zxx5058/OVRM. Xudong Tan, Mei Zhou, Menghan Hu, Zhanzhan Cheng, Nengfeng Qian, Changyin Wu, Guangtao Zhai, Xiao-Ping Zhang 0002 |
IEEE Trans. Multim. | 9 |
| 2025 | Explain Vision Focus: Blending Human Saliency Into Synthetic Face ImagesabstractSynthetic faces have been extensively researched and applied in various fields, such as face parsing and recognition. Compared to real face images, synthetic faces engender more controllable and consistent experimental stimuli due to the ability to precisely merge expression animations onto the facial skeleton. Accordingly, we establish an eye-tracking database with 780 synthetic face images and fixation data collected from 22 participants. The use of synthetic images with consistent expressions ensures reliable data support for exploring the database and determining the following findings: (1) A correlation study between saliency intensity and facial movement reveals that the variation of attention distribution within facial regions is mainly attributed to the movement of the mouth. (2) A categorized analysis of different demographic factors demonstrates that the bias towards salient regions aligns with differences in some demographic categories of synthetic characters. In practice, inference of facial saliency distribution is commonly used to predict the regions of interest for facial video-related applications. Therefore, we propose a benchmark model that accurately predicts saliency maps, closely matching the ground truth annotations. This achievement is made possible by utilizing channel alignment and progressive summation for feature fusion, along with the incorporation of Sinusoidal Position Encoding. The ablation experiment also demonstrates the effectiveness of our proposed model. We hope that this paper will contribute to advancing the photorealism of generative digital humans. Kaiwei Zhang, Dandan Zhu 0001, Xiongkuo Min, Huiyu Duan, Guangtao Zhai |
IEEE Trans. Multim. | 5 |
| 2025 | Elevating Mesh Saliency in VR: Introducing a Novel Prediction Network and DatasetabstractIn computer graphics, polygon meshes stand out as a popular representation providing effective delineation of delicate textures and complex geometries. When dealing with geometric processing tasks for critical regions of the mesh, it is necessary to consider the human visual perception related to saliency. Therefore, we establish a novel mesh saliency dataset, facilitated by a more comprehensive gathering pipeline of eye-tracking from subjects observing mesh models at arbitrary viewpoints in a virtual reality space with six degrees of freedom. Additionally, we propose a mesh saliency prediction model that accurately infers visual attention density maps for complex and irregular mesh surfaces. This model integrates surface curvature and triangular face shape information from multi-scale neighboring ranges as local geometric features, while also leveraging surface spatial positioning as a global feature. Our work aims to preserve critical areas and minimize visual loss in saliency-driven tasks such as mesh simplification, rendering, and texturing. We believe that our research can offer valuable insights for human-centered mesh computation applications. Kaiwei Zhang, Mohan He, Dandan Zhu 0001, Kun Zhu 0024, Xiongkuo Min, Guangtao Zhai |
ACM Trans. Multim. Comput. Commun. Appl. | 6 |
| 2025 | Benchmarking Multi-dimensional AIGC Video Quality Assessment: A Dataset and Unified ModelabstractIn recent years, AI-driven video generation has gained significant attention due to great advancements in visual and language generative techniques. Consequently, there is a growing need for accurate Video Quality Assessment (VQA) metrics to evaluate the perceptual quality of AI-generated content (AIGC) videos and optimize video generation models. However, assessing the quality of AIGC videos remains a significant challenge because these videos often exhibit highly complex distortions, such as unnatural actions and irrational objects. To address this challenge, we systematically investigate the AIGC-VQA problem in this article, considering both subjective and objective quality assessment perspectives. For the subjective perspective, we construct the L arge-scale G enerated V ideo Q uality Assessment (LGVQ) dataset, consisting of \(2,\!808\) AIGC videos generated by six video generation models using 468 carefully curated text prompts. Unlike previous subjective VQA experiments, we evaluate the perceptual quality of AIGC videos from three critical dimensions: spatial quality, temporal quality, and text-video alignment, which hold utmost importance for current video generation techniques. For the objective perspective, we establish a benchmark for evaluating existing quality assessment metrics on the LGVQ dataset. Our findings show that current metrics perform poorly on this dataset, highlighting a gap in effective evaluation tools. To bridge this gap, we propose the U nify G enerated V ideo Q uality Assessment (UGVQ) model, designed to accurately evaluate the multi-dimensional quality of AIGC videos. The UGVQ model integrates the visual and motion features of videos with the textual features of their corresponding prompts, forming a unified quality-aware feature representation tailored to AIGC videos. Experimental results demonstrate that UGVQ achieves state-of-the-art performance on the LGVQ dataset across all three quality dimensions, validating its effectiveness as an accurate quality metric for AIGC videos. We hope that our benchmark can promote the development of AIGC-VQA studies. Both the LGVQ dataset and the UGVQ model are publicly available on https://github.com/zczhang-sjtu/UGVQ.git . Wei Sun 0029, Xinyue Li 0001, Jun Jia, Xiongkuo Min, Chunyi Li 0001, Zijian Chen 0001, Puyi Wang, Fengyu Sun, Shangling Jui, Guangtao Zhai |
ACM Trans. Multim. Comput. Commun. Appl. | 12 |
| 2025 | Audio-Visual Saliency Prediction Model with Implicit Neural RepresentationabstractWith the remarkable advancement of deep learning techniques and the wide availability of large-scale datasets, the performance of audio-visual saliency prediction has been drastically improved. Actually, audio-visual saliency prediction is still at an early exploration stage due to the spatial-temporal signal complexity and dynamic continuity of video content. To our knowledge, most existing audio-visual saliency prediction approaches usually represent videos as 3D grid of RGB values using discrete convolutional neural networks (CNNs), which inevitably incurs video content-agnostic and ignores the dynamic continuity issues. This article proposes a novel parametric audio-visual saliency (PAVS) model with implicit neural representation (INR) to address the aforementioned problems. Specifically, by using the proposed parametric neural network, we can effectively encode the space-time coordinates of video frames into corresponding saliency values, which can significantly enhance the compact feature representation ability. Meanwhile, a parametric feature fusion method is developed to achieve intrinsic interactions between audio and visual information streams, which can adaptively fuse audio and visual features to obtain competitive performance. Notably, without resorting to any specific audio-visual feature fusion strategy, the proposed PAVS model outperforms other state-of-the-art saliency methods by a large margin. Dandan Zhu 0001, Kun Zhu 0024, Guangtao Zhai, Xiaokang Yang 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 5 |
| 2025 | MM-PCQA+: Advancing Multi-Modal Learning for Point Cloud Quality AssessmentabstractThe importance of visual quality in point clouds has been significantly underlined due to the rapid rise in 3D vision applications which aim to deliver affordable and superior user experiences. Reviewing the evolution of point cloud quality assessment (PCQA), it’s observed that visual quality evaluation typically employs single-modal data, either sourced from 2D projections or the 3D point clouds. The 2D projections possess abundant texture and semantic information while they are heavily reliant on viewpoints. In contrast, 3D point clouds are more reactive to geometric distortions and viewpoint-invariant. Consequently, to maximize the benefits of both point cloud and image modalities, we present an advanced no-reference Multi-Modal Point Cloud Quality Assessment (MM-PCQA+) metric. Specifically, we divide the point clouds into sub-models to reflect local geometric distortions such as point shifting and down-sampling. Afterwards, we render the point clouds using a cube-like projection setup and sample the projections of interest using a point-visible-ratio for image feature extraction. In order to fulfill these objectives, the sub-models and projected images are encoded using point-based and image-based neural networks. Lastly, we implement symmetric cross-modal attention to amalgamate multi-modal quality-aware features. Experimental results demonstrate that our metric surpasses all state-of-the-art methods and significantly advances beyond previous no-reference PCQA methods. Yingjie Zhou 0003, Chunyi Li 0001, Wei Sun 0029, Xiongkuo Min, Xiaohong Liu 0001, Guangtao Zhai |
ACM Trans. Multim. Comput. Commun. Appl. | 7 |
| 2025 | Quality Assessment in the Era of Large Models: A SurveyabstractQuality assessment, which evaluates the visual quality level of multimedia experiences, has garnered significant attention from researchers and has evolved substantially through dedicated efforts. Before the advent of large models, quality assessment typically relied on small expert models tailored for specific tasks. While these smaller models are effective at handling their designated tasks and predicting quality levels, they often lack explainability and robustness. With the advancement of large models, which align more closely with human cognitive and perceptual processes, many researchers are now leveraging the prior knowledge embedded in these large models for quality assessment tasks. This emergence of quality assessment within the context of large models motivates us to provide a comprehensive review focusing on two key aspects: (1) the assessment of large models and (2) the role of large models in assessment tasks. We begin by reflecting on the historical development of quality assessment. Subsequently, we move to detailed discussions of related works concerning quality assessment in the era of large models. Finally, we offer insights into the future progression and potential pathways for quality assessment in this new era. We hope that this survey will enable a rapid understanding of the development of quality assessment in the era of large models and inspire further advancements in the field. Yingjie Zhou 0003, Chunyi Li 0001, Baixuan Zhao, Xiaohong Liu 0001, Guangtao Zhai |
ACM Trans. Multim. Comput. Commun. Appl. | 6 |
| 2025 | Unified Approach to Mesh Saliency: Evaluating Textured and Non-Textured Meshes Through VR and Multifunctional PredictionabstractMesh saliency aims to empower artificial intelligence with strong adaptability to highlight regions that naturally attract visual attention. Existing advances primarily emphasize the crucial role of geometric shapes in determining mesh saliency, but it remains challenging to flexibly sense the unique visual appeal brought by the realism of complex texture patterns. To investigate the interaction between geometric shapes and texture features in visual perception, we establish a comprehensive mesh saliency dataset, capturing saliency distributions for identical 3D models under both non-textured and textured conditions. Additionally, we propose a unified saliency prediction model applicable to various mesh types, providing valuable insights for both detailed modeling and realistic rendering applications. This model effectively analyzes the geometric structure of the mesh while seamlessly incorporating texture features into the topological framework, ensuring coherence throughout appearance-enhanced modeling. Through extensive theoretical and empirical validation, our approach not only enhances performance across different mesh types, but also demonstrates the model's scalability and generalizability, particularly through cross-validation of various visual features. Kaiwei Zhang, Dandan Zhu 0001, Xiongkuo Min, Guangtao Zhai |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2025 | ESIQA: Perceptual Quality Assessment of Vision-Pro-based Egocentric Spatial ImagesabstractWith the development of eXtended Reality (XR), photo capturing and display technology based on head-mounted displays (HMDs) have experienced significant advancements and gained considerable attention. Egocentric spatial images and videos are emerging as a compelling form of stereoscopic XR content. The assessment for the Quality of Experience (QoE) of XR content is important to ensure a high-quality viewing experience. Different from traditional 2D images, egocentric spatial images present challenges for perceptual quality assessment due to their special shooting, processing methods, and stereoscopic characteristics. However, the corresponding image quality assessment (IQA) research for egocentric spatial images is still lacking. In this paper, we establish the Egocentric Spatial Images Quality Assessment Database (ESIQAD), the first IQA database dedicated for egocentric spatial images as far as we know. Our ESIQAD includes 500 egocentric spatial images and the corresponding mean opinion scores (MOSs) under three display modes, including 2D display, 3D-window display, and 3D-immersive display. Based on our ESIQAD, we propose a novel mamba2-based multi-stage feature fusion model, termed ESIQAnet, which predicts the perceptual quality of egocentric spatial images under the three display modes. Specifically, we first extract features from multiple visual state space duality (VSSD) blocks, then apply cross attention to fuse binocular view information and use transposed attention to further refine the features. The multi-stage features are finally concatenated and fed into a quality regression network to predict the quality score. Extensive experimental results demonstrate that the ESIQAnet outperforms 22 state-of-the-art IQA models on the ESIQAD under all three display modes. The database and code are available at https://github.com/IntMeGroup/ESIQA. Xilei Zhu, Huiyu Duan, Xiongkuo Min, Guangtao Zhai, Patrick Le Callet |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2024 | Q-Instruct: Improving Low-Level Visual Abilities for Multi-Modality Foundation ModelsabstractMulti-modality large language models (MLLMs), as represented by GPT-4V, have introduced a paradigm shift for visual perception and understanding tasks, that a variety of abilities can be achieved within one foundation model. While current MLLMs demonstrate primary low-level visual abilities from the identification of low-level visual attributes (e.g., clarity, brightness) to the evaluation on image quality, there's still an imperative to further improve the accuracy of MLLMs to substantially alleviate human burdens. To address this, we collect the first dataset consisting of human natural language feedback on low-level vision. Each feedback offers a comprehensive description of an image's low-level visual attributes, culminating in an overall quality assessment. The constructed Q-Pathway dataset includes 58K detailed human feedbacks on 18,973 multi-sourced images with diverse low-level appearance. To ensure MLLMs can adeptly handle diverse queries, we further propose a GPT-participated transformation to convert these feedbacks into a rich set of 200K instruction-response pairs, termed Q-Instruct. Experimental results indicate that the Q-Instruct consistently elevates various low-level visual capabilities across multiple base models. We anticipate that our datasets can pave the way for a future that foundation models can assist humans on low-level visual tasks. Haoning Wu 0001, Erli Zhang 0001, Chaofeng Chen, Annan Wang, Kaixin Xu, Chunyi Li 0001, Jingwen Hou, Guangtao Zhai, Geng Xue, Wenxiu Sun, Qiong Yan, Weisi Lin |
CVPR | 10 |
| 2024 | Text2QR: Harmonizing Aesthetic Customization and Scanning Robustness for Text-Guided QR Code GenerationabstractIn the digital era, QR codes serve as a linchpin connecting virtual and physical realms. Their pervasive integration across various applications highlights the demand for aesthetically pleasing codes without compromised scannability. However, prevailing methods grapple with the intrinsic challenge of balancing customization and scannability. Notably, stable-diffusion models have ushered in an epoch of high-quality, customizable content generation. This paper introduces Text2QR, a pioneering approach leveraging these advancements to address a fundamental challenge: concurrently achieving user-defined aesthetics and scanning robustness. To ensure stable generation of aesthetic QR codes, we introduce the QR Aesthetic Blueprint (QAB) module, generating a blueprint image exerting control over the entire generation process. Subsequently, the Scannability Enhancing Latent Refinement (SELR) process refines the output iteratively in the latent space, enhancing scanning robustness. This approach harnesses the potent generation capabilities of stable-diffusion models, navigating the trade-off between image aesthetics and QR code scannability. Our experiments demonstrate the seamless fusion of visual appeal with the practical utility of aesthetic QR codes, markedly outperforming prior methods. Codes are available at https://github.com/mulns/Text2QR Guangyang Wu, Xiaohong Liu 0001, Jun Jia, Xuehao Cui, Guangtao Zhai |
CVPR | 5 |
| 2024 | UniProcessor: A Text-Induced Unified Low-Level Image Processor
Huiyu Duan, Xiongkuo Min, Sijing Wu, Wei Shen 0002, Guangtao Zhai |
ECCV (67) | 5 |
| 2024 | Free-VSC: Free Semantics from Visual Foundation Models for Unsupervised Video Semantic Compression
Yuan Tian 0017, Guo Lu, Guangtao Zhai |
ECCV (49) | 3 |
| 2024 | Towards Open-Ended Visual Quality Comparison
Haoning Wu 0001, Hanwei Zhu, Erli Zhang 0001, Chaofeng Chen, Chunyi Li 0001, Annan Wang, Wenxiu Sun, Qiong Yan, Xiaohong Liu 0001, Guangtao Zhai, Shiqi Wang 0001, Weisi Lin |
ECCV (3) | 12 |
| 2024 | GLARE: Low Light Image Enhancement via Generative Latent Feature Based Codebook Retrieval
Han Zhou 0003, Wei Dong 0011, Xiaohong Liu 0001, Shuaicheng Liu, Xiongkuo Min, Guangtao Zhai, Jun Chen 0005 |
ECCV (48) | 6 |
| 2024 | AttentionLUT: Attention Fusion-Based Canonical Polyadic LUT for Real-Time Image EnhancementabstractRecently, many algorithms have employed image-adaptive lookup tables (LUTs) to achieve real-time image enhancement. Nonetheless, a prevailing trend among existing methods has been the employment of linear combinations of basic LUTs to formulate image-adaptive LUTs, which limits the generalization ability of these methods. To address this limitation, we propose a novel framework named AttentionLut for real-time image enhancement, which utilizes the attention mechanism to generate image-adaptive LUTs. Our proposed framework consists of three lightweight modules. We begin by employing the global image context feature module to extract image-adaptive features. Subsequently, the attention fusion module integrates the image feature with the priori attention feature obtained during training to generate image-adaptive canonical polyadic tensors. Finally, the canonical polyadic reconstruction module is deployed to reconstruct image-adaptive residual 3DLUT, which is subsequently utilized for enhancing input images. Experiments on the benchmark MIT-Adobe FiveK dataset demonstrate that the proposed method achieves better enhancement performance quantitatively and qualitatively than the state-of-the-art methods. Yicong Peng, Qihang Xu, Xiaohong Liu 0001, Jia Wang 0004, Guangtao Zhai |
ICASSP | 7 |
| 2024 | A Reduced-Reference Quality Assessment Metric for Textured Mesh Digital HumansabstractIn an era where 3D Digital Humans (DHs) are becoming increasingly prevalent in fields like gaming, automotive, and the metaverse, the demand for high DH visual quality is rising. This paper presents the first-ever reduced-reference (RR) quality assessment metric tailored specifically for textured mesh DHs, aiming to optimize transmission systems and improve Quality of Experience (QoE) for viewers in resource-constrained environments. Four critical geometric curvature-related attributes and two texture-related indicators are computed, which are then statistically analyzed and utilized in a Support Vector Regression (SVR) model for robust and efficient quality prediction. Experimental results confirm that our method outperforms existing full-reference (FR) metrics, making it an invaluable tool for the future of 3D DHs in various applications. The code is available at https://github.com/zzc-1998/RR-DHQA. Yingjie Zhou 0003, Chunyi Li 0001, Wei Sun 0029, Xiaohong Liu 0001, Xiongkuo Min, Guangtao Zhai |
ICASSP | 8 |
| 2024 | SG-JND: Semantic-Guided Just Noticeable Distortion Predictor for Image CompressionabstractJust noticeable distortion (JND), representing the threshold of distortion in an image that is minimally perceptible to the human visual system (HVS), is crucial for image compression algorithms to achieve a trade-off between transmission bit rate and image quality. However, traditional JND prediction methods only rely on pixel-level or sub-band level features, lacking the ability to capture the impact of image content on JND. To bridge this gap, we propose a Semantic-Guided JND (SG-JND) network to leverage semantic information for JND prediction. In particular, SG-JND consists of three essential modules: the image preprocessing module extracts semantic-level patches from images, the feature extraction module extracts multi-layer features by utilizing the cross-scale attention layers, and the JND prediction module regresses the extracted features into the final JND value. Experimental results show that SG-JND achieves the state-of-the-art performance on two publicly available JND datasets, which demonstrates the effectiveness of SG-JND and highlight the significance of incorporating semantic information in JND assessment. Linhan Cao, Wei Sun 0029, Xiongkuo Min, Jun Jia, Zijian Chen 0001, Yucheng Zhu, Lizhou Liu, Qiubo Chen, Guangtao Zhai |
ICIP | 11 |
| 2024 | DTSN: No-Reference Image Quality Assessment via Deformable Transformer and Semantic NetworkabstractFeature maps with varying resolutions usually serve different functions. High-resolution feature maps contain abundant texture and color information. Low-resolution feature maps provide significant semantic information. All of this information is crucial for Image Quality Assessment (IQA). Hence, it is challenging to evaluate an image’s quality using only one type of feature map. In this paper, a No-reference Image Quality Assessment (NR-IQA) based on Deformable Transformer and semantic network (DTSN) is proposed. DTSN efficiently learns distortion information across different feature layers in images. Simultaneously, it utilizes semantic information in images to identify which parts of an image significantly impact the image quality score. Experimental results indicate that excellent results can be achieved with only a few sampling points. Guoquan Zheng, Zesheng Wang 0004, Guangtao Zhai |
ICIP | 5 |
| 2024 | AIGCOIQA2024: Perceptual Quality Assessment of AI Generated Omnidirectional Imagesabstract[?]In recent years, the rapid advancement of Artificial Intelligence Generated Content (AIGC) has attracted widespread attention. Among the AIGC, AI generated omnidirectional images hold significant potential for Virtual Reality (VR) and Augmented Reality (AR) applications, hence omnidirectional AIGC techniques have also been widely studied. AI-generated omnidirectional images exhibit unique distortions compared to natural omnidirectional images, however, there is no dedicated Image Quality Assessment (IQA) criteria for assessing them. This study addresses this gap by establishing a large-scale AI generated omnidirectional image IQA database named AIGCOIQA2024 and constructing a comprehensive benchmark. We first generate 300 omnidirectional images based on 5 AIGC models utilizing 25 text prompts. A subjective IQA experiment is conducted subsequently to assess human visual preferences from three perspectives including quality, comfortability, and correspondence. Finally, we conduct a benchmark experiment to evaluate the performance of state-of-the-art IQA models on our database. The AIGCOIQA2024 database is released to facilitate future research on https://github.com/IntMeGroup/AIGCOIQA. Huiyu Duan, Yucheng Zhu, Xiaohong Liu 0001, Menghan Hu, Xiongkuo Min, Guangtao Zhai, Patrick Le Callet |
ICIP | 8 |
| 2024 | A Comparative Study of Perceptual Quality Metrics For Audio-Driven Talking Head VideosabstractThe rapid advancement of Artificial Intelligence Generated Content (AIGC) technology has propelled audio-driven talking head generation, gaining considerable research attention for practical applications. However, performance evaluation research lags behind the development of talking head generation techniques. Existing literature relies on heuristic quantitative metrics without human validation, hindering accurate progress assessment. To address this gap, we collect talking head videos generated from four generative methods and conduct controlled psychophysical experiments on visual quality, lip-audio synchronization, and head movement naturalness. Our experiments validate consistency between model predictions and human annotations, identifying metrics that align better with human opinions than widely-used measures. We believe our work will facilitate performance evaluation and model development, providing insights into AIGC in a broader context. Code is available at https://github.com/zwx8981/ADTH-QA. Weixia Zhang, Chengguang Zhu, Jingnan Gao, Yichao Yan, Guangtao Zhai, Xiaokang Yang 0001 |
ICIP | 5 |
| 2024 | Thqa: A Perceptual Quality Assessment Database for Talking HeadsabstractIn the realm of media technology, digital humans have gained prominence due to rapid advancements in computer technology. However, the manual modeling and control required for the majority of digital humans pose significant obstacles to efficient development. The speech-driven methods offer a novel avenue for manipulating the mouth shape and expressions of digital humans. Despite the proliferation of driving methods, the quality of many generated talking head (TH) videos remains a concern, impacting user visual experiences. To tackle this issue, this paper introduces the Talking Head Quality Assessment (THQA) database, featuring 800 TH videos generated through 8 diverse speechdriven methods. Extensive experiments affirm the THQA database’s richness in character and speech features. Subsequent subjective quality assessment experiments analyze correlations between scoring results and speech-driven methods, ages, and genders. In addition, experimental results show that mainstream image and video quality assessment methods have limitations for the THQA database, underscoring the imperative for further research to enhance TH video quality assessment. The THQA database is publicly accessible at https://github.com/zyj-2000/THQA. Yingjie Zhou 0003, Wei Sun 0029, Xiaohong Liu 0001, Xiongkuo Min, Zhihua Wang 0002, Xiao-Ping Zhang 0002, Guangtao Zhai |
ICIP | 8 |
| 2024 | Q-Bench: A Benchmark for General-Purpose Foundation Models on Low-level VisionabstractThe rapid evolution of Multi-modality Large Language Models (MLLMs) has catalyzed a shift in computer vision from specialized models to general-purpose foundation models. Nevertheless, there is still an inadequacy in assessing the abilities of MLLMs on **low-level visual perception and understanding**. To address this gap, we present **Q-Bench**, a holistic benchmark crafted to systematically evaluate potential abilities of MLLMs on three realms: low-level visual perception, low-level visual description, and overall visual quality assessment. **_a)_** To evaluate the low-level **_perception_** ability, we construct the **LLVisionQA** dataset, consisting of 2,990 diverse-sourced images, each equipped with a human-asked question focusing on its low-level attributes. We then measure the correctness of MLLMs on answering these questions. **_b)_** To examine the **_description_** ability of MLLMs on low-level information, we propose the **LLDescribe** dataset consisting of long expert-labelled *golden* low-level text descriptions on 499 images, and a GPT-involved comparison pipeline between outputs of MLLMs and the *golden* descriptions. **_c)_** Besides these two tasks, we further measure their visual quality **_assessment_** ability to align with human opinion scores. Specifically, we design a softmax-based strategy that enables MLLMs to predict *quantifiable* quality scores, and evaluate them on various existing image quality assessment (IQA) datasets. Our evaluation across the three abilities confirms that MLLMs possess preliminary low-level visual skills. However, these skills are still unstable and relatively imprecise, indicating the need for specific enhancements on MLLMs towards these abilities. We hope that our benchmark can encourage the research community to delve deeper to discover and enhance these untapped potentials of MLLMs. Haoning Wu 0001, Erli Zhang 0001, Chaofeng Chen, Annan Wang, Chunyi Li 0001, Wenxiu Sun, Qiong Yan, Guangtao Zhai, Weisi Lin |
ICLR | 10 |
| 2024 | Q-Refine: A Perceptual Quality Refiner for AI-Generated ImageabstractWith the rapid evolution of the Text-to-Image (T2I) model in recent years, their unsatisfactory generation result has become a challenge. However, uniformly refining AI-Generated Images (AIGIs) of different qualities not only limited optimization capabilities for low-quality AIGIs but also brought negative optimization to high-quality AIGIs. To address this issue, a quality-award refiner named Q-Refine is proposed. Based on the preference of the Human Visual System (HVS), Q-Refine uses the Image Quality Assessment (IQA) metric to guide the refining process for the first time, and modify images of different qualities through three adaptive pipelines. Experimental data shows that for mainstream T2I models, Q-Refine can perform effective optimization to AIGIs of different qualities. It can be a general refiner to optimize AIGIs from both fidelity and aesthetic quality levels, thus expanding the application of the T2I generation models. The code is released on https://github.com/Q-Future/Q-Refine. Chunyi Li 0001, Haoning Wu 0001, Hongkun Hao, Kaiwei Zhang, Lei Bai 0001, Xiaohong Liu 0001, Xiongkuo Min, Weisi Lin, Guangtao Zhai |
ICME | 10 |
| 2024 | Optimizing Projection-Based Point Cloud Quality Assessment with Human Preferred Viewpoints SelectionabstractViewpoint selection plays a pivotal role in projection-based point cloud quality assessment (PCQA). Generally speaking, sole reliance on a single projection fails to capture adequate quality information, leading to the prevalent use of multi-projection approaches. It is important to recognize that viewpoint selection is significantly influenced by human preferences and viewpoints that align with human predilections exert a greater impact on PCQA. Therefore, we introduce the first viewpoint selection database for PCQA, which comprises 405 distorted point clouds, accompanied by preferred viewpoints collected from humans. Then we propose a novel human preference index, devised from the Visible-Points Ratio and Visible-Color-Entropy Ratio, to guide the selection of viewpoints. Our experimental findings confirm that this human preference index correlates more closely with human preferences than traditional viewpoint selection settings. Moreover, the proposed PCQA method optimized with the human preference index demonstrates competitive performance as well. Wei Sun 0029, Xiongkuo Min, Xiaohong Liu 0001, Chunyi Li 0001, Haoning Wu 0001, Weisi Lin, Guangtao Zhai |
ICME | 10 |
| 2024 | Q-Align: Teaching LMMs for Visual Scoring via Discrete Text-Defined LevelsabstractThe explosion of visual content available online underscores the requirement for an accurate machine assessor to robustly evaluate scores across diverse types of visual contents. While recent studies have demonstrated the exceptional potentials of large multi-modality models (LMMs) on a wide range of related fields, in this work, we explore how to teach them for visual rating aligning with human opinions. Observing that human raters only learn and judge discrete text-defined levels in subjective studies, we propose to emulate this subjective process and teach LMMs with text-defined rating levels instead of scores. The proposed Q-Align achieves state-of-the-art accuracy on image quality assessment (IQA), image aesthetic assessment (IAA), as well as video quality assessment (VQA) under the original LMM structure. With the syllabus, we further unify the three tasks into one model, termed the OneAlign. Our experiments demonstrate the advantage of discrete levels over direct scores on training, and that LMMs can learn beyond the discrete levels and provide effective finer-grained evaluations. Code and weights will be released. Haoning Wu 0001, Weixia Zhang, Chaofeng Chen, Chunyi Li 0001, Annan Wang, Erli Zhang 0001, Wenxiu Sun, Qiong Yan, Xiongkuo Min, Guangtao Zhai, Weisi Lin |
ICML | 13 |
| 2024 | E2DAS: An Efficient Equivariant Dynamic Aggregation Saliency Model for Omnidirectional Images
Dandan Zhu 0001, Kun Zhu 0024, Guangtao Zhai, Xiaokang Yang 0001 |
ICPR (3) | 5 |
| 2024 | TDiffSal: Text-Guided Diffusion Saliency Prediction Model for Images
Dandan Zhu 0001, Kun Zhu 0024, Guangtao Zhai |
ICPR (8) | 5 |
| 2024 | DiffStega: Towards Universal Training-Free Coverless Image Steganography with Diffusion Models
Yiwei Yang 0007, Zheyuan Liu 0011, Jun Jia, Zhongpai Gao, Wei Sun 0029, Xiaohong Liu 0001, Guangtao Zhai |
IJCAI | 8 |
| 2024 | FS-BAND: A Frequency-Sensitive Banding DetectorabstractBanding artifact, as known as staircase-like contour, is a common quality annoyance that happens in compression, transmission, etc. scenarios, which largely affects the user’s quality of experience (QoE). The banding distortion typically appears as relatively small pixel-wise variations in smooth backgrounds, which is difficult to analyze in the spatial domain but easily reflected in the frequency domain. In this paper, we thereby study the banding artifact from the frequency aspect and propose a no-reference banding detection model to capture and evaluate banding artifacts, called the Frequency-Sensitive BANding Detector (FS-BAND). The proposed detector is able to generate a pixel-wise banding map with a perception correlated quality score. Experimental results show that the proposed FS-BAND method outperforms state-of-the-art image quality assessment (IQA) approaches with higher accuracy in banding classification task. Zijian Chen 0001, Wei Sun 0029, Ru Huang 0002, Fangfang Lu, Xiongkuo Min, Guangtao Zhai, Wenjun Zhang 0005 |
ISCAS | 7 |
| 2024 | Calculating Color Differences of Images via Siamese Neural NetworkabstractRecently, the color difference (CD) of standard dynamic range (SDR) images has attracted the attention of researchers. It is worth noting that due to the development of high dynamic range (HDR) image generation technology, the CD of the SDR and HDR images is also worth in-depth research. This is because the HDR image generated from an original SDR image may have changes in color. Some color changes can give people a comfortable impression, but this may also change the information originally expressed in the SDR image. Therefore, this paper researches the CD of the original SDR image and the generated HDR image, and proposes a network to predict the CDs of SDR and HDR image pairs. Specifically, we first build a SDR-HDR image CD dataset. The dataset contains 504 SDR and HDR image pairs, where HDR images are generated from the SDR images using five HDR image generation methods. Second, we propose a siamese neural network to predict the CDs of SDR and HDR image pairs, which consists of three parts: space conversion, feature extraction, and CD calculation. Finally, experiments prove that the proposed network has a superior ability to predict the CDs of SDR and HDR image pairs. Xiongkuo Min, Xiaohong Liu 0001, Lei Sun 0009, Yonglin Luo, Zuowei Cao, Guangtao Zhai |
ISCAS | 7 |
| 2024 | PrefIQA: Human Preference Learning for AI-generated Image Quality AssessmentabstractDespite recent advancements in generative models, the variation in image quality remains a significant concern. To tackle this issue, we propose PrefIQA, an effective human preference learning metric, which can better evaluate the quality of AI-generated images. PrefIQA consists of two units, namely Feature Extraction Unit and Feature Fusion Unit. In Feature Extraction Unit, we introduce a prompt-segmentation module to divide prompts into multiple phrases, enabling a more detailed evaluation of the alignment between images and texts. In Feature Fusion Unit, we introduce a modality-fusion module, which effectively mixes text features and image features to improve the overall performance. In the experiment part, extensive experiments are conducted, demonstrating that PrefIQA surpasses existing text-to-image alignment metrics. We believe that PrefIQA’s proposal would facilitate researches on AI-generated image quality assessment, and make a valuable contribution to the field of text-to-image generation. Hengjian Gao, Kaiwei Zhang, Wei Sun 0029, Chunyi Li 0001, Huiyu Duan, Xiaohong Liu 0001, Xiongkuo Min, Guangtao Zhai |
ISCAS | 8 |
| 2024 | Multidimensional Similarity Fusion for Speech Quality AssessmentabstractSince the perceptual quality of audio signals is easily to be affected by compression, transmission, noise adding, etc, it is of great significance to develop an effective audio quality assessment (AQA) method to measure end-user’s quality of experience. In this paper, we propose a full reference AQA model named Multidimensional Similarity Fusion for Audio Quality Assessment (MSF-AQA). We generalize the similarity-based image quality assessment methods for audio, then extract audio similarity features from multiple dimensions, and finally regress the multidimensional similarity features into the final quality score. The experimental results across three databases indicate that our MSF-AQA model outperforms the state-of-the-art AQA methods. Xiongkuo Min, Yuqin Cao, Xiao-Ping Zhang 0002, Guangtao Zhai |
ISCAS | 5 |
| 2024 | DSA-QoE: Quality of Experience Evaluation for Streaming Video Based on Dual-Stage AttentionabstractWith the rapid development of streaming media technology, the real-time streaming video Quality of Experience (QoE) assessment has become an important objective for creating new Adaptive Bitrate (ABR) algorithms. The QoE prediction on the client side is challenging considering the sophisticated perception mechanisms of humans, especially the human attention behaviors over time. To address this issue, we propose a learnable model based on the dual-stage attention mechanism to precisely predict continuous QoE which is not covered by most of the current related works. Given the close relationship between the continuous and overall QoE, we use a unified framework to predict these two indices. We have conducted comparison experiments on 6 open datasets, and our model shows superior performance. Ziheng Jia, Xiongkuo Min, Guangtao Zhai |
ISCAS | 3 |
| 2024 | PAPS-OVQA: Projection-Aware Patch Sampling for Omnidirectional Video Quality AssessmentabstractIn immersive multimedia systems, the perceptual quality model of omnidirectional video is indispensable. However, to cope with its resolution that is several times higher than ordinary video, the existing omnidirectional video quality assessment (OVQA) models require extremely high computational complexity and usually need to transcode the projection into a certain format. Therefore, to assess the perceptual quality of omnidirectional video effectively, we propose Projection-Aware Patch Sampling (PAPS)-OVQA to process its three common projection formats simultaneously while resizing high-resolution video into patches sampled from uniform grids and finally apply Fragment Attention Network (FANet) to perform quality regression. As a result, we avoid the overhead computational cost of projection transcoding and reduce the complexity of the quality model greatly. Experimental data show that PAPS-OVQA guarantees good performance while retaining high efficiency under different projection formats. Chunyi Li 0001, Haoning Wu 0001, Kaiwei Zhang, Lei Bai 0001, Xiaohong Liu 0001, Guangtao Zhai, Weisi Lin |
ISCAS | 7 |
| 2024 | T2I-Scorer: Quantitative Evaluation on Text-to-Image Generation via Fine-Tuned Large Multi-Modal ModelsabstractText-to-image (T2I) generation is a pivotal and core interest within the realm of AI content generation. Amid the swift advancements of both open-source (such as Stable Diffusion) and proprietary (for example, DALLE, MidJourney) T2I models, there is a notable absence of a comprehensive and robust quantitative framework for evaluating their output quality. Traditional methods of quality assessment overlook the textual prompts when judging images; meanwhile, the advent of large multi-modal models (LMMs) introduces the capability to incorporate text prompts in evaluations, yet the challenge of fine-tuning these models for precise T2I quality assessment remains unresolved. In our study, we introduce the T2I-Scorer, a novel two-stage training methodology aimed at fine-tuning LMMs for T2I evaluation. For the first stage, we collect 397K GPT-4V-labeled question-answer pairs related to T2I evaluation. Termed as T2I-ITD, the pseudo-labeled dataset is analyzed and examined by human, and used for instruction tuning to improve the LMM's low-level quality perception. The first stage model, T2I-Scorer-IT, has reached superior accuracy on T2I evaluation than all kinds of existing T2I metrics under zero-shot settings. For the second stage, we define an explicit multi-task training scheme to further align the LMM with human opinion scores, and the fine-tuned T2I-Scorer can reach state-of-the-art accuracy on both image quality and image-text alignment perspectives with significant improvements. We anticipate the proposed metrics can serve as a reliable metric to gauge the ability of T2I generation models in the future. We will make code, data, and weights publicly available. Haoning Wu 0001, Xiele Wu, Chunyi Li 0001, Chaofeng Chen, Xiaohong Liu 0001, Guangtao Zhai, Weisi Lin |
ACM Multimedia | 7 |
| 2024 | Subjective-Aligned Dataset and Metric for Text-to-Video Quality AssessmentabstractWith the rapid development of generative models, AI-Generated Content (AIGC) has exponentially increased in daily lives. Among them, Text-to-Video (T2V) generation has received widespread attention. Though many T2V models have been released for generating high perceptual quality videos, there is still lack of a method to evaluate the quality of these videos quantitatively. To solve this issue, we establish the largest-scale Text-to-Video Quality Assessment DataBase (T2VQA-DB) to date. The dataset is composed of 10,000 videos generated by 9 different T2V models, along with each video's corresponding mean opinion score. Based on T2VQA-DB, we propose a novel transformer-based model for subjective-aligned Text-to-Video Quality Assessment (T2VQA). The model extracts features from text-video alignment and video fidelity perspectives, then it leverages the ability of a large language model to give the prediction score. Experimental results show that T2VQA outperforms existing T2V metrics and SOTA video quality assessment models. Quantitative analysis indicates that T2VQA is capable of giving subjective-align predictions, validating its effectiveness. The dataset and code are available at https://github.com/QMME/T2VQA. Tengchuan Kou, Xiaohong Liu 0001, Chunyi Li 0001, Haoning Wu 0001, Xiongkuo Min, Guangtao Zhai |
ACM Multimedia | 7 |
| 2024 | G-Refine: A General Quality Refiner for Text-to-Image Generation
Chunyi Li 0001, Haoning Wu 0001, Hongkun Hao, Tengchuan Kou, Chaofeng Chen, Lei Bai 0001, Xiaohong Liu 0001, Weisi Lin, Guangtao Zhai |
ACM Multimedia | 10 |
| 2024 | Large Multi-modality Model Assisted AI-Generated Image Quality AssessmentabstractTraditional deep neural network (DNN)-based image quality assessment (IQA) models leverage convolutional neural networks (CNN) or Transformer to learn the quality-aware feature representation, achieving commendable performance on natural scene images. However, when applied to AI-Generated images (AGIs), these DNN-based IQA models exhibit subpar performance. This situation is largely due to the semantic inaccuracies inherent in certain AGIs caused by uncontrollable nature of the generation process. Thus, the capability to discern semantic content becomes crucial for assessing the quality of AGIs. Traditional DNN-based IQA models, constrained by limited parameter complexity and training data, struggle to capture complex fine-grained semantic features, making it challenging to grasp the existence and coherence of semantic content of the entire image. To address the shortfall in semantic content perception of current IQA models, we introduce a large Multi-modality model Assisted AI-Generated Image Quality Assessment (MA-AGIQA) model, which utilizes semantically informed guidance to sense semantic information and extract semantic vectors through carefully designed text prompts. Moreover, it employs a mixture of experts (MoE) structure to dynamically integrate the semantic information with the quality-aware features extracted by traditional DNN-based IQA models. Comprehensive experiments conducted on two AI-generated content datasets and two traditional IQA datasets show that MA-AGIQA achieves state-of-the-art performance, and demonstrate its superior generalization capabilities on assessing the quality of AGIs. The code is available at https://github.com/wangpuyi/MA-AGIQA. Puyi Wang, Wei Sun 0029, Jun Jia, Yanwei Jiang, Xiongkuo Min, Guangtao Zhai |
ACM Multimedia | 8 |
| 2024 | MMHead: Towards Fine-grained Multi-modal 3D Facial Animationabstract3D facial animation has attracted considerable attention due to its extensive applications in the multimedia field. Audio-driven 3D facial animation has been widely explored with promising results. However, multi-modal 3D facial animation, especially text-guided 3D facial animation is rarely explored due to the lack of multi-modal 3D facial animation dataset. To fill this gap, we first construct a large-scale multi-modal 3D facial animation dataset, MMHead, which consists of 49 hours of 3D facial motion sequences, speech audios, and rich hierarchical text annotations. Each text annotation contains abstract action and emotion descriptions, fine-grained facial and head movements (i.e., expression and head pose) descriptions, and three possible scenarios that may cause such emotion. Concretely, we integrate five public 2D portrait video datasets, and propose an automatic pipeline to 1) reconstruct 3D facial motion sequences from monocular videos; and 2) obtain hierarchical text annotations with the help of AU detection and ChatGPT. Based on the MMHead dataset, we establish benchmarks for two new tasks: text-induced 3D talking head animation and text-to-3D facial motion generation. Moreover, a simple but efficient VQ-VAE-based method named MM2Face is proposed to unify the multi-modal information and generate diverse and plausible 3D facial motions, which achieves competitive results on both benchmarks. Extensive experiments and comprehensive analysis demonstrate the significant potential of our dataset and benchmarks in promoting the development of multi-modal 3D facial animation. The dataset will be released at: https://wsj-sjtu.github.io/MMHead/. Sijing Wu, Yichao Yan, Huiyu Duan, Ziwei Liu 0002, Guangtao Zhai |
ACM Multimedia | 6 |
| 2024 | LMM-PCQA: Assisting Point Cloud Quality Assessment with LMMabstractAlthough large multi-modality models (LMMs) have seen extensive exploration and application in various quality assessment studies, their integration into Point Cloud Quality Assessment (PCQA) remains unexplored. Given LMMs' exceptional performance and robustness in low-level vision and quality assessment tasks, this study aims to investigate the feasibility of imparting PCQA knowledge to LMMs through text supervision. To achieve this, we transform quality labels into textual descriptions during the fine-tuning phase, enabling LMMs to derive quality rating logits from 2D projections of point clouds. To compensate for the loss of perception in the 3D domain, structural features are extracted as well. These quality logits and structural features are then combined and regressed into quality scores. Our experimental results affirm the effectiveness of our approach, showcasing a novel integration of LMMs into PCQA that enhances model understanding and assessment accuracy. We hope our contributions can inspire subsequent investigations into the fusion of LMMs with PCQA, fostering advancements in 3D visual quality analysis and beyond. The code is available at https://github.com/zzc-1998/LMM-PCQA. Haoning Wu 0001, Yingjie Zhou 0003, Chunyi Li 0001, Wei Sun 0029, Chaofeng Chen, Xiongkuo Min, Xiaohong Liu 0001, Weisi Lin, Guangtao Zhai |
ACM Multimedia | 10 |
| 2024 | Subjective and Objective Quality-of-Experience Assessment for 3D Talking HeadsabstractIn recent years, immersive communication has emerged as a compelling alternative to traditional video communication methods. One prospective avenue for immersive communication involves augmenting the user's immersive experience through the transmission of three-dimensional (3D) talking heads (THs). However, transmitting 3D THs poses significant challenges due to its complex and voluminous nature, often leading to pronounced distortion and a compromised user experience. Addressing this challenge, we introduce the 3D Talking Heads Quality Assessment (THQA-3D) dataset, comprising 1,000 sets of distorted and 50 original TH mesh sequences (MSs), to facilitate quality assessment in 3D TH transmission. A subjective experiment, characterized by a novel interactive approach, is conducted with recruited participants to assess the quality of MSs in THQA-3D dataset. Leveraging this dataset, we also propose a multimodal Quality-of-Experience (QoE) method incorporating a Large Quality Model (LQM). This method involves frontal projection of MSs and subsequent rendering into videos, with quality assessment facilitated by the LQM and a variable-length video memory filter (VVMF). Additionally, tone-lip coherence and silence detection techniques are employed to characterize audio-visual coherence in 3D MS streams. Experimental evaluation demonstrates the proposed method's superiority, achieving state-of-the-art performance on the THQA-3D dataset and competitiveness on other QoE datasets. Both the THQA-3D dataset and the QoE model have been publicly released at https://github.com/zyj-2000/THQA-3D Yingjie Zhou 0003, Wei Sun 0029, Xiaohong Liu 0001, Xiongkuo Min, Guangtao Zhai |
ACM Multimedia | 6 |
| 2024 | GAIA: Rethinking Action Quality Assessment for AI-Generated VideosabstractAssessing action quality is both imperative and challenging due to its significant impact on the quality of AI-generated videos, further complicated by the inherently ambiguous nature of actions within AI-generated video (AIGV). Current action quality assessment (AQA) algorithms predominantly focus on actions from real specific scenarios and are pre-trained with normative action features, thus rendering them inapplicable in AIGVs. To address these problems, we construct GAIA, a Generic AI-generated Action dataset, by conducting a large-scale subjective evaluation from a novel causal reasoning-based perspective, resulting in 971,244 ratings among 9,180 video-action pairs. Based on GAIA, we evaluate a suite of popular text-to-video (T2V) models on their ability to generate visually rational actions, revealing their pros and cons on different categories of actions. We also extend GAIA as a testbed to benchmark the AQA capacity of existing automatic evaluation methods. Results show that traditional AQA methods, action-related metrics in recent T2V benchmarks, and mainstream video quality methods perform poorly with an average SRCC of 0.454, 0.191, and 0.519, respectively, indicating a sizable gap between current models and human action perception patterns in AIGVs. Our findings underscore the significance of action quality as a unique perspective for studying AIGVs and can catalyze progress towards methods with enhanced capacities for AQA in AIGVs. Zijian Chen 0001, Wei Sun 0029, Yuan Tian 0017, Jun Jia, Ru Huang 0002, Xiongkuo Min, Guangtao Zhai, Wenjun Zhang 0005 |
NeurIPS | 9 |
| 2024 | Face2QR: A Unified Framework for Aesthetic, Face-Preserving, and Scannable QR Code GenerationabstractExisting methods to generate aesthetic QR codes, such as image and style transfer techniques, tend to compromise either the visual appeal or the scannability of QR codes when they incorporate human face identity. Addressing these imperfections, we present Face2QR—a novel pipeline specifically designed for generating personalized QR codes that harmoniously blend aesthetics, face identity, and scannability. Our pipeline introduces three innovative components. First, the ID-refined QR integration (IDQR) seamlessly intertwines the background styling with face ID, utilizing a unified SD-based framework with control networks. Second, the ID-aware QR ReShuffle (IDRS) effectively rectifies the conflicts between face IDs and QR patterns, rearranging QR modules to maintain the integrity of facial features without compromising scannability. Lastly, the ID-preserved Scannability Enhancement (IDSE) markedly boosts scanning robustness through latent code optimization, striking a delicate balance between face ID, aesthetic quality and QR functionality. In comprehensive experiments, Face2QR demonstrates remarkable performance, outperforming existing approaches, particularly in preserving facial recognition features within custom QR code designs. Xuehao Cui, Guangyang Wu, Zhenghao Gan, Guangtao Zhai, Xiaohong Liu 0001 |
NeurIPS | 4 |
| 2024 | On Learning Multi-Modal Forgery Representation for Diffusion Generated Video DetectionabstractLarge numbers of synthesized videos from diffusion models pose threats to information security and authenticity, leading to an increasing demand for generated content detection. However, existing video-level detection algorithms primarily focus on detecting facial forgeries and often fail to identify diffusion-generated content with a diverse range of semantics. To advance the field of video forensics, we propose an innovative algorithm named Multi-Modal Detection(MM-Det) for detecting diffusion-generated videos. MM-Det utilizes the profound perceptual and comprehensive abilities of Large Multi-modal Models (LMMs) by generating a Multi-Modal Forgery Representation (MMFR) from LMM's multi-modal space, enhancing its ability to detect unseen forgery content. Besides, MM-Det leverages an In-and-Across Frame Attention (IAFA) mechanism for feature augmentation in the spatio-temporal domain. A dynamic fusion strategy helps refine forgery representations for the fusion. Moreover, we construct a comprehensive diffusion video dataset, called Diffusion Video Forensics (DVF), across a wide range of forgery videos. MM-Det achieves state-of-the-art performance in DVF, demonstrating the effectiveness of our algorithm. Both source code and DVF are available at https://github.com/SparkleXFantasy/MM-Det. Xiufeng Song, Jiache Zhang, Lei Bai 0001, Xiaoming Liu 0002, Guangtao Zhai, Xiaohong Liu 0001 |
NeurIPS | 7 |
| 2024 | ResAD: A Simple Framework for Class Generalizable Anomaly DetectionabstractThis paper explores the problem of class-generalizable anomaly detection, where the objective is to train one unified AD model that can generalize to detect anomalies in diverse classes from different domains without any retraining or fine-tuning on the target data. Because normal feature representations vary significantly across classes, this will cause the widely studied one-for-one AD models to be poorly classgeneralizable (i.e., performance drops dramatically when used for new classes). In this work, we propose a simple but effective framework (called ResAD) that can be directly applied to detect anomalies in new classes. Our main insight is to learn the residual feature distribution rather than the initial feature distribution. In this way, we can significantly reduce feature variations. Even in new classes, the distribution of normal residual features would not remarkably shift from the learned distribution. Therefore, the learned model can be directly adapted to new classes. ResAD consists of three components: (1) a Feature Converter that converts initial features into residual features; (2) a simple and shallow Feature Constraintor that constrains normal residual features into a spatial hypersphere for further reducing feature variations and maintaining consistency in feature scales among different classes; (3) a Feature Distribution Estimator that estimates the normal residual feature distribution, anomalies can be recognized as out-of-distribution. Despite the simplicity, ResAD can achieve remarkable anomaly detection results when directly used in new classes. The code is available at https://github.com/xcyao00/ResAD. Xincheng Yao, Zixin Chen, Guangtao Zhai |
NeurIPS | 4 |
| 2024 | Adaptive Image Quality Assessment via Teaching Large Multimodal Model to CompareabstractWhile recent advancements in large multimodal models (LMMs) have significantly improved their abilities in image quality assessment (IQA) relying on absolute quality rating, how to transfer reliable relative quality comparison outputs to continuous perceptual quality scores remains largely unexplored. To address this gap, we introduce an all-around LMM-based NR-IQA model, which is capable of producing qualitatively comparative responses and effectively translating these discrete comparison outcomes into a continuous quality score. Specifically, during training, we present to generate scaled-up comparative instructions by comparing images from the same IQA dataset, allowing for more flexible integration of diverse IQA datasets. Utilizing the established large-scale training corpus, we develop a human-like visual quality comparator. During inference, moving beyond binary choices, we propose a soft comparison method that calculates the likelihood of the test image being preferred over multiple predefined anchor images. The quality score is further optimized by maximum a posteriori estimation with the resulting probability matrix. Extensive experiments on nine IQA datasets validate that the Compare2Score effectively bridges text-defined comparative levels during training with converted single image quality scores for inference, surpassing state-of-the-art IQA models across diverse scenarios. Moreover, we verify that the probability-matrix-based inference conversion not only improves the rating accuracy of Compare2Score but also zero-shot general-purpose LMMs, suggesting its intrinsic effectiveness. Hanwei Zhu, Haoning Wu 0001, Baoliang Chen, Lingyu Zhu 0006, Yuming Fang 0001, Guangtao Zhai, Weisi Lin, Shiqi Wang 0001 |
NeurIPS | 8 |
| 2024 | Perceptual Skin Tone Color Difference Measurement for Portrait PhotographyabstractIn portrait photography, measuring the perceptual color differences (CDs) of skin tone is significant. Many studies have documented that the perception of skin tone is characteristically different from that of other colors. However, most existing CD measures are proposed based on psychophysical data of uniform color patches or natural images, and do not generalize well to the measurement of skin tone. In this paper, we construct the first large-scale portrait dataset for perceptual skin tone CD assessment and conduct psychophysical experiments to collect 160,000 perceptual CD judgments for 40,000 image triplets. Based on this dataset, we propose a deep skin tone CD measure for portrait photography. Extensive experiments demonstrate that our measure substantially outperforms existing CD measures on the problem of assessing skin tone CDs. The constructed dataset and code will be released to facilitate future research. Shiqi Gao, Huiyu Duan, Qihang Xu, Jia Wang 0004, Xiongkuo Min, Guangtao Zhai, Patrick Le Callet |
VCIP | 6 |
| 2024 | ACIQA: A Dataset and Method for Assessing the Imaging Quality of Automotive CamerasabstractThe imaging quality of automotive cameras is crucial in complex driving environments. Therefore, it is essential to conduct subjective experiments that can realistically reflect drivers’ evaluation of the imaging quality of automotive cameras in real traffic scenarios. To accurately assess the imaging quality of automotive cameras, this paper proposes a no-reference quality assessment method with quality scores that are highly consistent with human subjective perception. Initially, this study constructs a new image quality assessment dataset and then obtains the subjective scores of image quality through subjective experiments. The dataset is constructed by using a variety of realistic props to simulate scene elements that might be captured by an automotive camera and are captured using a wide range of cameras with different sensor types, lens focus, and viewing angles, resulting in a dataset of diverse images. The objective quality assessment method proposed in this paper consists of an object detection network and a multi-branch quality evaluation network. The object detection network is responsible for identifying and classifying scene elements, while the multi-branch quality evaluation network performs feature extraction and score regression on various types of elements to effectively evaluate the imaging quality of the automotive cameras. In the experiments, this no-reference quality assessment method is tested on our built dataset, and the results show that the proposed method exhibits the best performance compared with the state-of-the-art image quality assessment methods. Haoyang Ni, Kaiwei Zhang, Ziheng Jia, Fangfang Lu, Xiongkuo Min, Guangtao Zhai |
VCIP | 7 |
| 2024 | End-to-end Prediction of Streaming Video Quality of Experience: Dataset and ApproachabstractWith the rapid development of video-on-demand (VOD) and real-time streaming video technologies, the accurate objective assessment of streaming video Quality of Experience (QoE) has become a focal point for optimizing streaming-related technologies. However, due to the inherent transmission distortions caused by poor Quality of Service (QoS) conditions in streaming videos, such as intermittent stalling, rebuffering, and drastic changes in video sharpness due to bitrate fluctuations, evaluating streaming video QoE presents numerous challenges. This paper introduces a large and diverse in-the-wild streaming video QoE evaluation dataset - the SJLIVE-1k dataset. This work addresses the limitations of corresponding datasets, which lack in-the-wild video sequences under real network conditions and whose amount of video content is insufficient. Furthermore, we propose an end-to-end objective QoE evaluation strategy that extracts video content and QoS features from the video itself without using any extra information. By implementing self-supervised contrastive learning as the "reminder" to bridge the gap between the different types of features, our approach achieves state-of-the-art results across three datasets. Our proposed dataset will be released to facilitate further research. Ziheng Jia, Xiongkuo Min, Guangtao Zhai |
VCIP | 3 |
| 2024 | MVBind: Self-Supervised Music Recommendation for Videos via Embedding Space BindingabstractRecent years have witnessed the rapid development of short videos, which usually contain both visual and audio modalities. Background music is important to the short videos, which can significantly influence the emotions of the viewers. However, at present, the background music of short videos is generally chosen by the video producer, and there is a lack of automatic music recommendation methods for short videos. This paper introduces MVBind, an innovative Music-Video embedding space Binding model for cross-modal retrieval. MVBind operates as a self-supervised approach, acquiring inherent knowledge of intermodal relationships directly from data, without the need of manual annotations. Additionally, to compensate the lack of a corresponding musical-visual pair dataset for short videos, we construct a dataset, SVM-10K (Short Video with Music-10K), which mainly consists of meticulously selected short videos. On this dataset, MVBind manifests significantly improved performance compared to other baseline methods. The database and code are available at: https://github.com/IntMeGroup/MVBind. Jiajie Teng, Huiyu Duan, Yucheng Zhu, Sijing Wu, Guangtao Zhai |
VCIP | 5 |
| 2024 | ReLI-QA: A Multidimensional Quality Assessment Dataset for Relighted Human HeadsabstractLighting conditions significantly affect the quality of both real and AI-generated images. Facial images are particularly sensitive to lighting due to their detailed nature and the importance of facial features in conveying identity. Poor lighting can easily obscure these critical details. To address this issue, various portrait relighting methods have been developed to adjust the lighting in improperly exposed images. However, these methods often encounter challenges such as overexposure, underexposure, and detail loss in the relighted portraits. Consequently, there is a need for effective quality assessment and control of relighted human heads (RHHs). In this study, one proposed simple baseline and three typical relighting methods are applied to six selected human head (HH) images, resulting in the creation of a quality assessment dataset named ReLI-QA, which comprises 840 RHHs. A multidimensional subjective quality assessment method based on visual guidance is proposed to accurately evaluate the visual quality of each RHH in the dataset. By analyzing the results of subjective experiments, the quality of RHHs is shown to be affected by multiple factors. Finally, based on ReLI-QA, some typical image quality assessment (IQA) methods are selected for benchmark experiments. The experimental results show the limitations of the existing methods in RHH quality assessment. The dataset and code for this research has been released at https://github.com/zyj-2000/ReLI-QA. Yingjie Zhou 0003, Farong Wen, Jun Jia, Xiongkuo Min, Jia Wang 0004, Guangtao Zhai |
VCIP | 7 |
| 2024 | Perceptual video quality assessment: a surveyabstractAbstract Perceptual video quality assessment plays a vital role in the field of video processing due to the existence of quality degradations introduced in various stages of video signal acquisition, compression, transmission and display. With the advancement of Internet communication and cloud service technology, video content and traffic are growing exponentially, which further emphasizes the requirement for accurate and rapid assessment of video quality. Therefore, numerous subjective and objective video quality assessment studies have been conducted over the past two decades for both generic videos and specific videos such as streaming, user-generated content, 3D, virtual and augmented reality, high dynamic range, high frame rate, audio-visual, etc. This survey provides an up-to-date and comprehensive review of these video quality assessment studies. Specifically, we first review the subjective video quality assessment methodologies and databases, which are necessary for validating the performance of video quality metrics. Second, the objective video quality assessment measures for general purposes are categorized and surveyed according to the methodologies utilized in the quality measures. Third, we overview the objective video quality assessment measures for specific applications and emerging topics. Finally, the performance of the state-of-the-art video quality assessment measures is compared and analyzed. This survey provides a systematic overview of both classical works and recent progress in the realm of video quality assessment, which can help other researchers quickly access the field and conduct relevant research. Xiongkuo Min, Huiyu Duan, Wei Sun 0029, Yucheng Zhu, Guangtao Zhai |
Sci. China Inf. Sci. | 5 |
| 2024 | Physical layer signal processing for XR communications and systems
Yongpeng Wu 0001, Mai Xu, Guangtao Zhai, Wenjun Zhang 0001 |
Sci. China Inf. Sci. | 3 |
| 2024 | Analysis of Video Quality Datasets via Design of Minimalistic Video Quality ModelsabstractBlind video quality assessment (BVQA) plays an indispensable role in monitoring and improving the end-users' viewing experience in various real-world video-enabled media applications. As an experimental field, the improvements of BVQA models have been measured primarily on a few human-rated VQA datasets. Thus, it is crucial to gain a better understanding of existing VQA datasets in order to properly evaluate the current progress in BVQA. Towards this goal, we conduct a first-of-its-kind computational analysis of VQA datasets via designing minimalistic BVQA models. By minimalistic, we restrict our family of BVQA models to build only upon basic blocks: a video preprocessor (for aggressive spatiotemporal downsampling), a spatial quality analyzer, an optional temporal quality analyzer, and a quality regressor, all with the simplest possible instantiations. By comparing the quality prediction performance of different model variants on eight VQA datasets with realistic distortions, we find that nearly all datasets suffer from the easy dataset problem of varying severity, some of which even admit blind image quality assessment (BIQA) solutions. We additionally justify our claims by comparing our model generalization capabilities on these VQA datasets, and by ablating a dizzying set of BVQA design choices related to the basic building blocks. Our results cast doubt on the current progress in BVQA, and meanwhile shed light on good practices of constructing next-generation VQA datasets and models. Wei Sun 0029, Wen Wen 0007, Xiongkuo Min, Long Lan, Guangtao Zhai, Kede Ma |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2024 | A Coding Framework and Benchmark Towards Low-Bitrate Video UnderstandingabstractVideo compression is indispensable to most video analysis systems. Despite saving the transportation bandwidth, it also deteriorates downstream video understanding tasks, especially at low-bitrate settings. To systematically investigate this problem, we first thoroughly review the previous methods, revealing that three principles, i.e., task-decoupled, label-free, and data-emerged semantic prior, are critical to a machine-friendly coding framework but are not fully satisfied so far. In this paper, we propose a traditional-neural mixed coding framework that simultaneously fulfills all these principles, by taking advantage of both traditional codecs and neural networks (NNs). On one hand, the traditional codecs can efficiently encode the pixel signal of videos but may distort the semantic information. On the other hand, highly non-linear NNs are proficient in condensing video semantics into a compact representation. The framework is optimized by ensuring that a transportation-efficient semantic representation of the video is preserved w.r.t. the coding procedure, which is spontaneously learned from unlabeled data in a self-supervised manner. The videos collaboratively decoded from two streams (codec and NN) are of rich semantics, as well as visually photo-realistic, empirically boosting several mainstream downstream video analysis task performances without any post-adaptation procedure. Furthermore, by introducing the attention mechanism and adaptive modeling scheme, the video semantic modeling ability of our approach is further enhanced. Fianlly, we build a low-bitrate video understanding benchmark with three downstream tasks on eight datasets, demonstrating the notable superiority of our approach. All codes, data, and models will be open-sourced for facilitating future research. Yuan Tian 0017, Guo Lu, Yichao Yan, Guangtao Zhai, Li Chen 0021 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2024 | Q-Bench$^+$+: A Benchmark for Multi-Modal Foundation Models on Low-Level Vision From Single Images to PairsabstractThe rapid development of Multi-modality Large Language Models (MLLMs) has navigated a paradigm shift in computer vision, moving towards versatile foundational models. However, evaluating MLLMs in low-level visual perception and understanding remains a yet-to-explore domain. To this end, we design benchmark settings to emulate human language responses related to low-level vision: the low-level visual perception (A1) via visual question answering related to low-level attributes (e.g. clarity, lighting); and the low-level visual description (A2), on evaluating MLLMs for low-level text descriptions. Furthermore, given that pairwise comparison can better avoid ambiguity of responses and has been adopted by many human experiments, we further extend the low-level perception-related questionanswering and description evaluations of MLLMs from single images to image pairs. Specifically, for perception (A1), we carry out the LLVisionQA+ dataset, comprising 2,990 single images and 1,999 image pairs each accompanied by an open-ended question about its low-level features; for description (A2), we propose the LLDescribe+ dataset, evaluating MLLMs for low-level descriptions on 499 single images and 450 pairs. Additionally, we evaluate MLLMs on assessment (A3) ability, i.e. predicting score, by employing a softmax-based approach to enable all MLLMs to generate quantifiable quality ratings, tested against human opinions in 7 image quality assessment (IQA) datasets. With 24 MLLMs under evaluation, we demonstrate that several MLLMs have decent low-level visual competencies on single images, but only GPT-4V exhibits higher accuracy on pairwise comparisons than single image evaluations (like humans). We hope that our benchmark will motivate further research into uncovering and enhancing these nascent capabilities of MLLMs. Datasets will be available at https://github.com/Q-Future/Q-Bench. Haoning Wu 0001, Erli Zhang 0001, Guangtao Zhai, Weisi Lin |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2024 | Duration-aware and mode-aware micro-expression spotting for long video sequences
Jing Liu 0001, Guangtao Zhai, Yuting Su 0001 |
Signal Process. Image Commun. | 4 |
| 2024 | Vision-Language Consistency Guided Multi-Modal Prompt Learning for Blind AI Generated Image Quality AssessmentabstractRecently, textual prompt tuning has shown inspirational performance in adapting Contrastive Language-Image Pre-training (CLIP) models to natural image quality assessment. However, such uni-modal prompt learning method only tunes the language branch of CLIP models. This is not enough for adapting CLIP models to AI generated image quality assessment (AGIQA) since AGIs visually differ from natural images. In addition, the consistency between AGIs and user input text prompts, which correlates with the perceptual quality of AGIs, is not investigated to guide AGIQA. In this letter, we propose vision-language consistency guided multi-modal prompt learning for blind AGIQA, dubbed CLIP-AGIQA. Specifically, we introduce learnable textual and visual prompts in language and vision branches of CLIP models, respectively. Moreover, we design a text-to-image alignment quality prediction task, whose learned vision-language consistency knowledge is used to guide the optimization of the above multi-modal prompts. Experimental results on two public AGIQA datasets demonstrate that the proposed method outperforms state-of-the-art quality assessment models. Jun Fu 0007, Wei Zhou 0021, Qiuping Jiang, Hantao Liu, Guangtao Zhai |
IEEE Signal Process. Lett. | 5 |
| 2024 | Channel Attention for No-Reference Image Quality Assessment in DCT DomainabstractAttention mechanism, especially self-attention, has gained great success in image quality assessment. The advent of Transformer has led to a substantial enhancement in noreference image quality assessment (NR-IQA). Existing works focus on leveraging the global perceptual capability of Transformer encoders to perceive image quality. In this work, we start from a different view and propose a novel multi-frequency channel attention framework for Transformer encoder. Through frequency analysis, we demonstrate mathematically that traditional global average pooling (GAP) is a specific instance of feature decomposition in the frequency domain. With the proof, we use the discrete cosine transform to compress channels, which optimally compresses channels by efficiently utilizing frequency components overlooked by GAP. The experimental results show that the proposed method leads to improvements of performance over the state-of-the-art methods. Zesheng Wang 0004, Guangtao Zhai |
IEEE Signal Process. Lett. | 3 |
| 2024 | BAND-2k: Banding Artifact Noticeable Database for Banding Detection and Quality AssessmentabstractBanding, also known as staircase-like contours, frequently occurs in flat areas of images/videos processed by compression or quantization algorithms. As undesirable artifacts, banding destroys the original image structure, thus inevitably degrading users’ quality of experience (QoE). In this paper, we systematically investigate the banding image quality assessment (IQA) problem, aiming to detect the image banding artifacts and evaluate their perceptual visual quality. Considering that the existing image banding databases only contain limited content sources and banding generation methods, and lack perceptual quality labels (i.e. mean opinion scores), we first build the largest banding IQA database so far, namedBanding Artifact Noticeable Database (BAND-2k), which consists of 2,000 banding images generated by 15 compression and quantization schemes. A total of 23 workers participated in the subjective IQA experiment, yielding over 214,000 patch-level banding class labels and 44,371 reliable image-level quality rating scores. Subsequently, we develop an effective no-reference (NR) banding evaluator for banding detection and quality assessment by leveraging frequency characteristics of banding artifacts. To be more specific, a dual convolutional neural network (CNN) is employed to concurrently learn the feature representation from the high-frequency and low-frequency maps, thereby enhancing the ability to discern banding artifacts. The quality score of a banding image is generated by pooling the banding detection maps masked by the spatial frequency filters. The experimental results demonstrate that our banding evaluator achieves remarkably high accuracy in banding detection and also exhibits high SRCC and PLCC results with the perceptual quality labels, even without directly learning a regression model for banding quality evaluation. These findings unveil the strong correlations between the intensity of banding artifacts and the perceptual visual quality, thus validating the necessity of banding quality assessment. The BAND-2k database and the proposed banding evaluator are available at https://github.com/zijianchen98/BAND-2k. Zijian Chen 0001, Wei Sun 0029, Jun Jia, Fangfang Lu, Jing Liu 0002, Ru Huang 0002, Xiongkuo Min, Guangtao Zhai |
IEEE Trans. Circuits Syst. Video Technol. | 9 |
| 2024 | Blind Image Quality Assessment: A Fuzzy Neural Network for Opinion Score Distribution PredictionabstractImage quality assessment (IQA) has always been a popular research topic. There have been many methods proposed for predicting image quality, also known as the mean opinion score (MOS). However, it is worth noting that different people may assign different opinion scores to the same image. Image quality described by all subjective opinion scores can express rich subjective information about the image, such as diversity and uncertainty, which cannot be accurately described by a single MOS. Therefore, this paper proposes a fuzzy neural network to predict the opinion score distribution (OSD) of image quality. The fuzzy neural network includes three sub-networks: a feature extraction network, a feature fuzzification network, and a fuzzy learning network. First, a novel network is designed to extract image features. The extracted features are then fuzzified by fuzzy theory to model the epistemic uncertainty in the feature extraction process. Finally, the OSD of image quality is predicted using the fuzzy learning network by learning the mapping from fuzzy features to fuzzy uncertainty when rating image quality. In addition, to train the proposed fuzzy neural network, we employ a new loss function based on the quantile and the cumulative density function. We experimentally validate the feasibility and superiority of the proposed method in two aspects. On the one hand, we demonstrate the performance of the proposed method in predicting the OSD of image quality on the SJTU IQSD and KonIQ-10K databases. On the other hand, we also prove the feasibility of the proposed method in predicting the MOS of image quality on several popular IQA databases, including CSIQ, TID2013, LIVE MD, and LIVE Challenge. Xiongkuo Min, Yucheng Zhu, Xiao-Ping Zhang 0002, Guangtao Zhai |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2024 | Continuous and Overall Quality of Experience Evaluation for Streaming Video Based on Rich Features Exploration and Dual-Stage AttentionabstractWith the rapid development of streaming media technology, the Quality of Experience (QoE) of streaming videos becomes crucial to optimize the video compression and transmission algorithms, such as adaptive bitrate (ABR). However, the complexity of human perceptual mechanisms, particularly in relation to temporal distortions, poses substantial challenges to effective QoE monitoring. In recent years, many efforts in video quality assessment (VQA) and video QoE evaluation have highlighted the influence of a broad spectrum of features—from Quality of Service (QoS) metrics to video content understanding—on viewer experience. On this basis, we believe that there is also a dynamic relationship among these features varying with the broadcasting content. Furthermore, research indicates a significant correlation between real-time and retrospective assessments of QoE for individual videos. In response to these insights, we introduce a novel approach leveraging a unified learnable network that incorporates dual-stage attention, the temporal and cross-feature attention, to accurately predict both continuous and overall QoE for streaming videos. The results of experiments conducted on several publicly available databases demonstrate the superiority of our proposed method over the state-of-the-art metrics. Ziheng Jia, Xiongkuo Min, Wei Sun 0029, Guangtao Zhai |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2024 | AGIQA-3K: An Open Database for AI-Generated Image Quality AssessmentabstractWith the rapid advancements of the text-to-image generative model, AI-generated images (AGIs) have been widely applied to entertainment, education, social media, etc. However, considering the large quality variance among different AGIs, there is an urgent need for quality models that are consistent with human subjective ratings. To address this issue, we extensively consider various popular AGI models, generated AGI through different prompts and model parameters, and collected subjective scores at the perceptual quality and text-to-image alignment, thus building the most comprehensive AGI subjective quality database AGIQA-3K so far. Furthermore, we conduct a benchmark experiment on this database to evaluate the consistency between the current Image Quality Assessment (IQA) model and human perception, while proposing StairReward that significantly improves the assessment performance of subjective text-to-image alignment. We believe that the fine-grained subjective scores in AGIQA-3K will inspire subsequent AGI quality models to fit human subjective perception mechanisms at both perception and alignment levels and to optimize the generation result of future AGI models. The database is released on https://github.com/lcysyzxdxc/AGIQA-3k-Database. Chunyi Li 0001, Haoning Wu 0001, Wei Sun 0029, Xiongkuo Min, Xiaohong Liu 0001, Guangtao Zhai, Weisi Lin |
IEEE Trans. Circuits Syst. Video Technol. | 7 |
| 2024 | Un-Gaze: A Unified Transformer for Joint Gaze-Location and Gaze-Object DetectionabstractThis paper proposes an efficient and effective method for joint gaze location detection (GL-D) and gaze object detection (GO-D), i.e., gaze following detection. Current approaches frame GL-D and GO-D as two separate tasks, employing a multi-stage framework where human head crops must first be detected and then be fed into a subsequent GL-D sub-network, which is further followed by an additional object detector for GO-D. In contrast, we reframe the gaze following detection task as detecting human head locations and their gaze followings simultaneously, aiming at jointly detect human gaze location and gaze object in a unified and single-stage pipeline. To this end, we propose GTR, short for Gaze following detection TRansformer, streamlining the gaze following detection pipeline by eliminating all additional components, leading to the first unified paradigm that unites GL-D and GO-D in a fully end-to-end manner. GTR enables an iterative interaction between holistic semantics and human head features through a hierarchical structure, inferring the relations of salient objects and human gaze from the global image context and resulting in an impressive accuracy. Concretely, GTR achieves a 12.1 mAP gain ($\mathbf {25.1}\%$) on GazeFollowing and a 18.2 mAP gain ($\mathbf {43.3\%}$) on VideoAttentionTarget for GL-D, as well as a 19 mAP improvement ($\mathbf {45.2\%}$) on GOO-Real for GO-D. Meanwhile, unlike existing systems detecting gaze following sequentially due to the need for a human head as input, GTR has the flexibility to comprehend any number of people’s gaze followings simultaneously, resulting in high efficiency. Specifically, GTR introduces over a$\times 9$improvement in FPS and the relative gap becomes more pronounced as the human number grows. Danyang Tu, Wei Shen 0002, Wei Sun 0029, Xiongkuo Min, Guangtao Zhai, Chang Wen Chen |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2024 | Synergetic Assessment of Quality and Aesthetic: Approach and Comprehensive Benchmark DatasetabstractQuantifications of image quality and aesthetic have been regarded as two independent fields in computer vision. Generally, image quality assessment aims at measuring image distortions and image aesthetic is judged by commonly established photography rules. However, either measuring image quality or aesthetic alone is not sufficient to qualitatively rank images. Therefore, this paper puts forward the synergetic assessment of quality and aesthetic to help understand the subjective human preferences of digital pictures more comprehensively. Specifically, considering that the images of existing benchmark datasets are only labeled with single attribute, we first establish a new dataset which contains 9042 real-world images with the corresponding human rated pair-wise quality-aesthetic scores. Previously, these images are only labeled with aesthetic score, and we evaluate the subjective quality score of them, so that it can make up the lack of image dataset with double attributes. Moreover, since the existing methods are mostly designed for individual attribute prediction. We then propose a two-stream learning network to assess both quality and aesthetic of images in parallel. This network follows the top-down perception mechanism which learns from both fined grained details and holistic image layout simultaneously. Furthermore, we introduce a Channel-Diversity loss, which can be deployed in grouped convolution operation, and can constrain channels to be mutually exclusive across the spatial dimensions. To some extent, this contributes to spotlight different local discriminative regions with a finer granularity. Finally, experiments demonstrate that our method outperforms the state-of-the-art methods on our established benchmark dataset and other benchmark datasets in terms of image quality and aesthetic assessment. We hope this paper could serve as a potent reference and be useful for future research on the study of image ranking. Both the benchmark dataset and the code will be publicly available to facilitate further research. Kaiwei Zhang, Dandan Zhu 0001, Xiongkuo Min, Zhongpai Gao, Guangtao Zhai |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2024 | How is Visual Attention Influenced by Text Guidance? Database and ModelabstractThe analysis and prediction of visual attention have long been crucial tasks in the fields of computer vision and image processing. In practical applications, images are generally accompanied by various text descriptions, however, few studies have explored the influence of text descriptions on visual attention, let alone developed visual saliency prediction models considering text guidance. In this paper, we conduct a comprehensive study on text-guided image saliency (TIS) from both subjective and objective perspectives. Specifically, we construct a TIS database named SJTU-TIS, which includes 1200 text-image pairs and the corresponding collected eye-tracking data. Based on the established SJTU-TIS database, we analyze the influence of various text descriptions on visual attention. Then, to facilitate the development of saliency prediction models considering text influence, we construct a benchmark for the established SJTU-TIS database using state-of-the-art saliency models. Finally, considering the effect of text descriptions on visual attention, while most existing saliency models ignore this impact, we further propose a text-guided saliency (TGSal) prediction model, which extracts and integrates both image features and text features to predict the image saliency under various text-description conditions. Our proposed model significantly outperforms the state-of-the-art saliency models on both the SJTU-TIS database and the pure image saliency databases in terms of various evaluation metrics. The SJTU-TIS database and the code of the proposed TGSal model will be released at: https://github.com/IntMeGroup/TGSal. Xiongkuo Min, Huiyu Duan, Guangtao Zhai |
IEEE Trans. Image Process. | 4 |
| 2024 | Task-Specific Normalization for Continual Learning of Blind Image Quality ModelsabstractIn this paper, we present a simple yet effective continual learning method for blind image quality assessment (BIQA) with improved quality prediction accuracy, plasticity-stability trade-off, and task-order/-length robustness. The key step in our approach is to freeze all convolution filters of a pre-trained deep neural network (DNN) for an explicit promise of stability, and learn task-specific normalization parameters for plasticity. We assign each new IQA dataset (i.e., task) a prediction head, and load the corresponding normalization parameters to produce a quality score. The final quality estimate is computed by a weighted summation of predictions from all heads with a lightweight K -means gating mechanism. Extensive experiments on six IQA datasets demonstrate the advantages of the proposed method in comparison to previous training techniques for BIQA. Weixia Zhang, Kede Ma, Guangtao Zhai, Xiaokang Yang 0001 |
IEEE Trans. Image Process. | 3 |
| 2024 | UIQI: A Comprehensive Quality Evaluation Index for Underwater ImagesabstractDue to the light absorption and scattering in waterbodies, acquired underwater images frequently suffer from color cast, blur, low contrast, noise, etc., which seriously degrade the image quality and affect their subsequent applications. Therefore, it is necessary to propose a reliable and practical underwater image quality assessment (IQA) model that can faithfully evaluate underwater image quality. To this end, in this article, we establish a novel quality assessment model for underwater images by in-depth analysis and characterization of multiple image properties. Specifically, we propose characterizing the image luminance, color cast, sharpness, contrast, fog density and noise to comprehensively describe the image quality to evaluate the underwater image quality more accurately. Dedicated features are elaborately investigated to characterize those quality-aware image properties. After feature extraction, we employ support vector regression (SVR) to integrate all the quality-aware features and regress them onto the underwater image quality score. Extensive tests performed on standard underwater image quality databases demonstrate the superior prediction performance of the proposed underwater IQA model to state-of-the-art congeneric quality assessment models. Yutao Liu 0002, Ke Gu 0001, Jingchao Cao, Shiqi Wang 0001, Guangtao Zhai, Junyu Dong, Sam Kwong |
IEEE Trans. Multim. | 5 |
| 2024 | Pixel-Learnable 3DLUT With Saturation-Aware Compensation for Image EnhancementabstractThe 3D Lookup Table (3DLUT)-based methods are gaining popularity due to their satisfactory and stable performance in achieving automatic and adaptive real time image enhancement. In this paper, we present a new solution to the intractability in handling continuous color transformations of 3DLUT due to the lookup via three independent color channel coordinates in RGB space. Inspired by the inherent merits of the HSV color space, we separately enhance image intensity and color composition. The Transformer-based Pixel-Learnable 3D Lookup Table is proposed to undermine contouring artifacts, which enhances images in a pixel-wise manner with non-local information to emphasize the diverse spatially variant context. In addition, noticing the underestimation of composition color component, we develop the Saturation-Aware Compensation (SAC) module to enhance the under-saturated region determined by an adaptive SA map with Saturation-Interaction block, achieving well balance between preserving details and color rendition. Our approach can be applied to image retouching and tone mapping tasks with fairly good generality, especially in restoring localized regions with weak visibility. The performance in both theoretical analysis and comparative experiments manifests that the proposed solution is effective and robust. Jing Liu 0002, Xiongkuo Min, Yuting Su 0001, Guangtao Zhai, Xiaokang Yang 0001 |
IEEE Trans. Multim. | 5 |
| 2024 | Underwater Image Quality Assessment: Benchmark Database and Objective MethodabstractUnderwater image quality assessment (UIQA) plays a crucial role in monitoring and detecting the quality of acquired underwater images in underwater imaging systems. Currently, the investigation of UIQA encounters two major challenges. First, a lack of large-scale UIQA databases for benchmarking UIQA algorithms remains, which greatly restricts the development of UIQA research. The other limitation is that there is a shortage of effective UIQA methods that can faithfully predict underwater image quality. To alleviate these two challenges, in this paper, we first construct a large-scale UIQA database (UIQD). Specifically, UIQD contains a total of 5369 authentic underwater images that span abundant underwater scenes and typical quality degradation conditions. Extensive subjective experiments are executed to annotate the perceived quality of the underwater images in UIQD. Based on an in-depth analysis of underwater image characteristics, we further establish a novel baseline UIQA metric that integrates channel and spatial attention mechanisms and a transformer. Channel- and spatial attention modules are used to capture the image channel and local quality degradations, while the transformer module characterizes the image quality from a global perspective. Multilayer perception is employed to fuse the local and global feature representations and yield the image quality score. Extensive experiments conducted on UIQD demonstrate that the proposed UIQA model achieves superior prediction performance compared with the state-of-the-art UIQA and IQA methods. The proposed UIQD and UIQA models will be released athttps://github.com/YT2015?tab=repositories. Yutao Liu 0002, Baochao Zhang, Runze Hu, Ke Gu 0001, Guangtao Zhai, Junyu Dong |
IEEE Trans. Multim. | 5 |
| 2024 | Consistent GT-Proposal Assignment for Challenging Pedestrian DetectionabstractAccurate pedestrian classification and localization has garnered significant attention due to their extensive applications in various multimedia applications such as security monitoring, autonomous driving, and more. We have observed that the commonly employed Intersection over Union (IoU) metric in many pedestrian detectors is susceptible to an inconsistent GT-Proposal assignment issue. This issue arises when spatially adjacent proposals, which have highly similar features, are assigned to distinct ground-truth boxes, leading to confusion during the training process and an increased number of false positives during inference. To address this challenge, our work presents a novel algorithm namedDirectionalAssignmentStrategy (DAS). Firstly, in conjunction with depth distribution, our approach transforms the assignment metric from a two-dimensional (2D) view into a three-dimensional (3D) space, enabling the optimization of the regression head under the constraint of depth direction. Secondly, in contrast to the conventional IoU-basedone-to-oneassignment of one proposal to one ground-truth box, our method aims to establish a more reasoned matching between sets of proposals and ground-truth boxes. By doing so, the detector is less reliant on the setting of a specific threshold. Leveraging this strategy as a plug-in module within state-of-the-art pedestrian detectors, we demonstrate a notable improvement in performance. Yan Luo 0003, Muming Zhao, Jun Sun 0005, Guangtao Zhai |
IEEE Trans. Multim. | 4 |
| 2024 | Lightweight Video-Based Respiration Rate Detection Algorithm: An Application Case on Intensive CareabstractThe video-based non-contact respiration detection technology can be used in many application scenarios to unobtrusively and ubiquitously monitor the physical state of living beings, and various researchers are currently working on this technology. The optical flow method in tandem with crossover point method is rather effective for respiration rate extraction. However, each method has one disadvantage: 1) the redundant feature points in the traditional optical flow method increase the computational effort and reduce the estimation accuracy; and 2) the traditional crossover point method suffers from crossover points unrelated to breathing movements. For these two challenges, two optimization points are proposed in this work: 1) optimize feature point space by combining spatio-temporal information; and 2) use negative feedback design to adaptively remove crossovers that are not related to respiratory movements. The performance of the proposed algorithm is validated by the Large-scale Bedside Respiration Dataset for Intensive Care (LBRD-IC), which is established using the actual surveillance videos acquired from ICU wards. The validity of the above two optimization points is verified by the ablation experiments. The influential analysis of computation time and video resolution on the performance of the proposed algorithm demonstrates that the proposed algorithm can be deployed to various application terminals to monitor the respiration rate of living organisms in real-time and with high accuracy. In addition, field measurements in the ICU ward have shown that our algorithm can measure respiratory signals of the single patient and multiple patients when only one surveillance camera is present. Xudong Tan, Menghan Hu, Guangtao Zhai, Wenfang Li, Xiao-Ping Zhang 0002 |
IEEE Trans. Multim. | 3 |
| 2024 | Unified Audio-Visual Saliency Model for Omnidirectional Videos With Spatial AudioabstractSpatial audio is a crucial component of omnidirectional videos (ODVs), which can provide an immersive experience by enabling viewers to perceive sound sources in all directions. However, most visual attention modeling works for ODVs focus only on visual cues, and audio modality is rather rarely considered. Additionally, the existing audio-visual saliency models for ODVs lack spatial audio location-awareness (i.e. sound source location-agnostic) and audio content attributes discriminability (i.e. audio content attributes-agnostic). To this end, we propose a novel audio-visual perception saliency (AVPS) model with spatial audio location-awareness and audio content attributes-adaptive to efficiently address the problem of fixation prediction in ODVs. Specifically, we first utilize the improved group equivariant convolutional neural network (G-CNN) with eidetic 3D LSTM (E3D-LSTM) to extract spatial-temporal visual features. Then we perceive sound source locations by computing the audio energy map (AEM) of the audio information in ODVs. Subsequently, we introduce SoundNet to extract audio features with multiple attributes. Finally, we develop an audio-visual feature fusion module to adaptively integrate spatial-temporal visual features and spatial auditory information to generate the final audio-visual saliency map. Extensive experiments in three audio modalities validate the effectiveness of the proposed model. Meanwhile, the performance of the proposed model is superior to the other 10 state-of-the-art saliency models. Dandan Zhu 0001, Kaiwei Zhang, Qiangqiang Zhou, Xiongkuo Min, Guangtao Zhai, Xiaokang Yang 0001 |
IEEE Trans. Multim. | 6 |
| 2024 | Hidden Barcode in Sub-Images with Invisible Locating MarkerabstractThe prevalence of the Internet of Things (IoT) has led to the widespread adoption of 2D barcodes as a means of offline-to-online communication. Whereas, 2D barcodes are not ideal for publicity materials, due to their space-consuming nature. Recent works have proposed 2D image barcodes that contain invisible codes or hyperlinks to transmit hidden information from offline to online. However, these methods undermine the purpose of the codes being invisible, due to the the requirement of markers to locate them. The conference version of this work has presents a novel imperceptible information embedding framework for display or print-camera scenarios, which includes not only hiding and recvoery but also locating and correcting. With the assistance of learned invisible markers, hidden codes can be rendered truly imperceptible. A highly effective multi-stage training scheme is proposed to achieve high visual fidelity and retrieval resiliency, wherein information is concealed in a sub-region rather than the entire image. However, our conference version does not address the optimal sub-region for hiding, which is crucial when dealing with local region concealment problems. In this paper extension, we consider human perceptual characteristics and introduce an optimal hiding region recommendation algorithm that comprehensively incorporates Just Noticeable Difference (JND) and visual saliency factors into consideration. Extensive experiments demonstrate superior visual quality and robustness compared to state-of-the-art methods. With the assistance of our proposed hiding region recommendation algorithm, concealed information becomes even less visible than the results of our conference version without compromising robustness. Jun Jia, Zhongpai Gao, Yiwei Yang 0007, Wei Sun 0029, Dandan Zhu 0001, Xiaohong Liu 0001, Xiongkuo Min, Guangtao Zhai |
ACM Trans. Multim. Comput. Commun. Appl. | 8 |
| 2024 | Psychology-Guided Environment Aware Network for Discovering Social Interaction Groups from VideosabstractSocial interaction is a common phenomenon in human societies. Different from discovering groups based on the similarity of individuals’ actions, social interaction focuses more on the mutual influence between people. Although people can easily judge whether or not there are social interactions in a real-world scene, it is difficult for an intelligent system to discover social interactions. Initiating and concluding social interactions are greatly influenced by an individual’s social cognition and the surrounding environment, which are closely related to psychology. Thus, converting the psychological factors that impact social interactions into quantifiable visual representations and creating a model for interaction relationships poses a significant challenge. To this end, we propose a Psychology-Guided Environment Aware Network (PEAN) that models social interaction among people in videos using supervised learning. Specifically, we divide the surrounding environment into scene-aware visual-based and human-aware visual-based descriptions. For the scene-aware visual clue, we utilize 3D features as global visual representations. For the human-aware visual clue, we consider instance-based location and behaviour-related visual representations to map human-centred interaction elements in social psychology: distance, openness, and orientation. In addition, we design an environment aware mechanism to integrate features from visual clues, with a Transformer to explore the relation between individuals and construct pairwise interaction strength features. The interaction intensity matrix reflecting the mutual nature of the interaction is obtained by processing the interaction strength features with the interaction discovery module. An interaction constrained loss function composed of interaction critical loss function and smoothFβloss function is proposed to optimize the whole framework to improve the distinction of the interaction matrix and alleviate class imbalance caused by pairwise interaction sparsity. Given the diversity of real-world interactions, we collect a new dataset named Social Basketball Activity Dataset (Soical-BAD), covering complex social interactions. Our method achieves the best performance among social-CAD, social-BAD, and their combined dataset named Video Social Interaction Dataset (VSID). Jinhai Yang 0001, Hua Yang 0001, Renjie Pan 0001, Pingrui Lai, Guangtao Zhai |
ACM Trans. Multim. Comput. Commun. Appl. | 6 |
| 2024 | GMS-3DQA: Projection-Based Grid Mini-patch Sampling for 3D Model Quality AssessmentabstractNowadays, most three-dimensional model quality assessment (3DQA) methods have been aimed at improving accuracy. However, little attention has been paid to the computational cost and inference time required for practical applications. Model-based 3DQA methods extract features directly from the 3D models, which are characterized by their high degree of complexity. As a result, many researchers are inclined towards utilizing projection-based 3DQA methods. Nevertheless, previous projection-based 3DQA methods directly extract features from multi-projections to ensure quality prediction accuracy, which calls for more resource consumption and inevitably leads to inefficiency. Thus, in this article, we address this challenge by proposing a no-reference (NR) projection-based G rid M ini-patch S ampling 3D Model Q uality A ssessment (GMS-3DQA) method. The projection images are rendered from six perpendicular viewpoints of the 3D model to cover sufficient quality information. To reduce redundancy and inference resources, we propose a multi-projection grid mini-patch sampling strategy (MP-GMS), which samples grid mini-patches from the multi-projections and forms the sampled grid mini-patches into one quality mini-patch map (QMM). The Swin-Transformer tiny backbone is then used to extract quality-aware features from the QMMs. The experimental results show that the proposed GMS-3DQA outperforms existing state-of-the-art NR-3DQA methods on the point cloud quality assessment databases for both accuracy and efficiency. The efficiency analysis reveals that the proposed GMS-3DQA requires far less computational resources and inference time than other 3DQA competitors. The code is available at https://github.com/zzc-1998/GMS-3DQA . Wei Sun 0029, Haoning Wu 0001, Yingjie Zhou 0003, Chunyi Li 0001, Zijian Chen 0001, Xiongkuo Min, Guangtao Zhai, Weisi Lin |
ACM Trans. Multim. Comput. Commun. Appl. | 8 |
| 2024 | Subjective and Objective Quality Assessment for in-the-Wild Computer Graphics ImagesabstractComputer graphics images (CGIs) are artificially generated by means of computer programs and are widely perceived under various scenarios, such as games, streaming media, etc. In practice, the quality of CGIs consistently suffers from poor rendering during production, inevitable compression artifacts during the transmission of multimedia applications, and low aesthetic quality resulting from poor composition and design. However, few works have been dedicated to dealing with the challenge of computer graphics image quality assessment (CGIQA). Most image quality assessment (IQA) metrics are developed for natural scene images (NSIs) and validated on databases consisting of NSIs with synthetic distortions, which are not suitable for in-the-wild CGIs. To bridge the gap between evaluating the quality of NSIs and CGIs, we construct a large-scale in-the-wild CGIQA database consisting of 6,000 CGIs (CGIQA-6k) and carry out the subjective experiment in a well-controlled laboratory environment to obtain the accurate perceptual ratings of the CGIs. Then, we propose an effective deep learning–based no-reference (NR) IQA model by utilizing both distortion and aesthetic quality representation. Experimental results show that the proposed method outperforms all other state-of-the-art NR IQA methods on the constructed CGIQA-6k database and other CGIQA-related databases. The database is released at https://github.com/zzc-1998/CGIQA6K . Wei Sun 0029, Yingjie Zhou 0003, Jun Jia, Jing Liu 0002, Xiongkuo Min, Guangtao Zhai |
ACM Trans. Multim. Comput. Commun. Appl. | 8 |
| 2024 | Blind Quality Assessment of Dense 3D Point Clouds with Structure Guided ResamplingabstractObjective quality assessment of three-dimensional (3D) point clouds is essential for the development of immersive multimedia systems in real-world applications. Despite the success of perceptual quality evaluation for 2D images and videos, blind/no-reference metrics are still scarce for 3D point clouds with large-scale irregularly distributed 3D points. Therefore, in this article, we propose an objective point cloud quality index with Structure Guided Resampling (SGR) to automatically evaluate the perceptually visual quality of dense 3D point clouds. The proposed SGR is a general-purpose blind quality assessment method without the assistance of any reference information. Specifically, considering that the human visual system is highly sensitive to structure information, we first exploit the unique normal vectors of point clouds to execute regional pre-processing that consists of keypoint resampling and local region construction. Then, we extract three groups of quality-related features, including (1) geometry density features, (2) color naturalness features, and (3) angular consistency features. Both the cognitive peculiarities of the human brain and naturalness regularity are involved in the designed quality-aware features that can capture the most vital aspects of distorted 3D point clouds. Extensive experiments on several publicly available subjective point cloud quality databases validate that our proposed SGR can compete with state-of-the-art full-reference, reduced-reference, and no-reference quality assessment algorithms. Wei Zhou 0021, Qi Yang 0003, Qiuping Jiang, Guangtao Zhai, Weisi Lin |
ACM Trans. Multim. Comput. Commun. Appl. | 5 |
| 2023 | GANHead: Towards Generative Animatable Neural Head AvatarsabstractTo bring digital avatars into people's lives, it is highly demanded to efficiently generate complete, realistic, and animatable head avatars. This task is challenging, and it is difficult for existing methods to satisfy all the requirements at once. To achieve these goals, we propose GANHead (Generative Animatable Neural Head Avatar), a novel generative head model that takes advantages of both the fine-grained control over the explicit expression parameters and the realistic rendering results of implicit representations. Specifically, GANHead represents coarse geometry, fine-gained details and texture via three networks in canonical space to obtain the ability to generate complete and realistic head avatars. To achieve flexible animation, we define the deformation filed by standard linear blend skinning (LBS), with the learned continuous pose and expression bases and LBS weights. This allows the avatars to be directly animated by FLAME [22] parameters and generalize well to unseen poses and expressions. Compared to state-of-the-art (SOTA) methods, GANHead achieves superior performance on head avatar generation and raw scan fitting. Sijing Wu, Yichao Yan, Yuhao Cheng, Wenhan Zhu, Ke Gao 0012, Guangtao Zhai |
CVPR | 8 |
| 2023 | CASP-Net: Rethinking Video Saliency Prediction from an Audio-Visual Consistency Perceptual PerspectiveabstractIncorporating the audio stream enables Video Saliency Prediction (VSP) to imitate the selective attention mechanism of human brain. By focusing on the benefits of joint auditory and visual information, most VSP methods are capable of exploiting semantic correlation between vision and audio modalities but ignoring the negative effects due to the temporal inconsistency of audio-visual intrinsics. Inspired by the biological inconsistency-correction within multi-sensory information, in this study, a consistency-aware audio-visual saliency prediction network (CASP-Net) is proposed, which takes a comprehensive consideration of the audio-visual semantic interaction and consistent perception. In addition a two-stream encoder for elegant association between video frames and corresponding sound source, a novel consistency-aware predictive coding is also designed to improve the consistency within audio and visual representations iteratively. To further aggregate the multi-scale audio-visual information, a saliency decoder is introduced for the final saliency map generation. Substantial experiments demonstrate that the proposed CASP-Net outperforms the other state-of-the-art methods on six challenging audio-visual eye-tracking datasets. For a demo of our system please see our project webpage. Junwen Xiong, Ganglai Wang, Peng Zhang 0005, Wei Huang 0013, Yufei Zha, Guangtao Zhai |
CVPR | 6 |
| 2023 | MD-VQA: Multi-Dimensional Quality Assessment for UGC Live VideosabstractUser-generated content (UGC) live videos are often bothered by various distortions during capture procedures and thus exhibit diverse visual qualities. Such source videos are further compressed and transcoded by media server providers before being distributed to end-users. Because of the flourishing of UGC live videos, effective video quality assessment (VQA) tools are needed to monitor and perceptually optimize live streaming videos in the distributing process. In this paper, we address UGC Live VQA problems by constructing a first-of-a-kind subjective UGC Live VQA database and developing an effective evaluation tool. Concretely, 418 source UGC videos are collected in real live streaming scenarios and 3,762 compressed ones at different bit rates are generated for the subsequent subjective VQA experiments. Based on the built database, we develop a Multi-12imensional VQA (MD-VQA) evaluator to measure the visual quality of UGC live videos from semantic, distortion, and motion aspects respectively. Extensive experimental results show that MD-VQA achieves state-of-the-art performance on both our UGC Live VQA database and existing compressed UGC VQA databases. Wei Wu 0002, Wei Sun 0029, Danyang Tu, Wei Lu 0021, Xiongkuo Min, Ying Chen 0011, Guangtao Zhai |
CVPR | 8 |
| 2023 | Blind Image Quality Assessment via Vision-Language Correspondence: A Multitask Learning PerspectiveabstractWe aim at advancing blind image quality assessment (BIQA), which predicts the human perception of image quality without any reference information. We develop a general and automated multitask learning scheme for BIQA to exploit auxiliary knowledge from other tasks, in a way that the model parameter sharing and the loss weighting are determined automatically. Specifically, we first describe all candidate label combinations (from multiple tasks) using a textual template, and compute the joint probability from the cosine similarities of the visual-textual embeddings. Predictions of each task can be inferred from the joint distribution, and optimized by carefully designed loss functions. Through comprehensive experiments on learning three tasks - BIQA, scene classification, and distortion type identification, we verify that the proposed BIQA method 1) benefits from the scene classification and distortion type identification tasks and outperforms the state-of-the-art on multiple IQA datasets, 2) is more robust in the group maximum differentiation competition, and 3) realigns the quality annotations from different IQA datasets more effectively. The source code is available at https://github.com/zwx8981/LIQE. Weixia Zhang, Guangtao Zhai, Ying Wei 0001, Xiaokang Yang 0001, Kede Ma |
CVPR | 2 |
| 2023 | Unobtrusive Respiratory Monitoring System for Intensive CareabstractThe video-based non-contact respiration detection technology can be used in many application scenarios to unobtrusively and ubiquitously monitor the physical state of living beings, and various researchers are currently working on this technology. The optical flow method in tandem with crossover point method is rather effective for respiration rate extraction. However, each method has one disadvantage: 1) the redundant feature points in the traditional optical flow method increase the computational effort and reduce the estimation accuracy; and 2) the traditional crossover point method suffers from crossover points unrelated to breathing movements. For these two challenges, two optimization points are proposed 1) optimize feature point space by combining spatio-temporal information; and 2) use negative feedback design to adaptively remove crossovers unrelated to respiratory movements. The performance of the proposed algorithm is validated by the Large-scale Bedside Respiration Dataset for Intensive Care (LBRD-IC), which is established using the actual surveillance videos acquired from ICU wards. In addition, field measurements in the ICU ward have shown that our algorithm can measure respiratory signals of the single patient and multiple patients when only one surveillance camera is present. Xudong Tan, Menghan Hu, Guangtao Zhai, Wenfang Li, Xiao-Ping Zhang 0002 |
ICASSP | 3 |
| 2023 | Perceptual Quality Assessment for Digital Human HeadsabstractDigital humans are attracting more and more research interest during the last decade, the generation, representation, rendering, and animation of which have been put into large amounts of effort. However, the quality assessment of digital humans has fallen behind. Therefore, to tackle the challenge of digital human quality assessment issues, we propose the first large-scale quality assessment database for three-dimensional (3D) scanned digital human heads (DHHs). The constructed database consists of 55 reference DHHs and 1,540 distorted DHHs along with the subjective perceptual ratings. Then, a simple yet effective full-reference (FR) projection-based method is proposed to evaluate the visual quality of DHHs. The pretrained Swin Transformer tiny is employed for hierarchical feature extraction and the multi-head attention module is utilized for feature fusion. The experimental results reveal that the proposed method exhibits state-of-the-art performance among the mainstream FR metrics. The database is released at https://github.com/zzc-1998/DHHQA. Yingjie Zhou 0003, Wei Sun 0029, Xiongkuo Min, Guangtao Zhai |
ICASSP | 6 |
| 2023 | Non-Semantics Suppressed Mask Learning for Unsupervised Video Semantic CompressionabstractMost video compression methods aim to improve the decoded video visual quality, instead of particularly guaranteeing the semantic-completeness, which deteriorates downstream video analysis tasks, e.g., action recognition. In this paper, we focus on a novel unsupervised video semantic compression problem, where video semantics is compressed in a downstream task-agnostic manner. To tackle this problem, we first propose a Semantic-Mining-then-Compensation (SMC) framework to enhance the plain video codec with powerful semantic coding capability. Then, we optimize the framework with only unlabeled video data, by masking out a proportion of the compressed video and reconstructing the masked regions of the original video, which is inspired by recent masked image modeling (MIM) methods. Although the MIM scheme learns generalizable semantic features, its inner generative learning paradigm may also facilitate the coding framework memorizing non-semantic information with extra bit costs. To suppress this deficiency, we explicitly decrease the non-semantic information entropy of the decoded video features, by formulating it as a parametrized Gaussian Mixture Model conditioned on the mined video semantics. Comprehensive experimental results demonstrate the proposed approach shows remarkable superiority over previous traditional, learnable, and perceptual quality-oriented video codecs, on three video analysis tasks and seven datasets. Yuan Tian 0017, Guo Lu, Guangtao Zhai |
ICCV | 3 |
| 2023 | Agglomerative Transformer for Human-Object Interaction DetectionabstractWe propose an agglomerative Transformer (AGER) that enables Transformer-based human-object interaction (HOI) detectors to flexibly exploit extra instance-level cues in a single-stage and end-to-end manner for the first time. AGER acquires instance tokens by dynamically clustering patch tokens and aligning cluster centers to instances with textual guidance, thus enjoying two benefits: 1) Integrality: each instance token is encouraged to contain all discriminative feature regions of an instance, which demonstrates a significant improvement in the extraction of different instance-level cues and subsequently leads to a new state-of-the-art performance of HOI detection with 36.75 mAP on HICO-Det. 2) Efficiency: the dynamical clustering mechanism allows AGER to generate instance tokens jointly with the feature learning of the Transformer encoder, eliminating the need of an additional object detector or instance decoder in prior methods, thus allowing the extraction of desirable extra cues for HOI detection in a single-stage and end-to-end pipeline. Concretely, AGER reduces GFLOPs by 8.5% and improves FPS by 36%, even compared to a vanilla DETR-like pipeline without extra cue extraction. The code will be available at https://github.com/six6607/AGER.git. Danyang Tu, Wei Sun 0029, Guangtao Zhai, Wei Shen 0002 |
ICCV | 3 |
| 2023 | AccFlow: Backward Accumulation for Long-Range Optical FlowabstractRecent deep learning-based optical flow estimators have exhibited impressive performance in generating local flows between consecutive frames. However, the estimation of long-range flows between distant frames, particularly under complex object deformation and large motion occlusion, remains a challenging task. One promising solution is to accumulate local flows explicitly or implicitly to obtain the desired long-range flow. Nevertheless, the accumulation errors and flow misalignment can hinder the effectiveness of this approach. This paper proposes a novel recurrent framework called AccFlow, which recursively backward accumulates local flows using a deformable module called as AccPlus. In addition, an adaptive blending module is designed along with AccPlus to alleviate the occlusion effect by backward accumulation and rectify the accumulation error. Notably, we demonstrate the superiority of backward accumulation over conventional forward accumulation, which to the best of our knowledge has not been explicitly established before. To train and evaluate the proposed AccFlow, we have constructed a large-scale high-quality dataset named CVO, which provides ground-truth optical flow labels between adjacent and distant frames. Extensive experiments validate the effectiveness of AccFlow in handling long-range optical flow estimation. Codes are available at https://github.com/mulns/AccFlow. Guangyang Wu, Xiaohong Liu 0001, Kunming Luo, Qingqing Zheng, Shuaicheng Liu, Xinyang Jiang, Guangtao Zhai, Wenyi Wang 0005 |
ICCV | 8 |
| 2023 | Audio-Visual Saliency for Omnidirectional Videos
Xilei Zhu, Huiyu Duan, Kaiwei Zhang, Yucheng Zhu, Li Chen 0021, Xiongkuo Min, Guangtao Zhai |
ICIG (5) | 9 |
| 2023 | Audio-Visual Quality Assessment for User Generated Content: Database and MethodabstractWith the explosive increase of User Generated Content (UGC), UGC video quality assessment (VQA) becomes more and more important for improving users’ Quality of Experience (QoE). However, most existing UGC VQA studies only focus on the visual distortions of videos, ignoring that the user’s QoE also depends on the accompanying audio signals. In this paper, we conduct the first study to address the problem of UGC audio and video quality assessment (AVQA). Specifically, we construct the first UGC AVQA database named the SJTU-UAV database, which includes 520 in-the-wild UGC audio and video (A/V) sequences, and conduct a user study to obtain the mean opinion scores of the A/V sequences. The content of the SJTU-UAV database is then analyzed from both the audio and video aspects to show the database characteristics. We also design a family of AVQA models, which fuse the popular VQA methods and audio features via support vector regressor (SVR). We validate the effectiveness of the proposed models on the three databases. The experimental results show that with the help of audio signals, the VQA models can evaluate the perceptual quality more accurately. The database will be released to facilitate further research. Yuqin Cao, Xiongkuo Min, Wei Sun 0029, Xiao-Ping Zhang 0002, Guangtao Zhai |
ICIP | 5 |
| 2023 | Hierarchical Feature Fusion Transformer for No-Reference Image Quality AssessmentabstractRecently, increasing interest has been drawn in Transformer-based models for No-reference Image Quality Assessment (NR-IQA), especially for the hybrid approach. The hybrid approach tend to apply Transformer to aggregate quality information from feature maps extracted by Convolutional Neural Networks (CNN). However, existing methods cannot fully utilize the information of hierarchical features extracted by the deep neural network, resulting in the limited performance of image quality evaluation. In this work, we propose a novel Hierarchical Feature Fusion Transformer for NR-IQA (HiFFTiq), which is able to effectively exploit complementary strengths of features extracted by different layers. Further, we propose a new Uniform Partition Pooling (UPP) which can reduce the resolution of input features via uniform partitions and can well retain the quality-related information compared to the traditional pooling method Sliding Window Pooling (SWP). The results of experiment demonstrate that HiFFTiq leads to improvements of performance over the state-of-the-art methods on three large scale NR-IQA datasets. Zesheng Wang 0004, Wei Wu 0002, Wei Sun 0029, Ying Chen 0011, Kai Li 0012, Guangtao Zhai |
ICIP | 7 |
| 2023 | Geometry-Aware Video Quality Assessment for Dynamic Digital HumanabstractDynamic Digital Humans (DDHs) are 3D digital models that are animated using predefined motions and are inevitably bothered by noise/shift during the generation process and compression distortion during the transmission process, which needs to be perceptually evaluated. Usually, DDHs are displayed as 2D rendered animation videos and it is natural to adapt video quality assessment (VQA) methods to DDH quality assessment (DDH-QA) tasks. However, the VQA methods are highly dependent on viewpoints and less sensitive to geometry-based distortions. Therefore, in this paper, we propose a novel no-reference (NR) geometry-aware video quality assessment method for DDH-QA challenge. Geometry characteristics are described by the statistical parameters estimated from the DDHs’ geometry attribute distributions. Spatial and temporal features are acquired from the rendered videos. Finally, all kinds of features are integrated and regressed into quality values. Experimental results show that the proposed method achieves state-of-the-art performance on the DDH-QA database. Yingjie Zhou 0003, Wei Sun 0029, Xiongkuo Min, Guangtao Zhai |
ICIP | 5 |
| 2023 | A No-Reference Quality Assessment Method for Digital Human HeadabstractIn recent years, digital humans have been widely applied in augmented/virtual reality (A/VR), where viewers are allowed to freely observe and interact with the volumetric content. However, the digital humans may be degraded with various distortions during the procedure of generation and transmission. Moreover, little effort has been put into the perceptual quality assessment of digital humans. Therefore, it is urgent to carry out objective quality assessment methods to tackle the challenge of digital human quality assessment (DHQA). In this paper, we develop a novel no-reference (NR) method based on Transformer to deal with DHQA in a multi-task manner. Specifically, the front 2D projections of the digital humans are rendered as inputs and the vision transformer (ViT) is employed for the feature extraction. Then we design a multi-task module to jointly classify the distortion types and predict the perceptual quality levels of digital humans. The experimental results show that the proposed method well correlates with the subjective ratings and outperforms the state-of-the-art quality assessment methods. Yingjie Zhou 0003, Wei Sun 0029, Xiongkuo Min, Xianghe Ma, Guangtao Zhai |
ICIP | 6 |
| 2023 | BH-VQA: Blind High Frame Rate Video Quality AssessmentabstractHigh frame rate (HFR) videos can provide consumers with a more immersive viewing experience in motion-rich scenes. However, they also pose a great challenge for video compression and transmission due to the increase in frame rates. Therefore, it is very important to choose proper frame rates and bit rates to achieve a trade-off between transmission bandwidth and visual quality. In this paper, we propose a novel Blind HFR Video Quality Assessment (BH-VQA) model by exploring the efficient and effective motion representation from the deep neural network (DNN). Concretely, we first train a baseline VQA model (i.e. a backbone network and a regressor) on a large-scale VQA database to derive a powerful quality-aware feature extractor for the spatial and motion feature extraction. Then, the HFR video is split into a sequence of video clips and the spatial features of each video clip are extracted just using the first frame of the video clip. To capture temporal distortions caused by frame rate variations and object and camera motion, we calculate deep structural similarities between continuous frames of each video clip as the motion features. Finally, the temporal quality dependencies between video clips are learned through a gated recurrent unit (GRU) network to obtain the perceptual video quality score. Experimental results show that BH-VQA achieves the best performance on two publicly available HFR VQA databases. The code of BH-VQA will be released. Wei Lu 0021, Wei Sun 0029, Danyang Tu, Xiongkuo Min, Guangtao Zhai |
ICME | 6 |
| 2023 | EEP-3DQA: Efficient and Effective Projection-Based 3D Model Quality AssessmentabstractCurrently, great numbers of efforts have been put into improving the effectiveness of 3D model quality assessment (3DQA) methods. However, little attention has been paid to the computational costs and inference time, which is also important for practical applications. Unlike 2D media, 3D models are represented by more complicated and irregular digital formats, such as point cloud and mesh. Thus it is normally difficult to perform an efficient module to extract quality-aware features of 3D models. In this paper, we address this problem from the aspect of projection-based 3DQA and develop a no-reference (NR) Efficient and Effective Projection-based 3D Model Quality Assessment (EEP-3DQA) method. The input projection images of EEP-3DQA are randomly sampled from the six perpendicular viewpoints of the 3D model and are further spatially downsampled by the grid-mini patch sampling strategy. Further, the lightweight Swin-Transformer tiny is utilized as the backbone to extract the quality-aware features. Finally, the proposed EEP-3DQA and EEP-3DQA-t (tiny version) achieve the best performance than the existing state-of-the-art NR-3DQA methods and even outperforms most full-reference (FR) 3DQA methods on the point cloud and mesh quality assessment databases while consuming less inference time than the compared 3DQA methods. Wei Sun 0029, Yingjie Zhou 0003, Wei Lu 0021, Yucheng Zhu, Xiongkuo Min, Guangtao Zhai |
ICME | 7 |
| 2023 | DDH-QA: A Dynamic Digital Humans Quality Assessment DatabaseabstractIn recent years, large amounts of effort have been put into pushing forward the real-world application of dynamic digital human (DDH). However, most current quality assessment research focuses on evaluating static 3D models and usually ignores motion distortions. Therefore, in this paper, we construct a large-scale dynamic digital human quality assessment (DDH-QA) database with diverse motion content as well as multiple distortions to comprehensively study the perceptual quality of DDHs. Both model-based distortion (noise, compression) and motion-based distortion (binding error, motion unnaturalness) are taken into consideration. Ten types of common motion are employed to drive the DDHs and a total of 800 DDHs are generated in the end. Afterward, we render the video sequences of the distorted DDHs as the evaluation media and carry out a well-controlled subjective experiment. Then a benchmark experiment is conducted with the state-of-the-art video quality assessment (VQA) methods and the experimental results show that existing VQA methods are limited in assessing the perceptual loss of DDHs. The database is available at https://github.com/zzc-1998/DDH-QA. Yingjie Zhou 0003, Wei Sun 0029, Wei Lu 0021, Xiongkuo Min, Yu Wang 0002, Guangtao Zhai |
ICME | 7 |
| 2023 | MM-PCQA: Multi-Modal Learning for No-reference Point Cloud Quality AssessmentabstractThe visual quality of point clouds has been greatly emphasized since the ever-increasing 3D vision applications are expected to provide cost-effective and high-quality experiences for users. Looking back on the development of point cloud quality assessment (PCQA), the visual quality is usually evaluated by utilizing single-modal information, i.e., either extracted from the 2D projections or 3D point cloud. The 2D projections contain rich texture and semantic information but are highly dependent on viewpoints, while the 3D point clouds are more sensitive to geometry distortions and invariant to viewpoints. Therefore, to leverage the advantages of both point cloud and projected image modalities, we propose a novel no-reference Multi-Modal Point Cloud Quality Assessment (MM-PCQA) metric. In specific, we split the point clouds into sub-models to represent local geometry distortions such as point shift and down-sampling. Then we render the point clouds into 2D image projections for texture feature extraction. To achieve the goals, the sub-models and projected images are encoded with point-based and image-based neural networks. Finally, symmetric cross-modal attention is employed to fuse multi-modal quality-aware information. Experimental results show that our approach outperforms all compared state-of-the-art methods and is far ahead of previous no-reference PCQA methods, which highlights the effectiveness of the proposed method. The code is available at https://github.com/zzc-1998/MM-PCQA. Wei Sun 0029, Xiongkuo Min, Qiyuan Wang 0002, Guangtao Zhai |
IJCAI | 7 |
| 2023 | The Influence of Text-guidance on Visual AttentionabstractVisual attention analysis and prediction have long been important tasks in computer vision and image processing. However, images often come along with various text descriptions in real applications, while the influence of these text-guidances on the visual saliency of corresponding images have rarely been studied. Therefore, in this paper, we mainly focus on the problem of whether and how the text-guidance influences the visual attention during image viewing, and perform subjective experiments, qualitative and quantitative comparisons as well as model evaluations on this new task. Specifically, we first conduct eye tracking experiments on 300 images under text-visual (TV) and visual (V) test conditions, respectively. Based on the subjective experiments, we perform qualitative and quantitative comparisons between the visual attention data collected under TV and V conditions, and conclude that the text-guidance can significantly influence the visual attention, especially when the text-described target is a non-salient object. Finally, we evaluate the existing saliency models on our database, and find that existing models cannot well handle this text-induced saliency prediction task. Our constructed database will be publicly available to facilitate future research. Xiongkuo Min, Huiyu Duan, Guangtao Zhai |
ISCAS | 4 |
| 2023 | Light-VQA: A Multi-Dimensional Quality Assessment Model for Low-Light Video EnhancementabstractRecently, Users Generated Content (UGC) videos becomes ubiquitous in our daily lives. However, due to the limitations of photographic equipments and techniques, UGC videos often contain various degradations, in which one of the most visually unfavorable effects is the underexposure. Therefore, corresponding video enhancement algorithms such as Low-Light Video Enhancement (LLVE) have been proposed to deal with the specific degradation. However, different from video enhancement algorithms, almost all existing Video Quality Assessment (VQA) models are built generally rather than specifically, which measure the quality of a video from a comprehensive perspective. To the best of our knowledge, there is no VQA model specially designed for videos enhanced by LLVE algorithms. To this end, we first construct a Low-Light Video Enhancement Quality Assessment (LLVE-QA) dataset in which 254 original low-light videos are collected and then enhanced by leveraging 8 LLVE algorithms to obtain 2,060 videos in total. Moreover, we propose a quality assessment model specialized in LLVE, named Light-VQA. More concretely, since the brightness and noise have the most impact on low-light enhanced VQA, we handcraft corresponding features and integrate them with deep-learning-based semantic features as the overall spatial information. As for temporal information, in addition to deep-learning-based motion features, we also investigate the handcrafted brightness consistency among video frames, and the overall temporal information is their concatenation. Subsequently, spatial and temporal information is fused to obtain the quality-aware representation of a video. Extensive experimental results show that our Light-VQA achieves the best performance against the current State-Of-The-Art (SOTA) on LLVE-QA and public dataset. Dataset and Codes can be found at https://github.com/wenzhouyidu/Light-VQA. Yunlong Dong, Xiaohong Liu 0001, Xunchu Zhou, Tao Tan 0002, Guangtao Zhai |
ACM Multimedia | 6 |
| 2023 | StableVQA: A Deep No-Reference Quality Assessment Model for Video StabilityabstractVideo shakiness is an unpleasant distortion of User Generated Content (UGC) videos, which is usually caused by the unstable hold of cameras. In recent years, many video stabilization algorithms have been proposed, yet no specific and accurate metric enables comprehensively evaluating the stability of videos. Indeed, most existing quality assessment models evaluate video quality as a whole without specifically taking the subjective experience of video stability into consideration. Therefore, these models cannot measure the video stability explicitly and precisely when severe shakes are present. In addition, there is no large-scale video database in public that includes various degrees of shaky videos with the corresponding subjective scores available, which hinders the development of Video Quality Assessment for Stability (VQA-S). To this end, we build a new database named StableDB that contains 1,952 diversely-shaky UGC videos, where each video has a Mean Opinion Score (MOS) on the degree of video stability rated by 34 subjects. Moreover, we elaborately design a novel VQA-S model named StableVQA, which consists of three feature extractors to acquire the optical flow, semantic, and blur features respectively, and a regression layer to predict the final stability score. Extensive experiments demonstrate that the StableVQA achieves a higher correlation with subjective opinions than the existing VQA-S models and generic VQA models. The database and codes are available at https://github.com/QMME/StableVQA. Tengchuan Kou, Xiaohong Liu 0001, Wei Sun 0029, Jun Jia, Xiongkuo Min, Guangtao Zhai |
ACM Multimedia | 6 |
| 2023 | Perceptual Quality Assessment for Video Frame InterpolationabstractThe quality of frames is significant for both research and application of video frame interpolation (VFI). In recent VFI studies, the methods of full-reference image quality assessment have generally been used to evaluate the quality of VFI frames. However, high frame rate reference videos, necessities for the full-reference methods, are difficult to obtain in most applications of VFI. To evaluate the quality of VFI frames without reference videos, a no-reference perceptual quality assessment method is proposed in this paper. This method is more compatible with VFI application and the evaluation scores from it are consistent with human subjective opinions. A new quality assessment dataset for VFI was constructed through subjective experiments firstly, to assess the opinion scores of interpolated frames. The dataset was created from triplets of frames extracted from high-quality videos using 9 state-of-the-art VFI algorithms. The proposed method evaluates the perceptual coherence of frames incorporating the original pair of VFI inputs. Specifically, the method applies a triplet network architecture, including three parallel feature pipelines, to extract the deep perceptual features of the interpolated frame as well as the original pair of frames. Coherence similarities of the two-way parallel features are jointly calculated and optimized as a perceptual metric. In the experiments, both full-reference and no-reference quality assessment methods were tested on the new quality dataset. The results show that the proposed method achieves the best performance among all compared quality assessment methods on the dataset. Jinliang Han, Xiongkuo Min, Jun Jia, Lei Sun 0009, Zuowei Cao, Yonglin Luo, Guangtao Zhai |
VCIP | 8 |
| 2023 | Split-Conv: A Resource-efficient Compression Method for Image Quality Assessment ModelsabstractBlind Image Quality Assessment (BIQA) models based on deep neural networks (DNNs) have achieved state-of-the-art performance recently. However, the heavyweight architecture makes them hard to deploy on resource-constrained devices. Filter pruning is one of the most effective ways to compress the DNN model. However, most pruning methods need complete retraining after pruning which is too resource-consuming for IQA models with complex training processes. In this paper, we propose a resource-efficient structural pruning method for IQA models called Split-Conv, where the model only needs to be retrained on small IQA databases. Specifically, we split convolutional layers into sub-convolutional kernels and decorators, which can measure the importance of convolutional channels more precisely, thus reducing the requirement for retraining conditions. The experiments on several IQA models demonstrate the effectiveness of our IQA model pruning approach. For instance, DBCNN, StairIQA, and HyperIQA’s performances can be kept at the baseline level when 50% parameters are pruned with the model being retrained on the authentically distort IQA database containing only 586 images, and the performances of StairIQA and HyperIQA on KonIQ-10k database are still acceptable after 80% of the parameters have been pruned. Besides, due to Split-Conv’s capability to identify the important, useless, and harmful filters, the performances of some BIQA models can even be boosted after model compression with a proper pruning ratio. Xiongkuo Min, Guangtao Zhai |
VCIP | 3 |
| 2023 | Audio-visual aligned saliency model for omnidirectional video with implicit neural representation learning
Dandan Zhu 0001, Xuan Shao, Kaiwei Zhang, Xiongkuo Min, Guangtao Zhai, Xiaokang Yang 0001 |
Appl. Intell. | 5 |
| 2023 | Human attention based movie summarization: Dataset and baseline modelabstractA movie summarization model can automatically edit a condensed version of a movie by selecting keyframes . Some previous works have proposed some movie summarizers based on traditional methods or recent neural networks and achieved some progress. Despite the demonstrated successes, there are some limitations: (1) previous works mainly resort to hand-crafted heuristics and most of them are unsupervised; (2) currently there is no publicly suitable dataset available for the supervised movie summarization; (3) existing works only focus on the movies themselves while neglecting the audiences, who have the most to say in which part of the movie is more attractive. To break through the aforementioned limitations, we establish a movie summarization dataset Movie50 and propose a novel human attention based annotation pipeline. Furthermore, we propose the A/V-MSNet, an audiovisual neural network that takes advantage of spatio-temporal visual and auditory information to better simulate human attention as well as exploit more plentiful information. The network is designed, trained end-to-end, and evaluated on the public dataset and our dataset. Extensive experiments demonstrate the superiority of the proposed method. Defang Zhao, Dandan Zhu 0001, Xiongkuo Min, Jiaomin Yue, Kaiwei Zhang, Qiangqiang Zhou, Guangtao Zhai, Xiaokang Yang 0001 |
Neurocomputing | 7 |
| 2023 | Decoupled dynamic group equivariant filter for saliency prediction on omnidirectional image
Dandan Zhu 0001, Kaiwei Zhang, Qiangqiang Zhou, Xiongkuo Min, Guangtao Zhai, Xiaokang Yang 0001 |
Neurocomputing | 6 |
| 2023 | Application of QR Code Watermarking and Encryption in the Protection of Data Privacy of Intelligent Mouth-Opening TrainerabstractQuick response (QR) codes are widely used in offline to online channels to transfer information from promotional materials to mobile devices. Self-service medical equipment can record the data of each test, so the use of QR codes can realize the data exchange between patients and doctors, medical institutions, and self-service medical equipment, and create a medical information platform for health files. However, since anyone can easily read the information in the QR code, it is not conducive to the protection of patient privacy. Therefore, we propose a QR code encryption and decryption model based on robust digital watermarking. We implement digital watermarking through the generative adversarial networks and increase the robustness of the watermark by adding noise to the model. At the same time, we encrypt and decrypt the QR code information through advanced encryption standards. Experimental results show that the proposed method can well protect the privacy of patients without affecting the data acquisition by patients and doctors. Jiannan Liu, Jun Jia, Dandan Zhu 0001, Guangtao Zhai |
IEEE Internet Things J. | 6 |
| 2023 | Computed Tomography and 3-D Face Scan Fusion for IoT-Based Diagnostic SolutionsabstractIn clinical diagnosis, multimodal medical image fusion is meaningful and necessary, for the reason that some diseases need to be diagnosed in combination with the situation of different tissues of patients. Spiral computed tomography (CT) realizes the high precision and smooth reconstruction of bone tissue, while it can not represent the color and texture information in soft tissue reconstruction with high accuracy. The face scan precisely records the color and shape of the maxillofacial region. The diagnosis of some diseases (like cavernous hemangioma and jaw deformity caused by idiopathic condylar resorption) needs to combine the information of maxillofacial soft tissue and bone, so it is of great significance to fuse spiral CT and face scan images. In this article, a novel intelligent Internet of Things scene is proposed: a multimodal medical images acquisition and fusion system, and by combining CT machine and face scan equipment, the CT and face scan of patients can be synchronously collected. Deep point neural networks are used to extract feature points and a threshold iterative closest point algorithm performing registration with deep feature points and contributed region segmentation is applied. Finally, high-precision fused modal data is output at the mobile terminal to facilitate diagnosis and analysis and improve the efficiency of doctor–patient communication. Quantitative experiments show promising results, and clinical experiments prove that our method enables patients and doctors to better understand the state of an illness and improves the efficiency of doctor–patient communication. Zhiyuan Qu, Hongyi Jing, Guo Bai, Zhongpai Gao, Leilei Yu, Guangtao Zhai, Chi Yang |
IEEE Internet Things J. | 7 |
| 2023 | Perceptual quality assessment for fine-grained compressed images
Wei Sun 0029, Wei Wu 0002, Xiongkuo Min, Guangtao Zhai |
J. Vis. Commun. Image Represent. | 6 |
| 2023 | Continual Learning for Blind Image Quality AssessmentabstractThe explosive growth of image data facilitates the fast development of image processing and computer vision methods for emerging visual applications, meanwhile introducing novel distortions to processed images. This poses a grand challenge to existing blind image quality assessment (BIQA) models, which are weak at adapting to subpopulation shift. Recent work suggests training BIQA methods on the combination of all available human-rated IQA datasets. However, this type of approach is not scalable to a large number of datasets and is cumbersome to incorporate a newly created dataset as well. In this paper, we formulate continual learning for BIQA, where a model learns continually from a stream of IQA datasets, building on what was learned from previously seen data. We first identify five desiderata in the continual setting with three criteria to quantify the prediction accuracy, plasticity, and stability, respectively. We then propose a simple yet effective continual learning method for BIQA. Specifically, based on a shared backbone network, we add a prediction head for a new dataset and enforce a regularizer to allow all prediction heads to evolve with new data while being resistant to catastrophic forgetting of old data. We compute the overall quality score by a weighted summation of predictions from all heads. Extensive experiments demonstrate the promise of the proposed continual learning method in comparison to standard training techniques for BIQA, with and without experience replay. We made the code publicly available at https://github.com/zwx8981/BIQA_CL. Weixia Zhang, Dingquan Li, Chao Ma 0004, Guangtao Zhai, Xiaokang Yang 0001, Kede Ma |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2023 | A weakly supervised inpainting-based learning method for lung CT image segmentation
Fangfang Lu, Zhihao Zhang 0005, Chi Tang, Hualin Bai, Guangtao Zhai, Jingjing Chen 0002, Xiaoxin Wu 0005 |
Pattern Recognit. | 6 |
| 2023 | Surprise-based JND estimation for perceptual quantization in H.265/HEVC codecs
Hongkui Wang, Li Yu 0003, Hailang Yang, Haibing Yin, Guangtao Zhai, Tianzong Li, Zhuo Kuang |
Signal Process. Image Commun. | 6 |
| 2023 | A Deep Learning-Based Multidimensional Aesthetic Quality Assessment Method for Mobile Game ImagesabstractMobile games have played an increasingly significant role in people's leisure lives in recent years, thanks to the fast expansion of the gaming industry and the widespread use of mobile devices. The aesthetic quality of game pictures is a very important factor that attracts users' interest. However, evaluating the aesthetic quality of mobile game pictures is difficult since the painting styles of games vary greatly and the evaluation criteria are also diversified. In this article, we propose a multitask deep learning-based method, which is able to predict the aesthetic quality of mobile game images in multiple dimensions. The proposed model consists of two modules, a feature extraction module and a quality regression module. We extract quality-aware features from intermediate layers of the deep convolution neural network and then incorporate them into the final feature representation in the feature extraction module, allowing the model to fully use visual information from low to high levels. The quality regression module uses fully connected layers to map quality-aware features into quality scores across multiple dimensions. The multidimensional aesthetic quality scores are trained using a multitask learning approach, in which quality-aware features are shared across multiple dimensional quality prediction tasks. Finally, several key factors which help the proposed model perform better are analyzed. The experimental results indicate that our proposed method not only achieves the greatest performance on mobile game images, but also is applicable to natural scene images. Tao Wang 0078, Wei Sun 0029, Wei Wu 0002, Ying Chen 0011, Xiongkuo Min, Wei Lu 0021, Guangtao Zhai |
IEEE Trans. Games | 8 |
| 2023 | Image Quality Score Distribution Prediction via Alpha Stable ModelabstractBased on potentially subjective and diverse image quality scores given by a group of subjects, we propose to predict the distribution of image quality scores rather than the mean opinion score (MOS) of image quality. Therefore, in this paper, we use an alpha stable model to parameterize the image quality score distribution (IQSD), and propose an objective method to predict the alpha-stable-model-based IQSD. First, the LIVE database is re-recorded. Specifically, we invite a large group of subjects (187 valid subjects) to evaluate the quality of all 808 images in the LIVE database, with their scores forming reliable IQSDs. All images in the LIVE database and their collected subjective quality scores form a new image quality assessment database, named the SJTU IQSD database. We then propose a framework and algorithm to predict the alpha-stable-model-based IQSD, in which quality features are extracted from the structural and natural statistical information of each image, and support vector regressors are trained to predict the alpha stable model parameters. Experiments carried out on the SJTU IQSD database verify the feasibility of using the alpha stable model to describe the IQSD, and the experimental results show that the alpha-stable-model-based IQSD can reflect a large amount of subjective information on image quality. We also prove that the objective alpha-stable-model-based IQSD prediction method is effective. The code and the SJTU IQSD database can be downloaded at ‘https://github.com/YixuanGao98/Image-Quality-Score-Distribution-Prediction-via-Alpha-Stable-Model.git’. Xiongkuo Min, Wenhan Zhu, Xiao-Ping Zhang 0002, Guangtao Zhai |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2023 | Toward a No-Reference Quality Metric for Camera-Captured ImagesabstractExisting no-reference (NR) image quality assessment (IQA) metrics are still not convincing for evaluating the quality of the camera-captured images. Toward tackling this issue, we, in this article, establish a novel NR quality metric for quantifying the quality of the camera-captured images reliably. Since the image quality is hierarchically perceived from the low-level preliminary visual perception to the high-level semantic comprehension in the human brain, in our proposed metric, we characterize the image quality by exploiting both the low-level image properties and the high-level semantics of the image. Specifically, we extract a series of low-level features to characterize the fundamental image properties, including the brightness, saturation, contrast, noiseness, sharpness, and naturalness, which are highly indicative of the camera-captured image quality. Correspondingly, the high-level features are designed to characterize the semantics of the image. The low-level and high-level perceptual features play complementary roles in measuring the image quality. To infer the image quality, we employ the support vector regression (SVR) to map all the informative features to a single quality score. Thorough tests conducted on two standard camera-captured image databases demonstrate the effectiveness of the proposed quality metric in assessing the image quality and its superiority over the state-of-the-art NR quality metrics. The source code of the proposed metric for camera-captured images is released at https://github.com/YT2015?tab=repositories. Runze Hu, Yutao Liu 0002, Ke Gu 0001, Xiongkuo Min, Guangtao Zhai |
IEEE Trans. Cybern. | 5 |
| 2023 | Implicit Neural Representation Learning for Hyperspectral Image Super-ResolutionabstractHyperspectral image (HSI) super-resolution (SR) without additional auxiliary image remains a constant challenge due to its high-dimensional spectral patterns, where learning an effective spatial and spectral representation is a fundamental issue. Recently, implicit neural representations (INRs) are making strides as a novel and effective representation, especially in the reconstruction task. Therefore, in this work, we propose a novel HSI reconstruction model based on INR which represents HSI by a continuous function mapping a spatial coordinate to its corresponding spectral radiance values. In particular, as a specific implementation of INR, the parameters of the parametric model are predicted by a hypernetwork that operates on feature extraction using a convolution network. It makes the continuous functions map the spatial coordinates to pixel values in a content-aware manner. Moreover, periodic spatial encoding is deeply integrated with the reconstruction procedure, which makes our model capable of recovering more high-frequency details. To verify the efficacy of our model, we conduct experiments on three HSI datasets (CAVE, NUS, and NTIRE2018). Experimental results show that the proposed model can achieve competitive reconstruction performance in comparison with the state-of-the-art methods. In addition, we provide an ablation study on the effect of individual components of our model. We hope this article could serve as a potent reference for future research. Kaiwei Zhang, Dandan Zhu 0001, Xiongkuo Min, Guangtao Zhai |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2023 | Attention-Guided Neural Networks for Full-Reference and No-Reference Audio-Visual Quality AssessmentabstractWith the popularity of mobile Internet, audio and video (A/V) have become the main way for people to entertain and socialize daily. However, in order to reduce the cost of media storage and transmission, A/V signals will be compressed by service providers before they are transmitted to end-users, which inevitably causes distortions in the A/V signals and degrades the end-user's Quality of Experience (QoE). This motivates us to research the objective audio-visual quality assessment (AVQA). In the field of AVQA, most previous works only focus on single-mode audio or visual signals, which ignores that the perceptual quality of users depends on both audio and video signals. Therefore, we propose an objective AVQA architecture for multi-mode signals based on attentional neural networks. Specifically, we first utilize an attention prediction model to extract the salient regions of video frames. Then, a pre-trained convolutional neural network is used to extract short-time features of the salient regions and the corresponding audio signals. Next, the short-time features are fed into Gated Recurrent Unit (GRU) networks to model the temporal relationship between adjacent frames. Finally, the fully connected layers are utilized to fuse the temporal related features of A/V signals modeled by the GRU network into the final quality score. The proposed architecture is flexible and can be applied to both full-reference and no-reference AVQA. Experimental results on the LIVE-SJTU Database and UnB-AVC Database demonstrate that our model outperforms the state-of-the-art AVQA methods. The code of the proposed method will be publicly available to promote the development of the field of AVQA. Yuqin Cao, Xiongkuo Min, Wei Sun 0029, Guangtao Zhai |
IEEE Trans. Image Process. | 4 |
| 2023 | Subjective and Objective Audio-Visual Quality Assessment for User Generated ContentabstractIn recent years, User Generated Content (UGC) has grown dramatically in video sharing applications. It is necessary for service-providers to use video quality assessment (VQA) to monitor and control users' Quality of Experience when watching UGC videos. However, most existing UGC VQA studies only focus on the visual distortions of videos, ignoring that the perceptual quality also depends on the accompanying audio signals. In this paper, we conduct a comprehensive study on UGC audio-visual quality assessment (AVQA) from both subjective and objective perspectives. Specially, we construct the first UGC AVQA database named SJTU-UAV database, which includes 520 in-the-wild UGC audio and video (A/V) sequences collected from the YFCC100m database. A subjective AVQA experiment is conducted on the database to obtain the mean opinion scores (MOSs) of the A/V sequences. To demonstrate the content diversity of the SJTU-UAV database, we give a detailed analysis of the SJTU-UAV database as well as other two synthetically-distorted AVQA databases and one authentically-distorted VQA database, from both the audio and video aspects. Then, to facilitate the development of AVQA fields, we construct a benchmark of AVQA models on the proposed SJTU-UAV database and other two AVQA databases, of which the benchmark models consist of AVQA models designed for synthetically distorted A/V sequences and AVQA models built through combining the popular VQA methods and audio features via support vector regressor (SVR). Finally, considering benchmark AVQA models perform poorly in assessing in-the-wild UGC videos, we further propose an effective AVQA model via jointly learning quality-aware audio and visual feature representations in the temporal domain, which is seldom investigated by existing AVQA models. Our proposed model outperforms the aforementioned benchmark AVQA models on the SJTU-UAV database and two synthetically distorted AVQA databases. The SJTU-UAV database and the code of the proposed model will be released to facilitate further research. Yuqin Cao, Xiongkuo Min, Wei Sun 0029, Guangtao Zhai |
IEEE Trans. Image Process. | 4 |
| 2023 | CLSA: A Contrastive Learning Framework With Selective Aggregation for Video RescalingabstractVideo rescaling has recently drawn extensive attention for its practical applications such as video compression. Compared to video super-resolution, which focuses on upscaling bicubic-downscaled videos, video rescaling methods jointly optimize a downscaler and a upscaler. However, the inevitable loss of information during downscaling makes the upscaling procedure still ill-posed. Furthermore, the network architecture of previous methods mostly relies on convolution to aggregate information within local regions, which cannot effectively capture the relationship between distant locations. To address the above two issues, we propose a unified video rescaling framework by introducing the following designs. First, we propose to regularize the information of the downscaled videos via a contrastive learning framework, where, particularly, hard negative samples for learning are synthesized online. With this auxiliary contrastive learning objective, the downscaler tends to retain more information that benefits the upscaler. Second, we present a selective global aggregation module (SGAM) to efficiently capture long-range redundancy in high-resolution videos, where only a few representative locations are adaptively selected to participate in the computationally-heavy self-attention (SA) operations. SGAM enjoys the efficiency of the sparse modeling scheme while preserving the global modeling capability of SA. We refer to the proposed framework as Contrastive Learning framework with Selective Aggregation (CLSA) for video rescaling. Comprehensive experimental results show that CLSA outperforms video rescaling and rescaling-based video compression methods on five datasets, achieving state-of-the-art performance. Yuan Tian 0017, Yichao Yan, Guangtao Zhai, Li Chen 0021 |
IEEE Trans. Image Process. | 3 |
| 2023 | GridDehazeNet+: An Enhanced Multi-Scale Network With Intra-Task Knowledge Transfer for Single Image DehazingabstractAdverse weather conditions such as haze can deteriorate the performance of autonomous driving and intelligent transport systems. As a potential remedy, we propose an enhanced multi-scale network, dubbed GridDehazeNet+, for single image dehazing. The proposed dehazing method does not rely on the Atmosphere Scattering Model (ASM), and an explanation as to why it is not necessarily performing the dimension reduction offered by this model is provided. GridDehazeNet+ consists of three modules: pre-processing, backbone, and post-processing. The trainable pre-processing module can generate learned inputs with better diversity and more pertinent features as compared to those derived inputs produced by hand-selected pre-processing methods. The backbone module implements multi-scale estimation with two major enhancements: 1) a novel grid structure that effectively alleviates the bottleneck issue via dense connections across different scales; 2) a spatial-channel attention block that can facilitate adaptive fusion by consolidating dehazing-relevant features. The post-processing module helps to reduce the artifacts in the final output. Due to domain shift, the model trained on synthetic data may not generalize well on real data. To address this issue, we shape the distribution of synthetic data to match that of real data, and use the resulting translated data to finetune our network. We also propose a novel intra-task knowledge transfer mechanism that can memorize and take advantage of synthetic domain knowledge to assist the learning process on the translated data. Experimental results demonstrate that the proposed method outperforms the state-of-the-art on several synthetic dehazing datasets, and achieves the superior performance on real-world hazy images after finetuning. Xiaohong Liu 0001, Zhihao Shi, Jun Chen 0005, Guangtao Zhai |
IEEE Trans. Intell. Transp. Syst. | 5 |
| 2023 | Blind Image Quality Assessment for Pathological Microscopic Image Under Screen and Immersion ScenariosabstractThe high-quality pathological microscopic images are essential for physicians or pathologists to make a correct diagnosis. Image quality assessment (IQA) can quantify the visual distortion degree of images and guide the imaging system to improve image quality, thus raising the quality of pathological microscopic images. Current IQA methods are not ideal for pathological microscopy images due to their specificity. In this paper, we present deep learning-based blind image quality assessment model with saliency block and patch block for pathological microscopic images. The saliency block and patch block can handle the local and global distortions, respectively. To better capture the area of interest of pathologists when viewing pathological images, the saliency block is fine-tuned by eye movement data of pathologists. The patch block can capture lots of global information strongly related to image quality via the interaction between different image patches from different positions. The performance of the developed model is validated by the home-made Pathological Microscopic Image Quality Database under Screen and Immersion Scenarios (PMIQD-SIS) and cross-validated by the five public datasets. The results of ablation experiments demonstrate the contribution of the added blocks. The dataset and the corresponding code are publicly available at: https://github.com/mikugyf/PMIQD-SIS. Yifei Guo, Menghan Hu, Xiongkuo Min, Yan Wang 0036, Guangtao Zhai, Xiao-Ping Zhang 0002, Xiaokang Yang 0001 |
IEEE Trans. Medical Imaging | 6 |
| 2023 | Develop Then Rival: A Human Vision-Inspired Framework for Superimposed Image DecompositionabstractA single superimposed image containing two image views causes visual confusion for both human vision and computer vision. Human vision needs a “develop-then-rival” process to decompose the superimposed image into two individual images, which effectively suppresses visual confusion. However, separating individual image views from a single superimposed image has been an important but challenging task in computer vision area for a long time. In this paper, we propose a human vision-inspired framework for single superimposed image decomposition. We first propose a network to simulate the development stage, which tries to understand and distinguish the semantic information of the two layers of a single superimposed image. To further simulate the rivalry activation/suppression process in human brains, we carefully design a rivalry stage, which incorporates the original mixed input (superimposed image), the activated visual information (outputs of the development stage) together, and then rivals to get images without ambiguity. Experimental results show that our novel framework effectively separates the superimposed images and significantly improves the performance with better output quality compared with state-of-the-art methods. The proposed method also achieves state-of-the-art results on related applications including single image reflection removal, single image rain removal, single image shadow removal, and illumination correction,etc., which validates the generalization of the framework. Huiyu Duan, Wei Shen 0002, Xiongkuo Min, Yuan Tian 0017, Jae-Hyun Jung, Xiaokang Yang 0001, Guangtao Zhai |
IEEE Trans. Multim. | 7 |
| 2023 | RIVIE: Robust Inherent Video Information EmbeddingabstractImagine an interesting situation when watching a movie, we can scan the screen using our smartphones to get some extra information about this movie such as the cast, the release date, the movie's homepage, etc. Our prospect is a world where each video contains invisible information that can be delivered to us through mobile devices with cameras. This paper proposes the first deep learning-based information hiding method for videos to achieve information transmission from screens to cameras. Compared with hiding information in single images, the methods for videos need to maintain visual quality in both spatial and temporal domains. Furthermore, the training of video models builds on a large video dataset, which needs much more computational resources than training models for images. To reduce the computational complexity, we propose to simulate data on-the-fly to generate simulated sequences from single images. Then, we use the simulated data to train a spatio-temporal generator that hides information in videos while maintaining visual quality. During training, a temporal loss function based on the simulated data is exploited to ensure the temporal consistency of generated videos. After embedding, we use a decoder to recover the hidden information. To simulate the imaging pipeline from screens to cameras in the real world, we insert a distortion network between the generator and decoder. The distortion network is based on differentiable 3D rendering to cover possible distortions introduced in the procedure of camera imaging. Experimental results show that the hidden information in videos can be extracted by cameras without impacting the visual quality. Our work can be applied to many fields, such as advertisement, entertainment, and education. Jun Jia, Zhongpai Gao, Dandan Zhu 0001, Xiongkuo Min, Menghan Hu, Guangtao Zhai |
IEEE Trans. Multim. | 6 |
| 2023 | Angel's Girl for Blind Painters: An Efficient Painting Navigation System Validated by Multimodal Evaluation ApproachabstractFor people who ardently love painting but unfortunately have visual impairments, holding a paintbrush to create a work is a very difficult task. People in this special group are eager to pick up the paintbrush, like Leonardo da Vinci, to create and make full use of their own talents. Therefore, to maximally bridge this gap, we propose a painting navigation system called “Angle’s Eyes” to assist blind people in artistic creation. The proposed system is composed of cognitive system and guidance system. The system adopts drawing board positioning based on QR code, brush navigation based on target detection and bush real-time positioning. Meanwhile, we design a simple yet efficient position information coding rule to remind the user of the current brush tip position. In addition, we design a criterion to efficiently judge whether the brush reaches the target or not. The numerous experiments are conducted to optimize and test the performance of the system. The results of real-world scenario experiments demonstrate that the developed system has great potential to help blind people with painting. This work also demonstrates that it is practicable for the blind people to feel the world through the brush in their hands. In the future, we plan to deploy “Angle’s Eyes” on the phone to make it more portable. The demo video of the proposed painting navigation system is available athttps://doi.org/10.6084/m9.figshare.9760004.v1. Menghan Hu, Qingli Li, Guangtao Zhai, Simon X. Yang, Xiao-Ping Zhang 0002, Xiaokang Yang 0001 |
IEEE Trans. Multim. | 5 |
| 2023 | Blind Image Quality Assessment via Cross-View ConsistencyabstractImage quality assessment (IQA) is very important for both end-users and service-providers since a high-quality image can significantly improve the user's quality of experience (QoE). Most existing blind image quality assessment (BIQA) models were developed for synthetically distorted images, however, they perform poorly on in-the-wild images, which are widely existed in various practical applications. In this paper, a BIQA model is proposed that consists of a desirable self-supervised feature learning approach to mitigate the data shortage problem and learn comprehensive feature representations, and a self-attention-based feature fusion module to introduce self-attention mechanism. We develop the image quality assessment model under the framework of contrastive learning with multi views. Since human visual system perceives signals through multiple channels, the most important visual information should exist among all views of the channels. So we design the cross-view consistent information mining (CVC-IM) module to extract compact mutual information between different views. Color information and pseudo-reference image (PRI) of different distortion types are employed to formulate rich feature embeddings and preserve the quality-aware fidelity of learned representations. We employ the Transformer as the self-attention-based architecture to integrate feature embeddings. Extensive experiments show that our model achieves remarkable image quality assessment results on in-the-wild IQA datasets. Yucheng Zhu, Wei Sun 0029, Xiongkuo Min, Guangtao Zhai, Xiaokang Yang 0001 |
IEEE Trans. Multim. | 5 |
| 2023 | Robust Mesh Representation Learning via Efficient Local Structure-Aware Anisotropic ConvolutionabstractMesh is a type of data structure commonly used for 3-D shapes. Representation learning for 3-D meshes is essential in many computer vision and graphics applications. The recent success of convolutional neural networks (CNNs) for structured data (e.g., images) suggests the value of adapting insights from CNN for 3-D shapes. However, 3-D shape data are irregular since each node's neighbors are unordered. Various graph neural networks for 3-D shapes have been developed with isotropic filters or predefined local coordinate systems to overcome the node inconsistency on graphs. However, isotropic filters or predefined local coordinate systems limit the representation power. In this article, we propose a local structure-aware anisotropic convolutional operation (LSA-Conv) that learns adaptive weighting matrices for each template's node according to its neighboring structure and performs shared anisotropic filters. In fact, the learnable weighting matrix is similar to the attention matrix in the random synthesizer-a new Transformer model for natural language processing (NLP). Since the learnable weighting matrices require large amounts of parameters for high-resolution 3-D shapes, we introduce a matrix factorization technique to notably reduce the parameter size, denoted as LSA-small. Furthermore, a residual connection with a linear transformation is introduced to improve the performance of our LSA-Conv. Comprehensive experiments demonstrate that our model produces significant improvement in 3-D shape reconstruction compared to state-of-the-art methods. Zhongpai Gao, Junchi Yan, Guangtao Zhai, Juyong Zhang, Xiaokang Yang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2023 | Feedforward and Feedback Modulations Based Foveated JND Estimation for ImagesabstractThe just noticeable difference (JND) reveals the key characteristic of visual perception, which has been widely used in many perception-based image and video applications. Nevertheless, the modulatory mechanism of the human visual system (HVS) has not been fully exploited in JND threshold estimation, which results in the existing JND models not being accurate enough. In this article, by analyzing the feedforward and feedback modulatory behaviors in the HVS, an enhanced foveated JND (FJND) estimation model is proposed considering modulatory effects and masking effects in visual perception. The contributions of this article are mainly twofold. On the one hand, by analyzing the modulatory behaviors in the HVS, the modulatory mechanism is incorporated into JND estimation and a hierarchical modulation-based JND estimation framework is proposed for the first time. On the other hand, according to the response characteristics of visual neurons, modulatory effects on visual sensitivity are formulated as several modulatory factors to modulate the estimated JND threshold properly. Compared with existing models, the proposed model is developed in view of not only the masking effects but also the modulatory effects, which makes our model more consistent with the HVS. For different complex input images, experimental results show that the proposed FJND model tolerates more distortion at the same perceptual quality in comparison with other existing models. Haibing Yin, Hongkui Wang, Li Yu 0003, Junhui Liang, Guangtao Zhai |
ACM Trans. Multim. Comput. Commun. Appl. | 5 |
| 2023 | Toward Visual Behavior and Attention Understanding for Augmented 360 Degree VideosabstractAugmented reality (AR) overlays digital content onto reality. In an AR system, correct and precise estimations of user visual fixations and head movements can enhance the quality of experience by allocating more computational resources for analyzing, rendering, and 3D registration on the areas of interest. However, there is inadequate research to help in understanding the visual explorations of the users when using an AR system or modeling AR visual attention. To bridge the gap between the saliency prediction on real-world scenes and on scenes augmented by virtual information, we construct the ARVR saliency dataset. The virtual reality (VR) technique is employed to simulate the real-world. Annotations of object recognition and tracking as augmented contents are blended into omnidirectional videos. The saliency annotations of head and eye movements for both original and augmented videos are collected and together constitute the ARVR dataset. We also design a model that is capable of solving the saliency prediction problem in AR. Local block images are extracted to simulate the viewport and offset the projection distortion. Conspicuous visual cues in the local block images are extracted to constitute the spatial features. The optical flow information is estimated as an important temporal feature. We also consider the interplay between virtual information and reality. The composition of the augmentation information is distinguished, and the joint effects of adversarial augmentation and complementary augmentation are estimated. The Markov chain is constructed with block images as graph nodes. In the determination of the edge weights, both the characteristics of the viewing behaviors and the visual saliency mechanisms are considered. The order of importance for block images is estimated through the state of equilibrium of the Markov chain. Extensive experiments are conducted to demonstrate the effectiveness of the proposed method. Yucheng Zhu, Xiongkuo Min, Dandan Zhu 0001, Guangtao Zhai, Xiaokang Yang 0001, Wenjun Zhang 0001, Ke Gu 0001, Jiantao Zhou 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2023 | A Novel Lightweight Audio-visual Saliency Model for VideosabstractAudio information has not been considered an important factor in visual attention models regardless of many psychological studies that have shown the importance of audio information in the human visual perception system. Since existing visual attention models only utilize visual information, their performance is limited but also requires high-computational complexity due to the limited information available. To overcome these problems, we propose a lightweight audio-visual saliency (LAVS) model for video sequences. To the best of our knowledge, this article is the first trial to utilize audio cues for an efficient deep-learning model for the video saliency estimation. First, spatial-temporal visual features are extracted by the lightweight receptive field block (RFB) with the bidirectional ConvLSTM units. Then, audio features are extracted by using an improved lightweight environment sound classification model. Subsequently, deep canonical correlation analysis (DCCA) aims at capturing the correspondence between audio and spatial-temporal visual features, thus obtaining a spatial-temporal auditory saliency. Lastly, the spatial-temporal visual and auditory saliency are fused to obtain the audio-visual saliency map. Extensive comparative experiments and ablation studies validate the performance of the LAVS model in terms of effectiveness and complexity. Dandan Zhu 0001, Xuan Shao, Qiangqiang Zhou, Xiongkuo Min, Guangtao Zhai, Xiaokang Yang 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 5 |
| 2022 | Learning Invisible Markers for Hidden Codes in Offline-to-online PhotographyabstractQR (quick response) codes are widely used as an offline-to-online channel to convey information (e.g., links) from publicity materials (e.g., display and print) to mobile devices. However, QR codes are not favorable for taking up valuable space of publicity materials. Recent works propose invisible codes/hyperlinks that can convey hidden information from offline to online. However, they require markers to locate invisible codes, which fails the purpose of invisible codes to be visible because of the markers. This paper proposes a novel invisible information hiding architecture for display/print-camera scenarios, consisting of hiding, locating, correcting, and recovery, where invisible markers are learned to make hidden codes truly invisible. We hide information in a sub-image rather than the entire image and include a localization module in the end-to-end framework. To achieve both high visual quality and high recovering robustness, an effective multi-stage training strategy is proposed. The experimental results show that the proposed method outperforms the state-of-the-art information hiding methods in both visual quality and robustness. In addition, the automatic localization of hidden codes significantly reduces the time of manually correcting geometric distortions for photos, which is a revolutionary innovation for information hiding in mobile applications. Jun Jia, Zhongpai Gao, Dandan Zhu 0001, Xiongkuo Min, Guangtao Zhai, Xiaokang Yang 0001 |
CVPR | 5 |
| 2022 | End-to-End Human-Gaze-Target Detection with TransformersabstractIn this paper, we propose an effective and efficient method for Human-Gaze-Target (HGT) detection, i.e., gaze following. Current approaches decouple the HGT detection task into separate branches of salient object detection and human gaze prediction, employing a two-stage framework where human head locations must first be detected and then be fed into the next gaze target prediction sub-network. In contrast, we redefine the HGT detection task as detecting human head locations and their gaze targets, simultaneously. By this way, our method, named Human-Gaze-Target detection TRansformer or HGTTR, streamlines the HGT detection pipeline by eliminating all other additional components. HGTTR reasons about the relations of salient objects and human gaze from the global image context. Moreover, unlike existing two-stage methods that require human head locations as input and can predict only one human's gaze target at a time, HGTTR can directly predict the locations of all people and their gaze targets at one time in an end-to-end manner. The effectiveness and robustness of our proposed method are verified with extensive experiments on the two standard benchmark datasets, GazeFollowing and VideoAttentionTarget. Without bells and whistles, HGTTR outperforms existing state-of-the-art methods by large margins (6.4 mAP gain on GazeFollowing and 10.3 mAP gain on VideoAttentionTarget) with a much simpler architecture. Danyang Tu, Xiongkuo Min, Huiyu Duan, Guodong Guo, Guangtao Zhai, Wei Shen 0002 |
CVPR | 5 |
| 2022 | Iwin: Human-Object Interaction Detection via Transformer with Irregular Windows
Danyang Tu, Xiongkuo Min, Huiyu Duan, Guodong Guo, Guangtao Zhai, Wei Shen 0002 |
ECCV (4) | 5 |
| 2022 | A Unified Two-Stage Model for Separating Superimposed ImagesabstractA single superimposed image containing two image views causes visual confusion for both human vision and computer vision. Human vision needs a "develop-then-rival" process to decompose the superimposed image into two individual images, which effectively suppresses visual confusion. In this paper, we propose a human vision-inspired framework for separating superimposed images. We first propose a network to simulate the development stage, which tries to understand and distinguish the semantic information of the two layers of a single superimposed image. To further simulate the rivalry activation/suppression process in human brains, we carefully design a rivalry stage, which incorporates the original mixed input (superimposed image), the activated visual information (outputs of the development stage) together, and then rivals to get images without ambiguity. Experimental results show that our novel framework effectively separates the superimposed images and significantly improves the performance with better output quality compared with state-of-the-art methods. Huiyu Duan, Xiongkuo Min, Wei Shen 0002, Guangtao Zhai |
ICASSP | 4 |
| 2022 | How Sound Affects Visual Attention in Omnidirectional VideosabstractIn this paper, we propose a new audio-visual attention dataset that records eye movement for omnidirectional videos with and without sound. We classify the videos into three types according to the number of salient objects and sound sources and analyze the impact of sound on visual attention distribution and inter-observer consistency of viewing area in different types of videos. From the quantitative and qualitative analysis, we find that visual attention will be drawn to and concentrated on the sound source with the presence of sound, especially when there are several visually salient objects and only one sound source. Also, the sound will enhance the consistency of observation areas among viewers to some extent. For more investigations on the impact of sound on visual attention and prospective audio-visual saliency model, we still need further study. Guangtao Zhai, Yucheng Zhu, Jun Zhou 0007, Xiao-Ping Zhang 0002 |
ICIP | 2 |
| 2022 | Enhanced Deep Animation Video InterpolationabstractExisting learning-based frame interpolation algorithms extract consecutive frames from high-speed natural videos to train the model. Compared to natural videos, cartoon videos are usually in a low frame rate. Besides, the motion between consecutive cartoon frames is typically nonlinear, which breaks the linear motion assumption of interpolation algorithms. Thus, it is unsuitable for generating a training set directly from cartoon videos. For better adapting frame interpolation algorithms from nature video to animation video, we present AutoFI, a simple and effective method to automatically render training data for deep animation video interpolation. AutoFI takes a layered architecture to render synthetic data, which ensures the assumption of linear motion. Experimental results show that AutoFI performs favorably in training both DAIN and ANIN. However, most frame interpolation algorithms will still fail in error-prone areas, such as fast motion or large occlusion. Besides AutoFI, we also propose a plug-and-play sketch-based post-processing module, named SktFI, to refine the final results using user-provided sketches manually. With AutoFI and SktFI, the interpolated animation frames show high perceptual quality. Wang Shen, Wenbo Bao, Guangtao Zhai, Li Chen 0021 |
ICIP | 4 |
| 2022 | Surveillance Video Quality Assessment Based on Quality Related RetrainingabstractSurveillance videos have been widely used in many vision-based systems, supporting intelligent tasks such as object detection and tracking. However, the quality of surveillance videos suffers from poor weather conditions and inevitable compression error, which may have a negative influence on the performance of such tasks. Therefore, accurately distinguishing distortions and predicting severity levels are crucial. In this paper, we propose a quality related retraining framework as well as a no-reference (NR) multi-task video quality assessment (VQA) model to tackle the challenge of surveillance videos quality assessment. The quality related retraining framework operates on a synthetic VQA database. The proposed NR VQA method utilizes both spatial and temporal information by using ResNet50 and SlowFast. Then multiple distortion detection heads are applied to predict the severity levels for corresponding distortions. The experimental results show that the proposed method gains competitive performance on the Video Surveillance Quality Assessment Dataset (VSQuAD). The ablation study further confirms the contributions of the quality related retraining framework, spatial information, and temporal information. Wei Lu 0021, Wei Sun 0029, Xiongkuo Min, Tao Wang 0078, Guangtao Zhai |
ICIP | 6 |
| 2022 | Learning a Blind Quality Evaluator for UGC Videos in Perceptually Relevant DomainsabstractThe absence of reference videos with pristine quality and the complexity of authentic distortions pose great challenges for blind video quality assessment (BVQA) towards user-generated content (UGC) videos. Although it is straightfor-ward to leverage transfer learning techniques to learn effective BVQA models, it is nontrivial to explore how to bridge the domain shifts for better video representation learning. In this work, we propose to transfer meaningful knowledge from perceptually relevant domains, i.e., image quality assessment (IQA) with authentic distortions and video classification with rich motion patterns. We develop a promising strategy to use both groups of data to learn the feature extractors. We train the proposed model on the target VQA databases using a mixed list-wise ranking loss function. Extensive experiments on six VQA databases demonstrate that our method performs very competitively under both individual database and mixed database training settings. Codes and models are available at https://github.com/zwx8981/BVQA-2021. Bowen Li 0018, Weixia Zhang, Jiu Jiang, Guangtao Zhai, Xianpei Wang |
ICME | 5 |
| 2022 | Implicit Neural Representation Learning for Hyperspectral Image Super-ResolutionabstractHyperspectral image (HSI) super-resolution without additional auxiliary image remains a constant challenge due to its high-dimensional spectral patterns, where learning an effective spatial and spectral representation is a fundamental issue. Recently, Implicit Neural Representations (INRs) are making strides as a novel and effective representation, especially in the reconstruction task. Therefore, in this work, we propose a novel HSI reconstruction model based on INR which represents HSI by a continuous function mapping a spatial coordinate to its corresponding spectral radiance values. In particular, as a specific implementation of INR, the parameters of parametric model are predicted by a hypernetwork. It makes the continuous functions map the spatial coordinates to pixel values in a content-aware manner. Moreover, periodic spatial encoding are deeply integrated with the reconstruction procedure, which makes our model capable of recovering more high frequency details. Experimental results on CAVE, NUS, and NTIRE2018 datasets demonstrate the superiority of our model. Kaiwei Zhang, Dandan Zhu 0001, Xiongkuo Min, Guangtao Zhai |
ICME | 4 |
| 2022 | Human Attention Based Movie Summarization: Dataset and Baseline ModelabstractThe movie summarization model can automatically edit a condensed and succinct version of the movie by selecting the keyframes. Previous works mainly resort to hand-crafted heuristics and most of them are unsupervised. Supervised movie summarization is a new research field and, there is currently no publicly suitable dataset available. Moreover, existing works only focus on the movies themselves while neglecting the audiences, who have the most say in which part of the movie is more attractive. To deal with the aforementioned limitations, we establish a human attention based movie summarization dataset Movie50. Specifically, we explore the human attention variations when watching videos and have the following findings: (1) The attention of humans is concentrated when watching keyframes. (2) The attention of humans is distracted when watching non-keyframes. Inspired by these findings, we collect the eye fixations of 20 participants when watching 50 movies and propose a novel human attention based annotation pipeline. In addition, we introduce A/V-MSNet, an audiovisual neural network that takes advantage of spatio-temporal visual and auditory information to better model human attention as well as exploit more plentiful information. Extensive experiments demonstrate the superiority of the proposed method. Defang Zhao, Dandan Zhu 0001, Xiongkuo Min, Jiaomin Yue, Kaiwei Zhang, Qiangqiang Zhou, Guangtao Zhai, Xiaokang Yang 0001 |
ICME | 7 |
| 2022 | A No-Reference Deep Learning Quality Assessment Method for Super-Resolution Images Based on Frequency MapsabstractTo support the application scenarios where high-resolution (HR) images are urgently needed, various single image super-resolution (SISR) algorithms are developed. However, SISR is an ill-posed inverse problem, which may bring artifacts like texture shift, blur, etc. to the reconstructed images, thus it is necessary to evaluate the quality of super-resolution images (SRIs). Note that most existing image quality assessment (IQA) methods were developed for synthetically distorted images, which may not work for SRIs since their distortions are more diverse and complicated. Therefore, in this paper, we propose a no-reference deep-learning image quality assessment method based on frequency maps because the artifacts caused by SISR algorithms are quite sensitive to frequency information. Specifically, we first obtain the high-frequency map (HM) and low-frequency map (LM) of SRI by using Sobel operator and piecewise smooth image approximation. Then, a two-stream network is employed to extract the quality-aware features of both frequency maps. Finally, the features are regressed into a single quality value using fully connected layers. The experimental results show that our method outperforms all compared IQA models on the selected three super-resolution quality assessment (SRQA) databases. Wei Sun 0029, Xiongkuo Min, Wenhan Zhu, Tao Wang 0078, Wei Lu 0021, Guangtao Zhai |
ISCAS | 7 |
| 2022 | Saliency in Augmented RealityabstractWith the rapid development of multimedia technology, Augmented Reality (AR) has become a promising next-generation mobile platform. The primary theory underlying AR is human visual confusion, which allows users to perceive the real-world scenes and augmented contents (virtual-world scenes) simultaneously by superimposing them together. To achieve good Quality of Experience (QoE), it is important to understand the interaction between two scenarios, and harmoniously display AR contents. However, studies on how this superimposition will influence the human visual attention are lacking. Therefore, in this paper, we mainly analyze the interaction effect between background (BG) scenes and AR contents, and study the saliency prediction problem in AR. Specifically, we first construct a Saliency in AR Dataset (SARD), which contains 450 BG images, 450 AR images, as well as 1350 superimposed images generated by superimposing BG and AR images in pair with three mixing levels. A large-scale eye-tracking experiment among 60 subjects is conducted to collect eye movement data. To better predict the saliency in AR, we propose a vector quantized saliency prediction method and generalize it for AR saliency prediction. For comparison, three benchmark methods are proposed and evaluated together with our proposed method on our SARD. Experimental results demonstrate the superiority of our proposed method on both of the common saliency prediction problem and the AR saliency prediction problem over benchmark methods. Our dataset and code are available at: https://github.com/DuanHuiyu/ARSaliency. Huiyu Duan, Wei Shen 0002, Xiongkuo Min, Danyang Tu, Jing Li 0026, Guangtao Zhai |
ACM Multimedia | 6 |
| 2022 | Image Quality Assessment: From Mean Opinion Score to Opinion Score DistributionabstractRecently, many methods have been proposed to predict the image quality which is generally described by the mean opinion score (MOS) of all subjective ratings given to an image. However, few efforts focus on predicting the opinion score distribution of the image quality ratings. In fact, the opinion score distribution reflecting subjective diversity, uncertainty, etc., can provide more subjective information about the image quality than a single MOS, which is worthy of in-depth study. In this paper, we propose a convolutional neural network based on fuzzy theory to predict the opinion score distribution of image quality. The proposed method consists of three main steps: feature extraction, feature fuzzification and fuzzy transfer. Specifically, we first use the pre-trained VGG16 without fully-connected layers to extract image features. Then, the extracted features are fuzzified by fuzzy theory, which is used to model epistemic uncertainty in the process of feature extraction. Finally, a fuzzy transfer network is used to predict the opinion score distribution of image quality by learning the mapping from epistemic uncertainty to the uncertainty existing in the image quality ratings. In addition, a new loss function is designed based on the subjective uncertainty of the opinion score distribution. Extensive experimental results prove the superior prediction performance of our proposed method. Xiongkuo Min, Yucheng Zhu, Jing Li 0026, Xiao-Ping Zhang 0002, Guangtao Zhai |
ACM Multimedia | 6 |
| 2022 | Skeleton2Humanoid: Animating Simulated Characters for Physically-plausible Motion In-betweeningabstractHuman motion synthesis is a long-standing problem with various applications in digital twins and the Metaverse. However, modern deep learning based motion synthesis approaches barely consider the physical plausibility of synthesized motions and consequently they usually produce unrealistic human motions. In order to solve this problem, we propose a system "Skeleton2Humanoid" which performs physics-oriented motion correction at test time by regularizing synthesized skeleton motions in a physics simulator. Concretely, our system consists of three sequential stages: (I) test time motion synthesis network adaptation, (II) skeleton to humanoid matching and (III) motion imitation based on reinforcement learning (RL). Stage I introduces a test time adaptation strategy, which improves the physical plausibility of synthesized human skeleton motions by optimizing skeleton joint locations. Stage II performs an analytical inverse kinematics strategy, which converts the optimized human skeleton motions to humanoid robot motions in a physics simulator, then the converted humanoid robot motions can be served as reference motions for the RL policy to imitate. Stage III introduces a curriculum residual force control policy, which drives the humanoid robot to mimic complex converted reference motions in accordance with the physical law. We verify our system on a typical human motion synthesis task, motion-in-betweening. Experiments on the challenging LaFAN1 dataset show our system can outperform prior methods significantly in terms of both physical plausibility and accuracy. Code will be released for research purposes at: https://github.com/michaelliyunhao/Skeleton2Humanoid. Zhenbo Yu, Yucheng Zhu, Bingbing Ni, Guangtao Zhai, Wei Shen 0002 |
ACM Multimedia | 5 |
| 2022 | A Deep Learning based No-reference Quality Assessment Model for UGC VideosabstractQuality assessment for User Generated Content (UGC) videos plays an important role in ensuring the viewing experience of end-users. Previous UGC video quality assessment (VQA) studies either use the image recognition model or the image quality assessment (IQA) models to extract frame-level features of UGC videos for quality regression, which are regarded as the sub-optimal solutions because of the domain shifts between these tasks and the UGC VQA task. In this paper, we propose a very simple but effective UGC VQA model, which tries to address this problem by training an end-to-end spatial feature extraction network to directly learn the quality-aware spatial feature representation from raw pixels of the video frames. We also extract the motion features to measure the temporal-related distortions that the spatial features cannot model. The proposed model utilizes very sparse frames to extract spatial features and dense frames (i.e. the video chunk) with a very low spatial resolution to extract motion features, which thereby has low computational complexity. With the better quality-aware features, we only use the simple multilayer perception layer (MLP) network to regress them into the chunk-level quality scores, and then the temporal average pooling strategy is adopted to obtain the video-level quality score. We further introduce a multi-scale quality fusion strategy to solve the problem of VQA across different spatial resolutions, where the multi-scale weights are obtained from the contrast sensitivity function of the human visual system. The experimental results show that the proposed model achieves the best performance on five popular UGC VQA databases, which demonstrates the effectiveness of the proposed model. Wei Sun 0029, Xiongkuo Min, Wei Lu 0021, Guangtao Zhai |
ACM Multimedia | 4 |
| 2022 | A No-reference Quality Assessment Metric for Point Cloud Based on Captured Video SequencesabstractPoint cloud is one of the most widely used digital formats of 3D models, the visual quality of which is quite sensitive to distortions such as downsampling, noise, and compression. To tackle the challenge of point cloud quality assessment (PCQA) in scenarios where reference is not available, we propose a no-reference quality assessment metric for colored point cloud based on captured video sequences. Specifically, three video sequences are obtained by rotating the camera around the point cloud through three specific orbits. The video sequences not only contain the static views but also include the multi-frame temporal information, which greatly helps understand the human perception of the point clouds. Then we modify the ResNet3D as the feature extraction model to learn the correlation between the capture videos and corresponding subjective quality scores. The experimental results show that our method outperforms most of the state-of-the-art full-reference and no-reference PCQA metrics, which validates the effectiveness of the proposed method. Wei Sun 0029, Xiongkuo Min, Qiyuan Wang 0002, Guangtao Zhai |
MMSP | 9 |
| 2022 | A Full- Reference Quality Assessment Metric for Cartoon ImagesabstractCartoon images are illustrations that are typically drawn, sometimes animated, in an unrealistic or semi-realistic style, which are widely applied in multimedia services. However, in some post-production processes as well as transmission systems, cartoon images are inevitably distorted by wrong color arrangement and compression. Therefore, it is urgent to carry out image quality assessment (IQA) metrics to automatically predict the perceptual quality levels of distorted cartoon images. Nevertheless, the existing mainstream IQA metrics are specially developed for natural scene images (NSIs). Due to the statistical difference in structure and color aspects between cartoon images and NSIs, the scores predicted by such metrics are often inconsistent with the human vision system (HVS) for cartoon images. To further improve the performance of cartoon image quality assessment (C-IQA) methods and provide guidance for practical applications, we propose a full-reference (FR) IQA method to tackle the challenge of C-IQA. Specifically, the proposed method extracts edge and texture features to analyze the structural error. Then the moment and entropy of various color spaces are computed to reflect color distortions. Then the features are regressed into quality scores with the assistance of a support vector regression (SVR) model. Experimental results show that our metric outperforms the mainstream FR-IQA metrics, which indicates that the proposed method is more capable of modeling the visual quality loss of cartoon images. Chunyi Li 0001, Wei Sun 0029, Xiongkuo Min, Guangtao Zhai |
MMSP | 5 |
| 2022 | Subjective Quality Assessment for Images Generated by Computer GraphicsabstractWith the development of rendering techniques, computer graphics generated images (CGIs) have been widely used in practical application scenarios such as architecture design, video games, simulators, movies, etc. Different from natural scene images (NSIs), the distortions of CGIs are usually caused by poor rending settings and limited computation resources. What's more, some CGIs may also suffer from compression distortions in transmission systems like cloud gaming and stream media. However, limited work has been put forward to tackle the problem of computer graphics generated images' quality assessment (CG-IQA). Therefore, in this paper, we establish a large-scale subjective CG-IQA database to deal with the challenge of CG-IQA tasks. We collect 25,454 in-the-wild CGIs through previous databases and personal collection. After data cleaning, we carefully select 1,200 CGIs to conduct the subjective experiment. Several popular no-reference image quality assessment (NR-IQA) methods are tested on our database. The experimental results show that the handcrafted-based methods achieve low correlation with subjective judgment and deep learning-based methods obtain relatively better performance. The current NR-IQA models are not suitable for CG-IQA tasks and more effective models are urgently needed. Tao Wang 0078, Wei Sun 0029, Xiongkuo Min, Wei Lu 0021, Guangtao Zhai |
MMSP | 6 |
| 2022 | CageNeRF: Cage-based Neural Radiance Field for Generalized 3D Deformation and AnimationabstractWhile implicit representations have achieved high-fidelity results in 3D rendering, it remains challenging to deforming and animating the implicit field. Existing works typically leverage data-dependent models as deformation priors, such as SMPL for human body animation. However, this dependency on category-specific priors limits them to generalize to other objects. To solve this problem, we propose a novel framework for deforming and animating the neural radiance field learned on \textit{arbitrary} objects. The key insight is that we introduce a cage-based representation as deformation prior, which is category-agnostic. Specifically, the deformation is performed based on an enclosing polygon mesh with sparsely defined vertices called \textit{cage} inside the rendering space, where each point is projected into a novel position based on the barycentric interpolation of the deformed cage vertices. In this way, we transform the cage into a generalized constraint, which is able to deform and animate arbitrary target objects while preserving geometry details. Based on extensive experiments, we demonstrate the effectiveness of our framework in the task of geometry editing, object animation and deformation transfer. Yicong Peng, Yichao Yan, Shengqi Liu, Yuhao Cheng, Shanyan Guan, Bowen Pan, Guangtao Zhai, Xiaokang Yang 0001 |
NeurIPS | 7 |
| 2022 | Video-based Human-Object Interaction Detection from Tubelet TokensabstractWe present a novel vision Transformer, named TUTOR, which is able to learn tubelet tokens, served as highly-abstracted spatial-temporal representations, for video-based human-object interaction (V-HOI) detection. The tubelet tokens structurize videos by agglomerating and linking semantically-related patch tokens along spatial and temporal domains, which enjoy two benefits: 1) Compactness: each token is learned by a selective attention mechanism to reduce redundant dependencies from others; 2) Expressiveness: each token is enabled to align with a semantic instance, i.e., an object or a human, thanks to agglomeration and linking. The effectiveness and efficiency of TUTOR are verified by extensive experiments. Results show our method outperforms existing works by large margins, with a relative mAP gain of $16.14\%$ on VidHOI and a 2 points gain on CAD-120 as well as a $4 \times$ speedup. Danyang Tu, Wei Sun 0029, Xiongkuo Min, Guangtao Zhai, Wei Shen 0002 |
NeurIPS | 4 |
| 2022 | Perceptual Attacks of No-Reference Image Quality Models with Human-in-the-LoopabstractNo-reference image quality assessment (NR-IQA) aims to quantify how humans perceive visual distortions of digital images without access to their undistorted references. NR-IQA models are extensively studied in computational vision, and are widely used for performance evaluation and perceptual optimization of man-made vision systems. Here we make one of the first attempts to examine the perceptual robustness of NR-IQA models. Under a Lagrangian formulation, we identify insightful connections of the proposed perceptual attack to previous beautiful ideas in computer vision and machine learning. We test one knowledge-driven and three data-driven NR-IQA methods under four full-reference IQA models (as approximations to human perception of just-noticeable differences). Through carefully designed psychophysical experiments, we find that all four NR-IQA models are vulnerable to the proposed perceptual attack. More interestingly, we observe that the generated counterexamples are not transferable, manifesting themselves as distinct design flows of respective NR-IQA methods. Source code are available at https://github.com/zwx8981/PerceptualAttack_BIQA. Weixia Zhang, Dingquan Li, Xiongkuo Min, Guangtao Zhai, Guodong Guo, Xiaokang Yang 0001, Kede Ma |
NeurIPS | 4 |
| 2022 | MRIQA: Subjective Method and Objective Model for Magnetic Resonance Image Quality AssessmentabstractMagnetic Resonance Imaging (MRI) is widely used for medical diagnosis, staging and follow-up of disease. However, MRI images may have artifacts due to various reasons such as patient movement or machine distortion, which may be unintentionally introduced during the procedure of medical image acquisition, processing, etc. These artifacts may affect the effectiveness of diagnosis or even cause false diagnosis. To solve this problem, we propose a general medical image quality assessment (MIQA) methodology, including subjective MIQA procedures and objective MIQA algorithms. We further apply this methodology to MRI images in this paper due to its widespread use in practical applications. We first establish a magnetic resonance imaging quality assessment (MRIQA) database, which contains 3809 MRI images. Then a subjective image quality assessment experiment is conducted by expert doctors according to the diagnostic value of these images, which split all MRI images into 1285 low quality images and 2524 high quality images. We then conduct a baseline deep learning experiment, and propose an attention based MIQANet model to automatically separate MRI images into high quality and low quality based on their diagnosis value. Our proposed method achieves a great quality assessment accuracy of 96.59%. The constructed MRIQA database and proposed MIQA model will be public available to further promote medical IQA research. Fang Liu 0001, Huiyu Duan, Xiongkuo Min, Guangtao Zhai |
VCIP | 7 |
| 2022 | Portable Eye Movement Feature Collection Device for Children with AutismabstractEye movement data has become an important char-acterization in the analysis of children with autism spectrum disorder (ASD). Current eye movement measurement meth-ods require specialized expensive equipment, calibration, and trained personnel, limiting their use in general ASD screening, especially in resource-scarce environments. Therefore, collecting eye movement features based on the standard RGB camera of a mobile phone or tablet has many advantages over professional equipment. The system design is based on the Android tablet design, and the screen is divided into two parts to display the normal children and the ASD children paintings. The eye movement data of children is obtained through the front camera, so as to provide data support for future data analysis. Taking the different cooperation degrees of children into account, two collection modes are designed: 1) directly displaying the stimuli in a loop (image mode); and 2) providing the background video interspersed with the stimulus display (video mode). The demo video of the proposed system is available at: https://doi.org/10.6084/m9.figshare.21346806.v1. Xinding Xia, Menghan Hu, Xiaojuan Xue, Qiaoyun Liu, Jian Zhang 0060, Guangtao Zhai |
VCIP | 6 |
| 2022 | Distinguishing Computer-Generated Images from Photographic Images: a Texture-Aware Deep Learning-Based MethodabstractWith the rapid development of computer graphics and generative models, computers are capable of generating images containing non-existent objects and scenes. Moreover, the computer-generated (CG) images may be indistinguishable from photographic (PG) images due to the strong representation ability of neural network and huge advancement of 3D rendering technologies. The abuse of such CG images may bring potential risks for personal property and social stability. Therefore, in this paper, we propose a dual-stream neural network to extract features enhanced by texture information to deal with the CG and PG image classification task. First, the input images are first converted to texture maps using the rotation-invariant uniform local binary patterns. Then we employ an attention-based texture-aware feature enhancement module to fuse the features extracted from each stage of the dual-stream neural network. Finally, the features are pooled and regressed into the predicted results by fully connected layers. The experimental results show that the proposed method achieves the best performance among all three popular CG and PG classification databases. The ablation study and cross-database validation experiments further confirm the effectiveness and generalization ability of the proposed algorithm. Wei Sun 0029, Xiongkuo Min, Tao Wang 0078, Wei Lu 0021, Guangtao Zhai |
VCIP | 6 |
| 2022 | EAN: Event Adaptive Network for Enhanced Action Recognition
Yuan Tian 0017, Yichao Yan, Guangtao Zhai, Guodong Guo |
Int. J. Comput. Vis. | 3 |
| 2022 | Graph-Based Denoising for Respiration and Heart Rate Estimation During Sleep in Thermal VideoabstractQuality sleep is a basic human need for well-being, yet sleep deprivation has been a long-term global problem. A common type of sleep deprivation is obstrucive sleep apnea, where people repeatedly stop breathing during sleep with subsequent abnormal vital signs, namely, respiration rate and heart rate. While tremendous effort has been made for vital signs monitoring systems during sleep, existing works still lack portability for bulky and intrusive systems and reliability for consumer-level, nonintrusive systems. To bridge the gap between practicability and accuracy and facilitate Internet of Things for smart healthcare, in this article, we propose a vital signs estimation system during sleep via a thermal camera. The system first captures thermal image sequences of a sleeping subject and then processes the facial regions within the thermal images for vital signs signal extraction. Specifically, leveraging on the inherent graph structure among subregions of the facial area, we propose a graph-based, spatial–temporal signal denoising scheme. Experimental results show that the graph-based denoising scheme in our system effectively reduces the noise level introduced by cameras and subjects, and our proposed system outperforms state-of-the-art nonintrusive vital signs monitoring systems. Since the algorithm components in our system have relatively low time complexity and no model training is required, our system can be deployed efficiently at the edge devices in a smart home setting. The extracted vital signs can then be used for sleep abnormality detection and disease screening. Cheng Yang 0003, Menghan Hu, Guangtao Zhai, Xiao-Ping Zhang 0002 |
IEEE Internet Things J. | 3 |
| 2022 | Multi-scale dilated convolution of feature Fusion Network for Crowd counting
Donghua Liu, Guodong Wang 0001, Guangtao Zhai |
Multim. Tools Appl. | 3 |
| 2022 | Calculation of ophthalmic diagnostic parameters on a single eye image based on deep neural network
Xuefei Song, Xiongkuo Min, Huifang Zhou, Wei Sun 0029, Jia Wang 0004, Guangtao Zhai |
Multim. Tools Appl. | 7 |
| 2022 | Blindly Assess Quality of In-the-Wild Videos via Quality-Aware Pre-Training and Motion PerceptionabstractPerceptual quality assessment of the videos acquired in the wilds is of vital importance for quality assurance of video services. The inaccessibility of reference videos with pristine quality and the complexity of authentic distortions pose great challenges for this kind of blind video quality assessment (BVQA) task. Although model-based transfer learning is an effective and efficient paradigm for the BVQA task, it remains to be a challenge to explorewhatandhowto bridge the domain shifts for better video representation. In this work, we propose to transfer knowledge from image quality assessment (IQA) databases with authentic distortions and large-scale action recognition with rich motion patterns. We rely on both groups of data to learn the feature extractor and use a mixed list-wise ranking loss function to train the entire model on the target VQA databases. Extensive experiments on six benchmarking databases demonstrate that our method performs very competitively under both individual database and mixed databases training settings. We also verify the rationality of each component of the proposed method and explore a simple ensemble trick for further improvement. Bowen Li 0018, Weixia Zhang, Guangtao Zhai, Xianpei Wang |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2022 | Spatial Temporal Video Enhancement Using Alternating ExposuresabstractHigh-speed video acquisition under poor illumination conditions is a challenging task. Imaging using long exposure can ensure brightness and suppress noise. However, the captured images may be blurry due to fast object movements or camera shakes. Imaging with short exposure can record sharp textures, but the high camera gain may cause noticeable noise. To alleviate this dilemma, we design a camera system using alternating exposures, where frames expose cyclically in a short-long way. The system consists of restoration and interpolation modules to reconstruct sharp, noise-reduced, high-frame-rate frames from low-frame-rate alternate-exposed input images. We design an optical-flow-based alternate-complementary alignment architecture for spatial enhancement, which effectively aligns the short-exposed and long-exposed images in a two-stage progressive way. Moreover, it explores complementary information from short-exposed and long-exposed inputs to ensure consistency between outputs. We propose a flow-enhanced frame interpolation module for temporal enhancement, which refines the intermediate flows and reconstructs the intermediate images based on the restored images of the alignment network and warped input neighboring frames. The whole network with two modules is end-to-end jointly learnable. We first evaluate the algorithm on simulation data. To demonstrate practicality, we then test it on real data by setting up a prototype camera. We propose an effective spatial degradation regularization strategy to reduce the domain gap between simulation and real data. Besides, we extend our method by integrating multi-frame exposure fusion technology to reduce overexposure areas in real scenarios. Experimental results show that our method performs favorably against state-of-the-art methods on both synthetic data and real-world data. Wang Shen, Guo Lu, Guangtao Zhai, Li Chen 0021, Muhammad Salman Asif |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2022 | Poxture: Human Posture Imitation Using Neural TextureabstractHuman pose imitation, which aims to generate an image with a source character’s appearance, the source character’s shape, and a target character’s posture, has many potential applications in virtual reality, augmented reality, games, movies, etc. It is incredibly challenging due to non-rigid human body motions, significant variations in clothing textures, and self-occluded human bodies in 2D images. In this paper, we propose Poxture, a novel human posture imitation method with neural texture, to address the challenges mentioned above. Concretely, first, we build a dense mapping between a source SMPL human body model (shape and posture) and its corresponding texture (appearance). Then, we apply a neural texture generator to recover the complete texture of the source character. At last, we wrap the source neural texture to the source SMLP model with a target pose to generate the desired image by a GAN model. Poxture does not require any annotations, and our framework can fully disentangle the source character’s appearance, shape, and pose, which enjoys several advantages: 1) It can synthesize high-resolution images with detailed textures, thanks to the learned neural textures containing both visible and invisible parts and high-frequency information; 2) It can imitate complex actions with various appearances and body figures since the complete texture of the source character is acquired. We compare our method with previous methods, showing state-of-the-art results on two challenging benchmarks. Extensive experiments demonstrate that, given any character, our method can manipulate this avatar imitating arbitrary posture. Chen Yang 0023, Zanwei Zhou, Bin Ji 0004, Guangtao Zhai, Wei Shen 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2022 | No-Reference Quality Assessment for 3D Colored Point Cloud and Mesh ModelsabstractTo improve the viewer’s Quality of Experience (QoE) and optimize computer graphics applications, 3D model quality assessment (3D-QA) has become an important task in the multimedia area. Point cloud and mesh are the two most widely used digital representation formats of 3D models, the visual quality of which is quite sensitive to lossy operations like simplification and compression. Therefore, many related studies such as point cloud quality assessment (PCQA) and mesh quality assessment (MQA) have been carried out to measure the visual quality of distorted 3D models. However, most previous studies utilize full-reference (FR) metrics, which indicates they can not predict the quality level in the absence of the reference 3D model. Furthermore, few 3D-QA metrics consider color information, which significantly restricts their effectiveness and scope of application. In this paper, we propose a no-reference (NR) quality assessment metric for colored 3D models represented by both point cloud and mesh. First, we project the 3D models from 3D space into quality-related geometry and color feature domains. Then, the 3D natural scene statistics (3D-NSS) and entropy are utilized to extract quality-aware features. Finally, a support vector regression (SVR) model is employed to regress the quality-aware features into visual quality scores. Our method is validated on the colored point cloud quality assessment database (SJTU-PCQA), the Waterloo point cloud assessment database (WPC), and the colored mesh quality assessment database (CMDM). The experimental results show that the proposed method outperforms most compared NR 3D-QA metrics with competitive computational resources and greatly reduces the performance gap with the state-of-the-art FR 3D-QA metrics. The code of the proposed model is publicly available now athttps://github.com/zzc-1998/NR-3DQA. Wei Sun 0029, Xiongkuo Min, Tao Wang 0078, Wei Lu 0021, Guangtao Zhai |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2022 | Viewing Behavior Supported Visual Saliency Predictor for 360 Degree VideosabstractIn virtual reality (VR), correct and precise estimations of user’s visual fixations and head movements can enhance the quality of experience by allocating more computation resources for analysing and rendering on the areas of interest. However, there is insufficient research about understanding the visual exploration of users when modeling VR visual attention. To bridge the gap between the saliency prediction for traditional 2D content and omnidirectional content, we construct the visual attention dataset and propose the visual saliency prediction framework for panoramic videos. Around the instantaneous viewing behavior, we propose a traditional method to adapt 2D saliency models and design a CNN-based model to better predict visual saliency. In the proposed traditional model, mechanism of visual attention and viewing behaviors are considered in the computation of edge weights on graphs which are interpreted as Markov chains. The fraction of the visual attention that is diverted to each high-clarity vision (HCV) area is estimated through equilibrium distribution of this chain. We also propose the Graph-Based CNN model. The RGB channel and optical flow form the spatial-temporal units of HCVs, from which node feature vectors are extracted. Graph convolution is used to learn the mutual information between node feature vectors of HCVs and retain geometric information. Then feature vectors are aligned according to geometry structure of equirectangular format, and the feature decoder maps the aligned feature maps to the data distribution. We also construct the dynamic omnidirectional monocular (DOM) saliency dataset with 64 diverse videos evaluated by 28 people. The subjective results show that the instantaneous viewing behavior is important in the VR experience. Extensive experiments are conducted on the dataset and the results demonstrate the effectiveness of the proposed framework. The dataset will be released to facilitate the future studies related to visual saliency prediction for 360-degree contents. Yucheng Zhu, Guangtao Zhai, Yiwei Yang 0007, Huiyu Duan, Xiongkuo Min, Xiaokang Yang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2022 | RIHOOP: Robust Invisible Hyperlinks in Offline and Online PhotographsabstractIn the era of multimedia and Internet, the quick response (QR) code helps people obtain information from offline to online quickly. However, the QR code is often limited in many scenarios because of its random and dull appearance. Therefore, this article proposes a novel approach to embed hyperlinks into common images, making the hyperlinks invisible for human eyes but detectable for mobile devices equipped with a camera. Our approach is an end-to-end neural network with an encoder to hide messages and a decoder to extract messages. To maintain the hidden message resilient to cameras, we build a distortion network between the encoder and the decoder to augment the encoded images. The distortion network uses differentiable 3-D rendering operations, which can simulate the distortion introduced by camera imaging in both printing and display scenarios. To maintain the visual attraction of the image with hyperlinks, a loss function conforming to the human visual system (HVS) is used to supervise the training of the encoder. Experimental results show that the proposed approach outperforms the previous work on both robustness and quality. Based on the proposed approach, many applications become possible, for example, "image hyperlinks" for advertisement on TV, website, or poster, and "invisible watermark" for copyright protection on digital resources or product packagings. Jun Jia, Zhongpai Gao, Menghan Hu, Xiongkuo Min, Guangtao Zhai, Xiaokang Yang 0001 |
IEEE Trans. Cybern. | 6 |
| 2022 | Confusing Image Quality Assessment: Toward Better Augmented Reality ExperienceabstractWith the development of multimedia technology, Augmented Reality (AR) has become a promising next-generation mobile platform. The primary value of AR is to promote the fusion of digital contents and real-world environments, however, studies on how this fusion will influence the Quality of Experience (QoE) of these two components are lacking. To achieve better QoE of AR, whose two layers are influenced by each other, it is important to evaluate its perceptual quality first. In this paper, we consider AR technology as the superimposition of virtual scenes and real scenes, and introduce visual confusion as its basic theory. A more general problem is first proposed, which is evaluating the perceptual quality of superimposed images, i.e., confusing image quality assessment. A ConFusing Image Quality Assessment (CFIQA) database is established, which includes 600 reference images and 300 distorted images generated by mixing reference images in pairs. Then a subjective quality perception experiment is conducted towards attaining a better understanding of how humans perceive the confusing images. Based on the CFIQA database, several benchmark models and a specifically designed CFIQA model are proposed for solving this problem. Experimental results show that the proposed CFIQA model achieves state-of-the-art performance compared to other benchmark models. Moreover, an extended ARIQA study is further conducted based on the CFIQA study. We establish an ARIQA database to better simulate the real AR application scenarios, which contains 20 AR reference images, 20 background (BG) reference images, and 560 distorted images generated from AR and BG references, as well as the correspondingly collected subjective quality ratings. Three types of full-reference (FR) IQA benchmark variants are designed to study whether we should consider the visual confusion when designing corresponding IQA algorithms. An ARIQA metric is finally proposed for better evaluating the perceptual quality of AR images. Experimental results demonstrate the good generalization ability of the CFIQA model and the state-of-the-art performance of the ARIQA model. The databases, benchmark models, and proposed metrics are available at: https://github.com/DuanHuiyu/ARIQA. Huiyu Duan, Xiongkuo Min, Yucheng Zhu, Guangtao Zhai, Xiaokang Yang 0001, Patrick Le Callet |
IEEE Trans. Image Process. | 4 |
| 2022 | HazDesNet: An End-to-End Network for Haze Density PredictionabstractVision-based intelligent systems such as driver assistance systems and transportation systems should take into account weather conditions. The presence of haze in images can be a critical threat to driving scenarios. Haze density measures the visibility and usability of hazy images captured in real-world conditions. The prediction of haze density can be valuable in various vision-based intelligent systems, especially in those systems deployed in outdoor environments. Haze density prediction is a challenging task since the haze and many scene contents have a lot in common in appearance. Existing methods generally utilize different priors and design complex handcrafted features to predict the visibility or haze density of the image. In this article, we propose a novel end-to-end convolutional neural network (CNN) based method to predict haze density, named as HazDesNet. Our HazDesNet takes a hazy image as input and predicts a pixel-level haze density map. The density map is then refined and smoothed, and the average of the refined map is calculated as the global haze density of the image. To verify the performance of HazDesNet, a subjective human study is performed to build a Human Perceptual Haze Density (HPHD) database, which includes 500 real-world hazy images and 100 synthetic hazy images, and the corresponding human-rated perceptual haze density scores. Experimental results show that our method achieves the best haze density prediction performance on our built HPHD database and existing databases. Besides the global quantitative results, our HazDesNet is capable of predicting a continuous, stable, fine, and high-resolution haze density map. We will make the database and code publicly available athttps://github.com/JiaheZhang/HazDesNet. Xiongkuo Min, Yucheng Zhu, Guangtao Zhai, Jiantao Zhou 0001, Xiaokang Yang 0001, Wenjun Zhang 0001 |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2022 | Dynamic Backlight Scaling Considering Ambient Luminance for Mobile Videos on LCD DisplaysabstractThe power consumption of mobile devices is always a concern for users due to the constraints on the size and weight of mobile devices. Up to now, mobile video traffic has accounted for the majority of total traffic, which implies that viewing mobile videos has been the major activity when using mobile devices. Among all the subsystems involving in mobile video playback, the display is the most power consuming subsystem. To reduce the power consumption, dynamic backlight scaling (DBS) technique is developed by adjusting the backlight magnitude when playing the mobile video. However, the convenience of mobile devices makes lots of people watch mobile videos in various luminance environments, which makes the existing DBS methods ineffective since ambient luminance varies greatly. In this paper, we propose a novel DBS strategy to maximally enhance the battery power performance under various ambient luminance conditions through backlight magnitude adjusting, while without negatively impacting users’ quality of experience (QoE). In particular, we conduct a series of subject quality assessment experiments to uncover the quantitative relationship among QoE, ambient luminance, video content luminance, and backlight luminance. We then investigate whether the continuous playback of backlight-scaled videos using the proposed scaling magnitude under various luminance environments would cause flicker effect or not. Motivated by the findings of these studies, we implement a novel DBS strategy for mobile energy saving which is suitable for various ambient luminance conditions. The experimental results demonstrate that the proposed DBS strategy can save more than 40 percent power at most and can save 10 percent power even at a very high ambient luminance condition. We also show that the proposed DBS strategy can be easily adapted to different user preferences and different devices, and can be conveniently integrated into practical applications. Wei Sun 0029, Xiongkuo Min, Guangtao Zhai, Ke Gu 0001, Siwei Ma 0001, Xiaokang Yang 0001 |
IEEE Trans. Mob. Comput. | 3 |
| 2022 | SMGEA: A New Ensemble Adversarial Attack Powered by Long-Term Gradient MemoriesabstractDeep neural networks are vulnerable to adversarial attacks. More importantly, some adversarial examples crafted against an ensemble of source models transfer to other target models and, thus, pose a security threat to black-box applications (when attackers have no access to the target models). Current transfer-based ensemble attacks, however, only consider a limited number of source models to craft an adversarial example and, thus, obtain poor transferability. Besides, recent query-based black-box attacks, which require numerous queries to the target model, not only come under suspicion by the target model but also cause expensive query cost. In this article, we propose a novel transfer-based black-box attack, dubbed serial-minigroup-ensemble-attack (SMGEA). Concretely, SMGEA first divides a large number of pretrained white-box source models into several "minigroups." For each minigroup, we design three new ensemble strategies to improve the intragroup transferability. Moreover, we propose a new algorithm that recursively accumulates the "long-term" gradient memories of the previous minigroup to the subsequent minigroup. This way, the learned adversarial information can be preserved, and the intergroup transferability can be improved. Experiments indicate that SMGEA not only achieves state-of-the-art black-box attack ability over several data sets but also deceives two online black-box saliency prediction systems in real world, i.e., DeepGaze-II (https://deepgaze.bethgelab.org/) and SALICON (http://salicon.net/demo/). Finally, we contribute a new code repository to promote research on adversarial attack and defense over ubiquitous pixel-to-pixel computer vision tasks. We share our code together with the pretrained substitute model zoo at https://github.com/CZHQuality/AAA-Pix2pix. Zhaohui Che, Ali Borji, Guangtao Zhai, Suiyi Ling, Jing Li 0026, Xiongkuo Min, Guodong Guo, Patrick Le Callet |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2022 | QoE Driven VR 360° Video Massive MIMO TransmissionabstractMassive multiple-input and multiple-output (MIMO) enables ultra-high throughput and low latency for tile-based adaptive virtual reality (VR) 360° video transmission in wireless network. In this paper, we consider a massive MIMO system where multiple users in a single-cell theater watch an identical VR 360° video. Based on tile prediction, base station (BS) deliveries the tiles in predicted field of view (FoV) to users. By introducing practical supplementary transmission for missing tiles and unacceptable VR sickness, we propose the first stable transmission scheme for VR video. we formulate an integer non-linear programming (INLP) problem to maximize users’ average quality of experience (QoE) score. Moreover, we derive the achievable spectral efficiency (SE) expression of predictive tile groups and the approximately achievable SE expression of missing tile groups, respectively. Analytically, the overall throughput is related to the number of tile groups and the length of pilot sequences. By exploiting the relationship between the structure of viewport tiles and SE expression, we propose a multi-lattice multi-stream grouping method aimed at improving the overall throughput for VR video transmission. Moreover, we analyze the relationship between QoE objective and number of predictive tile. We transform the original INLP problem into an integer linear programming problem by setting the predictive tiles groups as some constants. With variable relaxation and recovery, we obtain the optimal average QoE. Extensive simulation results validate that the proposed algorithm effectively improves QoE. Guangtao Zhai, Yongpeng Wu 0001, Xiongkuo Min, Wenjun Zhang 0001, Zhi Ding 0001, Chengshan Xiao |
IEEE Trans. Wirel. Commun. | 2 |
| 2021 | Learning Local Neighboring Structure for Robust 3D Shape RepresentationabstractMesh is a powerful data structure for 3D shapes. Representation learning for 3D meshes is important in many computer vision and graphics applications. The recent success of convolutional neural networks (CNNs) for structured data (e.g., images) suggests the value of adapting insight from CNN for 3D shapes. However, 3D shape data are irregular since each node's neighbors are unordered. Various graph neural networks for 3D shapes have been developed with isotropic filters or predefined local coordinate systems to overcome the node inconsistency on graphs. However, isotropic filters or predefined local coordinate systems limit the representation power. In this paper, we propose a local structure-aware anisotropic convolutional operation (LSA-Conv) that learns adaptive weighting matrices for each node according to the local neighboring structure and performs shared anisotropic filters. In fact, the learnable weighting matrix is similar to the attention matrix in random synthesizer -- a new Transformer model for natural language processing (NLP). Comprehensive experiments demonstrate that our model produces significant improvement in 3D shape reconstruction compared to state-of-the-art methods. Zhongpai Gao, Junchi Yan, Guangtao Zhai, Juyong Zhang, Yiyan Yang, Xiaokang Yang 0001 |
AAAI | 3 |
| 2021 | PRN: Psychology-Inspired Relation Network for Detecting Social Interaction Groups from Single Images
Jinhai Yang 0001, Hua Yang 0001, Guangtao Zhai |
BMVC | 4 |
| 2021 | Dual Attention Guided Gaze Target Detection in the WildabstractGaze target detection aims to infer where each person in a scene is looking. Existing works focus on 2D gaze and 2D saliency, but fail to exploit 3D contexts. In this work, we propose a three-stage method to simulate the human gaze inference behavior in 3D space. In the first stage, we introduce a coarse-to-fine strategy to robustly estimate a 3D gaze orientation from the head. The predicted gaze is decomposed into a planar gaze on the image plane and a depth-channel gaze. In the second stage, we develop a Dual Attention Module (DAM), which takes the planar gaze to produce the filed of view and masks interfering objects regulated by depth information according to the depth-channel gaze. In the third stage, we use the generated dual attention as guidance to perform two sub-tasks: (1) identifying whether the gaze target is inside or out of the image; (2) locating the target if inside. Extensive experiments demonstrate that our approach performs favorably against state-of-the-art methods on GazeFollow and VideoAttentionTarget datasets. Yi Fang 0009, Jiapeng Tang, Wang Shen, Wei Shen 0002, Xiao Gu 0001, Li Song 0001, Guangtao Zhai |
CVPR | 7 |
| 2021 | Identification of Deep Breath While Moving Forward Based on Multiple Body Regions and Graph Signal AnalysisabstractThis paper presents an unobtrusive solution that can automatically identify deep breath when a person is walking past the global depth camera. Existing non-contact breath assessments achieve satisfactory results under restricted conditions when human body stays relatively still. When someone moves forward, the breath signals detected by depth camera are hidden within signals of trunk displacement and deformation, and the signal length is short due to the short stay time, posing great challenges for us to establish models. To over-come these challenges, multiple region of interests (ROIs) based signal extraction and selection method is proposed to automatically obtain the signal informative to breath from depth video. Subsequently, graph signal analysis (GSA) is adopted as a spatial-temporal filter to wipe the components unrelated to breath. Finally, a classifier for identifying deep breath is established based on the selected breath-informative signal. In validation experiments, the proposed approach outperforms the comparative methods with the accuracy, precision, recall and F1 of 75.5%, 76.2%, 75.0% and 75.2%, respectively. This system can be extended to public places to provide timely and ubiquitous help for those who may have or are going through physical or mental trouble. Yunlu Wang, Cheng Yang 0003, Menghan Hu, Jian Zhang 0060, Qingli Li, Guangtao Zhai, Xiao-Ping Zhang 0002 |
ICASSP | 6 |
| 2021 | Perceptual Quality Assessment for Recognizing True and Pseudo 4k ContentabstractTo meet the imperative demand for monitoring the quality of Ultra High-Definition (UHD) content in multimedia industries, we propose an efficient no-reference (NR) image quality assessment (IQA) metric to distinguish original and pseudo 4K contents and measure the quality of their quality in this paper. First, we establish a database including more than 3000 4K images composed of natural 4K images together with upscaled versions interpolated from 1080p and 720p images by fourteen algorithms. To improve computing efficiency, our model segments the input image and selects three representative patches by local variances. Then, we extract the histogram features and cut-off frequency features in the frequency domain as well as the natural scenes statistic (NSS) based features from the representative patches. Finally, we employ support vector regressor (SVR) to aggregate these extracted features as an overall quality metric to predict the quality score of the target image. Extensive experimental comparisons using seven common evaluation indicators demonstrate that the proposed model outperforms the competitive NR IQA methods and has a great ability to distinguish true and pseudo 4K images. Wenhan Zhu, Guangtao Zhai, Xiongkuo Min, Xiaokang Yang 0001, Xiao-Ping Zhang 0002 |
ICASSP | 2 |
| 2021 | Looking here or there? Gaze Following in 360-Degree ImagesabstractGaze following, i.e., detecting the gaze target of a human subject, in 2D images has become an active topic in computer vision. However, it usually suffers from the out of frame issue due to the limited field-of-view (FoV) of 2D images. In this paper, we introduce a novel task, gaze following in 360-degree images which provide an omnidirectional FoV and can alleviate the out of frame issue. We collect the first dataset, "GazeFollow360"1, for this task, containing around 10,000 360-degree images with complex gaze behaviors under various scenes. Existing 2D gaze following methods suffer from performance degradation in 360degree images since they may use the assumption that a gaze target is in the 2D gaze sight line. However, this assumption is no longer true for long-distance gaze behaviors in 360-degree images, due to the distortion brought by sphere-to-plane projection. To address this challenge, we propose a 3D sight line guided dual-pathway framework, to detect the gaze target within a local region (here) and from a distant region (there), parallelly. Specifically, the local region is obtained as a 2D cone-shaped field along the 2D projection of the sight line starting at the human subject’s head position, and the distant region is obtained by searching along the sight line in 3D sphere space. Finally, the location of the gaze target is determined by fusing the estimations from both the local region and the distant region. Experimental results show that our method achieves significant improvements over previous 2D gaze following methods on our GazeFollow360 dataset. Wei Shen 0002, Zhongpai Gao, Yucheng Zhu, Guangtao Zhai, Guodong Guo |
ICCV | 5 |
| 2021 | Self-Conditioned Probabilistic Learning of Video RescalingabstractBicubic downscaling is a prevalent technique used to reduce the video storage burden or to accelerate the downstream processing speed. However, the inverse upscaling step is non-trivial, and the downscaled video may also deteriorate the performance of downstream tasks. In this paper, we propose a self-conditioned probabilistic framework for video rescaling to learn the paired downscaling and upscaling procedures simultaneously. During the training, we decrease the entropy of the information lost in the downscaling by maximizing its probability conditioned on the strong spatial-temporal prior information within the downscaled video. After optimization, the downscaled video by our framework preserves more meaningful information, which is beneficial for both the upscaling step and the downstream tasks, e.g., video action recognition task. We further extend the framework to a lossy video compression system, in which a gradient estimator for non-differential industrial lossy codecs is proposed for the end-to-end training of the whole system. Extensive experimental results demonstrate the superiority of our approach on video rescaling, video compression, and efficient action recognition tasks. Yuan Tian 0017, Guo Lu, Xiongkuo Min, Zhaohui Che, Guangtao Zhai, Guodong Guo |
ICCV | 5 |
| 2021 | Deep Neural Networks For Full-Reference And No-Reference Audio-Visual Quality AssessmentabstractIn the field of audio and visual quality assessment, most of previous works only focused on the single-mode visual or audio signal. However, for multi-mode signals, such as video and the accompanying audio, the overall perceptual quality depends on both video and audio. In this paper, we proposed an objective audio-visual quality assessment (AVQA) architecture for multi-mode signals based on deep neural networks. We first use a pretrained convolutional neural network to extract features of the single video frames and the concurrent short audio segments. Then, the extracted features are fed into Gated Recurrent Unit networks for time sequence modeling. Finally, we utilize the fully connected layers to fuse the qualities of audio and visual signals into the final quality score. The proposed architecture can be applied to both full-reference and no-reference AVQA. Experimental results on the LIVE-SJTU Database prove that our model outperforms the state-of-the-art AVQA methods. Yuqin Cao, Xiongkuo Min, Wei Sun 0029, Guangtao Zhai |
ICIP | 4 |
| 2021 | Muiqa: Image Quality Assessment Database And Algorithm For Medical Ultrasound ImagesabstractIn the process of medical image acquisition, medical images may be blurred or ghosted due to machine noise, electromagnetic interference, man-made disturbance, etc. This can result in poor image quality and severely affect the diagnosis accuracy and confidence of doctors. IntraVascular UltraSound (IVUS) is an important supplementary method for the diagnosis of coronary angiography. IVUS images can be distorted for many reasons and some severe distortions can affect diagnosis confidence. However, existing manual medical image quality control method is extremely time-consuming and requires a lot of manpower. To solve this problem, we first construct an Medical UltraSound Image Quality Assessment (MUIQA) database, which consists of 10766 IVUS images with quality labels given by professional doctors. Then we propose a deep-learning network to automatically distinguish the low, medium and high level images from each other. We achieve good classification accuracy of 96.34% on the testing set. Xiongkuo Min, Huiyu Duan, Yucheng Zhu, Guangtao Zhai |
ICIP | 5 |
| 2021 | Modeling Image Quality Score Distribution Using Alpha Stable ModelabstractIn recent years, image quality is generally described by a mean opinion score (MOS). However, we observe that an image’s quality ratings given by a group of subjects may not follow a Gaussian distribution and the image quality can not be fully described by a MOS. In this paper, we propose to describe the image quality using a parameterized distribution rather than a MOS, and an objective method is also proposed to predict the image quality score distribution (IQSD). Specifically, we selected 100 images from the LIVE database and invited a large group of subjects to evaluate the quality of these images. By analyzing the subjective quality ratings, we find that the IQSD can be well modeled by an alpha stable model and this model can reflect much more information than MOS. Therefore, we propose an algorithm to model the IQSD described by an alpha stable model. Features are extracted from images based on natural scene statistics and support vector regressors are trained to predict the IQSD described by an alpha stable model. We validate the proposed IQSD prediction model on the collected subjective quality ratings. Experimental results verify the effectiveness of the proposed algorithm in modeling the IQSD. Xiongkuo Min, Wenhan Zhu, Xiao-Ping Zhang 0002, Guangtao Zhai |
ICIP | 5 |
| 2021 | Accurate Compensation Makes the World More Clear for the Visually ImpairedabstractVisual impairment is one of the most serious social and public health problems in the world, therefore, it is of great theoretical and practical significance to study the image enhancement algorithms for the visually impaired, which is the basis for the development of assistive devices. In this paper, a general deep learning based image enhancement framework for the visually impaired is proposed, which can be used to enhance images to compensate for any visually impaired symptom that can be modeled. Take central vision loss as an example, we first model the central vision loss based on the contrast sensitivity function (CSF) specified by clinical indicator Pelli-Robson score and logMAR visual acuity, and then use the proposed framework to generate an image enhancement method aiming at compensating for the central vision loss. Both the simulation experiment and the patient experiment show the superiority of the proposed image enhancement method designed for the central vision loss, which also validates the effectiveness of the proposed framework. Sijing Wu, Huiyu Duan, Xiongkuo Min, Danyang Tu, Guangtao Zhai |
ICIP | 5 |
| 2021 | Deep Audio-Visual Fusion Neural Network for Saliency EstimationabstractIn this work, we propose a deep audio-visual fusion model to estimate the saliency of videos. The model extracts visual and audio features with two separate branches and fuses them to generate the saliency map. We design a novel temporal attention module to utilize the temporal information and a spatial feature pyramid module to fuse the spatial information. Then a multi-scale audio-visual fusion method is used to integrate different modalities. Furthermore, we propose a new dataset for audio-visual saliency estimation. The proposed dataset consists of 202 high quality video squences with a large range of motions, scenes and object types. Many of the videos have high audio-visual correspondence. Several experiments are conducted on different datasets. The results demonstrate that our model outperforms the previous state-of-the-art methods by a large margin and the proposed dataset can serve as a new benchmark for the audio-visual saliency estimation task. Xiongkuo Min, Guangtao Zhai |
ICIP | 3 |
| 2021 | Attention Based Network For No-Reference UGC Video Quality AssessmentabstractThe quality assessment of user-generated content (UGC) videos is a challenging problem due to the absence of reference videos and their complex distortions. Traditional no-reference video quality assessment (NR-VQA) algorithms mainly target specific synthetic distortions. Less attention has been paid to authentic distortions in UGC videos, which are not distributed evenly in both the spatial and temporal domains. In this paper, we propose an end-to-end neural network model for UGC video quality assessment based on the attention mechanism. The key step in our approach is to embed the attention modules in the feature extraction network, which effectively extracts local distortion information. In addition, to exploit the temporal perception mechanism of the human visual system (HVS), the gated recurrent unit (GRU) and temporal pooling layer are integrated into the proposed model. We validate the proposed model on three public in-the-wild VQA databases: KoNViD-1k, CVD2014, and LIVE-Qualcomm. Experimental results demonstrate that the proposed method outperforms state-of-the-art NR-VQA models. The implementation of our method is released at https://github.com/qingshangithub/AB-VQA. Fuwang Yi, Mianyi Chen, Wei Sun 0029, Xiongkuo Min, Yuan Tian 0017, Guangtao Zhai |
ICIP | 6 |
| 2021 | Key Facial Components Guided Micro-Expression Recognition Based on First & Second-Order MotionabstractAlthough there have been many successful attempts in the field of micro-expression recognition, plenty of challenges remain due to the subtle spatio-temporal changes and high locality of micro-expressions. In this paper, to tackle such issues, we propose a novel key facial components guided micro-expression recognition approach (KFC-MER). Face semantic segmentation probability maps involving several key components provide a guidance for feature learning. With the Components-Aware Attention (CAA) module, expression-related areas are highlighted and the relationship between components will also be learned. To cope with the limited size of micro-expression datasets, we design a parallel shallow residual network as the MER network. Both the first- and second-order motion are exploited as the input data, for capturing motive information and non-rigid deformation, respectively. Extensive experiments on three benchmark datasets demonstrate that our method outperforms previous works and achieves state-of-the-art performance. The code is publicly available on GitHub: https://github.com/TJUMMG/KFC-MER. Yuting Su 0001, Jing Liu 0002, Guangtao Zhai |
ICME | 4 |
| 2021 | A No-Reference Evaluation Metric for Low-Light Image EnhancementabstractLow-light images, which are usually taken in dark or back-lighting conditions, are hard to perceive due to the low visibility and low contrast. To improve viewers’ Quality of Experience (QoE) and support the application of vision-based systems, various low-light image enhancement algorithms (LIEAs) have been proposed to lighten low-light images. However, some LIEAs may amplify the hidden distortions in the dark like noise and even further, introduce new distortions such as structural damage, color shift, etc, which severely affect the quality of light-enhanced images and need to be evaluated quantificationally. However, in the literature, few measures are proposed to assess the quality of light-enhanced images. Therefore, in this paper, we develop a no-reference low-light image enhancement evaluation (NLIEE) metric to predict the quality of light-enhanced images. The image quality is mainly assessed from four key aspects: light enhancement, color comparison, noise measurement, and structure evaluation. The experiment results show that NLIEE achieves the best performance among the general no-reference image quality assessment (NR IQA) models and quality descriptors for light enhancement. Wei Sun 0029, Xiongkuo Min, Wenhan Zhu, Tao Wang 0078, Wei Lu 0021, Guangtao Zhai |
ICME | 7 |
| 2021 | A Lightweight Saliency Prediction Model for Omnidirectional ImagesabstractAt present, most high-performing saliency prediction models for omnidirectional images (ODIs) depend on deeper or wider convolutional neural networks (CNNs), benefiting from their superior feature representation capability but suffering from high computational costs. To address this issue, we propose a novel lightweight saliency prediction model to predict the eye fixations on ODIs. Specifically, our proposed model consists of three modules: a lightweight feature representation module, a supervised attention module, and a dynamic convolution aggregation module. Different from the existing saliency prediction models, our proposed model is the first to introduce the dynamic convolution into the saliency prediction and aggregate multiple parallel convolution kernels dynamically based on their attention. Such a dynamic convolution operation is not only computationally efficient (small kernel size), but also increases the feature representation capability since these convolution kernels are aggregated in a non-linear manner via attention. Experimental results on two benchmark datasets show that our model is lightweight and outperforms other state-of-the-art methods. Dandan Zhu 0001, Yongqing Chen, Defang Zhao, Xiongkuo Min, Qiangqiang Zhou, Shaobo Yu, Guangtao Zhai, Xiaokang Yang 0001 |
ICME | 7 |
| 2021 | Lavs: A Lightweight Audio-Visual Saliency Prediction ModelabstractAudio information is essential for guiding human attention and visual perception, which has been verified by many comprehensive psychological studies. However, the audio modality has been rather neglected in modeling visual attention, most of the current visual attention models heavily depend on visual information. Additionally, current existing high-performing visual attention models rely on deeper convolution neural networks (CNNs), benefiting from their extraordinary feature learning ability but incurring high computational cost. To this end, we propose a novel lightweight audio-visual saliency (LAVS) model to efficiently address the problem of fixation prediction in videos. To the best of our knowledge, our proposed model constitutes the first attempt to exploit a lightweight network and combines the visual and audio cues to perform saliency estimation in videos. Specifically, our proposed model consists of four modules, which are spatial-temporal visual saliency estimation module, audio features extraction module, source sound localization module, and audio-visual saliency fusion module. Extensive experiments across datasets validate the effectiveness and real-time performance of the proposed LAVS model, which outperforms the other state-of-the-art methods. Dandan Zhu 0001, Defang Zhao, Xiongkuo Min, Tian Han 0001, Qiangqiang Zhou, Shaobo Yu, Yongqing Chen, Guangtao Zhai, Xiaokang Yang 0001 |
ICME | 8 |
| 2021 | Learning Spectral Dictionary for Local Representation of MeshabstractFor meshes, sharing the topology of a template is a common and practical setting in face-, hand-, and body-related applications. Meshes are irregular since each vertex's neighbors are unordered and their orientations are inconsistent with other vertices. Previous methods use isotropic filters or predefined local coordinate systems or learning weighting matrices for each vertex of the template to overcome the irregularity. Learning weighting matrices for each vertex to soft-permute the vertex's neighbors into an implicit canonical order is an effective way to capture the local structure of each vertex. However, learning weighting matrices for each vertex increases the parameter size linearly with the number of vertices and large amounts of parameters are required for high-resolution 3D shapes. In this paper, we learn spectral dictionary (i.e., bases) for the weighting matrices such that the parameter size is independent of the resolution of 3D shapes. The coefficients of the weighting matrix bases for each vertex are learned from the spectral features of the template's vertex and its neighbors in a weight-sharing manner. Comprehensive experiments demonstrate that our model produces state-of-the-art results with a much smaller model size. Zhongpai Gao, Junchi Yan, Guangtao Zhai, Xiaokang Yang 0001 |
IJCAI | 3 |
| 2021 | Low-Cost and Unobtrusive Respiratory Condition Monitoring Based on Raspberry Pi and Recurrent Neural NetworkabstractThis paper presents a low-cost and unobtrusive intelligent respiratory monitoring system. To achieve low-cost and remote measurement of respiratory signal, an RGB camera collaborated with marker tracking is used as data acquisition sensor, and a Raspberry Pi is used as data processing platform. To overcome challenges in actual applications, the signal processing algorithms are designed for removing sudden body movements and smoothing the raw signal. To discover more specific information in the respiratory signal, respiratory rate is estimated by a translational cross point algorithm, and respiratory pattern is identified by recurrent neural network. Finally, the obtained decision-making information and some original information are sent to user's smartphone via a cloud service platform. For estimating respiratory rate, the Bland-Altman plot demonstrates the satisfactory results with agreement ranges of -0.13 ± 5.85 bpm. With respect to the classification of breathing patterns, the results validate that the system has the good performance with the accuracy, precision, recall, and F1 of 92.5%, 92.5%, 93.3%, and 92.9%, respectively. This work may contribute to the development of low-cost and non-contact respiratory monitoring products specific to home or work health care. Yunlu Wang, Menghan Hu, Jian Zhang 0060, Qingli Li, Guangtao Zhai, Simon X. Yang |
ISCAS | 7 |
| 2021 | Blindly Predict Image and Video Quality in the WildabstractEmerging interests have been brought to blind quality assessment for images/videos captured in the wild, known as in-the-wild I/VQA. Prior deep learning based approaches have achieved considerable progress in I/VQA, but are intrinsically troubled with two issues. Firstly, most existing methods fine-tune the image-classification-oriented pre-trained models for the absence of large-scale I/VQA datasets. However, the task misalignment between I/VQA and image classification leads to degraded generalization performance. Secondly, existing VQA methods directly conduct temporal pooling on the predicted frame-wise scores, resulting in ambiguous inter-frame relation modeling. In this work, we propose a two-stage architecture to separately predict image and video quality in the wild. In the first stage, we resort to supervised contrastive learning to derive quality-aware representations that facilitate the prediction of image quality. Specifically, we propose a novel quality-aware contrastive loss to pull together samples of similar quality and push away quality-different ones in embedding space. In the second stage, we develop a Relation-Guided Temporal Attention (RTA) module for video quality prediction, which captures global inter-frame dependencies in embedding space to learn frame-wise attention weights for frame quality aggregation. Extensive experiments demonstrate that our approach performs favorably against state-of-the-art methods on both authentically distorted image benchmarks and video benchmarks. Jiapeng Tang, Yi Fang 0009, Rong Xie 0004, Xiao Gu 0001, Guangtao Zhai, Li Song 0001 |
MMAsia | 6 |
| 2021 | Dual-Layer Barcodes
Jun Jia, Guangtao Zhai |
PRCV (2) | 3 |
| 2021 | A Multi-dimensional Aesthetic Quality Assessment Model for Mobile Game ImagesabstractWith the development of the game industry and the popularization of mobile devices, mobile games have played an important role in people's entertainment life. The aesthetic quality of mobile game images determines the users' Quality of Experience (QoE) to a certain extent. In this paper, we propose a multi-task deep learning based method to evaluate the aesthetic quality of mobile game images in multiple dimensions (i.e. the fineness, color harmony, colorfulness, and overall quality). Specifically, we first extract the quality-aware feature representation through integrating the features from all intermediate layers of the convolution neural network (CNN) and then map these quality-aware features into the quality score space in each dimension via the quality regressor module, which consists of three fully connected (FC) layers. The proposed model is trained through a multi-task learning manner, where the quality-aware features are shared by different quality dimension prediction tasks, and the multi-dimensional quality scores of each image are regressed by multiple quality regression modules respectively. We further introduce an uncertainty principle to balance the loss of each task in the training stage. The experimental results show that our proposed model achieves the best performance on the Multi-dimensional Aesthetic assessment for Mobile Game image database (MAMG) among state-of-the-art image quality assessment (IQA) algorithms and aesthetic quality assessment (AQA) algorithms. Tao Wang 0078, Wei Sun 0029, Xiongkuo Min, Wei Lu 0021, Guangtao Zhai |
VCIP | 6 |
| 2021 | LRS-Net: invisible QR Code embedding, detection, and restorationabstractQR code is a powerful tool to bridge the offline and online worlds. It has been widely used because it can store a large amount of information in a small space. However, the black-and-white style of QR codes is not attractive to the human eyes when embedded in videos, which greatly affects the viewing experience. Invisible QR code has proposed based on temporal psycho-visual modulation (TPVM) to embed invisible hyperlinks in shopping websites, copyright watermarks in movies, etc. However, existing embedding and detection methods are not robust enough. In this paper, we adopt a novel embedding method to greatly improve the visual quality of the embedded video. Furthermore, we build a new dataset of invisible QR codes named 'IQRCodes' to train deep neural networks. At last, we propose localization, refinement, and segmentation neural netowrks (LRS-Net) to efficiently detect and restore invisible QR codes that are captured by mobile phones. Yiyan Yang, Zhongpai Gao, Guangtao Zhai |
VCIP | 3 |
| 2021 | SalGFCN: Graph Based Fully Convolutional Network for Panoramic Saliency PredictionabstractThe saliency prediction of panoramic images is dramatically affected by the distortion caused by non-Euclidean geometry characteristic. Traditional CNN based saliency pre-diction algorithms for 2D images are no longer suitable for 360-degree images. Intuitively, we propose a graph based fully convolutional network for saliency prediction of 360-degree images, which can reasonably map panoramic pixels to spherical graph data structures for representation. The saliency prediction network is based on residual U-Net architecture, with dilated graph convolutions and attention mechanism in the bottleneck. Furthermore, we design a fully convolutional layer for graph pooling and unpooling operations in spherical graph space to retain node-to-node features. Experimental results show that our proposed method outperforms other state-of-the-art saliency models on the large-scale dataset. Yiwei Yang 0007, Yucheng Zhu, Zhongpai Gao, Guangtao Zhai |
VCIP | 4 |
| 2021 | Inter-Observer Visual Congruency in Video-ViewingabstractThere are individual differences in human visual attention between observers when viewing the same scene. Inter-observer visual congruency (IOVC) describes the dispersion between different people's visual attention areas when they observe the same stimulus. Research on the IOVC of video is interesting but lacking. In this paper, we first introduce the measurement to calculate the IOVC of video. And an eye-tracking experiment is conducted in a realistic movie-watching environment to establish a movie scene dataset. Then we propose a method to predict the IOVC of video, which employs a dual-channel network to extract and integrate content and optical flow features. The effectiveness of the proposed prediction model is validated on our dataset. And the correlation between inter-observer congruency and video emotion is analyzed. Jiaomin Yue, Dandan Zhu 0001, Xiongkuo Min, Xiao-Ping Zhang 0002, Guangtao Zhai |
VCIP | 6 |
| 2021 | A Full-Reference Quality Assessment Metric for Fine-Grained Compressed ImagesabstractCompressed image quality assessment (IQA) has been a crucial part of a wide range of image services such as storage and transmission. Due to the effect of different bit rates and compression methods, the compressed images usually have different levels of quality. Nowadays, the mainstream full-reference (FR) metrics are effective to predict the quality of compressed images at coarse-grained levels, however, they may perform poorly when quality differences of the compressed images are quite subtle. To better improve the Quality of Experience (QoE) and provide useful guidance for compression algorithms, we propose an FR-IQA metric for fine-grained compressed images, which estimates the image quality by analyzing the difference of structure and texture. Our metric is mainly validated on the fine-grained compression IQA (FGIQA) database and is tested on other commonly used compression IQA databases as well. The experimental results show that our metric outperforms mainstream FR-IQA metrics on the fine-grained compression IQA database and also obtains competitive performance on the coarse-grained compression IQA databases. Wei Sun 0029, Xiongkuo Min, Tao Wang 0078, Wei Lu 0021, Guangtao Zhai |
VCIP | 6 |
| 2021 | Psycho-visual modulation based information display: introduction and survey
Zhongpai Gao, Jia Wang 0004, Guangtao Zhai |
Frontiers Comput. Sci. | 4 |
| 2021 | RANSP: Ranking attention network for saliency prediction on omnidirectional images
Dandan Zhu 0001, Yongqing Chen, Xiongkuo Min, Yucheng Zhu, Qiangqiang Zhou, Guangtao Zhai, Xiaokang Yang 0001 |
Neurocomputing | 7 |
| 2021 | An Accurate and Efficient 1-D Barcode Detector for Medium of Deployment in IoT SystemsabstractCamera-based 1-D barcode detectors have a lot of applications in Internet-of-Things (IoT) systems (e.g., retail, air travel, post and parcel services, and manufacturing). Based on the observation that 1-D barcodes always come with a 12-digit product code, this article proposes an end-to-end trainable and fully convoluted model that can detect and output accurate localization results of 1-D barcode and product code simultaneously. Our method uses dilated convolutions-based feature extractors which are then combined with systematic feature merging layers to create a U-shaped network. It predicts multichannel feature maps which later yield localization results after thresholding with a confidence map generated by the model and nonmaximum suppression. Furthermore, we use a Taylor series expansion-based criterion to rank and eliminate a subset of least important convolutional filters of the model, which further increases the inference speed to a great extent. This model can act as a preprocessing module for camera-based barcode decoders. Experimental results on the data set combined from public data sets and self-collected retail products data set obtained in more challenging environment conditions demonstrate that our strategy is effective in increasing the decoding rate of existing commercial barcode decoders efficiently. Adnan Sharif, Guangtao Zhai, Jun Jia, Xiongkuo Min |
IEEE Internet Things J. | 2 |
| 2021 | Enhancing Decoding Rate of Barcode Decoders in Complex Scenes for IoT SystemsabstractCamera-based multiclass (1-D and 2-D) barcode detectors that can help in decoding barcodes in different complex scenes have huge potential applications in situations where Internet of Things (IoT) is combined with artificial intelligence (AI) and augmented reality (AR). The decoding rate in such applications under real-life complex scenes is greatly affected by two major factors: first, we cannot accurately localize the barcodes, and second, we cannot decode the blur samples. In this article, we first propose a barcode localization algorithm that is capable of regressing four vertices of barcodes accurately. Our localization method comprises an anchor-free approach that outputs multiscale output prediction maps. These segmentation-like maps of each scale are then further divided into three types of maps (classification, centerness, and 8-D regression). Eight-dimensional localization result of barcodes along with classification result is then obtained after postprocessing. Second, we propose a conditional generative adversarial network-based model for deblurring blur QR codes. Extensive decoding experiments on a challenging complex scene data set show that our localization and deblurring methods can contribute to improving the decoding rate of existing barcode decoders. Adnan Sharif, Guangtao Zhai, Xiongkuo Min, Jun Jia, Kashif Munir |
IEEE Internet Things J. | 2 |
| 2021 | Respiratory Consultant by Your Side: Affordable and Remote Intelligent Respiratory Rate and Respiratory Pattern Monitoring SystemabstractThe aim of this study is to develop an affordable and remote intelligent respiratory monitoring system. To achieve low-cost and remote measurement of respiratory signal, an RGB camera collaborated with marker tracking is used as a data acquisition sensor, and a Raspberry Pi is used as a data processing platform. To overcome challenges in actual applications, the signal processing algorithms are designed for removing sudden body movements and smoothing the raw signal. Subsequently, respiratory rate (RR) is estimated by a translational cross-point algorithm, and the respiratory pattern is identified by the recurrent neural network. For estimating RR, the translational cross-point algorithm performs better than other methods with root-mean-square error (RMSE) of 3.29 bpm. With respect to the classification of breathing patterns, the established neural network performs better than support vector machine-based classifiers with the accuracy, precision, recall, and F1 of 89.0%, 89.0%, 90.5%, and 89.0%, respectively. The obtained decision-making information and some original information are sent to the user’s smartphone via a cloud service platform. In a way, due to its low-price, noncontact, and portable merits, the established system can be seen as a “respiratory consultant” by your side. Yunlu Wang, Menghan Hu, Jian Zhang 0060, Qingli Li, Guangtao Zhai, Simon X. Yang, Xiao-Ping Zhang 0002, Xiaokang Yang 0001 |
IEEE Internet Things J. | 7 |
| 2021 | Fine localization and distortion resistant detection of multi-class barcode in complex environments
Xiongkuo Min, Jun Jia, Zehao Zhu, Jia Wang 0004, Guangtao Zhai |
Multim. Tools Appl. | 6 |
| 2021 | Saliency4ASD: Challenge, dataset and tools for visual attention modeling for autism spectrum disorder
Jesús Gutiérrez 0001, Zhaohui Che, Guangtao Zhai, Patrick Le Callet |
Signal Process. Image Commun. | 3 |
| 2021 | Adversarial Attack Against Deep Saliency Models Powered by Non-Redundant PriorsabstractSaliency detection is an effective front-end process to many security-related tasks, e.g. automatic drive and tracking. Adversarial attack serves as an efficient surrogate to evaluate the robustness of deep saliency models before they are deployed in real world. However, most of current adversarial attacks exploit the gradients spanning the entire image space to craft adversarial examples, ignoring the fact that natural images are high-dimensional and spatially over-redundant, thus causing expensive attack cost and poor perceptibility. To circumvent these issues, this paper builds an efficient bridge between the accessible partially-white-box source models and the unknown black-box target models. The proposed method includes two steps: 1) We design a new partially-white-box attack, which defines the cost function in the compact hidden space to punish a fraction of feature activations corresponding to the salient regions, instead of punishing every pixel spanning the entire dense output space. This partially-white-box attack reduces the redundancy of the adversarial perturbation. 2) We exploit the non-redundant perturbations from some source models as the prior cues, and use an iterative zeroth-order optimizer to compute the directional derivatives along the non-redundant prior directions, in order to estimate the actual gradient of the black-box target model. The non-redundant priors boost the update of some "critical" pixels locating at non-zero coordinates of the prior cues, while keeping other redundant pixels locating at the zero coordinates unaffected. Our method achieves the best tradeoff between attack ability and perturbation redundancy. Finally, we conduct a comprehensive experiment to test the robustness of 18 state-of-the-art deep saliency models against 16 malicious attacks, under both of white-box and black-box settings, which contributes a new robustness benchmark to the saliency community for the first time. Zhaohui Che, Ali Borji, Guangtao Zhai, Suiyi Ling, Jing Li 0026, Yuan Tian 0017, Guodong Guo, Patrick Le Callet |
IEEE Trans. Image Process. | 3 |
| 2021 | Quality Assessment of Free-Viewpoint Videos by Quantifying the Elastic Changes of Multi-Scale Motion TrajectoriesabstractVirtual viewpoints synthesis is an essential process for many immersive applications including Free-viewpoint TV (FTV). A widely used technique for viewpoints synthesis is Depth-Image-Based-Rendering (DIBR) technique. However, such technique may introduce challenging non-uniform spatial-temporal structure-related distortions. Most of the existing state-of-the-art quality metrics fail to handle these distortions, especially the temporal structure inconsistencies observed during the switch of different viewpoints. To tackle this problem, an elastic metric and multi-scale trajectory based video quality metric (EM-VQM) is proposed in this paper. Dense motion trajectory is first used as a proxy for selecting temporal sensitive regions, where local geometric distortions might significantly diminish the perceived quality. Afterwards, the amount of temporal structure inconsistencies and unsmooth viewpoints transitions are quantified by calculating 1) the amount of motion trajectory deformations with elastic metric and, 2) the spatial-temporal structural dissimilarity. According to the comprehensive experimental results on two FTV video datasets, the proposed metric outperforms the state-of-the-art metrics designed for free-viewpoint videos significantly and achieves a gain of 12.86% and 16.75% in terms of median Pearson linear correlation coefficient values on the two datasets compared to the best one, respectively. Suiyi Ling, Jing Li 0026, Zhaohui Che, Xiongkuo Min, Guangtao Zhai, Patrick Le Callet |
IEEE Trans. Image Process. | 5 |
| 2021 | Video Frame Interpolation and Enhancement via Pyramid Recurrent FrameworkabstractVideo frame interpolation aims to improve users' watching experiences by generating high-frame-rate videos from low-frame-rate ones. Existing approaches typically focus on synthesizing intermediate frames using high-quality reference images. However, the captured reference frames may suffer from inevitable spatial degradations such as motion blur, sensor noise, etc. Few studies have approached the joint video enhancement problem, namely synthesizing high-frame-rate and high-quality results from low-frame-rate degraded inputs. In this paper, we propose a unified optimization framework for video frame interpolation with spatial degradations. Specifically, we develop a frame interpolation module with a pyramid structure to cyclically synthesize high-quality intermediate frames. The pyramid module features adjustable spatial receptive field and temporal scope, thus contributing to controllable computational complexity and restoration ability. Besides, we propose an inter-pyramid recurrent module to connect sequential models to exploit the temporal relationship. The pyramid module integrates the recurrent module, thus can iteratively synthesize temporally smooth results. And the pyramid modules share weights across iterations, thus it does not expand the model's parameter size. Our model can be generalized to several applications such as up-converting the frame rate of videos with motion blur, reducing compression artifacts, and jointly super-resolving low-resolution videos. Extensive experimental results demonstrate that our method performs favorably against state-of-the-art methods on various video frame interpolation and enhancement tasks. Wang Shen, Wenbo Bao, Guangtao Zhai, Li Chen 0021, Xiongkuo Min |
IEEE Trans. Image Process. | 3 |
| 2021 | Uncertainty-Aware Blind Image Quality Assessment in the Laboratory and WildabstractPerformance of blind image quality assessment (BIQA) models has been significantly boosted by end-to-end optimization of feature engineering and quality regression. Nevertheless, due to the distributional shift between images simulated in the laboratory and captured in the wild, models trained on databases with synthetic distortions remain particularly weak at handling realistic distortions (and vice versa). To confront the cross-distortion-scenario challenge, we develop a unified BIQA model and an approach of training it for both synthetic and realistic distortions. We first sample pairs of images from individual IQA databases, and compute a probability that the first image of each pair is of higher quality. We then employ the fidelity loss to optimize a deep neural network for BIQA over a large number of such image pairs. We also explicitly enforce a hinge constraint to regularize uncertainty estimation during optimization. Extensive experiments on six IQA databases show the promise of the learned method in blindly assessing image quality in the laboratory and wild. In addition, we demonstrate the universality of the proposed training strategy by using it to improve existing BIQA models. Weixia Zhang, Kede Ma, Guangtao Zhai, Xiaokang Yang 0001 |
IEEE Trans. Image Process. | 3 |
| 2021 | Exploiting Local Degradation Characteristics and Global Statistical Properties for Blind Quality Assessment of Tone-Mapped HDR ImagesabstractTone mapping operators (TMOs) are developed to convert a high dynamic range (HDR) image into a low dynamic range (LDR) one for display with the goal of preserving as much visual information as possible. However, image quality degradation is inevitable due to the dynamic range compression during the tone-mapping process. This accordingly raises an urgent demand for effective quality evaluation methods to select a high-quality tone-mapped image (TMI) from a set of candidates generated by distinct TMOs or the same TMO with different parameter settings. A key element to the success of TMI quality evaluation is to extract effective features that are highly consistent with human perception. Towards this end, this paper proposes a novel blind TMI quality metric by exploiting both local degradation characteristics and global statistical properties for feature extraction. Several image attributes including texture, structure, colorfulness and naturalness are considered either locally or globally. The extracted local and global features are aggregated into an overall quality via regression. Experimental results on two benchmark databases demonstrate the superiority of the proposed metric over both the state-of-the-art blind quality models designed for synthetically distorted images (SDIs) and the blind quality models specifically developed for TMIs. Xuejin Wang, Qiuping Jiang, Feng Shao 0001, Ke Gu 0001, Guangtao Zhai, Xiaokang Yang 0001 |
IEEE Trans. Multim. | 5 |
| 2021 | Comparative Perceptual Assessment of Visual Signals Using Free Energy FeaturesabstractIn this paper, we put forward the concept of comparative perceptual quality assessment (C-PQA), which refers to the judgment of relative qualities of two visual signals of the same content, but subject to different types and levels of distortions. While it is straightforward for human observers to fulfill the CPQA task in daily lives, it remains a difficult challenge for the current research of perceptual quality assessment (PQA). Among the existing PQA algorithms, the full-reference (FR) and reducedreference (RR) methods both need prior knowledge of the original images while the no-reference (NR) algorithms usually work with a single input image. C-PQA is inherently different from those existing methods in that it takes an image pair as input and predicts their relative quality without using any knowledge about the original image. In this paper, we propose a brain theory inspired approach to C-PQA that emulates the process of comparing the relative quality of two visual stimuli as performed by the human visual system (HVS) within the framework of free energy minimization. The brain's internal generative models initialized on the inputs are then used to explain both images. During the internal generative modeling, a group of features are extracted and then integrated to determine the relative quality of two images. We designed a dedicated image database to test the proposed C-PQA algorithm. Experimental results show that the proposed method achieves up to 98% prediction accuracy in line with the subjective ratings, outperforming many state of the art PQA algorithms. Guangtao Zhai, Yucheng Zhu, Xiongkuo Min |
IEEE Trans. Multim. | 1 |
| 2021 | Perceptual Quality Assessment of Low-light Image EnhancementabstractLow-light image enhancement algorithms (LIEA) can light up images captured in dark or back-lighting conditions. However, LIEA may introduce various distortions such as structure damage, color shift, and noise into the enhanced images. Despite various LIEAs proposed in the literature, few efforts have been made to study the quality evaluation of low-light enhancement. In this article, we make one of the first attempts to investigate the quality assessment problem of low-light image enhancement. To facilitate the study of objective image quality assessment (IQA), we first build a large-scale low-light image enhancement quality (LIEQ) database. The LIEQ database includes 1,000 light-enhanced images, which are generated from 100 low-light images using 10 LIEAs. Rather than evaluating the quality of light-enhanced images directly, which is more difficult, we propose to use the multi-exposure fused (MEF) image and stack-based high dynamic range (HDR) image as a reference and evaluate the quality of low-light enhancement following a full-reference (FR) quality assessment routine. We observe that distortions introduced in low-light enhancement are significantly different from distortions considered in traditional image IQA databases that are well-studied, and the current state-of-the-art FR IQA models are also not suitable for evaluating their quality. Therefore, we propose a new FR low-light image enhancement quality assessment (LIEQA) index by evaluating the image quality from four aspects: luminance enhancement, color rendition, noise evaluation, and structure preserving, which have captured the most key aspects of low-light enhancement. Experimental results on the LIEQ database show that the proposed LIEQA index outperforms the state-of-the-art FR IQA models. LIEQA can act as an evaluator for various low-light enhancement algorithms and systems. To the best of our knowledge, this article is the first of its kind comprehensive low-light image enhancement quality assessment study. Guangtao Zhai, Wei Sun 0029, Xiongkuo Min, Jiantao Zhou 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2021 | Learning a Deep Agent to Predict Head Movement in 360-Degree ImagesabstractVirtual reality adequately stimulates senses to trick users into accepting the virtual environment. To create a sense of immersion, high-resolution images are required to satisfy human visual system, and low latency is essential for smooth operations, which put great demands on data processing and transmission. Actually, when exploring in the virtual environment, viewers only perceive the content in the current field of view. Therefore, if we can predict the head movements that are important behaviors of viewers, more processing resources can be allocated to the active field of view. In this article, we propose a model to predict the trajectory of head movement. Deep reinforcement learning is employed to mimic the decision making. In our framework, to characterize each state, features for viewport images are extracted by convolutional neural networks. In addition, the spherical coordinate maps and visited maps are generated for each viewport image, which facilitate the multiple dimensions of the state information by considering the impact of historical head movement and position information. To ensure the accurate simulation of visual behaviors during the watching of panoramas, we stipulate that the model imitates the behaviors of human demonstrators. To allow the model to generalize to more conditions, the intrinsic motivation is employed to guide the agent’s action toward reducing uncertainty, which can enhance robustness during the exploration. The experimental results demonstrate the effectiveness of the proposed stepwise head movement predictor. Yucheng Zhu, Guangtao Zhai, Xiongkuo Min, Jiantao Zhou 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2020 | A New Ensemble Adversarial Attack Powered by Long-Term Gradient MemoriesabstractDeep neural networks are vulnerable to adversarial attacks. More importantly, some adversarial examples crafted against an ensemble of pre-trained source models can transfer to other new target models, thus pose a security threat to black-box applications (when the attackers have no access to the target models). Despite adopting diverse architectures and parameters, source and target models often share similar decision boundaries. Therefore, if an adversary is capable of fooling several source models concurrently, it can potentially capture intrinsic transferable adversarial information that may allow it to fool a broad class of other black-box target models. Current ensemble attacks, however, only consider a limited number of source models to craft an adversary, and obtain poor transferability. In this paper, we propose a novel black-box attack, dubbed Serial-Mini-Batch-Ensemble-Attack (SMBEA). SMBEA divides a large number of pre-trained source models into several mini-batches. For each single batch, we design 3 new ensemble strategies to improve the intra-batch transferability. Besides, we propose a new algorithm that recursively accumulates the “long-term” gradient memories of the previous batch to the following batch. This way, the learned adversarial information can be preserved and the inter-batch transferability can be improved. Experiments indicate that our method outperforms state-of-the-art ensemble attacks over multiple pixel-to-pixel vision tasks including image translation and salient region prediction. Our method successfully fools two online black-box saliency prediction systems including DeepGaze-II (Kummerer 2017) and SALICON (Huang et al. 2017). Finally, we also contribute a new repository to promote the research on adversarial attack and defense over pixel-to-pixel tasks: https://github.com/CZHQuality/AAA-Pix2pix. Zhaohui Che, Ali Borji, Guangtao Zhai, Suiyi Ling, Jing Li 0026, Patrick Le Callet |
AAAI | 3 |
| 2020 | Blurry Video Frame InterpolationabstractExisting works reduce motion blur and up-convert frame rate through two separate ways, including frame deblurring and frame interpolation. However, few studies have approached the joint video enhancement problem, namely synthesizing high-frame-rate clear results from low-frame-rate blurry inputs. In this paper, we propose a blurry video frame interpolation method to reduce motion blur and up-convert frame rate simultaneously. Specifically, we develop a pyramid module to cyclically synthesize clear intermediate frames. The pyramid module features adjustable spatial receptive field and temporal scope, thus contributing to controllable computational complexity and restoration ability. Besides, we propose an inter-pyramid recurrent module to connect sequential models to exploit the temporal relationship. The pyramid module integrates a recurrent module, thus can iteratively synthesize temporally smooth results without significantly increasing the model size. Extensive experimental results demonstrate that our method performs favorably against state-of-the-art methods. The source code and pre-trained model are available at https://github.com/laomao0/BIN. Wang Shen, Wenbo Bao, Guangtao Zhai, Li Chen 0021, Xiongkuo Min |
CVPR | 3 |
| 2020 | Self-supervised Motion Representation via Scattering Local Motion Cues
Yuan Tian 0017, Zhaohui Che, Wenbo Bao, Guangtao Zhai |
ECCV (14) | 4 |
| 2020 | Identifying Children with Autism Spectrum Disorder Based on Gaze-FollowingabstractThis paper presents a novel method to identify children with Autism Spectrum Disorder (ASD) based on the stimuli with gaze-following. Individuals with ASD are characterized by having atypical visual attention patterns, especially in social scenes. Gaze-following is considered to be a key element in understanding social scenarios, and it is reasonable to use stimuli with gaze-following to identify the children with ASD. Thus in this paper, we first construct a dataset of eye movements in gaze-following scenes for children with ASD (i.e., GazeFollow4ASD dataset), including 300 images with gaze-following information inside them and the corresponding eye movement data collected from 8 children with ASD and 10 healthy controls. We propose a novel deep neural network (DNN) model to extract discriminative features and classify children with ASD and healthy controls on single images. The proposed model shows the best performance among all compared methods on all datasets. Yi Fang 0009, Huiyu Duan, Fangyu Shi, Xiongkuo Min, Guangtao Zhai |
ICIP | 5 |
| 2020 | Automatic Region Selection For Objective Sharpness Assessment Of Mobile Device PhotosabstractMobile devices are the source of a vast majority of digital photos today. Photos taken by mobile devices generally have fairly good visual quality. When evaluating high-quality mobile device photos, people have to manually zoom in to local regions to discern the subtle difference. Understandably, a global objective quality assessment method cannot perform well on such task. Therefore, local region selection is widely recognized as a prerequisite for the following quality evaluation. Clearly, subjective regions selection suffers from the drawbacks in terms of productivity, reproducibility and optimality. In this paper, we propose an automatic local region selection algorithm for sharpness measurement of mobile device photos. Specifically, local texture statistics, depth, saliency, as well as inter-pictures difference, are used as main features to select an optimal local region, in which the sharpness is then measured. For validation, we have built a largescale database for sharpness evaluation of mobile device photos, with 100 different scenes shot by several flagship mobile phones. The experimental results show that the performance of classic sharpness evaluation algorithms can be substantially improved with the region selected by the proposed algorithm. Guangtao Zhai, Wenhan Zhu, Yucheng Zhu, Xiongkuo Min, Xiao-Ping Zhang 0002, Hua Yang 0001 |
ICIP | 2 |
| 2020 | Learning To Blindly Assess Image Quality In The Laboratory And WildabstractComputational models for blind image quality assessment (BIQA) are typically trained in well-controlled laboratory environments with limited generalizability to realistically distorted images. Similarly, BIQA models optimized for images captured in the wild cannot adequately handle synthetically distorted images. To face the cross-distortion-scenario challenge, we develop a BIQA model and an approach of training it on multiple IQA databases (of different distortion scenarios) simultaneously. A key step in our approach is to create and combine image pairs within individual databases as the training set, which effectively bypasses the issue of perceptual scale realignment. We compute a continuous quality annotation for each pair from the corresponding human opinions, indicating the probability of one image having better perceptual quality. We train a deep neural network for BIQA over the training set of massive image pairs by minimizing the fidelity loss. Experiments on six IQA databases demonstrate that the optimized model by the proposed training strategy is effective in blindly assessing image quality in the laboratory and wild, outperforming previous BIQA methods by a large margin. Weixia Zhang, Kede Ma, Guangtao Zhai, Xiaokang Yang 0001 |
ICIP | 3 |
| 2020 | A Multiple Attributes Image Quality Database for Smartphone Camera Photo Quality AssessmentabstractSmartphone is the superstar product in digital device market and the quality of smartphone camera photos (SCPs) is becoming one of the dominant considerations when consumers purchase smartphones. How to evaluate the quality of smartphone cameras and the taken photos is urgent issue to be solved. To bridge the gap between academic research accomplishment and industrial needs, in this paper, we establish a new Smartphone Camera Photo Quality Database (SCPQD2020) including 1800 images with 120 scenes taken by 15 smartphones. Exposure, color, noise and texture which are four dominant factors influencing the quality of SCP are evaluated in the subjective study, respectively. Ten popular no-reference (NR) image quality assessment (IQA) algorithms are tested and analyzed on our database. Experimental results demonstrate that the current objective models are not suitable for SCPs, and quality metrics having high correlation with human visual perception are highly needed. Wenhan Zhu, Guangtao Zhai, Zongxi Han, Xiongkuo Min, Tao Wang 0078, Xiaokang Yang 0001 |
ICIP | 2 |
| 2020 | Blind Stereoscopic Image Quality Assessment By Deep Neural Network Of Multi-Level Feature FusionabstractIn this paper, we propose an effective blind image quality assessment (BIQA) method for stereoscopic images by deep neural network (DNN) of multi-level feature fusion (MLFF) inspired by the multi-scale characteristics and binocular properties of the human visual system (HVS). Specifically, we firstly feed the left- and right-view images into a weight sharing convolutional neural network (CNN) for jointly feature extraction. To aggregate multi-level features, we concatenate the low-, middle-, and high-level feature maps of stereoscopic images to simulate the complicated visual interaction processing in the HVS. Two fully connected layers are used to build the nonlinear mapping from the highly abstract features to the quality scores of stereoscopic images. The experiments conducted on two public databases prove the validity of the proposed MLFF method. Jiebin Yan, Yuming Fang 0001, Xiongkuo Min, Yiru Yao, Guangtao Zhai |
ICME | 6 |
| 2020 | Ransp: Ranking Attention Network For Saliency Prediction On Omnidirectional ImagesabstractVarious convolutional neural network (CNN)-based methods have shown the ability to boost the performance of saliency prediction on omnidirectional images (ODIs). However, these methods are limited by sub-optimal accuracy, because not all the features extracted by the CNN model are not useful for the final fine-grained saliency prediction. Features are redundant and have negative impact on the final fine-grained saliency prediction. To tackle this problem, we propose a novel Ranking Attention Network for saliency prediction (RANSP) of head fixations on ODIs. Specifically, the part-guided attention (PA) module and channel-wise feature (CF) extraction module are integrated in a unified framework and are trained in an end-to-end manner for fine-grained saliency prediction. To better utilize the channel-wise feature map, we further propose a new Ranking Attention Module (RAM), which automatically ranks and selects these maps based on scores for fine-grained saliency prediction. Extensive experiments are conducted to show the effectiveness of our method for saliency prediction of ODIs. Dandan Zhu 0001, Yongqing Chen, Tian Han 0001, Defang Zhao, Yucheng Zhu, Qiangqiang Zhou, Guangtao Zhai, Xiaokang Yang 0001 |
ICME | 7 |
| 2020 | Joint Barcode and Text Orientation Detection Model for Unmanned Retail SystemabstractThe 1D barcode and specification text on the package of retail products contain rich information, such as date of manufacture and source of product. Acquiring this information quickly and accurately can improve the efficiency of unmanned retail system. However, traditional OCR methods are sensitive to text orientations. Based on the fact that 1D barcodes are usually aligned with text, this paper proposes a joint barcode and text orientation detection method. Our approach first determines the four vertices of arbitrarily aligned 1D barcode by a CNN-based barcode localization network. Then, a post processing module calculates the angle of alignment of text with the help of it. Experiments on combined public and self-collected dataset verify that our approach can localize barcode regions accurately and enhance the robustness for OCR in unmanned retail systems. Adnan Sharif, Jun Jia, Guangtao Zhai |
ISCAS | 4 |
| 2020 | An Improved Algorithm for Real-Time Dual-View DisplayabstractDual-view display based on spatial psychovisual modulation (SPVM) aims to present two different views on a single screen. Users with special glasses see a personal view. Users without special glasses see a shared view which can be completely irrelevant to the personal view. This dual-view display technology can be used for information security, QR hiding, education, etc. In this paper, we propose an improved algorithm called pseudo-inv algorithm, using the pseudo-inverse method to solve the optimization problem. Moreover, a new Gaussian integration window and a new method to adjust the luminance range, are implemented for a better personal view. The experiments prove that the new algorithm results in a better image quality of both the shared view and the personal view. The time cost of the new algorithm is much less than that of the previous algorithms. Zhongpai Gao, Guangtao Zhai, Zhaodi Wang 0002, Yucheng Zhu |
ISCAS | 3 |
| 2020 | A Respiratory Monitoring System in Surgical Environment Based on Fiducial Markers TrackingabstractBreathing is an important indicator of human life. Respiratory monitoring is needed in many surgical environments. To build a respiratory monitoring system for data acquisition requires the aid of acoustic or infrared sensor devices. Adding relevant equipment to the existing surgical environment not only adds the corresponding technical costs, but also may conflict with the original surgical equipment. In this paper, a respiratory monitoring system based on Aruco, an open source reality enhancement library for fiducial marker tracking, was proposed. This system has the advantages of low cost and convenient construction. The test of this monitoring system in the environment of simulated pancreatic lithotripsy verified the feasibility of its respiratory monitoring capability. Based on this system, we further explored the functions of multi-camera calibrations, breathing data extraction and posture analysis. Yichen Yu, Lianghao Hu, Guangtao Zhai |
ISCAS | 5 |
| 2020 | Weakly Supervised Pedestrian Attribute Recognition with Attention in Latent Space
Mingjun Sun, Hua Yang 0001, Guangtao Zhai |
PRCV (2) | 3 |
| 2020 | Wearable Visually Assistive Device for Blind People to Appreciate Real-world Scene and Screen ImageabstractDue to the loss of vision, the appreciation of the realworld scene and the images displayed on the screen becomes almost impossible for blind people. In an effort to meet the needs of the blind community, we develop a wearable visually assistive device to help them perceive images. With the help of various multimedia information processing technologies, the proposed device can first acquire image information through a depth camera, then implement an image-to-text transformation using image caption technology, and finally the obtained text sequence is fed back to the user via voice. In this way, blind people are able to perceive the outside world, thus creating an unprecedented experience for them. The main technical specifications of the system are: distance perception range is 0.1m to 10m; RGB field of view is 69.4°×42.5°×77°; depth field of view is 91.2°×65.5°×100.6°; maximum weight is 3.05kg. Two demo videos of the proposed navigation system which are respectively recorded for real-world scene and screen image are available at: https://doi.org/10.6084/m9.figshare.12520499.v1. Jin Ai, Menghan Hu, Guangtao Zhai, Jian Zhang 0060, Qingli Li, Wendell Q. Sun |
VCIP | 4 |
| 2020 | Special Cane with Visual Odometry for Real-time Indoor Navigation of Blind PeopleabstractIndoor navigation is urgently needed by blind people in their everyday lives. In this paper, we design an assistive cane with visual odometry based on actual requirements of the blind to aid them in attaining safe indoor navigation. Compared to the state-of-the-art indoor navigation systems, the proposed device is portable, compact, and adaptable. The main specifications of the system are: the perception range is respectively from 0.10m to 2.10m, and 0.08m to 1.60m for width and length dimensions; the maximum weight is 2.1kg; the detection range is from 0.15m and 3.00m; the cruising ability is about 8h; and the objects whose heights are below 80cm can be detected. The demo video of the proposed navigation system is available at: https://doi.org/10.6084/m9.figshare.12399572.v1. Menghan Hu, Qingli Li, Jian Zhang 0060, Xiaofeng Zhou 0002, Guangtao Zhai |
VCIP | 7 |
| 2020 | Perceptual image quality assessment: a survey
Guangtao Zhai, Xiongkuo Min |
Sci. China Inf. Sci. | 1 |
| 2020 | Unobtrusive and Automatic Classification of Multiple People's Abnormal Respiratory Patterns in Real Time Using Deep Neural Network and Depth CameraabstractRespiratory pattern is a representation of human breathing activity, which can reflect people's physical and psychological condition. Capturing the unexpected abnormal respiratory pattern unobtrusively of the patient or the potential patient has great significance. In the current work, we attempt to capitalize on depth camera and deep learning architecture to achieve the accurate and unobtrusive measurement of abnormal respiratory patterns, and the whole system can classify multiple people's respiratory patterns in a real-time manner. The challenges in this task are threefold: 1) the real-time online system means that the Region of Interest (ROI) needs to be located and tracked automatically; 2) the amount of real-world data is not enough for training to obtain the robust deep neural network; and 3) the intraclass variation is large and the outer class variation is small. Consequently, human joints tracking is applied to determine the location of subjects shoulder and chest. Based on the characteristics of actual respiratory signals, a novel and efficient respiratory simulation model (RSM) is proposed to generate abundant and high-quality training data. Finally, we apply a gated recurrent unit (GRU) neural network with bidirectional and attentional mechanisms (BI-AT-GRU) to classify six clinically significant respiratory patterns (Eupnea, Tachypnea, Bradypnea, Biots, Cheyne-Stokes, and Central-Apnea). The performance of the obtained BI-AT-GRU is tested by the data that is actually measured by the depth camera. The experimental results demonstrate that the proposed model can classify six different respiratory patterns with the accuracy, precision, recall, and F1 of 94.5%, 94.4%, 95.1%, and 94.8%, respectively. In comparative experiments, the obtained BI-AT-GRU specific to respiratory pattern classification outperforms the existing state-of-the-art, viz., BI-AT-LSTM, GRU, long short-term memory (LSTM), and BI-AT-GRU. Moreover, other experimental results indicate that the proposed online measuring system, deep neural network, and the modeling ideas have the potential to be extended to the large-scale applications, such as public places, sleep scenario, and office environment. The demo videos of the proposed system are available at: https://doi.org/10.6084/m9.figshare.11493666.v1. Yunlu Wang, Menghan Hu, Yuwen Zhou, Qingli Li, Nan Yao, Guangtao Zhai, Xiao-Ping Zhang 0002, Xiaokang Yang 0001 |
IEEE Internet Things J. | 6 |
| 2020 | Fast stripe noise removal from hyperspectral image via multi-scale dilated unidirectional convolution
Ziying Wang, Guodong Wang 0001, Zhenkuan Pan 0001, Jiahua Zhang 0001, Guangtao Zhai |
Multim. Tools Appl. | 5 |
| 2020 | DevsNet: Deep Video Saliency Network using Short-term and Long-term Cues
Yuming Fang 0001, Chi Zhang 0027, Xiongkuo Min, Hanqin Huang, Yugen Yi, Guangtao Zhai, Chia-Wen Lin |
Pattern Recognit. | 6 |
| 2020 | Unsupervised Blind Image Quality Evaluation via Statistical Measurements of Structure, Naturalness, and PerceptionabstractMost existing blind image quality assessment (BIQA) methods belong to supervised methods, which always need a large number of image samples and expensive subjective scores for training a quality prediction model. In this paper, we focus our attention on the unsupervised BIQA methods and put forward a novel unsupervised approach. The main idea of our method is to quantify the image quality degradation through measuring the structure, naturalness, and the perception quality variations of the distorted image from the pristine natural images. In specific, the structure variation is captured by the deviations of the image phase congruency and gradients distributions. The naturalness variation is characterized through the distributions variations of the locally mean subtracted and contrast normalized (MSCN) coefficients and the products of pairs of the adjacent MSCN coefficients. Compared with existing unsupervised methods, we initiatively introduce the perception quality measurement into the construction of unsupervised BIQA method, which is conducted by characterizing the prediction discrepancy between the image and its brain prediction based on the free-energy principle in the newly revealed brain theory. After feature extraction, we learn a pristine multivariate Gaussian (MVG) model with the extracted features from a set of pristine natural images. The quality of a new image is finally defined as the distance between its MVG model and the learned pristine MVG model. The extensive experiments conducted on LIVE, TID2013, CSIQ, Toyama, CID2013, and the Waterloo Exploration databases demonstrate that the proposed method achieves comparative prediction performance with the state-of-the-art BIQA methods. Yutao Liu 0002, Ke Gu 0001, Yongbing Zhang 0002, Xiu Li 0001, Guangtao Zhai, Debin Zhao, Wen Gao 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2020 | How is Gaze Influenced by Image Transformations? Dataset and ModelabstractData size is the bottleneck for developing deep saliency models, because collecting eye-movement data is very time-consuming and expensive. Most of current studies on human attention and saliency modeling have used high-quality stereotype stimuli. In real world, however, captured images undergo various types of transformations. Can we use these transformations to augment existing saliency datasets? Here, we first create a novel saliency dataset including fixations of 10 observers over 1900 images degraded by 19 types of transformations. Second, by analyzing eye movements, we find that observers look at different locations over transformed versus original images. Third, we utilize the new data over transformed images, called data augmentation transformation (DAT), to train deep saliency models. We find that label-preserving DATs with negligible impact on human gaze boost saliency prediction, whereas some other DATs that severely impact human gaze degrade the performance. These label-preserving valid augmentation transformations provide a solution to enlarge existing saliency datasets. Finally, we introduce a novel saliency model based on generative adversarial networks (dubbed GazeGAN). A modified U-Net is utilized as the generator of the GazeGAN, which combines classic "skip connection" with a novel "center-surround connection" (CSC) module. Our proposed CSC module mitigates trivial artifacts while emphasizing semantic salient regions, and increases model nonlinearity, thus demonstrating better robustness against transformations. Extensive experiments and comparisons indicate that GazeGAN achieves state-of-the-art performance over multiple datasets. We also provide a comprehensive comparison of 22 saliency models on various transformed scenes, which contributes a new robustness benchmark to saliency community. Our code and dataset are available at. Zhaohui Che, Ali Borji, Guangtao Zhai, Xiongkuo Min, Guodong Guo, Patrick Le Callet |
IEEE Trans. Image Process. | 3 |
| 2020 | A Metric for Light Field Reconstruction, Compression, and Display Quality EvaluationabstractOwning to the recorded light ray distributions, light field contains much richer information and provides possibilities of some enlightening applications, and it has becoming more and more popular. To facilitate the relevant applications, many light field processing techniques have been proposed recently. These operations also bring the loss of visual quality, and thus there is need of a light field quality metric to quantify the visual quality loss. To reduce the processing complexity and resource consumption, light fields are generally sparsely sampled, compressed, and finally reconstructed and displayed to the users. We consider the distortions introduced in this typical light field processing chain, and propose a full-reference light field quality metric. Specifically, we measure the light field quality from three aspects: global spatial quality based on view structure matching, local spatial quality based on near-edge mean square error, and angular quality based on multi-view quality analysis. These three aspects have captured the most common distortions introduced in light field processing, including global distortions like blur and blocking, local geometric distortions like ghosting and stretching, and angular distortions like flickering and sampling. Experimental results show that the proposed method can estimate light field quality accurately, and it outperforms the state-of-the-art quality metrics which may be effective for light field. Xiongkuo Min, Jiantao Zhou 0001, Guangtao Zhai, Patrick Le Callet, Xiaokang Yang 0001, Xin-Ping Guan |
IEEE Trans. Image Process. | 3 |
| 2020 | Study of Subjective and Objective Quality Assessment of Audio-Visual SignalsabstractThe topics of visual and audio quality assessment (QA) have been widely researched for decades, yet nearly all of this prior work has focused only on single-mode visual or audio signals. However, visual signals rarely are presented without accompanying audio, including heavy-bandwidth video streaming applications. Moreover, the distortions that may separately (or conjointly) afflict the visual and audio signals collectively shape user-perceived quality of experience (QoE). This motivated us to conduct a subjective study of audio and video (A/V) quality, which we then used to compare and develop A/V quality measurement models and algorithms. The new LIVE-SJTU Audio and Video Quality Assessment (A/V-QA) Database includes 336 A/V sequences that were generated from 14 original source contents by applying 24 different A/V distortion combinations on them. We then conducted a subjective A/V quality perception study on the database towards attaining a better understanding of how humans perceive the overall combined quality of A/V signals. We also designed four different families of objective A/V quality prediction models, using a multimodal fusion strategy. The different types of A/V quality models differ in both the unimodal audio and video quality prediction models comprising the direct signal measurements and in the way that the two perceptual signal modes are combined. The objective models are built using both existing state-of-the-art audio and video quality prediction models and some new prediction models, as well as quality-predictive features delivered by a deep neural network. The methods of fusing audio and video quality predictions that are considered include simple product combinations as well as learned mappings. Using the new subjective A/V database as a tool, we validated and tested all of the objective A/V quality prediction models. We will make the database publicly available to facilitate further research. Xiongkuo Min, Guangtao Zhai, Jiantao Zhou 0001, Mylène C. Q. Farias, Alan C. Bovik |
IEEE Trans. Image Process. | 2 |
| 2020 | A Multimodal Saliency Model for Videos With High Audio-Visual CorrespondenceabstractAudio information has been bypassed by most of current visual attention prediction studies. However, sound could have influence on visual attention and such influence has been widely investigated and proofed by many psychological studies. In this paper, we propose a novel multi-modal saliency (MMS) model for videos containing scenes with high audio-visual correspondence. In such scenes, humans tend to be attracted by the sound sources and it is also possible to localize the sound sources via cross-modal analysis. Specifically, we first detect the spatial and temporal saliency maps from the visual modality by using a novel free energy principle. Then we propose to detect the audio saliency map from both audio and visual modalities by localizing the moving-sounding objects using cross-modal kernel canonical correlation analysis, which is first of its kind in the literature. Finally we propose a new two-stage adaptive audiovisual saliency fusion method to integrate the spatial, temporal and audio saliency maps to our audio-visual saliency map. The proposed MMS model has captured the influence of audio, which is not considered in the latest deep learning based saliency models. To take advantages of both deep saliency modeling and audio-visual saliency modeling, we propose to combine deep saliency models and the MMS model via a later fusion, and we find that an average of 5% performance gain is obtained. Experimental results on audio-visual attention databases show that the introduced models incorporating audio cues have significant superiority over state-of-the-art image and video saliency models which utilize a single visual modality. Xiongkuo Min, Guangtao Zhai, Jiantao Zhou 0001, Xiao-Ping Zhang 0002, Xiaokang Yang 0001, Xin-Ping Guan |
IEEE Trans. Image Process. | 2 |
| 2020 | The Prediction of Saliency Map for Head and Eye Movements in 360 Degree ImagesabstractBy recording the whole scene around the capturer, virtual reality (VR) techniques can provide viewers the sense of presence. To provide a satisfactory quality of experience, there should be at least 60 pixels per degree, so the resolution of panoramas should reach 21600 × 10800. The huge amount of data will put great demands on data processing and transmission. However, when exploring in the virtual environment, viewers only perceive the content in the current field of view (FOV). Therefore if we can predict the head and eye movements which are important behaviors of viewer, more processing resources can be allocated to the active FOV. But conventional saliency prediction methods are not fully adequate for panoramic images. In this paper, a new panorama-oriented model, to predict head and eye movements, is proposed. Due to the superiority of computation in the spherical domain, the spherical harmonics are employed to extract features at different frequency bands and orientations. Related low- and high-level features including the rare components in the frequency domain and color domain, the difference between center vision and peripheral vision, visual equilibrium, person and car detection, and equator bias are extracted to estimate the saliency. To predict head movements, visual mechanisms including visual uncertainty and equilibrium are incorporated, and the graphical model and functional representation for the switch of head orientation are established. Extensive experimental results on the publicly available database demonstrate the effectiveness of our methods. Yucheng Zhu, Guangtao Zhai, Xiongkuo Min, Jiantao Zhou 0001 |
IEEE Trans. Multim. | 2 |