EDBT 2026 Demo / reviewers in the wild / expert
Xiangjie Sui
dblp:244/7475
· DBLP profile ↗
14ranked-venue papers
4as first author
13since 2021 · last 2026
0009-0008-7604-3281ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 13 · 4 first-author · 13 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021Systems, architecture and hardware · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Mitigating Perception Bias: A Training-Free Approach to Enhance LMM for Image Quality AssessmentabstractDespite the impressive performance of large multimodal models (LMMs) in high-level visual tasks, their capacity for image quality assessment (IQA) remains limited. One main reason is that LMMs are primarily trained for high-level tasks (e.g., image captioning), emphasizing unified image semantics extraction under varied quality. Such semantic-aware yet quality-insensitive perception bias inevitably leads to a heavy reliance on image semantics when those LMMs are forced for quality rating. In this paper, instead of retraining or tuning an LMM costly, we propose a training-free debiasing framework, in which the image quality prediction is rectified by mitigating the bias caused by image semantics. Specifically, we first explore several semantic-preserving distortions that can significantly degrade image quality while maintaining identifiable semantics. By applying these specific distortions to the query/test images, we ensure that the degraded images are recognized as poor quality while their semantics remain. During quality inference, both a query image and its corresponding degraded version are fed to the LMM along with a prompt indicating that the query image quality should be inferred under the condition that the degraded one is deemed poor quality. This prior condition effectively aligns the LMM’s quality perception, as all degraded images are consistently rated as poor quality, regardless of their semantic difference. Finally, the quality scores of the query image inferred under different prior conditions (degraded versions) are aggregated using a conditional probability model. Extensive experiments on various IQA datasets show that our debiasing framework could consistently enhance the LMM performance and the code will be publicly available. Baoliang Chen, Siyi Pan, Dongxu Wu, Liang Xie 0013, Xiangjie Sui, Lingyu Zhu 0006, Hanwei Zhu |
AAAI | 5 |
| 2026 | Rate Control for 360$^{\circ }$ Versatile Video Coding Based on Visual Gaze MechanismabstractIn the past few years, 360° video has started to infiltrate various aspects of daily life. Although there have been significant developments in 360° video coding technology, understanding of the human visual gaze mechanism has been somewhat overlooked. In this paper, we propose a rate control scheme for 360° Versatile Video Coding (VVC) based on a human visual gaze mechanism, targeting at improving the coding performance and bitrate accuracy. More specifically, based on the Equi-rectangular Projection (ERP) format, latitude information is systematically analyzed and a stripe-level bit allocation scheme is established, to better mitigate the projection distortion. Subsequently, the Lagrange parameter λ is further optimized with distortion dependency and identification of the visual gaze guided key Coding Tree Units (CTUs). The proposed rate control scheme is implemented on the VVC Test Model for 360° video. Experimental results show that the proposed rate control scheme can achieve BD-rate savings in terms of Weighted to Spherically uniform-Peak Signal-to-Noise Ratio (WS-PSNR) and Sphere-Peak Signal-to-Noise Ratio (S-PSNR) under the various configurations, respectively. Meanwhile, a healthier buffer status and better visual quality can be observed, further demonstrating the advantages of the proposed scheme. Zeming Zhao, Meng Wang 0017, Xiangjie Sui, Peilin Chen 0001, Xiaohai He, Shiqi Wang 0001 |
IEEE Trans. Multim. | 3 |
| 2025 | CTU-Level Rate Control with λ Optimization Based on Visual Gaze Mechanism for 360-Degree Versatile Video CodingabstractUnderstanding the human visual gaze mechanism is crucial for enhancing 360° video coding technology. This paper presents a Coding Tree Unit (CTU)-level rate control scheme with λ optimization strategy for 360° Versatile Video Coding (VVC), with the aim of enhancing rate-distortion performance and bitrate accuracy. Specifically, the Lagrange parameter λ is optimized with consideration of distortion dependency and the identification of key CTUs guided by visual gaze, which are derived from a 360° video path generation network, thoroughly integrating the characteristics of the human visual gaze. Experimental results show that the proposed scheme achieves BD-rate savings in terms of Weighted to Spherically uniform-Peak Signal-to-Noise Ratio (WS-PSNR) and Sphere-Peak Signal-to-Noise Ratio (S-PSNR) across various coding configurations. Zeming Zhao, Meng Wang 0017, Xiangjie Sui, Xiaohai He, Shiqi Wang 0001 |
ICIP | 3 |
| 2025 | Perceptual Quality Assessment of 360° Images Based on Generative Scanpath RepresentationabstractDespite substantial efforts dedicated to the design of heuristic models for omnidirectional (i.e., 360°) image quality assessment (OIQA), a conspicuous gap remains due to the lack of consideration for the diversity of viewing behaviors that leads to the varying perceptual quality of 360° images. Two critical aspects underline this oversight: the neglect of viewing conditions that significantly sway user gaze patterns and the overreliance on a single viewport sequence from the 360° image for quality inference. To address these issues, we introduce a unique generative scanpath representation (GSR) for effective quality inference of 360° images, which aggregates varied perceptual experiences of multi-hypothesis users under a predefined viewing condition. More specifically, given a viewing condition characterized by the starting point of viewing and exploration time, a set of scanpaths consisting of dynamic visual fixations can be produced using an apt scanpath generator. Following this vein, we use the scanpaths to convert the 360° image into the unique GSR, which provides a global overview of gazed-focused contents derived from scanpaths. As such, the quality inference of the 360° image is swiftly transformed to that of GSR. We then propose an efficient OIQA computational framework by learning the quality maps of GSR. Comprehensive experimental results validate that the predictions of the proposed framework are highly consistent with human perception in the spatiotemporal domain, especially in the challenging context of locally distorted 360° images under varied viewing conditions. The code will be released at https://github.com/xiangjieSui/GSR. Xiangjie Sui, Hanwei Zhu, Xuelin Liu, Yuming Fang 0001, Shiqi Wang 0001, Zhou Wang 0001 |
IEEE Trans. Image Process. | 1 |
| 2024 | Perceptual Quality Assessment of Virtual Reality Videos in the WildabstractInvestigating how people perceive virtual reality (VR) videos in the wild (i.e., those captured by everyday users) is a crucial and challenging task in VR-related applications due to complexauthenticdistortionslocalizedinspaceandtime.Existingpanoramic video databases only consider synthetic distortions, assume fixed viewing conditions, and are limited in size. To overcome these shortcomings, we construct the VR Video Quality in the Wild (VRVQW) database, containing 502 user-generated videos with diverse content and distortion characteristics. Based on VRVQW, we conduct a formal psychophysical experiment to record the scanpaths and perceived quality scores from 139 participants under two different viewing conditions. We provide a thorough statistical analysis of the recordeddata, observing significantimpact of viewing conditions on both human scanpaths and perceived quality. Moreover, we develop an objective quality assessment model for VR videos based on pseudocylindrical representation and convolution. Results on the proposed VRVQW show that our method is superior to existing video quality assessment models.We have made the database and code available at https://github.com/ limuhit/VR-Video-Quality-in-the-Wild. Wen Wen 0007, Mu Li 0005, Yiru Yao, Xiangjie Sui, Yabin Zhang 0002, Long Lan, Yuming Fang 0001, Kede Ma |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2024 | 2AFC Prompting of Large Multimodal Models for Image Quality AssessmentabstractWhile abundant research has been conducted on improving high-level visual understanding and reasoning capabilities of large multimodal models (LMMs), their image quality assessment (IQA) ability has been relatively under-explored. Here we take initial steps towards this goal by employing the two-alternative forced choice (2AFC) prompting, as 2AFC is widely regarded as the most reliable way of collecting human opinions of visual quality. Subsequently, the global quality score of each image estimated by a particular LMM can be efficiently aggregated using the maximum a posteriori estimation. Meanwhile, we introduce three evaluation criteria: consistency, accuracy, and correlation, to provide comprehensive quantifications and deeper insights into the IQA capability of five LMMs. Extensive experiments show that existing LMMs exhibit remarkable IQA ability on coarse-grained quality comparison, but there is room for improvement on fine-grained quality discrimination. The proposed dataset sheds light on the future development of IQA models based on LMMs. The codes will be made publicly available athttps://github.com/h4nwei/2AFC-LMMs. Hanwei Zhu, Xiangjie Sui, Baoliang Chen, Xuelin Liu, Peilin Chen 0001, Yuming Fang 0001, Shiqi Wang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2023 | ScanDMM: A Deep Markov Model of Scanpath Prediction for 360° ImagesabstractScanpath prediction for 360° images aims to produce dynamic gaze behaviors based on the human visual perception mechanism. Most existing scanpath prediction methods for 360° images do not give a complete treatment of the time-dependency when predicting human scanpath, resulting in inferior performance and poor generalizability. In this paper, we present a scanpath prediction method for 360° images by designing a novel Deep Markov Model (DMM) architecture, namely ScanDMM. We propose a semantics-guided transition function to learn the nonlinear dynamics of time-dependent attentional landscape. Moreover, a state initialization strategy is proposed by considering the starting point of viewing, enabling the model to learn the dynamics with the correct “launcher”. We further demonstrate that our model achieves state-of-the-art performance on four 360° image databases, and exhibit its generalizability by presenting two applications of applying scanpath prediction models to other visual tasks - saliency detection and image quality assessment, expecting to provide profound insights into these fields. Xiangjie Sui, Yuming Fang 0001, Hanwei Zhu, Shiqi Wang 0001, Zhou Wang 0001 |
CVPR | 1 |
| 2023 | Study of Spatio-Temporal Modeling in Video Quality AssessmentabstractVideo quality assessment (VQA) has received remarkable attention recently. Most of the popular VQA models employ recurrent neural networks (RNNs) to capture the temporal quality variation of videos. However, each long-term video sequence is commonly labeled with a single quality score, with which RNNs might not be able to learn long-term quality variation well: What's the real role of RNNs in learning the visual quality of videos? Does it learn spatio-temporal representation as expected or just aggregating spatial features redundantly? In this study, we conduct a comprehensive study by training a family of VQA models with carefully designed frame sampling strategies and spatio-temporal fusion methods. Our extensive experiments on four publicly available in- the-wild video quality datasets lead to two main findings. First, the plausible spatio-temporal modeling module (i. e., RNNs) does not facilitate quality-aware spatio-temporal feature learning. Second, sparsely sampled video frames are capable of obtaining the competitive performance against using all video frames as the input. In other words, spatial features play a vital role in capturing video quality variation for VQA. To our best knowledge, this is the first work to explore the issue of spatio-temporal modeling in VQA. Yuming Fang 0001, Zhaoqian Li, Jiebin Yan, Xiangjie Sui, Hantao Liu |
IEEE Trans. Image Process. | 4 |
| 2022 | Benchmarking 360° Saliency Models by General-Purpose MetricsabstractHow to effectively evaluate a model's capability to predict the visual attention of observers in 360° scenes gains interest along with the advancement of saliency prediction modeling of omnidirectional images (ODIs). So far, many general-purpose metrics from 2D saliency literature have been adopted to evaluate the 360° saliency models. However, whether they are still effective when being adopted to evaluate the 360° saliency models has not been explored. In this paper, we testify several standard saliency evaluation metrics on the 360° saliency models and comprehensively analyze their behaviors in the omnidirectional scenario. We find that 1) most metrics under-penalize false positives; 2) existing 360° datasets involve severe equator bias that few metrics can effectively penalize. We hope this case study can provide a guideline for benchmarking 360° image/video saliency models. Xiangjie Sui, Jiebin Yan, Yuming Fang 0001 |
MMSP | 2 |
| 2022 | Perceptual Quality Assessment of Omnidirectional Images as Moving Camera VideosabstractOmnidirectional images (also referred to as static 360$^{\circ }$panoramas) impose viewing conditions much different from those of regular 2D images. How do humans perceive image distortions in immersive virtual reality (VR) environments is an important problem which receives less attention. We argue that, apart from the distorted panorama itself, two types of VR viewing conditions are crucial in determining the viewing behaviors of users and the perceived quality of the panorama: the starting point and the exploration time. We first carry out a psychophysical experiment to investigate the interplay among the VR viewing conditions, the user viewing behaviors, and the perceived quality of 360$^{\circ }$images. Then, we provide a thorough analysis of the collected human data, leading to several interesting findings. Moreover, we propose a computational framework for objective quality assessment of 360$^{\circ }$images, embodying viewing conditions and behaviors in a delightful way. Specifically, we first transform an omnidirectional image to several video representations using different user viewing behaviors under different viewing conditions. We then leverage advanced 2D full-reference video quality models to compute the perceived quality. We construct a set of specific quality measures within the proposed framework, and demonstrate their promises on three VR quality databases. Xiangjie Sui, Kede Ma, Yiru Yao, Yuming Fang 0001 |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2021 | Asymmetrically distorted 3D video quality assessment: From the motion variation to perceived quality
Yuming Fang 0001, Xiangjie Sui, Jiebin Yan, Yifan Zuo 0001, Jiheng Wang, Zhaoqian Li |
Signal Process. | 2 |
| 2021 | Objective quality assessment of synthesized images by local variation measurement
Xiangjie Sui, Mengna Ding, Jiebin Yan, Yuming Fang 0001, Yifan Zuo 0001, Zuowen Tan |
Signal Process. Image Commun. | 1 |
| 2021 | Perceptual Quality Assessment for Asymmetrically Distorted Stereoscopic Video by Temporal Binocular RivalryabstractIn this paper, we propose a two-stage weighting based perceptual quality assessment framework for asymmetrically distorted stereoscopic video (SV) sequences by temporal binocular rivalry. Firstly, a traditional 2D image quality assessment (IQA) method is employed to measure spatial distortion, and the temporal distortion is evaluated by the magnitude differences between motion vectors of distorted and reference video frames. Secondly, the structural strength (SS) computed by gradient map and the motion energy (ME) computed by frame difference map are used to estimate the intensity of visual stimulus in spatial and temporal domain respectively. Then, SS and ME are considered as the importance indexes to combine the quality scores of spatial and temporal distortion to estimate perceived distortion of single-view video sequences, which is denoted as the first-stage weighting. Finally, considering that the difference of intensity of visual stimulus between two eyes results in binocular rivalry, a novel temporal binocular rivalry inspired weighting method is designed to integrate the quality scores of left- and right-views for the final visual quality prediction of SV sequences, which is denoted as the second-stage weighting. Experimental results on Waterloo-IVC SV quality databases show that several specific examples of 2D-IQA methods within the proposed framework can obtain highly competitive performance over other existing ones. Yuming Fang 0001, Xiangjie Sui, Jiheng Wang, Jiebin Yan, Jianjun Lei 0001, Patrick Le Callet |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2019 | A Spatial-Temporal Weighted Method for Asymmetrically Distorted Stereo Video Quality AssessmentabstractWe propose a 2D-TO-3D video quality prediction model for assessing the perceptual quality of asymmetrically compressed stereoscopic 3D video. In the case that the distortions between left- and right-views are significantly different, directly averaging the qualities of single-view videos to estimate 3D video perceptual quality may lead a strong prediction bias. In order to eliminate or reduce the prediction bias, we design a two-stage approach for 3D video quality prediction. Firstly, we evaluate the perceptual quality of single-view videos with the state-of-the-art 2D image/video quality assessment approaches. Secondly, we design a binocular rivalry inspired model considering both spatial and temporal perceptual information to integrate 2D video perceptual quality of both views into the assessment of 3D video quality. We validate the highly competitive performance of the proposed approach on the Waterloo-IVC 3D video quality database. Yuming Fang 0001, Xiangjie Sui, Jiheng Wang |
ISCAS | 2 |