EDBT 2026 Demo / reviewers in the wild / expert
Xuemei Zhou
dblp:99/8497
· DBLP profile ↗
14ranked-venue papers
6as first author
10since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 9 · 5 first-author · 8 since 2021Computer networks · 3 · 1 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | UVG-CWI-DQPC: Dual-Quality Point Cloud Dataset for Volumetric Video ApplicationsabstractVolumetric video is a key enabler of immersive extended reality (XR) experiences and is often represented using point clouds for their structural simplicity. However, capturing volumetric content through multi-view acquisition and depth sensing poses many challenges, such as occlusions and depth mismatches. To foster research in this field, we introduce a unique dual-quality point cloud dataset, named UVG-CWI-DQPC, which is designed to support the development of point cloud enhancement, compression, and quality assessment. Our dataset includes 12 dynamic sequences captured simultaneously by: 1) a high-end capture system producing high-fidelity point clouds with extensive processing; and 2) a consumer-grade capture system relying on affordable RGB-D cameras, lightweight processing, and open-source tools. For each sequence, our dataset provides ground-truth point clouds from the high-end capture system and raw RGB-D footage from the consumer-grade capture system, along with calibration data and tools for point cloud generation. This dual-quality setup enables direct comparison and benchmarking of algorithms for densification, occlusion removal, registration, and quality enhancement. Our dataset is publicly available under a permissive license to support reproducible research and standardization work in Moving Picture Experts Group (MPEG) and 3rd Generation Partnership Project (3GPP). Guillaume Gautier, Xuemei Zhou, Jack Jansen 0001, Louis Fréneau, Marko Viitanen, Uyen Phan, Jani Käpylä, Irene Viola 0001, Alexandre Mercat, Pablo César, Jarno Vanne |
ACM Multimedia | 2 |
| 2025 | Large multimodal models evaluation: a survey
Farong Wen, Yijin Guo, Xinyu Fang, Shengyuan Ding, Ziheng Jia, Jiahao Xiao, Ye Shen, Yushuo Zheng, Xiaorong Zhu, Yalun Wu, Ziheng Jiao, Wei Sun 0029, Zijian Chen 0001, Kaiwei Zhang, Yuqin Cao, Yue Zhou 0005, Xuemei Zhou, Juntai Cao, Wei Zhou 0021, Jinyu Cao, Ronghui Li, Yuan Tian 0017, Chunyi Li 0001, Haoning Wu 0001, Xiaohong Liu 0001, Junjun He, Yu Zhou 0016, Zesheng Wang 0004, Huiyu Duan, Yingjie Zhou 0003, Xiongkuo Min, Dongzhan Zhou, Jiezhang Cao, Xue Yang 0005, Junzhi Yu 0001, Songyang Zhang 0001, Haodong Duan, Guangtao Zhai |
Sci. China Inf. Sci. | 22 |
| 2025 | PointPCA+: A full-reference Point Cloud Quality Assessment metric with PCA-based features
Xuemei Zhou, Evangelos Alexiou, Irene Viola 0001, Pablo César |
Signal Process. Image Commun. | 1 |
| 2025 | Subjective and Objective Quality Assessment for Dynamic Point Cloud with Visual Attention in 6 DoFabstractPerceptual quality assessment of Dynamic Point Cloud (DPC) contents plays an important role in various Virtual Reality (VR) applications that involve human beings as the end user. Understanding and modeling perceptual quality assessment is greatly enriched by insights from visual attention. However, incorporating aspects of visual attention in DPC quality models is largely unexplored, as ground-truth visual attention data are scarcely available. Besides, testing methods and procedures for collecting visual attention data are still to be agreed on. This article presents a dataset containing subjective opinion scores and visual attention maps of DPCs, collected in a VR environment using eye-tracking technology. Both the quality score and eye-tracking data were collected during a subjective quality assessment experiment, in which subjects were instructed to watch and rate DPCs at various degradation levels under 6 Degrees of Freedom (DoF) inspection, using a head-mounted display. Qualitative interview analysis was also conducted after the experiment. The dataset consists of 50 DPCs, including 5 reference DPCs, with each reference encoded at 3 distortion levels using 3 different codecs (namely G-PCC, V-PCC, CWI-PCL), amounting to a total of 9 degraded version per reference. Additionally, it incorporates 1,000 gaze trials from 40 participants, yielding a total of 15,000 visual attention maps across all the DPCs. We additionally benchmark objective quality metrics originally designed for static point clouds, evaluating their performance in our dataset using two temporal pooling strategies. Furthermore, we employ the visual attention data that are retrieved during our experiment to evaluate whether the performance of widely used objective quality metrics is improved by considering subjective measurements of visual attention. This dataset establishes a link between quality assessment and visual attention within the context of DPC. Moreover, thematic analysis of the interviews helps uncover user behavior and factors impacting perceptual quality for DPC in 6 DoF. This work deepens our understanding of DPC quality assessment and visual attention, driving progress in the realm of VR experiences and perception. Xuemei Zhou, Irene Viola 0001, Evangelos Alexiou, Jack Jansen 0001, Pablo César |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2025 | Comparison of Visual Saliency for Dynamic Point Clouds: Task-free vs. Task-dependentabstractThis paper presents a Task-Free eye-tracking dataset for Dynamic Point Clouds (TF-DPC) aimed at investigating visual attention. The dataset is composed of eye gaze and head movements collected from 24 participants observing 19 scanned dynamic point clouds in a Virtual Reality (VR) environment with 6 degrees of freedom. We compare the visual saliency maps generated from this dataset with those from a prior task-dependent experiment (focused on quality assessment) to explore how high-level tasks influence human visual attention. To measure the similarity between these visual saliency maps, we apply the well-known Pearson correlation coefficient and an adapted version of the Earth Mover's Distance metric, which takes into account both spatial information and the degrees of saliency. Our experimental results provide both qualitative and quantitative insights, revealing significant differences in visual attention due to task influence. This work enhances our understanding of the visual attention for dynamic point cloud (specifically human figures) in VR from gaze and human movement trajectories, and highlights the impact of task-dependent factors, offering valuable guidance for advancing visual saliency models and improving VR perception. Xuemei Zhou, Irene Viola 0001, Silvia Rossi 0001, Pablo César |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2024 | Deciphering Perceptual Quality in Colored Point Cloud: Prioritizing Geometry or Texture Distortion?abstractPoint clouds represent one of the prevalent formats for 3D content. Distortions introduced at various stages in the point cloud processing pipeline affect the visual quality, altering their geometric composition, texture information, or both. Understanding and quantifying the impact of the distortion domain on visual quality is vital to driving rate optimization and guiding post-processing steps to improve the quality of experience. In this paper, we propose a multi-task guided multi-modality no reference metric (M3-Unity), which utilizes 4 types of modalities across attributes and dimensionalities to represent point clouds. An attention mechanism establishes inter/intra associations among 3D/2D patches, which can complement each other, yielding local and global features, to fit the highly nonlinear property of the human vision system. A multi-task decoder involving distortion type classification selects the best association among 4 modalities, aiding the regression task and enabling the in-depth analysis of the interplay between geometrical and textural distortions. Furthermore, our framework design and attention strategy enable us to measure the impact of individual attributes and their combinations, providing insights into how these associations contribute particularly in relation to distortion type. Extensive experimental results on 4 datasets consistently outperform the state-of-the-art metrics by a large margin. The code is available at https://github.com/cwi-dis/ACMMM2024-Oral. Xuemei Zhou, Irene Viola 0001, Yunlu Chen, Jiahuan Pei, Pablo César |
ACM Multimedia | 1 |
| 2024 | ComPEQ-MR: Compressed Point Cloud Dataset with Eye Tracking and Quality Assessment in Mixed RealityabstractPoint clouds (PCs) have attracted researchers and developers due to their ability to provide immersive experiences with six degrees of freedom (6DoF). However, there are still several open issues in understanding the Quality of Experience (QoE) and visual attention of end users while experiencing 6DoF volumetric videos. First, encoding and decoding point clouds require a significant amount of both time and computational resources. Second, QoE prediction models for dynamic point clouds in 6DoF have not yet been developed due to the lack of visual quality databases. Third, visual attention in 6DoF is hardly explored, which impedes research into more sophisticated approaches for adaptive streaming of dynamic point clouds. In this work, we provide an open-source Compressed Point cloud dataset with Eye-tracking and Quality assessment in Mixed Reality (ComPEQ--MR). The dataset comprises four compressed dynamic point clouds processed by Moving Picture Experts Group (MPEG) reference tools (i.e., VPCC and GPCC), each with 12 distortion levels. We also conducted subjective tests to assess the quality of the compressed point clouds with different levels of distortion. The rating scores are attached to ComPEQ--MR so that they can be used to develop QoE prediction models in the context of MR environments. Additionally, eye-tracking data for visual saliency is included in this dataset, which is necessary to predict where people look when watching 3D videos in MR experiences. We collected opinion scores and eye-tracking data from 41 participants, resulting in 2132 responses and 164 visual attention maps in total. The dataset is available at https://ftp.itec.aau.at/datasets/ComPEQ-MR/. Minh Nguyen 0006, Shivi Vats, Xuemei Zhou, Irene Viola 0001, Pablo César, Christian Timmerer, Hermann Hellwagner |
MMSys | 3 |
| 2024 | Enhancing Immersive Experiences through 3D Point Cloud Analysis: A Novel Framework for Applying 2D Visual Saliency Models to 3D Point CloudsabstractIn the new area of immersive multimedia environments, understanding and manipulating visual attention are crucial for enhancing user experience. This study introduces an innovative framework that extends traditional 2D saliency maps to the analysis of 3D point clouds, a step forward in adapting saliency prediction to more complex and immersive environments. Our framework centers on the orthographic projection of 3D point clouds onto 2D planes, enabling the application of established 2D saliency models to this novel context. We further delve into the evaluation of these models on a 3D point cloud eye-tracking dataset, exploring various projection settings and thresholding techniques to maintain the integrity of saliency information in the transition from 2D to 3D. This research not only bridges a gap in applying visual attention models to 3D data but also offers insights into the optimization of quality of experience in immersive multimedia systems. Marouane Tliba, Xuemei Zhou, Irene Viola 0001, Pablo César, Aladine Chetouani, Giuseppe Valenzise, Frédéric Dufaux |
QoMEX | 2 |
| 2023 | QAVA-DPC: Eye-Tracking Based Quality Assessment and Visual Attention Dataset for Dynamic Point Cloud in 6 DoFabstractPerceptual quality assessment of Dynamic Point Cloud (DPC) contents plays an important role in various Virtual Reality (VR) applications that involve human beings as the end user, understanding and modeling perceptual quality assessment is greatly enriched by insights from visual attention. However, incorporating aspects of visual attention in DPC quality models is largely unexplored, as ground-truth visual attention data is scarcely available. This paper presents a dataset containing subjective opinion scores and visual attention maps of DPCs, collected in a VR environment using eye-tracking technology. The data was collected during a subjective quality assessment experiment, in which subjects were instructed to watch and rate DPCs at various degradation levels under 6 degrees-of-freedom inspection, using a head-mounted display. The dataset comprises 5 reference DPC contents, with each reference encoded at 3 distortion levels using 3 different codecs, amounting to a total of 9 degraded DPC contents. Moreover, it includes 1,000 gaze trials from 40 participants, resulting in 15,000 visual attention maps in total. The curated dataset can serve as authentic benchmark data for assessing the performance of objective DPC quality metrics. Additionally, it establishes a link between quality assessment and visual attention within the context of DPC. This work deepens our understanding of DPC quality and visual attention, driving progress in the realm of VR experiences and perception. Xuemei Zhou, Irene Viola 0001, Evangelos Alexiou, Jack Jansen 0001, Pablo César |
ISMAR | 1 |
| 2022 | A Dynamic Resource Allocation Model Based on SMDP and DRL Algorithm for Truck Platoon in Vehicle NetworkabstractThe rapid development of self-driving cars and breakthroughs in key technologies have made the truck platoon possible. In addition to reducing truck fuel consumption and air pollution by reducing air resistance, effective platoon strategies can also maximize highway throughput while improving driving safety. However, the truck platoon strategy’s current resource allocation model is still in the preliminary research stage. Therefore, inspired by the successful experience of deep reinforcement learning (DRL) in solving resource allocation problems, this article proposes a dynamic resource allocation model for the truck platoon based on the semi-Markov decision process (SMDP) and DRL, which is used to maximize system revenue when considering the resource cost and income balance of the transportation system. Precisely, the proposed method first models the process of controlling the dynamic in and out of the truck platoon as SMDP. The action value in a specific state obtained by the planning algorithm is used as a DRL sample for model training. Finally, the SMDP is optimized through the trained model to obtain a truck platoon resource that approximates the optimal strategy distribution plan. The experimental results show that compared with the traditional greedy algorithm, value iteration, and${Q}$-learning scheme concerning solving the dynamic resource allocation model of the truck platoon, the Deep${Q}$-Network (DQN) used in this article can reduce the probability of request processing delay while causing the system to obtain higher rewards. Hongbin Liang, Shuya Zhou, Xiaobo Liu 0002, Fangfang Zheng, Xintao Hong, Xuemei Zhou, Lian Zhao |
IEEE Internet Things J. | 6 |
| 2020 | Content-aware Hybrid Equi-angular Cubemap Projection for Omnidirectional Video CodingabstractOmnidirectional video is required to be projected from the Three-Dimensional (3D) sphere to a Two-Dimensional (2D) plane before compression due to its spherical characteristics. Therefore, various projection formats have been proposed in recent years. However, these existing projection methods have problems of either oversampling or discontinuous boundary, which penalize the coding performance. Among them, Hybrid Equiangular Cubemap (HEC) projection has achieved significant coding gains by keeping boundary continuity when compared with Equi-Angular Cubemap (EAC) projection. However, the parameters of its mapping function are fixed and cannot adapt to the video contents, which results in non-uniform sampling in certain regions. To address this limitation, a projection method named Content-aware HEC (CHEC) is presented in this paper. In particular, these parameters of mapping function are adaptively achieved by minimizing the projection conversion distortion. Additionally, an omnidirectional video coding framework with adaptive parameters of mapping function is proposed to effectively improve the coding performance. Experimental results show that the proposed scheme achieves 8.57% and 0.11% bit rate reduction on average in terms of End-to-End Weighted to Spherically uniform Peak Signal to Noise Ratio (E2E WS-PSNR) when compared with Equi-Rectangular Projection (ERP) and HEC projections, respectively. Jinyong Pi, Yun Zhang 0002, Linwei Zhu, Xinju Wu, Xuemei Zhou |
VCIP | 5 |
| 2016 | Energy-Efficient Joint Sensing Duration, Detection Threshold, and Power Allocation Optimization in Cognitive OFDM SystemsabstractThis paper investigates an energy efficiency optimization problem in cognitive orthogonal frequency division multiplexing systems. The goal is to maximize the energy efficiency by adapting the sensing duration, detection threshold, and transmit power to the constraints of the energy consumption of the secondary network and the interference to the primary network in a statistical manner. First, the case of identical detection threshold for all subcarriers is considered. In order to circumvent the intractability of the resulting problem, an alternate iteration framework is proposed to iteratively solve the three decoupled subproblems: sensing duration optimization, detection threshold optimization, and power allocation optimization. By exploiting the characteristics of each subproblem, the proposed framework is proved to be convergent. Then, the case with individual detection threshold for each subcarrier is explored. By proving that the optimal detection threshold is the root of a quadratic equation with one unknown variable, the proposed framework can be applied with minor modification. Simulation results show that the proposed alternating optimization framework can approach rapidly to the optimal solution, with less than 1% gap. Compared with the existing schemes, both the cases with identical and individual detection thresholds can achieve a considerable energy efficiency gain, with the latter further outperforming the former. Wenjun Xu 0001, Xuemei Zhou, Chia-han Lee, Zhiyong Feng 0001, Jiaru Lin |
IEEE Trans. Wirel. Commun. | 2 |
| 2015 | Energy Efficiency Optimization in OFDM-Based Cognitive Radio Systems: Impact of Power AmplifiersabstractNowadays energy efficiency (EE) of wireless communication systems has become a hot issue, yet the nonlinear effect and inefficiency of power amplifier (PA) have posed practical challenges for system designs to achieve high EE. However, most of previous work only considered linear PA. In this paper, we studies EE optimization concerning with I-way Doherty PA which has widespread use with I = 1 and I = 2 in orthogonal frequency division multiplex (OFDM)-based cognitive radio (CR) systems. The aim is to maximize EE with nonlinear PA subject to the total power budget, the interference constraint and the minimum rate requirement. Other than traditional methods, the problem is hard to solve directly due to the nonlinearity of PA, and a bisection search method tailored for nonlinear PA is adopted to achieve the sub-optimal solution. Numerical results demonstrate that if the PA's nonlinearity is not considered, the EE performance of OFDM-based CR systems can be severely overestimated by 34% at P max out = 80 W for 1-way Doherty PA. Meanwhile, the EE performance can be enhanced by approximate 32% if 2-way Doherty PA is utilized instead of 1-way Doherty PA. Xinxin Shi, Wenjun Xu 0001, Xuemei Zhou, Jiaru Lin |
VTC Spring | 3 |
| 2015 | Energy-Efficient Power Loading with Intercarrier and Intersymbol Interference Considerations for Cognitive OFDM SystemsabstractThis paper investigates the energy-efficient power loading with intercarrier and intersymbol interference considerations for OFDM-based cognitive systems. The objective is to maximize the energy efficiency (EE) as well as to balance the tradeoff between intercarrier interference (ICI) and intersymbol interference (ISI) by jointly optimizing the subcarrier bandwidth and power allocation in a mobile scenario, under the power budget and the interference constraint. First, the primal problem is converted into a convex optimization problem by fractional programming. Then, the Lagrange dual function and the sub-gradient method are adopted to achieve the optimal power allocation and the golden section method is employed to search for the optimal subcarrier number (i.e, subcarrier bandwidth). Numerical results show that the proposed algorithm can realize ICI control by choosing an optimal subcarrier number, and simultaneously the EE can be significantly improved by 139%. Xuemei Zhou, Wenjun Xu 0001, Xinxin Shi, Jiaru Lin |
VTC Spring | 1 |