Jiarun Song

dblp:119/1372 · DBLP profile ↗
← Back
19ranked-venue papers
9as first author
12since 2021 · last 2026
0000-0001-6718-4201ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 16 · 9 first-author · 9 since 2021Systems, architecture and hardware · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Vision-Language Modality Prompt Classification Method for Image Quality Assessment
Yongkang Hou, Jiarun Song
ISCAS2
2026 Multiple Candidates Derivation of Dominant Intra Prediction Mode in Beyond VVC
Junyan Huo, Yanzhuo Ma, Wei Zhang 0072, Fuzheng Yang 0001, Jiarun Song
ISCAS6
2026 Query-Based Hierarchical Fusion of Multi-Level Visual Features for AIGC Image Quality Assessment
Linghe Meng, Jiarun Song
QoMEX2
2026 Perceptual Quality Assessment of Trisoup-Lifting Encoded 3D Point Clouds
abstract
No-reference bitstream-layer point cloud quality assessment (PCQA) can be deployed without full decoding at any network node to achieve real-time quality monitoring. In this work, we develop the first PCQA model dedicated to Trisoup-Lifting encoded 3D point clouds by analyzing bitstreams without full decoding. Specifically, we investigate the relationship among texture bitrate per point (TBPP), texture complexity (TC) and texture quantization parameter (TQP) while geometry encoding is lossless. Subsequently, we estimate TC by utilizing TQP and TBPP. Then, we establish a texture distortion evaluation model based on TC, TBPP and TQP. Ultimately, by integrating this texture distortion model with a geometry attenuation factor, a function of trisoupNodeSizeLog2 (tNSL), we acquire a comprehensive NR bitstream-layer PCQA model named streamPCQ-TL. In addition, this work establishes a database named WPC6.0, the first PCQA database dedicated to Trisoup-Lifting encoding mode, encompassing 400 distorted point clouds with 4 geometry multiplied by 5 texture distortion levels. Experiment results on M-PCCD, ICIP2020 and the proposed WPC6.0 database suggest that the proposed streamPCQ-TL model exhibits robust and notable performance in contrast to existing advanced PCQA metrics, particularly in terms of computational cost.
Juncheng Long, Honglei Su, Qi Liu 0029, Hui Yuan 0001, Wei Gao 0003, Jiarun Song, Zhou Wang 0001
IEEE Trans. Vis. Comput. Graph.6
2026 Latency Effects on Multi-Dimensional QoE in Networked VR Whiteboards
abstract
Networked virtual reality (NVR) whiteboards are increasingly important for enabling geographically dispersed users to engage in real-time idea sharing, collaborative design, and discussion. However, latency caused by network limitations, rendering delays, or synchronization issues can significantly degrade the Quality of Experience (QoE) in whiteboard collaboration. To systematically investigate the impact of latency, this study classified QoE into pragmatic and hedonic aspects, each comprising multiple sub-dimensions. Controlled experiments were conducted to identify the sub-dimensions most affected by latency, which were then adopted as the primary QoE indicators, with the aim of uncovering the processes and mechanisms through which latency shapes QoE. Building on this, we further examined how these impacts vary across different collaboration modes, namely sequential collaboration (SC) for structured design workflows and free collaboration (FC) for open discussion. We also compared two VR whiteboard types, one with avatars (VR+) and the other without avatars (VR), and included a traditional PC-based whiteboard as a baseline. This multidimensional design enables a comprehensive evaluation of latency's impact on QoE across collaboration modes and platforms, providing practical guidance for optimizing NVR whiteboard systems under real-world network and system constraints.
Jiarun Song, Yongkang Hou, Fuzheng Yang 0001
IEEE Trans. Vis. Comput. Graph.1
2025 Subjective Fidelity Assessment of Audio- and Video-Driven Talking Head Generation Methods
abstract
Audio- and Video-Driven Talking Head Generation methods have attracted considerable research interest due to recent advances in Artificial Intelligence Generated Content (AIGC) technologies. In such approaches, a single image is artificially animated by leveraging audio and/or motion features extracted from video sources. Despite notable progress, current performance assessments rely primarily on traditional objective metrics, often neglecting subjective evaluation aspects. To address this issue, we propose in this paper a subjective fidelity assessment of recent Audio- and/or Video-Driven Talking Head Generation methods. This study aims to assess how accurately and convincingly the generated video reproduces the visual and behavioral characteristics of a real human face, as well as how closely the video aligns with expected natural human expressions, movements, and/or audio synchronization. In order to provide a detailed assessment of the fidelity in the context of talking heads, our study focuses on six key criteria: Overall Fidelity, Gaze Fidelity, Audio-Video Sync Fidelity, Head Pose Fidelity, Expression Fidelity, and Overall Visual Quality. Experiments results reveal a nuanced picture of the fidelity in this context, where the performance varies significantly depending on the video content itself as well as how the animation is generated, highlighting the needs for further research. This research represents an initial step towards the evaluation of Audio- and Video-Driven generative image animation methods for Talking heads while offering insights for improving the accuracy and realism of those techniques. The dataset and corresponding results are available at https://github.com/a-trioux/Subjective-Fidelity-Assessment-Talking-Head.
Anthony Trioux, Yusong Gao, Jiarun Song, Faming Ma, Fuzheng Yang 0001
ICASSP3
2024 Perceptual Quality Evaluation for Faster Playback Videos
abstract
Faster playback, as a common feature of video services nowadays, enables users to watch a video in less time, while maintaining a complete and coherent viewing experience. However, it is still unclear what kind of perceptual quality users will have at different video playback speeds. In this paper, a number of experiments were carried out to analyze the relationship between user’s perceptual quality and the playback speed for different types of video content. Then, combining the characteristics of temporal complexity of video content, an objective evaluation model was proposed to predict the perceptual quality of faster playback videos. Experimental results show that the proposed method can effectively predict the perceptual quality of videos under different playback speeds. This proposed model can be used for video service providers to recommend optimal playback speed limits for videos across a variety of scenarios, so as to ensure user’s viewing experience.
Jiarun Song
ICASSP1
2024 Refined Chroma From Luma Prediction in AV1 Based on Color Component Grouping
abstract
Chroma from luma (CfL) in AOMedia Video 1 (AV1) utilizes the correlation between color components to derive the predicted chroma from the reconstructed luma. The mean of the predicted chroma, i.e., dc, is set to the chroma average of the neighboring regions. Since the neighboring regions have limited reference samples, the distribution of the coding block is hard to be predicted from the neighboring regions, resulting in an error in the predicted dc. A new feature of CfL is that the luma of the coding block has been reconstructed. Using the luma information, the ground-truth distribution of the coding block can be built. Based on this observation, a refined dc prediction is proposed based on color component grouping (GDC). We design an adaptive grouping scheme and use the number of luma and the chroma average in each group to derive a refined dc. The offline experiment verifies that the proposed method provides a dc with high prediction accuracy. Compared with the libaom of AV1, the proposed GDC achieves BD-rate reductions of 0.15%, 0.90%, and 0.98% for the Y, Cb, and Cr components. With the increased group numbers, additional coding gains can be provided. The proposed GDC can be implemented with a small complexity, which is hardware-friendly.
Junyan Huo, Zhenyao Zhang, Jiarun Song, Yanzhuo Ma, Fuzheng Yang 0001
IEEE Trans. Ind. Informatics3
2023 Effect of latency on social presence in traditional video conference and VR conference: a comparative study
abstract
Virtual reality (VR) conference, as a typical social VR application, has gained popularity in recent years. It offers users located at different locations a fully immersive experience and a sense of togetherness. However, the remote communication also introduces inevitable latencies, which may adversely affect the so-called social presence. There is still a lack of research on the effect of latency on social presence. To fill the gap, this paper aims to examine the impact of latency on social presence of VR conference and contrast it with that of traditional video conference. Here, the social presence is measured using the Networked Minds Social Presence Inventory (NMSPI). We design and conduct two conversation-based subjective tests for both types of conference and compare the impact of the latency based on the test results. The conclusions of these studies can be used as guidelines for VR service providers to optimize their conference systems.
Jiarun Song, Anthony Trioux, Yusong Gao, Fuzheng Yang 0001
VCIP2
2023 The Impact of Black Edge Artifact on QoE of the FOV-Based Cloud VR Services
abstract
Cloud virtual reality (Cloud VR) services usually introduce high latency in rendering and streaming, resulting in a mismatch between the visual and vestibular systems, causing user sickness and dizziness during the service. To solve this problem, asynchronous rendering technology is usually used to provide smooth viewing. The asynchronous solution, on the other hand, will introduce another “black edge” (BE) artifact in the service, which frequently appears at the viewport's boundary with a black area when users turn their heads. This unwanted BE artifact also has an impact on the user's quality of experience (QoE). In this paper, we investigated the impact of the BE artifact on the user's QoE of the field of view (FOV) Cloud VR gaming services. The appearance of the BE artifact during the playing period was regarded as a series of BE events and the impact of BE artifact on the users’ QoE was evaluated by accumulating the influence of the BE events during the whole playing period. More specifically, the user's QoE affected by a single BE event was first evaluated by combining the area ratio and duration of the BE artifact. Then, the QoE affected by multiple BE events was analyzed, where the cumulative influence of the previous BE events on the user's current QoE was evaluated. Finally, a unified event-based evaluation model was proposed to predict the user's time-varying QoE at any point in time. Experimental results showed that the proposed model performed exceptionally well in predicting the impact of BE artifact on the user's QoE.
Jiarun Song, Xionghui Mao, Fuzheng Yang 0001
IEEE Trans. Multim.1
2021 The Impact of Black Edge Artifact on User Experience for the Interactive Cloud VR Services
abstract
When user turns head in an interactive Cloud VR service, such as the Cloud VR games, the black edge artifact often appears at the boundary of the viewport due to the inevitable latency, which significantly affects the user experience. In this paper, we focused on the interactive Cloud VR gaming services and investigated the influence of the black edge artifact on the user experience. The appearance of the black edge artifact during the playing period was regarded as a series of black edge events. The user experience on a single black edge event was first evaluated combining the influential factors of the area ratio and duration of the black edge artifact. Then, a pooling strategy was proposed to evaluate the user’s overall experience. Experimental results showed that the proposed model can accurately predict the user experience that influenced by the black edge artifact for the Cloud VR services. This work can be served as a guideline for service providers and network operators to improve their services.
Jiarun Song, Jianquan Zhou, Xionghui Mao, Fuzheng Yang 0001
ICME1
2021 Parametric Model for Video Streaming Services With Different Spatial and Temporal Resolutions
abstract
Parametric models of video quality are designed for service and network planning as well as video quality monitoring. They are widely applied in a broad range of applications especially when video streams are encrypted or even unavailable at all. Designing metrics in these models remains challenge due to limited available information. In this paper, a spatio-temporal resolution-adaptive parametric (STRAP) model is proposed to evaluate the quality of video streaming services considering the spatial and temporal resolutions. This work serves as a follow-up study for ITU Rec. P.1203.1 that we were previously involved. The relationship between the content complexity and the spatial and temporal resolutions are analyzed and incorporated into the proposed model. Moreover, the effect of video up/down-scaling in display devices on the perceived video quality is further taken into consideration. Experimental results showed that the proposed model can be used as a reliable indicator for video streaming providers to improve their services performance.
Jiarun Song, Fuzheng Yang 0001, Wei Zhang 0072, Zhibin Ma
IEEE Trans. Circuits Syst. Video Technol.1
2020 Towards the Instant Tile-Switching for Dash-Based Omnidirectional Video Streaming: Random Access Reference Frame
abstract
In this work, we propose an instant tile-switching mechanism for the tile-based omnidirectional DASH video streaming, which solves the problem that the client cannot switch to the high quality tiles immediately as the viewport changes. More specifically, we present the random access reference frame (RARF) on the server for each P frame as the random access point. When users change the viewport, the corresponding P frame at the viewport change moment will be replaced by RARF to transferred to the client for an immediately quality switching. By using the proposed instant tile-switching mechanism, the client can acquire the high quality video in the new viewport at any time and provide users a much better quality of experience.
Mingyi Yang, Wenjie Zou, Jiarun Song, Fuzheng Yang 0001
ICME3
2020 A Fast FoV-Switching DASH System Based on Tiling Mechanism for Practical Omnidirectional Video Services
abstract
With the development of multimedia technologies and virtual reality display devices, omnidirectional videos have gained popularity nowadays. To reduce the bandwidth requirement for omnidirectional video transmission, tile-based viewport adaptive streaming methods have been proposed in the literatures. Challenges related to decoding the tiles simultaneously with limited number of decoders, and ensuring user's viewing experience during the viewport switch are still to be solved. In this paper, a two-layer fast viewport switching dynamic adaptive streaming over HTTP (DASH) system based on tiling mechanism is proposed, which incorporates the viewing trajectory of end users. To deal with the simultaneously decoding problem, an open group of picture (GOP) technique is proposed to enable merging different types of tiles into a composite stream at the client side. To reduce the quality recovery duration after the viewport change, a fast-switching strategy is also proposed. Moreover, considering the priorities of different types of chunks, a download strategy is further proposed to adapt the bandwidth fluctuations and viewport changes. Experimental results showed that the proposed system can significantly reduce the recovery duration of high quality video by approximately 90%, which can provide a better viewing experience to end users.
Jiarun Song, Fuzheng Yang 0001, Wei Zhang 0072, Wenjie Zou, Yuqun Fan, Peiyun Di
IEEE Trans. Multim.1
2018 Initial Perceived Quality Analysis for DASH Video Streaming
abstract
Dynamic Adaptive Streaming over HTTP (DASH) uses pre-buffering mechanism to ensure smoothly playback of video streaming. However, it will lead to an additional initial loading delay, which significantly affects the initial perceived quality and may make the users give up the video services. In this paper, two kinds of subjective test methods were designed to obtain the initial perceived quality from the perspective of user perception process and rating habit, respectively. The initial perceived quality was rated at different stages of viewing period in these two methods and the accuracy of the initial perceived quality obtained by these two methods was compared. Then, the relationship between the initial loading delay and user quit rate was analyzed, which can be served as a guideline for service providers and network operators to improve their services.
Jiarun Song, Ruihuan Wang, Fuzheng Yang 0001, Zhibin Ma, Qiyong Zhao
VCIP1
2018 Event-Based Perceptual Quality Assessment for HTTP-Based Video Streaming With Playback Interruption
abstract
The popularity of HTTP-based video streaming services has been increasing in recent years. The quality of HTTP-based video streaming services is measured by the user's quality of experience (QoE). Perceptual quality (PQ), a strong indicator of the QoE, has been widely studied in the literature. Specifically, previous studies primarily focused on modeling the overall perceptual quality at the end of the viewing process. The specific influence of each event during the viewing process was neglected. Since the PQ is a function of time, which models the human perception of audiovisual services, this contribution focuses on investigating the time-varying feature of human perception. In this paper, the viewing timeline of subjects is divided into rebuffering and playback. An event-based perceptual quality assessment (EPQ) framework is introduced, which can assess the PQ of interrupted HTTP-based video streaming at any point in time. To understand the human perception of video streams with playback interruptions, a subjective quality assessment experiment was designed to obtain the PQ at a series of points in time. A between-subjects design was adopted in which each subject can view video content without repetition. Based on the experimental results, a PQ assessment model was proposed to predict the change in PQ (i.e., ΔPQ) after each rebuffering and playback. The overall PQ at any point in time is calculated as the summation of the ΔPQs of all the previous events. The experimental results demonstrated that the EPQ-based model outperforms the rebuffering components contained in the ITU-T Rec. P.1201.1, P.1201 Amd. 2 and P.1203.3 in our database.
Wenjie Zou, Fuzheng Yang 0001, Jiarun Song, Shuai Wan, Wei Zhang 0072, Hong Ren Wu
IEEE Trans. Multim.3
2017 Visual Experience Analysis for Polygon Mesh on Different Display Devices
Youguang Yu, Jiarun Song, Fuzheng Yang 0001
DCC2
2017 Parametric Planning Model for Video Quality Evaluation of IPTV Services Combining Channel and Video Characteristics
abstract
Parametric planning models are designed for estimating the video quality, which can be applied to effective planning, implementation, and management of network video applications and communication networks. However, different from the bitstream-based evaluation models, the planning models are not allowed to exploit the video streams, with only limited information available for use, i.e., a few general parameters predetermined by the service providers and network operators. In this paper, a parametric planning model combining channel and video characteristics is proposed to estimate the video distortion caused by packet loss for Internet protocol television (IPTV) services. More specifically, the probability distribution of the channel states is determined by detailed analysis of the channel characteristics. Then, considering the influence of burst packet loss and the temporal dependence between frames, several sequence-level and frame-level parameters for video quality evaluation are derived from the perspective of the probability distribution of the channel states. Utilizing these parameters, the proposed model approximates the video quality considering the effects of direct packet loss and error propagation. Experimental results show that the proposed model has a superior performance for video quality estimation than the three commonly used parametric planning models.
Jiarun Song, Fuzheng Yang 0001, Yicong Zhou
IEEE Trans. Multim.1
2016 QoE Evaluation of Multimedia Services Based on Audiovisual Quality and User Interest
abstract
Quality of experience (QoE) has significant influence on whether or not a user will choose a service or product in the competitive era. For multimedia services, there are various factors in a communication ecosystem working together on users, which stimulate their different senses inducing multidimensional perceptions of the services, and inevitably increase the difficulty in measurement and estimation of the user's QoE. In this paper, a user-centric objective QoE evaluation model (QAVIC model for short) is proposed to estimate the user's overall QoE for audiovisual services, which takes account of perceptual audiovisual quality (QAV) and user interest in audiovisual content (IC) amongst influencing factors on QoE such as technology, content, context, and user in the communication ecosystem. To predict the user interest, a number of general viewing behaviors are considered to formulate the IC evaluation model. Subjective tests have been conducted for training and validation of the QAVIC model. The experimental results show that the proposed QAVIC model can estimate the user's QoE reasonably accurately using a 5-point scale absolute category rating scheme.
Jiarun Song, Fuzheng Yang 0001, Yicong Zhou, Shuai Wan, Hong Ren Wu
IEEE Trans. Multim.1