VLDB 2026 Research / reviewers in the wild / expert
Fang-Yi Chao
dblp:231/2651
· DBLP profile ↗
6ranked-venue papers
4as first author
5since 2021 · last 2023
0000-0002-1212-1139ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 6 · 4 first-author · 5 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | Symmetric Geometry Coding for Static MeshesabstractMesh compression plays an important role in the efficient storage and transmission of 3D models. While techniques like Video-based Dynamic Mesh Coding (V-DMC) mainly exploit the local structure of mesh for enhanced compression efficiency, they often overlook the global structure as well as a bias towards scanned 3D meshes. This work explores the unique characteristics of computer-generated (CG) meshes and presents a novel Symmetric Geometry Coding (SGC) tool. SGC takes advantage of the global symmetric property to encode just half of the symmetric mesh and predicts the other. We further extended SGC to accommodate partial symmetric meshes by processing submesh with multiple symmetries. Our results demonstrate that SGC achieves up to 40.90%/40.41% bitrate reduction compared to V-DMC 1.1 in D1/D2-PSNR metrics for symmetric CG meshes. Thuong Nguyen-Canh, Fang-Yi Chao, Xiaozhong Xu, Shan Liu 0001 |
VCIP | 2 |
| 2023 | PAV-SOD: A New Task towards Panoramic Audiovisual Saliency DetectionabstractObject-level audiovisual saliency detection in 360° panoramic real-life dynamic scenes is important for exploring and modeling human perception in immersive environments, also for aiding the development of virtual, augmented, and mixed reality applications in fields such as education, social network, entertainment, and training. To this end, we propose a new task, p anoramic a udio v isual s alient o bject d etection, ( PAV-SOD 1 ), which aims to segment the objects grasping most of the human attention in 360° panoramic videos reflecting real-life daily scenes. To support the task, we collect PAVS10K , the first p anoramic video dataset for a udio v isual s alient object detection, which consists of 67 4K-resolution equirectangular videos with per-video labels including hierarchical scene categories and associated attributes depicting specific challenges for conducting PAV-SOD , and 10,465 uniformly sampled video frames with manually annotated object-level and instance-level pixel-wise masks. The coarse-to-fine annotations enable multi-perspective analysis regarding PAV-SOD modeling. We further systematically benchmark 13 state-of-the-art salient object detection (SOD)/video object segmentation (VOS) methods based on our PAVS10K . Besides, we propose a new baseline network, which takes advantage of both visual and audio cues of 360° video frames by using a new conditional variational auto-encoder (CVAE). Our C VAE-based a udio v isual net work, namely, CAV-Net , consists of a spatial-temporal visual segmentation network, a convolutional audio-encoding network, and audiovisual distribution estimation modules. As a result, our CAV-Net outperforms all competing models and is able to estimate the aleatoric uncertainties within PAVS10K . With extensive experimental results, we gain several findings about PAV-SOD challenges and insights towards PAV-SOD model interpretability. We hope that our work could serve as a starting point for advancing SOD towards immersive media. Yi Zhang 0076, Fang-Yi Chao, Wassim Hamidouche, Olivier Déforges |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2022 | Privacy-Preserving Viewport Prediction using Federated Learning for 360° Live Video StreamingabstractPredicting the user's viewport scanpath is an essential task for 360° viewport-based adaptive streaming. It informs the system which parts of content should be streamed with high quality for bandwidth saving over the best-effort Internet. However, in light of growing privacy concerns among consumers and increasingly strict data privacy legislation, user data collection and storage have been constrained. This paper proposes a novel privacy-preserving framework employing Federated Learning (FL) for online viewport prediction in a live 360° video streaming scenario. In this framework, the user data is only collected and processed on the client-side in the current viewing session and not shared with external parties, e.g., servers, and other clients. We evaluated the framework over a widely-used dataset and measure the computation and transmission time of the proposed streaming system. The experiments show that our framework provides high prediction accuracy and achieves real-time computation requirements of live video streaming. On privacy preservation, our results demonstrate that in a tile-based 360° video streaming system, the user identification rate can be decreased by 18.11 percentage points in 4 x 3 tiles per frame and 9.65 percentage points in 16x9 tiles per frame. The code will be publicly available to further contribute to the community. Fang-Yi Chao, Cagri Ozcinar, Aljoscha Smolic |
MMSP | 1 |
| 2021 | Transformer-based Long-Term Viewport Prediction in 360° Video: Scanpath is All You NeedabstractVirtual Reality (VR) multimedia technology has dramatically advanced in recent years. Its immersive and interactive natures enable users to view any direction in 360° content freely. Users do not see the entire 360° content at a glance, but only a portion in the viewport. Viewport-based adaptive streaming, which streams only the user’s viewport of interest with high quality, has emerged as the primary technique to save bandwidth over the best-effort Internet. Thus, users’ viewport prediction in the forthcoming seconds becomes an essential task for informing the streaming decisions in the VR system. Various viewport prediction methods based on deep neural networks have been proposed. However, typically they are composed of complex Convolutional Neural Networks (CNN) and Recurrent Neural Networks (RNN) that require heavy computation. To achieve high prediction accuracy in limited computation time in a streaming system, we propose a new transformer-based architecture, named 360° Viewport Prediction Transformer (VPT360), that only leverages the past viewport scanpath to predict a user’s future viewport scanpath. We evaluate VPT360 over three widely-used datasets and compare the computation complexity with the state-of-the-art methods. The experiments show that our VPT360 provides the highest accuracy for short-term and long-term prediction and achieves the lowest computation complexity. The code is publicly available at https://github.com/FannyChao/VPT360 to further contribute to the community. Fang-Yi Chao, Cagri Ozcinar, Aljoscha Smolic |
MMSP | 1 |
| 2021 | A Multi-FoV Viewport-Based Visual Saliency Model Using Adaptive Weighting Losses for 360$^\circ$ Imagesabstract360$^\circ$media allows observers to explore the scene in all directions. The consequence is that the human visual attention is guided by not only the perceived area in the viewport but also the overall content in 360$^\circ$. In this paper, we propose a method to estimate the 360$^\circ$saliency map which extracts salient features from the entire 360$^\circ$image in each viewport in three different Field of Views (FoVs). Our model is first pretrained with a large-scale 2D image dataset to enable the interpretation of semantic contents, then fine-tuned with a relative small 360$^\circ$image dataset. A novel weighting loss function attached with stretch weighted maps is introduced to adaptively weight the losses of three evaluation metrics and attenuate the impact of stretched regions in equirectangular projection during training process. Experimental results demonstrate that our model achieves better performance with the integration of three FoVs and its diverse viewport images. Results also show that the adaptive weighting losses and stretch weighted maps effectively enhance the evaluation scores compared to the fixed weighting losses solutions. Comparing to other state of the art models, our method surpasses them on three different datasets and ranks the top using 5 performance evaluation metrics on the Salient360! benchmark set. The code is available athttps://github.com/FannyChao/MV-SalGAN360. Fang-Yi Chao, Lu Zhang 0037, Wassim Hamidouche, Olivier Déforges |
IEEE Trans. Multim. | 1 |
| 2020 | Towards Audio-Visual Saliency Prediction for Omnidirectional Video with Spatial AudioabstractOmnidirectional videos (ODVs) with spatial audio enable viewers to perceive 360° directions of audio and visual signals during the consumption of ODVs with head-mounted displays (HMDs). By predicting salient audio-visual regions, ODV systems can be optimized to provide an immersive sensation of audio-visual stimuli with high-quality. Despite the intense recent effort for ODV saliency prediction, the current literature still does not consider the impact of auditory information in ODVs. In this work, we propose an audio-visual saliency (AVS360) model that incorporates 360° spatial-temporal visual representation and spatial auditory information in ODVs. The proposed AVS360 model is composed of two 3D residual networks (ResNets) to encode visual and audio cues. The first one is embedded with a spherical representation technique to extract 360° visual features, and the second one extracts the features of audio using the log mel-spectrogram. We emphasize sound source locations by integrating audio energy map (AEM) generated from spatial audio description (i.e., ambisonics) and equator viewing behavior with equator center bias (ECB). The audio and visual features are combined and fused with AEM and ECB via attention mechanism. Our experimental results show that the AVS360 model has significant superiority over five state-of-the-art saliency models. To the best of our knowledge, it is the first w ork that develops the audio-visual saliency model in ODVs. The code will be publicly available to foster future research on audio-visual saliency in ODVs. Fang-Yi Chao, Cagri Ozcinar, Lu Zhang 0037, Wassim Hamidouche, Olivier Déforges, Aljoscha Smolic |
VCIP | 1 |