EDBT 2026 Demo / reviewers in the wild / expert
Huaizheng Zhang
dblp:218/5222
· DBLP profile ↗
15ranked-venue papers
6as first author
6since 2021 · last 2025
0000-0002-0153-6400ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 10 · 6 first-author · 4 since 2021Artificial intelligence and machine learning · 4 · 2 since 2021Computer networks · 2 · 1 since 2021Databases, data management, data science and information retrieval · 2
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
4 papers |
Efficient and distributed learning · 77% Learning paradigms · 10% Information extraction and text analysis · 7% | |
| Computer architecture, parallel and distributed computing, and storage systems
2 papers |
Cloud and datacenter computing · 86% GPUs and heterogeneous computing · 14% | |
| Computer graphics and multimedia
4 papers |
Multimedia systems and quality of experience · 58% Multimedia analysis and retrieval · 42% | |
| Computer networks
2 papers |
Edge and fog computing · 100% | |
| Interdisciplinary, comprehensive, and emerging computing
2 papers |
Energy systems and smart grids · 75% Computational social science and digital humanities · 25% |
Topics — the 20 heaviest of 26, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Efficient and distributed learning
active learning |
1.4 | 2 | 2025 | Multi-Edge Reinforced Collaborative Data Acquisition for Continuous Video Analytics by Prioritizing Quality over Quantity · AAAI 2025 Towards Data-Efficient Continuous Learning for Edge Video Analytics via Smart Caching · SenSys 2022 |
Edge and fog computing › video analytics
edge video analytics |
1.4 | 2 | 2025 | Multi-Edge Reinforced Collaborative Data Acquisition for Continuous Video Analytics by Prioritizing Quality over Quantity · AAAI 2025 Towards Data-Efficient Continuous Learning for Edge Video Analytics via Smart Caching · SenSys 2022 |
Cloud and datacenter computing
cluster resource management and scheduling |
0.9 | 2 | 2020 | MLModelCI: An Automatic Cloud Platform for Efficient MLaaS · ACM Multimedia 2020 Hysia: Serving DNN-Based Video-to-Retail Applications in Cloud · ACM Multimedia 2020 |
Cloud and datacenter computing
machine learning as a service |
0.9 | 2 | 2020 | MLModelCI: An Automatic Cloud Platform for Efficient MLaaS · ACM Multimedia 2020 Hysia: Serving DNN-Based Video-to-Retail Applications in Cloud · ACM Multimedia 2020 |
Multimedia systems and quality of experience
adaptive bitrate streaming |
0.8 | 2 | 2020 | DeepQoE: A Multimodal Learning Framework for Video Quality of Experience (QoE) Prediction · IEEE Trans. Multim. 2020 Optimizing Quality of Experience for Adaptive Bitrate Streaming via Viewer Interest Inference · IEEE Trans. Multim. 2018 |
Machine learning › Efficient and distributed learning › federated learning
data heterogeneity |
0.7 | 1 | 2023 | PRIOR: Personalized Prior for Reactivating the Information Overlooked in Federated Learning · NeurIPS 2023 |
Machine learning › Efficient and distributed learning
federated learning |
0.7 | 1 | 2023 | PRIOR: Personalized Prior for Reactivating the Information Overlooked in Federated Learning · NeurIPS 2023 |
Machine learning › Efficient and distributed learning › federated learning
personalized federated learning |
0.7 | 1 | 2023 | PRIOR: Personalized Prior for Reactivating the Information Overlooked in Federated Learning · NeurIPS 2023 |
Edge and fog computing
continuous learning |
0.6 | 1 | 2022 | Towards Data-Efficient Continuous Learning for Edge Video Analytics via Smart Caching · SenSys 2022 |
Energy systems and smart grids › energy forecasting
photovoltaic power forecasting |
0.5 | 1 | 2021 | Missing Data Imputation for Solar Yield Prediction using Temporal Multi-Modal Variational Auto-Encoder · ACM Multimedia 2021 |
Multimedia analysis and retrieval
advertisement analysis |
0.4 | 1 | 2020 | Look, Read and Feel: Benchmarking Ads Understanding with Multimodal Multitask Learning · ACM Multimedia 2020 |
Cloud and datacenter computing › resource allocation
elastic resource allocation |
0.4 | 1 | 2020 | MLModelCI: An Automatic Cloud Platform for Efficient MLaaS · ACM Multimedia 2020 |
GPUs and heterogeneous computing
GPU scheduling |
0.4 | 1 | 2020 | Hysia: Serving DNN-Based Video-to-Retail Applications in Cloud · ACM Multimedia 2020 |
Cloud and datacenter computing
inference serving |
0.4 | 1 | 2020 | Hysia: Serving DNN-Based Video-to-Retail Applications in Cloud · ACM Multimedia 2020 |
Machine learning › Reinforcement learning
policy learning |
0.3 | 1 | 2025 | Multi-Edge Reinforced Collaborative Data Acquisition for Continuous Video Analytics by Prioritizing Quality over Quantity · AAAI 2025 |
Machine learning › Learning paradigms › continual learning
catastrophic forgetting |
0.2 | 1 | 2022 | Towards Data-Efficient Continuous Learning for Edge Video Analytics via Smart Caching · SenSys 2022 |
Machine learning › Learning paradigms
continual learning |
0.2 | 1 | 2022 | Towards Data-Efficient Continuous Learning for Edge Video Analytics via Smart Caching · SenSys 2022 |
Machine learning › Efficient and distributed learning › data-efficient learning
label-efficient learning |
0.1 | 1 | 2018 | ResumeNet: A Learning-Based Framework for Automatic Resume Quality Assessment · ICDM 2018 |
Machine learning › Learning paradigms
semi-supervised learning |
0.1 | 1 | 2018 | ResumeNet: A Learning-Based Framework for Automatic Resume Quality Assessment · ICDM 2018 |
Multimedia analysis and retrieval › video classification
video scene classification |
0.1 | 1 | 2018 | Optimizing Quality of Experience for Adaptive Bitrate Streaming via Viewer Interest Inference · IEEE Trans. Multim. 2018 |
Methods — techniques the papers use, named apart from their topics
active learning · 2.9reinforcement learning · 1.7exemplar pool · 1.1deep neural network · 0.9GPU acceleration · 0.9semi-supervised learning · 0.7pair/triplet-based loss · 0.7neural network · 0.7mirror descent · 0.7convergence analysis · 0.7bregman divergence · 0.7variational autoencoder · 0.5time-series imputation · 0.5multimodal fusion · 0.5word embedding · 0.4unsupervised visual metaphor decoding · 0.4multimodal representation learning · 0.4multimodal attention · 0.4
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Multi-Edge Reinforced Collaborative Data Acquisition for Continuous Video Analytics by Prioritizing Quality over QuantityabstractEdge computing-based video analytics faces data drift issues due to the occurrence of unseen objects or scenes in ever-changing environments. To maintain accuracy, continuous learning (CL) retrains stale models periodically with newly obtained data. However, it leads to unaffordable costs, as we must keep labeling drift data and retraining models. Regarding this concern, we first investigate video patterns across multiple cameras within an area and reveal significant data redundancies. We find that many of the same objects can be captured by multiple edge cameras or appear many times on the same edges. Our quantitative findings suggest that selecting a subset of high-quality data for CL is preferable over using a larger quantity. Yet, existing efforts for data acquisition have only focused on a single static dataset. These methods are not suitable for multi-edge video analytics scenarios, where videos are captured from multiple sources with non-iid data distribution. Hence, we propose a multi-edge collaborative active video acquisition (AVA) framework to collaboratively learn a reinforced video acquisition strategy to identify informative video frames from multiple edge nodes that best enhance model accuracy, avoiding redundancy across edges. Extensive experiments on three video datasets demonstrate that, our method achieves comparable performance to full-set video training while utilizing only 20% of the data in classification tasks. In object detection tasks, our methods can maintain productive accuracy with a reduction of nearly 70% in training video frames. Guanyu Gao, Haiyan Yin, Huaizheng Zhang |
AAAI | 4 |
| 2025 | Spatial-Temporal Federated Learning for Lifelong Person Re-Identification on Distributed EdgesabstractData drift is a thorny challenge when deploying person re-identification (ReID) models into real-world devices, where the data distribution is significantly different from that of the training environment and keeps changing. To tackle this issue, we propose a federated spatial-temporal incremental learning approach, named FedSTIL, which leverages both lifelong learning and federated learning to continuously optimize models deployed on many distributed edge clients. Unlike previous efforts, FedSTIL aims to mine spatial-temporal correlations among the knowledge learnt from different edge clients. Specifically, the edge clients first periodically extract general representations of drifted data to optimize their local models. Then, the learnt knowledge from edge clients will be aggregated by centralized parameter server, where the knowledge will be selectively and attentively distilled from spatial- and temporal-dimension with carefully designed mechanisms. Finally, the distilled informative spatial-temporal knowledge will be sent back to correlated edge clients to further improve the recognition accuracy of each edge client with a lifelong learning method. Extensive experiments on a mixture of five real-world datasets demonstrate that our method outperforms others by nearly 4% in Rank-1 accuracy, while reducing communication cost by 62%. All implementation codes are publicly available on https://github.com/MSNLAB/Federated-Lifelong-Person-ReID. Guanyu Gao, Huaizheng Zhang |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2023 | Robust Video Object Segmentation with Restricted AttentionabstractThis paper focuses on the two problems of the similar objects distraction and the lack of robustness for unseen object categories in semi-supervised video object segmentation task. Existing methods have achieved great results on the benchmark dataset, but these two problems still have not been completely solved. We propose the Robust Video Object Segmentation With Restricted Attention (RVOSR), which can suppress the effects caused by similar objects and filter out noise confusion from other irrelevant regions. Meanwhile augmenting the semantic information of features, which makes the features more suitable for video object segmentation task. Extensive experiments demonstrate the effectiveness of our approach and achieve the state-of-the-art performance on the widely-used VOS benchmarks including DAVIS-2016 (92.1% $\mathcal{J}{{\& }}\mathcal{F}$), DAVIS-2017 (86.8% $\mathcal{J}{{\& }}\mathcal{F}$) and YouTubeVOS-2019 (84.8%). Huaizheng Zhang, Pinxue Guo, Zhongwen Le |
ICASSP | 1 |
| 2023 | PRIOR: Personalized Prior for Reactivating the Information Overlooked in Federated LearningabstractClassical federated learning (FL) enables training machine learning models without sharing data for privacy preservation, but heterogeneous data characteristic degrades the performance of the localized model. Personalized FL (PFL) addresses this by synthesizing personalized models from a global model via training on local data. Such a global model may overlook the specific information that the clients have been sampled. In this paper, we propose a novel scheme to inject personalized prior knowledge into the global model in each client, which attempts to mitigate the introduced incomplete information problem in PFL. At the heart of our proposed approach is a framework, the $\textit{PFL with Bregman Divergence}$ (pFedBreD), decoupling the personalized prior from the local objective function regularized by Bregman divergence for greater adaptability in personalized scenarios. We also relax the mirror descent (RMD) to extract the prior explicitly to provide optional strategies. Additionally, our pFedBreD is backed up by a convergence analysis. Sufficient experiments demonstrate that our method reaches the $\textit{state-of-the-art}$ performances on 5 datasets and outperforms other methods by up to 3.5% across 8 benchmarks. Extensive analyses verify the robustness and necessity of proposed designs. The code will be made public. Mingjia Shi, Yuhao Zhou 0004, Kai Wang 0036, Huaizheng Zhang, Shudong Huang, Jiancheng Lv 0001 |
NeurIPS | 4 |
| 2022 | Towards Data-Efficient Continuous Learning for Edge Video Analytics via Smart CachingabstractContinuous learning (CL) has recently been adopted into edge video analytics, gaining huge success in maintaining high accuracy without constantly retraining DNN models by human intervention. Though existing solutions offer optimized processing pipelines, the cost brought by CL should not be neglected. This vision paper starts an investigation by exploring two kinds of cost, human labeling and edge storage. The former comes from the need for CL's automatically tuning, and the latter is due to an exemplar pool (including both drift and historical data) maintained to prevent catastrophic forgetting caused by naive retraining. To alleviate the costs, we propose a new CL-based edge video analytics system by incorporating an active learner mechanism. Specifically, we revisit the current CL video system design and develop an active CL pipeline atop them. The pipeline first accepts the drift data stored in drift pool and utilizes an active learner to sample a small partition of them for labeling. Then it mixes up both small labeled drifted data and some historical data to send them to an exemplar pool for CL. Our preliminary benchmark studies exhibit that the new system can achieve competitive accuracy by spending only 30% labeling and storage cost compared to other baselines, showing a promising research direction for future study. Guanyu Gao, Huaizheng Zhang |
SenSys | 3 |
| 2021 | Missing Data Imputation for Solar Yield Prediction using Temporal Multi-Modal Variational Auto-EncoderabstractThe accurate and robust prediction of short-term solar power generation is significant for the management of modern smart grids, where solar power has become a major energy source due to its green and economical nature. However, the solar yield prediction can be difficult to conduct in the real world where hardware and network issues can make the sensors unreachable. Such data missing problem is so prevalent that it degrades the performance of deployed prediction models and even fails the model execution. In this paper, we propose a novel temporal multi-modal variational auto-encoder (TMMVAE) model, to enhance the robustness of short-term solar power yield prediction with missing data. It can impute the missing values in time-series sensor data, and reconstruct them by consolidating multi-modality data, which then facilitates more accurate solar power yield prediction. TMMVAE can be deployed efficiently with an end-to-end framework. The framework is verified at our real-world testbed on campus. The results of extensive experiments show that our proposed framework can significantly improve the imputation accuracy when the inference data is severely corrupted, and can hence dramatically improve the robustness of short-term solar energy yield forecasting. Meng Shen 0002, Huaizheng Zhang, Yixin Cao 0002, Fan Yang 0172, Yonggang Wen 0001 |
ACM Multimedia | 2 |
| 2020 | Look, Read and Feel: Benchmarking Ads Understanding with Multimodal Multitask LearningabstractGiven the massive market of advertising and the sharply increasing online multimedia content (such as videos), it is now fashionable to promote advertisements (ads) together with the multimedia content. However, manually finding relevant ads to match the provided content is labor-intensive, and hence some automatic advertising techniques are developed. Since ads are usually hard to understand only according to its visual appearance due to the contained visual metaphor, some other modalities, such as the contained texts, should be exploited for understanding. To further improve user experience, it is necessary to understand both the ads' topic and sentiment. This motivates us to develop a novel deep multimodal multitask framework that integrates multiple modalities to achieve effective topic and sentiment prediction simultaneously for ads understanding. In particular, in our framework termed Deep$M^2$Ad, we first extract multimodal information from ads and learn high-level and comparable representations. The visual metaphor of the ad is decoded in an unsupervised manner. The obtained representations are then fed into the proposed hierarchical multimodal attention modules to learn task-specific representations for final prediction. A multitask loss function is also designed to jointly train both the topic and sentiment prediction models in an end-to-end manner, where bottom-layer parameters are shared to alleviate over-fitting. We conduct extensive experiments on a large-scale advertisement dataset and achieve state-of-the-art performance for both prediction tasks. The obtained results could be utilized as a benchmark for ads understanding. Huaizheng Zhang, Yong Luo 0002, Qiming Ai, Yonggang Wen 0001, Han Hu 0003 |
ACM Multimedia | 1 |
| 2020 | Hysia: Serving DNN-Based Video-to-Retail Applications in CloudabstractCombining video streaming and online retailing (V2R) has been a growing trend recently. In this paper, we provide practitioners and researchers in multimedia with a cloud-based platform named Hysia for easy development and deployment of V2R applications. The system consists of: 1) a back-end infrastructure providing optimized V2R related services including data engine, model repository, model serving and content matching; and 2) an application layer which enables rapid V2R application prototyping. Hysia addresses industry and academic needs in large-scale multimedia by: 1) seamlessly integrating state-of-the-art libraries including NVIDIA video SDK, Facebook faiss, and gRPC; 2) efficiently utilizing GPU computation; and 3) allowing developers to bind new models easily to meet the rapidly changing deep learning (DL) techniques. On top of that, we implement an orchestrator for further optimizing DL model serving performance. Hysia has been released as an open source project on GitHub, and attracted considerable attention. We have published Hysia to DockerHub as an official image for seamless integration and deployment in current cloud environments. Huaizheng Zhang, Yuanming Li, Qiming Ai, Yong Luo 0002, Yonggang Wen 0001, Yichao Jin 0002, Ta Nguyen Binh Duong |
ACM Multimedia | 1 |
| 2020 | MLModelCI: An Automatic Cloud Platform for Efficient MLaaSabstractMLModelCI provides multimedia researchers and developers with a one-stop platform for efficient machine learning (ML) services. The system leverages DevOps techniques to optimize, test, and manage models. It also containerizes and deploys these optimized and validated models as cloud services (MLaaS). In its essence, MLModelCI serves as a housekeeper to help users publish models. The models are first automatically converted to optimized formats for production purpose and then profiled under different settings (e.g., batch size and hardware). The profiling information can be used as guidelines for balancing the trade-off between performance and cost of MLaaS. Finally, the system dockerizes the models for ease of deployment to cloud environments. A key feature of MLModelCI is the implementation of a controller, which allows elastic evaluation which only utilizes idle workers while maintaining online service quality. Our system bridges the gap between current ML training and serving systems and thus free developers from manual and tedious work often associated with service deployment. We release the platform as an open-source project on GitHub under Apache 2.0 license, with the aim that it will facilitate and streamline more large-scale ML applications and research projects. Huaizheng Zhang, Yuanming Li, Yizheng Huang 0001, Yonggang Wen 0001, Jianxiong Yin, Kyle Guan |
ACM Multimedia | 1 |
| 2020 | DeepQoE: A Multimodal Learning Framework for Video Quality of Experience (QoE) PredictionabstractRecently, many models have been developed to predict video Quality of Experience (QoE), yet the applicability of these models still faces significant challenges. Firstly, many models rely on features that are unique to a specific dataset and thus lack the capability to generalize. Due to the intricate interactions among these features, a unified representation that is independent of datasets with different modalities is needed. Secondly, existing models often lack the configurability to perform both classification and regression tasks. Thirdly, the sample size of the available datasets to develop these models is often very small, and the impact of limited data on the performance of QoE models has not been adequately addressed. To address these issues, in this work we develop a novel and end-to-end framework termed as DeepQoE. The proposed framework first uses a combination of deep learning techniques, such as word embedding and 3D convolutional neural network (C3D), to extract generalized features. Next, these features are combined and fed into a neural network for representation learning. A learned representation will then serve as input for classification or regression tasks. We evaluate the performance of DeepQoE with three datasets. The results show that for small datasets (e.g., WHU-MVQoE2016 and Live-Netflix Video Database), the performance of state-of-the-art machine learning algorithms is greatly improved by using the QoE representation from DeepQoE (e.g., 35.71% to 44.82%); while for the large dataset (e.g., VideoSet), our DeepQoE framework achieves significant performance improvement in comparison to the best baseline method (90.94% vs. 82.84%). In addition to the much improved performance, DeepQoE has the flexibility to fit different datasets, to learn QoE representation, and to perform both classification and regression problems. We also develop a DeepQoE based adaptive bitrate streaming (ABR) system to verify that our framework can be easily applied to multimedia communication service. The software package of the DeepQoE framework has been released to facilitate the current research on QoE. Huaizheng Zhang, Linsen Dong, Guanyu Gao, Han Hu 0003, Yonggang Wen 0001, Kyle Guan |
IEEE Trans. Multim. | 1 |
| 2019 | ResumeGAN: An Optimized Deep Representation Learning Framework for Talent-Job Fit via Adversarial LearningabstractNowadays, it is popular to utilize online recruitment services for talent recruitment and job recommendation. Given the vast amounts of online talent profiles and job-posts, it is labor-intensive and exhausted for recruiters to manually select only a few potential candidates for further consideration, and also nontrivial for talents to find the most matched job positions. Recently, some deep learning-based approaches are developed to automatically matching the talent resumes and job requirements, and have achieved encouraging performance. In this paper, we propose a novel framework that targets the same task, but integrate different types of information in a more sophisticated way and introduce adversarial learning to learn more expressive representation. In addition, we build a dataset for model evaluation and the effectiveness of our framework is demonstrated by extensive experiments. Yong Luo 0002, Huaizheng Zhang, Yonggang Wen 0001, Xinwen Zhang |
CIKM | 2 |
| 2019 | Content-Aware Personalised Rate Adaptation for Adaptive Streaming via Deep Video AnalysisabstractAdaptive bitrate (ABR) streaming is the de facto solution for achieving smooth viewing experiences under unstable network conditions. However, most of the existing rate adaptation approaches for ABR are content-agnostic, without considering the semantic information of the video content. Nevertheless, semantic information largely determines the informativeness and interestingness of the video content, and consequently affects the QoE for video streaming. One common case is that the user may expect higher quality for the parts of video content that are more interesting or informative so as to reduce overall subjective quality loss. This creates two main challenges for such a problem: First, how to determine which parts of the video content are more interesting? Second, how to allocate bitrate budgets for different parts of the video content with different significances? To address these challenges, we propose a Content-of-Interest (CoI) based rate adaptation scheme for ABR. We first design a deep learning approach for recognizing the interestingness of the video content, and then design a Deep Q-Network (DQN) approach for rate adaptation by incorporating video interestingness information. The experimental results show that our method can recognize video interestingness precisely, and the bitrate allocation for ABR can be aligned with the interestingness of video content while not compromising the performances on objective QoE metrics. Guanyu Gao, Linsen Dong, Huaizheng Zhang, Yonggang Wen 0001, Wenjun Zeng 0001 |
ICC | 3 |
| 2018 | ResumeNet: A Learning-Based Framework for Automatic Resume Quality AssessmentabstractRecruitment of appropriate people for certain positions is critical for any companies or organizations. Manually screening to select appropriate candidates from large amounts of resumes can be exhausted and time-consuming. However, there is no public tool that can be directly used for automatic resume quality assessment (RQA). This motivates us to develop a method for automatic RQA. Since there is also no public dataset for model training and evaluation, we build a dataset for RQA by collecting around 10K resumes, which are provided by a private resume management company. By investigating the dataset, we identify some factors or features that could be useful to discriminate good resumes from bad ones, e.g., the consistency between different parts of a resume. Then a neural-network model is designed to predict the quality of each resume, where some text processing techniques are incorporated. To deal with the label deficiency issue in the dataset, we propose several variants of the model by either utilizing the pair/triplet-based loss, or introducing some semi-supervised learning technique to make use of the abundant unlabeled data. Both the presented baseline model and its variants are general and easy to implement. Various popular criteria including the receiver operating characteristic (ROC) curve, F-measure and ranking-based average precision (AP) are adopted for model evaluation. We compare the different variants with our baseline model. Since there is no public algorithm for RQA, we further compare our results with those obtained from a website that can score a resume. Experimental results in terms of different criteria demonstrate effectiveness of the proposed method. We foresee that our approach would transform the way of future human resources management. Yong Luo 0002, Huaizheng Zhang, Yonggang Wen 0001, Xinwen Zhang |
ICDM | 2 |
| 2018 | Deepqoe: A Unified Framework for Learning to Predict Video QoEabstractMotivated by the prowess of deep learning (DL) based techniques in prediction, generalization, and representation learning, we develop a novel framework called DeepQoE to predict video quality of experience (QoE). The end-to-end framework first uses a combination of DL techniques (e.g., word embeddings) to extract generalized features. Next, these features are combined and fed into a neural network for representation learning. Such representations serve as inputs for classification or regression tasks. Evaluating the performance of DeepQoE with two datasets, we show that for the small dataset, the accuracy of all shallow learning algorithms is improved by using the representation derived from DeepQoE. For the large dataset, our DeepQoE framework achieves significant performance improvement in comparison to the best baseline method (90.94% vs. 82.84%). Moreover, DeepQoE, also released as an open source tool, provides video QoE research much-needed flexibility in fitting different datasets, extracting generalized features, and learning representations. Huaizheng Zhang, Han Hu 0003, Guanyu Gao, Yonggang Wen 0001, Kyle Guan |
ICME | 1 |
| 2018 | Optimizing Quality of Experience for Adaptive Bitrate Streaming via Viewer Interest InferenceabstractRate adaptation is widely adopted in video streaming to improve the quality of experience (QoE). However, most of the existing rate adaptation approaches neglect the underlying video semantic information. In fact, influenced by video semantics and viewer preferences, the viewer may have different degrees of interest on different parts of a video. The interesting parts of a video can draw more visual attention from the viewer and have higher visual importance. As such, delivering the parts of a video that are interesting to the viewer in a higher quality can improve the perceptual video quality, compared with the semantics-agnostic approaches that treat each part of a video equally. Thus, it is natural to wonder: how to allocate bitrate budgets temporally over a video session under time-varying bandwidth while considering viewer interest? As an exploratory study, we propose an interest-aware rate adaptation approach for improving QoE by inferring viewer interest based on video semantics. We adopt the deep learning method to recognize the scenes of video frames and leverage the term frequency-inverse document frequency method to analyze the degrees of an individual viewer's interest on different types of scenes. The bandwidth, buffer occupancy, and viewer interest are jointly considered under the model predictive control framework for selecting appropriate bitrates for maximizing QoE. The objective and subjective evaluations measured in a real environment show that our method can achieve a higher QoE compared with the semantics-agnostic approaches. Guanyu Gao, Huaizheng Zhang, Han Hu 0003, Yonggang Wen 0001, Jianfei Cai 0001, Chong Luo 0001, Wenjun Zeng 0001 |
IEEE Trans. Multim. | 2 |