VLDB 2026 Research / reviewers in the wild / expert
Houwei Cao
dblp:21/9232
· DBLP profile ↗
28ranked-venue papers
7as first author
11since 2021 · last 2025
0000-0002-2310-7682ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 13 · 6 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 12 · 3 first-author · 3 since 2021Computer networks · 6 · 5 since 2021Human-computer interaction and ubiquitous computing · 3 · 1 first-author · 1 since 2021Security and privacy · 2 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Decentralized Federated Learning with Model Caching on Mobile AgentsabstractFederated Learning (FL) trains a shared model using data and computation power on distributed agents coordinated by a central server. Decentralized FL (DFL) utilizes local model exchange and aggregation between agents to reduce the communication and computation overheads on the central server. However, when agents are mobile, the communication opportunity between agents can be sporadic, largely hindering the convergence and accuracy of DFL. In this paper, we propose Cached Decentralized Federated Learning (Cached-DFL) to investigate delay-tolerant model spreading and aggregation enabled by model caching on mobile agents. Each agent stores not only its own model, but also models of agents encountered in the recent past. When two agents meet, they exchange their own models as well as the cached models. Local model aggregation utilizes all models stored in the cache. We theoretically analyze the convergence of Cached-DFL, explicitly taking into account the model staleness introduced by caching. We design and compare different model caching algorithms for different DFL and mobility scenarios. We conduct detailed case studies in a vehicular network to systematically investigate the interplay between agent mobility, cache staleness, and model convergence. In our experiments, Cached-DFL converges quickly, and significantly outperforms DFL without caching. Xiaoyu Wang 0015, Guojun Xiong, Houwei Cao, Yong Liu 0013 |
AAAI | 3 |
| 2025 | Coffee: Cost-effective edge caching for live 360 degree video streaming
Chen Li 0043, Tingwei Ye, Tongyu Zong, Liyang Sun, Houwei Cao, Yong Liu 0013 |
Comput. Networks | 5 |
| 2023 | Predictive edge caching through deep mining of sequential patterns in user content retrievals
Chen Li 0043, Xiaoyu Wang 0015, Tongyu Zong, Houwei Cao, Yong Liu 0013 |
Comput. Networks | 4 |
| 2023 | Audio-Visual Emotion Recognition With Preference Learning Based on Intended and Multi-Modal Perceived LabelsabstractThis article introduces a novel preference learning framework that simultaneously considers both the intended and the perceived labels while addressing the mismatches between them. Based on analyzing the discrepancies and agreements between the intended and the perceived labels in different modalities of audio-only, visual-only, and audio-visual, as well as the consistency among the perceptual ratings of all raters, we propose three sets of pair-wise ranking rules to generate multi-scale relevant scores for preference learning, scaling from sketchy manner to detailed manner. Three ranking models with support vector machine (SVM), deep neural networks (DNN), and gradient boosting decision trees (GBDT) are developed. Our results demonstrate that all three preference learning models significantly outperform the conventional classifiers baselines, and the LambdaMART model with gradient boosting decision trees achieves the best performance. The improvement from the preference learning models confirm the benefits of complementary information provided by different types of labels. We also observe additional improvement from the detailed ‘complex ranking rules’, particular with the best LambdaMART model, which suggests that we should treat intended and perceived labels in single-model & multi-modal differently. We further discuss the complementary of different ranking models, and obtain the best overall accuracy of 85.06% on CREMA-D dataset when combining the two best ranking models–LambdaMART and RankNet–together, which is significantly better than the 76.19% accuracy attained by the baseline models. Finally, we perform the cross-corpus emotion recognition experiments by training emotion rankers on CREMA-D dataset and tested the ranking-based emotion classifier on the SAVEE dataset that do not have perceived labels annotated. Yuanyuan Lei 0001, Houwei Cao |
IEEE Trans. Affect. Comput. | 2 |
| 2023 | Cocktail Edge Caching: Ride Dynamic Trends of Content Popularity With Ensemble LearningabstractEdge caching will play a critical role in facilitating the emerging content-rich applications. However, it faces many new challenges, in particular, the highly dynamic content popularity and the heterogeneous caching configurations. In this paper, we propose Cocktail Edge Caching, that tackles the dynamic popularity and heterogeneity through ensemble learning. Instead of trying to find a single dominating caching policy for all the caching scenarios, we employ an ensemble of constituent caching policies and adaptively select the best-performing policy to control the cache. Towards this goal, we first show through formal analysis and experiments that different variations of the LFU and LRU policies have complementary performance in different caching scenarios. We further develop a novel caching algorithm that enhances LFU/LRU with deep recurrent neural network (LSTM) based time-series analysis. Finally, we develop a deep reinforcement learning agent that adaptively combines base caching policies according to their virtual hit ratios on parallel virtual caches. Through extensive experiments driven by real content requests from two large video streaming platforms, we demonstrate that CEC not only consistently outperforms all single policies, but also improves the robustness of them. CEC can be well generalized to different caching scenarios with low computation overheads for deployment. Tongyu Zong, Chen Li 0043, Yuanyuan Lei 0001, Houwei Cao, Yong Liu 0013 |
IEEE/ACM Trans. Netw. | 5 |
| 2022 | Multimodal Emotion Recognition with Surgical and Fabric MasksabstractIn this study, we investigate how different types of masks affect automatic emotion classification in different channels of audio, visual, and multimodal. We train emotion classification models for each modality with the original data without mask and the re-generated data with mask respectively, and investigate how muffled speech and occluded facial expressions change the prediction of emotions. Moreover, we conduct the contribution analysis to study how muffled speech and occluded face interplay with each other and further investigate the individual contribution of audio, visual, and audio-visual modalities to the prediction of emotion with and without mask. Finally, we investigate the cross-corpus emotion recognition across clear speech and re-generated speech with different types of masks, and discuss the robustness of speech emotion recognition. Ziqing Yang 0003, Katherine Nayan, Zehao Fan, Houwei Cao |
ICASSP | 4 |
| 2022 | Realtime mobile bandwidth and handoff predictions in 4G/5G networks
Lifan Mei, Jinrui Gou, Yujin Cai, Houwei Cao, Yong Liu 0013 |
Comput. Networks | 4 |
| 2021 | Analysis of Eye Fixations During Emotion Recognition in Talking FacesabstractResearch on emotion recognition from cues expressed in facial expression has a long-standing tradition. In this study, we investigate human’s visual attention and fixation patterns when identifying six basic emotions on expressive talking faces. Stimuli for the current experiments consisted of 92 video clips of facial expression during talking. The whole experiments were divided into two sessions. The video stimuli in the first session were presented in random order across different face identities, while in the second session the video from the same face identity were be played sequentially. The participants’ eye movements were recorded by the Tobii X3-120 screen-based eye-tracking system. We defined a set of area-of-interest (AOI) regions, including 4 AOIs of general face areas and 12 AOIs related to specific Action Units (AUs) involved in the coding of the six basic emotions. The gaze pattern analysis was done by looking at the fixation time on this predetermined set of AOIs. Based on the ANOVA analysis, we did not find significant differences in mean fixation time on any AOI for discriminating the six basic emotions, but a subset of significant AOIs was found when we sectioned the six basic emotions into positive, negative, and neutral. Next, we propose to develop a novel emotion perception classifier which can automatically classify an observer’s emotion perception based on her gaze patterns and fixation sequence when identifying the basic emotions on expressive talking faces. The fixation time on the 16 predetermined AOIs were used as features to train support vector machine (SVM) models. The proposed models achieved the overall classification accuracy of 84.1% on recognizing 3-way emotions of negative, positive and neutral, suggesting that the proposed eye gaze patterns - the fixation time on 16 predetermined AOIs, are very promising for automatic classification of the perceived emotions. Finally, we divided the data into different gender and race groups, and discussed the diversity in gaze patterns across genders and different race groups. Houwei Cao, Forest Elliott |
ACII | 1 |
| 2021 | Cocktail Edge Caching: Ride Dynamic Trends of Content Popularity with Ensemble LearningabstractEdge caching will play a critical role in facilitating the emerging content-rich applications. However, it faces many new challenges, in particular, the highly dynamic content popularity and the heterogeneous caching configurations. In this paper, we propose Cocktail Edge Caching, that tackles the dynamic popularity and heterogeneity through ensemble learning. Instead of trying to find a single dominating caching policy for all the caching scenarios, we employ an ensemble of constituent caching policies and adaptively select the best-performing policy to control the cache. Towards this goal, we first show through formal analysis and experiments that different variations of the LFU and LRU polices have complementary performance in different caching scenarios. We further develop a novel caching algorithm that enhances LFU/LRU with deep recurrent neural network (LSTM) based time-series analysis. Finally, we develop a deep reinforcement learning agent that adaptively combines base caching policies according to their virtual hit ratios on parallel virtual caches. Through extensive experiments driven by real content requests from two large video streaming platforms, we demonstrate that CEC not only consistently outperforms all single policies, but also improves the robustness of them. CEC can be well generalized to different caching scenarios with low computation overheads for deployment. Tongyu Zong, Chen Li 0043, Yuanyuan Lei 0001, Houwei Cao, Yong Liu 0013 |
INFOCOM | 5 |
| 2021 | User Behavior Fingerprinting With Multi-Item-Sets and Its Application in IPTV Viewer IdentificationabstractUser activities in cyberspace leave unique traces for user identification (UI). Individual users can be identified by their frequent activity items through statistical feature matching. However, such approaches face the data sparsity problem. In this paper, we propose to address this problem by multi-item-set fingerprinting that identifies users not only based on their frequent individual activity items, but also their frequent consecutive item sequences with different lengths. We also propose a new similarity metric between fingerprint vectors that combines the advantages of Jaccard distance and relative entropy distance. Furthermore, we develop a fusion decision scheme by consolidating matching candidates generated by different similarity metrics. It improves the precision at the price of extra rejection. Our proposed approaches can be used in both one-by-one matching and bipartite graph group matching. Through extensive experiments on three real user datasets, in particular a large-scale Internet Protocol Television (IPTV) viewer dataset, we demonstrate that the proposed approaches outperform the state-of-the-art methods. The average matching precision reaches 93.8% for a dataset of 1,000 users and 100% for a dataset of 100 users. This work is of significance for information forensics and raises a new challenge for human privacy protection in cyberspace. Houwei Cao, Qihu Yuan, Yong Liu 0013 |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2021 | IPTV Channel Zapping Recommendation With Attention MechanismabstractInternet Protocol TV (IPTV) normally has the advantage of providing far more TV channels than the traditional TV services, while as the other side of the coin it has the problem of information overload. Users of IPTV usually have difficulties finding channels matching their interests. In this paper, using a large IPTV dataset, we analyze channel zapping behaviors of IPTV users and discover various patterns that can be used to generate more accurate channel zapping recommendations. Based on user behavior analysis, we develop several base and fusion recommender systems that generate in real-time a short list of channels for users to consider whenever they want to switch channels. A deep neural network model that consists of a “Recommender System Attention (RS Attention)” module and a “Channel Attention” module capturing the static and dynamic user switching behaviors is also developed to further improve the recommendation accuracy. Evaluation on the IPTV dataset demonstrates that our fusion recommender can achieve 41% hit ratio with only three candidate channels, and our attention neural network model further pushes it up to 45%. Our recommender systems only take as input user channel zapping sequences, and can be easily adopted by IPTV systems with low data and computation overheads. Lina Qiu, Chenguang Yu, Houwei Cao, Yong Liu 0013 |
IEEE Trans. Multim. | 4 |
| 2020 | Exploration of Acoustic and Lexical Cues for the INTERSPEECH 2020 Computational Paralinguistic Challenge
Ziqing Yang 0003, Zifan An, Zehao Fan, Chengye Jing, Houwei Cao |
INTERSPEECH | 5 |
| 2020 | Realtime mobile bandwidth prediction using LSTM neural network and Bayesian fusion
Lifan Mei, Runchen Hu, Houwei Cao, Yong Liu 0013, Zifan Han |
Comput. Networks | 3 |
| 2019 | Development of Emotion Rankers Based on Intended and Perceived Emotion Labels
Zhenghao Jin, Houwei Cao |
INTERSPEECH | 2 |
| 2019 | Realtime Mobile Bandwidth Prediction Using LSTM Neural Network
Lifan Mei, Runchen Hu, Houwei Cao, Yong Liu 0013, Zifa Han |
PAM | 3 |
| 2017 | Follow Me: Personalized IPTV Channel Switching GuideabstractCompared with the traditional television services, Internet Protocol TV (IPTV) can provide far more TV channels to end users. However, it may also make users feel confused even painful to find channels of their interests from a large number of them. In this paper, using a large IPTV trace, we analyze user channel-switching behaviors to understand when, why and how they switch channels. Based on user behavior analysis, we develop several base and fusion recommender systems that generate in real-time a short list of channels for users to consider whenever they want to switch channels. Evaluation on the IPTV trace demonstrates that our recommender systems can achieve up to 45 percent hit ratio with only three candidate channels. Our recommender systems only need access to user channel watching sequences, and can be easily adopted by IPTV systems with low data and computation overheads. Chenguang Yu, Hao Ding 0006, Houwei Cao, Yong Liu 0013 |
MMSys | 3 |
| 2016 | Improving Cold Music Recommendation through Hierarchical Audio AlignmentabstractCollaborative filtering (CF) is the state-of-the-art approach to item recommendation. However, it can neither recommend new items with no user feedbacks, nor could it recommend "long-tail" items easily. Content-based filtering can solve both problems through content analysis. However, content-based filtering alone has a much worse performance than CF. In this paper, we fuse user feedbacks and content analysis into the probabilistic matrix factorization framework. In particular, we propose a recursive dynamic programming approach to computing item similarity matrix from item content. Item latent factors are predicted from the item similarity matrix when no usage data is available. We investigate how performances of recommendation algorithms vary on items with different popularities. Results show that our approach has better performance than the same hybrid model with naive item similarity measures and Matrix Factorization. Hao Ding 0006, Houwei Cao, Yong Liu 0013 |
ISM | 3 |
| 2015 | Acoustic and lexical representations for affect prediction in spontaneous conversations
Houwei Cao, Arman Savran, Ragini Verma, Ani Nenkova |
Comput. Speech Lang. | 1 |
| 2015 | Speaker-sensitive emotion recognition via ranking: Studies on acted and spontaneous speech
Houwei Cao, Ragini Verma, Ani Nenkova |
Comput. Speech Lang. | 1 |
| 2015 | Temporal Bayesian Fusion for Affect Sensing: Combining Video, Audio, and Lexical ModalitiesabstractThe affective state of people changes in the course of conversations and these changes are expressed externally in a variety of channels, including facial expressions, voice, and spoken words. Recent advances in automatic sensing of affect, through cues in individual modalities, have been remarkable; yet emotion recognition is far from a solved problem. Recently, researchers have turned their attention to the problem of multimodal affect sensing in the hope that combining different information sources would provide great improvements. However, reported results fall short of the expectations, indicating only modest benefits and occasionally even degradation in performance. We develop temporal Bayesian fusion for continuous real-value estimation of valence, arousal, power, and expectancy dimensions of affect by combining video, audio, and lexical modalities. Our approach provides substantial gains in recognition performance compared to previous work. This is achieved by the use of a powerful temporal prediction model as prior in Bayesian fusion as well as by incorporating uncertainties about the unimodal predictions. The temporal prediction model makes use of time correlations on the affect sequences and employs estimated temporal biases to control the affect estimations at the beginning of conversations. In contrast to other recent methods for combination of modalities our model is simpler, since it does not model relationships between modalities and involves only a few interpretable parameters to be estimated from the training data. Arman Savran, Houwei Cao, Ani Nenkova, Ragini Verma |
IEEE Trans. Cybern. | 2 |
| 2014 | CREMA-D: Crowd-Sourced Emotional Multimodal Actors DatasetabstractPeople convey their emotional state in their face and voice. We present an audio-visual data set uniquely suited for the study of multi-modal emotion expression and perception. The data set consists of facial and vocal emotional expressions in sentences spoken in a range of basic emotional states (happy, sad, anger, fear, disgust, and neutral). 7,442 clips of 91 actors with diverse ethnic backgrounds were rated by multiple raters in three modalities: audio, visual, and audio-visual. Categorical emotion labels and real-value intensity values for the perceived emotion were collected using crowd-sourcing from 2,443 raters. The human recognition of intended emotion for the audio-only, visual-only, and audio-visual data are 40.9%, 58.2% and 63.6% respectively. Recognition rates are highest for neutral, followed by happy, anger, disgust, fear, and sad. Average intensity levels of emotion are rated highest for visual-only perception. The accurate recognition of disgust and fear requires simultaneous audio-visual cues, while anger and happiness can be well recognized based on evidence from a single modality. The large dataset we introduce can be used to probe other questions concerning the audio-visual perception of emotion. Houwei Cao, David G. Cooper, Michael K. Keutmann, Ruben C. Gur, Ani Nenkova, Ragini Verma |
IEEE Trans. Affect. Comput. | 1 |
| 2013 | Action Unit Models of Facial Expression of Emotion in the Presence of SpeechabstractAutomatic recognition of emotion using facial expressions in the presence of speech poses a unique challenge because talking reveals clues for the affective state of the speaker but distorts the canonical expression of emotion on the face. We introduce a corpus of acted emotion expression where speech is either present (talking) or absent (silent). The corpus is uniquely suited for analysis of the interplay between the two conditions. We use a multimodal decision level fusion classifier to combine models of emotion from talking and silent faces as well as from audio to recognize five basic emotions: anger, disgust, fear, happy and sad. Our results strongly indicate that emotion prediction in the presence of speech from action unit facial features is less accurate when the person is talking. Modeling talking and silent expressions separately and fusing the two models greatly improves accuracy of prediction in the talking setting. The advantages are most pronounced when silent and talking face models are fused with predictions from audio features. In this multi-modal prediction both the combination of modalities and the separate models of talking and silent facial expression of emotion contribute to the improvement. Miraj Shah, David G. Cooper, Houwei Cao, Ruben C. Gur, Ani Nenkova, Ragini Verma |
ACII | 3 |
| 2012 | Combining video, audio and lexical indicators of affect in spontaneous conversation via particle filteringabstractWe present experiments on fusing facial video, audio and lexical indicators for affect estimation during dyadic conversations. We use temporal statistics of texture descriptors extracted from facial video, a combination of various acoustic features, and lexical features to create regression based affect estimators for each modality. The single modality regressors are then combined using particle filtering, by treating these independent regression outputs as measurements of the affect states in a Bayesian filtering framework, where previous observations provide prediction about the current state by means of learned affect dynamics. Tested on the Audio-visual Emotion Recognition Challenge dataset, our single modality estimators achieve substantially higher scores than the official baseline method for every dimension of affect. Our filtering-based multi-modality fusion achieves correlation performance of 0.344 (baseline: 0.136) and 0.280 (baseline: 0.096) for the fully continuous and word level sub challenges, respectively. Arman Savran, Houwei Cao, Miraj Shah, Ani Nenkova, Ragini Verma |
ICMI | 2 |
| 2012 | Combining Ranking and Classification to Improve Emotion Recognition in Spontaneous SpeechabstractWe introduce a novel emotion recognition approach which integrates ranking models. The approach is speaker independent, yet it is designed to exploit information from utterances from the same speaker in the test set before making predictions. It achieves much higher precision in identifying emotional utterances than a conventional SVM classifier. Furthermore we test several possibilities for combining conventional classification and predictions based on ranking. All combinations improve overall prediction accuracy. All experiments are performed on the FAU AIBO database which contains realistic spontaneous emotional speech. Our best combination system achieves 6.6 % absolute improvement over the Interspeech 2009 emotion challenge baseline system on the 5-class classification tasks. Index Terms: emotion classification, ranking models, spontaneous speech Houwei Cao, Ragini Verma, Ani Nenkova |
INTERSPEECH | 1 |
| 2010 | Cross-lingual speaker adaptation via Gaussian component mapping
Houwei Cao, Tan Lee, Pak-Chung Ching |
INTERSPEECH | 1 |
| 2009 | Effects of language mixing for automatic recognition of Cantonese-English code-mixing utterances
Houwei Cao, Pak-Chung Ching, Tan Lee |
INTERSPEECH | 1 |
| 2008 | Language modeling for speech recognition of spoken Cantonese
Yu Ting Yeung, Houwei Cao, Nengheng Zheng, Tan Lee, Pak-Chung Ching |
INTERSPEECH | 2 |
| 2006 | Automatic speech recognition of Cantonese-English code-mixing utterances
Joyce Y. C. Chan, Pak-Chung Ching, Tan Lee, Houwei Cao |
INTERSPEECH | 4 |