EDBT 2026 Demo / reviewers in the wild / expert
Liang Chen 0009
dblp:01/5394-9
· DBLP profile ↗
31ranked-venue papers
8as first author
10since 2021 · last 2024
0000-0002-3149-0239ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 15 · 7 first-author · 1 since 2021Databases, data management, data science and information retrieval · 8 · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 2 first-author · 1 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Rankability-enhanced Revenue Uplift Modeling Framework for Online MarketingabstractUplift modeling has been widely employed in online marketing by predicting the response difference between the treatment and control groups, so as to identify the sensitive individuals toward interventions like coupons or discounts. Compared with traditional conversion uplift modeling,revenue uplift modeling exhibits higher potential due to its direct connection with the corporate income. However, previous works can hardly handle the continuous long-tail response distribution in revenue uplift modeling. Moreover, they have neglected to optimize the uplift ranking among different individuals, which is actually the core of uplift modeling. To address such issues, in this paper, we first utilize the zero-inflated lognormal (ZILN) loss to regress the responses and customize the corresponding modeling network, which can be adapted to different existing uplift models. Then, we study the ranking-related uplift modeling error from the theoretical perspective and propose two tighter error bounds as the additional loss terms to the conventional response regression loss. Finally, we directly model the uplift ranking error for the entire population with a listwise uplift ranking loss. The experiment results on offline public and industrial datasets validate the effectiveness of our method for revenue uplift modeling. Furthermore, we conduct large-scale experiments on a prominent online fintech marketing platform, Tencent FiT, which further demonstrates the superiority of our method in real-world applications. Bowei He, Yunpeng Weng, Xing Tang 0007, Ziqiang Cui, Zexu Sun, Liang Chen 0009, Xiuqiang He 0001, Chen Ma 0001 |
KDD | 6 |
| 2024 | Treatment-Aware Hyperbolic Representation Learning for Causal Effect Estimation with Social NetworksabstractEstimating the individual treatment effect (ITE) from observational data is a crucial research topic that holds significant value across multiple domains. How to identify hidden confounders poses a key challenge in ITE estimation. Recent studies have incorporated the structural information of social networks to tackle this challenge, achieving notable advancements. However, these methods utilize graph neural networks to learn the representation of hidden confounders in Euclidean space, disregarding two critical issues: (1) the social networks often exhibit a scale-free structure, while Euclidean embeddings suffer from high distortion when used to embed such graphs, and (2) each ego-centric network within a social network manifests a treatment-related characteristic, implying significant patterns of hidden confounders. To address these issues, we propose a novel method called Treatment-Aware Hyperbolic Representation Learning (TAHyper). Firstly, TAHy-per employs the hyperbolic space to encode the social networks, thereby effectively reducing the distortion of confounder representation caused by Euclidean embeddings. Secondly, we design a treatment-aware relationship identification module that enhances the representation of hidden confounders by identifying whether an individual and her neighbors receive the same treatment. Extensive experiments on two benchmark datasets are conducted to demonstrate the superiority of our method. The code is available at https://github.com/ziqiangcui/TAHyper. Ziqiang Cui, Xing Tang 0007, Bowei He, Liang Chen 0009, Xiuqiang He 0001, Chen Ma 0001 |
SDM | 5 |
| 2023 | Self-Sampling Training and Evaluation for the Accuracy-Bias Tradeoff in Recommendation
Dugang Liu, Xing Tang 0007, Liang Chen 0009, Xiuqiang He 0001, Weike Pan, Zhong Ming 0001 |
DASFAA (4) | 4 |
| 2023 | Prior-Guided Accuracy-Bias Tradeoff Learning for CTR Prediction in Multimedia RecommendationabstractAlthough debiasing in multimedia recommendation has shown promising results, most existing work relies on the ability of the model itself to fully disentangle the biased and unbiased information and considers arbitrarily removing all the biases. However, in many business scenarios, it is usually possible to extract a subset of features associated with the biases by means of expert knowledge, i.e., the confounding proxy features. Therefore, in this paper, we propose a novel debiasing framework with confounding proxy priors for the accuracy-bias tradeoff learning in the multimedia recommendation, or CP2Rec for short, in which these confounding proxy features driven by the expert experience are integrated into the model as prior knowledge corresponding to the biases. Specifically, guided by these priors, we use a bias disentangling module with some orthogonal constraints to force the model to avoid encoding biased information in the feature embeddings. We then introduce an auxiliary unbiased loss to synergize with the original biased loss in an accuracy-bias tradeoff module, aiming at recovering the beneficial bias information from the above-purified feature embeddings to achieve a more reasonable accuracy-bias tradeoff recommendation. Finally, we conduct extensive experiments on a public dataset and a product dataset to verify the effectiveness of CR2Rec. In addition, CR2Rec is also deployed on a large-scale financial multimedia recommendation platform in China and achieves a sustained performance gain. Dugang Liu, Xing Tang 0007, Liang Chen 0009, Xiuqiang He 0001, Zhong Ming 0001 |
ACM Multimedia | 4 |
| 2023 | Towards Hybrid-grained Feature Interaction Selection for Deep Sparse NetworkabstractDeep sparse networks are widely investigated as a neural network architecture for prediction tasks with high-dimensional sparse features, with which feature interaction selection is a critical component. While previous methods primarily focus on how to search feature interaction in a coarse-grained space, less attention has been given to a finer granularity. In this work, we introduce a hybrid-grained feature interaction selection approach that targets both feature field and feature value for deep sparse networks. To explore such expansive space, we propose a decomposed space which is calculated on the fly. We then develop a selection algorithm called OptFeature, which efficiently selects the feature interaction from both the feature field and the feature value simultaneously. Results from experiments on three large real-world benchmark datasets demonstrate that OptFeature performs well in terms of accuracy and efficiency. Additional studies support the feasibility of our method. All source code are publicly available\footnote{https://anonymous.4open.science/r/OptFeature-Anonymous}. Fuyuan Lyu, Xing Tang 0007, Dugang Liu, Chen Ma 0001, Weihong Luo, Liang Chen 0009, Xiuqiang He 0001, Xue (Steve) Liu |
NeurIPS | 6 |
| 2023 | Curriculum Modeling the Dependence among Targets with Multi-task Learning for Financial MarketingabstractMulti-task learning for various real-world applications usually involves tasks with logical sequential dependence. For example, in online marketing, the cascade behavior pattern of impression \rightarrow click \rightarrow conversion is usually modeled as multiple tasks in a multi-task manner, where the sequential dependence between tasks is simply connected with an explicitly defined function or implicitly transferred information in current works. These methods alleviate the data sparsity problem for long-path sequential tasks as the positive feedback becomes sparser along with the task sequence. However, the error accumulation and negative transfer will be a severe problem for downstream tasks. Especially, at the beginning stage of training, the optimization for parameters of former tasks is not converged yet, and thus the information transferred to downstream tasks is negative. In this paper, we propose a prior information merged model (PIMM), which explicitly models the logical dependence among tasks with a novel prior information merged (PIM) module for multiple sequential dependence task learning in a curriculum manner. Specifically, the PIM randomly selects the true label information or the prior task prediction with a soft sampling strategy to transfer to the downstream task during the training. Following an easy-to-difficult curriculum paradigm, we dynamically adjust the sampling probability to ensure that the downstream task will get the effective information along with the training. The offline experimental results on both public and product datasets verify that PIMM outperforms state-of-the-art baselines. Moreover, we deploy the PIMM in a large-scale FinTech platform, and the online experiments also demonstrate the effectiveness of PIMM. Yunpeng Weng, Xing Tang 0007, Liang Chen 0009, Xiuqiang He 0001 |
SIGIR | 3 |
| 2023 | Optimizing Feature Set for Click-Through Rate PredictionabstractClick-through prediction (CTR) models transform features into latent vectors and enumerate possible feature interactions to improve performance based on the input feature set. Therefore, when selecting an optimal feature set, we should consider the influence of both features and their interaction. However, most previous works focus on either feature field selection or only select feature interaction based on the fixed feature set to produce the feature set. The former restricts search space to the feature field, which is too coarse to determine subtle features. They also do not filter useless feature interactions, leading to higher computation costs and degraded model performance. The latter identifies useful feature interaction from all available features, resulting in many redundant features in the feature set. In this paper, we propose a novel method named OptFS to address these problems. To unify the selection of features and their interaction, we decompose the selection of each feature interaction into the selection of two correlated features. Such a decomposition makes the model end-to-end trainable given various feature interaction operations. By adopting feature-level search space, we set a learnable gate to determine whether each feature should be within the feature set. Because of the large-scale search space, we develop a learning-by-continuation training scheme to learn such gates. Hence, OptFS generates the feature set containing features that improve the final prediction results. Experimentally, we evaluate OptFS on three public datasets, demonstrating OptFS can optimize feature sets which enhance the model performance and further reduce both the storage and computational cost. Fuyuan Lyu, Xing Tang 0007, Dugang Liu, Liang Chen 0009, Xiuqiang He 0001, Xue (Steve) Liu |
WWW | 4 |
| 2022 | Identifying User Relationship on WeChat Money-Gifting NetworkabstractWith the proliferation of online social networks, the identification or classification of real-life relationship between users has been very useful for many applications such as financial fraud detection. In real life, usually people with different relationships would present gifts with special meanings to each other on different dates. In many Asian cultures, especially in Chinese culture, the red packet is a traditional form of monetary gift. With the rapid development of the Internet, people gradually began to give electronic red packets instead of paper ones as the means of money gifting on social network platforms. As motivated, in this paper we advocate a novel approach that exploits users’ red packet interactions for users relationship identification on WeChat, one of the largest social platforms in China. Specifically, we analyze the WeChat red packets network, identify the real-life relationship types between users through mining the semantic information of the amount and sending time of each red packet. In order to better capture the red packet gifting behaviors between users for relationship identification, on one hand, we construct an Amount-Date Graph and apply the graph embedding method to learn embeddings of the amount and sending date of each red packet. On the other hand, we propose a novel sequential model, Cross & Attention Sequence Model (CASM), which explicitly learns the interactions between the latent semantic information of each red packet’s amount and sending date in the red packets sequence between two users. To validate our approach, we conduct comprehensive experiments on a real-world WeChat Users Red Packets dataset that involves 8 kinds of real-life relationships. The experiments show that our proposed approach performs significantly better than baselines and achieves 81.70 percent prediction accuracy. Yunpeng Weng, Liang Chen 0009, Xu Chen 0004 |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2022 | GAIN: Graph Attention & Interaction Network for Inductive Semi-Supervised Learning Over Large-Scale GraphsabstractGraph Neural Networks (GNNs) have led to state-of-the-art performance on a variety of machine learning tasks such as recommendation, node classification and link prediction. Graph neural network models generate node embeddings by merging nodes features with the aggregated neighboring nodes information. Most existing GNN models exploit a single type of aggregator (e.g., mean-pooling) to aggregate neighboring nodes information, and then add or concatenate the output of aggregator to the current representation vector of the center node. However, using only a single type of aggregator is difficult to capture the different aspects of neighboring information and the simple addition or concatenation update methods limit the expressive capability of GNNs. Not only that, existing supervised or semi-supervised GNN models are trained based on the loss function of the node label, which leads to the neglect of graph structure information. In this paper, we propose a novel graph neural network architecture, Graph Attention & Interaction Network (GAIN), for inductive learning on graphs. Unlike the previous GNN models that only utilize a single type of aggregation method, we use multiple types of aggregators to gather neighboring information in different aspects and integrate the outputs of these aggregators through the aggregator-level attention mechanism. Furthermore, we design a graph regularized loss to better capture the topological relationship of the nodes in the graph. Additionally, we first present the concept of graph feature interaction and propose a vector-wise explicit feature interaction mechanism to update the node embeddings. We conduct comprehensive experiments on two node-classification benchmarks and a real-world financial news dataset. The experiments demonstrate our GAIN model outperforms current state-of-the-art performances on all the tasks. Yunpeng Weng, Xu Chen 0004, Liang Chen 0009, Wei Liu 0208 |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2021 | Deep Reinforcement Learning With Spatio-Temporal Traffic Forecasting for Data-Driven Base Station Sleep ControlabstractTo meet the ever increasing mobile traffic demand in 5G era, base stations (BSs) have been densely deployed in radio access networks (RANs) to increase the network coverage and capacity. However, as the high density of BSs is designed to accommodate peak traffic, it would consume an unnecessarily large amount of energy if BSs are on during off-peak time. To save the energy consumption of cellular networks, an effective way is to deactivate some idle base stations that do not serve any traffic demand. In this paper, we develop a traffic-aware dynamic BS sleep control framework, named DeepBSC, which presents a novel data-driven learning approach to determine the BS active/sleep modes while meeting lower energy consumption and satisfactory Quality of Service (QoS) requirements. Specifically, the traffic demands are predicted by the proposed GS-STN model, which leverages the geographical and semantic spatial-temporal correlations of mobile traffic. With accurate mobile traffic forecasting, the BS sleep control problem is cast as a Markov Decision Process that is solved by Actor-Critic reinforcement learning methods. To reduce the variance of cost estimation in the dynamic environment, we propose a benchmark transformation method that provides robust performance indicator for policy update. To expedite the training process, we adopt a Deep Deterministic Policy Gradient (DDPG) approach, together with an explorer network, which can strengthen the exploration further. Extensive experiments with a real-world dataset corroborate that our proposed framework significantly outperforms the existing methods. Qiong Wu 0009, Xu Chen 0004, Zhi Zhou 0006, Liang Chen 0009, Junshan Zhang |
IEEE/ACM Trans. Netw. | 4 |
| 2020 | DeepCP: Deep Learning Driven Cascade Prediction-Based Autonomous Content Placement in Closed Social NetworkabstractOnline social networks (OSNs) are emerging as the most popular mainstream platform for content cascade diffusion. In order to provide satisfactory quality of experience (QoE) for users in OSNs, much research dedicates to proactive content placement by using the propagation pattern, user's personal profiles and social relationships in open social network scenarios (e.g., Twitter and Weibo). In this paper, we take a new direction of popularity-aware content placement in a closed social network (e.g., WeChat Moment) where user's privacy is highly enhanced. We propose a novel data-driven holistic deep learning framework, namely DeepCP, for joint diffusion-aware cascade prediction and autonomous content placement without utilizing users' personal and social information. We first devise a time-window LSTM model for content popularity prediction and cascade geo-distribution estimation. Accordingly, we further propose a novel autonomous content placement mechanism CP-GAN which adopts the generative adversarial network (GAN) for agile placement decision making to reduce the content access latency and enhance users' QoE. We conduct extensive experiments using cascade diffusion traces in WeChat Moment (WM). Evaluation results corroborate that the proposed DeepCP framework can predict the content popularity with a high accuracy, generate efficient placement decision in a real-time manner, and achieve significant content access latency reduction over existing schemes. Qiong Wu 0009, Muhong Wu, Xu Chen 0004, Zhi Zhou 0006, Kaiwen He 0001, Liang Chen 0009 |
IEEE J. Sel. Areas Commun. | 6 |
| 2019 | Mobile Social Data Learning for User-Centric Location Prediction With Application in Mobile Edge Service MigrationabstractRecently, location prediction has attracted considerable research effort because of the popularity of location-based services, such as mobile advertising and recommendations. With the unprecedented proliferation of mobile social networks, such as WeChat and Twitter, we are able to use location service to bridge the online and offline worlds, which is of great significance to many smart city applications. Different from existing studies, in this paper, we promote a user-centric location prediction approach by leveraging a user's local mobile social information without involving other users' location privacy. We propose a factor graph learning model that integrates not only user's social and network information but also the correlations between a user's locations into a unified framework. Furthermore, we use ReliefF algorithm to select user-specific significant features for location prediction and define the measure of location entropy to study the similarity between location, network status, and social behavior. To show the benefit of precise location prediction, we further apply it to personalized service migration in mobile edge computing (MEC) and accordingly propose prediction-based amortizing algorithm and lazy migration algorithm that can well balance the tradeoff between migration cost and non-migration latency in a cost-efficient manner. We conduct extensive experiments using a real-world data trace, which shows that our model performs much better in location prediction compared with several classic methods and the MEC service quality can be significantly enhanced by leveraging the location prediction. Qiong Wu 0009, Xu Chen 0004, Zhi Zhou 0006, Liang Chen 0009 |
IEEE Internet Things J. | 4 |
| 2018 | User-Centric Location Prediction in Mobile Social Networks: A Factor Graph Learning ApproachabstractRecently, location prediction has attracted considerable research effort because of the popularity of location- based services, such as mobile advertising and recommendations. With the unprecedented proliferation of mobile social networks, we are able to use location service to bridge the online and offline worlds. Different from existing studies, in this paper we promote a user-centric location prediction approach by leveraging a user's local mobile social information without involving other users' location privacy. We propose a factor graph learning model that integrates not only user's social and network information, but also the correlations between user's locations into a unified framework. Furthermore, we use ReliefF algorithm to select user-specific significant features for location prediction and define the measure of location entropy to study the similarity between location, network status and social behavior. We conduct extensive experiments using a real-world dataset, which shows that our model performs much better in location prediction compared with several classic methods. Qiong Wu 0009, Xu Chen 0004, Zhi Zhou 0006, Liang Chen 0009 |
GLOBECOM | 4 |
| 2018 | Performance Analysis of Thunder Crystal: A Crowdsourcing-Based Video Distribution PlatformabstractDelivering high-definition (HD) videos to a large number of Internet users is a challenging research problem due to its heavy bandwidth consumption and inelastic quality-of-service (QoS) requirement. Different from the traditional content delivery networks and overlay peer-to-peer networks, crowdsourcing-based platforms, e.g., Thunder Crystal, deliver HD videos by renting agents' bandwidth and storage resources. Cash will be rewarded to agents based on agents' upload traffic. Online video providers, i.e., Tencent Video and YouKu, pay Thunder Crystal for its video distribution service. Therefore, a critical problem for Thunder Crystal is to evaluate the performance it can achieve, which determines Thunder Crystal's competitiveness and its bargaining power with online video providers. Although previously studied proportional video replication and random request scheduling strategy are implemented by Thunder Crystal, the system performance cannot be evaluated by simply using existing models, because of its novel business model. To address this problem, this paper proposes a theoretical framework that can analyze the performance for synchronized streaming, video-on-demand (VoD) streaming, and video downloading, which are all supported by Thunder Crystal. A differentiated bandwidth allocation is designed to boost Thunder Crystal's streaming performance by assigning downloading users more fluctuating bandwidth, which only slightly degrades downloading performance. Finally, simulation is conducted to validate the accuracy of our theoretical results. Yipeng Zhou, Liang Chen 0009, Mi Jing, Zhong Ming 0001, Yuedong Xu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2017 | An Incentive-Based Mixed QoE Framework for Content Delivery to Smart HomesabstractSmart-home is becoming increasingly popular in recent years, and it introduces a new content retrieval paradigm - delay-insensitive downloading. In this new paradigm, users do not require the content retrieval task to finish as soon as possible, but only set a deadline for it. We study the role of this paradigm in the traffic engineering of a chunk-based cloud storage service. We propose that it could help to reduce the high intra-datacenter traffic at peak resulting from the chunk-based architecture, by delaying users' content requests when necessary. Because of the introduction of the new paradigm, we consider the co- existence of three applications (downloading, streaming, delay-insensitive downloading) in the service, and try to understand the best way to delay users' requests. We conduct an incentive-based study for the evaluation of schemes of delaying users' requests: the service provider pays incentives to users to promote this new paradigm. The incentive, as well as users' application-specific QoE on the service, is modelled, and a framework for the evaluation is presented. We then apply our framework to study several proposed delay schemes and present both quantitative and qualitative results. Suiming Guo, Liang Chen 0009, Dah-Ming Chiu |
ICCCN | 2 |
| 2017 | On Carrier Sensing Accuracy and Range Scaling Laws in Nakagami Fading ChannelsabstractWe make a detailed study on carrier sensing of 802.11 in Nakagami fading channels. We prove that to maximize sensing accuracy, the optimal channel accessing probability is solely determined by the path-loss SIR (Signal to Interference Ratio). We define pfail -interference range and pbusy -carrier sensing range for fading channels and prove that their scaling laws in Nakagami fading channels are similar to those in the static channel. The newly derived theoretical results show a unified property between the static and fading channels. By extensive simulations, we reveal that fading depresses the probability of a dominating transmission state, and therefore it can mitigate severe hidden and exposed terminal problems, but fading harms the average sensing accuracy for an optimally adjusted carrier sensing threshold. Yue Wang 0014, Liang Chen 0009, Haifeng Li 0006, Huaihu Cao |
Wirel. Commun. Mob. Comput. | 2 |
| 2016 | Social Group Based Video Recommendation Addressing the Cold-Start Problem
Yipeng Zhou, Liang Chen 0009, Dah-Ming Chiu |
PAKDD (2) | 3 |
| 2016 | Design, Implementation, and Measurement of a Crowdsourcing-Based Content Distribution PlatformabstractContent distribution, especially the distribution of video content, unavoidably consumes bandwidth resources heavily. Internet content providers invest heavily in purchasing content distribution network (CDN) services. By deploying tens of thousands of edge servers close to end users, CDN companies are able to distribute content efficiently and effectively, but at considerable cost. Thus, it is of great importance to develop a new system that distributes content at a lower cost but comparable service quality. In lieu of expensive CDN systems, we implement a crowdsourcing-based content distribution system, Thunder Crystal, by renting bandwidth for content upload/download and storage for content cache from agents. This is a large-scale system with tens of thousands of agents, whose resources significantly amplify Thunder Crystal’s content distribution capacity. The involved agents are either from ordinary Internet users or enterprises. Monetary rewards are paid to agents based on their upload traffic so as to motivate them to keep contributing resources. As far as we know, this is a novel system that has not been studied or implemented before. This article introduces the design principles and implementation details before presenting the measurement study. In summary, with the help of agent devices, Thunder Crystal is able to reduce the content distribution cost by one half and amplify the content distribution capacity by 11 to 15 times. Yipeng Zhou, Liang Chen 0009, Mi Jing, Shenglong Zou, Richard Tianbai Ma |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2015 | Distributing very-large content from cloud to smart home hubs: Measurement and implicationsabstractWith recent proliferation of smart home hubs, public cloud storage service is facing a significant challenge of distributing very large content to end users. In these services, many delivered files are of size of tens of GB (even up to hundreds of GB for 4K resolution videos), and the fraction of very-large files keeps increasing. By observing a commercial cloud storage and distribution service, we notice the issue of excessive intra-datacenter network traffic for distributing very-large files, which seriously affects the quality of service and scalability. Based on the measurement results and the analysis of the issue, we propose cache-based methods (in both static and dynamic schemes) to help reduce the large volume of intra-datacenter traffic. Our proposals are evaluated on the chunk-level traces from one of the biggest commercial cloud storage and distribution service providers in China. The cache-based approach can reduce the maximum intra-datacenter traffic significantly in a cost-effective manner. Liang Chen 0009, Suiming Guo |
ICC | 1 |
| 2015 | DST: Leveraging Delay-Insensitive Workload in Cloud Storage for Smart Home NetworkabstractWe study the problem of how to manage the high intra-datacenter traffic in a chunk-based public cloud storage service serving primarily smart home devices. The large volume of traffic is introduced by delivering very large content during busy hours in the cloud. Measurement of a commercial cloud service shows that the peak traffic volume (at its edge servers) overwhelms the network interface cards (NICs), resulting in serious congestion and packet losses. Since it can be expected the large content downloading requests in smart home environment could be delay-insensitive, we propose DST to keep the peak load under a specified upper bound, by delaying users' requests when necessary. By modelling DST as a queueing system, we derive the relation between the mean delay and the traffic upper bound. With trace-driven simulations, we evaluate the system performance and validate the analysis results. For the commercial cloud service we study, we show that it is possible to keep the traffic upper bound to about 80% of peak traffic rate by introducing a mean delay of around 48 minutes. Suiming Guo, Liang Chen 0009, Dah-Ming Chiu |
ICCCN | 2 |
| 2015 | Analyzing streaming performance in crowdsourcing-based video service systemsabstractCrowdsourcing-based video service systems, e.g., Thunder Crystal, are novel content distribution platforms composed by a large number of agent devices, acting like miniservers. Most agents are normal Internet users who would like to earn rewarded cash by uploading content through their devices. Compared with CDN, the bandwidth cost is much cheaper; while compared with Peer-to-Peer(P2P), the bandwidth supply is more stable. In this work, we create a stochastic model to analyze the live and VoD streaming performance in such crowdsourcingbased video service systems. Simulation is conducted to validate the accuracy of our analytical results. Yipeng Zhou, Liang Chen 0009, Mi Jing, Zhong Ming 0001, Dah-Ming Chiu |
LANMAN | 2 |
| 2015 | Thunder crystal: a novel crowdsourcing-based content distribution platformabstractContent distribution, especially the distribution of video content, unavoidably consumes bandwidth resource heavily. Internet content providers (ICP) spend lots of money to buy content distribution network (CDN) service. By deploying thousands of edge servers close to end users, CDN companies are able to distribute content efficiently. In lieu of traditional CDN systems, we implement a crowdsourcing-based content distribution system, Thunder Crystal, which utilizes agents' upload bandwidth to amplify the content distribution capacity. Agents are well motivated to contribute storage and upload bandwidth to the system by rebated cash. As far as we know, this is a novel system that has not been studied before. In this work, we will present its design principles first. Then, we study agent behavior and methods to evaluate system efficiency and user efficiency. We evaluate the system by simulations, and observe that agents are well motivated to keep online most of the time and amplify the content distribution capacity by 10~20 times. Liang Chen 0009, Yipeng Zhou, Mi Jing, Richard T. B. Ma |
NOSSDAV | 1 |
| 2015 | Turbocharged Video Distribution via P2PabstractThere are two types of P2P systems satisfying two different user demands: 1) file downloading and 2) video-on-demand (VoD) streaming. An example of file downloading is the original BitTorrent, and examples for VoD streaming include various commercial P2P-based VoD streaming systems such as that offered by PPLive. We have a hypothesis - by combining a type: 1) system and 2) system as a single P2P system, both the file downloading users and the streaming users of the same video will benefit in performance. The reasoning is that at any moment, only a subset of the file downloading peers can provide good service to VoD streaming peers and the VoD streaming peers are only good at providing service to a different subset of the file downloading peers. The former subset is the set of peers close to completing the downloading of the video file; whereas the latter subset is the set of peers starting to download a video. In this paper, we propose a novel design for a mesh-based video distribution system without depending on video replication on streaming peers. We produce simple back-of-the-envelop analysis to show its effectiveness. Then, we further validate our design and compare it with other designs through simulation and experiments in practical networking environment by implementing a prototype. Yipeng Zhou, Liang Chen 0009, Tom Z. J. Fu, Dah-Ming Chiu |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2015 | Smart Streaming for Online Video ServicesabstractBandwidth cost is a significant concern for online video service providers. Today's video streaming systems mostly use HTTP streaming, with users accessing video segments as HTTP requests. A frequently used strategy is to serve all user requests as fast as possible, as if the user is downloading a file. The downloading rate can often far exceed the playback rate, when the system is below the peak load. This is known as progressive downloading. Since users may quit before viewing the complete video, however, much of the downloaded video can be “wasted.” By studying and exploiting the predictability of users' departure behavior , the authors developed a smart streaming strategy that can significantly improve overall streaming service quality under given server bandwidth. The improvement is achieved by avoiding the waste based on predicted user departure behavior. The proposed smart streaming technique is evaluated by modeling, analysis, and simulation, as well as experimentation using a prototype implementation. Liang Chen 0009, Yipeng Zhou, Dah-Ming Chiu |
IEEE Trans. Multim. | 1 |
| 2015 | Video Popularity Dynamics and Its Implication for ReplicationabstractPopular online video-on-demand (VoD) services all maintain a large catalog of videos for their users to access. The knowledge of video popularity is very important for system operation , such as video caching on content distribution network (CDN) servers. The video popularity distribution at a given time is quite well understood. We study how the video popularity changes with time, for different types of videos, and apply the results to design video caching strategies. Our study is based on analyzing the video access levels over time, based on data provided by a large video service provider. Our main finding is, while there are variations, the glory days of a video’s popularity typically pass by quickly and the probability of replaying a video by the same user is low. The reason appears to be due to fairly regular number of users and view time per day for each user, and continuous arrival of new videos. All these facts will affect how video popularity changes, hence also affect the optimal video caching strategy. Based on the observation from our measurement study, we propose a mixed replication strategy (of LFU and FIFO) that can handle different kinds of videos. Offline strategy assuming tomorrow’s video popularity is known in advance is used as a performance benchmark. Through trace-driven simulation, we show that the caching performance achieved by the mixed strategy is very close to the performance achieved by the offline strategy. Yipeng Zhou, Liang Chen 0009, Dah-Ming Chiu |
IEEE Trans. Multim. | 2 |
| 2015 | Analysis and Detection of Fake Views in Online Video ServicesabstractOnline video-on-demand(VoD) services invariably maintain a view count for each video they serve, and it has become an important currency for various stakeholders, from viewers, to content owners, advertizers, and the online service providers themselves. There is often significant financial incentive to use a robot (or a botnet) to artificially create fake views. How can we detect fake views? Can we detect them (and stop them) efficiently? What is the extent of fake views with current VoD service providers? These are the questions we study in this article. We develop some algorithms and show that they are quite effective for this problem. Liang Chen 0009, Yipeng Zhou, Dah-Ming Chiu |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2014 | A lifetime model of online video popularityabstractPopular online Video-on-Demand (VoD) services all maintain a large catalog of videos for their users to access. The number of servers assigned to serve each video is directly related to the relative popularity of the video. The distribution of popularity at a given time is quite well understood. We study how the video popularity changes over its lifetime, for different types of videos. Our study is based on analyzing the video access levels over time, based on data provided by a large video service provider. Our main finding is, while there are variations, the glory days of a video typically pass by quickly and probability of replaying a video by the same user is low. The reason appears to be due to fairly regular number of users and view time per day for each user, and continuous arrival of new videos. We then discuss the implication of our findings for video replication and recommendation. Liang Chen 0009, Yipeng Zhou, Dah-Ming Chiu |
ICCCN | 1 |
| 2014 | A measurement study of the potential benefits for peer-assisted mobile VoDabstractRapid growth in users of smartphones and other mobile devices is increasing the demand for online video service to these mobile devices. The mobile video on demand (VoD) service is almost exclusively provided by content delivery network (CDN) servers since mobile devices are not powerful enough to provide any peer-to-peer (P2P) service. However, in a mixed VoD system with both non-mobile and mobile peers, an interesting question is to ask whether we can leverage non-mobile peers to provide better support for VoD. In this paper, we try to answer this question by a measurement study on Tencent's existing VoD systems. Our measurement results show various interesting characteristics and potential benefits of leveraging non-mobile users. Such results will guide us towards a practical system design and implementation in the near future. Jiqiang Wu, Liang Chen 0009, Dah-Ming Chiu, Youwei Hua, Zirong Zhu |
IWCMC | 2 |
| 2014 | Fake View Analytics in Online Video ServicesabstractOnline video-on-demand (VoD) services invariably maintain a view count for each video they serve, and it has become an important currency for various stakeholders, from viewers, to content owners, advertizers, and the online service providers themselves. There is often significant financial incentive to use a robot (or a botnet) to artificially create fake views. How can we detect the fake views? Can we detect them (and stop them) efficiently? What is the extent of fake views with current VoD service providers? These are the questions we study in this paper. We develop some algorithms and show their effectiveness for this problem. Liang Chen 0009, Yipeng Zhou, Dah-Ming Chiu |
NOSSDAV | 1 |
| 2014 | A study of user behavior in online VoD services
Liang Chen 0009, Yipeng Zhou, Dah-Ming Chiu |
Comput. Commun. | 1 |
| 2013 | Video Browsing - A Study of User Behavior in Online VoD ServicesabstractA big portion of Internet traffic nowadays is video. A good understanding of user behavior in online VoD systems can help us design, configure and manage video content distribution. With the help of a major video on demand (VoD) service provider, we conduct a detailed study of user behavior watching streamed videos over the Internet. We engineered the video player at the client side to collect user behavior reports for over 540 million sessions.In order to isolate the possible effect of session quality of experience (QoE) on user behavior, we focus on the sessions with perfect QoE, and leave out those sessions with QoE impairments (such as freezes).Our main finding is that users spend a lot of time browsing: viewing part of one video after another, and only occasionally (around 20% of the time) watching a video to its completion. We consider seek (jump to a new position of the video) as a special form of browsing - repeating partial viewing of the same video. Our analysis leads towards a user behavior model in which a user transitions through a random number of short views before a longer view, and repeats the process a random number of times. A purely abstract version of such user behavior model was proposed by Wu et al [1] as a closed queueing network formulation. Our study uncovers the parameters and distributions of such a stochastic behavior model based on observations in practice. Liang Chen 0009, Yipeng Zhou, Dah-Ming Chiu |
ICCCN | 1 |