VLDB 2026 Research / reviewers in the wild / expert
Xiaofeng Gao 0001
dblp:95/6947-1
· DBLP profile ↗
112ranked-venue papers in the field
7as first author
69since 2021 · last 2026
0000-0003-1776-8799ORCID · conflict
Domains — venue-derived; a paper can count in several
Database Systems & Data Management · 65 (5 first)Data Mining & Knowledge Discovery · 31 (2 first)Information Retrieval & Web Search · 14Big Data, Cloud & Distributed Data Systems · 1Other / Interdisciplinary · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | ZTune: Model-Assisted Reinforcement Learning for Executor Tuning on Ad Hoc Spark SQL Query
Dejun Kong 0001, Xiuqi Huang, Xiaofeng Gao 0001 |
DASFAA (2) | 3 |
| 2026 | PAPT: Periodic-Aware Transformer with Polynomial Trend Fitting for General Time Series Forecasting
Xiuyuan Wei, Jiadong Chen, Yinbo Sun, Xiaofeng Gao 0001, Lintao Ma, Guihai Chen |
DASFAA (5) | 5 |
| 2026 | Disentangled Parameter-Efficient Linear Model for Long-Term Time Series Forecasting
Yuang Zhao, Jiadong Chen, Shenrong Ye, Fuxin Jiang, Xiaofeng Gao 0001 |
DASFAA (5) | 6 |
| 2026 | When Complex Event Recognition Meets Cloud-Native Architectures
Shizhe Liu, Haipeng Dai 0001, Meng Li 0010, Yuemeng Zhang, Shaoxu Song, Zhifeng Bao, Hancheng Wang, Xiaofeng Gao 0001, Guihai Chen |
ICDE | 8 |
| 2026 | Cross-Domain Interest Representation Learning for Scenario- and Task-Aware RecommendationabstractMany internet companies operate multiple flagship applications, each of which can be regarded as a distinct business domain, covering areas such as video, reading, and gaming. Within each domain, diverse recommendation scenarios coexist, and users engage in various tasks with heterogeneous behaviors. In our industrial setting, we observe three key phenomena that existing methods rarely address: (i) users' Cross-Domain Interests (CDI) are weakly exploited, since behavior sequences are often pooled for efficiency, losing transferable cross-domain dependencies; (ii) multi-domain, scenario, and task variations are under-modeled, making it difficult to capture fine-grained complementarities; and (iii) multimodal features remain misaligned with ID features, especially when different domains emphasize different modalities. These gaps hinder cross-product collaboration. Bokai Lin, Naijun Gao, Yucen Gao, Heng Chang, Cheng Hu 0002, Zhinan Zhang, Xiaofeng Gao 0001 |
SIGIR | 7 |
| 2026 | Decoupling User Features for User Cold-Start App Recommendation: Static Attributes versus Behavioral SequencesabstractIn app recommendation, user cold-start remains a fundamental challenge in recommender systems. Existing approaches primarily focus on efficiently leveraging limited data or transferring knowledge from active users to alleviate the user cold-start problem, yet they often overlook the influence of feature interactions on user cold-start. We group features according to their semantic types and identify an interesting phenomenon: user attribute features and behavioral sequence features interfere with each other, thereby constraining the model's ability to represent cold-start users effectively. We attribute this issue to differences in the latent space distributions and learning complexities of the two feature types, which hinder the model from accurately capturing cold-start users' interests. To address this challenge, we propose the AFIM architecture, which decouples the learning of user attribute and behavior sequential features. AFIM leverages a lightweight attention module to explicitly capture user interests from behavioral sequences, thereby reducing the learning burden on downstream recommendation networks. Additionally, it incorporates feature decoupling and dynamic fusion modules to mitigate learning bias arising from heterogeneous feature spaces. Extensive experiments on two public datasets and two industrial datasets demonstrate that AFIM consistently outperforms SOTA baselines, highlighting its effectiveness in user cold-start scenarios. Li Ma 0012, Yue Ding 0001, Xiaofeng Gao 0001 |
WSDM | 5 |
| 2026 | Macro-Micro Collaborative Learning for Logical Data Center Microservice Indicators ForecastingabstractAs microservice architecture is evolving toward Logical Data Center (LDC), accurate forecasting of the microservices indicators can support reasonable resource allocation, thereby ensuring the availability and reliability of cloud service. From a macro perspective, due to the architecture hierarchy, microservices exhibit: 1) collaborative relationships derived from shared functionalities, 2) backup relationships between replicas, and 3) dynamic correlation driven by cooperation. From a micro perspective, there exist causal relationships among indicators within a microservice. That is, workload will first impact system consumption, such as CPU and memory usage, then affect service quality like system latency. Based on these insights, we propose MaMiClif, a macro-micro collaborative learning framework for LDC microservice indicators forecasting. MaMiClif constructs Macro Graph and Micro Matrix to model the microservices dependencies and the causality of indicators. To learn fine-grained indicator dependencies, Indicator-Centric Embedding is leveraged to generate representations for indicator series. We use Heterogeneous Graph Convolution to update workload representations based on the Macro Graph, and adopt Causal Sparse Self-attention to integrate causal strength into the self-attention calculation, enabling a comprehensive exploration of dependencies among indicators. Experiments on two datasets, including LDC_MS, which was collected from the LDC system of Ant Group, demonstrate the effectiveness of MaMiClif. Mohan Gao, Zhemeng Yu, Yang Luo 0004, Lintao Ma, Yinbo Sun, Xiaofeng Gao 0001 |
WWW | 7 |
| 2026 | Encoder-decoder-based workload forecasting framework for database-as-a-service
Yunlong Cheng, Xiuqi Huang, Xiaofeng Gao 0001, Guihai Chen |
Knowl. Inf. Syst. | 3 |
| 2026 | Uncertainty-Aware Online Time Series Multi-Step Forecasting Framework in Cloud Systems
Jiadong Chen, Yang Luo 0004, Xiuqi Huang, Fuxin Jiang, Yangguang Shi, Tieying Zhang, Xiaofeng Gao 0001 |
IEEE Trans. Knowl. Data Eng. | 7 |
| 2025 | Predicting Enterprise Users' Consuming Potential for Cloud Services
Yunlong Cheng, Tianyao Shi, Xiuyuan Wei, Yulong Song, Xiaofeng Gao 0001, Zhipeng Bian, Zhenli Sheng |
DASFAA (1) | 5 |
| 2025 | CloudChurn: Optimizing Enterprise Customer Churn Prediction in Cloud Services for Huawei Cloud
Hengyu Ye, Yulong Song, Zhipeng Bian, Xiaofeng Gao 0001, Guihai Chen, Xin Jin 0008, Zhenli Sheng |
DASFAA (6) | 4 |
| 2025 | CausalScaler: A Causality-Driven Autoscaling Framework for the Cloud
Zhemeng Yu, Yang Luo 0004, Yucen Gao, Yinbo Sun, Xiaofeng Gao 0001, Lintao Ma, Guihai Chen |
DASFAA (4) | 5 |
| 2025 | Towards Unifying Feature Interaction Models for Click-Through Rate Prediction
Junwei Pan, Jipeng Jin, Shudong Huang, Xiaofeng Gao 0001, Lei Xiao 0001 |
ECML/PKDD (5) | 5 |
| 2025 | DRNCS: Dual-Level Route Generation Model Based on Node Contraction and Shortcuts
Yucen Gao, Xinle Li, Xiaofeng Gao 0001, Guihai Chen |
ECML/PKDD (3) | 6 |
| 2025 | Towards Personalized Federated Multi-Scenario Multi-Task RecommendationabstractIn modern recommender systems, especially in e-commerce, predicting multiple targets such as click-through rate (CTR) and post-view conversion rate (CTCVR) is common. Multi-task recommender systems are increasingly popular in both research and practice, as they leverage shared knowledge across diverse business scenarios to enhance performance. However, emerging real-world scenarios and data privacy concerns complicate the development of a unified multi-task recommendation model. Yue Ding 0001, Yanbiao Ji, Xin Xin 0003, Suizhi Huang, Chang Liu 0078, Xiaofeng Gao 0001, Tsuyoshi Murata, Hongtao Lu 0001 |
WSDM | 8 |
| 2025 | Overlap-aware influence maximization with balanced replay deep Q-network
Yuxin Zuo, Xiuqi Huang, Tiantian Wei, Jianxiong Guo, Xiaofeng Gao 0001, Guihai Chen |
Knowl. Inf. Syst. | 5 |
| 2025 | Fremer: Lightweight and Effective Frequency Transformer for Workload Forecasting in Cloud ServicesabstractWorkload forecasting is pivotal in cloud service applications, such as auto-scaling and scheduling, with profound implications for operational efficiency. Although Transformer-based forecasting models have demonstrated remarkable success in general tasks, their computational efficiency often falls short of the stringent requirements in large-scale cloud environments. Given that most workload series exhibit complicated periodic patterns, addressing these challenges in the frequency domain offers substantial advantages. To this end, we propose Fremer, an efficient and effective deep forecasting model. Fremer fulfills three critical requirements: it demonstrates superior efficiency, outperforming most Transformer-based forecasting models; it achieves exceptional accuracy, surpassing all state-of-the-art (SOTA) models in workload forecasting; and it exhibits robust performance for multi-period series. Furthermore, we collect and open-source four high-quality, open-source workload datasets derived from ByteDance's cloud services, encompassing workload data from thousands of computing instances. Extensive experiments on both our proprietary datasets and public benchmarks demonstrate that Fremer consistently outperforms baseline models, achieving average improvements of 5.5% in MSE, 4.7% in MAE, and 8.6% in SMAPE over SOTA models, while simultaneously reducing parameter scale and computational costs. Additionally, in a proactive auto-scaling test based on Kubernetes, Fremer improves average latency by 18.78% and reduces resource consumption by 2.35%, underscoring its practical efficacy in real-world applications. Hengyu Ye, Jiadong Chen, Xiao He 0008, Fuxin Jiang, Tieying Zhang, Jianjun Chen 0001, Xiaofeng Gao 0001 |
Proc. VLDB Endow. | 7 |
| 2025 | Building Robust and Trustworthy HGNN Models: A Learnable Threshold Approach for Node ClassificationabstractMessage passing scheme is a general idea for Graph Neural Networks (GNNs) to learn node representations. During message passing, given a target node, we transform and aggregate the feature vectors of its neighbors and generate a representation vector for the target node. However, real-world graph data is usually constructed from complicated scenarios based on manually pre-defined rules; it is often the case that noisy information gets involved in message passing, thereby resulting in sub-optimal performance for GNNs and also impacting their trustworthiness and reliability. In this study, we present an effective learnable threshold technique that explicitly optimizes heterogeneous graph structure with the goal to maximize performance improvement of GNNs for downstream tasks. We give an explanation about the design of the learnable threshold and show the ability that our model can be applied to large-scale graphs. Experiments on seven datasets show that our model has a powerful ability to deal with homogeneous graphs with low homophily ratio and dense graphs. With the verification of robustness analysis, our model can resist the noisy information, which proves the robustness of our model. Li Ma 0012, Yongchao Liu 0004, Xiaofeng Gao 0001, Peng Zhang 0001, Chuntao Hong |
ACM Trans. Knowl. Discov. Data | 3 |
| 2024 | MSTEM: Masked Spatiotemporal Event Series Modeling for Urban Undisciplined Events ForecastingabstractUrban undisciplined events (UUE) are of increasing concern to urban officials because they reduce the quality of life and cause societal disorder. How to accurately predict future occurrences is a key point in preventing these events. However, existing supervised methods struggle to perform well on sparse UUEs while self-supervised MAE-based methods adopt a traditional random masking strategy which leads to limited performance on UUE forecasting. Fortunately, we have designed an innovative spatiotemporal masking strategy and its corresponding pre-training task called Masked Spatio-Temporal Event Series Modeling (MSTEM). Through Cluster-assisted region masking, MSTEM efficiently distributes masked regions evenly among different clusters, enhancing the model's ability to capture spatial correlation and heterogeneity while addressing sparse region distribution of UUEs. Frequency-enhanced patch masking helps the model to sufficiently extract the temporal features of UUEs by reconstructing multiple views. Additionally, we propose future merge and cluster label modeling to enhance the extraction of spatiotemporal dependencies, thereby improving the performance of MSTEM on downstream prediction tasks. Experimental evaluations on four real-world datasets including crimes and disorderly conduct show that our masked autoencoder with MSTEM outperforms most of the state-of-the-art baselines. Zehao Gu, Yun Xiong, Yang Luo 0004, Hongrun Ren, Qiang Wang 0066, Xiaofeng Gao 0001, Philip S. Yu |
CIKM | 7 |
| 2024 | Pareto-based Multi-Objective Recommender System with Forgetting CurveabstractRecommender systems with cascading architecture play an increasingly significant role in online recommendation platforms, where the approach to dealing with negative feedback is a vital issue. For instance, in short video ad platforms, users tend to quickly slip away from ad candidates that they feel aversive, and recommender systems are expected to receive these explicit negative feedback and make adjustments to avoid these recommendations.Considering recency effect in memories, we propose a forgetting model based on Ebbinghaus Forgetting Curve to cope with negative feedback. In addition, we introduce a Pareto optimization solver to guarantee a better trade-off between recency and model performance.In conclusion, we propose Pareto-based Multi-Objective Recommender System with forgetting curve (PMORS), which can be applied to any multi-objective recommendation and show sufficiently superiority when facing explicit negative feedback.We have conducted evaluations of PMORS and achieved favorable outcomes in short-video scenarios on both public dataset and industrial dataset. After being deployed on an online short video ad platform named WeChat Channels Ads in May, 2023, PMORS has not only demonstrated promising results for both consistency and recency but also achieved an improvement of up to +1.45% Gross Merchandise Volume (GMV). Jipeng Jin, Zhaoxiang Zhang 0006, Xiaofeng Gao 0001, Xiongwen Yang, Lei Xiao 0001, Jie Jiang 0015 |
CIKM | 4 |
| 2024 | REDI: Recurrent Diffusion Model for Probabilistic Time Series ForecastingabstractTime series forecasting (TSF) consists of point prediction and probabilistic forecasting. Unlike point forecasting which predicts an expected value of a future target, probabilistic time series forecasting models the uncertainty in data by predicting the distribution of future values, which enhances decision-making flexibility and improves risk management. Traditional probabilistic forecasting methods usually assume a fixed distribution of data, which is not always true for time series. Recently, there have been efforts to adapt diffusion models for time series owing to their exceptional ability to model the distribution of data without prior assumptions. However, how to apply advantages of diffusion models to time series forecasting remains a substantial challenge due to specific issues in time series such as distribution drift and complex dynamic temporal patterns. Zehao Gu, Yun Xiong, Yang Luo 0004, Qiang Wang 0066, Xiaofeng Gao 0001 |
CIKM | 6 |
| 2024 | DeepMIN: Deep Multi-modal Interest Network with Cognitive Learning Modules
Zhaoxiang Zhang 0006, Jipeng Jin, Xiaofeng Gao 0001, Xiongwen Yang, Lei Xiao 0001 |
DASFAA (3) | 4 |
| 2024 | Corruption Robust Dynamic Pricing in Liner Shipping under Capacity ConstraintabstractThe shipping industry has irreplaceable importance in international trade and commerce. How to dynamically price different containers has long been a hot topic due to its direct connection to the final revenue. Two critical observations have been made after a comprehensive survey within a top liner company, China Ocean Shipping Company (COSCO). (1) Each type of container carried on a liner ship has its maximum capacity. (2) The sales volume is occasionally subject to huge fluctuations due to rare uncontrollable factors, such as COVID. Based on the above two points and the liner routine's periodic nature, we model the dynamic pricing problem as an episodic MDP model integrating with both capacity constraints and adversarial corruption, named C3-MDP. To maximize the cumulative revenue in the C3-MDP setting, we propose a programming framework, Bonus-Exploration based Episodic Programming (BEEP). This framework can directly accommodate the linear programming algorithm to form the algorithm BEEP-LP, which provides the episode-wise greedy optimal strategy. Furthermore, a detailed regret analysis is provided, showing that BEEP-LP has a regret that is sublinear in the number of episodes. Combining deep techniques, we also present an approximation algorithm BEEP-DQN in the case of large state-action space to strike a balance between the running time and the performance. Abundant experiments based on real container sales data exhibit the rationality of C3-MDP and the effectiveness of BEEP. Yongyi Hu, Xikai Wei, Yangguang Shi, Xiaofeng Gao 0001, Guihai Chen |
ICDE | 5 |
| 2024 | MetaSTC: A Backbone Agnostic Spatio-Temporal Framework for Traffic ForecastingabstractTraffic flow prediction is a critical issue in transportation engineering and presents distinct challenges when handling large-scale datasets in the real world. Existing complex spatio-temporal forecasting paradigms use the same parameters to fit traffic sequences with varying spatio-temporal features, and tend to train an average performance model over different time series. This approach greatly reduces their accuracy when applied to larger road networks. Moreover, the significant differences in traffic data distribution from one city to another can also pose great challenges. The same model may be excellent for one city and mediocre when applied to another. To this end, we propose a Meta Backbone Agnostic Spatio-Temporal Clustering Framework for Traffic Forecasting on Large-Scale Road Networks named MetaSTC. We tackle the disparities of spatio-temporal features of traffic flow through a spatio-temporal clustering-based strategy. We design meta-learner for large-scale road network that dynamically extracts the shared information across roads in the same sub-task. In this way, the model can represent task-specific details with a simpler model and make quick and accurate predictions. Our paradigm is backbone-agnostic and can be combined with different traffic prediction models, solving the problem caused by the difference in data distribution. Extensive experimental results conducted on real-world traffic dataset demonstrate the high accuracy and computational efficiency of our model over SOTA approaches. Zhemeng Yu, Yucen Gao, Songjian Zhang, Xiaofeng Gao 0001, Guihai Chen |
ICDM | 6 |
| 2024 | Weather Knows What Will Occur: Urban Public Nuisance Events Prediction and Control with Meteorological AssistanceabstractUrban public nuisance events, like garbage exposure, illegal parking, facilities damage, and etc., impair the quality of life for city residents. Predicting and controlling these nuisances is crucial but complicated due to their ties to subjective and psychological factors. In this study, we reveal a significant correlation between such nuisances and meteorological indicators, influenced by the impact of climate on people's psychological states. We employ meteorology predictions that are integrated in Hawkes processes to enhance the accuracy of predicting the category and timing of these nuisances. To this end, we propose Spatial-Temporal Two-Tower Transformer (ST-T3), which simultaneously considers spatial data and further improves the prediction accuracy. Evaluated by about three-year data from both downtown and suburban Shanghai, our method outperforms both traditional and advanced prediction systems. We share a portion of the de-identified dataset for open research. Yi Xie 0003, Yun Xiong, Xiuqi Huang, Xiaofeng Gao 0001, Chao Chen 0004, Qiang Wang 0066 |
KDD | 5 |
| 2024 | Online Preference Weight Estimation Algorithm with Vanishing Regret for Car-Hailing in Road NetworkabstractCar-hailing services play an important role in the modern transportation system, and the utilities of the service providers highly depend on the efficiency of route planning algorithms. A widely adopted route planning framework is to assign weights to roads and compute the routes with the shortest path algorithms. Existing techniques of weight-assigning often focus on the traveling time and length of the roads, but cannot incorporate with the preferences of the passengers (users). Yucen Gao, Zhehao Zhu, Mingqian Ma, Yangguang Shi, Xiaofeng Gao 0001 |
KDD | 7 |
| 2024 | Integrating System State into Spatio Temporal Graph Neural Network for Microservice Workload PredictionabstractMicroservice architecture has become a driving force in enhancing the modularity and scalability of web applications, as evidenced by the Alipay platform's operational success. However, a prevalent issue within such infrastructures is the suboptimal utilization of CPU resources due to inflexible resource allocation policies. This inefficiency necessitates the development of dynamic, accurate workload prediction methods to improve resource allocation. In response to this challenge, we present STAMP, a Spatio Temporal Graph Network for Microservice Workload Prediction. STAMP is designed to comprehensively address the multifaceted interdependencies between microservices, the temporal variability of workloads, and the critical role of system state in resource utilization. Through a graph-based representation, STAMP effectively maps the intricate network of microservice interactions. It employs time series analysis to capture the dynamic nature of workload changes and integrates system state insights to enhance prediction accuracy. Our empirical analysis, using three distinct real-world datasets, establishes that STAMP exceeds baselines by achieving an average boost of 5.72% in prediction precision, as measured by RMSE. Upon deployment in Alipay's microservice environment, STAMP achieves a 33.10% reduction in resource consumption, significantly outperforming existing online methods. This research solidifies STAMP as a validated framework, offering meaningful contributions to the field of resource management in microservice architecture-based applications. Yang Luo 0004, Mohan Gao, Zhemeng Yu, Haoyuan Ge, Xiaofeng Gao 0001, Tengwei Cai, Guihai Chen |
KDD | 5 |
| 2024 | A Dual-Embedding Based DQN for Worker Recruitment in Spatial Crowdsourcing with Social NetworkabstractSpatial Crowdsourcing (SC) is a promising service that incentives workers to finish location-based tasks with high quality by providing rewards. Worker recruitment is a core issue in SC, for which most state-of-the-art algorithms focus on designing incentive mechanisms based on the existing SC worker pool. However, they may fail when the number of SC workers is not enough, especially for the new SC platforms. In recent years, social networks have been found to be helpful for worker recruitment by selecting seed workers to spread the task information so as to inspire more social users to participate, but how to select seed workers remains a challenge. Existing methods typically require numerous iterative searches leading to inefficiency in facing the big picture and failing to cope with dynamic environments. Yucen Gao, Wei Liu 0189, Jianxiong Guo, Xiaofeng Gao 0001, Guihai Chen |
SIGIR | 4 |
| 2024 | Towards Resource Efficiency: Practical Insights into Large-Scale Spark Workloads at ByteDanceabstractAt ByteDance, where we execute over a million Spark jobs and handle 500PB of shuffled data daily, ensuring resource efficiency is paramount for cost savings. However, achieving optimization of resource efficiency in large-scale production environments poses significant challenges. Drawing from our practical experiences, we have identified three key issues critical to addressing resource efficiency in real-world production settings: 1 slow I/Os leading to excessive CPU and memory idleness, 2 coarse-grained resource control causing wastage, and 3 sub-optimal job configurations resulting in low utilization. To tackle these issues, we propose a resource efficiency governance framework for Spark workloads. Specifically, 1 we devise the multi-mechanism shuffle services, including Enhanced External Shuffle Service (ESS) and Cloud Shuffle Service (CSS), where CSS employs a push-based approach to enhance I/O efficiency through sequential reading. 2 We modify the Spark configuration parameter protocol, allowing for fine-grained resource control by introducing several new parameters such as milliCores and memoryBurst, as well as supporting operators with additional spill modes. 3 We design a two-stage configuration autotuning method, comprising rule-based and algorithm-based tuning, providing more reliable Spark configuration optimizations. By deploying these techniques on millions of Spark jobs in production over the last two years, we have achieved over 22% CPU utilization increase, 5% memory utilization increase, and 10% shuffle block time ratio decrease, effectively saving millions of CPU cores and petabytes of memory daily. Xiuqi Huang, Wei Zhongjia, Hang Cheng, Chaohui Xin, Zuzhi Chen, Binbin Chen 0005, Yufei Wu 0014, Hao Wang 0210, Tieying Zhang, Xiaofeng Gao 0001, Yuming Liang, Pengwei Zhao, Guihai Chen |
Proc. VLDB Endow. | 12 |
| 2023 | An Adaptive Data-Driven Imputation Model for Incomplete Event Series
Jiadong Chen, Hengyu Ye, Xiaofeng Gao 0001, Fan Wu 0006, Linghe Kong, Guihai Chen |
ADMA (1) | 3 |
| 2023 | SACA: An End-to-End Method for Dispatching, Routing, and Pricing of Online Bus-Booking
Yucen Gao, Yulong Song, Xikai Wei, Xiaofeng Gao 0001, Guihai Chen |
DASFAA (4) | 4 |
| 2023 | Attentive Hawkes Process Application for Sequential Recommendation
Shuodian Yu, Li Ma 0012, Xiaofeng Gao 0001, Jianxiong Guo, Guihai Chen |
DASFAA (2) | 3 |
| 2023 | DBAugur: An Adversarial-based Trend Forecasting System for Diversified WorkloadsabstractTrend forecasting is vital to optimize the workload performance. It becomes even more urgent with an increasing number of applications and database configurations. However, DBAs mainly target at historical workloads and may give suboptimal configuration advice when the workload trends have changed. Although there are some studies on trend forecasting, they have several limitations. First, they mainly predict the changes of query numbers, which do not combine other critical factors (e.g., disk utilization) and cannot fully reflect the future workload trends. Besides, there are numerous queries in the workloads and exact clustering algorithms like K-means cannot effectively merge similar queries which contain noises like time shifts. Second, basic machine learning models like RNN may have relatively low prediction accuracy on complex workloads (e.g., no cycles but random bursts). Third, real-world workloads may have diverse patterns, while previous models cannot efficiently and reliably predict for all the different workload patterns.To address these challenges, we propose a trend forecasting system (DBAugur) that utilizes adversarial neural networks to predict the trends of different workloads. First, DBAugur collects the important features (e.g., queries, resource metrics) to characterize workloads, and reduces the number of involved queries by separately merging similar queries based on the SQL semantics and trend patterns. Second, DBAugur utilizes Generative Adversarial Networks (GANs) to capture the latent patterns, correlations between different metrics, and occasional bursts within the complicated and time-varying workloads. Moreover, we further propose a time-sensitive ensemble algorithm that takes advantage of various machine learning models (e.g., generative models, convolutional models, feed-forward models) to accommodate the various workload patterns. The experimental results show that DBAugur outperformed state-of-the-art methods on various real-world workloads. Yuanning Gao, Xiuqi Huang, Xuanhe Zhou, Xiaofeng Gao 0001, Guoliang Li 0001, Guihai Chen |
ICDE | 4 |
| 2023 | Online Shipping Container Pricing Strategy Achieving Vanishing Regret with Limited InventoryabstractWith the growing demand for global trade transportation, the shipping container market has gained an increasingly important position. As a key issue of the market, container pricing is regarded as an important indicator to adjust the market supply and demand as well as the revenue of liner enterprises. Although various methods aimed at increasing enterprise revenue, such as expert pricing and dynamic pricing, have been proposed by industry and academia in recent years, these approaches rarely yield worst-case performance guarantee for the double-sided online scenarios of commodities and buyers.To cater to the double-sided online scenario and provide theoretical performance guarantee, we propose an online learning-based pricing framework named Balancing Inventory and Revenue with -chasing Decider (BIRD). BIRD determines container price by combining advantages of given multiple online pricing strategies. We utilize a strategy selector A to select a proper target strategy and use an ϵ-chasing decider ${{{\mathfrak{D}}}^{Cha\operatorname{s} ing}}$ to determine the price. BIRD is proven to combine the advantages of multiple online pricing strategies to achieve the performance close to the posterior optimal strategy for any sequence of online buyers on realistic sales platforms with inventory limitation. BIRD is proved to yield a vanishing regret for the online posted pricing problem with the features of limited inventory and multi-unit demand. Based on the historical data provided by COSCO, one of the largest liner enterprises in the world, we experimentally demonstrate the effectiveness of the proposed algorithm. Yucen Gao, Xikai Wei, Xi Jing, Yangguang Shi, Xiaofeng Gao 0001, Guihai Chen |
ICDE | 5 |
| 2023 | Automatic Fusion Network for Cold-start CVR Prediction with Explicit Multi-Level RepresentationabstractEstimating conversion rate (CVR) accurately has been one of the most central problems in online advertising. Existing methods in production focus on learning effective interactions among features to boost the model performance. Despite great success, these methods treat all the features equally without distinction. However, different features suffer differently from cold-start issues. Tail elements in those high-cardinality features, which we denote as fine-grained features, tend to have inadequate samples and thus fail to obtain semantically meaningful embeddings. Interacting with those features leads astray and impairs the accuracy of new ads in a cold-start scenario. In this paper, we propose Automatic Fusion Network (AutoFuse) to better tackle the challenge. AutoFuse explicitly separates features into groups based on their granularity and learns multiple levels of representation conditioned on different combinations of feature groups. Concretely, AutoFuse learns an ad-level representation to depict the unique individual character and a group-level representation to portray the collective information by discarding the fine-grained features. The final robust and general ad representation is obtained by integrating these two level representations adaptively. Such a combination encompasses a wider amount of information, and thereby mitigates the cold-start issue. Extensive experiments on two industrial-scale datasets and three public datasets show that AutoFuse significantly and consistently outperforms a spectrum of competitive methods including our currently deployed model. Meanwhile, the remarkable improvement on new ads validates the effectiveness of our method in cold-start scenarios. We design AutoFuse as a generic approach and thus it can be seamlessly transferred into other domains. Our method has been deployed online to serve billions of users and ads and has achieved significant GMV gain of 2.84%. Jipeng Jin, Guangben Lu, Xiaofeng Gao 0001, Ao Tan |
ICDE | 4 |
| 2023 | Meteorology-Assisted Spatio-Temporal Graph Network for Uncivilized Urban Event PredictionabstractUncivilized urban events disrupt urban order and have a detrimental impact on daily life. Recognizing the significant implications of these events, urban managers strive to proactively prevent them by accurately predicting their future occurrence. However, existing methods overlook crucial contextual information within urban scenarios while mining spatio-temporal dependencies in single event series. Fortunately, we discovered a connection between meteorological conditions and uncivilized events. To leverage this relationship, we propose a novel approach named the Meteorology-Assisted Spatio-Temporal Graph Neural Network (MAST) which integrates meteorological information into the spatio-temporal dependency modeling for predicting urban uncivilized events. Additionally, our approach captures latent regularities in human behavior by explicitly modeling individuals’ psychological states based on meteorological information. We also adopt cross-view contrastive learning between urban regions to dynamically capture the informative components of meteorological information for precise prediction of urban uncivilized events. Experimental evaluations on a real-world dataset demonstrate the superiority of MAST over state-of-theart baselines in terms of predictive performance. Yang Luo 0004, Zehao Gu, Yun Xiong, Xiaofeng Gao 0001 |
ICDM | 5 |
| 2023 | Interactive Activities Initiation through Retrieving Hidden Social Information NetworksabstractThe rise of social platforms based on online social networks has greatly enriched people’s lives, resulting in various applications. Traditional research mainly focuses on users but pays less attention to the edges between users, and they all assume the topology of social networks is known in advance. Indeed, obtaining the network topology is challenging because of privacy protection and business competition. In this paper, we propose an activity initiation problem inspired by real business applications, such as Pinduoduo and Tencent, where each edge can be abstracted as an activity in which both ends (users) of the edge participate together, and the edge can be initiated by one of them. At this time, we hope to select as few users as possible to initiate activities and make the users of the whole network participate together. This problem can be reduced to the classic vertex cover problem, but the network information is hidden by social platforms as much as possible. To address this challenge, we put forward a solver-detector model. In each round of interaction, the solver uses a detector to obtain a small amount of edge information and achieves vertex coverage. This is a model in which a solver has limited access to input, but still gives a 2-approximation that is as good as the conventional model. Finally, we can cover the whole network with very few edge samples, which is a brand-new research perspective. Yulong Song, Jianxiong Guo, Xiaofeng Gao 0001 |
ICDM | 4 |
| 2023 | SCRIPT: Sequential Cross-Meta-Information Recommendation in Pretrain and Prompt ParadigmabstractExisting online advertising systems employ separate models for each task and site, resulting in a large number of models that require significant computing power and human effort to train and deploy. Moreover, separate models have limitations in sharing cross-scenario information. To address these issues, we propose a unified sequential recommendation model called SCRIPT. It takes cross-scenario user behavior sequences as input and explicitly incorporates meta information that characterizes scenario features, such as domain, site, and behavior types. Inspired by the advances of the pretrain and prompt paradigm, we generate scenario-aware and personalized prompts based on the user profile and meta information of candidate items. This allows the model to leverage the knowledge learned during pre-training and adapt it to serve different downstream tasks. Extensive experiments on two public dataset and a production dataset demonstrate that our model achieves state-of-the-art performance on multiple downstream recommendation tasks. Xinyi Zhou 0006, Jipeng Jin, Li Ma 0012, Xiaofeng Gao 0001, Jianbo Yang, Xiongwen Yang, Lei Xiao 0001 |
ICDM | 4 |
| 2023 | IPOC: An Adaptive Interval Prediction Model based on Online Chasing and Conformal Inference for Large-Scale SystemsabstractIn large-scale systems, due to system complexity and demand volatility, diverse and dynamic workloads make accurate predictions difficult. In this work, we address an online interval prediction problem (OnPred-Int) and adopt ensemble learning to solve it. We depict that the ensemble learning for OnPred-Int is a dynamic deterministic Markov Decision Process (Dd-MDP) and convert it into a stateful online learning task. Then we propose IPOC, a lightweight and flexible model able to produce effective confidence intervals, adapting the dynamics of real-time workload streams. At each time, IPOC selects a target model and executes chasing for it by a designed chasing oracle, during which process IPOC produces accurate confidence intervals. The effectiveness of IPOCis theoretically validated through sublinear regret analysis and satisfaction of confidence interval requirements. Besides, we conduct extensive experiments on 4 real-world datasets comparing with 19 baselines. To the best of our knowledge, we are the first to apply the frontier theory of online learning to time series prediction tasks. Jiadong Chen, Yang Luo 0004, Xiuqi Huang, Fuxin Jiang, Yangguang Shi, Tieying Zhang, Xiaofeng Gao 0001 |
KDD | 7 |
| 2023 | RBNets: A Reinforcement Learning Approach for Learning Bayesian Network Structure
Zuowu Zheng, Xiaofeng Gao 0001, Guihai Chen |
ECML/PKDD (3) | 3 |
| 2023 | Cross-Platform Event Popularity Analysis via Dynamic Time Warping and Neural PredictionabstractNowadays, the primary media for information dissemination is shifting to online media. Events usually burst online through multiple modern online media. Therefore, predicting event popularity trends becomes crucial for online platforms to track pubic concerns and make appropriate decisions. However, few researches focus on events popularity prediction from a cross-platform perspective. Challenges origin from vast diversity from events and media, limited access to aligned datasets across different platforms and the great deal of noise in datasets. In this paper, we solve the cross-platform event popularity prediction problem by proposing a model named DancingLines, which is mainly composed of the following three parts. First, we propose TF-SW, a semantic-aware popularity quantification model based on Term Frequency with Semantic Weight, obtaining the event popularity based on Word2Vec and TextRank and generating Event Popularity Time Series(EPTS). Then, we propose DTW-CD, a pairwise time series alignment model derived from DTW with Compound Distance, aligning the EPTS on several platforms. Finally, we aggregate two time series and propose a neural based prediction model implementing Long Short-Term Memory with attention mechanism to obtain accurate predictions. Evaluation results based on large scale real-world datasets demonstrate that DancingLines can efficiently characterize, align, and predict event popularity on cross-platform. Xiaofeng Gao 0001, Guihai Chen |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2022 | MetisRL: A Reinforcement Learning Approach for Dynamic Routing in Data Center Networks
Yuanning Gao, Xiaofeng Gao 0001, Guihai Chen |
DASFAA (2) | 2 |
| 2022 | TEALED: A Multi-Step Workload Forecasting Approach Using Time-Sensitive EMD and Auto LSTM Encoder-Decoder
Xiuqi Huang, Yunlong Cheng, Xiaofeng Gao 0001, Guihai Chen |
DASFAA (2) | 3 |
| 2022 | TROP: Task Ranking Optimization Problem on Crowdsourcing Service Platform
Haozhen Lu, Xiaofeng Gao 0001, Ailun Song, Guihai Chen |
DASFAA (1) | 3 |
| 2022 | MDKE: Multi-level Disentangled Knowledge-Based Embedding for Recommender Systems
Haolin Zhou, Qingmin Liu, Xiaofeng Gao 0001, Guihai Chen |
DASFAA (2) | 3 |
| 2022 | CAKE: A Context-Aware Knowledge Embedding Model of Knowledge Graph
Jiadong Chen, Hua Ke, Haijian Mo, Xiaofeng Gao 0001, Guihai Chen |
DEXA (1) | 4 |
| 2022 | Intelligent Air Traffic Management System Based on Knowledge Graph
Jiadong Chen, Xiaofeng Gao 0001, Guihai Chen |
DEXA (2) | 3 |
| 2022 | Deep Active Learning Framework for Crowdsourcing-Enhanced Image Classification and Segmentation
Xiaofeng Gao 0001, Guihai Chen |
DEXA (1) | 2 |
| 2022 | KAPP: Knowledge-Aware Hierarchical Attention Network for Popularity Prediction
Shuodian Yu, Jianxiong Guo, Xiaofeng Gao 0001, Guihai Chen |
DEXA (1) | 3 |
| 2022 | AutoAttention: Automatic Field Pair Selection for Attention in User Behavior ModelingabstractIn Click-through rate (CTR) prediction models, a user’s interest is usually represented as a fixed-length vector based on her history behaviors. Recently, several methods are proposed to learn an attentive weight for each user behavior and conduct weighted sum pooling. However, these methods only manually select several fields from the target item side as the query to interact with the behaviors, neglecting the other target item fields, as well as user and context fields. Directly including all these fields in the attention may introduce noise and deteriorate the performance. In this paper, we propose a novel model named AutoAttention, which includes all item/user/context side fields as the query, and assigns a learnable weight for each field pair between behavior fields and query fields. Pruning on these field pairs via these learnable weights lead to automatic field pair selection, so as to identify and remove noisy field pairs. Though including more fields, the computation cost of AutoAttention is still low due to using a simple attention function and field pair selection. Extensive experiments on the public dataset and Tencent’s production dataset demonstrate the effectiveness of the proposed approach. Zuowu Zheng, Xiaofeng Gao 0001, Junwei Pan, Guihai Chen, Jie Jiang 0015 |
ICDM | 2 |
| 2022 | Multi-Objective Actor-Critics for Real-Time Bidding in Display Advertising
Haolin Zhou, Chaoqi Yang, Xiaofeng Gao 0001, Gongshen Liu, Guihai Chen |
ECML/PKDD (4) | 3 |
| 2022 | HIEN: Hierarchical Intention Embedding Network for Click-Through Rate PredictionabstractClick-through rate (CTR) prediction plays an important role in online advertising and recommendation systems, which aims at estimating the probability of a user clicking on a specific item. Feature interaction modeling and user interest modeling methods are two popular domains in CTR prediction, and they have been studied extensively in recent years. However, these methods still suffer from two limitations. First, traditional methods regard item attributes as ID features, while neglecting structure information and relation dependencies among attributes. Second, when mining user interests from user-item interactions, current models ignore user intents and item intents for different attributes, which lacks interpretability. Based on this observation, in this paper, we propose a novel approach Hierarchical Intention Embedding Network (HIEN), which considers dependencies of attributes based on bottom-up tree aggregation in the constructed attribute graph. HIEN also captures user intents for different item attributes as well as item intents based on our proposed hierarchical attention mechanism. Extensive experiments on both public and production datasets show that the proposed model significantly outperforms the state-of-the-art methods. In addition, HIEN can be applied as an input module to state-of-the-art CTR prediction methods, bringing further performance lift for these existing models that might already be intensively used in real systems. Zuowu Zheng, Changwang Zhang, Xiaofeng Gao 0001, Guihai Chen |
SIGIR | 3 |
| 2022 | Which is better? A modularized evaluation for topic popularity prediction
Jiacheng Luo, Xiaofeng Gao 0001, Guihai Chen |
Knowl. Inf. Syst. | 3 |
| 2022 | MAB-Based Reinforced Worker Selection Framework for Budgeted Spatial CrowdsensingabstractSpatial crowdsensing is a special kind of crowdsourcing which allocates tasks to workers in some special places where workers can sense data for them. Due to the lack of priori information about the quality of workers and the ground truth, selecting the most suitable workers, which can guarantee the quality of the sensing tasks, remains a great challenge. In this paper, we propose a novel framework which can choose the most reliable workers among available workers under constraint budget. We model the quality of workers through two factors, bias and variance, which describe the continuous feature of sensing tasks. Our framework first allocate some calibration tasks to calibrate the bias and then iteratively estimate the workers variance more and more accurately. To choose more reliable workers, we face the exploration and exploitation dilemma. Therefore, we design a novel Multi-Armed Bandit (MAB) algorithm which based on Upper Confidence Bounds (UCB) scheme and combined with a weighted data aggregation scheme to estimate a more accurate ground truth of a sensing task. Futhermore, a dynamic budget allocation algorithm is designed to achieve global optimization. Then, we prove the expected sensing error can be bounded according to the regret bound of the MAB. In simulation experiments, we compare our algorithm with several baselines with real-world data set and it shows the effectiveness in inferring the ground truth with limited budget. Xiaofeng Gao 0001, Shenwei Chen, Guihai Chen |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2022 | Using Survival Theory in Early Pattern Detection for Viral CascadesabstractIn recent years, social networks have developed rapidly and become an indispensable part of people’s everyday life. Many models try to predict whether some reshare cascades are going to be popular or not, but most of their performances are limited due to the lack of cascades’ information in the early stage. In this paper, we proposeEarly Pattern detection model for Outbreak Cascades(in abbreviation, EPOC) inspired by the survival theory. We use three features to predict cascades’ virality: retweet sequence, follower number sequence, and timestamps of the first tweet which includes both the static and dynamic characteristics of cascades. Utilizing the theory that distributions of both viral and non-viral cascades are Gaussian, we get the boundary between these two kinds of cascades with sufficient proof to testify its rationality. To detect the virality more precisely and earlier, based on hazard functions in the survival theory, we propose two different hazard ceilings to capture the bursting of the cascades. We also provide a series of numerical experiments to analyze impacts of different factors to performance of our model measured by three practical metrics. The results shows that our model could stably outperforms several state-of-art baselines. Xiaofeng Gao 0001, Xiaosong Jia, Chaoqi Yang, Guihai Chen |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2022 | Predicting Hot Events in the Early Period through Bayesian Model for Social NetworksabstractPredicting emerging hot events in an early stage is essential for various applications, including information dissemination mining, ads recommendation and etc. Existing techniques either require a long-term observation over the event or features that are expensive to extract. However, given limited data at the early stage of an emerging event, the temporal features of hot events and non-hot events are not distinctive enough yet. In this work, we introduce BEEP, a Bayesian perspective Early stage Event Prediction model, that tackles this dilemma. We formulate the hot event prediction problem by two Semi-Naive Bayes Classifiers, where we consider both the temporal features and structural features and perform distribution test for the selected features. Theoretical analysis and extensive empirical evaluations on two real datasets demonstrate the effectiveness of our methods. Zuowu Zheng, Xiaofeng Gao 0001, Xiao Ma 0006, Guihai Chen |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2021 | Seq2Bubbles: Region-Based Embedding Learning for User Behaviors in Sequential RecommendersabstractUser behavior sequences contain rich information about user interests and are exploited to predict user's future clicking in sequential recommendation. Existing approaches, especially recently proposed deep learning models, often embed a sequence of clicked items into a single vector, i.e., a point in vector space, which suffer from limited expressiveness for complex distributions of user interests with multi-modality and heterogeneous concentration. In this paper, we propose a new representation model, named as Seq2Bubbles, for sequential user behaviors via embedding an input sequence into a set of bubbles each of which is represented by a center vector and a radius vector in embedding space. The bubble embedding can effectively identify and accommodate multi-modal user interests and diverse concentration levels. Furthermore, we design an efficient scheme to compute distance between a target item and the bubble embedding of a user sequence to achieve next-item recommendation. We also develop a self-supervised contrastive loss based on our bubble embeddings as an effective regularization approach. Extensive experiments on four benchmark datasets demonstrate that our bubble embedding can consistently outperform state-of-the-art sequential recommendation models. Qitian Wu, Chenxiao Yang, Shuodian Yu, Xiaofeng Gao 0001, Guihai Chen |
CIKM | 4 |
| 2021 | An Attention-Based Bi-GRU for Route Planning and Order Dispatch of Bus-Booking Platform
Yucen Gao, Yuanning Gao, Xiaofeng Gao 0001, Xiang Li 0006, Guihai Chen |
DASFAA (1) | 4 |
| 2021 | Gated Sequential Recommendation System with Social and Textual Information Under Dynamic Contexts
Haoyu Geng, Shuodian Yu, Xiaofeng Gao 0001 |
DASFAA (3) | 3 |
| 2021 | SRecGAN: Pairwise Adversarial Training for Sequential Recommendation
Guangben Lu, Xiaofeng Gao 0001, Guihai Chen |
DASFAA (3) | 3 |
| 2021 | GCAN: A Group-Wise Collaborative Adversarial Networks for Item Recommendation
Xuehan Sun, Tianyao Shi, Xiaofeng Gao 0001, Xiang Li 0006, Guihai Chen |
DASFAA (3) | 3 |
| 2021 | A Reinforcement Learning Model for Influence Maximization in Social Networks
Xiaofeng Gao 0001, Guihai Chen |
DASFAA (2) | 3 |
| 2021 | FORM: Follow the Online Regularized Meta-Leader for Cold-Start RecommendationabstractMeta-learning based recommendation systems alleviate the cold-start problem through a bi-level meta-optimization process. Recommendation borrows prior experience from pre-trained static system-level parameters and fine-tunes the model in user-level for new users. However, it is more natural for the system to sample users in a dynamic online sequence in most real-world recommendation systems, which brings further challenges for existing meta-learning based recommendation: system-level updates begins before user-level recommendation models have converged on the whole time series; stable and randomness-resistant bi-level gradient descent approaches are missing in the current meta-learning framework; evaluation on learning abilities across different users are lacked for exploring the diversities of different users. Xuehan Sun, Tianyao Shi, Xiaofeng Gao 0001, Yanrong Kang, Guihai Chen |
SIGIR | 3 |
| 2021 | Trust Prediction for Online Social Networks with Integrated Time-Aware SimilarityabstractOnline social networks gain increasing popularity in recent years. In online social networks, trust prediction is significant for recommendations of high reputation users as well as in many other applications. In the literature, trust prediction problem can be solved by several strategies, such as matrix factorization, trust propagation, and -NN search. However, most of the existing works have not considered the possible complementarity among these mainstream strategies to optimize their effectiveness and efficiency. In this article, we propose a novel trust prediction approach named iSim : an integrated time-aware similarity-based collaborative filtering approach leveraging on user similarity, which integrates three kinds of factors to measure user similarity, including vector space similarity, time-aware matrix factorization, and propagated trust. This article is the first work in the literature employing time-aware matrix factorization and propagated trust in the study of similarity. Additionally, we use several methods like adding inverted index to reduce the time complexity of iSim , and provide its theoretical time bound. Moreover, we also provide the detailed overview and theoretical analysis of the existing works. Finally, the extensive experiments with real-world datasets show that iSim achieves great improvement for both efficiency and effectiveness over the state-of-the-art approaches. Xiaofeng Gao 0001, Mingding Liao, Guihai Chen |
ACM Trans. Knowl. Discov. Data | 1 |
| 2021 | Parallel Greedy Algorithm to Multiple Influence Maximization in Social NetworkabstractInfluence Maximization (IM) problem is to select influential users to maximize the influence spread, which plays an important role in many real-world applications such as product recommendation, epidemic control, and network monitoring. Nowadays multiple kinds of information can propagate in online social networks simultaneously, but current literature seldom discuss about this phenomenon. Accordingly, in this article, we propose Multiple Influence Maximization (MIM) problem where multiple information can propagate in a single network with different propagation probabilities. The goal of MIM problems is to maximize the overall accumulative influence spreads of different information with the limit of seed budget . To solve MIM problems, we first propose a greedy framework to solve MIM problems which maintains an -approximate ratio. We further propose parallel algorithms based on semaphores, an inter-thread communication mechanism, which significantly improves our algorithms efficiency. Then we conduct experiments for our framework using complex social network datasets with 12k, 154k, 317k, and 1.1m nodes, and the experimental results show that our greedy framework outperforms other heuristic algorithms greatly for large influence spread and parallelization of algorithms reduces running time observably with acceptable memory overhead. Guanhao Wu, Xiaofeng Gao 0001, Ge Yan 0001, Guihai Chen |
ACM Trans. Knowl. Discov. Data | 2 |
| 2021 | Flexible Aggregate Nearest Neighbor Queries and its Keyword-Aware Variant on Road NetworksabstractAggregate nearest neighbor (Ann) query in both the euclidean space and road networks has been extensively studied, and the flexible aggregate nearest neighbor (Fann) problem further generalizesAnnby introducing an extra flexibility parameter$\phi$that ranges in$(0, 1]$. In this article, we focus onFannon road networks, denoted asFann$_\mathcal {R}$, and its keyword-aware variant, denoted asKFann$_\mathcal {R}$. To solve these problems, we propose a series of universal (i.e., suitable for bothmaxandsum) algorithms, including a Dijkstra-based algorithm that enumerates$P$instead of$\phi |Q|$-combinations of$Q$, a queue-based approach that processes data points from-near-to-far, and a framework that combinesincremental euclidean restriction(IER) and$k$NN. We also propose a specific exact solution tomax-Fann$_\mathcal {R}$and a constant-factor ratio approximate solution tosum-Fann$_\mathcal {R}$. These specific algorithms are easy to implement and can achieve excellent performance in some scenarios. Besides, we further extend this problem to top-$k$and multipleFann$_\mathcal {R}$(resp.,KFann$_\mathcal {R}$) queries. We conduct a comprehensive experimental evaluation for the proposed algorithms on real datasets to demonstrate their superior efficiency and high quality. Zhongpu Chen, Bin Yao 0002, Zhi-Jie Wang 0009, Xiaofeng Gao 0001, Shuo Shang, Shuai Ma 0001, Minyi Guo |
IEEE Trans. Knowl. Data Eng. | 4 |
| 2021 | Quality Inference Based Task Assignment in Mobile CrowdsensingabstractWith the increase of mobile devices, Mobile Crowdsensing (MCS) has become an efficient way to ubiquitously sense and collect environment data. Comparing to traditional sensor networks, MCS has a vital advantage that workers play an active role in collecting and sensing data. However, due to the openness of MCS, workers and sensors are of different qualities. Low quality sensors and workers may yield noisy data or even inaccurate data. Which gives the importance of inferring the quality of workers and sensors and seeking a valid task assignment with enough total qualities for MCS. To solve the problem, we adopt truth inference methods to iteratively infer the truth and qualities. Based on the quality inference, this paper proposes a task assignment problem called quality-bounded task assignment with redundancy constraint (QTAR). Different from traditional task assignment problem, redundancy constraint is added to satisfy the preliminaries of truth inference, which requires that each task should be assigned a certain or more amount of workers. We prove that QTAR is NP-complete and propose a (2+ε) - approximation algorithm for QTAR, called QTA. Finally, experiments are conducted on both synthesis data and real dataset. The results of the experiments prove the efficiency and effectiveness of our algorithms. Xiaofeng Gao 0001, Haowei Huang, Chenlin Liu, Fan Wu 0006, Guihai Chen |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2021 | Popularity Prediction for Single Tweet Based on Heterogeneous Bass ModelabstractPredicting the popularity of a single tweet is useful for both users and enterprises. However, adopting existing topic or event prediction models cannot obtain satisfactory results. The reason is that one topic or event that consists of multiple tweets, has more features and characteristics than a single tweet. In this article, we propose two variations of Heterogeneous Bass models (HBass), originally developed in the field of marketing science, namely Spatial-Temporal Heterogeneous Bass Model (ST-HBass) and Feature-Driven Heterogeneous Bass Model (FD-HBass), to predict the popularity of a single tweet at the early stage and the stable stage. We further design an Interaction Enhancement to improve the performance, which considers the competition and cooperation from different tweets with the common topic. In addition, it is often difficult to depict popularity quantitatively. We design an experiment to get the weight of favorite, retweet and reply, and apply the linear regression to calculate the popularity. Furthermore, we design a clustering method to bound the popular threshold. Once the weight and popular threshold are determined, the status whether a tweet will be popular or not can be justified. Our model is validated by conducting experiments on real-world Twitter data, and the results show the efficiency and accuracy of our model, with less absolute percent error and the best Precision and F-score. In all, we introduce Bass model into social network single-tweet prediction to show it can achieve excellent performance. Xiaofeng Gao 0001, Zuowu Zheng, Quanquan Chu, Shaojie Tang 0001, Guihai Chen, Qianni Deng |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2021 | ERATO: Trading Noisy Aggregate Statistics over Private Correlated DataabstractWith the commoditization of personal privacy, pricing private data has become an intriguing problem. In this paper, we study noisy aggregate statistics trading from the perspective of a data broker in data markets. We thus propose ERATO, which enables aggrEgate statistics pRicing over privATe cOrrelated data. On one hand, ERATO guarantees arbitrage freeness against cunning data consumers. On the other hand, ERATO compensates data owners for their privacy losses using both bottom-up and top-down designs. We further apply ERATO to three practical aggregate statistics, namely weighted sum, probability distribution fitting, and degree distribution, and extensively evaluate their performances on MovieLens dataset, 2009 RECS dataset, and two SNAP large social network datasets, respectively. Our analysis and evaluation results reveal that ERATO well balances utility and privacy, achieves arbitrage freeness, and compensates data owners more fairly than differential privacy based approaches. Chaoyue Niu, Zhenzhe Zheng 0001, Fan Wu 0006, Shaojie Tang 0001, Xiaofeng Gao 0001, Guihai Chen |
IEEE Trans. Knowl. Data Eng. | 5 |
| 2020 | SentiMem: Attentive Memory Networks for Sentiment Classification in User Review
Xiaosong Jia, Qitian Wu, Xiaofeng Gao 0001, Guihai Chen |
DASFAA (1) | 3 |
| 2020 | KPML: A Novel Probabilistic Perspective Kernel Mahalanobis Distance Metric Learning Model for Semi-supervised Clustering
Yongyi Hu, Xiaofeng Gao 0001, Guihai Chen |
DEXA (2) | 3 |
| 2020 | FB2vec: A Novel Representation Learning Model for Forwarding Behaviors on Online Social Networks
Li Ma 0012, Mingding Liao, Xiaofeng Gao 0001, Guoze Zhang, Guihai Chen |
ECML/PKDD (1) | 3 |
| 2020 | Skia: Scalable and Efficient In-Memory Analytics for Big Spatial-Textual DataabstractIn recent years, spatial-keyword queries have attracted much attention with the fast development of location-based services. However, current spatial-keyword techniques are disk-based, which cannot fulfill the requirements of high throughput and low response time. With the surging data size, people tend to process data in distributed in-memory environments to achieve low latency. In this paper, we present the distributed solution, i.e., Skia (Spatial-Keyword In-memory Analytics), to provide a scalable backend for spatial-textual analytics. Skia introduces a two-level index framework for big spatial-textual data including: (1) efficient and scalable global index, which prunes the candidate partitions a lot while achieving small space budget; and (2) four novel local indexes, that further support low latency services for exact and approximate spatial-keyword queries. Skia can support common spatial-keyword queries via traditional SQL programming interfaces. The experiments conducted on large-scale real datasets have demonstrated the promising performance of the proposed indexes and our distributed solution. Yang Xu 0031, Bin Yao 0002, Zhi-Jie Wang 0009, Xiaofeng Gao 0001, Jiong Xie, Minyi Guo |
IEEE Trans. Knowl. Data Eng. | 4 |
| 2019 | Reinforcement Learning with Sequential Information Clustering in Real-Time BiddingabstractDisplay advertising is a billion dollar business which is the primary income of many companies. In this scenario, real-time bidding optimization is one of the most important problems, where the bids of ads for each impression are determined by an intelligent policy such that some global key performance indicators are optimized. Due to the highly dynamic bidding environment, many recent works try to use reinforcement learning algorithms to train the bidding agents. However, as the probability of the occurrence of a particular state is typically low and the state representation in current work lacks sequential information, the convergence speed and performance of deep reinforcement algorithms are disappointing. To tackle these two challenges in the real-time bidding scenario, we propose ClusterA3C, a novel Advantage Asynchronous Actor-Critic (A3C) variant integrated with a sequential information extraction scheme and a clustering based state aggregation scheme. We conduct extensive experiments to validate the proposed scheme on a real-world commercial dataset. Experimental results show that the proposed scheme outperforms the state of the art methods in terms of either performance or convergence speed. Chaoqi Yang, Xiaofeng Gao 0001, Liubin Wang, Guihai Chen |
CIKM | 3 |
| 2019 | Reinforced Reliable Worker Selection for Spatial Crowdsensing Networks
Yang Wang 0019, Jingxiao Chen, Xiaofeng Gao 0001, Guihai Chen |
DASFAA (1) | 4 |
| 2019 | Real-Time Route Planning and Online Order Dispatch for Bus-Booking Platforms
Hao Zhou 0016, Yucen Gao, Xiaofeng Gao 0001, Guihai Chen |
DASFAA (2) | 3 |
| 2019 | SENTI2POP: Sentiment-Aware Topic Popularity Prediction on Social MediaabstractTopic popularity prediction is an important task on social media, which aims at predicting the ongoing trends of topics according to logged historical text-based records. However, only limited existing approaches apply sentiment analysis to facilitate popularity prediction. Public sentiment is worth taking into consideration because the topics with strong sentiment tend to spread faster and broader on social media. In this paper, we propose a novel framework, SENTI2POP, to predict topic popularity utilizing sentiment information. We first adapt a state-of-art popularity quantification method to capture the topic popularity, and then design a novel tree-like network (Tree-Net) combining Long Short-Term Memory (LSTM) and Convolutional Neural Network (CNN) for sentiment analysis. In addition, we propose a sentiment-aware time series prediction approach based on Dynamic Time Warping (DTW) and Autoregressive Integrated Moving Average model (ARIMA) to predict topic popularity. We prove by experiments that SENTI2POP outperforms the existing popularity prediction models on a real-world Twitter dataset by reducing the prediction error. Experimental results also show that SENTI2POP could be applied to improve the accuracy of most non-sentiment popularity prediction models. Jinning Li 0001, Yirui Gao, Xiaofeng Gao 0001, Yan Shi 0009, Guihai Chen |
ICDM | 3 |
| 2019 | Dual Sequential Prediction Models Linking Sequential Recommendation and Information DisseminationabstractSequential recommendation and information dissemination are two traditional problems for sequential information retrieval. The common goal of the two problems is to predict future user-item interactions based on past observed interactions. The difference is that the former deals with users' histories of clicked items, while the latter focuses on items' histories of infected users.In this paper, we take a fresh view and propose dual sequential prediction models that unify these two thinking paradigms. One user-centered model takes a user's historical sequence of interactions as input, captures the user's dynamic states, and approximates the conditional probability of the next interaction for a given item based on the user's past clicking logs. By contrast, one item-centered model leverages an item's history, captures the item's dynamic states, and approximates the conditional probability of the next interaction for a given user based on the item's past infection records. To take advantage of the dual information, we design a new training mechanism which lets the two models play a game with each other and use the predicted score from the opponent to design a feedback signal to guide the training. We show that the dual models can better distinguish false negative samples and true negative samples compared with single sequential recommendation or information dissemination models. Experiments on four real-world datasets demonstrate the superiority of proposed model over some strong baselines as well as the effectiveness of dual training mechanism between two models. Qitian Wu, Yirui Gao, Xiaofeng Gao 0001, Paul Weng, Guihai Chen |
KDD | 3 |
| 2019 | Dual Graph Attention Networks for Deep Latent Representation of Multifaceted Social Effects in Recommender SystemsabstractSocial recommendation leverages social information to solve data sparsity and cold-start problems in traditional collaborative filtering methods. However, most existing models assume that social effects from friend users are static and under the forms of constant weights or fixed constraints. To relax this strong assumption, in this paper, we propose dual graph attention networks to collaboratively learn representations for two-fold social effects, where one is modeled by a user-specific attention weight and the other is modeled by a dynamic and context-aware attention weight. We also extend the social effects in user domain to item domain, so that information from related items can be leveraged to further alleviate the data sparsity problem. Furthermore, considering that different social effects in two domains could interact with each other and jointly influence users' preferences for items, we propose a new policy-based fusion strategy based on contextual multi-armed bandit to weigh interactions of various social effects. Experiments on one benchmark dataset and a commercial dataset verify the efficacy of the key components in our model. The results show that our model achieves great improvement for recommendation accuracy compared with other state-of-the-art social recommendation methods. Qitian Wu, Xiaofeng Gao 0001, Paul Weng, Guihai Chen |
WWW | 3 |
| 2019 | An efficient and scalable multi-dimensional indexing scheme for modular data centers
Yuanning Gao, Xiaofeng Gao 0001, Yichen Zhu 0002, Guihai Chen |
Data Knowl. Eng. | 2 |
| 2019 | Taxonomy and Evaluation for Microblog Popularity PredictionabstractAs social networks become a major source of information, predicting the outcome of information diffusion has appeared intriguing to both researchers and practitioners. By organizing and categorizing the joint efforts of numerous studies on popularity prediction, this article presents a hierarchical taxonomy and helps to establish a systematic overview of popularity prediction methods for microblog. Specifically, we uncover three lines of thoughts: the feature-based approach, time-series modelling, and the collaborative filtering approach and analyse them, respectively. Furthermore, we also categorize prediction methods based on their underlying rationale: whether they attempt to model the motivation of users or monitor the early responses. Finally, we put these prediction methods to test by performing experiments on real-life data collected from popular social networks Twitter and Weibo. We compare the methods in terms of accuracy, efficiency, timeliness, robustness, and bias. As far as we are concerned, there is no precedented survey aimed at microblog popularity prediction at the time of submission. By establishing a taxonomy and evaluation for the first time, we hope to provide an in-depth review of state-of-the-art prediction methods and point out directions for further research. Our evaluations show that time-series modelling has the advantage of high accuracy and the ability to improve over time. The feature-based methods using only temporal features performs nearly as well as using all possible features, producing average results. This suggests that temporal features do have strong predictive power and that power is better exploited with time-series models. On the other hand, this implies that we know little about the future popularity of an item before it is posted, which may be the focus of further research. Xiaofeng Gao 0001, Zhenhao Cao, Bin Yao 0002, Guihai Chen, Shaojie Tang 0001 |
ACM Trans. Knowl. Discov. Data | 1 |
| 2019 | Achieving Data Truthfulness and Privacy Preservation in Data MarketsabstractAs a significant business paradigm, many online information platforms have emerged to satisfy society's needs for person-specific data, where a service provider collects raw data from data contributors, and then offers value-added data services to data consumers. However, in the data trading layer, the data consumers face a pressing problem, i.e., how to verify whether the service provider has truthfully collected and processed data? Furthermore, the data contributors are usually unwilling to reveal their sensitive personal data and real identities to the data consumers. In this paper, we propose TPDM, which efficiently integrates Truthfulness and Privacy preservation in Data Markets. TPDM is structured internally in an Encrypt-then-Sign fashion, using partially homomorphic encryption and identity-based signature. It simultaneously facilitates batch verification, data processing, and outcome verification, while maintaining identity preservation and data confidentiality. We also instantiate TPDM with a profile matching service and a data distribution service, and extensively evaluate their performances on Yahoo! Music ratings dataset and 2009 RECS dataset, respectively. Our analysis and evaluation results reveal that TPDM achieves several desirable properties, while incurring low computation and communication overheads when supporting large-scale data markets. Chaoyue Niu, Zhenzhe Zheng 0001, Fan Wu 0006, Xiaofeng Gao 0001, Guihai Chen |
IEEE Trans. Knowl. Data Eng. | 4 |
| 2018 | Adaptive General Event Popularity Analysis on Streaming DataabstractIn today's big data era, breaking events are disseminating on various platforms like social media and search engine at a high velocity, generating large volumes of event-related data. Numerous studies focus on the event dissemination trend analysis on a particular media platform, while very few works are conducted to adaptively identify hot words and analyze an event's popularity in a platform independent manner from streaming data. In this work, we propose GAPES(General Analysis and Prediction of Event popularity on Streaming data), a novel and unified approach to analyze and predict the event popularity on different media platforms. GAPES first processes event streaming data and adaptively selects event hot words by considering word's importance, correlation and completeness aspects, then it uses an unified event popularity analysis model and an probability distribution based event popularity prediction model to accurately analyze and forecast event trend with a flexible timescale. Experimental results based on 16 real-world event datasets produce several interesting findings and validate the capability of GAPES in adaptively selecting event's representative words, effectively analyzing and accurately predicting event popularity on both social network and search engine platforms. Qidong Zhang, Zequan Guo, Xiaofeng Gao 0001 |
BDCAT | 6 |
| 2018 | Adversarial Training Model Unifying Feature Driven and Point Process Perspectives for Event Popularity PredictionabstractThis paper targets a general popularity prediction problem for event sequence, which has recently gained great attention due to its extensive applications in various domains. Feature driven method and point process method are two basic thinking paradigms to tackle the prediction problem, but both of them suffer from limitations. In this paper, we propose PreNets unifying the two thinking paradigms in an adversarial manner. On one side, feature driven model acts like a 'critic' who aims to discriminate the predicted popularity from the real one based on a set of temporal features from the sequence. On the other side, point process model acts like an 'interpreter' who recognizes the dynamic patterns in sequence to generate a predicted popularity that can fool the 'critic'. Through a Wasserstein learning based two-player game, the training loss of the 'critic' guides the 'interpreter' to better exploit the sequence patterns and enhance prediction, while the 'interpreter' pushes the 'critic' to select effective early features that helps discrimination. This mechanism enables the framework to absorb the advantages of both feature driven and point process methods. Empirical results show that PreNets achieves significant MAPE improvement for both Twitter cascade and Amazon review prediction. Qitian Wu, Chaoqi Yang, Xiaofeng Gao 0001, Paul Weng, Guihai Chen |
CIKM | 4 |
| 2018 | DancingLines: An Analytical Scheme to Depict Cross-Platform Event Popularity
Tianxiang Gao, Weiming Bao, Jinning Li 0001, Xiaofeng Gao 0001, Boyuan Kong, Guihai Chen |
DEXA (1) | 4 |
| 2018 | CROP: An Efficient Cross-Platform Event Popularity Prediction Model for Online Media
Mingding Liao, Xiaofeng Gao 0001, Xuezheng Peng, Guihai Chen |
DEXA (2) | 2 |
| 2018 | R^2 -Tree: An Efficient Indexing Scheme for Server-Centric Data Center Networks
Yin Lin, Xinyi Chen 0004, Xiaofeng Gao 0001, Bin Yao 0002, Guihai Chen |
DEXA (1) | 3 |
| 2018 | EPOC: A Survival Perspective Early Pattern Detection Model for Outbreak Cascades
Chaoqi Yang, Qitian Wu, Xiaofeng Gao 0001, Guihai Chen |
DEXA (1) | 3 |
| 2018 | QDR-Tree: An Efficient Index Scheme for Complex Spatial Keyword Query
Xinshi Zang, Peiwen Hao, Xiaofeng Gao 0001, Bin Yao 0002, Guihai Chen |
DEXA (1) | 3 |
| 2018 | Flexible Aggregate Nearest Neighbor Queries in Road NetworksabstractAggregate nearest neighbor (ANN) query has been studied in both the Euclidean space and road networks. The flexible aggregate nearest neighbor (FANN) problem further generalizes ANN by introducing an extra flexibility. Given a set of data points P, a set of query points Q, and a user-defined flexibility parameter φ that ranges in (0, 1], an FANN query returns the best candidate from P, which minimizes the aggregate (usually max or sum) distance to any φ |Q| objects in Q. In this paper, we focus on the problem in road networks (denoted as FANNR), and present a series of universal (i.e., suitable for both max and sum) algorithms to answer FANNRqueries in road networks, including a Dijkstra-based algorithm enumerating P, a queue-based approach that processes data points from-near-to-far, and a framework that combines Incremental Euclidean Restriction (IER) and kNN. We also propose a specific exact solution to max-FANNRand a specific approximate solution to sum-FANNRwhich can return a near-optimal result with a guaranteed constant-factor approximation. These specific algorithms are easy to implement and can achieve excellent performance in some scenarios. Besides, we further extend the FANNRto k-FANNR, and successfully adapt most of the proposed algorithms to answer k-FANNRqueries. We conduct a comprehensive experimental evaluation for the proposed algorithms on real road networks to demonstrate their superior efficiency and high quality. Bin Yao 0002, Zhongpu Chen, Xiaofeng Gao 0001, Shuo Shang, Shuai Ma 0001, Minyi Guo |
ICDE | 3 |
| 2018 | EPAB: Early Pattern Aware Bayesian Model for Social Content Popularity PredictionabstractThe boom of information technology enables social platforms (like Twitter) to disseminate social content (like news) in an unprecedented rate, which makes early-stage prediction for social content popularity of great practical significance. However, most existing studies assume a long-term observation before prediction and suffer from limited precision for early-stage prediction due to insufficient observation. In this paper, we take a fresh perspective, and propose a novel early pattern aware Bayesian model. The early pattern representation, which stands for early time series normalized on future popularity, can address what we call early-stage indistinctiveness challenge. Then we use an expressive evolving function to fit the time series and estimate three interpretable coefficients characterizing temporal effect of observed series on future evolution. Furthermore, Bayesian network is leveraged to model the probabilistic relations among features, early indicators and early patterns. Experiments on three real-world social platforms (Twitter, Weibo and WeChat) show that under different evaluation metrics, our model outperforms other methods in early-stage prediction and possesses low sensitivity to observation time. Qitian Wu, Chaoqi Yang, Xiaofeng Gao 0001, Guihai Chen |
ICDM | 3 |
| 2018 | Unlocking the Value of Privacy: Trading Aggregate Statistics over Private Correlated DataabstractWith the commoditization of personal privacy, pricing private data has become an intriguing problem. In this paper, we study noisy aggregate statistics trading from the perspective of a data broker in data markets. We thus propose ERATO, which enables aggrEgate statistics pRicing over privATe cOrrelated data. On one hand, ERATO guarantees arbitrage freeness against cunning data consumers. On the other hand, ERATO compensates data owners for their privacy losses using both bottom-up and top-down designs. We further apply ERATO to three practical aggregate statistics, namely weighted sum, probability distribution fitting, and degree distribution, and extensively evaluate their performances on MovieLens dataset, 2009 RECS dataset, and two SNAP large social network datasets, respectively. Our analysis and evaluation results reveal that ERATO well balances utility and privacy, achieves arbitrage freeness, and compensates data owners more fairly than differential privacy based approaches. Chaoyue Niu, Zhenzhe Zheng 0001, Fan Wu 0006, Shaojie Tang 0001, Xiaofeng Gao 0001, Guihai Chen |
KDD | 5 |
| 2018 | Top-kCritical Vertices Query on Shortest PathabstractShortest path query is one of the most fundamental and classic problems in graph analytics, which returns the complete shortest path between any two vertices. However, in many real-life scenarios, only critical vertices on the shortest path are desirable and it is unnecessary to search for the complete path. This paper investigates the shortest path sketch by defining a top- $k$ critical vertices ( $k$ CV) query on the shortest path. Given a source vertex $s$ and target vertex $t$ in a graph, $k$ CV query can return the top- $k$ significant vertices on the shortest path $SP(s,t)$ . The significance of the vertices can be predefined. The key strategy for seeking the sketch is to apply off-line preprocessed distance oracle to accelerate on-line real-time queries. This allows us to omit unnecessary vertices and obtain the most representative sketch of the shortest path directly. We further explore a series of methods and optimizations to answer $k$ CV query on both centralized and distributed platforms, using exact and approximate approaches, respectively. We evaluate our methods in terms of time, space complexity and approximation quality. Experiments on large-scale real-world networks validate that our algorithms are of high efficiency and accuracy. Jing Ma 0002, Bin Yao 0002, Xiaofeng Gao 0001, Yanyan Shen, Minyi Guo |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2017 | FM-Hawkes: A Hawkes Process Based Approach for Modeling Online Activity CorrelationsabstractUnderstanding and predicting user behavior on online platforms has proved to be of significant value, with applications spanning from targeted advertising, political campaigning, anomaly detection to user self-monitoring. With the growing functionality and flexibility of online platforms, users can now accomplish a variety of tasks online. This advancement has rendered many previous works that focus on modeling a single type of activity obsolete. In this work, we target this new problem by modeling the interplay between the time series of different types of activities and apply our model to predict future user behavior. Our model, FM-Hawkes, stands for Fourier-based kernel multi-dimensional Hawkes process. Specifically, we model the multiple activity time series as a multi-dimensional Hawkes process. The correlations between different types of activities are then captured by the influence factor. As for the temporal triggering kernel, we observe that the intensity function consists of numerous kernel functions with time shift. Thus, we employ a Fourier transformation based non-parametric estimation. Our model is not bound to any particular platform and explicitly interprets the causal relationship between actions. By applying our model to real-life datasets, we confirm that the mutual excitation effect between different activities prevails among users. Prediction results show our superiority over models that do not consider action types and flexible kernels Xiaofeng Gao 0001, Weiming Bao, Guihai Chen |
CIKM | 2 |
| 2017 | AngleCut: A Ring-Based Hashing Scheme for Distributed Metadata Management
Renxuan Wang, Xiaofeng Gao 0001, Xiaochun Yang 0001, Guihai Chen |
DASFAA (1) | 3 |
| 2017 | Trading Data in Good Faith: Integrating Truthfulness and Privacy Preservation in Data MarketsabstractAs a significant business paradigm, many online information platforms have emerged to satisfy society's needs for person-specific data, where a service provider collects raw data from data contributors, and then offers value-added data services to data consumers. However, in the data trading layer, the data consumers face a pressing problem, i.e., how to verify whether the service provider has truthfully collected and processed data? Furthermore, the data contributors are usually unwilling to reveal their sensitive personal data and real identities to the data consumers. In this paper, we propose TPDM, which efficiently integrates Truthfulness and Privacy preservation in Data Markets. TPDM is structured internally in an Encrypt-then-Sign fashion, using somewhat homomorphic encryption and identitybased signature. It simultaneously facilitates batch verification, data processing, and outcome verification, while maintaining identity preservation and data confidentiality. We also instantiate TPDM with a profile-matching service, and extensively evaluate its performance on Yahoo! Music ratings dataset. Our evaluation results show that TPDM achieves several desirable properties, while incurring low computation and communication overheads when supporting a large-scale data market. Chaoyue Niu, Zhenzhe Zheng 0001, Fan Wu 0006, Xiaofeng Gao 0001, Guihai Chen |
ICDE | 4 |
| 2017 | BEEP: A Bayesian Perspective Early Stage Event Prediction Model for Online Social NetworksabstractIn recent years, predicting future hot events in online social networks is becoming increasingly meaningful in marketing, advertisement, and recommendation systems to support companies' strategy making. Currently, most prediction models require long-term observations over the event or depend a lot on other features which are expensive to extract. However, at the early stage of an event, the temporal features of hot events and non-hot events are not distinctive yet. Besides, given the small amount of available data, high noise and complex network structure, those state-of-art models are unable to give an accurate prediction at the very early stage of an event. Hence, we propose two Bayesian perspective models to handle this dilemma. We first mathematically define the hot event prediction problem and introduce the general early stage event prediction framework, then model the five selected features into several continuous distributions, and present two Semi-Naive Bayes Classifier based prediction models, BEEP and SimBEEP, which is the simplified version of BEEP. Extensive experiments on real dataset have demonstrated that our model significantly outperforms the baseline methods. Xiao Ma 0006, Xiaofeng Gao 0001, Guihai Chen |
ICDM | 2 |
| 2016 | STH-Bass: A Spatial-Temporal Heterogeneous Bass Model to Predict Single-Tweet Popularity
Zhaowei Tan, Xiaofeng Gao 0001, Shaojie Tang 0001, Guihai Chen |
DASFAA (2) | 3 |
| 2016 | FR-Index: A Multi-dimensional Indexing Framework for Switch-Centric Data Centers
Yatao Zhang, Jialiang Cao, Xiaofeng Gao 0001, Guihai Chen |
DEXA (2) | 3 |
| 2016 | ESAP: A Novel Approach for Cross-Platform Event Dissemination Trend Analysis Between Social Network and Search Engine
Pengju Ma, Boyuan Kong, Wenqian Ji, Xiaofeng Gao 0001, Xuezheng Peng |
WISE (1) | 5 |
| 2016 | SMe: explicit & implicit constrained-space probabilistic threshold range queries for moving objects
Zhi-Jie Wang 0009, Bin Yao 0002, Reynold Cheng, Xiaofeng Gao 0001, Lei Zou 0001, Haibing Guan, Minyi Guo |
GeoInformatica | 4 |
| 2016 | Efficient R-Tree Based Indexing Scheme for Server-Centric Cloud Storage SystemabstractCloud storage system poses new challenges to the community to support efficient concurrent querying tasks for various data-intensive applications, where indices always hold important positions. In this paper, we explore a practical method to construct a two-layer indexing scheme for multi-dimensional data in diverse server-centric cloud storage system. We first propose RT-HCN, an indexing scheme integrating R-tree based indexing structure and HCN-based routing protocol. RT-HCN organizes storage and compute nodes into an HCN overlay, one of the newly proposed sever-centric data center topologies. Based on the properties of HCN, we design a specific index mapping technique to maintain layered global indices and corresponding query processing algorithms to support efficient query tasks. Then, we expand the idea of RT-HCN onto another server-centric data center topology DCell, discovering a potential generalized and feasible way of deploying two-layer indexing schemes on other server-centric networks. Furthermore, we prove theoretically that RT-HCN is both space-efficient and query-efficient, by which each node actually maintains a tolerable number of global indices while high concurrent queries can be processed within accepted overhead. We finally conduct targeted experiments on Amazon's EC2 platforms, comparing our design with RT-CAN, a similar indexing scheme for traditional P2P network. The results validate the query efficiency, especially the speedup of point query of RT-HCN, depicting its potential applicability in future data centers. Qiwei Tang, Xiaofeng Gao 0001, Bin Yao 0002, Guihai Chen, Shaojie Tang 0001 |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2015 | Indexing Multi-dimensional Data in Modular Data Centers
Libo Gao, Yatao Zhang, Xiaofeng Gao 0001, Guihai Chen |
DEXA (2) | 3 |
| 2015 | A Universal Distributed Indexing Scheme for Data Centers with Tree-Like Topologies
Yuang Liu, Xiaofeng Gao 0001, Guihai Chen |
DEXA (1) | 2 |
| 2014 | Efficient R-Tree Based Indexing for Cloud Storage System with Dual-Port Servers
Wanchao Liang, Xiaofeng Gao 0001, Bin Yao 0002, Guihai Chen |
DEXA (2) | 3 |
| 2014 | Evaluation and comparison of various indexing schemes in single-channel broadcast communication environment
Jiaofei Zhong, Weili Wu 0001, Xiaofeng Gao 0001, Yan Shi 0009 |
Knowl. Inf. Syst. | 3 |
| 2013 | Distributed AH-Tree Based Index Technology for Multi-channel Wireless Data Broadcast
Yongtian Yang, Xiaofeng Gao 0001, Jiaofei Zhong, Guihai Chen |
DASFAA (1) | 2 |
| 2011 | A Novel Hash-Based Streaming Scheme for Energy Efficient Full-Text Search in Wireless Data Broadcast
Yan Shi 0009, Weili Wu 0001, Xiaofeng Gao 0001, Jiaofei Zhong |
DASFAA (1) | 4 |
| 2011 | Energy-Efficient Tree-Based Indexing Schemes for Information Retrieval in Wireless Data Broadcast
Jiaofei Zhong, Weili Wu 0001, Yan Shi 0009, Xiaofeng Gao 0001 |
DASFAA (2) | 4 |
| 2010 | Efficient Parallel Data Retrieval Protocols with MIMO Antennae for Data Broadcast in 4G Wireless Communications
Yan Shi 0009, Xiaofeng Gao 0001, Jiaofei Zhong, Weili Wu 0001 |
DEXA (2) | 2 |
| 2009 | Three Approximation Algorithms for Energy-Efficient Query Dissemination in Sensor Database System
Zhao Zhang 0002, Xiaofeng Gao 0001, Weili Wu 0001, Hui Xiong 0001 |
DEXA | 2 |
| 2009 | Multi-focal learning and its application to customer service supportabstractIn this study, we formalize a multi-focal learning problem, where training data are partitioned into several different focal groups and the prediction model will be learned within each focal group. The multi-focal learning problem is motivated by numerous real-world learning applications. For instance, for the same type of problems encountered in a customer service center, the problem descriptions from different customers can be quite different. The experienced customers usually give more precise and focused descriptions about the problem. In contrast, the inexperienced customers usually provide more diverse descriptions. In this case, the examples from the same class in the training data can be naturally in different focal groups. As a result, it is necessary to identify those natural focal groups and exploit them for learning at different focuses. The key developmental challenge is how to identify those focal groups in the training data. As a case study, we exploit multi-focal learning for profiling problems in customer service centers. The results show that multifocal learning can significantly boost the learning accuracies of existing learning algorithms, such as Support Vector Machines (SVMs), for classifying customer problems. Yong Ge 0001, Hui Xiong 0001, Wenjun Zhou 0001, Ramendra K. Sahoo, Xiaofeng Gao 0001, Weili Wu 0001 |
KDD | 5 |