EDBT 2026 Demo / reviewers in the wild / expert
Weixiong Rao
dblp:79/5179
· DBLP profile ↗
44ranked-venue papers in the field
9as first author
16since 2021 · last 2026
0000-0001-6644-0349ORCID · verified
Domains — venue-derived; a paper can count in several
Database Systems & Data Management · 21 (6 first)Information Retrieval & Web Search · 14 (2 first)Data Mining & Knowledge Discovery · 7 (1 first)Other / Interdisciplinary · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Not All Imputations are Trustworthy: An Uncertainty-aware Multi-modal Entity Alignment FrameworkabstractMulti-modal Entity Alignment (MMEA) aims to identify equivalent entities across diverse knowledge graphs by leveraging structural, attribute, and visual information. However, real-world datasets frequently suffer from missing modalities, necessitating feature imputation. A critical yet underexplored issue is that not all imputed modalities are inherently trustworthy. Ignoring the aleatoric uncertainty of such modalities introduces severe noise which propagates through the fusion process and degrades alignment performance. To address this challenge, we propose a novel MMEA framework, namely SURE, to Suppress Uncertainty for tRustworthy Entity alignment. Specifically, SURE introduces an uncertainty-aware variational imputation module to estimate the aleatoric uncertainty of generated features. Crucially, rather than using these imputation features blindly, SURE leverages the estimated uncertainty to suppress noise propagation via a confidence-gated multi-modal fusion and an adaptive contrastive learning objective. Extensive experiments on DBP15K datasets demonstrate that SURE significantly outperforms state-of-the-art baselines, exhibiting exceptional robustness particularly in scenarios with high modality missing rates. Weijie Wang 0003, Shijie Luo 0001, Xinyuan Lu, Qinpei Zhao, Weixiong Rao |
SIGIR | 5 |
| 2026 | Combining Structural and Textual Knowledge for Knowledge Graph Link Prediction via Large Language ModelsabstractIn recent years, large language models (LLMs) have emerged as powerful tools for link prediction in knowledge graphs (KGs) due to their strong capabilities in understanding and generation. However, many LLM-based methods still heavily rely on textual descriptions of KGs, limiting their ability to capture structural information and to model complex relational patterns. Although some methods integrate structural embeddings into LLMs, their ability to harness the complementary strengths of both modalities and dynamically prioritize candidate entities based on query context remains limited. In this paper, we propose ST-KGLP, a novel framework that improves link prediction by aligning structural knowledge with textual knowledge and employing query-aware adaptive weighting for candidate selection. Specifically, our proposed ST-KGLP employs a knowledge aligner to bridge the information gap between structural and textual knowledge, and then utilizes a query-aware adaptive weighting strategy that dynamically computes attention weights between query representations and candidate entities, enabling contextually relevant candidate re-ranking for more accurate prediction. Extensive experiments on various datasets show that our ST-KGLP outperforms state-of-the-art approaches, achieving average improvements of 3.81%, 11.52%, 2.22%, and 1.55% across four evaluation metrics. Our code and datasets are available at https://github.com/shijielaw/ST-KGLP. Shijie Luo 0001, Xinyuan Lu, Qinpei Zhao, Weixiong Rao |
WSDM | 4 |
| 2025 | Real-Time Femoral Von Mises Stress Distribution Prediction via Graph Neural Networks
Jiasheng Shi, Chenwei Wu 0008, Qinpei Zhao, Wenxin Niu, Weixiong Rao, Shitan Wang, Shi Zhan 0001, Yanmei Jia |
ADMA (3) | 7 |
| 2025 | Bridging the Gap between Knowledge Graphs and LLMs for Multi-hop Question AnsweringabstractTo achieve multi-hop question answering over knowledge graphs (KGQA), many studies have explored converting retrieved subgraphs into textual form and feeding them into large language models (LLMs) to leverage their reasoning capabilities. However, due to the linear and discrete nature of text sequences, model performance may degrade when handling complex questions. To this end, we propose a novel structure-text knowledge synergistic method, BrikQA, which bridges the knowledge gap between knowledge graphs (KGs) and LLMs for multi-hop KGQA. LLMs and KGs complement each other by leveraging explicit topological patterns and implicit knowledge mining to enhance knowledge understanding and address sparsity issues. Experimental results on various datasets demonstrate that BrikQA outperforms state-of-the-art baselines. Our source code is available at https://github.com/shijielaw/BrikQA. Shijie Luo 0001, Xinyuan Lu, Qinpei Zhao, Weixiong Rao |
CIKM | 4 |
| 2025 | SimFormer: Multilevel Transformer on Learnable Mesh Graphs for Engineering SimulationabstractNumerical simulation is important in real-world engineering systems, such as solid mechanics and aero-dynamics. Hierarchical GNNs can learn engineering simulation with low simulation time and acceptable accuracy, but fail to represent complex interactions in simulation systems. In this paper, we propose a novel multilevel Transformer on learnable clusters, namely SimFormer. The key novelty of SimFormer is to interweave the learning of a learnable soft-cluster assignment algorithm and the inter-cluster/cluster-to-node attention. In form of a closed-loop, SimFormer learns the soft cluster assignment possibility by the feedback signals provided by the attention, and the attention can leverage the learnable clusters to better represent long-range interactions. In this way, the learnable clusters can adaptively match actual simulation results, and the multilevel attention modules can also effectively represent node embeddings. Experiments on four datasets demonstrate the superiority of SimFormer over seven baseline approaches. For example, on the real dataset, ours outperforms the recent work Eagle by 17.36% lower RMSE and 27.03% smaller FLOPs. The code and datasets are available at: https://github.com/pro-orp/SimFormer. Jiasheng Shi, Weixiong Rao, Ze Gao 0001 |
CIKM | 3 |
| 2025 | MMKG-RAG: Retrieval-Augmented Generation with Multi-modal Knowledge Graph
Shuaitao Zhao, Shijie Luo 0001, Xinyuan Lu, Weixiong Rao |
DASFAA (6) | 4 |
| 2025 | Towards Smarter and Safer Traffic Signal Control via Multiagent Deep Reinforcement LearningabstractRecently, deep reinforcement learning (DRL) has been employed for intelligent traffic‐light control and demonstrated promising results. However, state‐of‐the‐art DRL‐based systems still rely on discrete decision‐making, which can lead to unsafe driving practices. Additionally, existing feature representations of the environment often fail to capture the complex dynamics of traffic flows, resulting in imprecise predictions of traffic conditions. To address these issues, we propose a novel DRL framework based on the multiagent deep deterministic policy gradient algorithm. Our method offers several key innovations: it suggests employing a transitional phase before changing the current phase for safer traffic management, integrates local road network topology into feature representation to enhance the accuracy of traffic flow predictions, and uses two‐layer regional features to improve coordination among agents within the region. Our extensive evaluations using simulation of urban mobility, a widely used multimodal traffic simulation package, demonstrated that the proposed method outperformed previous methods and reduced the number of emergency stops, queue lengths, and waiting times. Jiajing Shen, Bingquan Yu, Qinpei Zhao, Weixiong Rao |
Int. J. Intell. Syst. | 4 |
| 2023 | STIP: A Seasonal Trend Integrated Predictor for Blood Glucose Level in Time Series
Weixiong Rao, Guangda Yang, Qinpei Zhao, Hongming Zhu, Xuefeng Li 0001, Yinjia Zhang |
ADMA (5) | 1 |
| 2023 | TIGAN: Trajectory Imputation via Generative Adversarial Network
Hongye Gao, Weixiong Rao |
ADMA (5) | 3 |
| 2023 | Learning to Simulate Complex Physical Systems: A Case StudyabstractComplex physical system simulation is important in many real world applications. We study the general simulation scenario to generate the response result when a physical object is applied by external factors. Traditional solvers on Partial Differential Equations (PDEs) suffer from significantly high computational cost. Many recent learning-based approaches focus on multivariate time series alike simulation prediction problem and do not work for our case. In this paper, we propose a novel two-level graph neural networks (GNNs) to learn the simulation result of a physical object applied by external factors. The key is a two-level graph structure where one fine mesh graph is mapped to multiple coarse one. Our preliminary evaluation on both synthetic and real datasets demonstrates that our work outperforms three state-of-the-arts by much lower errors. Jiasheng Shi, Weixiong Rao |
CIKM | 3 |
| 2023 | Scalable Communication for Mobile Multi-Agent Cooperative DetectionabstractCommunication is an effective mechanism to coordinate the behavior of mobile multi-agent systems. We propose a general mobile multi-agent cooperative detection framework, which provides a detection system with enhanced collaboration capabilities based on graph neural networks and reinforcement learning. Each agent gains information about the surrounding area by heuristic KNN, and then exchanges this part of subgraph information through communication so as to obtain global information in a decentralized way. At the same time, to avoid the computation overhead caused by the increase of the number of agents, attention mechanism is used to filter and optimize the communication process to improve scalability. We evaluate our method on large-scale detection tasks. Our approach is able to outperform the baselines, while making superior communication efficiency1. Hongye Gao, Tianlong Zhou, Weixiong Rao |
MDM | 3 |
| 2023 | Category tree distance: a taxonomy-based transaction distance for web user analysis
Yinjia Zhang, Qinpei Zhao, Yang Shi 0002, Weixiong Rao |
Data Min. Knowl. Discov. | 5 |
| 2023 | Outdoor Position Recovery From Heterogeneous Telco Cellular DataabstractRecent years have witnessed unprecedented amounts of data generated by telecommunication (Telco) cellular networks. For example, measurement records (MRs) are generated to report the connection states between mobile devices and Telco networks, e.g., received signal strength. MR data have been widely used to localize outdoor mobile devices for human mobility analysis, urban planning, and traffic forecasting. Existing works using first-order sequence models such as the Hidden Markov Model (HMM) attempt to capture spatio-temporal locality in underlying mobility patterns for lower localization errors. The HMM approaches typically assume stable mobility patterns of the underlying mobile devices. Yet real MR datasets exhibit heterogeneous mobility patterns due to mixed transportation modes of the underlying mobile devices and uneven distribution of the positions associated with MR samples. Thus, the existing solutions cannot handle these heterogeneous mobility patterns. To this end, we propose a multi-task learning-based deep neural network (DNN) framework, namely${{\sf PRNet}}$$^+$, to incorporate outdoor position recovery and transportation mode detection. To make sure that${{\sf PRNet}}$$^+$can work, we develop a feature extraction module to precisely learn local-, short- and long-term spatio-temporal locality from heterogeneous MR samples. Extensive evaluation on eight datasets collected at three representative areas in Shanghai indicates that${{\sf PRNet}}$${^{+}}$greatly outperforms state-of-the-arts by lower localization errors. Yige Zhang, Weixiong Rao, Kun Zhang 0001, Lei Chen 0002 |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2021 | Learning Shortest Paths on Large Dynamic GraphsabstractThe shortest path problem (SPP) in graph theory has wide applications in daily travels, transportation and network routing. Existing works do not work well on large dynamic graphs and suffer from either low scalability or gealization issue. To overcome these issues, in this paper, we propose an efficient and effective learning framework, namely SPP-GS, to solve the SPP problem on large dynamic graph. The key components of SPP- GS include the techniques to decompose a large SPP instance into multiple small instance and the developed GCN-DQN model to solve small SPP instances. The evaluation result on 7 real road network graphs indicates that our approach SPP-GS performs well on large dynamic graphs by rather high quality and reasonable running time. Jiaming Yin, Weixiong Rao, Chenxi Zhang 0001 |
MDM | 2 |
| 2021 | aHCQ: Adaptive Hierarchical Clustering Based Quantization Framework for Deep Neural Networks
Weixiong Rao, Qinpei Zhao |
PAKDD (2) | 2 |
| 2021 | A Data-Driven Sequential Localization Framework for Big Telco DataabstractThe proliferation of telco networks and mobile terminals brings the accumulation of tremendous amounts of measure report(MR) data at a rapid pace. The MR data is generated by mobile objects while connecting to data services and is stored in backend data centers. To geo-tag or localize such MR data is believed to have a profound effect on the analytics and optimizations of telco and traffic networks. However, MR records are of noisy and partial observations regarding to mobile objects' geo-locations and hence pose challenges to accurate telco data localization. There have been quite a few attempts. Single-point localization methods map a MR record to a location, but come out with limited accuracies due to the ignorance of spatiotemporal coherence of successive MR records. Recent efforts on sequential localization techniques alleviate this by mapping a sequence of MR records to a trajectory. However, existing solutions are often with assumptions on specific models, e.g., mobility and signal strength distributions, or priori knowledge on topology space, e.g., road networks, limiting the deployment in practice. To this end, we propose a data-driven framework to tackle the challenges in sequential telco localization. We solely use raw MR records and a public third-party GPS dataset for the learning of the correlations between mobile objects' locations and MR records, requiring no model assumptions and priori knowledge. To handle the data-intensive workloads during the learning process, we use materialized views for efficient online localization and light-weighted indexing techniques for periodical parameters tuning, in order to improve the efficiency and scalability. Results on real data show that our solution achieves 58.8 percent improvement in median localization errors compared with state-of-art sequential localization techniques that require hypothesis models and priori knowledge, making our solution superior in terms of effectiveness, efficiency, and employability. Fangzhou Zhu, Mingxuan Yuan, Xike Xie, Shenglin Zhao, Weixiong Rao |
IEEE Trans. Knowl. Data Eng. | 6 |
| 2020 | Smarter and Safer Traffic Signal Controlling via Deep Reinforcement LearningabstractRecently deep reinforcement learning (DRL) has been used for intelligent traffic light control. Unfortunately, we find that state-of-the-art on DRL-based intelligent traffic light essentially adopts discrete decision making and would suffer from the issue of unsafe driving. Moreover, existing feature representation of environment may not capture dynamics of traffic flow and thus cannot precisely predict future traffic flows. To overcome these issues, in this paper, we propose a DDPG-based DRL framework to learn a continuous time duration of traffic signal phases by introducing 1) a transit phase before the change of current phase for better safety, and 2) vehicle moving speed into feature representation for more precise estimation of traffic flow in next phase. Our preliminary evaluation on a well-known simulator SUMO indicates that our work significantly outperforms a recent work by much smaller number of emergency stops, queue length and waiting time. Bingquan Yu, Jinqiu Guo, Qinpei Zhao, Weixiong Rao |
CIKM | 5 |
| 2020 | Towards Accurate Retail Demand Forecasting Using Deep Neural Networks
Shanhe Liao, Jiaming Yin, Weixiong Rao |
DASFAA (3) | 3 |
| 2020 | Accurate Demand Forecasting for Retails with Deep Neural Networks
Shanhe Liao, Weixiong Rao |
EDBT | 2 |
| 2020 | IFLoc: Indoor Height Estimation by Telco DataabstractUnderstanding the fine-grained distribution of telecommunication (Telco) signals in terms of a three-dimensional (3D) space is important for Telco operators to manage, operate and optimize Telco networks. It is particularly true in nowadays urban cities with a large number of high buildings. One of the key tasks is to infer the location height of mobile devices, e.g., the floor within a high building where mobile devices are located. However, precise height estimation is challenging due to complex Telco signal propagation within an indoor 3D space, sparse cell tower deployment and scarce training samples. To tackle these issues, in this paper, we propose an indoor MR height estimation framework, namely IFLoc, via a machine learning model. IFLoc first builds a training MR database via a pre-processing step to comfortably tag raw MR samples by precisely inferred height from auxiliary data such as GPS and barometer readings. Next, IFLoc trains a regression model for height estimation by a set of developed techniques including 3D space division, post-processing techniques, feature augmentation and an improved SVR (Supported Vector Regression) model. Our evaluation on eight real datasets collected within five representative high buildings in Shanghai validates that IFLoc outperforms state-of-the-art counterparts in particularly with scarce training data. Jinhua Lv, Yige Zhang, Weixiong Rao, Jiehua Chen 0005, Xiaofeng Hu, Qinglin Chen |
MDM | 3 |
| 2020 | SST: Synchronized Spatial-Temporal Trajectory Similarity Search
Weixiong Rao, Chengxi Zhang, Gong Su, Qi Zhang 0009 |
GeoInformatica | 2 |
| 2019 | Experimental Study of Multivariate Time Series Forecasting ModelsabstractMultivariate time series forecasting has wide applications such as traffic flow prediction, supermarket commodity demand forecasting and etc. In literature, Due to the complex temporal patterns and inter-dependencies among multivariate time series, a large number of forecasting models have been developed. However, one question still remains unclear: how these models perform on a certain forecasting task, and there is lack of comprehensive performance comparison of these models on different tasks. To this end, in this paper, we conduct a systematic evaluation of eight representative forecasting models over eight multivariate time series datasets, and have the following findings: 1) When the datasets exhibit strong periodic patterns, deep learning models perform best. Otherwise on the datasets in a non-periodic manner, the statistical models such as ARIMA perform best. 2) For the long term prediction involving a high horizon value, the direct prediction strategy could lead to lower errors than the recursive one, but at the cost of higher training time. 3) For the multivariate time series explicitly involving graphic inter-dependencies among the multivariates, e.g., the road network topology in the spatio-temporal time series of traffic volumes in multiple routes, the Graph Convolution Network can incorporate the graphic inter-dependencies into their forecasting models for smaller prediction errors. Jiaming Yin, Weixiong Rao, Mingxuan Yuan, Kai Zhao 0011, Chenxi Zhang 0001, Qinpei Zhao |
CIKM | 2 |
| 2019 | PRNet: Outdoor Position Recovery for Heterogenous Telco Data by Deep Neural NetworkabstractRecent years have witnessed unprecedented amounts of telecommunication (Telco) data generated by Telco networks. For example, measurement records (MRs) are generated to report the connection states, e.g., received signal strength, between mobile devices and Telco networks. MR data have been widely used to precisely recover outdoor locations of mobile devices for the applications e.g., human mobility, urban planning and traffic forecasting. Existing works using first-order sequence models such as the Hidden Markov Model (HMM) attempt to capture the spatio-temporal locality in underlying mobility patterns for lower localization errors. Such HMM approaches typically assume stable mobility pattern of underlying mobile devices. Yet real MR datasets frequently exhibit heterogeneous mobility patterns due to mixed transportation modes of underlying mobile devices and uneven distribution of the positions associated with MR samples. To address this issue, we propose a deep neural network (DNN)-based position recovery framework, namely PRNet, which can ensemble the power of CNN, sequence model LSTM, and two attention mechanisms to learn local, short- and long-term spatio-temporal dependencies from input MR samples. Extensive evaluation on six datasets collected at three representative areas (core, urban, and suburban areas in Shanghai, China) indicates that PRNet greatly outperforms seven counterparts. Yige Zhang, Weixiong Rao, Kun Zhang 0001, Mingxuan Yuan |
CIKM | 2 |
| 2019 | Traffic Congestion Prediction by Spatiotemporal Propagation PatternsabstractAccurate prediction of traffic congestion at the granularity of road segment is important for planning travel routes and optimizing traffic control in urban areas. Previous works often calculated only the average congestion levels of a large region covering many road segments and did not take into account spatial correlation between road segments, resulting in inaccurate and coarse-grained prediction. To overcome these issues, we propose in this paper CPM-ConvLSTM, a spatiotemporal model for short-term prediction of congestion level in each road segment. Our model is built on a spatial matrix which incorporates both the congestion propagation pattern and the spatial correlation between road segments. The preliminary experiments on the traffic data set collected from Helsinki, Finland prove that CPM-ConvLSTM greatly outperforms 6 counterparts in terms of prediction accuracy. Xiaolei Di, Yu Xiao 0001, Chao Zhu 0002, Qinpei Zhao, Weixiong Rao |
MDM | 6 |
| 2019 | CLEAN: Frequent Pattern-Based Trajectory Spatial-Temporal Compression on Road NetworksabstractThe volume of trajectory data has become tremendously large in recent years. How to efficiently maintain and compute such trajectory data becomes a challenging task. In this paper, we propose a trajectory spatial and temporal compression framework, namely CLEAN. The key of spatial compression is to mine meaningful trajectory frequent patterns on road networks. By treating the mined patterns as dictionary items, we have the chance to encode a long trajectory by shorter paths, thus leading to smaller space cost. Meanwhile, we design an error-bounded temporal compression on top of the identified spatial patterns for much low space cost. Extensive experiments on real trajectory datasets validate that CLEAN significantly outperforms existing state-of-art approaches in terms of both space saving and runtime of trajectory compression. Qinpei Zhao, Chenxi Zhang 0001, Gong Su, Qi Zhang 0009, Weixiong Rao |
MDM | 6 |
| 2018 | Frequent Pattern-Based Map-Matching on Low Sampling Rate TrajectoriesabstractMap-matching is an important preprocessing task for many location-based services (LBS). It projects each GPS point in trajectory data onto digital maps. The state of art work typically employed the Hidden Markov model (HMM) by shortest path computation. Such shortest path computation may not work very well for very low sampling rate trajectory data, leading to low matching precision and high running time. To solve this problem, this paper, we first identify the frequent patterns from historical trajectory data and next perform the map matching for higher precision and faster running time. Since the identified frequent patterns indicate the mobility behaviours for the majority of trajectories, the map matching thus has chance to satisfy the matching precision with high confidence. Moreover, the proposed FP-forest structure can greatly speedup the lookup of frequent paths and lead to high computation efficiency. Our experiments on real world data set validate that the proposed FP-matching outperforms state of arts in terms of effectiveness and efficiency. Weixiong Rao, Mingxuan Yuan |
MDM | 2 |
| 2017 | Experimental Study of Telco Localization MethodsabstractTelecommunication (Telco) localization is a technique to accurately locate mobile devices (MDs) using measurement report (MR) data, and has been widely used in Telco industry. Many techniques have been proposed, including measurement-based statistical algorithms, fingerprinting algorithms and different machine learning-based algorithms. However, it has not been well studied yet on how these algorithms perform on various Telco MR data sets. In this paper, we conduct a comprehensive experimental study of five state-of-art algorithms for Telco localization. Based on real data sets from two Telco networks, we study the localization performance of such algorithms. We find that a Random Forest-based machine learning algorithm performs best in most experiments due to high localization accuracy and insensitivity to data volume. The experimental result and observation in this paper may inspire and enhance future research in Telco localization. Weixiong Rao, Fangzhou Zhu, Mingxuan Yuan |
MDM | 2 |
| 2017 | Confidence Model-Based Data Repair for Telco LocalizationabstractTelecommunication (Telco) localization is a technique to localize mobile devices (MDs) outdoor by using measurement report (MR) data. Unfortunately, existing Telco localization approaches (with localization error > 50 meters) cannot achieve comparable localization accuracy as GPS (with localization error ±10 meters). In particular, due to signal interference and attenuation caused by high buildings in urban cities, it is hard to achieve high localization accuracy if MR records contain unstable signal data. To this end, we propose to first detect those MR records incurring high localization errors and next repair the predicted location (with high errors). Our experiments on two real MR data sets from 2G GSM and 4G LTE Telco networks verify that a Telco localization approach, enhanced by the proposed detection and repair algorithms, can greatly improve localization accuracy. For example, the enhancement algorithm on 2G and 4G data sets can achieve 29.5 and 13.2 meters of median errors, around 160.68% and 201.51% better than previous results. Such result indicates that our work can achieve nearly comparable localization accuracy as GPS. Yige Zhang, Weixiong Rao, Mingxuan Yuan |
MDM | 2 |
| 2017 | Topic Model-Based Road Network Inference from Massive TrajectoriesabstractRecent years witnessed popular use of various mobile devices, e.g., smart phones, vehicle networks and wearable watches. Such mobile devices generate massive trajectory data, and literature have proposed various algorithms to leverage the trajectory data for map inference. Unfortunately, such algorithms are hard to achieve both high map quality and computation efficiency. In this paper, we propose a solution framework to infer road network maps with high quality and efficiency. The key of our map inference is to divide map extent into smaller cells and maintain a binary cell-trajectory matrix. The binary matrix determines whether or not a trajectory passes a cell. We infer the importance of each cell from the matrix using a popular topic model (e.g., LDA [13] and pLSA [8]). Based on such computed importance, we next infer representative points and road segments to derive a road network map. Our extensive experiments on real data sets verify that the proposed inference algorithm can achieve higher map quality and meanwhile 1.5 ×, 6.8 × and 280 × shorter running time, when compared with three state of the arts including three representative work [4], [7], [14]. Renjie Zheng, Qin Liu 0004, Weixiong Rao, Mingxuan Yuan, Zhongxiao Jin |
MDM | 3 |
| 2016 | LDA Revisited: Entropy, Prior and ConvergenceabstractInference algorithms of latent Dirichlet allocation (LDA), either for small or big data, can be broadly categorized into expectation-maximization (EM), variational Bayes (VB) and collapsed Gibbs sampling (GS). Looking for a unified understanding of these different inference algorithms is currently an important open problem. In this paper, we revisit these three algorithms from the entropy perspective, and show that EM can achieve the best predictive perplexity (a standard performance metric for LDA accuracy) by minimizing directly the cross entropy between the observed word distribution and LDA's predictive distribution. Moreover, EM can change the entropy of LDA's predictive distribution through tuning priors of LDA, such as the Dirichlet hyperparameters and the number of topics, to minimize the cross entropy with the observed word distribution. Finally, we propose the adaptive EM (AEM) algorithm that converges faster and more accurate than the current state-of-the-art SparseLDA [20] and AliasLDA [12] from small to big data and LDA models. The core idea is that the number of active topics, measured by the residuals between E-steps at successive iterations, decreases significantly, leading to the amortized σ(1) time complexity in terms of the number of topics. The open source code of AEM is available at GitHub. Mingxuan Yuan, Weixiong Rao, Jianfeng Yan |
CIKM | 4 |
| 2016 | City-Scale Localization with Telco Big DataabstractIt is still challenging in telecommunication (telco) industry to accurately locate mobile devices (MDs) at city-scale using the measurement report (MR) data, which measure parameters of radio signal strengths when MDs connect with base stations (BSs) in telco networks for making/receiving calls or mobile broadband (MBB) services. In this paper, we find that the widely-used location based services (LBSs) have accumulated lots of over-the-top (OTT) global positioning system (GPS) data in telco networks, which can be automatically used as training labels for learning accurate MR-based positioning systems. Benefiting from these telco big data, we deploy a context-aware coarse-to-fine regression (CCR) model in Spark/Hadoop-based telco big data platform for city-scale localization of MDs with two novel contributions. First, we design map-matching and interpolation algorithms to encode contextual information of road networks. Second, we build a two-layer regression model to capture coarse-to-fine contextual features in a short time window for improved localization performance. In our experiments, we collect 108 GPS-associated MR records in the centroid of Shanghai city with 12 x 11 square kilometers for 30 days, and measure four important properties of real-world MR data related to localization errors: stability, sensitivity, uncertainty and missing values. The proposed CCR works well under different properties of MR data and achieves a mean error of 110m and a median error of $80m$, outperforming the state-of-art range-based and fingerprinting localization methods. Fangzhou Zhu, Chen Luo 0003, Mingxuan Yuan, Yijian Zhu, Zhengqing Zhang, Tao Gu 0001, Weixiong Rao |
CIKM | 8 |
| 2016 | BMF: An Indexing Structure to Support Multi-element Check
Weixiong Rao |
WAIM (1) | 2 |
| 2014 | Cost-Based Optimization of Logical Partitions for a Query Workload in a Hadoop Data Warehouse
Shu Peng, Xiaoyang Sean Wang, Weixiong Rao, Min Yang 0002, Yu Cao 0004 |
APWeb | 4 |
| 2014 | Cost-Based Join Algorithm Selection in Hadoop
Shu Peng, Xiaoyang Sean Wang, Weixiong Rao, Min Yang 0002, Yu Cao 0004 |
WISE (2) | 4 |
| 2013 | Subscription Privacy Protection in Topic-Based Pub/Sub
Weixiong Rao, Lei Chen 0002, Mingxuan Yuan, Sasu Tarkoma, Hong Mei 0001 |
DASFAA (1) | 1 |
| 2013 | Bitlist: New Full-text Index for Low Space Cost and Efficient Keyword SearchabstractNowadays Web search engines are experiencing significant performance challenges caused by a huge amount of Web pages and increasingly larger number of Web users. The key issue for addressing these challenges is to design a compact structure which can index Web documents with low space and meanwhile process keyword search very fast. Unfortunately, the current solutions typically separate the space optimization from the search improvement. As a result, such solutions either save space yet with search inefficiency, or allow fast keyword search but with huge space requirement. In this paper, to address the challenges, we propose a novel structure bitlist with both low space requirement and supporting fast keyword search. Specifically, based on a simple and yet very efficient encoding scheme, bitlist uses a single number to encode a set of integer document IDs for low space, and adopts fast bitwise operations for very efficient boolean-based keyword search. Our extensive experimental results on real and synthetic data sets verify that bitlist outperforms the recent proposed solution, inverted list compression [23, 22] by spending 36.71% less space and 61.91% faster processing time, and achieves comparable running time as [8] but with significantly lower space. Weixiong Rao, Lei Chen 0002, Pan Hui 0001, Sasu Tarkoma |
Proc. VLDB Endow. | 1 |
| 2013 | Toward Efficient Filter Privacy-Aware Content-Based Pub/Sub SystemsabstractIn recent years, the content-based publish/subscribe [12], [22] has become a popular paradigm to decouple information producers and consumers with the help of brokers. Unfortunately, when users register their personal interests to the brokers, the privacy pertaining to filters defined by honest subscribers could be easily exposed by untrusted brokers, and this situation is further aggravated by the collusion attack between untrusted brokers and compromised subscribers. To protect the filter privacy, we introduce an anonymizer engine to separate the roles of brokers into two parts, and adapt the k-anonymity and `-diversity models to the contentbased pub/sub. When the anonymization model is applied to protect the filter privacy, there is an inherent tradeoff between the anonymization level and the publication redundancy. By leveraging partial-order-based generalization of filters to track filters satisfying k-anonymity and ℓ-diversity, we design algorithms to minimize the publication redundancy. Our experiments show the proposed scheme, when compared with studied counterparts, has smaller forwarding cost while achieving comparable attack resilience. Weixiong Rao, Lei Chen 0002, Sasu Tarkoma |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2012 | A General Framework for Publishing Privacy Protected and Utility Preserved GraphabstractThe privacy protection of graph data has become more and more important in recent years. Many works have been proposed to publish a privacy preserving graph. All these works prefer publishing a graph, which guarantees the protection of certain privacy with the smallest change to the original graph. However, there is no guarantee on how the utilities are preserved in the published graph. In this paper, we propose a general fine-grained adjusting framework to publish a privacy protected and utility preserved graph. With this framework, the data publisher can get a trade-off between the privacy and utility according to his customized preferences. We used the protection of a weighted graph as an example to demonstrate the implementation of this framework. Mingxuan Yuan, Lei Chen 0002, Weixiong Rao, Hong Mei 0001 |
ICDM | 3 |
| 2012 | Distributed top-k full-text content dissemination
Weixiong Rao, Lei Chen 0002 |
Distributed Parallel Databases | 1 |
| 2011 | STAIRS: Towards efficient full-text filtering and dissemination in DHT environments
Weixiong Rao, Lei Chen 0002, Ada Wai-Chee Fu |
VLDB J. | 1 |
| 2009 | STAIRS: Towards Efficient Full-Text Filtering and Dissemination in a DHT EnvironmentabstractNowadays contents in Internet like weblogs, wikipedia and news sites become "live". How to notify and provide users with the relevant contents becomes a challenge. Unlike conventional Web search technology or the RSS feed, this paper envisions a personalized full-text content filtering and dissemination system in a highly distributed environment such as a Distributed Hash Table (DHT). Users can subscribe to their interested contents by specifying some terms and threshold values for filtering. Then, published contents will be disseminated to the associated subscribers. We propose a novel and simple framework of filter registration and content publication, STAIRS. By the new framework, we propose three algorithms (default forwarding, dynamic forwarding and adaptive forwarding) to reduce the forwarding cost and false dismissal rate; meanwhile, the subscriber can receive the desired contents with no duplicates. In particular, the adaptive forwarding utilizes the filter information to significantly reduce the forwarding cost. Experiments based on two real query logs and two real datasets show the effectiveness of our proposed framework. Weixiong Rao, Ada Wai-Chee Fu, Lei Chen 0002, Hanhua Chen |
ICDE | 1 |
| 2007 | Optimal proactive caching in peer-to-peer network: analysis and applicationabstractAs a promising new technology with the unique properties like high efficiency, scalability and fault tolerance, Peer-to-Peer (P2P) technology is used as the underlying network to build new Internet-scale applications. However, one of the well known issues in such an application (for example WWW) is that the distribution of data popularities is heavily tailed with a Zipf-like distribution. With consideration of the skewed popularity we adopt a proactive caching approach to handle the challenge, and focus on two key problems: where (i.e. the placement strategy: where to place the replicas) and how (i.e. the degree problem: how many replicas are assigned to one specific content)? For the where problem, we propose a novel approach which can be generally applied to structured P2P networks. Next, we solve two optimization objectives related to the how problem: MAX_PERF and MIN_COST. Our solution is called PoPCache, and we discover two interesting properties: (1) the number of replicas assigned to each content is proportional to its popularity; (2) the derived optimal solutions are related to the entropy of popularity. To our knowledge, none of the previous works has mentioned such results. Finally, we apply the results of PoPCache to propose a P2P base web caching, called as Web-PoPCache. By means of web cache trace driven simulation, our extensive evaluation results demonstrate the advantages of PoPCache and Web-PoPCache. Weixiong Rao, Lei Chen 0002, Ada Wai-Chee Fu, Yingyi Bu |
CIKM | 1 |
| 2004 | MTrie: A Scalable Filtering Engine of Well-Structured XML Message Stream
Weixiong Rao, Yingjian Chen, Xinquan Zhang, Fanyuan Ma |
APWeb | 1 |
| 2003 | DEBIZ: A Decentralized Lookup Service for E-commerce
Zengde Wu, Weixiong Rao, Fanyuan Ma |
APWeb | 2 |