VLDB 2026 Research / reviewers in the wild / expert
Wuman Luo
dblp:96/6576
· DBLP profile ↗
38ranked-venue papers
2as first author
24since 2021 · last 2026
0000-0002-2480-3997ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 11 · 2 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 11 · 9 since 2021Artificial intelligence and machine learning · 10 · 8 since 2021Systems, architecture and hardware · 6 · 1 since 2021Security and privacy · 3 · 3 since 2021Computer networks · 2Human-computer interaction and ubiquitous computing · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Dual feature-driven approach for partial multi-label learning
Yanqiang Tu, Gengyu Lyu, Wenbin Qian, Wuman Luo |
Pattern Recognit. | 5 |
| 2026 | Multi-Agent Learning for Precise and Collaborative Control of Anesthetics in TIVAabstractPrecise and collaborative control of multiple anesthetics in Total Intravenous Anesthesia (TIVA) is essential for ensuring patient safety and maintaining the target depth of anesthesia (DoA). However, existing automated anesthesia control methods often fail to effectively capture the complex synergistic interactions between anesthetics and lack adaptability to patient-specific physiological variability, thereby limiting their clinical applicability. To address these issues, we propose AnesMADRL, a novel Multi-Agent Deep Reinforcement Learning (MADRL)-based framework that leverages the Counterfactual Multi-Agent algorithm for effective credit assignment between agents controlling propofol and remifentanil, and adopts a continuous action space to enable fine-grained dose adjustments. Furthermore, AnesMADRL integrates comprehensive patient-specific physiological indicators and employs a random forest-based simulator to generate dynamic and diverse training environments. Experimental results show that AnesMADRL significantly outperforms baseline methods and human expertise in terms of anesthetic efficiency and total drug consumption. Relative to human expertise, AnesMADRL achieves roughly twofold efficiency while reducing total dose to about one-half, highlighting its potential to enhance patient safety and optimize clinical outcomes. Huijie Li, Yide Yu, Yuejing Zhai, Anmin Hu, Jian Huo, Wuman Luo |
IEEE J. Biomed. Health Informatics | 7 |
| 2026 | Value Decomposition-Based Multi-Agent Learning for Anesthetics Collaborative ControlabstractAutomated control of personalized multiple anesthetics in clinical Total Intravenous Anesthesia (TIVA) is crucial yet challenging. Current systems, including target-controlled infusion (TCI) and closed-loop systems, either rely on relatively static pharmacokinetic/pharmacodynamic (PK/PD) models or focus on single anesthetic control. So they limit both personalization and collaborative control. To address these issues, we propose a novel Value Decomposition Multi-Agent Deep Reinforcement Learning (VD-MADRL) framework based on Markov Game (MG) for Personalized Multiple Anesthetics Control in a Closed-Loop system (PMAC-CL). VD-MADRL optimizes the collaboration between two anesthetics propofol (Agent I) and remifentanil (Agent II) by leveraging a MG to identify optimal actions among heterogeneous agents. We employ various value function decomposition methods to resolve the credit allocation problem and enhance collaborative control. We also introduce a multivariate environment model based on random forest (RF) for anesthesia state simulation. To ensure data validity, we design a data resampling and alignment technique to synchronize trajectory data from different devices, avoiding gradient explosion and maintaining conformity to Markov property. Extensive experiments on general and thoracic surgery datasets demonstrate that VD-MADRL provides more refined dose adjustments and maintains multiple anesthesia state indicators more stably at target levels compared to human experience. Especially, the best-performing algorithm, VDN in general surgery with online training, achieved a 16.4% increase in cumulative reward (CR) and a 58.0% reduction in mean MDPE compared to human experience. This demonstrates its great clinical value. Huijie Li, Yide Yu, Anmin Hu, Jian Huo, Chaoran Wu, Wuman Luo |
IEEE J. Biomed. Health Informatics | 8 |
| 2025 | Robust Prototype-Driven Patient Representation Enhancement for Disease PredictionabstractElectronic health records (EHR) contain sequential patient visit data with critical features for disease prediction. Recent inter-patient modeling approaches leverage information from other patients to improve generalization, but they often face challenges like noisy predictions due to ambiguous latent spaces and neglect intra-class diversity. To address these issues, we propose RPPRE, a Robust Prototype-driven Patient Representation Enhancer for EHR-based disease prediction. RPPRE enhances both inter-class discrimination and intra-class diversity by first stratifying disease classes into core, intermediate, and peripheral prototypes, then using a contrastive loss to align patient representations with these prototypes. Experiments on the MIMIC-III Respiratory and PhysioNet Sepsis datasets show that RPPRE consistently improves AUROC, AUPRC, and F1-score across seven backbone models and outperforms existing inter-patient approaches. Ablation studies further validate the importance of prototype stratification and contrastive enhancement. Our code is released at https://github.com/cp3mvp-24/RPPRE. Hongxu Yuan, Xiaozhu Jing, Yuzheng Yan, Wuman Luo |
BIBM | 4 |
| 2025 | Modeling Frequency Correlation to Achieve Better Long-Term Series Forecasting
Yuejing Zhai, Wuman Luo |
IEEE Big Data | 3 |
| 2025 | The Dual-Focus Dynamic Multiple Imputation Approach For MNAR Missing Values In Medical DataabstractMissing value imputation in medical datasets is an important research topic. Most studies assume that missing values are Missing at Random (MAR), but verifying whether data are MAR or Missing Not at Random (MNAR) is challenging because it is impossible to evaluate whether the unobserved data is related to the missing data. Besides, the possible evaluation method sensitivity analysis has many limitations and shortcomings. Therefore, considering the extreme complexity of the human body, treating missing values as MNAR is preferable. Existing MNAR imputation methods require assumptions about the distribution of unobserved variables and establish joint probabilities, so they rely on experience and may lead to bias. In addition, these methods also struggle with complex relationships and distributions. To address these problems, this paper proposes the Dual-Focus Dynamic Multiple Imputation (DDMI) model for MNAR missing values in medical data. The DDMI uses piece-wise approximation to decompose complex relationships and directly calculates the impact of unobserved variables on patient indicators, avoiding distribution assumptions. In addition, the DDMI captures both population and individual-level information to predict missing values and then refines the results through multiple iterations, and the original information is combined in each iteration to mitigate information loss and improve convergence. We test DDMI on two real-world datasets. Results show that DDMI outperforms other methods. Yuejing Zhai, Huijie Li, Wuman Luo |
CEC | 4 |
| 2025 | Graph Attention Network and Dynamic Adjustment Mechanism for Drug Recommendation
Xionghui Lai, Wuman Luo |
ICIC (25) | 2 |
| 2025 | An Adaptive Multi-Indicator Contrastive Predictive Coding Framework for Patient Representation LearningabstractEffective patient representation learning from Electronic Health Records (EHR) is essential for improving disease prediction models, yet it faces critical challenges such as the scarcity of labeled data and the difficulty of capturing complex temporal and multi-indicator relationships. To address these limitations, we propose the Adaptive Multi-Indicator Contrastive Predictive Coding (AMCPC) framework, a self-supervised learning approach designed for EHR data. AMCPC incorporates two key innovations: first, it employs an adaptive optimal window size selection algorithm to segment patient visit sequences into temporal subwindows, which enables the model to focus on localized, context-specific health patterns; second, it extends Contrastive Predictive Coding (CPC) with a multi-indicator approach, leveraging a 2D convolutional neural network (CNN) to capture global correlations among diverse medical indicators within each subwindow. Through extensive experiments on real-world clinical datasets, we demonstrate that AMCPC outperforms both fully-supervised and existing self-supervised methods in disease prediction tasks, particularly when trained on limited labeled data. Our results establish AMCPC as an effective framework for leveraging unlabeled EHR data for self-supervised pretraining, which can then be fine-tuned with a small amount of labeled data to significantly enhance downstream prediction performance, reducing reliance on large-scale labeled datasets. Hongxu Yuan, Yuzheng Yan, Xiaozhu Jing, Wuman Luo |
IJCNN | 4 |
| 2025 | PatSimBoosting: Enhancing Patient Representations for Disease Prediction Through Similarity Analysis
Yuzheng Yan, Ziyue Yu, Wuman Luo |
IoTBDS | 3 |
| 2025 | A Highly Nonlinear Survival Network for Hospital Readmission Prediction of Cardiac Patients
Yuejing Zhai, Lihua He, Wuman Luo |
IoTBDS | 4 |
| 2025 | ClinCoCoOp: An Interpretable Prompt Learning Framework with Clinical Concept Guidance for Context Optimization
Jianjing Wei, Wuman Luo, Bidong Chen |
PRCV (6) | 2 |
| 2025 | PSformer: Periodic-aware Semantic Transformer for Traffic PredictionabstractTraffic prediction plays an important role in Intelligent Transportation Systems (ITS). The main challenge lies in effectively capturing the dynamic multiple temporal periodic correlations and the long-range spatial correlation of traffic data. Despite the significant progress of many existing works, these methods often have two major limitations: 1) They mined the dynamic multi-period properties by using raw traffic sequences or the fixed periodicity strategy (e.g., hours, days, weeks), which failed to capture the dynamic multi-period characteristics of temporal correlation. 2) They mined the long-range spatial correlation of traffic data by stacking multilayer networks or directly using traditional similarity algorithms (e.g., conventional DTW). However, DTW has its own limitations leading to sub-optimal similarity assessment. To address these issues, we propose a periodic-aware spatial semantic transformer called PSformer for traffic prediction. Specifically, we propose the Periodic-aware Embedding Module (PAEmbed) to capture the dynamic multi-period properties by decoupling the traffic sequence into the multilevel frequency components via Fast Fourier Transform (FFT). In addition, we propose a Semantic Spatial Attention Mechanism (SSAM) to capture the long-range spatial correlation. In SSAM, we propose Time-weighted Dynamic Time Warping (TDTW) to model spatial correlations in semantically identical but geographically distant regions, which avoids considering two traffic patterns with large time spans as similar. Finally, to evaluate the performance of PSformer, we conduct extensive experiments on four real datasets. Experimental results show that our model achieves better performance than other state-of-the-art methods. Lihua He, Ziyue Yu, Wuman Luo |
SMC | 3 |
| 2025 | ST-RLNet: Spatio-temporal representation learning for multi-step traffic flow predictionabstractTraffic flow prediction provides valuable traffic information to transportation agencies and individuals in advance. Compared to next-step prediction, multi-step prediction provides users with traffic information for a longer time horizon, allowing users to have a more comprehensive understanding of traffic conditions. So far, various methods have been proposed for multi-step traffic flow prediction. However, most of them become sub-optimal in effectively detecting the spatio-temporal correlations of traffic data. Furthermore, as the number of prediction steps increases, the input data used to predict the flow of the next step tends to deviate further from the ground truth value. This deviation leads to a rapid decrease in prediction accuracy as the number of prediction steps increases. To address these issues, in this paper, we propose a deep spatio-temporal representation learning network named ST-RLNet for multi-step traffic flow prediction. The goal is to effectively generate the traffic data representation by better capturing the complex correlations of the data. In particular, we design a network called 3D-ConvLSTMNet to effectively extract short-term and long-term spatio-temporal data correlations for the next step prediction. To solve the performance degradation problem, we propose a feedback mechanism called PS-Feedback to dynamically reconstruct temporal correlation representations of input traffic flow for each round of next-step prediction. To evaluate the performance of the ST-RLNet, we conduct extensive experiments on two real-world datasets. Experimental results show that the ST-RLNet outperforms the state-of-the-art methods in both next-step and multi-step predictions, and exhibits consistent high performance under different traffic flows. Lihua He, Dian Zhang 0001, Wuman Luo |
Neurocomputing | 4 |
| 2024 | A survey on personalized document-level sentiment analysis
Jiayue Qiu, Ziyue Yu, Wuman Luo |
Neurocomputing | 4 |
| 2024 | Multi-perspective patient representation learning for disease prediction on electronic health recordsabstractAbstract Patient representation learning based on electronic health records (EHR) is a critical task for disease prediction. This task aims to effectively extract useful information on dynamic features. Although various existing works have achieved remarkable progress, the model performance can be further improved by fully extracting the trends, variations, and the correlation between the trends and variations in dynamic features. In addition, sparse visit records limit the performance of deep learning models. To address these issues, we propose the multi-perspective patient representation Extractor (MPRE) for disease prediction. Specifically, we propose frequency transformation module (FTM) to extract the trend and variation information of dynamic features in the time–frequency domain, which can enhance the feature representation. In the 2D multi-extraction network (2D MEN), we form the 2D temporal tensor based on trend and variation. Then, the correlations between trend and variation are captured by the proposed dilated operation. Moreover, we propose the first-order difference attention mechanism (FODAM) to calculate the contributions of differences in adjacent variations to the disease diagnosis adaptively. To evaluate the performance of MPRE and baseline methods, we conduct extensive experiments on two real-world public datasets. The experiment results show that MPRE outperforms state-of-the-art baseline methods in terms of AUROC and AUPRC. Ziyue Yu, Wuman Luo, Rita Tse, Giovanni Pau 0001 |
Knowl. Inf. Syst. | 3 |
| 2023 | MPRE: Multi-perspective Patient Representation Extractor for Disease PredictionabstractPatient representation learning based on electronic health records (EHR) is a critical task for disease prediction. This task aims to effectively extract useful information on dynamic features. Although various existing works have achieved remarkable progress, the model performance can be further improved by fully extracting the trends, variations, and the correlation between the trends and variations in dynamic features. In addition, sparse visit records limit the performance of deep learning models. To address these issues, we propose the Multi-perspective Patient Representation Extractor (MPRE) for disease prediction. Specifically, we propose Frequency Transformation Module (FTM) to extract the trend and variation information of dynamic features in the time-frequency domain, which can enhance the feature representation. In the 2D Multi-Extraction Network (2D MEN), we form the 2D temporal tensor based on trend and variation. Then, the correlations between trend and variation are captured by the proposed dilated operation. Moreover, we propose the First-Order Difference Attention Mechanism (FODAM) to calculate the contributions of differences in adjacent variations to the disease diagnosis adaptively. To evaluate the performance of MPRE and baseline methods, we conduct extensive experiments on two real-world public datasets. The experiment results show that MPRE outperforms state-of-the-art baseline methods in terms of AUROC and AUPRC. Ziyue Yu, Wuman Luo, Rita Tse, Giovanni Pau 0001 |
ICDM | 3 |
| 2023 | UCM: Personalized Document-Level Sentiment Analysis Based on User Correlation Mining
Jiayue Qiu, Ziyue Yu, Wuman Luo |
ICIC (4) | 3 |
| 2023 | OOCL-DDQN: Online Evaluation and Offline Training-Based Clipped Double DQN for Automated Anesthesia ControlabstractAnesthesia control is critical in surgical procedures. Nowadays, various methods have been proposed for automated anesthesia control. However, traditional control model-based methods and PK/PD-based methods cannot fully capture the individual characteristics of patients. Machine learning-based methods are not suitable for dealing with large-scale clinical anaesthesia data. The existing deep reinforcement learning-based (DRL) methods cannot provide both stable and reliable anesthesia strategies while ensuring the agent’s adaptability to environmental changes. To address these issues, in this paper, we propose a deep reinforcement learning model named OOCL-DDQN for automated anesthesia control. Specifically, we propose an online evaluation and offline training mechanism to well avoid excessive reliance on the environment models. To enhance the model stability, we design a clip method to optimize the Bellman equation in Double DQN. Besides, we propose a fixed-order random sampling method to enable the agent to effectively learn the relationship between anesthesia state and anesthetic dosage in different patients. In data preprocessing, we design a same-frequency resampling method to ensures that the resampled data obey the Markov property. To evaluate the performance of our proposed method, we conduct extensive experiments on a real-world dataset. The experiment results show that our proposed model gets better performances than the other state-of-the-art methods. Huijie Li, Jian Huo, Wuman Luo |
ICPADS | 4 |
| 2023 | FGRL-Net: Fine-Grained Personalized Patient Representation Learning for Clinical Risk Prediction Based on EHRsabstractPersonalized patient representation learning (PPRL) is a critical element in clinical risk prediction. It aims to obtain a complete portrait of each patient based on Electronic Health Records (EHR). Although existing works have achieved remarkable progress in healthcare prediction, there are still three major issues. First, feature correlation is crucial for risk prediction, but it has not yet been fully exploited by existing works. Second, variation pattern of dynamic feature contains useful information about patient's physical status, but adaptive pattern recognition is still a challenge. Third, existing works usually adopt a two-stage embedding process to process each dimension of the EHR data. However, some useful low-level information for PPRL will be lost. To address these issues, in this paper, we propose a fine-grained PPRL architecture named FG RL- N et for clinical risk prediction based on EHR. Specifically, we propose a Medical Feature Correlation Detection Module (FCM) to effectively learn the feature correlations for each patient and a Temporal Variation Pattern Recognition Module (TVM) to effectively detect the variation patterns of each dynamic feature. Moreover, we design a Fine-Grained Representation Mechanism (FGRM) to preserve the low-level information (from both feature and visit dimensions) useful for risk prediction. In addition, in the stage of data preprocessing, We utilize generic medical classification knowledge to classify numerical dynamic data. We conduct the in-hospital mortality experiment and the decompensation experiment on a real-world dataset. The experiment results show that the FGRL-Net outperforms state-of-the-art approaches. The source code is provided in github https://github.com/JackyChio/FGRL-Net. KaKit Chio, Lihua He, Dian Zhang 0001, Xu Yang 0010, Wuman Luo |
SMC | 6 |
| 2023 | DMNet: A Personalized Risk Assessment Framework for Elderly People With Type 2 DiabetesabstractType 2 diabetes is the most common chronic disease for the elderly people. This disease is difficult to be cured and causes continued medical expenses. The early and personalized risk assessment of type 2 diabetes is necessary. So far, various type 2 diabetes risk prediction methods have been proposed. However, these methods have three major issues: 1) not fully considering the importance of personal information and rating information of healthcare system, 2) not adopting the long-term temporal information, and 3) not comprehensively capturing the correlation between the diabetes risk factor categories. To address these issues, the personalized risk assessment framework for elderly people with type 2 diabetes is needed. However, it is very challenging due to two reasons, namely imbalanced label distribution and high-dimensional features. In this paper, we propose diabetes mellitus network framework (DMNet) for type 2 diabetes risk assessment of elderly people. Specifically, we propose tandem long short-term memory to extract the long-term temporal information of different diabetes risk categories. In addition, the tandem mechanism is used to capture the correlation between the diabetes risk factor categories. To balance the label distribution, we adopt the method of synthetic minority over-sampling technique with Tomek links. To form the better feature representations, we utilize entity embedding to solve the problem of high-dimensional features. To evaluate the performance of our proposed method, we conduct the experiments on a real-world dataset called Research on Early Life and Aging Trends and Effects. The experiment results show that DMNet outperforms the baseline methods in terms of six evaluation metrics (i.e., accuracy of 0.94, balanced accuracy of 0.94, precision of 0.95, F1-score of 0.95, recall of 0.95 and AUC of 0.94). Ziyue Yu, Wuman Luo, Rita Tse, Giovanni Pau 0001 |
IEEE J. Biomed. Health Informatics | 2 |
| 2022 | 3D-ConvLSTMNet: A Deep Spatio-Temporal Model for Traffic Flow PredictionabstractSpatiotemporal correlations are crucial for traffic flow prediction. So far, various traffic flow prediction methods based on convolutional neural network (CNN) and long short-term memory (LSTM) network have been proposed. However, the common CNN - based models cannot preserve the temporal information after the first layer. Although the 3D CNN-based models can effectively capture short-term spatial and tempo-ral features, they are not suitable for long-term information capturing. LSTM is excellent at long-term features extraction. However, it alone cannot be used for spatial information extraction. To address these issues, we propose a deep architecture called 3D-ConvLSTMNet to better capture the spatiotemporal correlations among the traffic data. Specifically, we proposed a short-long term spatiotemporal feature extraction module called 3D-ConvLSTM, which uses 3D CNN to extract short-term spatiotemporal correlations, and uses ConvLSTM to extract the long-term spatiotemporal correlations. To get the long-distance spatial features, we adopt the residual neural network to develop the depth of 3D-ConvLSTMNet. Finally, we utilize a channel-wise attention mechanism to quantify the contribution of each grid in space domain. To evaluate the performances of ConvLSTMNet, we conduct extensive experiments on two real-world datasets. The experiment results show that our model gets better performances than the other state-of-the-art methods. Lihua He, Wuman Luo |
MDM | 2 |
| 2022 | Deep Learning Hybrid Models for COVID-19 PredictionabstractCOVID-19 is a highly contagious virus. Blood test is one of effective methods for COVID-19 diagnosis. However, the issues of blood test are time-consuming and lack of medical staff. In this paper, four deep learning hybrid models are proposed to address these issues (i.e., CNN+GRU, CNN+Bi-RNN, CNN+Bi-LSTM, CNN+Bi-GRU). In addition, two best models, CNN and CNN+LSTM, from Turabieh et al. and Alakus et al., are implemented, respectively. Blood test data from Hospital Israelita Albert Einstein is used to train and test six models. The proposed best model, CNN+Bi-GRU, is accuracy of 0.9415, precision of 0.9417, recall of 0.9417, F1-score of 0.9417, AUC of 0.91, which outperforms the best models from Turabieh et al. and Alakus et al. Furthermore, the proposed model can help patients to get blood test results faster than traditional manual tests without errors caused by fatigue. The authors can envisage a wide deployment of proposed model in hospitals to alleviate the testing pressure from medical workers, especially in developing and underdeveloped countries. Ziyue Yu, Lihua He, Wuman Luo, Rita Tse, Giovanni Pau 0001 |
J. Glob. Inf. Manag. | 3 |
| 2022 | Machine learning-driven credit risk: a systemic reviewabstractAbstract Credit risk assessment is at the core of modern economies. Traditionally, it is measured by statistical methods and manual auditing. Recent advances in financial artificial intelligence stemmed from a new wave of machine learning (ML)-driven credit risk models that gained tremendous attention from both industry and academia. In this paper, we systematically review a series of major research contributions (76 papers) over the past eight years using statistical, machine learning and deep learning techniques to address the problems of credit risk. Specifically, we propose a novel classification methodology for ML-driven credit risk algorithms and their performance ranking using public datasets. We further discuss the challenges including data imbalance, dataset inconsistency, model transparency, and inadequate utilization of deep learning models. The results of our review show that: 1) most deep learning models outperform classic machine learning and statistical algorithms in credit risk estimation, and 2) ensemble methods provide higher accuracy compared with single models. Finally, we present summary tables in terms of datasets and proposed models. Rita Tse, Wuman Luo, Stefano D'Addona, Giovanni Pau 0001 |
Neural Comput. Appl. | 3 |
| 2021 | Deep Learning for COVID-19 Prediction based on Blood Test
Ziyue Yu, Lihua He, Wuman Luo, Rita Tse, Giovanni Pau 0001 |
IoTBDS | 3 |
| 2018 | Diverse Mobile System for Location-Based Mobile DataabstractThe value of large amount of location‐based mobile data has received wide attention in many research fields including human behavior analysis, urban transportation planning, and various location‐based services. Nowadays, both scientific and industrial communities are encouraged to collect as much location‐based mobile data as possible, which brings two challenges: (1) how to efficiently process the queries of big location‐based mobile data and (2) how to reduce the cost of storage services, because it is too expensive to store several exact data replicas for fault‐tolerance. So far, several dedicated storage systems have been proposed to address these issues. However, they do not work well when the ranges of queries vary widely. In this work, we design a storage system based on diverse replica scheme which not only can improve the query processing efficiency but also can reduce the cost of storage space. To the best of our knowledge, this is the first work to investigate the data storage and processing in the context of big location‐based mobile data. Specifically, we conduct in‐depth theoretical and empirical analysis of the trade‐offs between different spatial‐temporal partitioning and data encoding schemes. Moreover, we propose an effective approach to select an appropriate set of diverse replicas, which is optimized for the expected query loads while conforming to the given storage space budget. The experiment results show that using diverse replicas can significantly improve the overall query performance and the proposed algorithms for the replica selection problem are both effective and efficient. Qing Liao 0001, Haoyu Tan, Wuman Luo, Ye Ding 0002 |
Wirel. Commun. Mob. Comput. | 3 |
| 2016 | Clockwise compression for trajectory data under road network constraintsabstractBig trajectory data introduces severe challenges for data storage and communication. In this paper, we propose a novel compression framework called Clockwise Compression Framework (CCF) for big trajectory data compression under road network constraints. In CCF, we design several new methods: 1) a spatial compression algorithm called Enhanced Clockwise Encoding (ECE), 2) a temporal compression algorithm called Fitting-based Temporal Simplification (FTS), and 3) a dedicated querier that processes queries based on the above spatial and temporal compression algorithms, without fully decompressing the trajectroy data. By leveraging the topological information of the road network, CCF is able to perform both spatial compression and temporal compression in on-line modes. We perform extensive experiments in a real big trajectory dataset to verify both effectiveness and efficiency of our methods. CCF shows promising performances in various metrics and outperforms the state-of-the-art methods. Yudian Ji, Yuda Zang, Wuman Luo, Xibo Zhou, Ye Ding 0002, Lionel M. Ni |
IEEE BigData | 3 |
| 2016 | A Comparison of Road-Network-Constrained Trajectory Compression MethodsabstractThe popularity of location-acquisition devices has led to a rapid increase in the amount of trajectory data collected. The large volume of trajectory data causes the difficulties of storing and processing the data. Various trajectory compression methods are therefore proposed to deal with these problems. In this paper, we overview the existing road-network-constrained trajectory compression methods and propose a novel classification based on the features leveraged by them. We also propose new methods that fill in the research blanks indicated by the classification. We conduct a thorough comparison among the existing and new road-network-constrained trajectory compression methods. The performances of the methods are studied via various metrics on real-world dataset. We make new discoveries regarding the performances and the scalability of existing methods, and provide guidelines of road-network-constrained trajectory compression for various scenarios. Yudian Ji, Hao Liu 0026, Ye Ding 0002, Wuman Luo |
ICPADS | 5 |
| 2014 | Inferring Road Type in Crowdsourced Map Services
Ye Ding 0002, Jiangchuan Zheng, Haoyu Tan, Wuman Luo, Lionel M. Ni |
DASFAA (2) | 4 |
| 2014 | Exploring the Use of Diverse Replicas for Big Location Tracking DataabstractThe value of large amount of location tracking data has received wide attention in many applications including human behavior analysis, urban transportation planning, and various location-based services (LBS). Nowadays, both scientific and industrial communities are encouraged to collect as much location tracking data as possible, which brings about two issues: 1) it is challenging to process the queries on big location tracking data efficiently, and 2) it is expensive to store several exact data replicas for fault-tolerance. So far, several dedicated storage systems have been proposed to address these issues. However, they do not work well when the query ranges vary widely. In this paper, we present the design of a storage system using diverse replica scheme which improves the query processing efficiency with reduced cost of storage space. To the best of our knowledge, we are the first to investigate the data storage and processing in the context of big location tracking data. Specifically, we conduct in-depth theoretical and empirical analysis of the trade-offs between different spatio-temporal partitioning schemes as well as data encoding schemes. Then we propose an effective approach to select an appropriate set of diverse replicas, which is optimized for the expected query loads while conforming to the given storage space budget. The experiment results confirm that using diverse replicas can significantly improve the overall query performance. The results also demonstrate that the proposed algorithms for the replica selection problem is both effective and efficient. Ye Ding 0002, Haoyu Tan, Wuman Luo, Lionel M. Ni |
ICDCS | 3 |
| 2014 | MRTune: A simulator for performance tuning of MapReduce jobs with skewed dataabstractMapReduce is a programming model designed by Google that has been widely used for both high performance computing and big data processing. Although the programming model is simple, it is very challenging to conduct performance tuning for a MapReduce job, considering the complexities of the configuration parameters and various tradeoffs between the performance gain of the optimization approaches and the extra overhead they bring about. One naive way to address this issue is to run the MapReduce jobs repeatedly using different combinations of configuration parameters and optimization methods, then select the one with the shortest running time. However, real execution is impractical because the combinations may be too many and the time of one run of each combination may be too long. Therefore, it is desirable if we can efficiently estimate the runtime of a job without real execution using only the input data and the configuration parameter settings of the cluster. In this paper, we propose a novel MapReduce simulator called MRTune for runtime estimation of MapReduce jobs. MRTune takes the key distribution of input data into consideration and can work well even when the key distribution of data is skewed. Moreover, MRTune can estimate the runtime of a job in the presence of unpredictable task failures. We evaluate MRTune implementing MapReduce jobs with Zipfian distributed input data. The result shows that MRTune can estimate the runtime of MapReduce jobs with high accuracy and efficiency while the key distribution of input data is skewed. We also conduct two case studies to analyse the impact of data skew and task failures on a MapReduce job. Xibo Zhou, Wuman Luo, Haoyu Tan |
ICPADS | 2 |
| 2014 | MR-DBSCAN: a scalable MapReduce-based DBSCAN algorithm for heavily skewed data
Yaobin He, Haoyu Tan, Wuman Luo, Shengzhong Feng, Jianping Fan 0002 |
Frontiers Comput. Sci. | 3 |
| 2013 | Finding time period-based most frequent path in big trajectory dataabstractThe rise of GPS-equipped mobile devices has led to the emergence of big trajectory data. In this paper, we study a new path finding query which finds the most frequent path (MFP) during user-specified time periods in large-scale historical trajectory data. We refer to this query as time period-based MFP (TPMFP). Specifically, given a time period T, a source v_s and a destination v_d, TPMFP searches the MFP from v_s to v_d during T. Though there exist several proposals on defining MFP, they only consider a fixed time period. Most importantly, we find that none of them can well reflect people's common sense notion which can be described by three key properties, namely suffix-optimal (i.e., any suffix of an MFP is also an MFP), length-insensitive (i.e., MFP should not favor shorter or longer paths), and bottleneck-free (i.e., MFP should not contain infrequent edges). The TPMFP with the above properties will reveal not only common routing preferences of the past travelers, but also take the time effectiveness into consideration. Therefore, our first task is to give a TPMFP definition that satisfies the above three properties. Then, given the comprehensive TPMFP definition, our next task is to find TPMFP over huge amount of trajectory data efficiently. Particularly, we propose efficient search algorithms together with novel indexes to speed up the processing of TPMFP. To demonstrate both the effectiveness and the efficiency of our approach, we conduct extensive experiments using a real dataset containing over 11 million trajectories. Wuman Luo, Haoyu Tan, Lei Chen 0002, Lionel M. Ni |
SIGMOD Conference | 1 |
| 2012 | CloST: a hadoop-based storage system for big spatio-temporal data analyticsabstractDuring the past decade, various GPS-equipped devices have generated a tremendous amount of data with time and location information, which we refer to as big spatio-temporal data. In this paper, we present the design and implementation of CloST, a scalable big spatio-temporal data storage system to support data analytics using Hadoop. The main objective of CloST is to avoid scan the whole dataset when a spatio-temporal range is given. To this end, we propose a novel data model which has special treatments on three core attributes including an object id, a location and a time. Based on this data model, CloST hierarchically partitions data using all core attributes which enables efficient parallel processing of spatio-temporal range scans. According to the data characteristics, we devise a compact storage structure which reduces the storage size by an order of magnitude. In addition, we proposes scalable bulk loading algorithms capable of incrementally adding new data into the system. We conduct our experiments using a very large GPS log dataset and the results show that CloST has fast data loading speed, desirable scalability in query processing, as well as high data compression ratio. Haoyu Tan, Wuman Luo, Lionel M. Ni |
CIKM | 2 |
| 2012 | Efficient Similarity Joins on Massive High-Dimensional Datasets Using MapReduceabstractHigh-dimensional similarity join (HDSJ) is critical for many novel applications in the domain of mobile data management. Nowadays, performing HDSJs efficiently faces two challenges. First, the scale of datasets is increasing rapidly, making parallel computing on a scalable platform a must. Second, the dimensionality of the data can be up to hundreds or even thousands, which brings about the issue of dimensionality curse. In this paper, we address these challenges and study how to perform parallel HDSJs efficiently in the MapReduce paradigm. Particularly, we propose a cost model to demonstrate that it is important to take both communication and computation costs into account as dimensionality and data volume increases. To this end, we propose DAA (Dimension Aggregation Approximation), an efficient compression approach that can help significantly reduce both these costs when performing parallel HDSJs. Moreover, we design DAA-based parallel HDSJ algorithms which can scale up to massive data sizes and very high dimensionality. We perform extensive experiments using both synthetic and real datasets to evaluate the speedup and the scale up of our algorithms. Wuman Luo, Haoyu Tan, Huajian Mao, Lionel M. Ni |
MDM | 1 |
| 2012 | On Packing Very Large R-treesabstractMany emerging mobile applications require analyzing large spatial datasets. In these applications, efficient query processing relies on spatial access methods such as R-trees. For datasets that are fairly static, R-trees are often built as a data loading process using packing techniques. However, traditional R-tree packing algorithms can only run on a single machine and thereby cannot scale to very large datasets. In this paper, we design and implement a general framework for parallel Rtree packing using MapReduce. This framework sequentially packs each R-tree level from bottom up. For lower levels that have a large number of rectangles, we propose a partition based algorithm for parallel packing. We also discuss two spatial partitioning methods that can efficiently handle heavily skewed datasets. To evaluate the performance, we conducted extensive experiments using large real datasets. The size of the datasets is up to 100GB and the number of spatial objects is up to 2 billion. Besides range queries, k-nearest neighbor searches and spatial joins are also used for evaluation. To the best of our knowledge, it is the first work that evaluates the query performance of packed R-trees on such large datasets with spatial queries other than range queries. The results confirm the scalability of our proposed framework and parallel packing algorithms. It is also shown that our packed R-trees have good query performance and optimal space utilization. Haoyu Tan, Wuman Luo, Huajian Mao, Lionel M. Ni |
MDM | 2 |
| 2011 | MR-DBSCAN: An Efficient Parallel Density-Based Clustering Algorithm Using MapReduceabstractData clustering is an important data mining technology that plays a crucial role in numerous scientific applications. However, it is challenging due to the size of datasets has been growing rapidly to extra-large scale in the real world. Meanwhile, MapReduce is a desirable parallel programming platform that is widely applied in kinds of data process fields. In this paper, we propose an efficient parallel density-based clustering algorithm and implement it by a 4-stages MapReduce paradigm. Furthermore, we adopt a quick partitioning strategy for large scale non-indexed data. We study the metric of merge among bordering partitions and make optimizations on it. At last, we evaluate our work on real large scale datasets using Hadoop platform. Results reveal that the speedup and scale up of our work are very efficient. Yaobin He, Haoyu Tan, Wuman Luo, Huajian Mao, Shengzhong Feng, Jianping Fan 0002 |
ICPADS | 3 |
| 2010 | Data Vitalization: A New Paradigm for Large-Scale Dataset AnalysisabstractNowadays, datasets grow enormously both in size and complexity. One of the key issues confronted by large-scale dataset analysis is how to adapt systems to new, unprecedented query loads. Existing systems nail down the data organization scheme once and for all at the beginning of the system design, thus inevitably will see the performance goes down when user requirements change. In this paper, we propose a new paradigm, Data Vitalization, for large-scale dataset analysis. Our goal is to enable high flexibility such that the system is adaptive to complex analytical applications. Specifically, data are organized into a group of vitalized cells, each of which is a collection of data coupled with computing power. As user requirements change over time, cells evolve spontaneously to meet the potential new query loads. Besides basic functionality of Data Vitalization, we also explore an envisioned architecture of Data Vitalization including possible approaches for query processing, data evolution, as well as its tight-coupled mechanism for data storage and computing. Zhang Xiong 0001, Wuman Luo, Lei Chen 0002, Lionel M. Ni |
ICPADS | 2 |
| 2008 | Multi-path GEM for Routing in Wireless Sensor Networks
Qiang Ye 0001, Yuxing Huang, Andrew Reddin, Lei Wang 0126, Wuman Luo |
WASA | 5 |