VLDB 2026 Research / reviewers in the wild / expert
Meng Ma 0001
dblp:15/4248-1
· DBLP profile ↗
45ranked-venue papers
13as first author
27since 2021 · last 2026
0000-0002-1963-2513ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 11 · 1 first-author · 8 since 2021Databases, data management, data science and information retrieval · 8 · 2 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 3 first-author · 4 since 2021Systems, architecture and hardware · 6 · 2 first-author · 2 since 2021Software engineering, systems software and programming languages · 6 · 2 first-author · 4 since 2021Computer networks · 4 · 2 first-authorGraphics, computer vision, multimedia, augmented reality and games · 4 · 4 since 2021Security and privacy · 3 · 1 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | CAVIAR: Disentangling Root Causes with an ICA-based VAE for Large-Scale Microservice SystemsabstractMicroservice architectures in modern software engineering generate vast quantities of heterogeneous metrics, making fault diagnosis notoriously difficult. Conventional root cause analysis (RCA) methods often struggle with high-dimensional, diverse data where only a small subset of metrics may truly drive the observed failures. In this paper, we propose CAVIAR (Causality-based Analysis via VAE and ICA for Anomaly Root-cause), a two-phase framework for interpretable RCA in large-scale microservice systems. First, we train a variational autoencoder (VAE) enhanced with Independent Component Analysis (ICA) principles to learn a low-dimensional, disentangled representation of normal microservice operation. By enforcing independence among latent variables, we discover semantically coherent factors, such as specific service loads or network-level conditions. Second, when a fault occurs, we treat anomalies as external interventions on some latent factor and optimize an interventional matrix to identify the culprit dimension. This factor is then mapped back to the original metrics for actionable diagnostics. Xinrui Jiang 0001, Tingzhu Bi, Meng Ma 0001, Ping Wang 0003 |
KDD (1) | 3 |
| 2026 | PowerCause: Leveraging Causality for Self-Recovery in Unbalanced Distribution NetworksabstractThis article proposes PowerCause, a method for phase unbalance positioning and active regulation in smart distribution networks. PowerCause can automatically detect anomaly intervals, locate the source of phase unbalance using Granger causality test, back-search, and generate a list of potential root cause buses. It takes regulation measures for specific buses to alleviate the impact of the unbalance. This implements a closed-loop solution that handles the entire process from the occurrence of phase unbalance to its positioning, and finally, to active regulation. This study builds a closed-loop simulation environment based on the open distribution system simulator (OpenDSS) to enable autonomous and controllable unbalance injection, collect multidimensional bus metrics, including voltage and phase angle. The environment also includes causal analysis and active regulation modules for method verification. The proposed method demonstrates high accuracy in root cause location, efficient performance, and robustness against environmental influences such as measurement noise and data loss errors. Yicheng Pan 0002, Meng Ma 0001, Ping Wang 0003 |
IEEE Trans. Reliab. | 3 |
| 2025 | GET-AID: Graph-Enhanced Transformer for Provenance-Based Advanced Persistent Threats Investigation and Detection
Fengyuan Xu, Jiahong Yang 0003, Wenting Li 0002, Zonghua Zhang, Chenbin Zhang, Meng Ma 0001, Ping Wang 0003 |
ESORICS (4) | 7 |
| 2025 | UnCLe: Towards Scalable Dynamic Causal Discovery in Non-linear Temporal SystemsabstractUncovering cause-effect relationships from observational time series is fundamental to understanding complex systems. While many methods infer static causal graphs, real-world systems often exhibit *dynamic causality*—where relationships evolve over time. Accurately capturing these temporal dynamics requires time-resolved causal graphs. We propose UnCLe, a novel deep learning method for scalable dynamic causal discovery. UnCLe employs a pair of Uncoupler and Recoupler networks to disentangle input time series into semantic representations and learns inter-variable dependencies via auto-regressive Dependency Matrices. It estimates dynamic causal influences by analyzing datapoint-wise prediction errors induced by temporal perturbations. Extensive experiments demonstrate that UnCLe not only outperforms state-of-the-art baselines on static causal discovery benchmarks but, more importantly, exhibits a unique capability to accurately capture and represent evolving temporal causality in both synthetic and real-world dynamic systems (e.g., human motion). UnCLe offers a promising approach for revealing the underlying, time-varying mechanisms of complex phenomena. Tingzhu Bi, Yicheng Pan 0002, Xinrui Jiang 0001, Huize Sun, Meng Ma 0001, Ping Wang 0003 |
NeurIPS | 5 |
| 2024 | G-Cause: Parameter-free Global Diagnosis for Hyperscale Web Service InfrastructuresabstractHyperscale web service infrastructures are becoming increasingly complex and facing a variety of threats, raising the demand for more sophisticated automated operations and diagnosis solutions. Existing anomaly root cause localization approaches often focus on Service-level components without drilling down to the lower-level resources where services are deployed, hindering the implementation of fine-grained failure fix measures. This paper introduces a challenging task called global diagnosis and addresses it by proposing a technique called G-Cause, which is applicable to both Service-level and host-level root cause analysis scenarios. G-Cause builds a highly adaptive diagnostic framework based on the frequency-domain and time-domain characteristics of monitoring metrics, allowing it to handle global diagnosis requirements from app to host with minimal parameter adjustments. We deploy and validate our approach in two typical scenarios: homogeneous metric diagnosis from app to microservice, and heterogeneous metric diagnosis for various host resources. The results demonstrate that G-Cause outperforms state-of-the-art diagnosis algorithms while providing strong interpretability. Our approach helps operators understand the core mechanism of anomaly propagation and adjust their management strategies more effectively. With these strengths, G-Cause successfully services our global product operations and also makes an impressive contribution in many other workflows. Xinrui Jiang 0001, Yang Zhang 0103, Tingzhu Bi, Xiangzhuang Shen, Yu Zhang 0209, Yicheng Pan 0002, Meng Ma 0001, Linlin Han, Feng Wang 0054, Ping Wang 0003 |
ICWS | 7 |
| 2024 | FaultInsight: Interpreting Hyperscale Data Center Host FaultsabstractOperating and maintaining hyperscale data centers involving millions of service hosts has been an extremely intricate task to tackle for top Internet companies.Incessant system failures cost operators countless hours of browsing through performance metrics to diagnose the underlying root cause to prevent the recurrence.Although many state-of-the-art (SOTA) methods have used time-series causal discovery to construct causal relationships among anomalous metrics, they only focus on homogeneous service-level performance metrics and fail to yield useful insights on heterogeneous host-level metrics.To address the challenge, this study presents FaultInsight, a highly interpretable deep causal host fault diagnosing framework that offers diagnostic insights from various perspectives to reduce human effort in troubleshooting.We evaluate FaultInsight using dozens of incidents collected from our production environment.FaultInsight provides markedly better root cause identification accuracy than SOTA baselines in our incident dataset.It also shows outstanding advantages in terms of deployability in real production systems.Our engineers are deeply impressed by FaultInsight's ability to interpret incidents from multiple perspectives, helping them quickly understand the mechanism behind the faults. Tingzhu Bi, Yang Zhang 0103, Yicheng Pan 0002, Yu Zhang 0209, Meng Ma 0001, Xinrui Jiang 0001, Linlin Han, Feng Wang 0054, Ping Wang 0003 |
KDD | 5 |
| 2024 | Hypergraph Multi-modal Large Language Model: Exploiting EEG and Eye-tracking Modalities to Evaluate Heterogeneous Responses for Video UnderstandingabstractUnderstanding of video creativity and content often varies among individuals, with differences in focal points and cognitive levels across different ages, experiences, and genders. There is currently a lack of research in this area, and most existing benchmarks suffer from several drawbacks: 1) a limited number of modalities and answers with restrictive length; 2) the content and scenarios within the videos are excessively monotonous, transmitting allegories and emotions that are overly simplistic. To bridge the gap to real-world applications, we introduce a large-scale Video Subjective Multi-modal Evaluation dataset, namely Video-SME. Specifically, we collected real changes in Electroencephalographic (EEG) and eye-tracking regions from different demographics while they viewed identical video content. Utilizing this multi-modal dataset, we developed tasks and protocols to analyze and evaluate the extent of cognitive understanding of video content among different users. Along with the dataset, we designed a Hypergraph Multi-modal Large Language Model (HMLLM) to explore the associations among different demographics, video elements, EEG and eye-tracking indicators. HMLLM could bridge semantic gaps across rich modalities and integrate information beyond different modalities to perform logical reasoning. Extensive experimental evaluations on Video-SME and other additional video-based generative performance benchmarks demonstrate the effectiveness of our method. The code and dataset are available at https://github.com/mininglamp-MLLM/HMLLM Anyang Su, Donglin Di, Tianyu Fu 0001, Da An, Meng Ma 0001, Kun Yan 0008, Ping Wang 0003 |
ACM Multimedia | 9 |
| 2024 | EffCause: Discover Dynamic Causal Relationships Efficiently from Time-SeriesabstractSince the proposal of Granger causality, many researchers have followed the idea and developed extensions to the original algorithm. The classic Granger causality test aims to detect the existence of the static causal relationship. Notably, a fundamental assumption underlying most previous studies is the stationarity of causality, which requires the causality between variables to keep stable. However, this study argues that it is easy to break in real-world scenarios. Fortunately, our paper presents an essential observation: if we consider a sufficiently short window when discovering the rapidly changing causalities, they will keep approximately static and thus can be detected using the static way correctly. In light of this, we develop EffCause, bringing dynamics into classic Granger causality. Specifically, to efficiently examine the causalities on different sliding window lengths, we design two optimization schemes in EffCause and demonstrate the advantage of EffCause through extensive experiments on both simulated and real-world datasets. The results validate that EffCause achieves state-of-the-art accuracy in continuous causal discovery tasks while achieving faster computation. Case studies from cloud system failure analysis and traffic flow monitoring show that EffCause effectively helps us understand real-world time-series data and solve practical problems. Yicheng Pan 0002, Yifan Zhang 0029, Xinrui Jiang 0001, Meng Ma 0001, Ping Wang 0003 |
ACM Trans. Knowl. Discov. Data | 4 |
| 2023 | CTSSeg: Consistent Teacher-Student model for magnetic resonance image SegmentationabstractSegmentation of magnetic resonance images is an essential way of measuring the volume of tissues and lesions, which can improve the efficiency of diagnosis. The mainstream image segmentation methods are based on deep learning, which requires a large amount of labeled data. However, labeling magnetic resonance images is expensive and time-consuming. Therefore, we propose a consistent teacher-student model for magnetic resonance image segmentation, which is abbreviated as CTSSeg. Specifically, the CTSSeg includes a student network and a teacher network, where the student network learns supervised from labeled data, while the teacher network utilizes unlabeled data to improve the student network via contrastive learning and pseudo-label learning. We evaluate the proposed CTSSeg on the Atrial Segmentation Challenge dataset and a local clinical dataset. The experimental results show that our method can make full use of both labeled and unlabeled data and yield state-of-the-art performance. Chenbin Zhang, Qingyuan He, Kun Yan 0008, Meng Ma 0001, Defeng Liu, Ping Wang 0003 |
ICME | 4 |
| 2023 | ECANodule: Accurate Pulmonary Nodule Detection and Segmentation with Efficient Channel AttentionabstractAccurate detection and segmentation of pulmonary nodules in low-dose CT images is essential for early screening and treatment of lung cancer. Previous methods have often overlooked the critical role of segmentation in nodule feature learning, relying on relatively simple region proposal networks and false positive reduction modules. To address this limitation’ we introduce an segmentation branch to fully utilize the additional information such as nodule shape and boundary. Our proposed 3D U-Net detection model based on multi-task learning is optimized through bottom-layer parameter sharing to enhance prediction performance by fully utilizing complementary information between tasks. As for challenging problem of large nodule scale variety and complex background, we add more skip connections between the encoder and decoder structures, enhancing the fusion of features from different levels and facili-tating gradient flow, thus reducing model training difficulty. We also incorporate an efficient channel attention module in residual block to improve model learning and representation capability. Our method, named ECANodule, achieves an average detection sensitivity of 91.1% and a segmentation Dice score of 83.4% on the LIDC-IDRI dataset, surpassing many previous detection methods. In addition, we provide in-depth discussions on the multi-task strategy, network structure, and channel attention mechanism, offering valuable insights for future research. Deng Luo, Qingyuan He, Meng Ma 0001, Kun Yan 0008, Defeng Liu, Ping Wang 0003 |
IJCNN | 3 |
| 2023 | A Bidirectional Tree Tagging Scheme for Joint Medical Relation ExtractionabstractJoint medical relation extraction refers to extracting triples, composed of entities and relations, from the medical text with a single model. One of the solutions is to convert this task into a sequential tagging task. However, in the existing works, the methods of representing and tagging the triples in a linear way failed to the overlapping triples, and the methods of organizing the triples as a graph faced the challenge of large computational effort. In this paper, inspired by the tree-like relation structures in the medical text, we propose a novel scheme called Bidirectional Tree Tagging (BiTT) to form the medical relation triples into two binary trees and convert the trees into a word-level tags sequence. Based on BiTT scheme, we develop a joint relation extraction model to predict the BiTT tags and further extract medical triples efficiently. Our model outperforms the best baselines by 2.0% and 2.5% in F1 score on two medical datasets. What's more, the models with our BiTT scheme also obtain promising results in three public datasets of other domains. Xukun Luo, Weijie Liu 0002, Meng Ma 0001, Ping Wang 0003 |
IJCNN | 3 |
| 2023 | Look Deep into the Microservice System Anomaly through Very Sparse LogsabstractIntensive monitoring and anomaly diagnosis have become a knotty problem for modern microservice architecture due to the dynamics of service dependency. While most previous studies rely heavily on ample monitoring metrics, we raise a fundamental but always neglected issue - the diagnostic metric integrity problem. This paper solves the problem by proposing MicroCU – a novel approach to diagnose microservice systems using very sparse API logs. We design a structure named dynamic causal curves to portray time-varying service dependencies and a temporal dynamics discovery algorithm based on Granger causal intervals. Our algorithm generates a smoother space of causal curves and designs the concept of causal unimodalization to calibrate the causality infidelities brought by missing metrics. Finally, a path search algorithm on dynamic causality graphs is proposed to pinpoint the root cause. Experiments on commercial system cases show that MicroCU outperforms many state-of-the-art approaches and reflects the superiorities of causal unimodalization to raw metric imputation. Xinrui Jiang 0001, Yicheng Pan 0002, Meng Ma 0001, Ping Wang 0003 |
WWW | 3 |
| 2023 | DyCause: Crowdsourcing to Diagnose Microservice Kernel FailureabstractToday many web applications in the cloud (apps) are built based on microservices. However, as the anomaly propagates in a highly dynamic and complex way, troubleshooting them becomes full of challenges. Existing diagnostic methods are mostly designed based on monitoring metrics retrieved from the microservice system kernel. Therefore, application owners and even site reliability engineers (SREs) cannot effectively resort to those methods when the microservice systems lack such a comprehensive monitoring infrastructure. In this article, we develop DyCause, a crowdsourcing solution to the asymmetric diagnostic information problem. Our solution collects the operational status of kernel services collaboratively from the user space and initiates diagnosis on demand. Without the requirement of any architectural or functional infrastructure, it is both fast and lightweight to deploy DyCause in a microservice system. In order to discover the fine-grained dynamic causalities between services during the anomaly, we also design an efficient algorithm based on statistical analysis. Based on this algorithm, we can also analyze the anomaly propagation paths within the microservice system and generate a better interpretable diagnosis. In our evaluation, we test DyCause in a controlled simulation environment and a real-world cloud system. Our results have shown that DyCause has the best accuracy and efficiency among several state-of-the-art methods and is more robust in terms of parameters. Yicheng Pan 0002, Meng Ma 0001, Xinrui Jiang 0001, Ping Wang 0003 |
IEEE Trans. Dependable Secur. Comput. | 2 |
| 2022 | Abnormal Situation Simulation and Dynamic Causality Discovery in Urban Traffic Networks (Short Paper)
Yicheng Pan 0002, Meng Ma 0001, Ping Wang 0003 |
COSIT | 3 |
| 2022 | When Dynamic Causality Comes to Graph-Temporal Neural NetworkabstractSpatial-temporal data forecasting is a core task in many applications, and traffic forecasting is a typical example. Researchers have proposed various methods to explore spatial and temporal characteristics to improve forecasting accuracy, including the recently emerging graph convolution networks. However, most of them only consider the road network's prior knowledge or the graph's static characteristics and thus ignore the dynamic and deeper information hidden in the data. This paper presents a novel module based on dynamic causality analysis and graph convolution to integrate statistical theories and deep learning for better capturing spatial dependencies. Then we apply the module to two specific models. In each model, we introduce the causality adjacency matrix computed by the proposed algorithm into the conventional graph convolution network to reveal the dynamic correlations between nodes in the road network. The temporal neural network is then applied to extract temporal correlations. Extensive experiments demonstrate the superiority of our method, which achieves state-of-the-art prediction accuarcy on two public transportation data sets. Yicheng Pan 0002, Meng Ma 0001, Ping Wang 0003 |
IJCNN | 3 |
| 2022 | VECROsim: A Versatile Metric-oriented Microservice Fault Simulation System (Tools and Artifact Track)abstractAutomated fault diagnosis of microservice systems has been a hot topic in recent years. As most incidents in real commercial cloud systems are not publicly available, we have witnessed researchers putting considerable effort into developing various experimental systems. However, previous tools cannot quickly refactor their functionality, scale the architecture, and customize fault characteristics. Given this, we develop VECROsim, a versatile metric-oriented microservice fault simulation system, and release the VECROsim benchmark dataset. VECROsim works delicately as a highly-customizable toolkit to generate abnormal performance metrics datasets of microservice systems on demand and automatically. Validation of representative services from the benchmark dataset confirms the capability of VECROsim to generate realistic performance metrics for diverse real-world systems. Our case studies on root cause analysis and dynamic correlation discovery demonstrated the superiority of VECROsim. We also witnessed that the VECROsim dataset brings new research challenges to state-of-the-art fault diagnosis schemes. VECROsim concretely supports microservice developers from the industry, as well as academic researchers working on fault diagnosis or broader research topics in many ways. Tingzhu Bi, Yicheng Pan 0002, Xinrui Jiang 0001, Meng Ma 0001, Ping Wang 0003 |
ISSRE | 4 |
| 2022 | Efficient Event Inference and Context-Awareness in Internet of Things Edge SystemsabstractInternet of Things (IoT) connects physical, cyber and human spaces. Event-based system is one of the cornerstones to help IoT achieve real-time monitoring, context-awareness and intelligent control. In the era of big data, the huge amount and high complexity of event inference rule pose a great challenge to traditional event-based system in its efficiency, especially resources-constrained IoT edge systems. This paper proposes a high-efficiency joint event inference model for real-time context-awareness and decision-making in IoT edge systems. We define different kinds of redundancy relations between event inference models and propose a description mechanism, named event containing graph, to support multi-pattern optimization. Three operations on single-pattern event inference models,Merge,FailureandOutputare defined respectively. The joint inference model is established by merging sharing patterns, constructing failure transitions and conditional output to eliminate inter-model redundancies. Experimental results prove that the joint model consumes less computational resources and provides higher performance than other benchmarks. It also verifies and proves that joint model has better optimization effect when processing large number of complex events. Especially in edge computing environment, joint inference model improves the real-time performance and significantly reduces the energy consumption in data transmission from edges to data center. Meng Ma 0001, Ping Wang 0003 |
IEEE Trans. Big Data | 1 |
| 2022 | ServiceRank: Root Cause Identification of Anomaly in Large-Scale Microservice ArchitecturesabstractNowadays, increasing business applications running in the cloud are embracing the microservice architecture. This article presents the challenges and implications of diagnosing root causes of anomalies in large-scale microservice architecture using real incidents in IBM Bluemix. We propose ServiceRank, a novel framework for anomaly detection and root cause identification in the microservice architecture to tackle these challenges. ServiceRank introduces an anomaly detector followed by a root cause analysis module, which detects the suspected abnormal service without pre-defined thresholds. To generalize our approach, we design a causal relationship extraction approach to construct impact graphs for root cause investigation according to specific anomalies. To eliminate cloud design-patterns' impact on anomaly diagnosis, we propose a correlation calibration mechanism in ServiceRank and present a calibration algorithm for the circuit breaker - A typical protection pattern in the microservice architecture. Finally, we design a heuristic investigation algorithm based on the second-order random walk to identify the anomaly's root cause. Experimental results in a simulated environment and the IBM Bluemix platform show that ServiceRank outperforms selected approaches in accuracy and offers fast identification of root cause service when an anomaly occurs. Moreover, we can deploy ServiceRank rapidly and easily in various systems without any pre-defined knowledge. Meng Ma 0001, Weilan Lin, Disheng Pan, Ping Wang 0003 |
IEEE Trans. Dependable Secur. Comput. | 1 |
| 2022 | DePo: Dynamically Offload Expensive Event Processing to the Edge of Cyber-Physical SystemsabstractEvent processing is one of the cornerstones to manage massive data streams in Cyber-Physical Systems (CPS). Due to CPS applications' increasing complexity, detecting highly complicated events (aka. “expensive” events) leads to significant performance degradation, particularly harmful to mission-critical systems. To tackle this challenge, we define a new task - dynamic event processing offloading to CPS-edges. This paper proves the problem NP-hard and proposes a solution -DePo.DePosplits the expensive events into sub-models and offloads them to CPS edges. We design a long and short-term event memory mechanism inDePothat enables the edges and server to process expensive events collaboratively within their capabilities. Besides, we propose a concept called Edge Utility to measure the optimality of offloading schemes. A heuristic algorithm is presented in this study to guide how to dispatch events to edges, thereby helpingDePogenerate a sub-optimal solution in polynomial computational complexity. Our extensive experiments show that the performance gap betweenDePoand the optimal benchmark is less than 5%.DePoeffectively reduces more than 40% redundant states and provides over 100% higher throughput than state-of-the-art approaches. Experimental results verified the high stability and scalability ofDePo, especially when dealing with a large number of expensive events. Meng Ma 0001, Ping Wang 0003 |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2022 | Self-Adaptive Root Cause Diagnosis for Large-Scale Microservice ArchitectureabstractThe emergence of microservice architecture in Cloud systems poses a new challenges for the reliability operation and maintenance. Due to numerous services and diverse types of metrics, it is time-consuming and challenging to identify the root cause of anomaly in large-scale microservice architecture. To solve this issue, this article presents a multi-metric and self-adaptive root cause diagnosis framework, named MS-Rank. MS-Rank decomposes the task into four phases: impact graph construction, random walk diagnosis, result precision evaluation, metrics weight update. Initially, we introduce the concept of implicit metrics and propose a composite impact graph construction algorithm, using multiple types of metrics to discover causal relationships between services. Afterwards, we propose a diagnostic algorithm in which forward, selfward and backward transitions are designed to heuristically identify the root cause services. In addition, we establish a self-adaptive mechanism to update the confidence of different metrics dynamically according to their diagnostic precision. Lastly, we develop a prototype system and integrate MS-Rank into real production system - IBM Cloud. Experimental results show that MS-Rank has a high diagnostic precision and its performance outperforms several selected benchmarks. Through multiple rounds of diagnosis, MS-Rank can optimize itself effectively. MS-Rank can be rapidly deployed in various microservice-based systems and applications, requiring no predefined knowledge. MS-Rank also allows us to introduce expert experiences into its framework to improve the diagnostic efficiency and precision. Meng Ma 0001, Weilan Lin, Disheng Pan, Ping Wang 0003 |
IEEE Trans. Serv. Comput. | 1 |
| 2021 | Hierarchical Refined Attention for Scene Text RecognitionabstractRecent years have witnessed increased interests in scene text recognition (STR). Current state-of-the-art (SOTA) approaches adopt sequence-to-sequence (Seq2Seq) structure to leverage the mutual interaction between images and textual information. However, these methods still struggle to recognize texts in arbitrary shapes. The leading cause is that it brings about information loss and negative noises when directly compressing two-dimension image features into one-dimension vectors. This paper proposes a novel framework named hierarchical refined attention network (HRAN) for STR. HRAN obtains refined representations with the hierarchical attention, which localizes the precise region of current character from two-dimension perspective. Two novel co-attention mechanisms, stacked and guided co-attention, explicitly leverage dependency between spatial-aware contextual features and region-aware visual features without extra character annotations. Experiments show that both on regular and irregular texts, HRAN achieves highly competitive performance compared to SOTA models. Meng Ma 0001, Ping Wang 0003 |
ICASSP | 2 |
| 2021 | Dyna-PTM: OD-enhanced GCN for Metro Passenger Flow PredictionabstractMetro transit is an important part of the public transportation infrastructure and provides convenience for people's daily travel. Due to the limitation of capacity, under certain conditions, such as peak hours and severe weather, the traffic of metro stations will increase rapidly and cause congestion. Precise prediction of the passenger flow guarantees the metro's stable operation and passengers' safety. Previous to our study, several models based on spatial-temporal graph convolutional networks have been designed to handle this problem. Still, most of them have not considered the Original-Destination (OD) information adequately. Some only used the metro traffic network as a station adjacency matrix to describe stations' correlation without OD information. Others treated the OD information as a static adjacency matrix. However, the matrix is actually changing over time. This paper presents a novel method that converts the time-varying OD information into dynamic probability transition matrixes to effectively extract the dynamic correlation of stations in OD information into dynamic probability transition matrixes (Dyna-PTM). Dyna-PTM is a supplement adjacency matrix in the spatial-temporal graph convolutional network to describe stations' hidden and dynamic correlation. We verify Dyna-PTM using real metro datasets collected from two megacities in China - Chongqing, and Hangzhou. Experimental results demonstrate the superior performance of our method. Congjie He, Xinrui Jiang 0001, Meng Ma 0001, Ping Wang 0003 |
IJCNN | 4 |
| 2021 | Faster, deeper, easier: crowdsourcing diagnosis of microservice kernel failure from user spaceabstractWith the widespread use of cloud-native architecture, increasing web applications (apps) choose to build on microservices. Simultaneously, troubleshooting becomes full of challenges owing to the high dynamics and complexity of anomaly propagation. Existing diagnostic methods rely heavily on monitoring metrics collected from the kernel side of microservice systems. Without a comprehensive monitoring infrastructure, application owners and even cloud operators cannot resort to these kernel-space solutions. This paper summarizes several insights on operating a top commercial cloud platform. Then, for the first time, we put forward the idea of user-space diagnosis for microservice kernel failures. To this end, we develop a crowdsourcing solution - DyCause, to resolve the asymmetric diagnostic information problem. DyCause deploys on the application side in a distributed manner. Through lightweight API log sharing, apps collect the operational status of kernel services collaboratively and initiate diagnosis on demand. Deploying DyCause is fast and lightweight as we do not have any architectural and functional requirements for the kernel. To reveal more accurate correlations from asymmetric diagnostic information, we design a novel statistical algorithm that can efficiently discover the time-varying causalities between services. This algorithm also helps us build the temporal order of the anomaly propagation. Therefore, by using DyCause, we can obtain more in-depth and interpretable diagnostic clues with limited indicators. We apply and evaluate DyCause on both a simulated test-bed and a real-world cloud system. Experimental results verify that DyCause running in the user-space outperforms several state-of-the-art algorithms running in the kernel on accuracy. Besides, DyCause shows superior advantages in terms of algorithmic efficiency and data sensitivity. Simply put, DyCause produces a significantly better result than other baselines when analyzing much fewer or sparser metrics. To conclude, DyCause is faster to act, deeper in analysis, and easier to deploy. Yicheng Pan 0002, Meng Ma 0001, Xinrui Jiang 0001, Ping Wang 0003 |
ISSTA | 2 |
| 2021 | Scene Text Recognition with Cascade Attention NetworkabstractScene text recognition (STR) has experienced increasing popularity both in academia and in industry. Regarding STR as a sequence prediction task, most state-of-the-art (SOTA) approaches employ the attention-based encoder-decoder architecture to recognize texts. However, these methods still struggle in localizing the precise alignment center associated with the current character, which is also named as the attention drift phenomenon. One major reason is that directly converting low-quality or distorted word images to sequential features may introduce confusing information and thus mislead the network. To address the problem, this paper proposes a cascade attention network. The model is composed of three novel attention modules: a vanilla attention module that attends to sequential features from the horizontal direction, a cross-network attention module to take advantage of both one-dimension contextual information and two-dimension visual distributions, and an aspects fusion attention module to fuse spatial and channel-wise information. Accordingly, the network manages to yield distinguished and refined representations correlated to the target sequence. Compared to SOTA methods, experimental results on seven benchmarks demonstrate the superiority of our framework in recognizing scene texts on various conditions. Meng Ma 0001, Ping Wang 0003 |
ICMR | 2 |
| 2021 | RAGA: Relation-Aware Graph Attention Networks for Global Entity Alignment
Renbo Zhu, Meng Ma 0001, Ping Wang 0003 |
PAKDD (1) | 2 |
| 2021 | PrePCT: Traffic congestion prediction in smart cities with relative position congestion tensor
Mengting Bai, Yangxin Lin, Meng Ma 0001, Ping Wang 0003, Lihua Duan |
Neurocomputing | 3 |
| 2021 | Middleware for the Internet of Things: A survey on requirements, enabling technologies, and solutions
Meng Ma 0001, Ping Wang 0003 |
J. Syst. Archit. | 2 |
| 2020 | Lead Time-Aware Proactive Adaptation for Service-Oriented SystemsabstractMany service-oriented systems (SoS) operate in uncertain and changing environments. Hence, SoSs should be able to adapt itself during runtime to ensure that they maintain user-expected quality indicators. In real-world environments, some adaptations may have non-negligible latency, and take some lead time to produce their effect. Adapting reactively is an after-the-fact approach, which starts when the system deviates from the expected indicators. It can result in inefficiency and instability due to without anticipating the subsequent adaptation needs. To solve this problem, we propose a novel proactive adaptation solution - LetPa, which makes decisions based on predictions about how adaptations will unfold up to its completion. LetPa divides control parameters into three levels according to the SoS architecture and rates the adaptations considering both goal satisfaction and action penalties. We design a dynamic programming based decision mechanism in LetPa that enables SoS to determine which adaptations need be performed that can prevent and mitigate upcoming problems in the near-future time series. Simulation result implies that LetPa shows good stability and efficiency in SoSs. Meng Ma 0001, Ping Wang 0003 |
ICWS | 2 |
| 2020 | AutoMAP: Diagnose Your Microservice-based Web Applications AutomaticallyabstractThe high complexity and dynamics of the microservice architecture make its application diagnosis extremely challenging. Static troubleshooting approaches may fail to obtain reliable model applies for frequently changing situations. Even if we know the calling dependency of services, we lack a more dynamic diagnosis mechanism due to the existence of indirect fault propagation. Besides, algorithm based on single metric usually fail to identify the root cause of anomaly, as single type of metric is not enough to characterize the anomalies occur in diverse services. In view of this, we design a novel tool, named AutoMAP, which enables dynamic generation of service correlations and automated diagnosis leveraging multiple types of metrics. In AutoMAP, we propose the concept of anomaly behavior graph to describe the correlations between services associated with different types of metrics. Two binary operations, as well as a similarity function on behavior graph are defined to help AutoMAP choose appropriate diagnosis metric in any particular scenario. Following the behavior graph, we design a heuristic investigation algorithm by using forward, self, and backward random walk, with an objective to identify the root cause services. To demonstrate the strengths of AutoMAP, we develop a prototype and evaluate it in both simulated environment and real-work enterprise cloud system. Experimental results clearly indicate that AutoMAP achieves over 90% precision, which significantly outperforms other selected baseline methods. AutoMAP can be quickly deployed in a variety of microservice-based systems without any system knowledge. It also supports introduction of various expert knowledge to improve accuracy. Meng Ma 0001, Jingmin Xu, Pengfei Chen 0002, Zonghua Zhang, Ping Wang 0003 |
WWW | 1 |
| 2020 | Extract interpretability-accuracy balanced rules from artificial neural networks: A review
Congjie He, Meng Ma 0001, Ping Wang 0003 |
Neurocomputing | 2 |
| 2020 | On-demand deployment for IoT applications
Meng Ma 0001, Ping Wang 0003 |
J. Syst. Archit. | 2 |
| 2019 | MS-Rank: Multi-Metric and Self-Adaptive Root Cause Diagnosis for Microservice ApplicationsabstractThis paper presents a self-adaptive root cause diagnosis framework, named MS-Rank, to analyze multiple metrics collected from micro-service architecture. MS-Rank decomposes the task into four phases: impact graph construction, random walk diagnosis, result precision calculation and metrics weight update. First, we introduce a series of basic and implied metrics into MS-Rank, and design an impact graph construction algorithm to discover causal relationship between services during anomalies. Second, we propose a random walk algorithm with forward, selfward and backward transitions to heuristically identify the root cause service. Third, we establish a self-optimizing mechanism to dynamically update the confidence weight of different metrics according to their diagnosis precision. We develop a prototype system and integrate MS-Rank into IBM Cloud, to validate and compare it with selected benchmarks. Experimental results show that MS-Rank offers fast identification and precise diagnosis result. In multiple rounds of diagnosis, MS-Rank optimizes itself effectively. Meng Ma 0001, Weilan Lin, Disheng Pan, Ping Wang 0003 |
ICWS | 1 |
| 2019 | Gate Decorator: Global Filter Pruning Method for Accelerating Deep Convolutional Neural NetworksabstractFilter pruning is one of the most effective ways to accelerate and compress convolutional neural networks (CNNs). In this work, we propose a global filter pruning algorithm called Gate Decorator, which transforms a vanilla CNN module by multiplying its output by the channel-wise scaling factors (i.e. gate). When the scaling factor is set to zero, it is equivalent to removing the corresponding filter. We use Taylor expansion to estimate the change in the loss function caused by setting the scaling factor to zero and use the estimation for the global filter importance ranking. Then we prune the network by removing those unimportant filters. After pruning, we merge all the scaling factors into its original module, so no special operations or structures are introduced. Moreover, we propose an iterative pruning framework called Tick-Tock to improve pruning accuracy. The extensive experiments demonstrate the effectiveness of our approaches. For example, we achieve the state-of-the-art pruning ratio on ResNet-56 by reducing 70% FLOPs without noticeable loss in accuracy. For ResNet-50 on ImageNet, our pruned model with 40% FLOPs reduction outperforms the baseline model by 0.31% in top-1 accuracy. Various datasets are used, including CIFAR-10, CIFAR-100, CUB-200, ImageNet ILSVRC-12 and PASCAL VOC 2011. Zhonghui You, Kun Yan 0008, Jinmian Ye, Meng Ma 0001, Ping Wang 0003 |
NeurIPS | 4 |
| 2018 | CloudRanger: Root Cause Identification for Cloud Native SystemsabstractAs more and more systems are migrating to cloud environment, the cloud native system becomes a trend. This paper presents the challenges and implications when diagnosing root causes for cloud native systems by analyzing some real incidents occurred in IBM Bluemix (a large commercial cloud). To tackle these challenges, we propose CloudRanger, a novel system dedicated for cloud native systems. To make our system more general, we propose a dynamic causal relationship analysis approach to construct impact graphs amongst applications without given the topology. A heuristic investigation algorithm based on second-order random walk is proposed to identify the culprit services which are responsible for cloud incidents. Experimental results in both simulation environment and IBM Bluemix platform show that CloudRanger outperforms some state-of-the-art approaches with a 10% improvement in accuracy. It offers a fast identification of culprit services when an anomaly occurs. Moreover, this system can be deployed rapidly and easily in multiple kinds of cloud native systems without any predefined knowledge. Ping Wang 0003, Jingmin Xu, Meng Ma 0001, Weilan Lin, Disheng Pan, Pengfei Chen 0002 |
CCGrid | 3 |
| 2018 | FacGraph: Frequent Anomaly Correlation Graph Mining for Root Cause Diagnose in Micro-Service ArchitectureabstractMicro-service architecture is a promising paradigm to develop, deploy and maintain applications using independent and autonomous cloud services. Nowadays, increasingly applications are embracing this model. However, it is difficult and time-consuming to diagnose and identify the actual root cause when anomalies occurs in micro-service architecture due to various factors. This paper introduces a novel framework for anomaly investigation and root cause identification in micro-service architecture. The novelty in our work lies on: (1) Different from existing solutions, in our framework, we propose a frequent pattern mining algorithm on anomaly correlation graph, named FacGraph, to discover root cause services. (2) We leverage breadth first ordered string (BFOS) to reduce the time-consumption of the frequent graph mining (FSM) (3) We further develop a distributed version of FacGraph to improve its paralleled computing efficiency. We evaluate our framework in real production environment IBM Bluemix. Result demonstrate that FacGraph outperforms other methods in diagnosis accuracy and offers a fast identification of root cause service when an anomaly occurs. Weilan Lin, Meng Ma 0001, Disheng Pan, Ping Wang 0003 |
IPCCC | 2 |
| 2018 | Redundant Reader Elimination in Large-Scale Distributed RFID NetworksabstractRadio frequency identification (RFID) is a key enabling technology for large-scale Internet of Things (IoT). Its deployment and management impact significantly on the operational effectiveness and scalability of IoT. Redundant reader elimination is of great importance to reduce the system's overhead and prolong the lifetime of RFID networks. It helps to reduce unnecessary reader-tag interactions and the cost of collision avoidance algorithm in distributed data collection of RFID network. In large-scale distributed RFID networks, one of the most challenging tasks for redundant reader elimination is to improve the performance of distributed algorithms. This paper proposes a novel distributed redundant reader elimination algorithm based on the threshold selection process, named threshold selection algorithm (TSA), for RFID networks. TSA algorithm selects effective reader iteratively based on the threshold sequence determined by expected tag coverage. This paper also introduces an optimization mechanism into TSA based on detected movement, named TSA with movement detection algorithm. By preliminary simulation, we determine the suggested parameter of linear multiplier for TSA. Our experiments show that TSA algorithm can provide 30%-60% better performance than other major distributed algorithms and also multiphase schemes in both dense and sparse environments. The overhead of TSA algorithm is 30%-50% lower than other selected algorithms, especially in tag-write operations. Meng Ma 0001, Ping Wang 0003, Chao-Hsien Chu |
IEEE Internet Things J. | 1 |
| 2018 | Long-Term Event Processing over Data Streams in Cyber-Physical SystemsabstractEvent processing is a crucial cornerstone supporting the revolution of Internet of Things (IoT) and Cyber-Physical Systems (CPS) by integrating physical-layer networking and providing intelligent computation and real-time control abilities. In various IoT and CPS application scenarios, the event processing systems are required to detect complex event patterns using large time window, namely long-term events. The detection of long-term event usually leads to a large number of redundant runtime instances and calculations that significantly deteriorates the system efficiency. In this article, we propose an efficient long-term event processing model, named Long-Term Complex Event Processing (LTCEP). It leverages the semantic constraints calculus to split long-term event into sub-models. We establish a long-term query and intermediate result buffering mechanism to optimize the real-time response ability and throughput performance. Experimental results show that LTCEP can effectively reduce more than 50% redundant runtime states, which provides over 60% faster response performance and around 30% higher system throughput comparing to other selected benchmarks. The results also imply that LTCEP model has better stability and scalability in large-scale event processing applications. Ping Wang 0003, Meng Ma 0001, Chao-Hsien Chu |
ACM Trans. Cyber Phys. Syst. | 2 |
| 2018 | Toward Energy-Awareness Smart Building: Discover the Fingerprint of Your Electrical AppliancesabstractEnergy efficiency raises significant concerns as it is one of the most promising ways to mitigate climate change. Disaggregation and identification of individual electrical appliances activities are one of the essentials for energy preservation especially for smart buildings. This paper proposes a lightweight electrical appliance activity detection approach for smart building, which leverages a single smart metering device to establish a learning and detection processing for multiple appliances. In this system, data interpolation and transition detection algorithm are proposed to effectively reduce the cost of model training and optimize the detection accuracy. The concept of appliance fingerprint is proposed and a variety of fingerprints, including appliance-based and context-based, are defined to depict fine-grained appliance characteristics. Based on these fingerprints, the paper proposes a multisource fingerprint-weighting KNN (FWKNN) classification algorithm and presents a boosting framework for continuous online learning and detection. A prototype system is implemented and demonstrated in IBM Bluemix PaaS cloud platform. Experimental result and analysis prove that FWKNN outperforms other benchmark methods in detection accuracy. Meng Ma 0001, Weilan Lin, Ping Wang 0003, Xiaoxing Liang |
IEEE Trans. Ind. Informatics | 1 |
| 2017 | MODE: A Context-Aware IoT Middleware Supporting On-Demand Deployment for Mobile DevicesabstractWith the development of Internet of Things (IoT), various mobile sensing devices emerge in the market, which brings great convenience to people's life. Middleware, as the connection platform of sensors and applications, has become increasingly important to solve the problem caused by diverse devices and changing environments. In this paper, we proposed MODE, a middleware that can dynamically change its deployment of function modules based on context awareness to adapt to environment changes. As a highly scalable middleware, MODE not only has basic tasks, but also provides developers with user-specific tasks, based on a highly-abstracted scripting language, which supports the application development in different scenarios and improves the efficiency. Besides, we design various function libraries in MODE that developers can load or unload them dynamically. MODE improves the efficiency, adaptability and scalability of IoT systems and applications. Meng Ma 0001, Ping Wang 0003 |
ICPADS | 3 |
| 2017 | Event Description and Detection in Cyber-Physical Systems: An Ontology-Based Language and ApproachabstractIn this paper, we propose an ontology-based language, OntoEvent, for semantic complex event modeling and detection in Cyber-Physical Systems (CPS). We divide the core concepts of OntoEvent model into two levels: general concepts and domain-specific instances, to promise it can be dynamic extended for different CPS application domains. Complex events are modeled based on event ontology with logical and temporal operators, and these operators are extended by nature language synonymies. OntoEvent language is of rich expressiveness compared to other traditional languages. Based on OntoEvent, We propose an event detection model and elaborate its construction procedures. Experimental results prove that OntoEvent-based event detection model outperforms other selected models in processing efficiency, especially when processing multiple complex event ontologies. Meng Ma 0001, Yangxin Lin, Disheng Pan, Ping Wang 0003 |
ICPADS | 1 |
| 2017 | On the consistency of event processing: A semantic approach
Meng Ma 0001, Ping Wang 0003 |
Knowl. Based Syst. | 1 |
| 2015 | Class-based delta-encoding for high-speed train data streamabstractRailway transportation plays an important role in both economic and social development. The requirements of the railway traffic increase in recent decades. In order to meet the growing demand, a new generation control system of railway transportation emerges. It consists of collection, transmission, analysis and scheduling module. In such a context, an information transmission system is built to connect trains and scheduling center. However, the infrastructure of the railway system cannot provide enough bandwidth for such amount of data. As a result, the efficiency of data transmission cannot be ensured. In this paper, we focus on the compression algorithm that reduce the amount of transmitted data and improve the system performance. Based on the analysis of the common algorithms, an efficient compression algorithm, named delta-encoding, is proposed. It consists of two steps: preprocessing and compression. Delta-encoding utilizes a class-based difference model, which reduces the data redundancy, to realize a preprocessing algorithm. With the combination of preprocessing algorithm and a regular compression algorithm, delta-encoding has better performance on compression ratio, and becomes a universal hybrid algorithm for structured data in IoT system rather than a specific algorithm in high-speed train system. Finally, several experiments are provided to prove that delta-encoding have advantages in both compression ratio and compression time. Yangxin Lin, Ping Wang 0003, Jinlong Lin, Meng Ma 0001 |
IPCCC | 4 |
| 2015 | OntoEvent: An Ontology-Based Event Description Language for Semantic Complex Event Processing
Meng Ma 0001, Ping Wang 0003 |
WAIM | 1 |
| 2015 | Efficient Multipattern Event Processing Over High-Speed Train Data StreamsabstractBig data is becoming a key basis for productivity growth, innovation, and consumer surplus, but also bring us great challenges in its volume, velocity, variety, value, and veracity. The notion of event is an important cornerstone to manage big data. High-speed railway is one of the most typical application domains for event-based system, especially for the train onboard system. There are usually numerous complex event patterns subscribed in system sharing the same prefix, suffix, or subpattern; consequently, multipattern complex event detection often results in plenty of redundant detection operations and computations. In this paper, we propose a multipattern complex event detection model, multipattern event processing (MPEP), constructed by three parts: 1) multipattern state transition; 2) failure transition; and 3) state output. Based on MPEP, an intelligent onboard system for high-speed train is preliminarily implemented. The system logic is described using our proposed complex event description model and compiled into a multipattern event detection model. Experimental results show that MPEP can effectively optimize the complex event detection process and improve its throughput by eliminating duplicate automata states and redundant computations. This intelligent onboard system also provides better detection ability than other models when processing real-time events stored in high-speed train Juridical Recording Unit (JRU). Meng Ma 0001, Ping Wang 0003, Chao-Hsien Chu |
IEEE Internet Things J. | 1 |
| 2012 | User-driven cloud transportation system for smart drivingabstractIntelligent transportation systems (ITS) have emerged as an efficient and effective way of alleviating the traffic congestion and improving the performance of transportation systems. Key challenges of ITS in recent years include the pervasive data collection, data security, privacy preserving, large volume data processing, and intelligent analytics. These challenges lead to a revolution in ITS development by leveraging the crowdsourcing scheme and cloud computing architecture. In this paper, we propose a user-driven Cloud Transportation system (CTS) which employs a scheme of user-driven crowdsourcing to collect user data for traffic model construction and congestion prediction including data collection, filtering, modeling, intelligent computation and publish. We describe in details the application scenario, system architecture, and core CTS services model. To verify the feasibility of our approach, we have developed a prototype system which elaborated the cloud architecture and other implementation details. This paper aims to inspire further research of user-driven CTS on intelligent data processing model for smarter utilization of transportation infrastructure. Meng Ma 0001, Chao-Hsien Chu, Ping Wang 0003 |
CloudCom | 1 |