EDBT 2026 Demo / reviewers in the wild / expert
Long Cheng 0003
dblp:49/225-3
· DBLP profile ↗
85ranked-venue papers
26as first author
59since 2021 · last 2026
0000-0003-1638-059XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 37 · 18 first-author · 20 since 2021Computer networks · 15 · 1 first-author · 14 since 2021Artificial intelligence and machine learning · 11 · 2 first-author · 9 since 2021Software engineering, systems software and programming languages · 11 · 1 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 9 · 2 first-author · 7 since 2021Databases, data management, data science and information retrieval · 5 · 2 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Deep reinforcement learning for energy-efficient workflow scheduling in edge computing
Mengyao Wen, Xiufeng Liu 0001, Xin Ning 0001, Cong Liu 0012, Jiawei Nian, Long Cheng 0003 |
Comput. Networks | 7 |
| 2026 | HybridEditDif: Text and exemplar guided image editing with diffusion models
Xuemei Fu, Long Cheng 0003, Jungong Han, Catarina Moreira, Xin Ning 0001, Xiao Bai 0001 |
Pattern Recognit. | 4 |
| 2026 | HydraPIM: A Heterogeneous PIM Architecture for Efficient Attention in Long-Context LLMs
Xiangwen An, Yutian Zhou, Yintao He, Long Cheng 0003, Mengdi Wang 0004, Xiaowei Li 0001, Yinhe Han 0001, Ying Wang 0001 |
IEEE Trans. Computers | 6 |
| 2026 | Prototype Retrieval-Augmented Federated Learning System for Robust Intrusion DetectionabstractDetecting malicious attacks is essential for protecting computer systems and ensuring device security. Federated Learning (FL)-based Intrusion Detection Systems (IDS) have emerged as promising solutions, enabling multiple clients (i.e., data owners) to collaboratively train intrusion detection models without sharing private data. However, current FL studies typically assume that each client’s training and test label distribution is identical. This assumption is overly idealistic and rarely holds in real-world scenarios, leading to suboptimal performance when label distribution shifts occur between the training and testing data. To address this challenge, we propose FedPRO, a plug-and-play framework designed to improve the test-time performance of existing FL methods, without modifying their original training pipelines or fine-tuning the trained FL models. Specifically, we develop a unique prototype generation and optimization mechanism to produce semantically meaningful class prototypes. These prototypes constitute a prototype memory bank, serving as an external knowledge repository. At test time, a prototype retrieval-augmented inference strategy is employed to query relevant prototypes and refine predictions on each client, effectively alleviating the label distribution shift issues and boosting prediction accuracy. We evaluate FedPRO by integrating it with various off-the-shelf FL methods on benchmark datasets. Extensive results consistently demonstrate its effectiveness in diverse settings. Notably, applying FedPRO to the state-of-the art method FedDBE improves its test accuracy from 79.25% to 86.66% on the CICIDS-2018 dataset, while introducing only approximately 32KB of additional communication overhead. Hanlin Zhou, Huiru Yan, Jiawei Nian, Cong Liu 0012, Ying Wang 0001, Georgios Theodoropoulos 0001, Long Cheng 0003 |
IEEE Trans. Computers | 8 |
| 2026 | MCU-MixQ: A HW/SW Co-optimized Mixed-precision Neural Network Design Framework for MCUsabstractMixed-precision neural network (MPNN) that utilizes just enough data width for the neural network processing is an effective approach to meet the stringent resources constraints including memory and computing of MCUs. Nevertheless, there is still a lack of sub-byte and mixed-precision SIMD operations in MCU-class ISA and the limited computing capability of MCUs remains underutilized, which further aggravates the computing bound encountered in neural network processing. As a result, the benefits of MPNNs cannot be fully unleashed. In this work, we propose to pack multiple low-bitwidth arithmetic operations within a single instruction multiple data (SIMD) instructions in typical MCUs, and then develop an efficient convolution operator by exploring both the data parallelism and computing parallelism in convolution along with the proposed SIMD packing. Finally, we further leverage Neural Architecture Search (NAS) to build a HW/SW co-designed MPNN design framework, namely MCU-MixQ. This framework can optimize both the MPNN quantization and MPNN implementation efficiency, striking an optimized balance between neural network performance and accuracy. According to our experiment results, MCU-MixQ achieves 2.1× and 1.4× speedup over CMix-NN and MCUNet respectively under the same resource constraints. MCU-MixQ is also open sourced on GitHub. 1 Junfeng Gong, Long Cheng 0003, Jiawei Nian, Cheng Liu 0008, Huawei Li 0001 |
ACM Trans. Embed. Comput. Syst. | 2 |
| 2026 | A Hierarchical GNN-Based Multi-Agent Framework for Workflow Scheduling in Hybrid Clouds Considering Privacy Constraints
Hanlin Zhou, Cong Liu 0012, Fang Fang 0007, Zhiming Zhao, Georgios Theodoropoulos 0001, Long Cheng 0003 |
IEEE Trans. Serv. Comput. | 7 |
| 2026 | Detecting Root Causes for Process Performance Anomalies Using Causal InferenceabstractProcess execution time is a key performance indicator for evaluating bottlenecks in business processes. Cases and activities that exceed the specified time constraints can be seen as anomalies, affecting process performance and leading to risks such as delays and customer complaints. Identifying the root causes of these anomalies can help formulate effective intervention measures. However, this task is inherently complex, and conducting incomplete or inaccurate analysis can result in misguided interventions that inadvertently exacerbate process inefficiencies. To address these challenges, this paper proposes a traceability-based root cause analysis approach for process performance anomalies using causal inference. Specifically, the approach begins by extracting hidden contextual information from the event log to enrich the pool of potential causal factors. Then formulates causal hypotheses linking these factors to observed performance anomalies (at both the case and activity level) and establishes potential causal relations through a traceability mechanism. A meta-learning based causal inference approach is used to estimate the strength of causal effects. The proposed approach is evaluated against a state-of-the-art approach using four synthetic event logs with known root causes and nine public real-life event logs. Experimental results demonstrate that the proposed approach delivers accurate insights into the root causes of process performance anomalies in synthetic event logs, while maintaining high efficiency in the comprehensive analysis of potential causal factors. Cong Liu 0012, Qingtian Zeng, Youxi Wu, Jinglin Zhang 0001, Xixi Lu 0001, Long Cheng 0003 |
IEEE Trans. Serv. Comput. | 7 |
| 2026 | Toward Efficient Support for Business Process Event Log SamplingabstractLarge volumes of event logs have been accumulated by business information systems. Accompanied by that, various process discovery techniques are invented to uncover underlying business processes based on event logs. Event log sampling, recognized as one of the most effective techniques for accelerating discovery efficiency, has gained significant attention in recent days. However, achieving high performance in sampling while maintaining superior sample log quality remains a challenge for current techniques. To tackle the problem, a novel event log sampling technique, denoted assigRank, is introduced to improve both the sampling efficiency and the quality of the sample log by quantifying the significance of each trace. The proposed sampling technique has been implemented as a publicly available tool in the open-source process mining platform ProM. Compared with state-of-the-art techniques using 12 public event logs, we experimentally illustrate that the proposed approach can significantly accelerate sampling efficiency while guaranteeing superior sample log quality for process discovery. Xuan Su, Cong Liu 0012, Shuaipeng Zhang, Qingtian Zeng, Long Cheng 0003 |
IEEE Trans. Serv. Comput. | 6 |
| 2026 | Enhancing Process Discovery by Optimizing Imprecise Sub-ProcessesabstractProcess discovery aims to derive a process model that accurately represents the observed behavior in an event log. As a state-of-the-art process discovery technique, Inductive Miner (IM) generates sound process models (i.e., free of deadlocks) while ensuring optimal replay fitness. However, IM may sometimes produce over-generalized process models with locally imprecise structures, often resulting in the creation of so-called flower structures. To address this limitation, this paper presents a novel technique that refines the process model generated by IM by optimizing its imprecise sub-processes. Specifically, the technique begins by identifying and extracting sub-logs corresponding to imprecise sub-processes in the initial IM-generated process model. Then, these imprecise sub-processes are iteratively optimized using a frequency-based filtering mechanism applied to the sub-logs. Once optimized, the imprecise sub-processes in the initial process model are replaced by the optimized ones, generating a set of candidates process models. Finally, the candidate with the best quality, in terms of fitness and precision, is selected as the final optimized process models. The proposed technique has been implemented as a plugin for the open-source process mining platform ProM. Through comparisons with state-of-the-art process discovery techniques using 10 publicly available real-life event logs, the experimental results demonstrate that the proposed method achieves an average absolute improvement of 0.173 in F-measure over its IMi variant, while also exhibiting competitive performance relative to other state-of-the-art approaches. Jiaxin Yan, Cong Liu 0012, Qingtian Zeng, Jian Cao 0001, Youxi Wu, Chun Ouyang 0001, Long Cheng 0003 |
IEEE Trans. Serv. Comput. | 7 |
| 2025 | COMET: Towards Practical W4A4KV4 LLMs ServingabstractQuantization is a widely-used compression technology to reduce the overhead of serving large language models (LLMs) on terminal devices and in cloud data centers. However, prevalent quantization methods, such as 8-bit weight-activation or 4-bit weight-only quantization, achieve limited performance improvements due to poor support for low-precision (e.g., 4-bit) activation. This work, for the first time, realizes practical W4A4KV4 serving for LLMs, fully utilizing the INT4 tensor cores on modern GPUs and reducing the memory bottleneck caused by the KV cache. Specifically, we propose a novel fine-grained mixed-precision quantization algorithm (FMPQ) that compresses most activations into 4-bit with negligible accuracy loss. To support mixed-precision matrix multiplication for W4A4 and W4A8, we develop a highly optimized W4Ax kernel. Our approach introduces a novel mixed-precision data layout to facilitate access and fast dequantization for activation and weight tensors, utilizing the GPU's software pipeline to hide the overhead of data loading and conversion. Additionally, we propose fine-grained streaming multiprocessor (SM) scheduling to achieve load balance across different SMs. We integrate the optimized W4Ax kernel into our inference framework, COMET, and provide efficient management to support popular LLMs such as LLaMA-3-70B. Extensive evaluations demonstrate that, when running LLaMA family models on a single A100-80G-SMX4, COMET achieves a kernel-level speedup of 2.88x over cuBLAS and a 2.02x throughput improvement compared to TensorRT-LLM from an end-to-end framework perspective. Long Cheng 0003, Haimeng Ren, Zhaohui Xu, Yudong Pan, Mengdi Wang 0004, Xiaowei Li 0001, Yinhe Han 0001, Ying Wang 0001 |
ASPLOS (2) | 2 |
| 2025 | EdgeIM: An Efficient Edge-Based Process Model Discovery TechniqueabstractThe rapid expansion of Internet of Things (IoT) devices has led to an explosion of event data, posing significant challenges for traditional process model discovery techniques in terms of scalability and discovery accuracy. These techniques rely on centralized storage and processing, which are hindered by data transfer limitations, storage capacity, and computational overhead in distributed IoT environments. Edge-based model discovery techniques offer a promising solution for analyzing large-scale IoT data. However, existing techniques suffer from low efficiency and an inability to handle complex process structures. To address these challenges, we propose EdgeIM, an efficient edge-based process model discovery technique that enhances efficiency and model accuracy. EdgeIM operates in three key stages: preprocessing and feature-preserving sampling to eliminate redundant data, local processing at edge nodes to extract key structural features, and global feature aggregation at a central node for model discovery. EdgeIM has been implemented on the open-source process mining platform PM4Py, and experimental results on nine public event logs demonstrate that, compared to existing edge-based model discovery techniques, EdgeIM significantly improves discovery efficiency while maintaining high model quality. Xuan Su, Cong Liu 0012, Faming Lu, Long Cheng 0003, Qingtian Zeng, Shouli Zhang |
ICWS | 4 |
| 2025 | Enhancing Healthcare Process Model Discovery Through Duplicate Task Identification
Xuan Su, Cong Liu 0012, Faming Lu, Long Cheng 0003, Qingtian Zeng, Jiehan Zhou |
ICWS | 4 |
| 2025 | Enhancing Manufacturing Process Discovery Through Sub-Process OptimizationabstractManufacturing process discovery extracts insights from event logs recorded by Manufacturing Information Systems (MISs) to optimize operational processes. However existing process discovery techniques struggle with complex concurrency relations, resulting in imprecise sub-processes that compromise model accuracy. This paper proposes a novel enhancement to Inductive Miner (IM)-generated models by optimizing local imprecise structures in manufacturing process models. The method first identifies imprecise sub-processes and extracts their corresponding sub-logs. Then imprecise sub-processes are incrementally optimized using a frequency-based filter mechanism, generating multiple candidate models. Finally, the best-quality candidate model based on evaluation metrics is selected as the final output. The proposed technique has been implemented as an open source process mining toolkit ProM plugin and evaluated on six real-life manufacturing event logs. Experimental results demonstrate that it outperforms state-of-the-art techniques, producing higher quality process models, making it particularly suited for manufacturing process discovery. Jiaxin Yan, Cong Liu 0012, Long Cheng 0003, Jiujun Cheng, Weijian Ni, Qingtian Zeng |
ICWS | 3 |
| 2025 | Job Scheduling in Hybrid Clouds With Privacy Constraints: A Deep Reinforcement Learning ApproachabstractABSTRACT With the proliferation of cloud computing and the escalating demand for extensive data processing capabilities, an increasing number of enterprises are embracing hybrid cloud solutions. However, as more businesses move toward hybrid clouds, the need for effective solutions to privacy and security concerns becomes increasingly important. Although current scheduling approaches for cloud computing have addressed privacy protection to some extent, few have adequately considered the unique challenges posed by hybrid clouds. To address this gap, we propose a novel approach for scheduling jobs in hybrid clouds that prioritizes privacy protection. Our approach, called PH‐DRL, leverages Deep Reinforcement Learning (DRL) to intelligently allocate jobs to virtual machines, optimizing both privacy and Quality of Service (QoS), while minimizing response time. We present the detailed implementation of our approach and our experimental results demonstrate the superior performance of PH‐DRL in terms of privacy protection compared to existing methods. Haoyang He, Qingzhi Liu, Hao Wu 0017, Long Cheng 0003 |
Concurr. Comput. Pract. Exp. | 5 |
| 2025 | Real-time workflow scheduling in hybrid clouds with privacy and security constraints: A deep reinforcement learning approach
Haoyang He, Yang Hu 0009, Fang Fang 0007, Xin Ning 0001, Long Cheng 0003 |
Expert Syst. Appl. | 7 |
| 2025 | Dual Guidance Enabled Fuzzy Inference for Enhanced Fine-Grained RecognitionabstractIn the field of fine-grained visual recognition (FGVR), the ability to resolve minute and often subtle differences between highly similar object categories is paramount. The advent of vision transformers (ViTs) has marked a significant advancement in this domain, primarily due to their capacity to model the intricate interdependencies among object parts represented as image patches. However, their inherent single-scale processing limitation hampers their effectiveness in FGVR tasks. Furthermore, the challenge of uncertainty inherent in FGVR tasks remains unresolved, necessitating the development of methods that bolster the robustness of these models, particularly across varying scales of visual features. We introduce a new plug-in module that can be seamlessly integrated into ViT, called dual guidance enabled fuzzy inference (DGEFI), which combines fuzzy inference with dual guidance mechanisms. Dual guidance includes scale-aware guidance and probability guidance. The former strengthens the model's focus on salient scales, and the latter refines the distinction between similar categories by optimizing intraclass compactness and interclass separability. Fuzzy inference enables the model to adaptively tweak the influence of distinct scales in the final decision-making phase, thereby enhancing the overall accuracy of recognition tasks. We demonstrate the versatility and efficacy of our DGEFI module by integrating it into several leading ViT backbones, including ViT, Swin, Mvitv2, and EVA-02. Empirical results exhibit exceptional performance gains, with the integration of DGEFI into EVA-02 remarkable accuracy improvements, reaching 93.6% on the CUB-200-2011 dataset and 94.5% on the NA-Birds dataset, respectively, improving over the state-of-the-art method 0.5% and 1.5%. Qiupu Chen, Feng He 0008, Gang Wang 0023, Xiao Bai 0001, Long Cheng 0003, Xin Ning 0001 |
IEEE Trans. Fuzzy Syst. | 5 |
| 2024 | MARS: Multi-Agent Deep Reinforcement Learning for Real-Time Workflow Scheduling in Hybrid Clouds with Privacy ProtectionabstractScheduling workflows in hybrid cloud environments presents significant challenges due to the inherent complexity of workflows and the dynamic nature of cloud resources. This complexity is further increased when attempting to balance workflow performance with privacy protection. Recent efforts have leveraged deep reinforcement learning (DRL) to address these challenges. However, most of these approaches rely on single-agent models, which can lead to security issues and scalability problems due to their centralized processing. Specifically, the properties of workflows are transferred to the single agent, which risks leaking privacy information. Our paper addresses these issues by introducing MARS, a real-time workflow scheduling method that prioritizes privacy protection in hybrid clouds. MARS leverages multi-agent deep reinforcement learning (MADRL) to optimize the workflow scheduling of cloud virtual machines (VMs). The benefit of our solution is that it relies on the collaborative learning of multi-agents on multiple VMs, which could assign user data to specific cloud servers for privacy protection while sharing training experiences between agents. In our implementation, MARS aims to reduce workflow completion time and operational costs while complying with strict privacy protection guidelines. The experimental results demonstrate that MARS can significantly surpass existing methods, reducing makespan by an average of $53.18 \%$ and costs by $61.98 \%$ compared to basic techniques, and achieving $20.26 \%$ and $25.71 \%$ improvements over the latest advanced methods, respectively. Long Cheng 0003, Haoyang He, Qingzhi Liu, Zhiming Zhao, Fang Fang 0007 |
ICPADS | 1 |
| 2024 | CoTV: Cooperative Control for Traffic Light Signals and Connected Autonomous Vehicles Using Deep Reinforcement LearningabstractThe target of reducing travel time only is insufficient to support the development of future smart transportation systems. To align with the United Nations Sustainable Development Goals (UN-SDG), a further reduction in fuel consumption and emissions, improvements in traffic safety, and the ease of infrastructure deployment and maintenance should also be considered. Most existing research in sustainable urban traffic control adjusts either traffic light signals or vehicle speed. Adaptive traffic light signal control can increase the intersection throughput and reduce travel time as well as energy consumption and emissions. Connected Autonomous Vehicles (CAVs) can proactively control vehicle acceleration to achieve more stable traffic nearby with relatively higher driving velocity (i.e., lower fuel consumption and CO 2 emissions) and maintain a safe distance from the surrounding traffic (i.e., longer time-to-collision). Jiaying Guo, Long Cheng 0003, Shen Wang 0006 |
IV | 2 |
| 2024 | BD-TTS: A blockchain and DRL-based framework for trusted task scheduling in edge computing
Hengyang Zhang, Shike Li, Long Cheng 0003, Yiguo Guo, Sixing Wu |
Comput. Networks | 4 |
| 2024 | Approximate data mapping in refresh-free DRAM for energy-efficient computing in modern mobile systems
Yingke Gao, Ying Wang 0001, Shuhong Dai, Yongjun Xu 0001, Long Cheng 0003 |
Comput. Commun. | 7 |
| 2024 | Preface of special issue on Artificial Intelligence for time-critical computing systems
Long Cheng 0003, Zhiming Zhao |
Future Gener. Comput. Syst. | 1 |
| 2024 | Deep Reinforcement Learning for Efficient IoT Data Compression in Smart Railroad ManagementabstractModern smart railroad management relies heavily on Internet of Things (IoT)-enabled sensors to monitor train performance, resulting in the generation of extensive data streams. The ensuing data influx poses critical challenges in storage, processing, and transmission. This paper presents a novel data compression method specifically designed for IoT-generated data within railroad management scenarios. Leveraging deep reinforcement learning (DRL), our approach intelligently compresses data from onboard IoT sensors. This not only streamlines data streams, ensuring essential information is retained and redundant data is pruned, but also alleviates strain on resource-limited devices such as edge servers tasked with complex computations or aggregating large datasets for long-term trend analysis. Experimental results demonstrate that our approach can outperform current baselines with achieving an enhancement of more than 18% on compression rate in real-time configurations. This signifies a transformative solution for handling IoT-induced big data in contemporary railroad management systems. Qixuan Yu 0003, Shuhong Dai, Pengfei Sun 0002, Haichuan Tang, Long Cheng 0003 |
IEEE Internet Things J. | 6 |
| 2024 | MARP: A Cooperative Multiagent DRL System for Connected Autonomous Vehicle PlatooningabstractIn modern urban areas, inefficiency traffic management is one of the main causes of road congestion, leading to reduced fuel efficiency and increased traffic safety hazards. Traditional researches typically focus only on enhancing the throughput of intersections by optimizing traffic signals or individual vehicle trajectories. However, these methods often overlook the dynamic nature of the traffic system and the potential benefits of vehicle platooning, limiting their effectiveness in complex traffic environments. Addressing this challenge, this article presents MARP, a Cooperative Multiagent deep reinforcement learning (DRL) System for connected autonomous vehicle (CAV) Platooning. Utilizing vehicle to vehicle (V2I) and vehicles to infrastructure (V2V) technologies, MARP integrates sensing, computing, and communication to collect and process real-time data on traffic conditions, thereby achieving dynamic synchronization between traffic signal controllers and CAV platoons. By constructing platoons that collaborates with the infrastructure through a multiagent DRL collaboration model, MARP adapts to real-time traffic flow changes, significantly optimizing the fluidity and efficiency of the entire traffic network. Detailed experiments show that MARP effectively reduces traffic congestion, shortens intersection travel times, and cuts fuel consumption and emissions, surpassing the state-of-the-art approach. Shuhong Dai, Shike Li, Haichuan Tang, Xin Ning 0001, Fang Fang 0007, Yunxiao Fu, Qingle Wang, Long Cheng 0003 |
IEEE Internet Things J. | 8 |
| 2024 | CASA: cost-effective EV charging scheduling based on deep reinforcement learning
Qingzhi Liu, Long Cheng 0003 |
Neural Comput. Appl. | 4 |
| 2024 | Privacy and Integrity Protection for IoT Multimodal Data Using Machine Learning and BlockchainabstractWith the wide application of Internet of Things (IoT) technology, large volumes of multimodal data are collected and analyzed for various diagnoses, analyses, and predictions to help in decision-making and management. However, the research on protecting data integrity and privacy is quite limited, while the lack of proper protection for sensitive data may have significant impacts on the benefits and gains of data owners. In this research, we propose a protection solution for data integrity and privacy. Specifically, our system protects data integrity through distributed systems and blockchain technology. Meanwhile, our system guarantees data privacy using differential privacy and Machine Learning (ML) techniques. Our system aims to maintain the usability of the data for further data analytical tasks of data users, while encrypting the data according to the requirements of data owners. We implement our solution with smart contracts, distributed file systems, and ML models. The experimental results show that our proposed solution can effectively encrypt source IoT data according to the requirements of data users while data integrity can be protected under the blockchain. Qingzhi Liu, Chenglu Jin, Xiaohan Zhou, Ying Mao 0001, Cagatay Catal, Long Cheng 0003 |
ACM Trans. Multim. Comput. Commun. Appl. | 7 |
| 2024 | Discovering Hierarchical Multi-Instance Business Processes From Event LogsabstractProcess discovery aims to extract descriptive process models from event logs. To date, various process discovery algorithms have been proposed for different application settings. However, most of them meet challenges in handling event logs produced from hierarchical multi-instance business processes, in which multiple sub-process instances are invoked by the execution of a parent process. To address the problem, a novel approach is presented to support the discovery of hierarchical multi-instance process models. Specifically, taking event logs with multi-instance information as input, the detailed implementation of our method can be generally divided into four steps: nesting relation detection, hierarchical event log construction, sub-process case identification, and hierarchical multi-instance model discovery. We have implemented our approach properly as plugins in the openly accessible ProM toolkit, and compared its performance against the state-of-the-art process discovery approaches over six publicly available event logs. Based on the experimental result, it is demonstrated that the proposed approach can effectively discover hierarchical multi-instance process models with better quality. Cong Liu 0012, Ying Wang 0001, Lijie Wen 0001, Jiujun Cheng, Long Cheng 0003, Qingtian Zeng |
IEEE Trans. Serv. Comput. | 5 |
| 2024 | Advancements in Accelerating Deep Neural Network Inference on AIoT Devices: A SurveyabstractThe amalgamation of artificial intelligence with Internet of Things (AIoT) devices have seen a rapid surge in growth, largely due to the effective implementation of deep neural network (DNN) models across various domains. However, the deployment of DNNs on such devices comes with its own set of challenges, primarily related to computational capacity, storage, and energy efficiency. This survey offers an exhaustive review of techniques designed to accelerate DNN inference on AIoT devices, addressing these challenges head-on. We delve into critical model compression techniques designed to adapt to the limitations of devices and hardware optimization strategies that aim to boost efficiency. Furthermore, we examine parallelization methods that leverage parallel computing for swift inference, as well as novel optimization strategies that fine-tune the execution process. This survey also casts a future-forward glance at emerging trends, including advancements in mobile hardware, the co-design of software and hardware, privacy and security considerations, and DNN inference on AIoT devices with constrained resources. All in all, this survey aspires to serve as a holistic guide to advancements in the acceleration of DNN inference on AIoT devices, aiming to provide sustainable computing for upcoming IoT applications driven by artificial intelligence. Long Cheng 0003, Qingzhi Liu, Lei Yang 0018, Cheng Liu 0008, Ying Wang 0001 |
IEEE Trans. Sustain. Comput. | 1 |
| 2024 | A Deep Reinforcement Learning-Based Preemptive Approach for Cost-Aware Cloud Job SchedulingabstractWith some specific characteristics such as elastics and scalability, cloud computing has become the most promising technology for online business nowadays. However, how to efficiently perform real-time job scheduling in cloud still poses significant challenges. The reason is that those jobs are highly dynamic and complex, and it is always hard to allocate them to computing resources in an optimal way, such as to meet the requirements from both service providers and users. In recent years, various works demonstrate that deep reinforcement learning (DRL) can handle real-time cloud jobs well in scheduling. However, to our knowledge, none of them has ever considered extra optimization opportunities for the allocated jobs in their scheduling frameworks. Given this fact, in this work, we introduce a novel DRL-based preemptive method for further improve the performance of the current studies. Specifically, we try to improve the training of scheduling policy with effective job preemptive mechanisms, and on that basis to optimize job execution cost while meeting users' expected response time. We introduce the detailed design of our method, and our evaluations demonstrate that our approach can achieve better performance than other scheduling algorithms under different real-time workloads, including the DRL approach. Long Cheng 0003, Yue Wang 0073, Cheng Liu 0008, Zhiming Zhao, Ying Wang 0001 |
IEEE Trans. Sustain. Comput. | 1 |
| 2023 | Efficient Supernet Training Using Path ParallelismabstractCompared to conventional neural networks, training a supernet for Neural Architecture Search (NAS) is very time consuming. Although current works have demonstrated that parallel computing can significantly speed up the training process, almost all of their parallelism still follow the conventional data- and model-based paradigms, which actually face performance issues in both computation and inter-node communication of the supernet training. To further improve the performance of current methods, we discover the unique path-parallelism that exists in supernet training, and proposed a novel training approach designed specifically for supernet. In detail, we focus on analyzing path correlations between subnets in a supernet and exploiting effective path-merging methods to reduce redundant computations and communications raised by concurrent subnets. Moreover, we also try to combine the proposed path parallelism with traditional intra-subnet parallelism to perform multi-level parallelization to further optimize the parallel performance. We present the detailed design and implementation of our method, and our experimental results show that our proposed approach can achieve up to 3.2x end-to-end speedup over conventional parallel training solutions, and 1.46x–5.78x speedup compared to the state-of-art supernet training frameworks. Long Cheng 0003, Xuyi Cai, Lei Zhang 0008, Ying Wang 0001 |
HPCA | 2 |
| 2023 | DeepHealth: Geospatial and ML-Based Approach to Identify Health Disparities and Determinants for Improving Pandemic Health CareabstractThe COVID-19 pandemic has exacerbated existing health disparities, and its impact has fallen disproportionately on disadvantaged and vulnerable communities. Racial and ethnic minorities such as Black Americans who are at a particular disadvantage are more likely to be the potential target of COVID-19 infection and are dying at alarmingly high rates. Despite a promising solution of the COVID-19 vaccination offers hope, equitable access to COVID-19 vaccines remains a challenge in the US, which has compounded the existing disparities in cases, hospitalizations, and deaths among racial and ethnic minority groups. The deep and pervasive history of medical racism in the US has led to the vaccine hesitancy in racial and ethnic minorities, and thereby caused the disparities. Although some studies examine determinants of health disparities (e.g., social health determinants), there is a shortage of studies examining the social, structural and constructural health determinants, either alone or in tandem with other determinants. Little research paid attention to leveraging geographic information to trace the social, structural and constructural health determinants, which can provide a lower level of granularity. In this paper, we propose DeepHealth, a geospatial and ML-based (machine learning based) approach to identify diverse determinants (including the social, structural, and constructural determinants) of health disparities in COVID-19 pandemic, which provides a lower level of granularity. We provide a thorough analysis of health disparities based on multiple COVID-19 datasets and examine the social, structural, and constructural health determinants to assist in ascertaining why disparities (in racial and ethnic minorities who are particularly disadvantaged) occur in incidence and mortality rates due to COVID-19 pandemic. Extensive experimental results show the effectiveness of our approach. This research provides new strategies for health disparity identification and determinant tracking with a goal of mitigating health disparities and improving pandemic health care. The research suggests that policymakers should give attention to initiatives that will protect the health of populations (i.e., an upstream approach to reducing health disparities) rather than solely focusing only on providing health and social services. Long Cheng 0003, Richard A. Aló |
ICCCN | 3 |
| 2023 | Cost-aware scheduling systems for real-time workflows in cloud: An approach based on Genetic Algorithm and Deep Reinforcement Learning
Long Cheng 0003, Cong Liu 0012, Zhiming Zhao, Ying Mao 0001 |
Expert Syst. Appl. | 2 |
| 2023 | Optimal alignments between large event logs and process models over distributed systems: An approach based on Petri nets
Long Cheng 0003, Cong Liu 0012, Qingtian Zeng |
Inf. Sci. | 1 |
| 2023 | Cross-Department Collaborative Healthcare Process Model Discovery From Event LogsabstractHealthcare plays an increasingly essential role in our daily life. Modern Hospital Information Systems (HISs) record and store detailed medical treatment process information for all patients as event logs. By taking event logs as input, process mining techniques have been widely applied to extract valuable insights to improve medical treatment processes and deliver better healthcare services. However, considering the complexity of collaborations among different medical departments, existing model discovery techniques cannot be applied directly. To handle this limitation, this paper proposes a novel approach to support the discovery of Cross-department Collaborative Healthcare Process (CCHP) models from medical event logs. Specifically, an extension of classical Petri Nets with message and resource attributes is first introduced to formalize CCHPs. Then, a novel discovery algorithm is proposed to discover Intra-department Healthcare Process (IHP) models. Next, collaboration patterns among medical departments are formalized and corresponding discovery algorithms are given on that basis. Finally, a global CCHP model is obtained by integrating all discovered collaboration patterns and IHP models. By using four public medical event logs, we quantitatively compare our approach with the state-of-the-art process mining techniques in terms of model quality, and our experimental results demonstrate that the proposed approach can discover more accurate healthcare process models.Note to Practitioners—The recorded medical event logs by HISs can be used to extract valuable insights for the analysis of healthcare processes. However, existing process model discovery techniques cannot be applied for the analysis directly due to the complex collaborations among different medical departments of a hospital. This paper introduces a novel approach for cross-department collaborative healthcare process model discovery from medical event logs. All proposed techniques are fully implemented and publicly available. Using four public medical event logs, we show the applicability and advantages of our approach against existing ones. The proposed techniques are applicable to the model discovery and behavior understanding of real-life operational healthcare processes. Cong Liu 0012, Shuaipeng Zhang, Long Cheng 0003, Qingtian Zeng |
IEEE Trans Autom. Sci. Eng. | 4 |
| 2023 | Statistical Modeling of Soft Error Influence on Neural NetworksabstractSoft errors in large VLSI circuits have a significant impact on computing- and memory-intensive neural network (NN) processing. Understanding the influence of soft errors on NNs is critical to protect against soft errors for reliable NN processing. Prior work mainly relies on fault simulation to analyze the influence of soft errors on NN processing. They are accurate but usually specific to limited configurations of errors and NN models due to the prohibitively slow simulation speed especially for large NN models and datasets. With the observation that the influence of soft errors propagates across a large number of neurons and accumulates as well, we propose to characterize the soft error-induced data disturbance on each neuron with a normal distribution model using the central limit theorem and develop a series of statistical models to analyze the behavior of NN models under soft errors in general. The statistical models reveal not only the correlation between soft errors and the accuracy of NN models but also how NN parameters, such as quantization and architecture affect the reliability of NNs. The proposed models are compared with fault simulations and verified comprehensively. In addition, we observe that the statistical models that characterize the soft error influence can also be utilized to predict fault simulation results in many cases and we explore the use of the proposed statistical models to accelerate fault simulations of NNs. Our experiments show that the proposed accelerated fault simulation provides almost two orders of magnitude speedup with negligible loss of simulation accuracy compared to the baseline fault simulations. Haitong Huang, Xinghua Xue, Cheng Liu 0008, Ying Wang 0001, Tao Luo 0014, Long Cheng 0003, Huawei Li 0001, Xiaowei Li 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 6 |
| 2023 | Elastic Resource Management for Deep Learning Applications in a Container ClusterabstractThe increasing demand for learning from massive datasets is restructuring our economy. Effective learning, however, involves nontrivial computing resources. Most businesses utilize commercial infrastructure providers (e.g., AWS) to host their computing clusters in the cloud, where various jobs compete for available resources. While cloud resource management is a fruitful research field that has made many advances in production, such as Kubernetes and YARN, few efforts have been invested to further optimize the system performance, especially for Deep Learning (DL) training jobs in a container cluster. This work introduces FlowCon, a system that is able to monitor the individual evaluation functions of DL jobs at runtime, and thus to make placement decisions and resource allocations elastically. We present a detailed design and implementation of FlowCon and conduct intensive experiments over various DL models. The results demonstrate that FlowCon significantly improves DL job completion time and resource utilization efficiency when compared to default systems. According to the results, FlowCon can improve the completion time by up to 68.8% and meanwhile, reduce the makespan by 18.0%, in the presence of various DL job workloads. Ying Mao 0001, Vaishali Sharma, Wenjia Zheng, Long Cheng 0003, Qiang Guan, Ang Li 0006 |
IEEE Trans. Cloud Comput. | 4 |
| 2023 | Differentiate Quality of Experience Scheduling for Deep Learning Inferences With Docker Containers in the CloudabstractWith the prevalence of big-data-driven applications, such as face recognition on smartphones and tailored recommendations from Google Ads, we are on the road to a lifestyle with significantly more intelligence than ever before. Various neural network powered models are running at the back end of their intelligence to enable quick responses to users. Supporting those models requires lots of cloud-based computational resources, e.g., CPUs and GPUs. The cloud providers charge their clients by the amount of resources that they occupy. Clients have to balance the budget and quality of experiences (e.g., response time). The budget leans on individual business owners, and the required Quality of Experience (QoE) depends on usage scenarios of different applications. For instance, an autonomous vehicle requires an real-time response, but unlocking your smartphone can tolerate delays. However, cloud providers fail to offer a QoE-based option to their clients. In this paper, we proposeDQoES, differentiated quality of experience scheduler for deep learning inferences.DQoESaccepts clients’ specifications on targeted QoEs, and dynamically adjusts resources to approach their targets. Through the extensive cloud-based experiments,DQoESdemonstrates that it can schedule multiple concurrent jobs with respect to various QoEs and achieve up to 8x times more satisfied models when compared to the existing system. Ying Mao 0001, Weifeng Yan, Yun Song, Long Cheng 0003, Qingzhi Liu |
IEEE Trans. Cloud Comput. | 6 |
| 2023 | CoTV: Cooperative Control for Traffic Light Signals and Connected Autonomous Vehicles Using Deep Reinforcement LearningabstractThe target of reducing travel time only is insufficient to support the development of future smart transportation systems. To align with the United Nations Sustainable Development Goals (UN-SDG), a further reduction of fuel and emissions, improvements of traffic safety, and the ease of infrastructure deployment and maintenance should also be considered. Different from existing work focusing on optimizing the control in either traffic light signal (to improve the intersection throughput), or vehicle speed (to stabilize the traffic), this paper presents a multi-agent Deep Reinforcement Learning (DRL) system called CoTV, which Cooperatively controls both Traffic light signals and Connected Autonomous Vehicles (CAV). Therefore, our CoTV can well balance the reduction of travel time, fuel, and emissions. CoTV is also scalable to complex urban scenarios by cooperating with only one CAV that is nearest to the traffic light controller on each incoming road. This avoids costly coordination between traffic light controllers and all possible CAVs, thus leading to the stable convergence of training CoTV under the large-scale multi-agent scenario. We describe the system design of CoTV and demonstrate its effectiveness in a simulation study using SUMO under various grid maps and realistic urban scenarios with mixed-autonomy traffic. Jiaying Guo, Long Cheng 0003, Shen Wang 0006 |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2023 | Eco-CSAS: A Safe and Eco-Friendly Speed Advisory System for Autonomous Vehicle Platoon Using Consortium BlockchainabstractFuture worldwide 6G research will drive the evolution of emerging intelligent control technologies, such as intelligent speed advisory systems (ISA), to a more advanced generation. As a special type of ISA, consensus-based speed advisory systems (CSAS) can be widely used to recommend a consensus speed for a vehicle platoon, enabling minimizing energy consumption or emissions over a planned route. Recently, speed recommendation services that protect data privacy (i.e., how to obtain an optimal speed in a privacy-preserving way) have drawn tremendous attention. However, current approaches could still encounter service trust issues with central servers and the malicious behavior of vehicles. Furthermore, existing research lacks considering road safety constraints (i.e., safe distance between adjacent vehicles and road speed limits) that are essential for the practical deployment of CSAS. To address the above issues, this paper proposes Eco-CSAS, a safe and eco-friendly consensus speed advisory system using blockchain. We formulate an optimization problem subject to the minimum following distance and maximum road speed limit to minimize the energy consumption of the automatic vehicle platoon. In addition, we introduce a consortium blockchain and cryptographic primitives to ensure service trust and data privacy. We implement the system on the Hyperledger platform, and experimental results show that the system can achieve speed recommendations in a trustworthy and privacy-preserving manner while ensuring a secure platoon. Shike Li, Jiaming Pei, Sixing Wu, Shen Wang 0006, Long Cheng 0003 |
IEEE Trans. Intell. Transp. Syst. | 6 |
| 2023 | Toward Network-Aware Query Execution Systems in Large DatacentersabstractHow to efficiently process concurrent data tasks such as online analytical queries in datacenter environments is still a big challenge for current computing techniques. One of the fundamental reasons is that their task execution normally involves large numbers of distributed data operators, which are always expensive in terms of communication time. To improve the general performance, various advanced approaches on the execution optimization of data operators have been proposed in the past years. However, most of them focus on application-level optimization, such as using data locality scheduling to reduce network traffic. Moreover, few of them has considered the optimization opportunities for concurrent execution of multiple data operators. In this paper, we propose a novel coflow-based scheduling system called CoFlop, which aims to improve network communication time for multiple distributed operators at a query level, and on that basis to lay a solid foundation for the development of a network-aware query execution system in datacenter networks. We introduce the detailed system design of CoFlop and conduct a simulation-based evaluation with large concurrent distributed join operations. Compared to existing methods, the experimental results show that CoFlop can perform better in the presence of different large workloads. Long Cheng 0003, Ying Wang 0001, Rutvij H. Jhaveri, Qingle Wang, Ying Mao 0001 |
IEEE Trans. Netw. Serv. Manag. | 1 |
| 2023 | Formal Modeling and Discovery of Hierarchical Business Processes: A Petri Net-Based ApproachabstractBusiness processes are critical for information systems to control workflows and deliver services. Although existing process discovery techniques can generate flat process models from business event logs, few of them have investigated the notion of hierarchy (i.e., subprocesses) yet. To fill the gap, this article first defines the concept of hierarchical Petri nets (HPNs), which can support the formal modeling and correctness verification of processes with subprocesses. Followed by that, we propose an approach which can effectively discover HPNs from event logs with lifecycle information. Moreover, to quantify the quality of discovered HPNs, details on how to transform an HPN to a classical Petri net are given such that existing metrics can be applied. All proposed approaches have been fully implemented in ProM, and experiments over both synthetic and real-life event logs demonstrate that our approach can effectively discover hierarchical process models. Specifically, compared to exiting approaches on processes discovery, our approach can generally perform better in terms of model quality. Cong Liu 0012, Long Cheng 0003, Qingtian Zeng, Lijie Wen 0001 |
IEEE Trans. Syst. Man Cybern. Syst. | 2 |
| 2022 | CrossCas: A Novel Cross-Platform Approach for Predicting Cascades in Online Social Networks with Hidden Markov ModelabstractInformation sharing through online social networks (OSNs) facilitates quick discovery and consumption of information online. Many OSNs such as Facebook, Twitter provide resharing or reposting features, which allows users to share others' content with their own friends or followers. As content is shared from person to person, cascades of information-sharing can occur. There are many existing works focusing on analyzing and characterizing the cascades in OSNs. However, previous works focus on the analysis and characterization of cascades without providing a solution to accurately predict cascades. Although some methods for cascade prediction have been proposed recently, their methods work in social networks such as Facebook (or Twitter), and do not work well simultaneously in multiple OSNs such as Software Social Network (SSN) GitHub, Twitter and Reddit because GitHub, Twitter and Reddit have different social activity patterns. In this paper, we first perform a thorough analysis of cascades in multiple OSNs: GitHub, Twitter and Reddit, and identify the cascades of information-sharing. We then propose CrossCas, a novel cross-platform approach for predicting cascades in multiple OSNs with Hidden Markov Model (HMM). The experimental results show that our proposed method achieves high performance. Xiaonan Zhang 0001, Richard A. Aló, Xiuzhen Huang, Long Cheng 0003, Feng Deng |
GLOBECOM | 5 |
| 2022 | Performance Evaluation of Resource Management Schemes for Cloud Native Platforms with Computing ContainersabstractBusinesses have made increasing adoption and incorporation of cloud technology into internal processes in the last decade. The cloud-based deployment provides on-demand availability without active management. More recently, the concept of cloud-native application has been proposed and represents an invaluable step toward helping organizations develop software faster and update it more frequently to achieve dramatic business outcomes. Cloud-native is an approach to build and run applications that exploit the cloud computing delivery model’s advantages. It is more about how applications are created and deployed than where. The container-based virtualization technology, e.g., Docker and Kubernetes, serves as the foundation for cloud-native applications. This paper evaluates the performance of deep learning applications in a cloud-native environment. Yuqi Fu, Naseem Machlovi, Ying Mao 0001, Long Cheng 0003, Qingzhi Liu |
IPCCC | 5 |
| 2022 | Iterative Qubits Management for Quantum Index Searching in a Hybrid SystemabstractRecent advances in quantum computing systems attract tremendous attention. Commercial companies, such as IBM, Amazon, and IonQ, have started to provide access to noisy intermediate-scale quantum computers. Researchers and entrepreneurs attempt to deploy their applications that aim to achieve a quantum speedup. Grover’s algorithm and quantum phase estimation are the foundations of many applications with the potential for such a speedup. While these algorithms, in theory, obtain marvelous performance, deploying them on existing quantum devices is a challenging task. For example, quantum phase estimation requires extra qubits and a large number of controlled operations, which are impractical due to low-qubit and noisy hardware. To fully utilize the limited onboard qubits, we propose IQuCS, which aims at index searching and counting in a quantum-classical hybrid system. IQuCS is based on Grover’s algorithm. From the problem size perspective, it analyzes results and tries to filter out unlikely data points iteratively. A reduced data set is fed to the quantum computer in the next iteration. With a reduction in the problem size, IQuCS requires fewer qubits iteratively, which provides the potential for a shared computing environment. We implement IQuCS with Qiskit and conduct intensive experiments. The results demonstrate that it reduces qubits consumption by up to 66.2%. Wenrui Mu, Ying Mao 0001, Long Cheng 0003, Qingle Wang, Weiwen Jiang |
IPCCC | 3 |
| 2022 | Cost-aware real-time job scheduling for hybrid cloud using deep reinforcement learning
Long Cheng 0003, Archana Kalapgar, Amogh Jain, Yue Wang 0073, Yongtai Qin, Yuancheng Li 0005, Cong Liu 0012 |
Neural Comput. Appl. | 1 |
| 2022 | Measuring Similarity for Data-Aware Business ProcessesabstractBusiness process similarity measures are of vital importance for process repository management applications, such as process query, process recommendation, and process clustering. Most existing approaches measure process similarity by relying on control-flow structures only. This article investigates the role of data in process similarity measure. To incorporate data-flow information into business process control flow, it proposes a data-aware workflow net (DWF-net) by extending the classical workflow net with data reading and writing semantics. Then, we introduce three types of similarity measures, i.e., data item set-based similarity, data operation set-based similarity, and data-aware behavior-based similarity, to quantify the similarity of data-aware business processes from different perspectives. Next, a methodology is introduced to help process analysts apply these three measures in a systematical way. Finally, we evaluate the effectiveness and applicability of the proposed similarity measures by a group of comparative experiments. Cong Liu 0012, Qingtian Zeng, Long Cheng 0003, Hua Duan, Jiujun Cheng |
IEEE Trans Autom. Sci. Eng. | 3 |
| 2022 | CAP: Communication-Aware Automated Parallelization for Deep Learning Inference on CMP ArchitecturesabstractReal-time inference of deep learning models on embedded and energy-efficient devices becomes increasingly desirable with the rapid growth of artificial intelligence on edge. Specifically, to achieve superb energy-efficiency and scalability, efficient parallelization of single-pass deep neural network (DNN) inference on chip multiprocessor (CMP) architectures is urgently required by many time-sensitive applications. However, as the number of processing cores scales up and the performance of cores has grown much fast, the on-chip inter-core data movement is prone to be a performance bottleneck for computation. To remedy this problem and further improve the performance of network inference, in this work, we introduce a communication-aware DNN parallelization technique called CAP, by exploiting the elasticity and noise-tolerance of deep learning algorithms on CMP. Moreover, in the hope that the conducted studies can provide new design values for real-time neural network inference on embedded chips, we also have evaluated the proposed approach on both multi-core Neural Network Accelerators (NNA) chips and general-purpose chip-multiprocessors. Our experimental results show that the proposed CAP can achieve 1.12×-1.65× system speedups and 1.14×-2.70× energy efficiency for different neural networks while maintaining the inference accuracy, compared to baseline approaches. Kaiwei Zou, Ying Wang 0001, Long Cheng 0003, Songyun Qu, Huawei Li 0001, Xiaowei Li 0001 |
IEEE Trans. Computers | 3 |
| 2022 | A Fast Precision Tuning Solution for Always-On DNN AcceleratorsabstractDue to the nonvolatility nature of resistive RAM (ReRAM), dynamic operations in the arrays contribute to a much larger portion of power in ReRAM-based neural networks than static power. To reduce the dynamic power ofin-situoperations with neural parameters, precision-tuning is considered a viable approach of approximate computing to tradeoff excessive computation exactness for power and efficiency gains. However, the switching overhead of precision tuning in hardware severely impacts its effectiveness when the systems need to quickly react to the change of environment, user constraint or input quality. This work for the first time investigates the feasibility of agile precision tuning for neural network accelerators to benefit from approximate computing. The proposed computing in memory (CiM) CNN accelerators fully utilize the normally off characteristics of memristor crossbars to achieve instant network precision tuning without worrying about the model reloading penalty. The ReRAM-based accelerator, with the proposed neural parameter mapping policy and the novel mixed-model training method, induces negligible precision-switching latency and power consumption when compared with traditional variable precision accelerators. In evaluation with state-of-the-art workloads, the proposed ReRAM deep learning and neural network architecture saves 58.3%–62.47% area overhead over the baseline design. We also leverage the proposed ReRAM accelerator architecture to build a novel always-on key-word spotting (KWS) system. The KWS design can switch between different precision modes to capture the relevant sound with high accuracy. The experimental results show the precision-adjustable KWS architecture saves considerable operating energy when fed with realistic test-sets of audio data. Ying Wang 0001, Yintao He, Long Cheng 0003, Huawei Li 0001, Xiaowei Li 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2022 | MPC-CSAS: Multi-Party Computation for Real-Time Privacy-Preserving Speed Advisory SystemsabstractAs a part of Advanced Driver Assistance Systems (ADASs), Consensus-based Speed Advisory Systems (CSAS) have been proposed to recommend a common speed to a group of vehicles for specific application purposes, such as emission control and energy management. With Vehicle-to-Vehicle (V2V), Vehicle-to-Infrastructure (V2I) technologies and advanced control theories in place, state-of-the-art CSAS can be designed to get an optimal speed in a privacy-preserving and decentralized manner. However, the current method only works for specific cost functions of vehicles, and its execution usually involves many algorithm iterations leading long convergence time. Therefore, the state-of-the-art design method is not applicable to a CSAS design which requires real-time decision making. In this article, we address the problem by introducing MPC-CSAS, a Multi-Party Computation (MPC) based design approach for privacy-preserving CSAS. Our proposed method is simple to implement and applicable to all types of cost functions of vehicles. Moreover, our simulation results show that the proposed MPC-CSAS can achieve very promising system performance in just one algorithm iteration without using extra infrastructure for a typical CSAS. Mingming Liu 0001, Long Cheng 0003, Yingqi Gu, Ying Wang 0001, Qingzhi Liu, Noel E. O'Connor |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2022 | Deep Reinforcement Learning for Load-Balancing Aware Network Control in IoT Edge SystemsabstractLoad balancing is directly associated with the overall performance of a parallel and distributed computing system. Although the relevant problems in communication and computation have been well studied in data center environments, few works have considered the issues in an Internet of Things (IoT) edge scenario. In fact, processing data in a load balancing way for the latter case is more challenging. The main reason is that, unlike a data center, both the data sources and the network infrastructure in an IoT edge system can be dynamic. Moreover, with different performance requirements from IoT networks and edge servers, it will be hard to characterize the performance model and to perform runtime optimization for the whole system. To tackle this problem, in this work, we propose a load-balancing aware networking approach for efficient data processing in IoT edge systems. Specifically, we introduce an IoT network dynamic clustering solution using the emerging deep reinforcement learning (DRL), which can both fulfill the communication balancing requirements from IoT networks and the computation balancing requirements from edge servers. Moreover, we implement our system with a long short term memory (LSTM) based Dueling Double Deep Q-Learning Network (D3QN) model, and our experiments with real-world datasets collected from an autopilot vehicle demonstrate that our proposed method can achieve significant performance improvement compared to benchmark solutions. Qingzhi Liu, Tiancong Xia, Long Cheng 0003, Merijn van Eijk, Tanir Ozcelebi, Ying Mao 0001 |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2021 | DeepTrack: An ML-based Approach to Health Disparity Identification and Determinant Tracking for Improving Pandemic Health CareabstractThe Coronavirus disease 2019 (COVID-19) pandemic has severely impacted countries around the world with unprecedented mortality and economic devastation and has disproportionately and negatively impacted different communities—especially racial and ethnic minorities who are at a particular disadvantage. Black Americans have a long-standing history of disadvantage (e.g., long-standing disparities in health outcomes) and are in a vulnerable position to experience the impact of this pandemic. Some studies indicate high-risk and vulnerability of the elderly and patients with underlying co-morbidities, however, little research paid attention to leveraging geographic information and machine learning (ML) to track the social and structural health determinants, which can provide a lower level of granularity. In this paper, we propose DeepTrack, a geospatial and ML-based approach to identify diverse determinants (including the structural, social, and constructural determinants) of health disparities in COVID-19 pandemic, which provides a lower level of granularity. We provide a thorough analysis of health disparities and diets based on multiple COVID-19 datasets and examine the structural, social, and constructural health determinants to assist in ascertaining why disparities (in racial and ethnic minorities who are particularly disadvantaged) occur in infection and death rates due to COVID-19 pandemic. We track determinants of nutrition and obesity through diet examination. Extensive experimental results show the effectiveness of our approach. The research provides new strategies for health disparity identification and determinant tracking with a goal to improve pandemic health care. Long Cheng 0003, Ankur Sarker, Li Yan 0004, Richard A. Aló |
IEEE BigData | 2 |
| 2021 | CDetector: Extracting Textual Features of Financial Social Media to Detect Cyber AttacksabstractWith the proliferation of social media, cyber threats and attacks have significantly increased in complexity and quantity in financial market. Malicious hackers leverage the influence of social media to spread deceptive information with an intent to gain abnormal profits illegally or to cause losses. Measuring information content in financial social media helps identify these threats and attacks. In this paper, we propose CDetector, an ML-based approach to identifying social media features that correlate with abnormal returns of the stocks of companies vulnerable to be targets of cyber attacks (e.g., cognitive hacking). To test our approach, we collected price data and the social media messages on multiple technology companies, and extracted features that contributed to abnormal stock movements. Preliminary results show that the top social media features associated with abnormal price movements are the terms that are simple, motivate actions, incite emotion, and use exaggeration, and the selected features correlate with abnormal messages and abnormal returns of the stocks of companies. Long Cheng 0003, Hongmei Chi, Cong Liu 0012, Richard A. Aló |
ICCCN | 2 |
| 2021 | TEBC-Net: An Effective Relation Extraction Approach for Simple Question Answering over Knowledge Graphs
Ketong Qu, Jingchen Yan, Liting Zhou, Long Cheng 0003 |
KSEM | 5 |
| 2021 | Cluster-based flow control in hybrid software-defined wireless sensor networksabstractSoftware-defined networking (SDN) is a cornerstone of next-generation networks and has already led to numerous advantages for data-center networks and wide-area networks. However, SDN is not widely adopted in constrained networks, such as Wireless Sensor Networks (WSN), due to excessive control overhead, lossy medium, and in-band control channels. Therefore, a key challenge to enable Software-Defined Wireless Sensor Networks (SD-WSN) is to reduce the number of control messages required to configure the data plane. In this paper, we propose a cluster-based flow control approach in hybrid SDNs. Our approach is hybrid in the sense that it takes advantage of distributed legacy routing and centralized SDN routing. In addition, it makes a trade-off between the granularity of flow control and the communication overhead induced by the SDN controller. The approach partitions a network into clusters with minimum number of border nodes. Instead of handling the individual flows of each node, the SDN controller only manages incoming and outgoing traffic flows of clusters through border nodes, while the flows inside each cluster are controlled by a distributed legacy WSN routing algorithm. Our proof-of-concept implementations in both software and hardware show that our approach is efficient with respect to reducing the number of nodes that must be managed and the number of control messages. In comparison to benchmark solutions with and without clustering, our solution reduces communication costs for flow configuration in an SD-WSN at least by 27% and at most by 88% respectively, without degrading packet delay nor delivery rate. Qingzhi Liu, Long Cheng 0003, Renan C. A. Alves, Tanir Ozcelebi, Fernando A. Kuipers, Johan J. Lukkien, Shanzhi Chen |
Comput. Networks | 2 |
| 2021 | Sampling business process event logs using graph-based ranking modelabstractSummary Modern information systems are continuously collecting and storing large volumes of business process event logs. The analysis of event logs can provide valuable insights for business process re‐engineering and enhancement. Process discovery, as one of the most challenging event log analysis techniques, aims to discover a business process model from an event log. Many process discovery approaches have been proposed in the past two decades, however, most of them suffer from efficiency problem when dealing with large‐scale event logs. Motivated by PageRank, we propose LogRank, a graph‐based ranking model, for event log sampling in this paper. The LogRank is capable of sampling a large‐scale event log to a smaller size that can be efficiently handled by existing discovery approaches. To support real‐life applications, we instantiate the LogRank model for two typical types of event logs, that is, simple event logs and lifecycle event logs. To quantify the quality of a sample log with respect to the original one, we introduce a general evaluation framework that can be instantiated for different quality metrics. The proposed sampling approach has been implemented in the open‐source process mining toolkit ProM. By experiments with both synthetic and real‐life event logs, we demonstrate that the proposed LogRank‐based sampling approach provides an effective means to improve process discovery efficiency as well as guaranteeing high quality of discovered models. Cong Liu 0012, Yulong Pei, Long Cheng 0003, Qingtian Zeng, Hua Duan |
Concurr. Comput. Pract. Exp. | 3 |
| 2021 | Privacy-Preserving Behavioral Correctness Verification of Cross-Organizational Workflow With Task Synchronization PatternsabstractWorkflow management technology has become a key means to improve enterprise productivity. More and more workflow systems are crossing organizational boundaries and may involve multiple interacting organizations. This article focuses on a type of loosely coupled workflow architecture with collaborative tasks, i.e., each business partner owns its private business process and is able to operate independently, and all involved organizations need to be synchronized at a certain point to complete certain public tasks. Because of each organization’s privacy consideration, they are unwilling to share the business details with others. In this way, traditional correctness verification approaches via reachability analysis are not practical as a global business process model is unavailable for privacy preservation. To ensure its globally correct execution, this work establishes a correctness verification approach for the cross-organizational workflow with task synchronization patterns. Its core idea is to use local correctness of each suborganizational workflow process to guarantee its global correctness. We prove that the proposed approach can be used to investigate the behavioral property preservation when synthesizing suborganizational workflows via collaborative tasks. A medical diagnosis running case is used to illustrate the applicability of the proposed approaches.Note to Practitioners—Cross-organizational workflow verification techniques play an increasingly important role in ensuring the correct execution of collaborative enterprise businesses. This work addresses the issue of correctness verification for loosely coupled interactive workflows with collaborative tasks. To ensure the globally correct execution, a behavioral correctness verification approach is established. All proposed concepts and techniques are supported by open-source tools, and evaluation over a medical diagnosis process case has shown their applicability. The proposed methodology is readily applicable to industrial-size workflow correctness verification problems. Cong Liu 0012, Qingtian Zeng, Long Cheng 0003, Hua Duan, MengChu Zhou, Jiujun Cheng |
IEEE Trans Autom. Sci. Eng. | 3 |
| 2021 | An Edge 3D CNN Accelerator for Low-Power Activity Recognitionabstract3D convolutional neural networks (CNNs) are gaining increasing popularity in the area of video-based action/activity analysis. Compared to 2D convolutions that share the filters in a 2D spatial domain, 3D convolutions further reuse filters in the temporal dimension to capture temporal-domain features in the video. How to exploit the data locality in the temporal dimension directly impacts the energy efficiency of specialized architectures for 3D CNN inference. Prior works on specialized 3D-CNN accelerators employ additional on-chip memories and multicluster architecture to reuse data among the process element (PE) arrays, which is very expensive for low-power chip implementation. Instead of harvesting in-memory data locality, we propose the architecture of systolic cube to exploit the spatial and temporal localities in 3D CNNs, which moves the reusable data in-between PEs connected via a 3D-cube network-on-chip. Furthermore, due to the existence of visual feature reappearance in the temporal domain, there exists a considerable portion of repetitive pixels and activations among the feature maps captured at adjacent time slots. To eliminate such temporal redundancy in 3D CNNs, the proposed accelerator architecture is equipped with a redundancy detection and elimination mechanism, capable of skipping the computations with the same activations and parameters when reusing the convolutional filters along the temporal dimension. In our evaluation, the experimental results show that the systolic-cube architecture contributes to a considerable energy-efficiency boost for state-of-the-art activity-recognition benchmarks and datasets. Ying Wang 0001, Yongchen Wang, Cong Shi 0003, Long Cheng 0003, Huawei Li 0001, Xiaowei Li 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2021 | A Low-Cost Multi-Failure Resilient Replication Scheme for High-Data Availability in Cloud StorageabstractData availability is one of the most important performance factors in cloud storage systems. To enhance data availability, replication is a common approach to handle the machine failures. However, previously proposed replication schemes cannot effectively handle both correlated and non-correlated machine failures, especially while increasing the data availability with limited resources. The schemes for correlated machine failures must create a constant number of replicas for each data object, which often neglects diverse data popularities and does not utilize the resource to maximize the expected data availability. Also, the previous schemes neglect the consistency maintenance cost and the storage cost caused by replication. It is critical for cloud providers to maximize data availability (hence minimize SLA violations) while minimizing costs caused by replication in order to maximize the revenue. In this paper, we build a nonlinear integer programming model to maximize data availability in both types of failures, and therefore minimize the cost caused by replication. Based on the model's solution for the replication degree of each data object, we propose a low-cost multi-failure (correlated and non-correlated machine failures) resilient replication scheme (MRR). MRR can effectively handle both correlated and non-correlated machine failures, considers data popularities to enhance data availability, and also tries to minimize consistency maintenance and storage cost. Extensive numerical results from trace parameters and experiments from real-world Amazon S3 demonstrate that MRR achieves high data availability, low data loss probability and low consistency maintenance and storage costs when compared to previous replication schemes. Haiying Shen, Hongmei Chi, Husnu S. Narman, Yongyi Yang, Long Cheng 0003, Wingyan Chung |
IEEE/ACM Trans. Netw. | 6 |
| 2021 | Network-Aware Locality Scheduling for Distributed Data Operators in Data CentersabstractLarge data centers are currently the mainstream infrastructures for big data processing. As one of the most fundamental tasks in these environments, the efficient execution of distributed data operators (e.g., join and aggregation) are still challenging current data systems, and one of the key performance issues is network communication time. State-of-the-art methods trying to improve that problem focus on either application-layer data locality optimization to reduce network traffic or on network-layer data flow optimization to increase bandwidth utilization. However, the techniques in the two layers are totally independent from each other, and performance gains from a joint optimization perspective have not yet been explored. In this article, we propose a novel approach called NEAL (NEtwork-Aware Locality scheduling) to bridge this gap, and consequently to further reduce communication time for distributed big data operators. We present the detailed design and implementation of NEAL, and our experimental results demonstrate that NEAL always performs better than current approaches for different workloads and network bandwidth configurations. Long Cheng 0003, Ying Wang 0001, Qingzhi Liu, Dick H. J. Epema, Cheng Liu 0008, Ying Mao 0001, John Murphy 0001 |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2021 | R2F: A Remote Retraining Framework for AIoT Processors With Computing ErrorsabstractArtificial Intelligence of Things (AIoT) processors fabricated with newer technology nodes suffer rising soft errors due to the shrinking transistor sizes and lower power supply. Soft errors on the AIoT processors particularly the deep learning accelerators (DLAs) with massive computing may cause substantial computing errors. These computing errors are difficult to be captured by the conventional training on general-purposed processors such as CPUs and GPUs in a server. Applying the offline trained neural network models to the edge accelerators with errors directly may lead to considerable prediction accuracy loss. To address the problem, we propose a remote retraining framework (R2F) for remote AIoT processors with computing errors. It takes the remote AIoT processor with soft errors in the training loop such that the on-site computing errors can be learned with the application data on the server and the retrained models can be resilient to the soft errors. Meanwhile, we propose an optimized partial triple modular redundancy (TMR) strategy to enhance the retraining. According to our experiments, R2F enables elastic design tradeoffs between the model accuracy and the performance penalty. The top-5 model accuracy can be improved by 1.93%–13.73% with 0%–200% performance penalty at high fault error rate. In addition, we notice that the retraining requires massive data transmission and even dominates the training time and propose a sparse increment compression approach for the data transmission optimization, which reduces the retraining time by 38%–88% on average with negligible accuracy loss over straightforward remote retraining. Dawen Xu 0002, Meng He 0012, Cheng Liu 0008, Ying Wang 0001, Long Cheng 0003, Huawei Li 0001, Xiaowei Li 0001, Kwang-Ting Cheng |
IEEE Trans. Very Large Scale Integr. Syst. | 5 |
| 2020 | Nonnegative Residual Matrix Factorization for Community Detection
Yulong Pei, Cong Liu 0012, Chuanyang Zheng, Long Cheng 0003 |
WISE (1) | 4 |
| 2020 | Rethinking Defeasible Reasoning: A Scalable ApproachabstractAbstract Recent technological advances have led to unprecedented amounts of generated data that originate from the Web, sensor networks, and social media. Analytics in terms of defeasible reasoning – for example, for decision making – could provide richer knowledge of the underlying domain. Traditionally, defeasible reasoning has focused on complex knowledge structures over small to medium amounts of data, but recent research efforts have attempted to parallelize the reasoning process over theories with large numbers of facts. Such work has shown that traditional defeasible logics come with overheads that limit scalability. In this work, we design a new logic for defeasible reasoning, thus ensuring scalability by design. We establish several properties of the logic, including its relation to existing defeasible logics. Our experimental results indicate that our approach is indeed scalable and defeasible reasoning can be applied to billions of facts. Michael J. Maher, Ilias Tachmazidis, Grigoris Antoniou, Stephen J. Wade, Long Cheng 0003 |
Theory Pract. Log. Program. | 5 |
| 2020 | Scalable Discovery of Hybrid Process Models in a Cloud Computing EnvironmentabstractProcess descriptions are used to create products and deliver services. To lead better processes and services, the first step is to learn a process model. Process discovery is such a technique which can automatically extract process models from event logs. Although various discovery techniques have been proposed, they focus on either constructing formal models which are very powerful but complex, or creating informal models which are intuitive but lack semantics. In this work, we introduce a novel method that returns hybrid process models to bridge this gap. Moreover, to cope with today's big event logs, we propose an efficient method, calledf-HMD, aims at scalable hybrid model discovery in a cloud computing environment. We present the detailed implementation of our approach over the Spark framework, and our experimental results demonstrate that the proposed method is efficient and scalable. Long Cheng 0003, Boudewijn F. van Dongen, Wil M. P. van der Aalst |
IEEE Trans. Serv. Comput. | 1 |
| 2019 | Resilient Neural Network Training for Accelerators with Computing ErrorsabstractWith the advancements of neural networks, customized accelerators are increasingly adopted in massive AI applications. To gain higher energy efficiency or performance, many hardware design optimizations such as near-threshold logic or overclocking can be utilized. In these cases, computing errors may happen and the computing errors are difficult to be captured by conventional training on general purposed processors (GPPs). Applying the offline trained neural network models to the accelerators with errors directly may lead to considerable prediction accuracy loss. To address this problem, we explore the resilience of neural network models and relax the accelerator design constraints to enable aggressive design options. First of all, we propose to train the neural network models using the accelerators' forward computing results such that the models can learn both the data and the computing errors. In addition, we observe that some of the neural network layers are more sensitive to the computing errors. With this observation, we schedule the most sensitive layer to the attached GPP to reduce the negative influence of the computing errors. According to the experiments, the neural network models obtained from the proposed training outperform the original models significantly when the CNN accelerators are affected by computing errors. Dawen Xu 0002, KouZi Xing, Cheng Liu 0008, Ying Wang 0001, Yulin Dai, Long Cheng 0003, Huawei Li 0001, Lei Zhang 0008 |
ASAP | 6 |
| 2019 | Deep Reinforcement Learning for IoT Network Dynamic Clustering in Edge ComputingabstractProcessing big data generated in large Internet of Things (IoT) networks is challenging current techniques. To date, a lot of network clustering approaches have been proposed to improve the performance of data collection in IoT. However, most of them focus on partitioning networks with static topologies, and thus they are not optimal in handling the case with moving objects in the networks. Moreover, to the best of our knowledge, none of them has ever considered the performance of computing in edge servers. To solve these problems, we propose a highly efficient IoT network dynamic clustering solution in edge computing using deep reinforcement learning (DRL). Our approach can both fulfill the data communication requirements from IoT networks and load-balancing requirements from edge servers, and thus provide a great opportunity for future high performance IoT data analytics. We implement our approach using a Deep Q-learning Network (DQN) model, and our preliminary experimental results show that the DQN solution can achieve higher scores in cluster partitioning compared with the current static benchmark solution. Qingzhi Liu, Long Cheng 0003, Tanir Ozcelebi, John Murphy 0001, Johan J. Lukkien |
CCGRID | 2 |
| 2019 | FlowCon: Elastic Flow Configuration for Containerized Deep Learning ApplicationsabstractAn increasing number of companies are using data analytics to improve their products, services, and business processes. However, learning knowledge effectively from massive data sets always involves nontrivial computational resources. Most businesses thus choose to migrate their hardware needs to a remote cluster computing service (e.g., AWS) or to an in-house cluster facility which is often run at its resource capacity. In such scenarios, where jobs compete for available resources utilizing resources effectively to achieve high-performance data analytics becomes desirable. Although cluster resource management is a fruitful research area having made many advances (e.g., YARN, Kubernetes), few projects have investigated how further optimizations can be made specifically for training multiple machine learning (ML) / deep learning (DL) models. In this work, we introduce FlowCon, a system which is able to monitor loss functions of ML/DL jobs at runtime, and thus to make decisions on resource configuration elastically. We present a detailed design and implementation of FlowCon, and conduct intensive experiments over various DL models. Our experimental results show that FlowCon can strongly improve DL job completion time and resource utilization efficiency, compared to existing approaches. Specifically, FlowCon can reduce the completion time by up to 42.06% for a specific job without sacrificing the overall makespan, in the presence of various DL job workloads. Wenjia Zheng, Michael Tynes, Henry Gorelick, Ying Mao 0001, Long Cheng 0003, Yantian Hou |
ICPP | 5 |
| 2019 | Learning Process Models in IoT EdgeabstractProcess models as knowledge graph representation have been widely used in various domains to create products and deliver services. Although different process model discovery approaches have been proposed in recent years, few of them are designed for distributed computing environments. Specifically, none of them has been studied in the emerging edge computing application scenarios. In this paper, based on the requirements of some real-time process services, we propose a system design for learning process models in IoT edge. We present the details of our solution and our preliminary results on a simulated IoT network show that our method can discover real-time process models in less than a second. Long Cheng 0003, Cong Liu 0003, Qingzhi Liu, Yucong Duan, John Murphy 0001 |
SERVICES | 1 |
| 2019 | CluFlow: Cluster-based Flow Management in Software-Defined Wireless Sensor NetworksabstractSoftware-defined networking (SDN) is a cornerstone of next-generation networks and has already led to numerous advantages for data-center networks and wide-area networks, for instance in terms of reduced management complexity and more fine-grained traffic engineering. However, the design and implementation of SDN within wireless sensor networks (WSN) have received far less attention. Unfortunately, because of the multi-hop type of communication in WSN, a direct reuse of the wired SDN architecture could lead to excessive communication overhead. In this paper, we propose a cluster-based flow management approach that makes a trade-off between the granularity of monitoring by an SDN controller and the communication overhead of flow management. A network is partitioned into clusters with a minimum number of border nodes. Instead of having to handle the individual flows of all nodes, the SDN controller only manages incoming and outgoing traffic flows of clusters through border nodes. Our proof-of-concept implementations in software and hardware show that, when compared with benchmark solutions, our approach is significantly more efficient with respect to the number of nodes that must be managed and the number of control messages exchanged. Qingzhi Liu, Tanir Ozcelebi, Long Cheng 0003, Fernando A. Kuipers, Johan J. Lukkien |
WCNC | 3 |
| 2019 | Load-balancing distributed outer joins through operator decomposition
Long Cheng 0003, Spyros Kotoulas, Qingzhi Liu, Ying Wang 0001 |
J. Parallel Distributed Comput. | 1 |
| 2019 | A QoS-QoR Aware CNN Accelerator Design ApproachabstractRecently powerful convolutional neural network (CNN) accelerators are emerging as energy-efficient solutions for real-time vision/speech processing, recognition and a wide spectrum of approximate computing applications. In addition to the broad applicability scope of such deep learning (DL) accelerators, we found that the fascinating feature of deterministic performance makes them ideal candidates as application-processors in embedded SoCs concerned with real-time processing. However, unlike traditional accelerator designs, DL accelerators introduce the new aspect of design tradeoff between real-time processing [quality of service (QoS)] and computation approximation [quality of result (QoR)] into embedded systems. This paper proposes an elastic CNN acceleration architecture that automatically adapts to the user-specified QoS constraint by exploiting the error-resilience in typical approximate computing workloads. For the first time, the proposed design, including the network tuning-and-mapping software and reconfigurable accelerator hardware, aims to reconcile the design constraint of QoS and QoR, which are respectively, the critical concerns in real-time and approximate computing. It is shown in experiments the proposed architecture enables the embedded system to work flexibly in an expanded operating space, significantly enhances its real-time ability, and maximizes the system energy-efficiency within the user-specified QoS-QoR constraint through self-reconfiguration. Also, we showcase the application of the proposed design approach to lower power image recognition challenge (LPIRC) and how it is employed to forge an energy-efficient solution to the LPIRC contest. Ying Wang 0001, Huawei Li 0001, Long Cheng 0003, Xiaowei Li 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2018 | Minimizing Network Traffic for Distributed Joins Using Lightweight Locality-Aware Scheduling
Long Cheng 0003, John Murphy 0001, Qingzhi Liu, Chunliang Hao, Georgios Theodoropoulos 0001 |
Euro-Par | 1 |
| 2018 | Efficient Skew Handling for Outer Joins in a Cloud Computing EnvironmentabstractOuter joins are ubiquitous in many workloads and Big Data systems. The question of how to best execute outer joins in large parallel systems is particularly challenging, as real world datasets are characterized by data skew leading to performance issues. Although skew handling techniques have been extensively studied for inner joins, there is little published work solving the corresponding problem for parallel outer joins, especially in the extremely popular Cloud computing environment. Conventional approaches to the problem such as ones based on hash redistribution often lead to load balancing problems while duplication-based approaches incur significant overhead in terms of network communication. In this paper, we propose a new approach for efficient skew handling in outer joins over a Cloud computing environment. We present an efficient implementation of our approach over the Spark framework. We evaluate the performance of our approach on a 192-core system with large test datasets in excess of 100 GB and with varying skew. Experimental results show that our approach is scalable and, at least in cases of high skew, significantly faster than the state-of-the-art. Long Cheng 0003, Spyros Kotoulas |
IEEE Trans. Cloud Comput. | 1 |
| 2017 | Efficient Event Correlation over Distributed SystemsabstractEvent correlation is a cornerstone for process discovery over event logs crossing multiple data sources. The computed correlation rules and process instances will greatly help us to unleash the power of process mining. However, exploring all possible event correlations over a log could be time consuming, especially when the log is large. State-of-the-art methods based on MapReduce designed to handle this challenge have offered significant performance improvements over standalone implementations. However, all existing techniques are still based on a conventional generating-and-pruning scheme. Therefore, event partitioning across multiple machines is often inefficient. In this paper, following the principle of filtering-and-verification, we propose a new algorithm, called RF-GraP, which provides a more efficient correlation over distributed systems. We present the detailed implementation of our approach and conduct a quantitative evaluation using the Spark platform. Experimental results demonstrate that the proposed method is indeed efficient. Compared to the state-of-the-art, we are able to achieve significant performance speedups with obviously less network communication. Long Cheng 0003, Boudewijn F. van Dongen, Wil M. P. van der Aalst |
CCGrid | 1 |
| 2017 | A Coflow-Based Co-Optimization Framework for High-Performance Data AnalyticsabstractEfficient execution of distributed database operators such as joining and aggregating is critical for the performance of big data analytics. With the increase of the compute speedup of modern CPUs, reducing the network communication time of these operators in large systems is becoming increasingly important, and also challenging current techniques. Significant performance improvements have been achieved by using state-of-the-art methods, such as reducing network traffic designed in the data management domain, and data flow scheduling in the data communications domain. However, the proposed techniques in both fields just view each other as a black box, and performance gains from a co-optimization perspective have not yet been explored. In this paper, based on current research in coflow scheduling, we propose a novel Coflow-based Co-optimization Framework (CCF), which can co-optimize application-level data movement and network-level data communications for distributed operators, and consequently contribute to their performance in large distributed environments. We present the detailed design and implementation of CCF, and conduct an experimental evaluation of CCF using large-scale simulations on large data joins. Our results demonstrate that CCF can always perform faster than current approaches on network communications in large-scale distributed scenarios. Long Cheng 0003, Ying Wang 0001, Yulong Pei, Dick H. J. Epema |
ICPP | 1 |
| 2017 | Improving the robustness and performance of parallel joins over distributed systems
Long Cheng 0003, Spyros Kotoulas, Tomás Ward, Georgios Theodoropoulos 0001 |
J. Parallel Distributed Comput. | 1 |
| 2017 | Design and evaluation of small-large outer joins in cloud computing environments
Long Cheng 0003, Ilias Tachmazidis, Spyros Kotoulas, Grigoris Antoniou |
J. Parallel Distributed Comput. | 1 |
| 2016 | Efficient Large Outer Joins over MapReduce
Long Cheng 0003, Spyros Kotoulas |
Euro-Par | 1 |
| 2016 | Efficient Data Redistribution to Speedup Big Data Analytics in Large SystemsabstractThe performance of parallel data analytics systems becomes increasingly important with the rise of Big Data. An essential operation in such environment is parallel join, which always incurs significant cost on network communication. State-of-the-art approaches have achieved performance improvements over conventional implementations through minimizing network traffic or communication time. However, these approaches still face performance issues in the presence of big data and/or large-scale systems, due to their heavy overhead of data redistribution scheduling. In this paper, we propose near-join, a network-aware redistribution approach targeting to efficiently reduce both network traffic and communication time of join executions. Particularly, near-join is lightweight and adaptable to processing large datasets over large systems. We present the details of our algorithm and its implementation. The experiments performed on a cluster of up to 400 nodes and datasets of about 100GB have demonstrated that our scheduling algorithm is much faster than the state-of-the-art methods. Moreover, our join implementation can also achieve speedups over the conventional approaches. Long Cheng 0003 |
HiPC | 1 |
| 2016 | Fast Compression of Large Semantic Web Data Using X10abstractThe Semantic Web comprises enormous volumes of semi-structured data elements. For interoperability, these elements are represented by long strings. Such representations are not efficient for the purposes of applications that perform computations over large volumes of such information. A common approach to alleviate this problem is through the use of compression methods that produce more compact representations of the data. The use of dictionary encoding is particularly prevalent in Semantic Web database systems for this purpose. However, centralized implementations present performance bottlenecks, giving rise to the need for scalable, efficient distributed encoding schemes. In this paper, we propose an efficient algorithm for fast encoding large Semantic Web data. Specially, we present the detailed implementation of our approach based on the state-of-art asynchronous partitioned global address space (APGAS) parallel programming model. We evaluate performance on a cluster of up to 384 cores and datasets of up to 11 billion triples (1.9 TB). Compared to the state-of-art approach, we demonstrate a speed-up of$2.6 - 7.4\times$and excellent scalability. In the meantime, these results also illustrate the significant potential of the APGAS model for efficient implementation of dictionary encoding and contributes to the engineering of more efficient, larger scale Semantic Web applications. Long Cheng 0003, Avinash Malik, Spyros Kotoulas, Tomás Ward, Georgios Theodoropoulos 0001 |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2015 | Scale-Out Processing of Large RDF DatasetsabstractDistributed RDF data management systems become increasingly important with the growth of the Semantic Web. Regardless, current methods meet performance bottlenecks either on data loading or querying when processing large amounts of data. In this work, we propose efficient methods for processing RDF using dynamic data re-partitioning to enable rapid analysis of large datasets. Our approach adopts a two-tier index architecture on each computation node: (1) a lightweight primary index, to keep loading times low, and (2) a series of dynamic, multi-level secondary indexes, calculated as a by-product of query execution, to decrease or remove inter-machine data movement for subsequent queries that contain the same graph patterns. In addition, we propose methods to replace some secondary indexes with distributed filters, so as to decrease memory consumption. Experimental results on a commodity cluster with 16 nodes show that the method presents good scale-out characteristics and can indeed vastly improve loading speeds while remaining competitive in terms of performance. Specifically, our approach can load a dataset of 1.1 billion triples at a rate of 2.48 million triples per second and provide competitive performance to RDF-3X and 4store for expensive queries. Long Cheng 0003, Spyros Kotoulas |
IEEE Trans. Big Data | 1 |
| 2014 | Investigating Distributed Approaches to Efficiently Extract Textual Evidences for Biomedical OntologiesabstractHeterogeneous data resources in biomedicine become available both in structured and unstructured formats, such as scientific publications, healthcare guidelines, controlled vocabularies, and formal ontologies. %Increasing researches have been conducted to Bridging the gaps among these heterogeneous data is useful to discovery implicit knowledge. To make this happen, efficient computational approaches are a necessity for applications in such a knowledge- and data-intensive domain. In this paper, we first define a particular task, relation alignment, which is to identify textual evidences for biomedical ontologies. Then, we investigate two parallel approaches for this task over distributed systems and present the details of their implementations. Moreover, we characterize the performance of our methods through extensive experiments, thereby allowing researchers to make a more informed choice in the presence of large-scale biomedical data. Long Cheng 0003 |
BIBE | 1 |
| 2014 | Efficiently Handling Skew in Outer Joins on Distributed SystemsabstractOuter joins are ubiquitous in databases and big data systems. The question of how best to execute outer joins in large parallel systems is particularly challenging as real world datasets are characterized by data skew leading to performance issues. Although skew handling techniques have been extensively studied for inner joins, there is little published work solving the corresponding problem for parallel outer joins. Conventional approaches to this problem such as ones based on hash redistribution often lead to load balancing problems while duplication-based approaches incurs significant overhead in terms of network communication. In this paper, we propose a new algorithm, query with counters (QC), for directly handling skew in outer joins on distributed architectures. We present an efficient implementation of our approach based on the asynchronous partitioned global address space (APGAS) parallel programming model. We evaluate the performance of our approach on a cluster of 192 cores (16 nodes) and datasets of 1 billion tuples with different skew. Experimental results show that our method is scalable and, in cases of high skew, faster than the state-of-the-art. Long Cheng 0003, Spyros Kotoulas, Tomás Ward, Georgios Theodoropoulos 0001 |
CCGRID | 1 |
| 2014 | Robust and Skew-resistant Parallel Joins in Shared-Nothing SystemsabstractThe performance of joins in parallel database management systems is critical for data intensive operations such as querying. Since data skew is common in many applications, poorly engineered join operations result in load imbalance and performance bottlenecks. State-of-the-art methods designed to handle this problem offer significant improvements over naive implementations. However, performance could be further improved by removing the dependency on global skew knowledge and broadcasting. In this paper, we propose PRPQ (partial redistribution & partial query), an efficient and robust join algorithm for processing large-scale joins over distributed systems. We present the detailed implementation and a quantitative evaluation of our method. The experimental results demonstrate that the proposed PRPQ algorithm is indeed robust and scalable under a wide range of skew conditions. Specifically, compared to the state-of-art PRPD method, we achieve 16% - 167% performance improvement and 24% - 54% less network communication under different join workloads. Long Cheng 0003, Spyros Kotoulas, Tomás Ward, Georgios Theodoropoulos 0001 |
CIKM | 1 |
| 2014 | Robust and Efficient Large-Large Table Outer Joins on Distributed Infrastructures
Long Cheng 0003, Spyros Kotoulas, Tomás Ward, Georgios Theodoropoulos 0001 |
Euro-Par | 1 |
| 2014 | Design and evaluation of parallel hashing over large-scale dataabstractHigh-performance analytical data processing systems often run on servers with large amounts of memory. A common data structure used in such environment is the hash tables. This paper focuses on investigating efficient parallel hash algorithms for processing large-scale data. Currently, hash tables on distributed architectures are accessed one key at a time by local or remote threads while shared-memory approaches focus on accessing a single table with multiple threads. A relatively straightforward “bulk-operation” approach seems to have been neglected by researchers. In this work, using such a method, we propose a high-level parallel hashing framework, Structured Parallel Hashing, targeting efficiently processing massive data on distributed memory. We present a theoretical analysis of the proposed method and describe the design of our hashing implementations. The evaluation reveals a very interesting result - the proposed straightforward method can vastly outperform distributed hashing methods and can even offer performance comparable with approaches based on shared memory supercomputers which use specialized hardware predicates. Moreover, we characterize the performance of our hash implementations through extensive experiments, thereby allowing system developers to make a more informed choice for their high-performance applications. Long Cheng 0003, Spyros Kotoulas, Tomás Ward, Georgios Theodoropoulos 0001 |
HiPC | 1 |
| 2014 | Massively Parallel Reasoning under the Well-Founded Semantics Using X10abstractAcademia and industry are investigating novel approaches for processing vast amounts of data coming from enterprises, the Web, social media and sensor readings in an area that has come to be known as Big Data. Logic programming has traditionally focused on complex knowledge structures/programs. The question arises whether and how it can be applied in the context of Big Data. In this paper, we study how the well-founded semantics can be computed over huge amounts of data using mass parallelization. Specifically, we propose and evaluate a parallel approach based on the X10 programming language. Our experiments demonstrate that our approach has the ability to process up to 1 billion facts within minutes. Ilias Tachmazidis, Long Cheng 0003, Spyros Kotoulas, Grigoris Antoniou, Tomás Ward |
ICTAI | 2 |