Liting Hu

dblp:10/6462 · DBLP profile ↗
← Back
33ranked-venue papers
5as first author
15since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 18 · 5 first-author · 6 since 2021Computer networks · 5 · 5 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 1 since 2021Software engineering, systems software and programming languages · 3 · 2 since 2021Artificial intelligence and machine learning · 2Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 since 2021Databases, data management, data science and information retrieval · 1
YearPublicationVenuePosition
2026 A Distributed Learned Hash Table
abstract
Distributed Hash Tables (DHTs) are pivotal in numerous high-impact key-value applications built on distributed networked systems, offering a decentralized architecture that avoids single points of failure and improves data availability. Despite their widespread utility, DHTs face substantial challenges in handling range queries, which are crucial for applications such as LLM serving, distributed storage, databases, content delivery networks, and blockchains. To address this limitation, we present LEAD, a novel system incorporating learned models within DHT structures to significantly optimize range query performance. LEAD utilizes recursive machine learning models as the Learned Hash Function to map and retrieve data across a distributed system while preserving the inherent order of data. LEAD includes the designs to minimize range query latency and message cost while maintaining high scalability and resilience to network churn. Our comprehensive evaluations, conducted in both testbed implementation and simulations, demonstrate that LEAD achieves tremendous advantages in system efficiency compared to existing range query methods in large-scale distributed systems, reducing query latency and message cost by 80% to 90%+. Furthermore, LEAD exhibits scalability and robustness against system churn, providing a robust, scalable structure for efficient data retrieval in distributed key-value systems.
Shengze Wang 0004, Yi Liu 0115, Xiaoxue Zhang 0001, Liting Hu, Chen Qian 0001
IEEE Trans. Netw.4
2026 Totoro+: An Adaptive and Scalable Edge Federated Learning System
abstract
Federated Learning (FL) is an emerging distributed machine learning (ML) technique that enables in-situ model training and inference on decentralized edge devices. We propose Totoro$^+$, a novel scalable FL system that enables massive FL applications to run simultaneously on edge networks. The key insight is to explore a distributed hash table (DHT)-based peer-to-peer (P2P) model to re-architect the centralized FL system design into a fully decentralized one. In contrast to previous studies where many FL applications shared one centralized parameter server, Totoro$^+$assigns a dedicated parameter server to each application. Any edge node can act as any application's coordinator, aggregator, client selector, worker (participant device), or any combination of the above, thereby radically improving scalability and adaptivity. Totoro$^+$introduces three innovations to realize its design: a locality-aware P2P multi-ring structure, a publish/subscribe-based forest abstraction, and a game-theoretic path planning model with a guarantee of an$\epsilon$-approximate Nash equilibrium. Real-world experiments on 500 Amazon EC2 servers show that Totoro$^+$scales gracefully with the number of FL applications and$N$edge nodes speeds up the total training time by$1.2\times -14.0\times$, achieves$\mathcal {O}(\log N)$hops for model dissemination and gradient aggregation with millions of nodes, and efficiently adapts to the practical edge networks and churns.
Cheng-Wei Ching, Xin Chen 0084, Taehwan Kim 0012, Jian-Jhih Kuo, Dilma Da Silva, Liting Hu
IEEE Trans. Parallel Distributed Syst.6
2025 DNN-Adapt: Reinforcement Learning-Based Hybrid Batching for Efficient DNN Serving
abstract
Deep Neural Network (DNN) inference serving presents significant challenges due to variable workloads, heterogeneous models, and strict Service Level Objectives (SLOs). While GPU acceleration enables high-throughput inference, the cost-effectiveness of these specialized resources depends on maintaining high utilization. Current DNN serving systems employ static or heuristic-based batching strategies that cannot adapt effectively to dynamic cloud workloads. We present DNN-Adapt, a novel system that employs reinforcement learning (RL) to dynamically optimize batching decisions in DNN serving environments. DNN-Adapt introduces a hybrid batching framework that combines traditional batching techniques with RL-based decision making, a multi-timescale architecture that operates at different levels of granularity, and a hybrid decision system that integrates rule-based safety constraints with learned policies.
Milind Varma, Sai Venkat Malreddy, Liting Hu
CLOUD3
2025 A Distributed Learned Hash Table
abstract
Distributed Hash Tables (DHTs) are pivotal in numerous high-impact key-value applications built on distributed networked systems, offering a decentralized architecture that avoids single points of failure and improves data availability. Despite their widespread utility, DHTs face substantial challenges in handling range queries, which are crucial for applications such as LLM serving, distributed storage, databases, content delivery networks, and blockchains. To address this limitation, we present LEAD, a novel system incorporating learned models within DHT structures to significantly optimize range query performance. LEAD utilizes a recursive machine learning model to map and retrieve data across a distributed system while preserving the inherent order of data. LEAD includes the designs to minimize range query latency and message cost while maintaining high scalability and resilience to network churn. Our comprehensive evaluations, conducted in both testbed implementation and simulations, demonstrate that LEAD achieves tremendous advantages in system efficiency compared to existing range query methods in large-scale distributed systems, reducing query latency and message cost by 80% to 90%+. Furthermore, LEAD exhibits remarkable scalability and robustness against system churn, providing a robust, scalable solution for efficient data retrieval in distributed key-value systems.
Shengze Wang 0007, Yi Liu 0115, Xiaoxue Zhang 0001, Liting Hu, Chen Qian 0001
ICNP4
2025 Ekko: Fully Decentralized Scheduling for Serverless Edge Computing
abstract
While originally designed for the cloud, the benefits of the serverless paradigm are vital in Edge/Fog computing environments. In this paper, we propose Ekko, a novel decentralized edge serverless scheduling system, which enables a large number of serverless applications to run simultaneously at the edge through the Functionas-a-Service (FaaS) model. The key insight is to re-architect the common centralized or hierarchical scheduling systems into a fully decentralized one by using the distributed hash table (DHT) based peer-to-peer (P2P) model, in which many distributed schedulers operate autonomously without any centralized state. In sharp contrast to existing studies, any edge node in our system can act as a scheduler, a function worker, a query forwarder, or a storage node, and flexibly switch between these roles, thereby significantly improving scalability and adaptivity. Ekko introduces three design innovations: a boundary-aware P2P organization, distributed shadow schedulers with a keychain scheduling algorithm, and a distributed locality-aware bucket image store. Our evaluation on 500 Amazon EC2 nodes shows that, compared to the state-of-the-art, Ekko reduces the 90-th percentile tail queue wait time by up to 96.6 %, the scheduling time by up to 38.5 %, and the total deployment time by up to 89.5 %, while efficiently scaling to millions of function invocation requests on thousands of edge nodes.
Xin Chen 0084, Manoj Prabhakar Paidiparthy, Dilma Da Silva, Liting Hu
IPDPS4
2025 Capybara: an Edge-Friendly Distributed Object Store for Diverse Serverless Functions
abstract
While originally designed for the cloud, the benefits of the serverless paradigm are also vital in Edge/Fog computing environments. This paper presents Capybara, a new scalable, programmable distributed object store for storing and sharing serverless function data objects (state) on edge infrastructures. The key innovations here are (1) achieving scalability and avoiding the significant DRAM cost of indexing metadata servers through a "game-theoretic" DHT-based P2P architecture; (2) providing edge users with a "programmable" handler abstraction to customize data management policies, such as different function image caching policies, warm container "keep-alive" durations, data access control methods, and data replication policies.
Xin Chen 0084, Manoj Prabhakar Paidiparthy, Chen Qian 0001, Liting Hu
Middleware4
2025 AgileDART: An Agile and Scalable Edge Stream Processing Engine
abstract
Edge applications generate a large influx of sensor data on massive scales, and these massive data streams must be processed shortly to derive actionable intelligence. However, traditional data processing systems are not well-suited for these edge applications as they often do not scale well with a large number of concurrent stream queries, do not support low-latency processing under limited edge computing resources, and do not adapt to the level of heterogeneity and dynamicity commonly present in edge computing environments. As such, we present AgileDart, an agile and scalable edge stream processing engine that enables fast stream processing of many concurrently running low-latency edge applications' queries at scale in dynamic, heterogeneous edge environments. The novelty of our work lies in a dynamic dataflow abstraction that leverages distributed hash table-based peer-to-peer overlay networks to autonomously place, chain, and scale stream operators to reduce query latencies, adapt to workload variations, and recover from failures and a bandit-based path planning model that re-plans the data shuffling paths to adapt to unreliable and heterogeneous edge networks. We show that AgileDart outperforms Storm and EdgeWise on query latency and significantly improves scalability and adaptability when processing many real-world edge stream applications' queries.
Cheng-Wei Ching, Xin Chen 0084, Chaeeun Kim, Tongze Wang, Dong Chen 0025, Dilma Da Silva, Liting Hu
IEEE Trans. Mob. Comput.7
2024 Totoro: A Scalable Federated Learning Engine for the Edge
abstract
Federated Learning (FL) is an emerging distributed machine learning (ML) technique that enables in-situ model training and inference on decentralized edge devices. We propose Totoro, a novel scalable FL engine, that enables massive FL applications to run simultaneously on edge networks. The key insight is to explore a distributed hash table (DHT)-based peer-to-peer (P2P) model to re-architect the centralized FL system design into a fully decentralized one. In contrast to previous studies where many FL applications shared one centralized parameter server, Totoro assigns a dedicated parameter server to each individual application. Any edge node can act as any application's coordinator, aggregator, client selector, worker (participant device), or any combination of the above, thereby radically improving scalability and adaptivity. Totoro introduces three innovations to realize its design: a locality-aware P2P multi-ring structure, a publish/subscribe-based forest abstraction, and a bandit-based exploitation-exploration path planning model. Real-world experiments on 500 Amazon EC2 servers show that Totoro scales gracefully with the number of FL applications and N edge nodes, speeds up the total training time by 1.2 × -14.0×, achieves O (logN) hops for model dissemination and gradient aggregation with millions of nodes, and efficiently adapts to the practical edge networks and churns.
Cheng-Wei Ching, Xin Chen 0084, Taehwan Kim 0012, Bo Ji 0001, Qingyang Wang 0001, Dilma Da Silva, Liting Hu
EuroSys7
2024 Poster: Distributed Learned Hash Table
abstract
Distributed Hash Tables (DHTs) are pivotal in numerous high-impact key-value applications built on distributed networked systems, offering a decentralized architecture that avoids single points of failure and improves data availability. Despite their widespread utility, DHTs face substantial challenges in handling range queries, which are crucial for applications such as storage systems, decentralized databases, content distribution networks, and blockchains. To address this limitation, we present LEAD, a novel system incorporating learned models within DHT structures to significantly optimize range query performance. LEAD utilizes a recursive machine learning model to map and retrieve data across a distributed system while preserving the inherent order of data. Preliminary results indicate LEAD achieves tremendous advantages in system efficiency compared to existing range query methods in large-scale distributed systems while maintaining high scalability and resilience to network churn.
Shengze Wang 0007, Yi Liu 0115, Xiaoxue Zhang 0001, Liting Hu, Chen Qian 0001
ICNP4
2024 Safeguarding User-Centric Privacy in Smart Homes
abstract
Internet of Things (IoT) devices have been increasingly deployed in smart homes to automatically monitor and control their environments. Unfortunately, extensive recent research has shown that on-path external adversaries can infer and further fingerprint people’s sensitive private information by analyzing IoT network traffic traces. In addition, most recent approaches that aim to defend against these malicious IoT traffic analytics cannot adequately protect user privacy with reasonable traffic overhead. In particular, these approaches often did not consider practical traffic reshaping limitations, user daily routine permitting, and user privacy protection preference in their design. To address these issues, we design a new low-cost, open source user-centric defense system—PrivacyGuard—that enables people to regain the privacy leakage control of their IoT devices while still permitting sophisticated IoT data analytics that is necessary for smart home automation. In essence, our approach employs intelligent deep convolutional generative adversarial network assisted IoT device traffic signature learning, long short-term memory based artificial traffic signature injection, and partial traffic reshaping to obfuscate private information that can be observed in IoT device traffic traces. We evaluate PrivacyGuard using IoT network traffic traces of 31 IoT devices from five smart homes and buildings. We find that PrivacyGuard can effectively prevent a wide range of state-of-the-art adversarial machine learning and deep learning based user in-home activity inference and fingerprinting attacks and help users achieve the balance between their IoT data utility and privacy preserving.
Keyang Yu, Qi Li 0046, Dong Chen 0025, Liting Hu
ACM Trans. Internet Techn.4
2023 Sora: A Latency Sensitive Approach for Microservice Soft Resource Adaptation
abstract
Fast response time for modern web services that include numerous distributed and lightweight microservices becomes increasingly important due to its business impact. While hardware-only resource scaling approaches (e.g., FIRM [47] and PARSLO [40]) have been proposed to mitigate response time fluctuations on critical microservices, the re-adaptation of soft resources (e.g., threads or connections) that control the concurrency of hardware resource usage has been largely ignored. This paper shows that the soft resource adaptation of critical microservices has a significant impact on system scalability because either under- or over-allocation of soft resources can lead to inefficient usage of underlying hardware resources. We present Sora, an intelligent, fast soft resource adaptation management framework for quickly identifying and adjusting the optimal concurrency level of critical microservices to mitigate service-level objective (SLO) violations. Sora leverages online fine-grained system metrics and the propagated deadline along the critical path of request execution to quickly and accurately provide optimal concurrency setting for critical microservices. Based on six real-world bursty workload traces and two representative microservices benchmarks (Sock Shop and Social Network), our experimental results show that Sora can effectively mitigate large response time fluctuations and reduce the 99th percentile latency by up to 2.5× compared to the hardware-only scaling strategy FIRM [47] and 1.5× to the state-of-the-art concurrency-aware system scaling strategy ConScale.
Jianshu Liu, Qingyang Wang 0001, Shungeng Zhang, Liting Hu, Dilma Da Silva
Middleware4
2023 A novel data hiding scheme based on improved diamond encoding in IWT domain
Shihuan Sun, Liting Hu, Fanli Meng
Multim. Tools Appl.3
2023 Adaptive Fragment-Based Parallel State Recovery for Stream Processing Systems
abstract
Today, large-scale cloud organizations are deploying datacenters and “edge” clusters globally to provide low-latency access to services. Running stream applications across geo-distributed sites are emerging as a daily requirement. However, existing efforts have dominantly centered aroundstateless stream processing, leaving another urgent trend-stateful stream processing-much less explored. A driving need is to store and update states during processing, and most importantly, successfully recover large distributed states when faults and failures happen. Existing studies exhibit major limitations including: (1) they mostly inherit MapReduce's “single master/many workers” architecture, where the central master can easily become ascalability bottleneck; (2) they offer state recovery mainly through three approaches: replication recovery, checkpointing recovery, and DStream-based lineage recovery, which are either slow, resource-expensive or failing to handle multiple failures; and (3) they are not adaptive to heterogeneous hardware settings. We present A-FP4S, a novel adaptive fragments-based parallel state recovery mechanism for stream processing systems. A-FP4S organizes stream operators into a distributed hash table based peer-to-peer overlay and divides each node's local state into many fragments. These fragments are periodically stored in node's multiple neighbors, ensuring different sets of available fragments can reconstruct failed states in parallel. This mechanism is extremely scalable to the lost state, significantly reduces failure recovery time, and can tolerate multiple node failures. A-FP4S is adaptive to heterogeneous hardware settings by automatic parameter tuning over phases. Compared to Apache Storm, A-FP4S achieves 31.8% to 50.5% reduction in recovery latency. Large-scale experiments using real-world datasets demonstrate A-FP4S's attractive scalability and adaptivity properties.
Hailu Xu, Pinchao Liu, Sarker Tanzir Ahmed, Dilma Da Silva, Liting Hu
IEEE Trans. Parallel Distributed Syst.5
2022 Towards low-latency I/O services for mixed workloads using ultra-low latency SSDs
abstract
Low-latency I/O services are essential for latency-sensitive workloads when they co-run with throughput-oriented workloads in cloud data centers. Although advanced SSDs such as Intel Optane SSDs can offer ultra-low latency at the device layer, I/O interference among various workloads through the I/O stack can still significantly enlarge I/O latency. It is still an open problem to best utilize ultra-low latency SSDs in cloud computing environments.
Haikun Liu, Chencheng Ye 0001, Xiaofei Liao, Hai Jin 0001, Yu Zhang 0027, Liting Hu
ICS8
2021 DART: A Scalable and Adaptive Edge Stream Processing Engine
Pinchao Liu, Dilma Da Silva, Liting Hu
USENIX ATC3
2020 FP4S: Fragment-based Parallel State Recovery for Stateful Stream Applications
abstract
Streaming computations are by nature long-running. They run in highly dynamic distributed environments where many stream operators may leave or fail at the same time. Most of them are stateful, in which stream operators need to store and maintain large-sized state in memory, resulting in expensive time and space costs to recover them. The state-of-the-art stream processing systems offer failure recovery mainly through three approaches: replication recovery, checkpointing recovery, and DStream-based lineage recovery, which are either slow, resource-expensive or fail to handle many simultaneous failures.We present FP4S, a novel fragment-based parallel state recovery mechanism that can handle many simultaneous failures for a large number of concurrently running stream applications. The novelty of FP4S is that we organize all the application's operators into a distributed hash table (DHT) based consistent ring to associate each operator with a unique set of neighbors. Then we divide each operator's in-memory state into many fragments and periodically save them in each node's neighbors, ensuring that different sets of available fragments can reconstruct lost state in parallel. This approach makes this failure recovery mechanism extremely scalable, and allows it to tolerate many simultaneous operator failures. We apply FP4S on Apache Storm and evaluate it using large-scale real-world experiments, which demonstrate its scalability, efficiency, and fast failure recovery features. When compared to the state-of-the-art solutions (Apache Storm), FP4S reduces 37.8% latency of state recovery and saves more than half of the hardware costs. It can scale to many simultaneous failures and successfully recover the states when up to 66.6% of states fail or get lost.
Pinchao Liu, Hailu Xu, Dilma Da Silva, Qingyang Wang 0001, Sarker Tanzir Ahmed, Liting Hu
IPDPS6
2020 Max orientation coverage: efficient path planning to avoid collisions in the CNC milling of 3D objects
abstract
Most path planning algorithms for covering a complex 3D object ignore physical limitations or constraints on a robot's motion. Adhering to such constraints for a given path can slow down the time to cover the path because the motion may need to be adjusted. This work considers a scenario in computer numerical control (CNC) milling applications, where the robot is a cutting tool that needs to cover the surface of a complex 3D object under the following constraint: for every point on the generated path, the robot must be assigned an accessible orientation to avoid collisions between it and other parts of the object. Our proposed approach, which we call max orientation coverage, employs a two-step optimization scheme. It can improve path efficiency with respect to both the length of the path and the cost of dealing with the collision-avoiding constraints. We evaluate our approach through extensive simulation studies on four CAD benchmarks against a state-of-the-art baseline. We show that our proposed approach can improve the efficiency of the path by 29.7% on average compared with the baseline and the improvement goes up to 46.5% for certain complex objects.
Xin Chen 0096, Thomas M. Tucker, Thomas R. Kurfess, Richard W. Vuduc, Liting Hu
IROS5
2020 SR3: Customizable Recovery for Stateful Stream Processing Systems
abstract
Modern stream processing applications need to store and update state along with their processing, and process live data streams in a timely fashion from massive and geo-distributed data sets. Since they run in a dynamic distributed environment and their workloads may change in unexpected ways, multiple stream operators can fail at the same time, causing severe state loss. However, the state-of-the-art stream processing systems are mainly designed for low-latency intra-datacenter settings and do not scale well for running stream applications that contain large distributed states, suffering a significantly centralized bottleneck and high latency to recover state. They offer failure recovery mainly through three approaches: replication recovery, checkpointing recovery, and DStream-based lineage recovery, which are either slow, resource-expensive or fail to handle multiple simultaneous failures.
Hailu Xu, Pinchao Liu, Susana Cruz-Diaz, Dilma Da Silva, Liting Hu
Middleware5
2020 The Impact of Event Processing Flow on Asynchronous Server Efficiency
abstract
Asynchronous event-driven server architecture has been considered as a superior alternative to the thread-based counterpart due to reduced multithreading overhead. In this paper, we conduct empirical research on the efficiency of asynchronous Internet servers, showing that an asynchronous server may perform significantly worse than a thread-based one due to two design deficiencies. The first one is the widely adopted one-event-one-handler event processing model in current asynchronous Internet servers, which could generate frequent unnecessary context switches between event handlers, leading to significant CPU overhead of the server. The second one is a write-spin problem (i.e., repeatedly making unnecessary I/O system calls) in asynchronous servers due to some specific runtime workload and network conditions (e.g., large response size and non-trivial network latency). To address these two design deficiencies, we present a hybrid solution by exploiting the merits of different asynchronous architectures so that the server is able to adapt to dynamic runtime workload and network conditions in the cloud. Concretely, our hybrid solution applies a lightweight runtime request checking and seeks for the most efficient path to process each request from clients. Our results show that the hybrid solution can achieve from 10 to 90 percent higher throughput than all the other types of servers under the various realistic workload and network conditions in the cloud.
Shungeng Zhang, Qingyang Wang 0001, Yasuhiko Kanemasa, Huasong Shan, Liting Hu
IEEE Trans. Parallel Distributed Syst.5
2019 A data hiding scheme based on multidirectional line encoding and integer wavelet transform
Liting Hu
Signal Process. Image Commun.2
2019 Integrating Concurrency Control in n-Tier Application Scaling Management in the Cloud
abstract
Scaling complex distributed systems such as e-commerce is an importance practice to simultaneously achieve high performance and high resource efficiency in the cloud. Most previous research focuses on hardware resource scaling to handle runtime workload variation. Through extensive experiments using a representative n-tier web application benchmark (RUBBoS), we demonstrate that scaling an n-tier system by adding or removing VMs without appropriately re-allocating soft resources (e.g., server threads and connections) may lead to significant performance degradation resulting from implicit change of request processing concurrency in the system, causing either over- or under-utilization of the critical hardware resource in the system. We build a concurrency-aware model that determines a near optimal soft resource allocation of each tier by combining some operational queuing laws and the fine-grained online measurement data of the system. We then develop a dynamic concurrency management (DCM) framework that integrates the concurrency-aware model to intelligently reallocate soft resources in the system during the system scaling process. We compare DCM with Amazon EC2-AutoScale, the state-of-the-art hardware only scaling management solution using six real-world bursty workload traces. The experimental results show that DCM achieves significantly shorter tail latency and higher throughput compared to Amazon EC2-AutoScale under all the workload traces.
Qingyang Wang 0001, Shungeng Zhang, Liting Hu, Balaji Palanisamy
IEEE Trans. Parallel Distributed Syst.4
2018 A Toolset for Detecting Containerized Application's Dependencies in CaaS Clouds
abstract
There has been a dramatic increase in the popularity of Container as a Service (CaaS) clouds. The CaaS multi-tier applications could be optimized by using network topology, link or server load knowledge to choose the best endpoints to run in CaaS cloud. However, it is difficult to apply those optimizations to the public datacenter shared by multi-tenants. This is because of the opacity between the tenants and the datacenter providers: Providers have no insight into tenant's container workloads and dependencies, while tenants have no clue about the underlying network topology, link, and load. As a result, containers might be booted at wrong physical nodes that lead to performance degradation due to bi-section bandwidth bottleneck or co-located container interference. We propose 'DocMan', a toolset that adopts a black-box approach to discover container ensembles and collect information about intra-ensemble container interactions. It uses a combination of techniques such as distance identification and hierarchical clustering. The experimental results demonstrate that DocMan enables optimized containers placement to reduce the stress on bi-section bandwidth of the datacenter's network. The method can detect container ensembles at low cost and with 92% accuracy and significantly improve performance for multi-tier applications under the best of circumstances.
Pinchao Liu, Liting Hu, Hailu Xu, Jason Liu 0001, Qingyang Wang 0001, Jai Dayal, Yuzhe Tang
IEEE CLOUD2
2018 Oases: An Online Scalable Spam Detection System for Social Networks
abstract
Web-based social networks enable new community-based opportunities for participants to engage, share their thoughts, and interact with each other. Theses related activities such as searching and advertising are threatened by spammers, content polluters, and malware disseminators. We propose a scalable spam detection system, termed Oases, for uncovering social spam in social networks using an online and scalable approach. The novelty of our design lies in two key components: (1) a decentralized DHT-based tree overlay deployment for harvesting and uncovering deceptive spam from social communities; and (2) a progressive aggregation tree for aggregating the properties of these spam posts for creating new spam classifiers to actively filter out new spam. We design and implement the prototype of Oases and discuss the design considerations of the proposed approach. Our large-scale experiments using real-world Twitter data demonstrate scalability, attractive load-balancing, and graceful efficiency in online spam detection for social networks.
Hailu Xu, Liting Hu, Pinchao Liu, Jai Dayal, Qingyang Wang 0001, Yuzhe Tang
IEEE CLOUD2
2018 Harnessing the Nature of Spam in Scalable Online Social Spam Detection
abstract
Disinformation in social networks has been a worldwide problem. Social users are surrounded by a huge volume of malicious links, biased comments, fake reviews, or fraudulent advertisements, etc. Traditional spam detection approaches propose a variety of statistical feature-based models to filter out social spam from a historical dataset. However, they omit the real word situation of social data, that is, social spam is fast changing with new topics or events. Therefore, traditional approaches cannot effectively achieve online detection of the "drifting" social spam with a fixed statistic feature set. In this paper, we present Sifter, a system which can detect online social spam in a scalable manner without the labor-intensive feature engineering. The Sifter system is two-fold: (1) a decentralized DHT-based overlay deployment for harnessing the group characteristics of social spam activities within a specific topic/event; (2) a social spam processing with the support of Recurrent Neural Network (RNN) to get rid of the traditional manual feature engineering. Results show that Sifter achieves graceful spam detection performances with the minimal size of data and good balance in group management.
Hailu Xu, Boyuan Guan, Pinchao Liu, William Escudero, Liting Hu
IEEE BigData5
2017 Breaking Down Hadoop Distributed File Systems Data Analytics Tools: Apache Hive vs. Apache Pig vs. Pivotal HWAQ
abstract
Apache Hive, Apache Pig and Pivotal HWAQ are very popular open source cluster computing frameworks for large scale data analytics. These frameworks hide the complexity of task parallelism and fault-tolerance, by exposing a simple programming API to users. In this paper, we discuss the major architectural component differences in them and conduct detailed experiments to compare their performances with different inputs. Furthermore, we attribute these performance differences to different components which are architected differently in the three frameworks and we show the detailed execution overheads of Apache Hive, Apache Pig and Pivotal HAWQ, in which the CPU utilization, memory utilization, and disk read/write during their runtime are analyzed. Finally, a discussion and summary of our findings and suggestions are presented.
Liting Hu, Liangqi Liu, Diana Leante Bone
CLOUD2
2017 RBAY: A Scalable and Extensible Information Plane for Federating Distributed Datacenter Resources
abstract
While many institutions, whether industrial, academic, or governmental, satisfy their computing needs through public cloud providers, many others still manage their own resources, often as geographically distributed datacenters. Spare capacity from these geographically distributed datacenters could be offered to others, provided there were a mechanism to discover, and then request these resources. Unfortunately, single datacenter administrators tend not to cooperate due to issues of scalability, diverse administrative policies, and site-specific monitoring infrastructure. This paper describes RBAY, an integrated information plane that enables secure and scalable sharing between geographically distributed datacenters. RBAY's key design features are twofold. First, RBAY employs a decentralized `hierarchical aggregation tree' structure to seamlessly aggregate spare resources from geographically distributed datacenters to a global information plane. Second, RBAY attaches to each participating server a `admin-customized' handler, which follows site-specific policy to expose, hide, add, remove resources to RBAY, and thus fulfill the task of `which resource to expose to whom, when, and how'. An experimental evaluation on eight real-world geo-distributed sites demonstrates RBAY's rapid response to composite queries, as well as its extensible, scalable, and lightweight nature.
Liting Hu, Douglas M. Blough, Michael A. Kozuch, Matthew Wolf
ICDCS2
2014 ELF: Efficient Lightweight Fast Stream Processing at Scale
Liting Hu, Karsten Schwan, Hrishikesh Amur
USENIX ATC1
2013 FastMR: fast processing for large distributed data streams
abstract
FastMR is a graph-style framework for steam-oriented applications to realize near real-time streaming data record processing, and more importantly, complex coordinations between those applications. We introduces two components --- compressed buffer trees (CBTs) and shared reducer trees (SRTs) to assist with this task. CBTs address the problem of maintaining a significant amount of application-specific "accumulator" state in memory so that streaming data processing can combine current data with historical data. They do so by employing a novel, batch-oriented approach to updating the accumulator state. SRTs are basically P2P-based reducer trees that enable fine-grained queries (both one-shot and continual) to be efficiently rolled up concurrently. CBT's intermediate results are aggregated to the root of SRT via network aggregation. The roots of SRTs are analogous to vertices and anycast/multicast message transmission between the vertices (roots of SRTs) are analogous to edges in the graph-style computation model.
Liting Hu, Karsten Schwan, Hrishikesh Amur
SoCC1
2012 v-Bundle: Flexible Group Resource Offerings in Clouds
abstract
Traditional Infrastructure-as-a-Service offerings provide customers with large numbers of fixed-size virtual machine (VM) instances with resource allocations that are designed to meet application demands. With application demands varying over time, cloud providers gain efficiencies through resource consolidation and over-commitment. For cloud customers, however, this leads to inefficient use of the cloud resources they have purchased. To address cloud customers' dynamic application requirements, we present a new cloud resource offering, called v-Bundle, which makes flexible the exchange of resource capacity among multiple VM instances belonging to the same customer. Specifically targeting network resources, for each customer application, we first use DHT-based techniques to achieve an initial VM placement that minimizes its use of the data center network's bi-section bandwidth. When VMs' networking requirements change, the customer can then use v-Bundle to trade the networking resources allocated to her application. v-Bundle maintains information about network resources with any-cast tree-based methods implemented as extensions of the Pastry pub-sub core. Experimental evaluations show that the approach can scale well to thousands of hosts and VMs, and that v-Bundle can provide customers with better bandwidth utilization and improved application quality of service through borrowing extra bandwidth when needed, at no additional cost in terms of the total resources allocated to the customer.
Liting Hu, Kyung Dong Ryu, Dilma Da Silva, Karsten Schwan
ICDCS1
2010 Towards virtualized desktop environment
abstract
Abstract Virtualization is being widely used now as an emerging trend. Rapid improvements in network bandwidth, ubiquitous security hazards and high total cost of ownership of personal computers have created a growing market for desktop virtualization. Much like server virtualization, virtualizing desktops involves separating the physical location of a client device from its logical interface. But, the performance and usability of some traditional desktop frameworks do not satisfy end‐users. Other solutions, including WebOS, which needs to rebuild all daily‐used applications into Client/Server mode, cannot be easily accepted by people in a short time. We present LVD, a system that combines the virtualization technology and inexpensive personal computers (PCs) to realize a lightweight virtual desktop system. Comparing to the previous desktop systems, LVD builds an integrated novel desktop environment, which can support the backup, mobility, suspending and resuming of per‐user's working environment, and support synchronous using of incompatible applications on different platforms and achieves great saving in power consumption. We have implemented LVD in a cluster with Xen and compared its performance against widely used commercial approaches, including Microsoft RDP, Citrix MetaFrameXP and Sun Ray. Experimental results demonstrate that LVD is effective in performing the functions while imposing little overhead. Copyright © 2009 John Wiley & Sons, Ltd.
Xiaofei Liao, Hai Jin 0001, Liting Hu, Haikun Liu
Concurr. Comput. Pract. Exp.3
2009 Live migration of virtual machine based on full system trace and replay
abstract
Live migration of virtual machines (VM) across distinct physical hosts provides a significant new benefit for administrators of data centers and clusters. Previous migration schemes focused on transferring the runtime memory state of the VM. Those approaches employed memory pre-copy algorithm to synchronize the migrating VM states, which make VM live migration cost much network traffic and application downtime, especially for memory intensive workloads. This paper describes the design and implementation of a novel approach CR/TR-Motion that adopts checkpointing/recovery and trace/replay technology to provide fast, transparent VM migration. With execution trace logged on the source host, a synchronization algorithm is performed to orchestrate the running source and target VM until they get a consistent state. We also give a formalized characterization about the migration evaluation metrics and make a mathematical analysis about our algorithm. Our scheme can greatly reduce the migration downtime and network bandwidth consumption. Experimental measurements show that our approach can drastically reduce migration overheads compared with pre-copy algorithm: up to 72.4% on application observed downtime, up to 31.5% on total migration time and up to 95.9% on the data to synchronize the VM state, while the application performance overhead due to migration is less than 8.54% on average.
Haikun Liu, Hai Jin 0001, Xiaofei Liao, Liting Hu, Chen Yu 0003
HPDC4
2009 Implementation of G.729 Codec Based on Mediastreamer Technology
abstract
Currently, the mediastreamer2 contains internal support for G.711u, G.711a, speex and gsm. However, it does not contain internal support for the popular G.729, which is known as the best audio codec ever. In this paper, we present the method of how to implement a G.729 codec filter on mediastreamer2. Then, we evaluate the performance of the G.729 filter in our platform. The experimental results show that our implementation of G.729 codec on mediastreamer2 is successful on the whole.
Liting Hu, Xiangping Kong, Lianfen Huang, Hezhi Lin, Xueyuan Jiang
NAS1
2008 Magnet: A novel scheduling policy for power reduction in cluster with virtual machines
abstract
The concept of green computing has attracted much attention recently in cluster computing. However, previous local approaches focused on saving the energy cost of the components in a single workstation without a global vision on the whole cluster, so it achieved undesirable power reduction effect. Other cluster-wide energy saving techniques could only be applied to homogeneous workstations and specific applications. This paper describes the design and implementation of a novel approach that uses live migration of virtual machines to transfer load among the nodes on a multilayer ring-based overlay. This scheme can reduce the power consumption greatly by regarding all the cluster nodes as a whole. Plus, it can be applied to both the homogeneous and heterogeneous servers. Experimental measurements show that the new method can reduce the power consumption by 74.8% over base at most with certain adjustably acceptable overhead. The effectiveness and performance insights are also analytically verified.
Liting Hu, Hai Jin 0001, Xiaofei Liao, Xianjie Xiong, Haikun Liu
CLUSTER1