EDBT 2026 Demo / reviewers in the wild / expert
Hai Huang 0002
dblp:51/944-2
· DBLP profile ↗
26ranked-venue papers
7as first author
3since 2021 · last 2021
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 20 · 5 first-author · 3 since 2021Security and privacy · 4Computer networks · 2 · 1 first-authorSoftware engineering, systems software and programming languages · 1 · 1 first-authorDatabases, data management, data science and information retrieval · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2021 | Iter8: Online Experimentation in the CloudabstractOnline experimentation is an agile software development practice that plays an essential role in enabling rapid innovation. Existing solutions for online experimentation in Web and mobile applications are unsuitable for cloud applications. There is a need for rethinking online experimentation in the cloud to advance the state-of-the-art by considering the unique challenges posed by cloud environments. Mert Toslali, Srinivasan Parthasarathy 0001, Fábio Oliveira, Hai Huang 0002, Ayse K. Coskun |
SoCC | 4 |
| 2021 | Virtual Machine Extrospection: A Reverse Information Retrieval in CloudsabstractIn a virtualized environment, it is not difficult to retrieve guest OS information from its hypervisor. However, it is very challenging to retrieve information in the reverse direction, i.e., retrieve the hypervisor information from within a guest OS, which remains an open problem and has not yet been comprehensively studied before. In this paper, we take the initiative and study this reverse information retrieval problem. In particular, we investigate how to determine the host OS kernel version from within a guest OS. We observe that modern commodity hypervisors introduce new features and bug fixes in almost every new release. Thus, by carefully analyzing the seven-year evolution of Linux KVM development (including 3,485 patches), we can identify 19 features and 20 bugs in the hypervisor detectable from within a guest OS. Building on our detection of these features and bugs, we present a novel framework called Hyperprobe that for the first time enables users in a guest OS to automatically detect the underlying host OS kernel version in a few minutes. We implement a prototype of Hyperprobe and evaluate its effectiveness in six real world clouds, including Google Compute Engine (a.k.a. Google Cloud), HP Helion Public Cloud, ElasticHosts, Joyent Cloud, CloudSigma, and VULTR, as well as in a controlled testbed environment, all yielding promising results. Jidong Xiao, Hai Huang 0002, Haining Wang 0001 |
IEEE Trans. Cloud Comput. | 3 |
| 2021 | Cryptomining Detection in Container Clouds Using System Calls and Explainable Machine LearningabstractThe use of containers in cloud computing has been steadily increasing. With the emergence of Kubernetes, the management of applications inside containers (or pods) is simplified. Kubernetes allows automated actions like self-healing, scaling, rolling back, and updates for the application management. At the same time, security threats have also evolved with attacks on pods to perform malicious actions. Out of several recent malware types, cryptomining has emerged as one of the most serious threats with its hijacking of server resources for cryptocurrency mining. During application deployment and execution in the pod, a cryptomining process, started by a hidden malware executable can be run in the background, and a method to detect malicious cryptomining software running inside Kubernetes pods is needed. One feasible strategy is to use machine learning (ML) to identify and classify pods based on whether or not they contain a running process of cryptomining. In addition to such detection, the system administrator will need an explanation as to the reason(s) of the ML's classification outcome. The explanation will justify and support disruptive administrative decisions such as pod removal or its restart with a new image. In this article, we describe the design and implementation of an ML-based detection system of anomalous pods in a Kubernetes cluster by monitoring Linux-kernel system calls (syscalls). Several types of cryptominers images are used as containers within an anomalous pod, and several ML models are built to detect such pods in the presence of numerous healthy cloud workloads. Explainability is provided using SHAP, LIME, and a novel auto-encoding-based scheme for LSTM models. Seven evaluation metrics are used to compare and contrast the explainable models of the proposed ML cryptomining detection engine. Rupesh Raj Karn, Prabhakar Kudva, Hai Huang 0002, Sahil Suneja, Ibrahim M. Elfadel |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2020 | Customizable Scale-Out Key-Value StoresabstractEnterprise KV stores are often not well suited for HPC applications, and thus cumbersome end-to-end KV design customization is required to meet the needs of modern HPC applications. To this end, in this article we present bespoKV, an adaptive, extensible, and scale-out KV store framework. bespoKV decouples the KV store design into the control plane for distributed management and the data plane for local data store. For the control plane, bespoKVprovides pre-built modules, called controlets, supporting common distributed functionalities (e.g., replication, consistency, and topology) and their various combinations. This decoupling allows bespoKV to take a user-provided single-server KV store, called a datalet, and transparently enables a scalable and fault-tolerant distributed KV store service. The resulting distributed stores are also adaptive to consistency or topology requirement changes and can be easily extended for new types of services. Such specializations enable innovative uses of KV stores in HPC applications, especially for emerging applications that utilize KV-friendly workloads. We evaluate bespoKV in a local testbed as well as in a public cloud settings. Experiments show that bespoKV-enabled distributed KV stores scale horizontally to a large number of nodes, and performs comparably and sometimes 1.2× to 2.6× better than the state-of-the-art systems. Ali Anwar 0001, Yue Cheng 0001, Hai Huang 0002, Jingoo Han, Hyogi Sim, Fred Douglis, Ali Raza Butt |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2019 | Meteor: Optimizing spark-on-yarn for short applications
Hong Zhang 0047, Hai Huang 0002, Liqiang Wang 0001 |
Future Gener. Comput. Syst. | 2 |
| 2018 | bespoKV: application tailored scale-out key-value stores
Ali Anwar 0001, Yue Cheng 0001, Hai Huang 0002, Jingoo Han, Hyogi Sim, Fred Douglis, Ali Raza Butt |
SC | 3 |
| 2017 | MRapid: An Efficient Short Job Optimizer on HadoopabstractData have been generated and collected at an accelerating pace. Hadoop has made analyzing large scale data much simpler to developers/analysts using commodity hardware. Interestingly, it has been shown that most Hadoop jobs have small input size and do not run for long time. For example, higher level query languages, such as Hive and Pig, would handle a complex query by breaking it into smaller adhoc ones. Although Hadoop is designed for handling complex queries with large data sets, we found that it is highly inefficient to operate at small scale data, despite a new Uber mode was introduced specifically to handle jobs with small input size. In this paper, we propose an optimized Hadoop extension called MRapid, which significantly speeds up the execution of short jobs. It is completely backward compatible to Hadoop, and imposes negligible overhead. Our experiments on Microsoft Azure public cloud show that MRapid can improve performance by up to 88% compared to the original Hadoop. Hong Zhang 0047, Hai Huang 0002, Liqiang Wang 0001 |
IPDPS | 2 |
| 2016 | ClusterOn: Building Highly Configurable and Reusable Clustered Data Services Using Simple Data Nodes
Ali Anwar 0001, Yue Cheng 0001, Hai Huang 0002, Ali Raza Butt |
HotStorage | 3 |
| 2016 | Hyperprobe: Towards Virtual Machine Extrospection
Jidong Xiao, Hai Huang 0002, Haining Wang 0001 |
USENIX ATC | 3 |
| 2015 | Hyperprobe: Towards Virtual Machine Extrospection
Jidong Xiao, Hai Huang 0002, Haining Wang 0001 |
LISA | 3 |
| 2015 | Defeating Kernel Driver Purifier
Jidong Xiao, Hai Huang 0002, Haining Wang 0001 |
SecureComm | 2 |
| 2015 | Kernel Data Attack Is a Realistic Security Threat
Jidong Xiao, Hai Huang 0002, Haining Wang 0001 |
SecureComm | 2 |
| 2014 | You Are How You Touch: User Verification on Smartphones via Tapping BehaviorsabstractSmartphone users have their own unique behavioral patterns when tapping on the touch screens. These personal patterns are reflected on the different rhythm, strength, and angle preferences of the applied force. Since smart phones are equipped with various sensors like accelerometer, gyroscope, and touch screen sensors, capturing a user's tapping behaviors can be done seamlessly. Exploiting the combination of four features (acceleration, pressure, size, and time) extracted from smart phone sensors, we propose a non-intrusive user verification mechanism to substantiate whether an authenticating user is the true owner of the smart phone or an impostor who happens to know the pass code. Based on the tapping data collected from over 80 users, we conduct a series of experiments to validate the efficacy of our proposed system. Our experimental results show that our verification system achieves high accuracy with averaged equal error rates of down to 3.65%. As our verification system can be seamlessly integrated with the existing user authentication mechanisms on smart phones, its deployment and usage are transparent to users and do not require any extra hardware support. Hai Huang 0002, Haining Wang 0001 |
ICNP | 3 |
| 2014 | SMARTH: Enabling Multi-pipeline Data Transfer in HDFSabstractHadoop is a popular open-source implementation of the MapReduce programming model to handle large data sets, and HDFS is one of Hadoop's most commonly used distributed file systems. Surprisingly, we found that HDFS is inefficient when handling upload of data files from client local file system, especially when the storage cluster is configured to use replicas. The root cause is HDFS's synchronous pipeline design. In this paper, we introduce an improved HDFS design called SMARTH. It utilizes asynchronous multi-pipeline data transfers instead of a single pipeline stop-and-wait mechanism. SMARTH records the actual transfer speed of data blocks and sends this information to the namenode along with periodic heartbeat messages. The namenode sorts datanodes according to their past performance and tracks this information continuously. When a client initiates an upload request, the namenode will send it a list of "high performance" datanodes that it thinks will yield the highest throughput for the client. By choosing higher performance datanodes relative to each client and by taking advantage of the multi-pipeline design, our experiments show that SMARTH significantly improves the performance of data write operations compared to HDFS. Specifically, SMARTH is able to improve the throughput of data transfer by 27-245% in a heterogeneous virtual cluster on Amazon EC2. Hong Zhang 0047, Liqiang Wang 0001, Hai Huang 0002 |
ICPP | 3 |
| 2013 | Security implications of memory deduplication in a virtualized environmentabstractMemory deduplication has been widely used in various commodity hypervisors. By merging identical memory contents, it allows more virtual machines to run concurrently on top of a hypervisor. However, while this technique improves memory efficiency, it has a large impact on system security. In particular, memory deduplication is usually implemented using a variant of copy-on-write techniques, for which, writing to a shared page would incur a longer access time than those non-shared. In this paper, we investigate the security implication of memory deduplication from the perspectives of both attackers and defenders. On one hand, using the artifact above, we demonstrate two new attacks to create a covert channel and detect virtualization, respectively. On the other hand, we also show that memory deduplication can be leveraged to safeguard Linux kernel integrity. Jidong Xiao, Zhang Xu, Hai Huang 0002, Haining Wang 0001 |
DSN | 3 |
| 2012 | A covert channel construction in a virtualized environmentabstractMemory deduplication has been widely used in various commodity hypervisors. However, while this technique improves memory efficiency, it has an impact on system security. In particular, memory deduplication is usually implemented using a variant of copy-on-write techniques, for which, writing to a shared page would incur a longer access time than those non-shared. By exploiting this artifact, we demonstrate a new covert channel can be built in a virtualized environment. Jidong Xiao, Zhang Xu, Hai Huang 0002, Haining Wang 0001 |
CCS | 3 |
| 2012 | Understanding performance implications of nested file systems in a virtualized environment
Duy Le 0005, Hai Huang 0002, Haining Wang 0001 |
FAST | 2 |
| 2010 | Splitter: a proxy-based approach for post-migration testing of web applicationsabstractThe benefits of virtualized IT environments, such as compute clouds, have drawn interested enterprises to migrate their applications onto new platforms to gain the advantages of reduced hardware and energy costs, increased flexibility and deployment speed, and reduced management complexity. However, the process of migrating a complex application takes a considerable amount of effort, particularly when performing post-migration testing to verify that the application still functions correctly in the target environment. The traditional approach of test case generation and execution can take weeks and synthetic test cases may not adequately reflect actual application usage. Xiaoning Ding, Hai Huang 0002, Yaoping Ruan, Anees Shaikh, Brian Peterson, Xiaodong Zhang 0001 |
EuroSys | 2 |
| 2009 | Building end-to-end management analytics for enterprise data centersabstractThe complexity of modern data centers has evolved significantly in recent years. One typically is comprised of a large number and types of middleware and applications that are hosted in a heterogeneous pool of both physical and virtual servers, connected by a complex web of virtual and physical networks. Therefore, to manage everything in a data center, system administrators usually need a plethora of management tools since one tool often manages only one type of devices. The boundaries between the different management tools can limit productivity of system administrators on their daily tasks as each tool only offers a partial view of the entire managed environment. As a result, advanced analytics such as impact analysis and problem determination are generally not achievable using the traditional management tools as they require a holistic view of the entire data center. In this paper, we describe an integrated management system for applications, servers, network and storage devices called DataGraph. Our system integrates data across heterogeneous point products and agents for management and monitoring to enable the above mentioned management analytics capabilities. A common data model is introduced to federate data collected by the different tools in multiple database repositories so no modifications are needed to existing management tools. A common integrated web user interface is implemented to facilitate management tasks that would otherwise require invoking multiple tools. We deployed this tool in a lab environment and demonstrated these analytics capabilities through several case studies. Hai Huang 0002, Yaoping Ruan, Anees Shaikh, Ramani Routray, Chung-Hao Tan, Sandeep Gopisetty |
Integrated Network Management | 1 |
| 2008 | Automatic Software Fault Diagnosis by Exploiting Application Signatures
Xiaoning Ding, Hai Huang 0002, Yaoping Ruan, Anees Shaikh, Xiaodong Zhang 0001 |
LISA | 2 |
| 2007 | PDA: A Tool for Automated Problem Determination
Hai Huang 0002, Raymond B. Jennings III, Yaoping Ruan, Ramendra K. Sahoo, Sambit Sahu, Anees Shaikh |
LISA | 1 |
| 2007 | Partial Disk Failures: Using Software to Analyze Physical Damage
Hai Huang 0002, Kang G. Shin |
MSST | 1 |
| 2005 | Improving energy efficiency by making DRAM less randomly accessedabstractExisting techniques manage power for the main memory by passively monitoring the memory traffic, and based on which, predict when to power down and into which low-power state to transition. However, passively monitoring the memory traffic can be far from being effective as idle periods between consecutive memory accesses are often too short for existing power-management techniques to take full advantage of the deeper power-saving state implemented in modern DRAM architectures. In this paper, we propose a new technique that will actively reshape the memory traffic to coalesce short idle periods --- which were previously unusable for power management --- into longer ones, thus enabling existing techniques to effectively exploit idleness in the memory Hai Huang 0002, Kang G. Shin, Charles Lefurgy, Tom W. Keller |
ISLPED | 1 |
| 2005 | FS2: dynamic data replication in free disk space for improving disk performance and energy consumptionabstractDisk performance is increasingly limited by its head positioning latencies, i.e., seek time and rotational delay. To reduce the head positioning latencies, we propose a novel technique that dynamically places copies of data in file system's free blocks according to the disk access patterns observed at runtime. As one or more replicas can now be accessed in addition to their original data block, choosing the "nearest" replica that provides fastest access can significantly improve performance for disk I/O operations.We implemented and evaluated a prototype based on the popular Ext2 file system. In our prototype, since the file system layout is modified only by using the free/unused disk space (hence the name Free Space File System, or FS2), users are completely oblivious to how the file system layout is modified in the background; they will only notice performance improvements over time. For a wide range of workloads running under Linux, FS2 is shown to reduce disk access time by 41--68% (as a result of a 37--78% shorter seek time and a 31--68% shorter rotational delay) making a 16--34% overall user-perceived performance improvement. The reduced disk access time also leads to a 40--71% energy savings per access. Hai Huang 0002, Wanda Hung, Kang G. Shin |
SOSP | 1 |
| 2003 | Design and Implementation of Power-Aware Virtual Memory
Hai Huang 0002, Padmanabhan Pillai, Kang G. Shin |
USENIX ATC, General Track | 1 |
| 2002 | Improving Wait-Free Algorithms for Interprocess Communication in Embedded Real-Time Systems
Hai Huang 0002, Padmanabhan Pillai, Kang G. Shin |
USENIX ATC, General Track | 1 |