EDBT 2026 Demo / reviewers in the wild / expert
Witawas Srisa-an
dblp:72/4170
· DBLP profile ↗
63ranked-venue papers
6as first author
10since 2021 · last 2025
0000-0003-0021-5696ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 30 · 2 first-author · 1 since 2021Systems, architecture and hardware · 15 · 4 first-author · 2 since 2021Computer networks · 7 · 1 first-author · 3 since 2021Security and privacy · 7 · 3 since 2021Artificial intelligence and machine learning · 4Databases, data management, data science and information retrieval · 2Human-computer interaction and ubiquitous computing · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Enabling Symbolic Execution for Hardware TCP/IP Stack based on AMD Vitis HLSabstractHardware TCP/IP stacks, which directly implement TCP/IP functionality in hardware, have gained increasing attention due to their ability to meet the performance requirements of rapidly growing network speeds while significantly reducing CPU overhead. However, comprehensively testing these hardware implementations remains challenging because of their prohibitively large test input spaces involving diverse packet contents and complex packet dynamics. Symbolic execution, a powerful program analysis technique, has successfully improved testing coverage in software TCP/IP stacks but has not yet been widely adopted for hardware TCP/IP stacks. This paper addresses this gap by enabling symbolic execution to systematically test hardware TCP/IP stacks based on AMD Vitis High-Level Synthesis (HLS). We identify key challenges in applying symbolic execution in this hardware context and propose methods to overcome them. Evaluations on a real-world open-source hardware TCP/IP stack demonstrate the effectiveness of our methods in achieving high test coverage and discovering previously undetected bugs. Nianhang Hu, Tate Koziol, Witawas Srisa-an, Lisong Xu |
ICCCN | 3 |
| 2025 | WIA-SZZ: Work item aware SZZ
Salomé Perez-Rosero, Robert Dyer 0001, Samuel W. Flint, Shane McIntosh, Witawas Srisa-an |
Empir. Softw. Eng. | 5 |
| 2025 | AVOID: Automated Void Detection in STL FilesabstractAdditive manufacturing is a multi-billion dollar industry 21.58 billion in 2024), so its processes should be dependable and secure. Malicious actors can inject negative spaces, known as voids, in STL files, which can have a devastating impact on a final product's quality. Current ways of detecting voids use machine sensor or simulation data, and physical verification measures after printing. However, to the best of our knowledge, no method exists for detecting hidden voids solely at the STL level. Void detection at this level is inexpensive, and has the potential to detect voids in designs en masse. In this work, we proposeAVOID, a new approach to detect hidden voids. It both detects voids, and assesses their risk of weakening the manufactured part.AVOIDperforms risk assessment based on the size and location of each detected void. We empirically evaluatedAVOIDusing several large datasets, in total thousands of STL files, each with multiple hidden voids. We found thatAVOIDis highly accurate, with 97.7% recall, 98.1% precision, and 99% F1 on average across six data sets.AVOIDis also robust, scalable, and efficient. We find that high risk voids account for approximately 7% of all detected voids inserted randomly. Sarah Roscoe, Logan Hellbusch, Chamath Gunawardena, Jitender S. Deogun, Witawas Srisa-an, Yi Qian 0001, Gabriela F. Ciocarlie |
IEEE Trans. Dependable Secur. Comput. | 5 |
| 2023 | Efficient Verification of Timing-Related Network Functions in High-Speed HardwareabstractTo achieve a line rate in the high-speed environment of modern networks, there is a continuing effort to offload network functions from software to programmable hardware (HW). Although the offloading effort has led to greater performance, it brings difficulty in the verification of timing-related network functions (Time-NFs) as well. Time-NFs use numerical timing values to perform various network tasks. For example, congestion control algorithm BBR uses round-trip time to improve throughput. Errors in Time-NFs could cause packet loss and poor throughput. However, verifying Time-NFs in HW often involves many clock cycles that can result in an exponentially increasing number of test cases. Current verification methods either do not scale or sacrifice soundness for scalability.In this paper, we propose an invariant-based method to improve the verification efficiency without losing soundness. Our method is motivated by an observation that most Time-NFs follow a few fixed patterns to use timing information. Based on these patterns, we develop a set of easy-to-validate invariants to constrain the examination space. According to experiments on real Time-NFs, our method can speed up verification by 7 times on average without losing the verification soundness. Tianqi Fang, Lisong Xu, Witawas Srisa-an |
INFOCOM | 3 |
| 2023 | Evaluation of the ProgHW/SW Architectural Design Space of Bandwidth Estimation
Tianqi Fang, Lisong Xu, Witawas Srisa-an |
PAM | 3 |
| 2022 | SAINTDroid: Scalable, Automated Incompatibility Detection for AndroidabstractWith the ever-increasing popularity of mobile devices over the last decade, mobile applications and the frameworks upon which they are built frequently change, leading to a confusing jumble of devices and applications utilizing differing features even within the same framework. For Android apps and devices—the largest such framework and marketplace— mismatches between the version of the app API installed on a device and the version targeted by the developers of an app running on that device can lead to run-time crashes, providing a poor user experience. This paper presents SAINTDroid, a holistic compatibility analysis approach that seamlessly examines both the application code and the framework code by gradually loading and analyzing classes as needed during the compatibility analysis to enable efficient and scalable identification of various types of crash-leading Android compatibility issues. We applied SAINTDroid to 3,590 real-world apps and compared the analysis results against the state-of-the-art techniques, which corroborates that SAINTDroid is up to 76% more successful in detecting compatibility issues while issuing significantly fewer false alarms. The experimental results also show that SAINTDroid is remarkably (up to 8.3 times and four times on average) faster than the state-of-the-art techniques. Bruno Vieira Resende e Silva, Clay Stevens, Niloofar Mansoor, Witawas Srisa-an, Tingting Yu 0001, Hamid Bagheri |
DSN | 4 |
| 2021 | ReHAna: An Efficient Program Analysis Framework to Uncover Reflective Code in Android
Shakthi Bachala, Yutaka Tsutano, Witawas Srisa-an, Gregg Rothermel, Jackson Dinh, Yuanjiu Hu |
MobiQuitous | 3 |
| 2021 | SEMEO: A Semantic Equivalence Analysis Framework for Obfuscated Android Applications
Bruno Vieira Resende e Silva, Hamid Bagheri, Witawas Srisa-an, Gregg Rothermel, Jackson Dinh |
MobiQuitous | 4 |
| 2021 | Kollector: Detecting Fraudulent Activities on Mobile Devices Using Deep LearningabstractWith the rapid growth in smartphone usage, preventing leakage of personal information and privacy has become a challenging task. One major consequence of such leakage is impersonation. This type of illegal usage is nearly impossible to prevent as existing preventive mechanisms (e.g., passcode and fingerprinting), are not capable of continuously monitoring usage and determining whether the user is authorized. Once unauthorized users can defeat the initial protection mechanisms, they would have full access to the devices including using stored passwords to access high-value websites. We present Kollector, a new framework to detect impersonation based on a multi-view bagging deep learning approach to capture sequential tapping information on the smart-phone's keyboard. We construct a sequential-tapping biometrics model to continuously authenticate the user while typing. We empirically evaluated our system using real-world phone usage sessions from 26 users over eight weeks. We then compared our model against commonly used shallow machine techniques and find that our system performs better than other approaches and can achieve an 8.42 percent equal error rate, a 94.24 percent accuracy and a 94.41 percent H-mean using only the accelerometer and only five keyboard taps. We also experiment with using only three keyboard taps and find that the system still yields high accuracy while giving additional opportunities to make more decisions that can result in more accurate final decisions. Lichao Sun 0001, Bokai Cao, Ji Wang 0002, Witawas Srisa-an, Philip S. Yu, Alex D. Leow, Stephen Checkoway |
IEEE Trans. Mob. Comput. | 4 |
| 2021 | Improving the Performance of Deduplication-Based Storage Cache via Content-Driven Cache Management MethodsabstractData deduplication, as a proven technology for effective data reduction in backup and archiving storage systems, is also showing promises in increasing the logical space capacity for storage caches by removing redundant data. However, our in-depth evaluation of the existing deduplication-aware caching algorithms reveals that they only work well when the cached block size is set to 4 KB. Unfortunately, modern storage systems often set the block size to be much larger than 4 KB, and in this scenario, the overall performance of these caching schemes drops below that of the conventional replacement algorithms without any deduplication. There are several reasons for this performance degradation. The first reason is the deduplication overhead, which is the time spent on generating the data fingerprints and their use to identify duplicate data. Such overhead offsets the benefits of deduplication. The second reason is the extremely low cache space utilization caused by read and write alignment. The third reason is that existing algorithms only exploit access locality to identify block replacement. There is a lost opportunity to effectively leverage the content usage patterns such as intensity of content redundancy and sharing in deduplication-based storage caches to further improve performance. We propose CDAC, a Content-driven Deduplication-Aware Cache, to address this problem. CDAC focuses on exploiting the content redundancy in blocks and intensity of content sharing among source addresses in cache management strategies. We have implemented CDAC based on LRU and ARC algorithms, called CDAC-LRU and CDAC-ARC respectively. Our extensive experimental results show that CDAC-LRU and CDAC-ARC outperform the state-of-the-art deduplication-aware caching algorithms, D-LRU, and D-ARC, by up to 23.83X in read cache hit ratio, with an average of 3.23X, and up to 53.3 percent in IOPS, with an average of 49.8 percent, under a real-world mixed workload when the cache size ranges from 20 to 50 percent of the workload size and the block size ranges from 4KB to 32 KB. Yujuan Tan, Congcong Xu, Zhichao Yan 0001, Hong Jiang 0001, Witawas Srisa-an, Xianzhang Chen, Duo Liu 0002 |
IEEE Trans. Parallel Distributed Syst. | 6 |
| 2020 | Automated Field-based Decomposition to Accelerate Model Checking FPGA-based TCP/IPabstractThere is a rising effort to move the full TCP/IP stack from the software to the hardware to improve the network performance and programmability further. The hardware-based TCP/IP stack must be utterly correct since TCP is the foundation of many critical applications. However, it is impractical to apply the conventional model checking methods or decomposition techniques to verify the correctness because there are numerous stateless and stateful functions involved in TCP/IP stack. Therefore, we propose an automated field-based decomposition method to make feasible the model checking of the hardware-based TCP/IP. Our method can significantly mitigate the state space explosion issue while maintaining the verification completeness. We choose FPGA, a popular programmable hardware, as our study object in the paper. Our method effectiveness is demonstrated by the verification of both stateful and stateless functions of the FPGA-based TCP/IP. Tianqi Fang, Lisong Xu, Witawas Srisa-an |
ICC | 3 |
| 2020 | DINA: Detecting Hidden Android Inter-App Communication in Dynamic Loaded CodeabstractAndroid inter-app communication (IAC) allows apps to request functionalities from other apps, which has been extensively used to provide a better user experience. However, IAC has also become an enticing target by attackers to launch malicious activities. Dynamic class loading (DCL) and reflection are effective features to enhance the functionality of the apps. In this paper, we expose a new attack that leverages these features in conjunction with inter-app communication to conceal malicious attacks with the ability to bypass existing security mechanisms. To counteract such attack, we present DINA, a novel hybrid analysis approach for identifying malicious IAC behaviors concealed within dynamically loaded code through reflective/DCL calls. DINA appends reflection and DCL invocations to control-flow graphs and continuously performs incremental dynamic analysis to detect the misuse of reflection and DCL that obfuscates malicious Intent communications. DINA utilizes string analysis and inter-procedural analysis to resolve hidden IAC and achieves superior detection performance. Our extensive evaluation on 49,000 real-world apps corroborates the prevalent usage of reflection and DCL, and reveals previously unknown and potentially harmful, hidden IAC behaviors in real-world apps. Mohannad Alhanahnah, Qiben Yan 0001, Hamid Bagheri, Hao Zhou 0043, Yutaka Tsutano, Witawas Srisa-an, Xiapu Luo |
IEEE Trans. Inf. Forensics Secur. | 6 |
| 2020 | APMigration: Improving Performance of Hybrid Memory Performance via An Adaptive Page Migration MethodabstractByte-addressable, non-volatile memory (NVRAM) combines the benefits of DRAM and flash memory. However, due to its slower speed than DRAM, it is best to deploy it in combination with typical DRAM. In such Hybrid NVRAM systems, frequently accessed, hotpages can be stored in DRAM while other cold pages can reside in NVRAM, providing the benefits of both high performance (from DRAM) and lower power consumption and cost/performance (from NVRAM). While the idea seems beneficial, realizing an efficient hybrid NVRAM system requires careful page migration and accurate data temperature measurement. Existing solutions, however, often cause invalid migrations due to inaccurate data temperature accounting, because hot and cold pages are separately identified in DRAM and NVRAM regions. Moreover, since a new NVRAM frame is always allocated for each page swapped back NVRAM, a large amount of unnecessary NVRAM writes are generated during each page migration. Based on these observations, we propose APMigrate, an adaptive data migration approach for hybrid NVRAM systems. APMigrate consist of two parts, UIMigrate and LazyWriteback. UIMigrate focuses on eliminating invalid page migrations by considering data temperature in the entire DRAM-NVRAM space, while LazyWriteback focus on rewriting only dirty data back when the page is swapped back to NVRAM. Our experiments using SPEC 2006 show that APMigrate can reduce the number of migrations and improves performance by up to 90 percent compared to existing state-of-the-art approaches. For some workloads, LazyWriteback can reduce unnecessary NVRAM writes for existing page migrations by up to 75 percent. Yujuan Tan, Baiping Wang, Zhichao Yan 0001, Witawas Srisa-an, Xianzhang Chen, Duo Liu 0002 |
IEEE Trans. Parallel Distributed Syst. | 4 |
| 2019 | Detecting Vulnerable Android Inter-App Communication in Dynamically Loaded CodeabstractJava reflection and dynamic class loading (DCL) are effective features for enhancing the functionalities of Android apps. However, these features can be abused by sophisticated malware to bypass detection schemes. Advanced malware can utilize reflection and DCL in conjunction with Android Inter-App Communication (IAC) to launch collusion attacks using two or more apps. Such dynamically revealed malicious behaviors enable a new type of stealthy, collusive attacks, bypassing all existing detection mechanisms. In this paper, we present DINA, a novel hybrid analysis approach for identifying malicious IAC behaviors concealed within dynamically loaded code through reflective/DCL calls. DINA continuously appends reflection and DCL invocations to control-flow graphs; it then performs incremental dynamic analysis on such augmented graphs to detect the misuse of reflection and DCL that may lead to malicious, yet concealed, IAC activities. Our extensive evaluation on 3,000 real-world Android apps and 14,000 malicious apps corroborates the prevalent usage of reflection and DCL, and reveals previously unknown and potentially harmful, hidden IAC behaviors in real-world apps. Mohannad Alhanahnah, Qiben Yan 0001, Hamid Bagheri, Hao Zhou 0043, Yutaka Tsutano, Witawas Srisa-an, Xiapu Luo |
INFOCOM | 6 |
| 2019 | Obfusifier: Obfuscation-Resistant Android Malware Detection System
Qiben Yan 0001, Witawas Srisa-an, Yutaka Tsutano |
SecureComm (1) | 4 |
| 2018 | Characterizing and optimizing hotspot parallel garbage collection on multicore systemsabstractThe proliferation of applications, frameworks, and services built on Java have led to an ecosystem critically dependent on the underlying runtime system, the Java virtual machine (JVM). However, many applications running on the JVM, e.g., big data analytics, suffer from long garbage collection (GC) time. The long pause time due to GC not only degrades application throughput and causes long latency, but also hurts overall system efficiency and scalability. Kun Suo, Jia Rao, Hong Jiang 0001, Witawas Srisa-an |
EuroSys | 4 |
| 2018 | Leverage Redundancy in Hardware Transactional Memory to Improve Cache ReliabilityabstractSoft error is a type of transient errors that occur due in part to reductions in capacitance and operating voltages in modern electronic components. Recently, the problem of soft errors has become more prevalent due to several design factors, including aggressive device scaling and newer energy-efficient designs, thus significantly threatening the reliability of computer systems. Since the occurrence of soft errors is non-deterministic, detecting them and recovering from them can be quite challenging. A common way to detect soft errors is to execute two identical program instances and then compare their results. Although this approach is effective, it is not efficient as both non-trivial computation and memory resources must be invested to support such redundant executions. Zhichao Yan 0001, Hong Jiang 0001, Witawas Srisa-an, Sharad C. Seth, Yujuan Tan |
ICPP | 3 |
| 2018 | GranDroid: Graph-Based Detection of Malicious Network Behaviors in Android Applications
Qiben Yan 0001, Witawas Srisa-an, Shakthi Bachala |
SecureComm (1) | 4 |
| 2018 | EvoIsolator: Evolving Program Slices for Hardware Isolation Based SecurityabstractTo provide strong security support for today’s applications, microprocessor manufacturers have introduced hardware isolation, an on-chip mechanism that provides secure accesses to sensitive data. Currently, hardware isolation is still difficult to use by software developers because the process to identify access points to sensitive data is error-prone and can lead to under and over protection of sensitive data. Under protection can lead to security vulnerabilities. Over protection can lead to an increased attack surface and excessive communication overhead. In this paper we describe EvoIsolator , a search-based framework to (i) automatically generate executable minimal slices that include all access points to a set of specified sensitive data; and (ii) automatically optimize (for small code block size and low communication overhead) the code modules for hardware isolation. We demonstrate, through a small feasibility study, the potential impact of our proposed code optimizer. Mengmei Ye, Myra B. Cohen, Witawas Srisa-an, Sheng Wei 0001 |
SSBSE | 3 |
| 2018 | A hybrid approach to testing for nonfunctional faults in embedded systems using genetic algorithmsabstractSummary Embedded systems are challenging to program correctly, because they use an interrupt‐driven programming paradigm and run in resource‐constrained environments. This leads to various classes of nonfunctional faults that can be detected only by customized verification techniques. These nonfunctional faults are specifically related to usage of resources such as time and memory. For example, the presence of interrupts can induce delays in interrupt servicing and in system execution time. Such delays can occur when multiple interrupt service routines and interrupts of different priorities compete for resources on a given CPU. As another example, stack overflows are caused when the combination of active methods and interrupt invocations on the stack grows too large, and these can lead to data loss and other significant device failures. To detect these types of nonfunctional faults, developers need to estimate worst‐case resource usage. Most existing approaches for calculating such estimates are based on static analysis; however, these have a tendency to overapproximate the resources needed. Dynamic techniques such as random testing, in contrast, often underapproximate resource usage. In this article, we presentSimEspresso, a framework that uses a combination of static analysis and a test case generation algorithm to estimate worst‐case resource usage. There are three different worst‐case resource usage scenarios that we consider: (1) worst‐case execution times, (2) worst‐case interrupt latencies, and (3) worst‐case stack usage.SimEspressofirst uses static analysis to identify program paths and interrupt interleavings that potentially lead to worst‐case scenarios. It then uses a genetic algorithm to generate test cases that guide program execution down these paths, using these particular interrupt interleavings. We performed an empirical study to evaluate the effectiveness ofSimEspresso; our results show thatSimEspressois more effective than static analysis approaches and improves significantly over the state of the art dynamic technique, random test case generation. We also find that when we use only the genetic algorithm, omitting the static analysis,SimEspressoperforms almost as effectively, but takes significantly longer to complete its task. Tingting Yu 0001, Witawas Srisa-an, Myra B. Cohen, Gregg Rothermel |
Softw. Test. Verification Reliab. | 2 |
| 2018 | Significant Permission Identification for Machine-Learning-Based Android Malware DetectionabstractThe alarming growth rate of malicious apps has become a serious issue that sets back the prosperous mobile ecosystem. A recent report indicates that a new malicious app for Android is introduced every 10 s. To combat this serious malware campaign, we need a scalable malware detection approach that can effectively and efficiently identify malware apps. Numerous malware detection tools have been developed, including system-level and network-level approaches. However, scaling the detection for a large bundle of apps remains a challenging task. In this paper, we introduce Significant Permission IDentification (SigPID), a malware detection system based on permission usage analysis to cope with the rapid increase in the number of Android malware. Instead of extracting and analyzing all Android permissions, we develop three levels of pruning by mining the permission data to identify the most significant permissions that can be effective in distinguishing between benign and malicious apps. SigPID then utilizes machine-learning-based classification methods to classify different families of malware and benign apps. Our evaluation finds that only 22 permissions are significant. We then compare the performance of our approach, using only 22 permissions, against a baseline approach that analyzes all permissions. The results indicate that when a support vector machine is used as the classifier, we can achieve over 90% of precision, recall, accuracy, and F-measure, which are about the same as those produced by the baseline approach while incurring the analysis times that are 4-32 times less than those of using all permissions. Compared against other state-of-the-art approaches, SigPID is more effective by detecting 93.62% of malware in the dataset and 91.4% unknown/new malware samples. Jin Li 0002, Lichao Sun 0001, Qiben Yan 0001, Witawas Srisa-an, Heng Ye |
IEEE Trans. Ind. Informatics | 5 |
| 2018 | Improving Restore Performance in Deduplication-Based Backup Systems via a Fine-Grained Defragmentation ApproachabstractIn deduplication-based backup systems, the removal of redundant data transforms the otherwise logically adjacent data chunks into physically scattered chunks on the disks. This, in effect, changes the retrieval operations from sequential to random and significantly degrades the performance of restoring data. These scattered chunks are called fragmented data and many techniques have been proposed to identify and sequentially rewrite such fragmented data to new address areas, trading off the increased storage space for reduced number of random reads (disk seeks) to improve the restore performance. However, existing solutions for backup workloads share a common assumption that every read operation involves a large fixed-size window of contiguous chunks, which restricts the fragment identification to a fixed-size read window. This can lead to inaccurate identifications due to false positives since the data fragments can vary in size and appear in any different and unpredictable address locations. Based on these observations, we propose FGdefrag , a Fine-Grained defragmentation approach that uses variable-sized and adaptively located data groups, instead of using fixed-size read windows, to accurately identify and effectively remove fragmented data. When we compare its performance to those of existing solutions, FGdefrag not only reduces the amount of rewritten data but also significantly improves the restore performance. Our experimental results show that FGdefrag can improve the restore performance by 14 to 329 percent, while simultaneously reducing the rewritten data by 25 to 87 percent. Yujuan Tan, Baiping Wang, Zhichao Yan 0001, Hong Jiang 0001, Witawas Srisa-an |
IEEE Trans. Parallel Distributed Syst. | 6 |
| 2017 | Contaminant removal for Android malware detection systemsabstractA recent report indicates that there is a new malicious app introduced every 4 seconds. This rapid malware distribution rate causes existing malware detection systems to fall far behind, allowing malicious apps to escape vetting efforts and be distributed by even legitimate app stores. When trusted downloading sites distribute malware, several negative consequences ensue. First, the popularity of these sites would allow such malicious apps to quickly and widely infect devices. Second, analysts and researchers who rely on machine learning based detection techniques may also download these apps and mistakenly label them as benign since they have not been disclosed as malware. These apps are then used as part of their benign dataset during model training and testing. The presence of contaminants in benign dataset can compromise the effectiveness and accuracy of their detection and classification techniques. To address this issue, we introduce PUDROID (Positive and Unlabeled learning-based malware detection for Android) to automatically and effectively remove contaminants from training datasets, allowing machine learning based malware classifiers and detectors to be more effective and accurate. To further improve the performance of such detectors, we apply a feature selection strategy to select pertinent features from a variety of features. We then compare the detection rates and accuracy of detection systems using two datasets; one using PUDROID to remove contaminants and the other without removing contaminants. The results indicate that once we remove contaminants from the datasets, we can significantly improve both malware detection rate and detection accuracy. Lichao Sun 0001, Xiaokai Wei, Jiawei Zhang 0001, Lifang He 0001, Philip S. Yu, Witawas Srisa-an |
IEEE BigData | 6 |
| 2017 | Energy-efficient I/O Thread Schedulers for NVMe SSDs on NUMAabstractNon-volatile memory express (NVMe) based SSDs and the NUMA platform are widely adopted in servers to achieve faster storage speed and more powerful processing capability. As of now, very little research has been conducted to investigate the performance and energy efficiency of the state-of-the-art NUMA architecture integrated with NVMe SSDs, an emerging technology used to host parallel I/O threads. As this technology continues to be widely developed and adopted, we need to understand the runtime behaviors of such systems in order to design software runtime systems that deliver optimal performance while consuming only the necessary amount of energy. This paper characterizes the runtime behaviors of a Linux-based NUMA system employing multiple NVMe SSDs. Our comprehensive performance and energy-efficiency study using massive numbers of parallel I/O threads shows that the penalty due to CPU contention is much smaller than that due to remote access of NVMe SSDs. Based on this insight, we develop a dynamic "lesser evil" algorithm called ESN, to minimize the impact of these two types of penalties. ESN is an energy-efficient profiling-based I/O thread scheduler for managing I/O threads accessing NVMe SSDs on NUMA systems. Our empirical evaluation shows that ESN can achieve optimal I/O throughput and latency while consuming up to 50% less energy and using fewer CPUs. Junjie Qian, Hong Jiang 0001, Witawas Srisa-an, Sharad C. Seth, Stan Skelton, Joseph Moore |
CCGrid | 3 |
| 2017 | An efficient, robust, and scalable approach for analyzing interacting android appsabstractWhen multiple apps on an Android platform interact, faults and security vulnerabilities can occur. Software engineers need to be able to analyze interacting apps to detect such problems. Current approaches for performing such analyses, however, do not scale to the numbers of apps that may need to be considered, and thus, are impractical for application to real-world scenarios. In this paper, we introduce JITANA, a program analysis framework designed to analyze multiple Android apps simultaneously. By using a classloader-based approach instead of a compiler-based approach such as SOOT, JITANA is able to simultaneously analyze large numbers of interacting apps, perform on-demand analysis of large libraries, and effectively analyze dynamically generated code. Empirical studies of JITANA show that it is substantially more efficient than a state-of-the-art approach, and that it can effectively and efficiently analyze complex apps including Facebook, Pokemon Go, and Pandora that the state-of-the-art approach cannot handle. Yutaka Tsutano, Shakthi Bachala, Witawas Srisa-an, Gregg Rothermel, Jackson Dinh |
ICSE | 3 |
| 2017 | Sequential Keystroke Behavioral Biometrics for Mobile User Identification via Multi-view Deep Learning
Lichao Sun 0001, Bokai Cao, Philip S. Yu, Witawas Srisa-an, Alex D. Leow |
ECML/PKDD (3) | 5 |
| 2016 | RRF: A Race Reproduction Framework for Use in Debugging Process-Level RacesabstractProcess-level races are endemic in modern systems. These races are difficult to debug because they are sensitive to execution events such as interrupts and scheduling. Unless a process interleaving that can result in the race can be found, it cannot be reproduced and cannot be corrected. In practice, however, the number of interleavings that can occur among processes in practice is large, and the patterns of interleavings can be complex. Thus, approaches for reproducing process-level races to date are often ineffective. In this paper, we present RRF, a race reproduction framework that can help software engineers reproduce reported process-level races, enabling them to potentially debug these races. RRF performs a hybrid analysis by leveraging existing static program analysis tools, dynamic kernel event reporting tools, and yield points to provide the observability and controllability needed to reproduce races. We conducted an empirical study to evaluate RRF, our results show that RRF can be effective for reproducing races. Supat Rattanasuksun, Tingting Yu 0001, Witawas Srisa-an, Gregg Rothermel |
ISSRE | 3 |
| 2016 | DroidClassifier: Efficient Adaptive Mining of Application-Layer Header for Classifying Android Malware
Lichao Sun 0001, Qiben Yan 0001, Witawas Srisa-an |
SecureComm | 4 |
| 2016 | Exploiting FIFO Scheduler to Improve Parallel Garbage Collection PerformanceabstractRecent studies have found that parallel garbage collection performs worse with more CPUs and more collector threads. As part of this work, we further investigate this enomenon and find that poor scalability is worst in highly scalable Java applications. Our investigation to find the causes clearly reveals that efficient multi-threading in an application can prolong the average object lifespan, which results in less effective garbage collection. We also find that prolonging lifespan is the direct result of Linux's Completely Fair Scheduler due to its round-robin like behavior that can increase the heap contention between the application threads. Instead, if we use pseudo first-in-first-out to schedule application threads in large multicore systems, the garbage collection scalability is significantly improved while the time spent in garbage collection is reduced by as much as 21%. The average execution time of the 24 Java applications used in our study is also reduced by 11%. Based on this observation, we propose two approaches to optimally select scheduling policies based on application scalability profile. Our first approach uses the profile information from one execution to tune the subsequent executions. Our second approach dynamically collects profile information and performs policy selection during execution. Junjie Qian, Witawas Srisa-an, Sharad C. Seth, Hong Jiang 0001, Du Li, Pan Yi |
VEE | 2 |
| 2015 | Factors affecting scalability of multithreaded Java applications on manycore systemsabstractModern Java applications employ multithreading to improve performance by harnessing execution parallelism available in today’s multicore processors. However, as the numbers of threads and processing cores are scaled up, many applications do not achieve the desired level of performance improvement. In this paper, we explore two factors, lock contention and garbage collection performance that can affect scalability of Java applications. Our initial result reveals two new observations. First, applications that are highly scalable may experience more instances of lock contention than those experienced by applications that are less scalable. Second, efficient multithreading can make garbage collection less effective, and therefore, negatively impacting garbage collection performance. Junjie Qian, Du Li, Witawas Srisa-an, Hong Jiang 0001, Sharad C. Seth |
ISPASS | 3 |
| 2014 | SimRT: an automated framework to support regression testing for data racesabstractConcurrent programs are prone to various classes of difficult-to-detect faults, of which data races are particularly prevalent. Prior work has attempted to increase the cost-effectiveness of approaches for testing for data races by employing race detection techniques, but to date, no work has considered cost-effective approaches for re-testing for races as programs evolve. In this paper we present SimRT, an automated regression testing framework for use in detecting races introduced by code modifications. SimRT employs a regression test selection technique, focused on sets of program elements related to race detection, to reduce the number of test cases that must be run on a changed program to detect races that occur due to code modifications, and it employs a test case prioritization technique to improve the rate at which such races are detected. Our empirical study of SimRT reveals that it is more efficient and effective for revealing races than other approaches, and that its constituent test selection and prioritization components each contribute to its performance. Tingting Yu 0001, Witawas Srisa-an, Gregg Rothermel |
ICSE | 2 |
| 2014 | SimLatte: A Framework to Support Testing for Worst-Case Interrupt Latencies in Embedded SoftwareabstractEmbedded systems tend to be interrupt-driven, yet the presence of interrupts can affect system dependability because there can be delays in servicing interrupts. Such delays can occur when multiple interrupt service routines and interrupts of different priorities compete for resources on a given CPU. For this reason, researchers have sought approaches by which to estimate worst-case interrupt latencies (WCILs) for systems. Most existing approaches, however, are based on static analysis. In this paper, we present SIMLATTE, a testing-based approach for finding WCILs. SIMLATTE uses a genetic algorithm for test case generation that converges on a set of inputs and interrupt arrival points that are likely to expose WCILs. It also uses an opportunistic interrupt invocation approach to invoke interrupts at a variety of feasible locations. Our evaluation of SIMLATTE on several non-trivial embedded systems reveals that it is considerably more effective and efficient than random testing. We also determine that the combination of the genetic algorithm and opportunistic interrupt invocation allows SIMLATTE to perform better than it can when using either one in isolation. Tingting Yu 0001, Witawas Srisa-an, Myra B. Cohen, Gregg Rothermel |
ICST | 2 |
| 2014 | An approach to testing commercial embedded systems
Tingting Yu 0001, Ahyoung Sung, Witawas Srisa-an, Gregg Rothermel |
J. Syst. Softw. | 3 |
| 2013 | An empirical comparison of the fault-detection capabilities of internal oraclesabstractModern computer systems are prone to various classes of runtime faults due to their reliance on features such as concurrency and peripheral devices such as sensors. Testing remains a common method for uncovering faults in these systems, but many runtime faults are difficult to detect using typical testing oracles that monitor only program output. In this work we empirically investigate the use of internal test oracles: oracles that detect faults by monitoring aspects of internal program and system states. We compare these internal oracles to each other and to output-based oracles for relative effectiveness and examine tradeoffs between oracles involving incorrect reports about faults (false positives and false negatives). Our results reveal several implications that test engineers and researchers should consider when testing for runtime faults. Tingting Yu 0001, Witawas Srisa-an, Gregg Rothermel |
ISSRE | 2 |
| 2013 | SimRacer: an automated framework to support testing for process-level racesabstractFaults introduced by races are difficult to detect because they usually occur only under specific execution interleavings. Numerous program analysis and testing techniques have been proposed to detect races between threads. Little work, however, has addressed the problem of detecting and testing for process-level races, in which two processes access a shared resource without proper synchronization. In this paper, we present SIMRACER, a novel testing-based framework that allows engineers to effectively test for process-level races. SIMRACER first computes potential races based on runtime traces obtained by running existing tests on target processes, and then it controls process scheduling relative to the potential races so that real races can be created. We implemented SIMRACER on a commercial virtual platform that is widely used to support hardware/software co-design. We then evaluated its effectiveness on sixteen real-world applications containing known process-level races. Our results show that SIMRACER is effective at detecting process-level races, and more effective than traditional stress testing techniques at detecting faults caused by those races. Tingting Yu 0001, Witawas Srisa-an, Gregg Rothermel |
ISSTA | 2 |
| 2012 | SimTester: a controllable and observable testing framework for embedded systemsabstractIn software for embedded systems, the frequent use of interrupts for timing, sensing, and I/O processing can cause concurrency faults to occur due to interactions between applications, device drivers, and interrupt handlers. This type of fault is considered by many practitioners to be among the most difficult to detect, isolate, and correct, in part because it can be sensitive to execution interleavings and often occurs without leaving any observable incorrect output. As such, commonly used testing techniques that inspect program outputs to detect failures are often ineffective at detecting them. To test for these concurrency faults, test engineers need to be able to control interleavings so that they are deterministic. Furthermore, they also need to be able to observe faults as they occur instead of relying on observable incorrect outputs. Tingting Yu 0001, Witawas Srisa-an, Gregg Rothermel |
VEE | 2 |
| 2011 | Using Property-Based Oracles when Testing Embedded System ApplicationsabstractEmbedded systems are becoming increasingly ubiquitous, controlling a wide variety of popular and safety-critical devices. Effective testing techniques could improve the dependability of these systems. In prior work we presented an approach for testing embedded systems, focusing on embedded system applications and the tasks that comprise them. In this work we focus on a second but equally important aspect of testing embedded systems, namely, the need to provide observability of system behavior sufficient to allow engineers to detect failures. We present several property-based oracles that can be instantiated in embedded systems through program analysis and instrumentation, and can detect failures for which simple output-based oracles are inadequate. An empirical study of our approach shows that it can be effective. Tingting Yu 0001, Ahyoung Sung, Witawas Srisa-an, Gregg Rothermel |
ICST | 3 |
| 2011 | SOS: saving time in dynamic race detection with stationary analysisabstractData races are subtle and difficult to detect errors that arise during concurrent program execution. Traditional testing techniques fail to find these errors, but recent research has shown that targeted dynamic analysis techniques can be developed to precisely detect races (i.e., no false race reports are generated) that occur during program execution. Unfortunately, precise race detection is still too expensive to be used in practice. State-of-the-art techniques still slow down program execution by a factor of eight or more. In this paper, we incorporate an optimization technique based on the observation that many thread-shared objects are written early in their lifetimes and then become read-only for the remainder of their lifetimes; these are known as stationary objects. The main contribution of our work is the insight that once a stationary object becomes thread-shared, races cannot occur. Therefore, our proposed approach does not monitor access to these objects. As such, our system only incurs an average overhead of 45% of that of an implementation of FastTrack, a low-overhead dynamic race detector. We then compared the effectiveness of our approach to de- tect races in deployed environments with that of Pacer, a sampling based race detector based on FastTrack. We found that our approach can detect over five times more races than Pacer when we budget 50% for runtime overhead. Du Li, Witawas Srisa-an, Matthew B. Dwyer |
OOPSLA | 2 |
| 2010 | Testing Inter-layer and Inter-task Interactions in RTES ApplicationsabstractReal-time embedded systems (RTESs) are becoming increasingly ubiquitous, controlling a wide variety of popular and safety-critical devices. Effective testing techniques could improve the dependability of these systems. In this paper we present an approach for testing RTESs, intended specifically to help RTES application developers detect faults related to functional correctness. Our approach consists of two techniques that focus on exercising the interactions between system layers and between the multiple user tasks that enact application behaviors. We present results of an empirical study that shows that our techniques are effective at detecting faults. Ahyoung Sung, Witawas Srisa-an, Gregg Rothermel, Tingting Yu 0001 |
APSEC | 2 |
| 2010 | A self-adjusting code cache manager to balance start-up time and memory usageabstractIn virtual machines for embedded devices that use just-in-time compilation, the management of the code cache can significantly impact performance in terms of both memory usage and start-up time. Although improving memory usage has been a common focus for system designers, start-up time is often overlooked. In systems with constrained resources, however, these two performance metrics are often at odds and must be considered together. In this paper, we present an adaptive self-adjusting code cache manager to improve performance with respect to both start-up time and memory usage. It balances these concerns by detecting changes in method compilation rates, resizing the cache after each pitching event. We conduct experiments to validate our proposed system and quantify the impacts that different code cache management techniques have on memory usage and start-up time through two oracle systems. Our results show that the proposed algorithm yields nearly the same start-up times as a hand-tuned oracle and shorter execution times than those of the SSCLI in eight out of ten applications. It also has lower memory usage over time in all but one application. Witawas Srisa-an, Myra B. Cohen, Mithuna Soundararaj |
CGO | 1 |
| 2009 | Investigating the effects of using different nursery sizing policies on performanceabstractIn this paper, we investigate the effects of using three different nursery sizing policies on overall and garbage collection performances. As part of our investigation, we modify the parallel generational collector in HotSpot to support a fixed ratio policy and heap availability policy (similar to that used in Appel collectors), in addition to its GC Ergonomics policy. We then compare the performances of 16 large and small multithreaded Java benchmarks; each is given a reasonably sized heap and utilizes all three policies. The result of our investigation indicates that many benchmarks are sensitive to heap sizing policies, resulting in overall performance differences that can range from 1 percent to 36 percents. We also find that in our server application benchmarks, more than one policy may be needed. As a preliminary study, we introduce a hybrid policy that uses one policy when the heap space is plentiful to yield optimal performance and then switches to a different policy to improve survivability and yield more graceful performance degradation under heavy memory pressure. Xiaohua Guan, Witawas Srisa-an, ChengHuan Jia |
ISMM | 2 |
| 2008 | Contention-aware scheduler: unlocking execution parallelism in multithreaded java programsabstractIn multithreaded programming, locks are frequently used as mechanism for synchronization. Because today's operating systems do not consider lock usage as scheduling criterion, scheduling decisions can be unfavorable to multithreaded applications, leading to performance issues such as convoying and heavy lock contention in systems with multiple processors. Previous efforts to address these issues (e.g., transactional memory, lock-free data structure) often treat scheduling decisions as a fact of life, and therefore these solutions try to cope with the consequences of undesirable scheduling instead of dealing with the problem directly. In this paper, we introduce Contention-Aware Scheduler (CA-Scheduler), which is designed to support efficient execution of large multithreaded Java applications in multiprocessor systems. Our proposed scheduler employs scheduling policy that reduces lock contention. As will be shown in this paper, our prototype implementation of the CA-Scheduler in Linux and Sun HotSpot virtual machine only incurs 3.5% runtime overhead, while the overall performance differences, when compared with system with no contention awareness, range from degradation of 3% in small multithreaded benchmark to an improvement of 15% in large Java application server benchmark. Feng Xian, Witawas Srisa-an, Hong Jiang 0001 |
OOPSLA | 2 |
| 2008 | Garbage collection: Java application servers' Achilles heel
Feng Xian, Witawas Srisa-an, Hong Jiang 0001 |
Sci. Comput. Program. | 2 |
| 2007 | AS-GC: An Efficient Generational Garbage Collector for Java Application Servers
Feng Xian, Witawas Srisa-an, ChengHuan Jia, Hong Jiang 0001 |
ECOOP | 2 |
| 2007 | Allocation-phase aware thread scheduling policies to improve garbage collection performanceabstractPast studies have shown that objects are created and then die in phases. Thus, one way to sustain good garbage collection efficiency is to have a large enough heap to allow many allocation phases to complete and most of the objects to die before invoking garbage collection. However, such an operating environment is hard to maintain in large multithreaded applications because most typical time-sharing schedulers are not allocation-phase cognizant; i.e., they often schedule threads in a way that prevents them from completing their allocation phases quickly. Thus, when garbage collection is invoked, most allocation phases have yet to be completed, resulting in poor collection efficiency. We introduce two new scheduling strategies, LARF (lower allocation rate first) and MQRR (memory-quantum round robin) designed to be allocation-phase aware by assigning higher execution priority to threads in computation-oriented phases. The simulation results show thatallthe reductions of the garbage collection time in a generational collector can range from 0%-27% when compare to a round robin scheduler. The reductions of the overall execution time and the average thread turnaround time range from -0.1%-3% and -0.1%-13%, respectively. Feng Xian, Witawas Srisa-an, Hong Jiang 0001 |
ISMM | 2 |
| 2007 | Microphase: an approach to proactively invoking garbage collection for improved performanceabstractTo date, the most commonly used criterion for invoking garbage collection (GC) is based on heap usage; that is, garbage collection is invoked when the heap or an area inside the heap is full. This approach can suffer from two performance shortcomings: untimely garbage collection invocations and large volumes of surviving objects. In this work, we explore a new GC triggering approach called MicroPhase that exploits two observations: (i) allocation requests occur in phases and (ii) phase boundaries coincide with times when most objects also die. Thus, proactively invoking garbage collection at these phase boundaries can yield high efficiency. We extended the HotSpot virtual machine from Sun Microsystems to support MicroPhase and conducted experiments using 20 benchmarks. The experimental results indicate that our technique can reduce the GC times in 19 applications. The differences in GC overhead range from an increase of 1% to a decrease of 26% when the heap is set to twice the maximum live-size. As a result, MicroPhase can improve the overall performance of 13 benchmarks. The performance differences range from a degradation of 2.5% to an improvement of 14%. Feng Xian, Witawas Srisa-an, Hong Jiang 0001 |
OOPSLA | 2 |
| 2006 | Evaluating Hardware Support for Reference Counting Using Software Configurable ProcessorsabstractReference counting is an incremental garbage collection technique that yields nearly unnoticeable pause time but can suffer from high processing overhead. Previous attempts to use hardware to reduce this overhead have shown successes but limited applicability. With recent discovery that reference counting can be well suited for Java embedded devices, it is worthwhile to rethink hardware solutions that can further improve its performance by leveraging current trends in embedded computing. In this paper, we introduce a custom instruction solution that is (i) more practical because it only provides hardware support for the expensive but straight-forward software function-reference count update; the existing complex runtime functions such as memory allocation and liberation remain unchanged and (ii) better positioned for widespread adoption because it is designed to leverage readily available configurable logics in many embedded processor cores. As a proof-of-concept, we implement two reference counting algorithms that utilize the proposed custom instructions on Stretch S5000 software reconfigurable processors. We then analyze the performance impacts on the execution time as well as the architectural behavior. The results show that we can achieve as much as 70% performance gain over pure software implementation. Feng Xian, Witawas Srisa-an, Hong Jiang 0001 |
ASAP | 2 |
| 2006 | Clustering the heap in multi-threaded applications for improved garbage collectionabstractGarbage collection can be a performance bottleneck in large distributed, multi-threaded applications. Applications may produce millions of objects during their lifetimes and may invoke hundreds or thousands of threads. When using a single shared heap, each time a garbage collection phase occurs all threads must be stopped, essentially halting all other processing. Attempts to fix this bottleneck include creating a single heap per thread, however this may not scale to large thread intensive applications. In this paper we explore the potential of clustering threads into related sub-heaps. We hypothesize that this will lead to a smaller shared heap, while maintaining good garbage collection parallelism. We leverage results from software module clustering to achieve this goal. Our results show that we can significantly reduce the number of sub-heaps created and reduce the number of objects in the shared heap in a representative application. This suggests that clustering may be a promising optimization technique for garbage collection in large multi-threaded systems with many shared objects. Myra B. Cohen, Shiu Beng Kooi, Witawas Srisa-an |
GECCO | 3 |
| 2005 | An energy efficient garbage collector for java embedded devicesabstractThis paper presents a detailed design and implementation of a power-efficient garbage collector for Java embedded systems. The proposed scheme is a hybrid between the standard mark-sweep-compact collector available in Sun's KVM and a limited-field reference counter. There are three benefits resulting from the proposed scheme. (a) the proposed scheme reclaims memory more efficiently and this results in less mark-sweep garbage collection invocations, (b) reduction in garbage collection invocations improves cache locality and reduces the number of main memory accesses, and (c) reduction in memory access ultimately results in lower energy consumption, since a memory access can consume a large amount of energy when compared with an instruction execution. The proposed scheme has been implemented into Sun's KVM, and has been shown to reduce the number of mark-sweep garbage collection invocations by up to 100% in some cases, and the number of level-1 cache misses by as much as 87% when compared to the default garbage collector. We also find that in some applications, the proposed scheme can reduce the power consumption by as much as 27% when compared to the default Sun's KVM. Paul A. Griffin, Witawas Srisa-an, J. Morris Chang |
LCTES | 2 |
| 2004 | Object allocation and memory contention study of Java multithreaded applicationsabstractJava has become a popular programming language used on different platforms, ranging from embedded systems to powerful servers. Since the memory management is one of the most time-consuming parts within Java virtual machine (JVM), various techniques have been developed to boost its performance. However, the JVM memory management still does not scale very well, especially for multithreaded server applications. In this paper, we study different aspects of JVM object allocation from thread's perspective, using the trace data we collected from Sun JDK 1.3.1. Additionally, we construct a heap simulator to study the potential memory contentions among different threads. The simulation results show that dividing heap into different subheaps is very effective in alleviating the memory contentions. The results imply the potential benefits of using subheaps in improving the Java memory management performance. Wei Huang 0032, Witawas Srisa-an, J. Morris Chang |
IPCCC | 3 |
| 2004 | Dynamic pretenuring schemes for generational garbage collectionabstractPrevious research efforts have shown that pretenuring can potentially reduce the copying cost by creating long lived objects into the mature memory regions directly. To date, researchers often employ profiling and static analysis to accurately select the objects that should be pretenured. However, little research efforts have been spent on dynamic approaches for pretenuring objects. In this paper, we propose a novel approach that dynamically predicts object lifespan to assist with pretenuring selection. The proposed scheme performs dynamic pretenuring selection based on a feedback mechanism that records lifespan of objects from each class during garbage collection invocations. This information is then used to pretenure objects in subsequent allocation requests. We experiment with two approaches, jumpstart feedback and continuous feedback, to collect tenuring information. The experimental results of selected benchmark programs show that our schemes can improve the garbage collection time of IBM's Jikes RVM by up to 37%, and improve the overall execution time by up to 28%. Wei Huang 0032, Witawas Srisa-an, J. Morris Chang |
ISPASS | 2 |
| 2004 | The design and analysis of a quantitative simulator for dynamic memory management
Dan Chia-Tien Lo, Witawas Srisa-an, J. Morris Chang |
J. Syst. Softw. | 2 |
| 2003 | Active Memory Processor: A Hardware Garbage Collector for Real-Time Java Embedded DevicesabstractJava possesses many advantages for embedded system development, including fast product deployment, portability, security, and a small memory footprint. As Java makes inroads into the market for embedded systems, much effort is being invested in designing real-time garbage collectors. The proposed garbage-collected memory module, a bitmap-based processor with standard DRAM cells is introduced to improve the performance and predictability of dynamic memory management functions that include allocation, reference counting, and garbage collection. As a result, memory allocation can be done in constant time and sweeping can be performed in parallel by multiple modules. Thus, constant time sweeping is also achieved regardless of heap size. This is a major departure from the software counterparts where sweeping time depends largely on the size of the heap. In addition, the proposed design also supports limited-field reference counting, which has the advantage of distributing the processing cost throughout the execution. However, this cost can be quite large and results in higher power consumption due to frequent memory accesses and the complexity of the main processor. By doing reference counting operation in a coprocessor, the processing is done outside of the main processor. Moreover, the hardware cost of the proposed design is very modest (about 8000 gates). Our study has shown that 3-bit reference counting can eliminate the need to invoke the garbage collector in all tested applications. Moreover, it also reduces the amount of memory usage by 77 percent. Witawas Srisa-an, Dan Chia-Tien Lo, J. Morris Chang |
IEEE Trans. Mob. Comput. | 1 |
| 2002 | Performance Enhancements to the Active Memory SystemabstractThe Active Memory System - a garbage collected memory module - was introduced as a way to provide hardware support for garbage collection in embedded systems. The major component in the design was the Active Memory Processor (AMP) that utilized a set of bit-maps and a combinational circuit to perform mark-sweep garbage collection. The design can achieve constant time for both allocation and sweeping. In this paper two enhancements are made to the design of AMP so that it can perform one-bit reference counting that postpones the need to perform garbage collection. Moreover, a caching mechanism is also introduced to reduce the hardware cost of the design. The experimental results show that the proposed modification can reduce the number of garbage collection invocations by 76%. The speed-up in marking time can be as much as 5.81. With the caching mechanism, the hardware cost can be as small as 27 K gates and 6 KB of SRAM. Witawas Srisa-an, Dan Chia-Tien Lo, J. Morris Chang |
ICCD | 1 |
| 2002 | DMMX: Dynamic memory management extensions
J. Morris Chang, Witawas Srisa-an, Dan Chia-Tien Lo, Edward F. Gehringer |
J. Syst. Softw. | 2 |
| 2001 | A Performance Analysis of the Active Memory SystemabstractOne major problem of using Java in real-time and embedded devices is the non-deterministic turnaround time of dynamic memory management systems (memory allocation and garbage collection). For the allocation, the nondeterminism is often contributed by the time to perform searching, splitting, and coalescing. For the garbage collection, the turnaround time is usually determined by the size of the heap, the number of live objects, the number of object collected, and the amount of garbage collected Even with the current state-of-the-art garbage collectors (generational and incremental schemes), they may or may not guarantee the worst case latency. Moreover such schemes often prolong overall garbage collection time. In this paper, the performance analysis of the proposed Active Memory Module (AMM) for embedded systems is presented Unlike the software counterparts, the AMM can perform a memory allocation in a predictable and hounded fashion (14 cycles). Moreover it can also yield a bounded sweeping time regardless of the number of live objects or heap size. By utilizing the proposed system, the overall speed-up can be as high as 23% over the JDK 1.2.2 running in classic mode. Witawas Srisa-an, Dan Chia-Tien Lo, J. Morris Chang |
ICCD | 1 |
| 2001 | Cycle accurate thread timer for linux environmentabstractDue to the increasing popularity of java in clienthemer environments, most of today's server applications are multithreaded. Thus, research focusing on the performance analysis of multi-threaded environments has become increasingly important. Since per-thread information can be crucial in such analysis, measuring tools are needed to provide perthread information that may include cycle-based timers and filters to eliminate tracing overhead. In this papel; a Cycle Accurate Thread Timing for Linun Environment (CAlTLE) is presented. This approach provides a cycle-accurate timer with functions to filter out tracing overhead by coordinating efforts from both kernel and user applications. In this scheme, the kernel keeps track of accurate thread timing, while applications inform the kernel which part of the execution is to be measured. To demonstrate the tool$functionality, two case studies are provided, which include measuring latencies incurred by malloc calls and monitoring potential memory heap contention in multithreaded-multiprocessor environments. Witawas Srisa-an, Therapon Skotiniotis, J. Morris Chang |
ISPASS | 2 |
| 2001 | A study of the allocation behavior of C++ programs
J. Morris Chang, Woo Hyong Lee, Witawas Srisa-an |
J. Syst. Softw. | 3 |
| 2001 | A study of page replacement performance in garbage collection heap
Dan Chia-Tien Lo, Witawas Srisa-an, J. Morris Chang |
J. Syst. Softw. | 2 |
| 2000 | Architectural Support for Dynamic Memory ManagementabstractRecent advances in software engineering, such as graphical user interfaces and object-oriented programming, have caused applications to become more memory intensive. These applications tend to allocate dynamic memory prolifically. Moreover, automatic dynamic memory reclamation (garbage collection, GC) has become a popular feature in modern programming languages. As a result, the time consumed by dynamic storage management can be up to one-third of the program execution time. This illustrates the need for a high-performance memory management scheme. This paper presents a top-level design and evaluation of the proposed instruction extensions to facilitate heap management. J. Morris Chang, Witawas Srisa-an, Dan Chia-Tien Lo |
ICCD | 2 |
| 2000 | A quantitative simulator for dynamic memory managersabstractIn the last thirty years, several dynamic memory management schemes have been proposed. Such schemes include first fit, best fit, segregated fit, and buddy systems. Because the performance (speed and memory utilization) of each scheme differs, software engineers often face difficult choices in selecting the most suitable approach for their applications. In this paper, a quantitative simulator for dynamic memory management and memory tracing techniques are presented. This simulator receives dynamic memory management traces and performs allocations according to schemes (first fit, best fit, buddy systems, and segregated fit) defined by the user. At the end of each simulation run, different performance metrics are reported to the users. By using this approach, software engineers can evaluate system performance and decide which algorithm is the most suitable for their applications. Dan Chia-Tien Lo, Witawas Srisa-an, J. Morris Chang |
ISPASS | 2 |
| 2000 | Do generational schemes improve the garbage collection efficiency?abstractRecently, most research efforts on garbage collection have concentrated on reducing pause times. However, very little effort has been spent on the study of garbage collection efficiency, especially generational garbage collection which was introduced as a way to reduce garbage collection pause times. In this paper a detailed study of garbage collection efficiency in generational schemes is presented. The study provides a mathematical model for the efficiency of generation garbage collection. Additionally, important issues such as write-barrier overhead, pause times, residency, and heap size are also addressed. We find that generational garbage collection often has lower garbage collection efficiency than other approaches (e.g. mark-sweep, copying) due to a smaller collected area and write-barrier overhead. Witawas Srisa-an, J. Morris Chang, Dan Chia-Tien Lo |
ISPASS | 1 |
| 2000 | A hardware implementation of realloc function
Witawas Srisa-an, Dan Chia-Tien Lo, J. Morris Chang |
Integr. | 1 |