VLDB 2026 Research / reviewers in the wild / expert
Yung-Hsiang Lu
dblp:12/5138
· DBLP profile ↗
100ranked-venue papers
8as first author
15since 2021 · last 2025
0000-0002-5491-7661ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 65 · 6 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 17 · 1 first-author · 3 since 2021Artificial intelligence and machine learning · 14 · 1 since 2021Computer networks · 11 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 2 first-author · 3 since 2021Software engineering, systems software and programming languages · 10 · 1 first-author · 2 since 2021Security and privacy · 2 · 1 since 2021Databases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Detecting Music Performance Errors with TransformersabstractBeginner musicians often struggle to identify specific errors in their performances, such as playing incorrect notes or rhythms. There are two limitations in existing tools for music error detection: (1) Existing approaches rely on automatic alignment; therefore, they are prone to errors caused by small deviations between alignment targets; (2) There is insufficient data to train music error detection models, resulting in over-reliance on heuristics. To address (1), we propose a novel transformer model, Polytune, that takes audio inputs and outputs annotated music scores. This model can be trained end-to-end to implicitly align and compare performance audio with music scores through latent space representations. To address (2), we present a novel data generation technique capable of creating large-scale synthetic music error datasets. Our approach achieves a 64.1% average Error Detection F1 score, improving upon prior work by 40 percentage points across 14 instruments. Additionally, our model can handle multiple instruments compared with existing transcription methods repurposed for music error detection. Benjamin Shiue-Hal Chou, Purvish Jajal, Nicholas Eliopoulos, Tim Nadolsky, Cheng-Yun Yang, Nikita Ravi, James C. Davis 0001, Kristen Yeon-Ji Yun, Yung-Hsiang Lu |
AAAI | 9 |
| 2025 | FlexGE: Towards Secure and Flexible Model Partition for Deep Neural Networks
Aravind Machiry, Yung-Hsiang Lu, Jing (Dave) Tian |
DIMVA (2) | 3 |
| 2025 | Pruning One More Token is Enough: Leveraging Latency-Workload Non-Linearities for Vision Transformers on the EdgeabstractThis paper investigates how to efficiently deploy vision transformers on edge devices for small workloads. Recent methods reduce the latency of transformer neural networks by removing or merging tokens with small accuracy degradation. However these methods are not designed with edge device deployment in mind: they do not leverage information about the latency-workload trends to improve efficiency. We address this shortcoming in our work. First we identify factors that affect ViT latency-workload relationships. Second we determine token pruning schedule by leveraging non-linear latency-workload relationships. Third we demonstrate a training-free token pruning method utilizing this schedule. We show other methods may increase latency by 2-30% while we reduce latency by 9-26%. For similar latency (within 5.2% or 7ms) across devices we achieve 78.6%-84.5% ImageNet1K classification accuracy while the state-of-the-art Token Merging achieves 45.8%-85.4%. Nicholas Eliopoulos, Purvish Jajal, James C. Davis 0001, Gaowen Liu, George K. Thiravathukal, Yung-Hsiang Lu |
WACV | 6 |
| 2025 | Token Turing Machines are Efficient Vision ModelsabstractWe propose Vision Token Turing Machines (ViTTM), an efficient, low-latency, memory-augmented Vision Transformer (ViT). Our approach builds on Neural Turing Machines (NTM) and Token Turing Machines (TTM), which were applied to NLP and sequential visual understanding tasks. ViTTMs are designed for non-sequential computer vision tasks such as image classification and segmentation. Our model creates two sets of tokens: process tokens and memory tokens; process tokens pass through encoder blocks and read-write from memory tokens at each encoder block in the network, allowing them to store and retrieve information from memory. By ensuring that there are fewer process tokens, we are able to reduce the inference time of the network while maintaining its accuracy. On ImageNet-1K, the state-of-the-art ViT-B has median latency of 529.5 ms and 81.0% accuracy, while our ViTTM-B is 56% faster (234.1 ms), with 2.4× fewer FLOPs, with an accuracy of 82.9%. On ADE20K semantic segmentation, ViT-B achieves 45.65 mIoU at 13.8 frames per second (FPS) whereas our ViTTM-B model achieves 45.17 mIoU with 26.8 FPS (+94%). Purvish Jajal, Nicholas Eliopoulos, Benjamin Shiue-Hal Chou, George K. Thiravathukal, James C. Davis 0001, Yung-Hsiang Lu |
WACV | 6 |
| 2024 | An automated approach for improving the inference latency and energy efficiency of pretrained CNNs by removing irrelevant pixels with focused convolutionsabstractComputer vision often uses highly accurate Convolutional Neural Networks (CNNs), but these deep learning models are associated with ever-increasing energy and computation requirements. Producing more energy-efficient CNNs often requires model training which can be cost-prohibitive. We propose a novel, automated method to make a pretrained CNN more energy-efficient without re-training. Given a pretrained CNN, we insert a threshold layer that filters activations from the preceding layers to identify regions of the image that are irrelevant, i.e. can be ignored by the following layers while maintaining accuracy. Our modified focused convolution operation saves inference latency (by up to 25%) and energy costs (by up to 22%) on various popular pretrained CNNs, with little to no loss in accuracy. Caleb Tung, Nicholas Eliopoulos, Purvish Jajal, Gowri Ramshankar, Cheng-Yun Yang, Nicholas Synovic, Xuecen Zhang, Vipin Chaudhary, George K. Thiruvathukal, Yung-Hsiang Lu |
ASPDAC | 10 |
| 2024 | Securing Deep Neural Networks on Edge from Membership Inference Attacks Using Trusted Execution EnvironmentsabstractPrivacy concerns arise from malicious attacks on Deep Neural Network (DNN) applications during sensitive data inference on edge devices. Membership Inference Attack (MIA) is developed by adversaries to determine whether sensitive data is used to train the DNN applications. Prior work uses Trusted Execution Environments (TEEs) to hide DNN model inference from adversaries on edge devices. Unfortunately, existing methods have two major problems. First, due to the restricted memory of TEEs, prior work cannot secure large-size DNNs from gradient-based MIAs. Second, prior work is ineffective on output-based MIAs. To mitigate the problems, we present a depth-wise layer partitioning method to run large sensitive layers inside TEEs. We further propose a model quantization strategy to improve the defense capability of DNNs against output-based MIAs and accelerate the computation. We also automate the process of securing PyTorch-based DNN models inside TEEs. Experiments on Raspberry Pi 3B+ show that our method can reduce the accuracy of gradient-based MIAs on AlexNet, VGG-16, and ResNet-20 evaluated on the CIFAR-100 dataset by 28.8%, 11%, and 35.3%. The accuracy of output-based MIAs on the three models is also reduced by 18.5%, 13.4%, and 29.6%, respectively. Cheng-Yun Yang, Gowri Ramshankar, Nicholas Eliopoulos, Purvish Jajal, Sudarshan Nambiar, Evan Miller, Jing (Dave) Tian, Shuo-Han Chen, Chiy-Ferng Perng, Yung-Hsiang Lu |
ISLPED | 11 |
| 2024 | Interoperability in Deep Learning: A User Survey and Failure Analysis of ONNX Model ConvertersabstractSoftware engineers develop, fine-tune, and deploy deep learning (DL) models using a variety of development frameworks and runtime environments. DL model converters move models between frameworks and to runtime environments. Conversion errors compromise model quality and disrupt deployment. However, the failure characteristics of DL model converters are unknown, adding risk when using DL interoperability technologies. This paper analyzes failures in DL model converters. We survey software engineers about DL interoperability tools, use cases, and pain points (N=92). Then, we characterize failures in model converters associated with the main interoperability tool, ONNX (N=200 issues in PyTorch and TensorFlow). Finally, we formulate and test two hypotheses about structural causes for the failures we studied. We find that the node conversion stage of a model converter accounts for ∼75% of the defects and 33% of reported failure are related to semantically incorrect models. The cause of semantically incorrect models is elusive, but models with behaviour inconsistencies share operator sequences. Our results motivate future research on making DL interoperability software simpler to maintain, extend, and validate. Research into behavioural tolerances and architectural coverage metrics would be fruitful. Purvish Jajal, Wenxin Jiang 0001, Arav Tewari, Erik Kocinare, Joseph Woo, Anusha Sarraf, Yung-Hsiang Lu, George K. Thiruvathukal, James C. Davis 0001 |
ISSTA | 7 |
| 2023 | Lightning Talk 6: Bringing Together Foundation Models and Edge DevicesabstractDeep learning models have been widely used in natural language processing and computer vision. These models require heavy computation, large memory, and massive amounts of training data. Deep learning models may be deployed on edge devices when transferring data to cloud is infeasible or undesirable. Running these models on edge devices require significant improvement in the efficiency by reducing the models’ resource demands. Existing methods to improve efficiency often require new architectures and retraining. The recent trend in machine learning is to create general-purpose models (called foundation models). These pre-trained models can be repurposed for different applications. This paper reviews the methods for improving efficiency of machine learning models, the rise of foundation models, challenges and possible solutions improving efficiency of pre-trained models. Future solutions for better efficiency should focus on improving existing trained models with no or limited training. Nicholas Eliopoulos, Yung-Hsiang Lu |
DAC | 2 |
| 2023 | An Empirical Study of Pre-Trained Model Reuse in the Hugging Face Deep Learning Model RegistryabstractDeep Neural Networks (DNNs) are being adopted as components in software systems. Creating and specializing DNNs from scratch has grown increasingly difficult as state-of-the-art architectures grow more complex. Following the path of traditional software engineering, machine learning engineers have begun to reuse large-scale pre-trained models (PTMs) and fine-tune these models for downstream tasks. Prior works have studied reuse practices for traditional software packages to guide software engineers towards better package maintenance and dependency management. We lack a similar foundation of knowledge to guide behaviors in pre-trained model ecosystems. In this work, we present the first empirical investigation of PTM reuse. We interviewed 12 practitioners from the most popular PTM ecosystem, Hugging Face, to learn the practices and challenges of PTM reuse. From this data, we model the decision-making process for PTM reuse. Based on the identified practices, we describe useful attributes for model reuse, including provenance, reproducibility, and portability. Three challenges for PTM reuse are missing attributes, discrepancies between claimed and actual performance, and model risks. We substantiate these identified challenges with systematic measurements in the Hugging Face ecosystem. Our work informs future directions on optimizing deep learning ecosystems by automated measuring useful attributes and potential attacks, and envision future research on infrastructure and standardization for model registries. Wenxin Jiang 0001, Nicholas Synovic, Matt Hyatt, Taylor R. Schorlemmer, Rohan Sethi, Yung-Hsiang Lu, George K. Thiruvathukal, James C. Davis 0001 |
ICSE | 6 |
| 2022 | Efficient Computer Vision on Edge Devices with Pipeline-Parallel Hierarchical Neural NetworksabstractComputer vision on low-power edge devices enables applications including search-and-rescue and security. State-of-the-art computer vision algorithms, such as Deep Neural Networks (DNNs), are too large for inference on low-power edge devices. To improve efficiency, some existing approaches parallelize DNN inference across multiple edge devices. How-ever, these techniques introduce significant communication and synchronization overheads or are unable to balance workloads across devices. This paper demonstrates that the hierarchical DNN architecture is well suited for parallel processing on multiple edge devices. We design a novel method that creates a parallel inference pipeline for computer vision problems that use hierarchical DNNs. The method balances loads across the collaborating devices and reduces communication costs to facilitate the processing of multiple video frames simultaneously with higher throughput. Our experiments consider a representative computer vision problem where image recognition is performed on each video frame, running on multiple Raspberry Pi 4Bs. With four collaborating low-power edge devices, our approach achieves 3.21× higher throughput, 68% less energy consumption per device per frame, and a 58% decrease in memory when compared with existing sinaledevice hierarchical DNNs. Abhinav Goel, Caleb Tung, Xiao Hu 0004, George K. Thiruvathukal, James C. Davis 0001, Yung-Hsiang Lu |
ASP-DAC | 6 |
| 2022 | Directed Acyclic Graph-based Neural Networks for Tunable Low-Power Computer VisionabstractProcessing visual data on mobile devices has many applications, e.g., emergency response and tracking. State-of-the-art computer vision techniques rely on large Deep Neural Networks (DNNs) that are usually too power-hungry to be deployed on resource-constrained edge devices. Many techniques improve DNN efficiency of DNNs by compromising accuracy. However, the accuracy and efficiency of these techniques cannot be adapted for diverse edge applications with different hardware constraints and accuracy requirements. This paper demonstrates that a recent, efficient tree-based DNN architecture, called the hierarchical DNN, can be converted into a Directed Acyclic Graph-based (DAG) architecture to provide tunable accuracy-efficiency tradeoff options. We propose a systematic method that identifies the connections that must be added to convert the tree to a DAG to improve accuracy. We conduct experiments on popular edge devices and show that increasing the connectivity of the DAG improves the accuracy to within 1% of the existing high accuracy techniques. Our approach requires 93% less memory, 43% less energy, and 49% fewer operations than the high accuracy techniques, thus providing more accuracy-efficiency configurations. Abhinav Goel, Caleb Tung, Nicholas Eliopoulos, Xiao Hu 0004, George K. Thiruvathukal, James C. Davis 0001, Yung-Hsiang Lu |
ISLPED | 7 |
| 2022 | Automated Discovery of Network Cameras in Heterogeneous Web PagesabstractReduction in the cost of Network Cameras along with a rise in connectivity enables entities all around the world to deploy vast arrays of camera networks. Network cameras offer real-time visual data that can be used for studying traffic patterns, emergency response, security, and other applications. Although many sources of Network Camera data are available, collecting the data remains difficult due to variations in programming interface and website structures. Previous solutions rely on manually parsing the target website, taking many hours to complete. We create a general and automated solution for aggregating Network Camera data spread across thousands of uniquely structured web pages. We analyze heterogeneous web page structures and identify common characteristics among 73 sample Network Camera websites (each website has multiple web pages). These characteristics are then used to build an automated camera discovery module that crawls and aggregates Network Camera data. Our system successfully extracts 57,364 Network Cameras from 237,257 unique web pages. Ryan Dailey, Aniesh Chawla, Sripath Mishra, Josh Majors, Yung-Hsiang Lu, George K. Thiruvathukal |
ACM Trans. Internet Techn. | 7 |
| 2021 | Low-Power Multi-Camera Object Re-Identification using Hierarchical Neural NetworksabstractLow-power computer vision on embedded devices has many applications. This paper describes a low-power technique for the object re-identification (reID) problem: matching a query image against a gallery of previously-seen images. State-of-the-art techniques rely on large, computationally-intensive Deep Neural Networks (DNNs). We propose a novel hierarchical DNN architecture that uses attribute labels in the training dataset to perform efficient object reID. At each node in the hierarchy, a small DNN identifies a different attribute of the query image. The small DNN at each leaf node is specialized to re-identify a subset of the gallery-only the images with the attributes identified along the path from the root to a leaf. Thus, a query image is re-identified accurately after processing with a few small DNNs. We compare our method with state-of-the-art object reID techniques. With a $\sim 4\%$ loss in accuracy, our approach realizes significant resource savings: 74% less memory, 72% fewer operations, and 67% lower query latency, yielding 65% less energy consumption. Abhinav Goel, Caleb Tung, Xiao Hu 0004, James C. Davis 0001, George K. Thiruvathukal, Yung-Hsiang Lu |
ISLPED | 7 |
| 2021 | Adaptive Resource Management for Analyzing Video Streams from Globally Distributed Network CamerasabstractThere has been tremendous growth in the amount of visual data available on the Internet in recent years. One type of visual data of particular interest is produced by network cameras providing real-time views. Millions of network cameras around the world continuously stream data to viewers connected to the Internet. This data may be used by a wide variety of applications such as enhancing public safety, urban planning, emergency response, and traffic management which are computationally intensive. Analyzing this data requires significant amounts of computational resources. Cloud computing can be a preferred solution for meeting the resource requirements for analyzing these data. There are many options when selecting cloud instances (amounts of memory, number of cores, locations, etc.). Inefficient provisioning of cloud resources may become costly in pay-per-use cloud computing. This paper presents a method to select cloud instances in order to meet the performance requirements for visual data analysis at a lower cost. We measure the frame rates when analyzing the data using different computer vision methods and model the relationships between frame rates and resource utilizations. We formulate the problem of managing cloud resources as a Variable Size Bin Packing Problem and use a heuristic solution. Experiments using Amazon EC2 validate the model and demonstrate that the proposed solution can reduce the cost up to 62 percent while meeting the performance requirements. Anup Mohan, Ahmed S. Kaseb, Yung-Hsiang Lu, Thomas J. Hacker |
IEEE Trans. Cloud Comput. | 3 |
| 2021 | Modular Neural Networks for Low-Power Image Classification on Embedded DevicesabstractEmbedded devices are generally small, battery-powered computers with limited hardware resources. It is difficult to run deep neural networks (DNNs) on these devices, because DNNs perform millions of operations and consume significant amounts of energy. Prior research has shown that a considerable number of a DNN’s memory accesses and computation are redundant when performing tasks like image classification. To reduce this redundancy and thereby reduce the energy consumption of DNNs, we introduce the Modular Neural Network Tree architecture. Instead of using one large DNN for the classifier, this architecture uses multiple smaller DNNs (called modules ) to progressively classify images into groups of categories based on a novel visual similarity metric. Once a group of categories is selected by a module, another module then continues to distinguish among the similar categories within the selected group. This process is repeated over multiple modules until we are left with a single category. The computation needed to distinguish dissimilar groups is avoided, thus reducing redundant operations, memory accesses, and energy. Experimental results using several image datasets reveal the effectiveness of our proposed solution to reduce memory requirements by 50% to 99%, inference time by 55% to 95%, energy consumption by 52% to 94%, and the number of operations by 15% to 99% when compared with existing DNN architectures, running on two different embedded systems: Raspberry Pi 3 and Raspberry Pi Zero. Abhinav Goel, Sarah Aghajanzadeh, Caleb Tung, Shuo-Han Chen, George K. Thiruvathukal, Yung-Hsiang Lu |
ACM Trans. Design Autom. Electr. Syst. | 6 |
| 2020 | A Real-Time Feature Indexing System on Live Video StreamsabstractMost of the existing video storage systems rely on offline processing to support the feature-based indexing on video streams. The feature-based indexing technique provides an effective way for users to search video content through visual features, such as object categories (e.g., cars and persons). However, due to the reliance on offline processing, video streams along with their captured features cannot be searchable immediately after video streams are recorded. According to our investigation, buffering and storing live video steams are more time-consuming than the YOLO v3 object detector. Such observation motivates us to propose a real-time feature indexing (RTFI) system to enable instantaneous feature-based indexing on live video streams after video streams are captured and processed through object detectors. RTFI achieves its real-time goal via incorporating the novel design of metadata structure and data placement, the capability of modern object detector (i.e., YOLO v3), and the deduplication techniques to avoid storing repetitive video content. Notably, RTFI is the first system design for realizing real-time feature-based indexing on live video streams. RTFI is implemented on a Linux server and can improve the system throughput by upto 10.60x, compared with the base system without the proposed design. In addition, RTFI is able to make the video content searchable within 20 milliseconds for 10 live video streams after the video content is received by the proposed system, excluding the network transfer latency. Aditya Chakraborty, Akshay Pawar, Hojoung Jang, Shunqiao Huang, Sripath Mishra, Shuo-Han Chen, Yuan-Hao Chang 0001, George K. Thiruvathukal, Yung-Hsiang Lu |
COMPSAC | 9 |
| 2020 | Camera Placement Meeting Restrictions of Computer VisionabstractIn the blooming era of smart edge devices, surveillance cameras have been deployed in many locations. Surveillance cameras are most useful when they are spaced out to maximize coverage of an area. However, deciding where to place cameras is an NP-hard problem and researchers have proposed heuristic solutions. Existing work does not consider a significant restriction of computer vision: in order to track a moving object, the object must occupy enough pixels. The number of pixels depends on many factors (How far away is the object? What is the camera resolution? What is the focal length?). In this study, we propose a camera placement method that identifies effective camera placement in arbitrary spaces and can account for different camera types as well. Our strategy represents spaces as polygons, then uses a greedy algorithm to partition the polygons and determine the cameras' locations to provide the desired coverage. Our solution also makes it possible to perform object tracking via overlapping camera placement. Our method is evaluated against complex shapes and real-world museum floor plans, achieving up to 85% coverage and 25% overlap. Sarah Aghajanzadeh, Roopasree Naidu, Shuo-Han Chen, Caleb Tung, Abhinav Goel, Yung-Hsiang Lu, George K. Thiruvathukal |
ICIP | 6 |
| 2020 | Low-power object counting with hierarchical neural networksabstractDeep Neural Networks (DNNs) achieve state-of-the-art accuracy in many computer vision tasks, such as object counting. Object counting takes two inputs: an image and an object query and reports the number of occurrences of the queried object. To achieve high accuracy, DNNs require billions of operations, making them difficult to deploy on resource-constrained, low-power devices. Prior work shows that a significant number of DNN operations are redundant and can be eliminated without affecting the accuracy. To reduce these redundancies, we propose a hierarchical DNN architecture for object counting. This architecture uses a Region Proposal Network (RPN) to propose regions-of-interest (RoIs) that may contain the queried objects. A hierarchical classifier then efficiently finds the RoIs that actually contain the queried objects. The hierarchy contains groups of visually similar object categories. Small DNNs at each node of the hierarchy classify between these groups. The RoIs are incrementally processed by the hierarchical classifier. If the object in an RoI is in the same group as the queried object, then the next DNN in the hierarchy processes the RoI further; otherwise, the RoI is discarded. By using a few small DNNs to process each image, this method reduces the memory requirement, inference time, energy consumption, and number of operations with negligible accuracy loss when compared with the existing techniques. Abhinav Goel, Caleb Tung, Sarah Aghajanzadeh, Isha Ghodgaonkar, Shreya Ghosh 0003, George K. Thiruvathukal, Yung-Hsiang Lu |
ISLPED | 7 |
| 2020 | Crowdsourcing Detection of Sampling Biases in Image DatasetsabstractDespite many exciting innovations in computer vision, recent studies reveal a number of risks in existing computer vision systems, suggesting results of such systems may be unfair and untrustworthy. Many of these risks can be partly attributed to the use of a training image dataset that exhibits sampling biases and thus does not accurately reflect the real visual world. Being able to detect potential sampling biases in the visual dataset prior to model development is thus essential for mitigating the fairness and trustworthy concerns in computer vision. In this paper, we propose a three-step crowdsourcing workflow to get humans into the loop for facilitating bias discovery in image datasets. Through two sets of evaluation studies, we find that the proposed workflow can effectively organize the crowd to detect sampling biases in both datasets that are artificially created with designed biases and real-world image datasets that are widely used in computer vision research and system development. Xiao Hu 0004, Anirudh Vegesana, Somesh Dube, Kaiwen Yu, Gore Kao, Shuo-Han Chen, Yung-Hsiang Lu, George K. Thiruvathukal, Ming Yin 0001 |
WWW | 8 |
| 2019 | Cloud Resource Management for Analyzing Big Real-Time Visual Data from Network CamerasabstractThousands of network cameras stream real-time visual data for different environments, such as streets, shopping malls, and natural scenes. The big visual data from these cameras can be useful for many applications, but analyzing the large quantities of data requires significant amounts of resources. These resources can be obtained from cloud vendors offering cloud instances (referred to as instances in this paper) with different capabilities and hourly costs. It is a challenging problem to manage cloud resources to reduce the cost for analyzing the big real-time visual data from network cameras while meeting the performance requirements. That is because the problem is affected by many factors related to the analysis programs, the cameras, and the instances. This paper proposes a cloud resource manager (referred to as manager in this paper) that aims at solving this problem. The manager estimates the resource requirements of analyzing the data stream from each camera, formulates the resource allocation problem as a 2D vector bin packing problem, and solves it using a heuristic algorithm. The resource manager monitors the allocated instances; it allocates more instances if needed and deallocates existing instances to reduce the cost if possible. The experiments show that the resource manager is able to reduce up to 60 percent of the overall cost. The experiments use multiple analysis programs, such as moving objects detection, feature tracking, and human detection. One experiment analyzes more than 97 million images (3.3 TB of data) from 5,310 cameras simultaneously over 24 hours using 15 Amazon EC2 instances costing $188. Ahmed S. Kaseb, Anup Mohan, Youngsol Koh, Yung-Hsiang Lu |
IEEE Trans. Cloud Comput. | 4 |
| 2018 | Three years of low-power image recognition challenge: Introduction to special sessionabstractReducing power consumption has been one of the most important goals since the creation of electronic systems. Energy efficiency is increasingly important as battery-powered systems (such as smartphones, drones, and body cameras) are widely used. It is desirable using the on-board computers to recognize objects in the images captured by these cameras. The Low-Power Image Recognition Challenge (LPIRC) is an annual competition started in 2015. The special session includes presentations given by the winners of the first three years of LPIRC. This paper explains the rules of the competition and the rationale, summarizes the teams' scores, and describes the lessons learned in the first three years. The paper suggests possible improvements of future challenges. Kent Gauen, Ryan Dailey, Yung-Hsiang Lu, Eunbyung Park, Wei Liu 0015, Alexander C. Berg, Yiran Chen 0001 |
DATE | 3 |
| 2017 | Low-power image recognition challengeabstractSignificant progress has been made in recent years using computer programs recognizing objects in images. Meanwhile, many cameras are embedded in battery-powered systems (such as mobile phones, wearable devices, and drones) and energy efficiency is essential. Even though many research papers have been published on the topics related to low power and image recognition, there does not exist a common metric for comparing different solutions in terms of (1) energy efficiency and (2) accuracy in recognition. Low-Power Image Recognition Challenge (LPIRC) is, to our knowledge, the only on-site competition that considers both energy consumption and recognition accuracy. LPIRC was held as one-day workshops in the Design Automation Conference in 2015 and 2016. Each participating team brought their own system to the workshops. The referee system of LPIRC includes (1) an intranet, (2) a power meter, and (3) an HTTP server that provided the images and accepted the answers from the contestants' systems. The scores were the ratio of recognition accuracy and the energy consumption. The winner of 2016 was able to analyze 7,347 images and achieve 9.44% normalized mAP (mean average precision) with average power consumption of 4.7 W. Another team analyzed 1,020 images and achieved 25.7% normalized mAP. Kent Gauen, Rohit Rangan, Anup Mohan, Yung-Hsiang Lu, Wei Liu 0015, Alexander C. Berg |
ASP-DAC | 4 |
| 2017 | Internet of video things in 2030: A world with many camerasabstractThe Internet of Things (IoT) is the internetworking of a variety of devices, including sensors. Among all sensors, visual sensors (i.e. cameras) are special because they can provide rich and versatile information. The world already has more than one billion cameras on mobile phones. We define the internet-working of visual sensors as the Internet of Video Things. This article estimates the number of cameras the world will see in 2030 and the implications of a large number of cameras. Transmitting, storing, and analyzing the data from cameras could impose significant challenges to existing technological infrastructures. This paper surveys recent progress in relevant technologies and suggests directions for future research. Anup Mohan, Kent Gauen, Yung-Hsiang Lu, Wei Wayne Li, Xuemin Chen |
ISCAS | 3 |
| 2017 | Privacy Protection in Online MultimediaabstractOnline multimedia has been growing rapidly due to ubiquitous mobile phones, widely deployed surveillance cameras, dashcams and mini-drones. When one takes photographs or videos at a public location, it is highly likely that some other people ("bystanders") also appear in the visual data. The data may be available online, such as shared by social media, and questions about privacy arise. This panel discusses the issues about privacy in online multimedia from legal, technological, and social aspects. Yung-Hsiang Lu, Andrea Cavallaro, Catherine Crump, Gerald Friedland, Keith Winstein |
ACM Multimedia | 1 |
| 2016 | Location Based Cloud Resource Management for Analyzing Real-Time Videos from Globally Distributed Network CamerasabstractThe use of video data for a variety of applications has gained immense popularity. These applications include traffic monitoring, surveillance, retail store management, etc. Thousands of publicly accessible network cameras distributed around the world are potential sources of video data. The applications analyzing data from network cameras have different resource requirements (CPU, memory, etc.) and performance requirements (video frame rate). A preferred way to meet these requirements is to use a cloud infrastructure and a pay-per-use model of the cloud. Cloud computing offers resources, referred to as cloud instances, with different capacities and at different locations. The cost of these instances are dependent on their capacities and locations. The frame rate of the video data obtained from a network camera impacts the accuracy of the analysis and is dependent on the location of the cloud instance. Hence it is important to select the locations of the instances to meet the application performance requirements. This paper presents a resource management approach to select the locations, types, and number of cloud instances to analyze real-time video data from the network cameras while meeting the performance requirements and reducing the analysis cost. We model the resource management problem as a variable size bin packing problem and describe a heuristic algorithm to find a solution. This paper uses Amazon EC2 to evaluate our new resource manager, and observes that our method can reduce the analysis cost up to 56% compared with two other strategies for selecting instance locations. Anup Mohan, Ahmed S. Kaseb, Yung-Hsiang Lu, Thomas J. Hacker |
CloudCom | 3 |
| 2016 | Predictive Model for Dynamically Provisioning Resources in Multi-Tier Web ApplicationsabstractIn the competitive world of modern web applications, performance plays a crucial role. An e-commerce company estimated that every 100ms delay reduces sales by 1 percent, and a popular search engine reported that every 500ms delay in search reduces earnings by 20 percent. The demands from users for these services can vary widely based on factors such as the time-of-day and unexpected events that can trigger flash crowds. To meet these demands web applications can be organized using a multi-tier architecture to make them modular and scalable in a cloud environment. However, a highly dynamic workload and different types of resource requirements in each tier can make it difficult to model the behavior of these applications. This presents two significant challenges to infrastructure providers: 1) to model the behavior of an application workload and provide responsive resources using dynamic resource provisioning, and 2) to maintain performance-based (response time) Service Level Agreements (SLAs). In this paper, we formulate a convex optimization problem for resource allocation, and offer a strict SLA for performance. We adopt an SLA violation cost model to formulate our optimization problem and derive the solution for dynamic resource provisioning. To achieve a strict SLA for the response time of an application, we propose a predictive model that seeks to dynamically provision resources using a Feedback-based Control System (FCS). Our model is applicable for a broad range of multi-tier applications. We demonstrate the effective use of our model through experiments that analyze the behavior of an online auction application using a common workload benchmark. Saurav Nanda, Thomas J. Hacker, Yung-Hsiang Lu |
CloudCom | 3 |
| 2016 | Hybrid Power Management for Office EquipmentabstractOffice machines (such as printers, scanners, facsimile machines, and copiers) can consume significant amounts of power. Most office machines have sleep modes to save power. Power management of these machines is usually timeout-based: a machine sleeps after being idle long enough. Setting the time-out duration can be difficult: if it is too long, the machine wastes power during idleness. If it is too short, the machine sleeps too soon and too often—the wake-up delay can significantly degrade productivity. Thus, power management is a tradeoff between saving energy and keeping response time short. Many power management policies have been published and one policy may outperform another in some scenarios. There is no definite conclusion regarding which policy is always better. This article describes two methods for office equipment power management. The first method adaptively reduces power based on a constraint of the wake-up delay. The second is a hybrid method with multiple candidate policies and it selects the most appropriate power management policy. Using 6 months of request traces from 18 different printers, we demonstrate that the hybrid policy outperforms individual policies. We also discover that power management based on business hours does not produce consistent energy savings. Ganesh Gingade, Wenyi Chen, Yung-Hsiang Lu, Jan P. Allebach, Hernan Ildefonso Gutierrez-Vazquez |
ACM Trans. Design Autom. Electr. Syst. | 3 |
| 2015 | Adaptive Cloud Resource Allocation for Analysing Many Video StreamsabstractThe visual data generated by network cameras can be valuable for a wide range of scientific studies such as weather, wildlife, and traffic. The resource demands for analysis of the data may fluctuate significantly for some of these studies (for example, seasonal or during only rush hours). Cloud computing's pay-per-use can be a preferred solution for analysing large amounts of data from these network cameras. There are few studies focusing on how to allocate cloud resources to analyse many video streams from network cameras. Existing autoscaling and load balancing methods are inapplicable due to the nature of the workloads. This paper presents a solution to manage cloud resources analysing multiple video streams. When an analysis program is launched, the resource manager adjusts the resource allocation by one of two methods: gradually scaling up the number of instances or predicting the needed number of instances. Then, the resource manager monitors the utilization and scales the resources in response to the analysis program's loads. The proposed resource manager has been implemented using Microsoft Azure and demonstrated that the system can adaptively adjust resource to cater to the workload. We explore the trade-offs inherent to both methods. Gradually scaling up can reach more efficient resource allocation, but may take multiple trials. The prediction method is faster in making allocation decisions, but may need additional adjustment. We also investigate the impact of the types of virtual machines and find that smaller sizes are usually better in terms of performance-cost ratios. Wenyi Chen, Yung-Hsiang Lu, Thomas J. Hacker |
CloudCom | 2 |
| 2015 | Rebooting Computing and Low-Power Image Recognition Challengeabstract“Rebooting Computing” (RC) is an effort in the IEEE to rethink future computers. RC started in 2012 by the co-chairs, Elie Track (IEEE Council on Superconductivity) and Tom Conte (Computer Society). RC takes a holistic approach, considering revolutionary as well as evolutionary solutions needed to advance computer technologies. Three summits have been held in 2013 and 2014, discussing different technologies, from emerging devices to user interface, from security to energy efficiency, from neuromorphic to reversible computing. The first part of this paper introduces RC to the design automation community and solicits revolutionary ideas from the community for the directions of future computer research. Energy efficiency is identified as one of the most important challenges in future computer technologies. The importance of energy efficiency spans from miniature embedded sensors to wearable computers, from individual desktops to data centers. To gauge the state of the art, the RC Committee organized the first Low Power Image Recognition Challenge (LPIRC). Each image contains one or multiple objects, among 200 categories. A contestant has to provide a working system that can recognize the objects and report the bounding boxes of the objects. The second part of this paper explains LPIRC and the solutions from the top two winners. Yung-Hsiang Lu, Alan M. Kadin, Alexander C. Berg, Thomas M. Conte, Erik DeBenedictis, Ganesh Gingade, Bichlien Hoang, Yongzhen Huang, Boxun Li, Jingyu Liu 0004, Wei Liu 0015, Huizi Mao, Junran Peng, Tianqi Tang 0001, Elie K. Track, Jingqiu Wang, Tao Wang 0004, Yu Wang 0002 |
ICCAD | 1 |
| 2015 | Opportunities and Challenges of Global Network CamerasabstractSince the introduction of consumer digital cameras, user-created multimedia content has become increasingly popular. Digital cameras, together with inexpensive editing tools, and free hosting sites have made multimedia an integral part of everyday life. Today, hundreds of hours video are uploaded to hosting sites every minute. Video-on-demand through wireless networks and smartphones have profoundly changed how people consume multimedia content. Meanwhile, the widely deployed network cameras can provide live views of many parts of the world. These cameras can provide rich sources creating multimedia content. This panel will explore the opportunities and discuss the challenges using global network cameras for creating multimedia contents and understanding the world. Every year, millions of network cameras are deployed. The data from some of these network cameras are publicly available, continuously streaming live views of national parks, city halls, streets, highways, and shopping malls. A person may see multiple tourist attractions through these cameras, without leaving home. Researchers may observe the weather in different cities. Using the data from the cameras, it is possible to observe natural disasters, such as volcano eruption or tsunami, at a safe distance. News reporters may obtain instant views of an unfolding riot without risking their lives. A spectator may watch a celebration parade from multiple locations using the street cameras. Despite the many promising applications, the opportunities of using global network cameras for creating multimedia content have not been fully exploited. Joanna Batstone, Touradj Ebrahimi, Tiejun Huang 0001, Yung-Hsiang Lu, Yonggang Wen 0001 |
ACM Multimedia | 4 |
| 2015 | LiTMaS: Live road traffic maps for smartphonesabstractA smartphone application that displays a live view of road traffic can provide drivers with real-time information on traffic jams, flash floods or accidents along their planned routes. This information aids them in avoiding delays and uncertainties on the road, enhancing their navigation experience. Several US departments of transportation provide information from street cameras in the form of a snapshot of a particular location on a Google map. However, it is difficult for drivers to manually select cameras on the map to obtain consolidated information about road conditions along a particular route. In this paper, we design an energy-efficient mobile service that provides drivers with a convenient interface to observe live camera coverage along a route. Our proposed solution comprises an Android application and a cloud-based proxy service between smartphones and traffic cameras. The proxy provides an abstraction layer over different communication protocols adopted by traffic cameras, and transfers camera images for a specified route to the smartphone as a batch, thereby reducing communication overhead and improving user response time. By employing an in-memory cache, the proxy server maintains the most recently accessed camera images and reduces camera polling. We integrate our solution with traffic cameras from cities in Massachusetts, New York, and Washington DC, and demonstrate its reduced energy consumption and reduced response time. S. M. Iftekharul Alam, Sonia Fahmy, Yung-Hsiang Lu |
WOWMOM | 3 |
| 2014 | An Instructional Cloud-Based Testbed for Image and Video AnalyticsabstractThis paper describes a cloud-based software infrastructure for teaching big data analytics. Using this infrastructure, students can retrieve and analyze real-time visual data retrieved from globally distributed network cameras. Thomas J. Hacker, Yung-Hsiang Lu |
CloudCom | 2 |
| 2013 | A Survey of Computation Offloading for Mobile Systems
Karthik Kumar, Jibang Liu, Yung-Hsiang Lu, Bharat K. Bhargava |
Mob. Networks Appl. | 3 |
| 2013 | Energy-efficient data dissemination using beamforming in wireless sensor networksabstractEnergy conservation is essential in Wireless Sensor Networks (WSNs) because of limited energy in nodes' batteries. Collaborative beamforming uses multiple transmitters to form antenna arrays; the electromagnetic waves from these antenna arrays can create constructive interferences at the receiver and increase the transmission distance. Each transmitter can use lower power and save energy, since the energy consumption is spread over multiple transmitters. Each beamforming transmission requires multiple collaborative transmitters. Repetitively using the same transmitters will deplete their energy and create coverage holes in the sensing area. To prevent holes, energy should be balanced by using different transmitters. This article investigates the factors that can affect the energy consumption and network lifetime when using beamforming. We present an algorithm for selecting the transmitters in order to prolong the network lifetime. Compared with an existing beamforming transmitters scheduling algorithm, our algorithm doubles the network lifetime. When compared with direct transmission or multihop transmissions towards a receiver far away from the sensing area, our approach can increase the network's lifetime substantially. Yung-Hsiang Lu, Byunghoo Jung, Dimitrios Peroulis, Y. Charlie Hu |
ACM Trans. Sens. Networks | 2 |
| 2012 | A game theoretic resource allocation for overall energy minimization in mobile cloud computing systemabstractCloud computing and virtualization techniques provide mobile devices with battery energy saving opportunities by allowing them to offload computation and execute code remotely. When the cloud infrastructure consists of heterogeneous servers, the mapping between mobile devices and servers plays an important role in determining the energy dissipation on both sides. From an environmental impact perspective, any energy dissipation related to computation should be counted. To achieve energy sustainability, it is important reducing the overall energy consumption of the mobile systems and the cloud infrastructure. Furthermore, reducing cloud energy consumption can potentially reduce the cost of mobile cloud users because the pricing model of cloud services is pay-by-usage. In this paper, we propose a game-theoretic approach to optimize the overall energy in a mobile cloud computing system. We formulate the energy minimization problem as a congestion game, where each mobile device is a player and his strategy is to select one of the servers to offload the computation while minimizing the overall energy consumption. We prove that the Nash equilibrium always exists in this game and propose an efficient algorithm that could achieve the Nash equilibrium in polynomial time. Experimental results show that our approach is able to reduce the total energy of mobile devices and servers compared to a random approach and an approach which only tries to reduce mobile devices alone. Yukan Zhang, Qinru Qiu, Yung-Hsiang Lu |
ISLPED | 4 |
| 2012 | Energy Conservation for Image Retrieval on Mobile SystemsabstractMobile systems such as PDAs and cell phones play an increasing role in handling visual contents such as images. Thousands of images can be stored in a mobile system with the advances in storage technology: this creates the need for better organization and retrieval of these images. Content Based Image Retrieval (CBIR) is a method to retrieve images based on their visual contents. In CBIR, images are compared by matching their numerical representations called features; CBIR is computation and memory intensive and consumes significant amounts of energy. This article examines energy conservation for CBIR on mobile systems. We present three improvements to save energy while performing the computation on the mobile system: selective loading, adaptive loading, and caching features in memory. Using these improvements adaptively reduces the features to be loaded into memory for each search. The reduction is achieved by estimating the difficulty of the search. If the images in the collection are dissimilar, fewer features are sufficient; less computation is performed and energy can be saved. We also consider the effect of consecutive user queries and show how features can be cached in memory to save energy. We implement a CBIR algorithm on an HP iPAQ hw6945 and show that these improvements can save energy and allow CBIR to scale up to 50,000 images on a mobile system. We further investigate if energy can be saved by migrating parts of the computation to a server, called computation offloading . We analyze the impact of the wireless bandwidth, server speed, number of indexed images, and the number of image queries on the energy consumption. Using our scheme, CBIR can be made energy efficient under all conditions. Karthik Kumar, Yamini Nimmagadda, Yung-Hsiang Lu |
ACM Trans. Embed. Comput. Syst. | 3 |
| 2012 | Introduction to the special section on adaptive power management for energy and temperature-aware computing systemsabstractNo abstract available. Ayse K. Coskun, Yung-Hsiang Lu, Qinru Qiu |
ACM Trans. Design Autom. Electr. Syst. | 2 |
| 2011 | Resource Allocation for Real-Time Tasks Using Cloud ComputingabstractThis paper presents a method to allocate resources for real-time tasks using the "Infrastructure as a Service" model offered by cloud computing. Real-time tasks have to be completed before deadlines, and cloud computing offers selection of resources with different speeds and costs. In cloud computing, resource allocations can be scaled up based on the requirements; this is called elasticity and is the key difference from existing multiprocessor task allocation. Scalable resources make economical allocation of resources an important problem. We analyze the problem of allocating resources for a set of realtime tasks such that the economic cost is minimized and all the deadlines are met. We formulate the problem as a constrained optimization problem and propose a polynomial-time solution to allocate resources efficiently. We compare the economic costs and performance provided by our solution with the optimal solution and an EDF (earliest deadline first) method. We show how the cost varies based on the distribution of the tasks. Karthik Kumar, Yamini Nimmagadda, Yung-Hsiang Lu |
ICCCN | 4 |
| 2011 | Memory energy management for an enterprise decision support system
Karthik Kumar, Kshitij A. Doshi, Martin Dimitrov, Yung-Hsiang Lu |
ISLPED | 4 |
| 2011 | Dependence-based multi-level tracing and replay for wireless sensor networks debuggingabstractDue to resource constraints and unreliable communication, wireless sensor network (WSN) programming and debugging remain to be a challenging task. Runtime errors must be constantly monitored, often by checking for violations of certain invariants. Once an error is detected, diagnosis must be performed to identify the origin of the error. Deterministic replay is an error diagnosis method which has long been proposed for distributed systems. However, one of the significant hurdles for applying deterministic replay on WSN is posed by the small program memory on typical sensor nodes. This paper proposes a dependence-based multi-level method for memory-efficient tracing and replay. In the interest of portability across different hardware platforms, the method is implemented as a source-level tracing and replaying tool. To further reduce the code size after tracing instrumentation, a cost model is used for making the decision on which functions to in-line. A prototype for the tool targets C programs is developed on top of the Open64 compiler and is tested using several TinyOS applications running on TelosB motes. Preliminary experimental results show that the test programs, which do not fit the program memory after straightforward instrumentation, can be successfully accommodated in memory using the new method such that the injected errors can be found. Zhiyuan Li 0001, Feng Li 0045, Xiaobing Feng 0002, Saurabh Bagchi, Yung-Hsiang Lu |
LCTES | 6 |
| 2010 | Analysis of Energy Consumption on Data Sharing in Beamforming for Wireless Sensor NetworksabstractCollaborative beamforming is a transmission technique by using multiple transmitters to form antenna arrays and creating highly directional beams to a distant receiver. Collaborative beamforming can enhance energy efficiency in wireless sensor networks. However, the same information needs to be shared among the transmitters in advance. As a result, communication among the transmitters is required; this consumes energy and shortens the lifetime of the sensor network. In this paper, we propose a procedure for data sharing when multiple sensing nodes and multiple transmitters are used and examine the energy consumption for beamforming. We show that beamforming is energy-efficient when the sensor nodes are deployed far away from the base station and the energy consumed for data sharing has negligible effects on the network's lifetime. Yamini Nimmagadda, Yung-Hsiang Lu, Byunghoo Jung, Dimitrios Peroulis, Y. Charlie Hu |
ICCCN | 3 |
| 2010 | Real-time moving object recognition and tracking using computation offloadingabstractMobile robots are widely used for computation-intensive tasks such as surveillance, moving object recognition and tracking. Existing studies perform the computation entirely on robot processors or on dedicated servers. The robot processors are limited by their computation capability; real-time performance may not be achieved. Even though servers can perform tasks faster, the communication time between robots and servers is affected by variations in wireless bandwidths. In this paper, we present a system for realtime moving object recognition and tracking using computation offloading. Offloading migrates computation to servers to reduce the computation time on the robots. However, the migration consumes additional time, referred as communication time in this paper. The communication time is dependent on data size exchanged and the available wireless bandwidth. We estimate the computation and communication needed for the tasks and choose to execute them on robot processors or servers to minimize the total execution time, in order to satisfy real-time constraints. Yamini Nimmagadda, Karthik Kumar, Yung-Hsiang Lu, C. S. George Lee |
IROS | 3 |
| 2010 | Phase difference and frequency offset estimation for collaborative beamforming in sensor networksabstractA fast and power efficient phase difference and frequency offset estimation technique for collaborative beamforming in a wireless sensor network is presented. Common radio blocks are used to implement the phase difference estimation technique and we present an analytical expression for the phase difference estimation accuracy in an AWGN channel. The analysis including the effect of noise and multipath on the estimation accuracy shows the effectiveness of the proposed estimation technique for collaborative beamforming applications. Serkan Sayilir, Yung-Hsiang Lu, Dimitrios Peroulis, Y. Charlie Hu, Byunghoo Jung |
ISCAS | 2 |
| 2010 | Tradeoff between energy savings and privacy protection in computation offloadingabstractOffloading can save energy on mobile systems for computation-intensive applications. The mobile systems send programs and data to grid-powered servers where computation is performed. Offloading, however, causes privacy concerns because sensitive data may be sent to servers. This paper investigates how to protect privacy in computation offloading. We use steganography to hide data before sending them to servers. This paper evaluates the tradeoff between energy savings and privacy protection for content-based image retrieval with different steganographic techniques. We implement these methods on a PDA and compare their energy consumption, performance, and effectiveness of protecting privacy. Jibang Liu, Karthik Kumar, Yung-Hsiang Lu |
ISLPED | 3 |
| 2010 | Energy-Efficient Transmission for Beamforming in Wireless Sensor NetworksabstractEnergy conservation is essential in wireless sensor networks (WSNs) because of limited energy in batteries. Collaborative beamforming uses multiple transmitters to form antenna arrays; the electromagnetic waves from these antenna arrays can create constructive interferences at the receiver and increase the transmission distance. Each transmitter can use lower power and save energy, since the energy consumption is spread over multiple transmitters. However, if the same nodes are always used for collaborative beamforming, these nodes would deplete their energy much sooner and this sensing area will no longer be monitored. To avoid this situation, energy consumption for collaborative beamforming needs to be balanced over the whole network by assigning the transmitters in turns. The transmitters in each round are selected by a scheduler and the energy carried in each node is balanced to increase the number of transmissions. We define the lifetime of a network as the number of transmissions until a certain percentage of the nodes depletes their energy. This paper proposes an algorithm to calculate energy-efficient schedules based on the remaining energy and the phase differences of their signals arriving at the receiver. Compared with an existing algorithm, our algorithm can extend the network lifetime by more than 60%. Serkan Sayilir, Yung-Hsiang Lu, Byunghoo Jung, Dimitrios Peroulis, Y. Charlie Hu |
SECON | 4 |
| 2010 | Adaptation of Multimedia Presentations for Different Display Sizes in the Presence of Preferences and Temporal ConstraintsabstractMultimedia content adaptation has become important due to many devices with different amounts of resources like display sizes, memories, and computation capabilities. Existing studies perform content adaptation on web pages and other media files that have the same start times and durations. In this paper, we present a content adaptation method for multimedia presentations constituting media files with different start times and durations. We perform adaptation based on preferences and temporal constraints specified by authors and generate an order of importance among media files. Our method can automatically generate layouts by computing the locations, start times, and durations of the media files. We compare three solutions to generate layouts: 1) exhaustive search, 2) dynamic programming, and 3) greedy algorithms. We analyze the presentations by varying screen resolutions, media files, preferences, and temporal constraints. Our analysis shows that screen utilizations are 92%, 85%, and 80% for the three methods, respectively. The time to generate layouts for a presentation with 100 media files is 1200, 17, and 10 s, respectively, for the three methods. Yamini Nimmagadda, Karthik Kumar, Yung-Hsiang Lu |
IEEE Trans. Multim. | 3 |
| 2009 | Establishing Trust for Computation OffloadingabstractComputation offloading can extend the battery lifetime of a portable device by migrating computation to grid- powered servers. The server may charge the portable device for the computation performed, thus providing offloading as a service. The portable device has to trust that the server has indeed performed the computation as claimed. We propose a protocol to establish trust for computation offloading. The protocol uses two keys to establish trust. The first key is part of the input for the computation; executing the program with the first key generates the second key. Without executing the program, it is difficult to generate the correct second key. We implement an algorithm that meets the requirements of the protocol on an HP iPAQ PDA and a server. We demonstrate through experiments how the algorithm establishes trust and quantify the overhead required to establish trust. Karthik Kumar, Yamini Nimmagadda, Yung-Hsiang Lu |
ICCCN | 3 |
| 2009 | Preference-based adaptation of multimedia presentations for different display sizesabstractThe increased availability of electronic devices with multimedia capabilities has resulted in a dramatic growth in the usage of multimedia presentations. In this paper, we present a framework for automatic generation of layouts for multimedia presentations for different screen resolutions. We develop a heuristic to find the layout for a screen of given resolution in polynomial time. The heuristic achieves 92-98% performance compared with the exhaustive search for different screen sizes. The execution time of the heuristic is as low as 15 seconds for a presentation of 100 media files whereas the exhaustive search takes an average of 6 hours. Yamini Nimmagadda, Karthik Kumar, Yung-Hsiang Lu |
ICME | 3 |
| 2009 | Energy-efficient image compression in mobile devices for wireless transmissionabstractThis paper presents an adaptive compression algorithm for mobile devices to reduce the transmission energy of images through wireless networks. Our algorithm selects different compression schemes based on the amount of details in the images and the available transmission bandwidths. At low bandwidths, significant amounts of energy can be saved because the reduction in transmission energy outweighs the additional processing energy. We implement this adaptive algorithm on an HP iPAQ hw6945 Personal Digital Assistant (PDA). The algorithm is invoked before transmitting images and is transparent to users. The average energy reduction is 20% with an average delay of 1.1 seconds.respectively. Yamini Nimmagadda, Karthik Kumar, Yung-Hsiang Lu |
ICME | 3 |
| 2009 | Energy Efficient Collaborative Beamforming in Wireless Sensor NetworksabstractReducing energy consumption is critical for wireless sensor networks. The dominant factor is the energy for data transmission. Collaborative beamforming can save transmission energy by improving the directivity of electromagnetic waves so that the signal at the receiver is stronger. Compared with a single transmitter, collaborative beamforming spreads the long distance transmission energy over multiple transmitters. This balances the battery lifetime on individual nodes because each transmitter can use lower power. However, beamforming depends on proper coordination of phases among the participating sensor nodes. This requires communication among the sensor nodes and consumes energy. This paper investigates whether the transmission energy can be saved by using beamforming based on the number of nodes and the total size of data needed to transmit. We determine the minimum size of data to balance the communication overhead when given the total number of nodes. Yung-Hsiang Lu, Byunghoo Jung, Dimitrios Peroulis |
ISCAS | 2 |
| 2009 | Energy Efficient Content-based Image Retrieval for Mobile SystemsabstractThis paper presents a method to save energy for mobile systems that perform content-based image retrieval (CBIR). CBIR is a computation-intensive application and a resource-constrained mobile system may save energy by offloading computation to a grid-powered server. We develop three offloading schemes based on different conditions to save energy for CBIR on mobile systems. We implement a CBIR algorithm on an HP iPAQ hw6945 and physically measure the energy savings. Yu-Ju Hong, Karthik Kumar, Yung-Hsiang Lu |
ISCAS | 3 |
| 2009 | Ranking servers based on energy savings for computation offloadingabstractOffloading may save energy for battery-powered devices by migrating computation to grid-powered servers. Offloading can be provided as a service and the servers charge the devices' users based on the consumed resources. In this paper, we propose a scheme to rank the servers based on the amounts of energy savings. The ranking depends on two factors: (1) the energy saved due to offloading and (2) the energy consumed while waiting for the results. We instrument the offloaded programs to estimate the amounts of computation performed by the servers, and use this information to determine the amounts of saved energy. When the servers perform the offloaded computation, the battery-powered devices wait for the results and consume energy. The ratio of the two factors determines the rank of a server. If a server performs more computation within a shorter duration, the server is ranked higher. We implement our method on an HP iPAQ and demonstrate that our method can effectively rank servers based on energy savings. Karthik Kumar, Yamini Nimmagadda, Yung-Hsiang Lu |
ISLPED | 3 |
| 2009 | Multigrade Security Monitoring for Ad-Hoc Wireless NetworksabstractAd-hoc wireless networks are being deployed in critical applications that require protection against sophisticated adversaries. However, wireless routing protocols, such as the widely-used AODV, are often designed with the assumption that nodes are benign. Cryptographic extensions such as Secure AODV (SAODV) protect against some attacks but are still vulnerable to easily-performed attacks using colluding adversaries, such as the wormhole attack. In this paper, we make two contributions to securing routing protocols. First, we present a protocol called Route Verification (RV) that can detect and isolate malicious nodes involved in routing-based attacks with very high likelihood. However, RV is expensive in terms of energy consumption due to its radio communications. To remedy the high energy cost of RV, we make our second contribution. We propose a multigrade monitoring (MGM) approach. The MGM approach employs a previously developed lightweight local monitoring technique to detect any necessary condition for an attack to succeed. However, local monitoring suffers from false positives due to collisions on the wireless channel. When a necessary condition is detected, the heavy-weight RV protocol is triggered. We show through simulation that MGM applied to AODV generally requires little extra energy compared to baseline AODV, under the common case where there is no attack present. It is also more resource-efficient and powerful than SAODV in detecting attacks. Our work, for the first time, lays out the framework of multigrade monitoring, which we believe fundamentally addresses the tension between security and resource consumption in ad-hoc wireless networks. Matthew Tan Creti, Matthew Beaman, Saurabh Bagchi, Zhiyuan Li 0001, Yung-Hsiang Lu |
MASS | 5 |
| 2009 | A Homogeneous Architecture for Power Policy Integration in Operating SystemsabstractA significant volume of research has concentrated on operating system (OS)-directed power management. The primary focus of previous research has been the development of better policies. In this paper, we provide evidence that one policy may outperform another under different conditions. Hence, it is difficult, or even impossible, to design the "best" policy for all computers. We explain how to select the best policies at runtime without user or administrator intervention by using a software framework called the homogeneous architecture for power policy integration (HAPPI). This architecture is portable across different platforms running Linux. HAPPI specifies common requirements for policies and provides an interface to simplify the implementation of policies in a commodity OS. Our approach allows these policies to be compared simultaneously to select the best policy among a set of distinct policies at runtime. Experimental results indicate that HAPPI achieves energy savings within 4 percent of the best individual policy for each device in several computing systems without a priori knowledge of workloads. Nathaniel Pettis, Yung-Hsiang Lu |
IEEE Trans. Computers | 2 |
| 2008 | Energy conservation by adaptive feature loading for mobile content-based image retrievalabstractWe present an adaptive loading scheme to save energy for content based image retrieval (CBIR) in a mobile system. In CBIR, images are represented and compared by high-dimensional vectors called features. Loading these features into memory and comparing them consumes a significant amount of energy. Our method adaptively reduces the features to be loaded into memory for each query image. The reduction is achieved by estimating the difficulty of the query and by reusing cached features in memory for subsequent queries. We implement our method on a PDA and obtain overall energy reduction of 61.3% compared with an existing CBIR implementation. Karthik Kumar, Yamini Nimmagadda, Yu-Ju Hong, Yung-Hsiang Lu |
ISLPED | 4 |
| 2008 | SeNDORComm: An Energy-Efficient Priority-Driven Communication Layer for Reliable Wireless Sensor NetworksabstractIn many reliable Wireless Sensor Network (WSN) applications, messages have different priorities depending on urgency or importance. For example, a message reporting the failure of all nodes in a region is more important than that for a single node. Moreover, traffic can be bursty in nature, such as when a correlated error is reported by multiple nodes running identical code. Current communication layers in WSNs lack efficient support for these two requirements. We present a priority-driven communication layer, called SeNDORComm, which schedules transmission of packets driven by application-specified priority, buffers and packs multiple messages in a packet, and honors the latency guarantee for a message. We show that SeNDORComm improves energy efficiency, message reliability, and network utilization and delays congestion in a network. We extensively evaluate SeNDORComm using analysis, simulation, and testbed experiments. We demonstrate the improvement in goodput of SeNDORComm over the default communication layer, GenericComm in TinyOS (134.78% for a network of 20 nodes). Vinaitheerthan Sundaram, Saurabh Bagchi, Yung-Hsiang Lu, Zhiyuan Li 0001 |
SRDS | 3 |
| 2008 | Dynamic Voltage Scaling for Multitasking Real-Time Systems With Uncertain Execution TimeabstractDynamic voltage and frequency scaling can save energy for real-time systems. Frequencies are generally assumed proportional to voltages. Previous studies consider the probabilistic distributions of tasks' execution time to assist dynamic voltage scaling in task scheduling. These studies use probability information for intratask voltage scheduling but do not sufficiently explore the opportunities for intertask scheduling to save more energy. This paper presents a new approach to combine intra- and intertask voltage scheduling for better energy savings in hard real-time systems with uncertain task execution time. Our approach takes three steps: 1) We calculate statistically the optimal voltage schedules for multiple concurrent tasks, using earliest deadline first scheduling for an ideal processor that can change the frequency continuously; 2) we then adapt the solution to a processor with a limited range of discrete frequencies, using a polynomial-time heuristic algorithm; and 3) finally, we improve our solution, considering the time and energy overheads of frequency switching for schedulability and energy reduction. Our simulation shows that the new approach can save more energy than existing solutions while meeting hard deadlines. Changjiu Xian, Yung-Hsiang Lu, Zhiyuan Li 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2007 | Energy-Aware Scheduling for Real-Time Multiprocessor Systems with Uncertain Task Execution TimeabstractThis paper presents an energy-aware method to schedule multiple real-time tasks in multiprocessor systems that support dynamic voltage scaling (DVS). The key difference from existing approaches is that we consider the probabilistic distributions of the tasks' execution time to partition the workload for better energy reduction. We analyze the problem of energy-aware scheduling for multiprocessor with probabilistic workload information and derive its mathematical formulation. As the problem is NP-hard, we present a polynomial-time heuristic method to transform the problem into a probability-based load balancing problem that is then solved with worst-fit decreasing bin-packing heuristic. Simulation results with synthetic, multimedia, and stereo-vision tasks show that our method saves significantly more energy than existing methods. Changjiu Xian, Yung-Hsiang Lu, Zhiyuan Li 0001 |
DAC | 2 |
| 2007 | Improving quality-of-service of file migration policies in high-performance serversabstractPrevious work demonstrates that migrating files between disks may induce idleness in servers and facilitate power management. However, these studies focus on steady-state power savings and do not consider the energy consumption and performance overhead of migrating files between disks. In this study, we construct a constrained optimization problem of file location for high-bandwidth applications by considering file migration energy, as well as steady-state energy consumption. We present an algorithm that improves quality-of-service by 63% for a streaming video server compared with existing techniques while maintaining comparable energy savings. Nathaniel Pettis, Yung-Hsiang Lu |
ICPADS | 2 |
| 2007 | Adaptive computation offloading for energy conservation on battery-powered systemsabstractThis paper considers the problem of extending the battery lifetime for a portable computer by offloading its computation to a server. Depending on the inputs, computation time for different instances of a program can vary significantly and they are often difficult to predict. Different from previous studies on computation offloading, our approach does not require estimating the computation time before the execution. We execute the program initially on the portable client with a timeout. If the computation is not completed after the timeout, it is offloaded to the server. We first set the timeout to be the minimum computation time that can benefit from offloading. This method is proved to be 2-competitive. We further consider collecting online statistics of the computation time and and the statistically optimal timeout. Finally, we provide guidelines to construct programs with computation offloading. Experiments show that our methods can save up to 17% more energy than existing approaches. Changjiu Xian, Yung-Hsiang Lu, Zhiyuan Li 0001 |
ICPADS | 2 |
| 2007 | Multi-robot SLAM with topological/metric mapsabstractIn recent years, the success of single-robot SLAM has led to more multi-robot SLAM (MR-SLAM) research. A team of robots with MR-SLAM can explore an environment more efficiently and reliably; however, MR-SLAM also raises many challenging problems, including map fusion, unknown robot poses and scalability issues. The first two problems can be considered as an optimization problem of finding a consistent joint map based on robots’ relative poses and sensory data. This optimization problem exhibits a similar property of a singlerobot topological/metric mapping. To exploit this property, we propose a multi-robot SLAM (MR-SLAM) algorithm, which builds a graph-like topological map with vertices representing local metric maps and edges describing relative positions of adjacent local maps. In this MR-SLAM algorithm, the map fusion between two robots can be naturally done by adding an edge that connects two topological maps, and the estimation of relative robot pose is simply performed by optimizing this edge. For the third scalable problem, the proposed algorithm is also scalable to the number of robots and the size of an environment. Computer simulations with a public data set and experimental work on Pioneer 3-DX robots have been conducted to validate the performance of the proposed MR-SLAM algorithm. H. Jacky Chang, C. S. George Lee, Y. Charlie Hu, Yung-Hsiang Lu |
IROS | 4 |
| 2007 | A programming environment with runtime energy characterization for energy-aware applicationsabstractSystem-level power management has been studied extensively. For further energy reduction, the collaboration from user applications becomes critical. This paper presents a programming environment to ease the construction of energy-aware applications. We observe that energy-aware programs may identify different ways (called options) to achieve the desired functionalities and choose the most energy-efficient option at runtime. Our framework provides a programming interface to obtain the estimated energy consumption for choosing a particular option. The energy is estimated based on runtime energy characterization that records a set of runtime conditions correlated with the energy consumption of the options. We provide the procedure and general guidelines for using the environment to construct energy-aware programs. The prototype demonstrates that (a) energy-aware applications can be programmed easily with our interface, (b) accurate estimates are achieved by integrating multiple runtime conditions, and (c) the framework can make multiple devices collaborate for significant energy savings (15% to 41%) with negligible time and energy overhead (<0.35%). Changjiu Xian, Yung-Hsiang Lu, Zhiyuan Li 0001 |
ISLPED | 2 |
| 2007 | Sensor replacement using mobile robots
Yongguo Mei, Changjiu Xian, Saumitra M. Das, Y. Charlie Hu, Yung-Hsiang Lu |
Comput. Commun. | 5 |
| 2007 | Adaptive correctness monitoring for wireless sensor networks using hierarchical distributed run-time invariant checkingabstractThis article presents a hierarchical approach for detecting faults in wireless sensor networks (WSNs) after they have been deployed. The developers of WSNs can specify “invariants” that must be satisfied by the WSNs. We present a framework, Hierarchical SEnsor Network Debugging (H-SEND), for lightweight checking of invariants. H-SEND is able to detect a large class of faults in data-gathering WSNs, and leverages the existing message flow in the network by buffering and piggybacking messages. H-SEND checks as closely to the source of a fault as possible, pinpointing the fault quickly and efficiently in terms of additional network traffic. Therefore, H-SEND is suited to bandwidth or communication energy constrained networks. A specification expression is provided for specifying invariants so that a protocol developer can write behavioral level invariants. We hypothesize that data from sensor nodes does not change dramatically, but rather changes gradually over time. We extend our framework for the invariants that includes values determined at run-time in order to detect data trends. The value range can be based on information local to a single node or the surrounding nodes' values. Using our system, developers can write invariants to detect data trends without prior knowledge of correct values. Automatic value detection can be used to detect anomalies that cannot be detected in existing WSNs. To demonstrate the benefits of run-time range detection and fault checking, we construct a prototype WSN using CO 2 and temperature sensors coupled to Mica2 motes. We show that our method can detect sudden changes of the environments with little overhead in communication, computation, and storage. Douglas Herbert, Vinaitheerthan Sundaram, Yung-Hsiang Lu, Saurabh Bagchi, Zhiyuan Li 0001 |
ACM Trans. Auton. Adapt. Syst. | 3 |
| 2007 | Low-Power Buffer Management for Streaming DataabstractIn this paper, a method for reducing the power consumption in a pipeline of streaming data is presented. The pipeline is divided into multiple stages and buffer memories are inserted between adjacent stages. By dynamically changing the power state of each stage, the overall power consumption of the pipeline, including the power consumed by each stage and by each buffer, can be reduced. For example, when a buffer is full, the stage immediately before it can be turned off for a certain time period to save power. In this paper, the problem of finding the optimal strategy for managing the power state to minimize average power consumption is studied for a pipeline of data consisting of three stages with two buffers in between. The problem is formulated as an optimal control problem of a hybrid system, namely, a dynamical system with both continuous dynamics and discrete transitions. Several important properties of the optimal solutions are derived. Using these properties, the optimal periodic solutions are obtained analytically for the special case when each stage is allowed to turn on and off at most once during each period. Simulation results using practical data are presented. The results show substantial power savings of the proposed method over heuristic buffering strategies Jason Ridenour, Jianghai Hu, Nathaniel Pettis, Yung-Hsiang Lu |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2007 | P-SLAM: Simultaneous Localization and Mapping With Environmental-Structure PredictionabstractTraditionally, simultaneous localization and mapping (SLAM) algorithms solve the localization and mapping problem in explored regions. This paper presents a prediction-based SLAM algorithm (called P-SLAM), which has an environmental-structure predictor to predict the structure inside an unexplored region (i.e., look-ahead mapping). The prediction process is based on the observation of the surroundings of an unexplored region and comparing it with the built map of explored regions. If a similar environment/structure is matched in the map of explored regions, a hypothesis is generated to indicate that a similar structure has been explored before. If the environment has repeated structures, the mobile robot can use the predicted structure as a virtual mapping, and decide whether or not to explore the unexplored region to save the exploration time. If the mobile robot decides to explore the unexplored region, a correct prediction can be used to speed up the SLAM process and build a more accurate map. We have also derived the Bayesian formulation of P-SLAM to show its compact recursive form for real-time operation. We have experimentally implemented the proposed P-SLAM on a Pioneer 3-DX mobile robot using a Rao-Blackwellized particle filter in real time. Computer simulations and experimental results validated the performance of the proposed P-SLAM and its effectiveness in indoor environments H. Jacky Chang, C. S. George Lee, Yung-Hsiang Lu, Y. Charlie Hu |
IEEE Trans. Robotics | 3 |
| 2006 | Automatic run-time selection of power policies for operating systemsabstractA significant volume of research has concentrated on operating-system directed power management (OSPM). The primary focus of previous research has been the development of OSPM policies. Under different conditions, one policy may outperform another and vice versa. In this paper, we explain how to select the best policies at run-time without user or administrator intervention. We present a hardware-neutral architecture portable across different platforms running Linux. Our experiments reveal that changing policies at run-time can adapt to workloads more quickly than using any of the policies individually. Nathaniel Pettis, Jason Ridenour, Yung-Hsiang Lu |
DATE | 3 |
| 2006 | Energy reduction by workload adaptation in a multi-process environmentabstractReducing energy consumption is an important issue in modern computers. Dynamic power management (DPM) has been extensively studied in recent years. One approach for DPM is to adjust workloads, such as clustering or eliminating requests, as a way to trade-off energy consumption and quality of services. Previous studies focus on single processes. However, when multiple concurrently running processes are considered, workload adjustment must be determined based on the interleaving of the processes' requests. When multiple processes share the same hardware component, adjusting one process may not save energy. This paper presents an approach to assign energy responsibility to individual processes based on how they affect power management. The assignment is used to estimate potential energy reduction by adjusting the processes. We use the estimation to guide runtime adaptation of workload behavior. Experiments demonstrate that our approach can save more energy and improve energy efficiency. Changjiu Xian, Yung-Hsiang Lu |
DATE | 2 |
| 2006 | Dynamic voltage scaling for multitasking real-time systems with uncertain execution timeabstractDynamic voltage scaling (DVS) for real-time systems has been extensively studied to save energy. Previous studies consider the probabilistic distributions of tasks' execution time to assist DVS in task scheduling. These studies use probability information for intra-task frequency scheduling but do not sufficiently explore the opportunities for inter-task scheduling to save more energy. This paper presents a new approach to integrate intra-task and inter-task frequency scheduling for better energy savings in hard real-time systems with uncertain task execution time. Our approach has two steps: (a) We calculate statistically optimal frequency schedules for multiple periodic tasks using earliest deadline first (EDF) scheduling for processors that can change frequencies continuously. (b) For processors with a limited range of discrete frequencies, we further present a heuristic algorithm to construct frequency schedules. Our evaluation shows that our approach saves up to 23% more energy than existing solutions. Changjiu Xian, Yung-Hsiang Lu |
ACM Great Lakes Symposium on VLSI | 2 |
| 2006 | Simultaneous Localization and Mapping with Environmental Structure PredictionabstractTraditionally, the SLAM problem solves the localization and mapping problem in explored and sensed regions. This paper presents a prediction-based SLAM algorithm (called P-SLAM), which has an environmental structure predictor to predict the structure inside an unexplored region (i.e., look-ahead mapping). The prediction process is based on the observation of the surroundings of an unexplored region and comparing it with the built map of explored regions. If a similar structure is matched in the map of explored regions, a hypothesis is generated to indicate that a similar structure has been explored before. If the environment has repeated structures, the mobile robot can utilize the predicted structure as a virtual mapping, and decide whether or not to explore the unexplored region to save exploration time. If the mobile robot decides to explore the unexplored region, a correct prediction can be utilized to localize the robot and speed up the SLAM process. We also derive the Bayesian formulation of P-SLAM to show its compact recursive form for real-time operation. We have experimentally implemented the proposed P-SLAM in a Pioneer 3-DX mobile robot using a Rao-Blackwellized particle filter in real-time. Computer simulations and experimental results validated the performance of the proposed P-SLAM and its effectiveness in an indoor environment H. Jacky Chang, C. S. George Lee, Yung-Hsiang Lu, Y. Charlie Hu |
ICRA | 3 |
| 2006 | Energy-efficient Mobile Robot ExplorationabstractMobile robots can be used in many applications, including exploration in an unknown area. Robots usually carry limited energy so energy conservation is vital. This paper presents an approach for energy-efficient robot exploration. Our approach determines the next target for the robot to visit based upon orientation information. The robot plans the path between the current position to the next target in an energy-efficient way. Our method reduces repeated coverage, a common problem for most existing utility-based target selecting methods. We conduct simulations for both random and structured environments, and compare our method with a utility-based method that chooses the middle cell from the widest opening. Results show that our method can reduce energy consumption by 42% and traveling distance by 41% Yongguo Mei, Yung-Hsiang Lu, C. S. George Lee, Y. Charlie Hu |
ICRA | 2 |
| 2006 | Power reduction of multiple disks using dynamic cache resizing and speed controlabstractThis paper presents an energy-conservation method for multiple disks and their cache memory. Our method periodically resizes the cache memory and controls the rotation speeds under performance constraints. The cache memory stores the data from the disks for reuse. Enlarging the cache memory reduces disk accesses and disk utilization. This allows the disks to reduce their speeds and conserve energy because the disks' power consumption is quadratic to their speeds. However, the cache memory itself consumes power to retain data. Shrinking cache memory can save memory power while increasing disk accesses and degrading performance. Choosing proper cache sizes and rotation speeds can reduce the energy consumption of both memory and disks with satisfactory performance. We model cache resizing and speed setting as an optimization problem with minimizing the power consumption as objective and limiting disk utilization as constraints. We compare our method with the methods resizing cache based on request rates. The simulation results show that our method achieves better energy savings while limiting disk access latency. Le Cai, Yung-Hsiang Lu |
ISLPED | 2 |
| 2006 | Energy-Effcient Scheduling for Autonomous Mobile RobotsabstractAn autonomous robot has to sense its environment and prevent collision. Stereovision is a widely used technology for distance calculation and collision prevention. While a robot is moving, it has to detect an obstacle before a collision. This result in a real-time constraint and it may be adjusted by slowing down the motor. Since mobile robots usually carry limited energy, energy conservation is crucial. Stereovision requires substantial amounts of computation and causes much energy consumption. This paper presents a new approach for energy conservation for a mobile robot. We consider the energy consumed by both the on-board processor and the motor. Our method controls both the processor's frequency and the motor's speed to reduce energy while preventing collision. We formulate the problem as non-linear optimization and demonstrate that more energy can be saved by adjusting both the frequency and the speed simultaneously Jeff Brateman, Changjiu Xian, Yung-Hsiang Lu |
VLSI-SoC | 3 |
| 2006 | Program Counter-Based Prediction Techniques for Dynamic Power ManagementabstractReducing energy consumption has become one of the major challenges in designing future computing systems. This paper proposes a novel idea of using program counters to predict I/O activities in the operating system. It presents a complete design of program-counter access predictor (PCAP) that dynamically learns the access patterns of applications and predicts when an I/O device can be shut down to save energy. PCAP uses path-based correlation to observe a particular sequence of program counters leading to each idle period and predicts future occurrences of that idle period. PCAP differs from previously proposed shutdown predictors in its ability to: 1) correlate I/O operations to particular behavior of the applications and users, 2) carry prediction information across multiple executions of the applications, and 3) attain higher energy savings while incurring lower mispredictions. We perform an extensive evaluation study of PCAP using a detailed trace-driven simulation and an actual Linux implementation. Our results show that PCAP achieves lower average mispredictions and higher energy savings than the simple timeout scheme and the state-of-the-art learning tree scheme. Chris Gniady, Ali Raza Butt, Y. Charlie Hu, Yung-Hsiang Lu |
IEEE Trans. Computers | 4 |
| 2006 | Statistically Optimal Dynamic Power Management for Streaming DataabstractThis paper presents a method that uses data buffers to create long periods of idleness to exploit power management. This method considers the power consumed by the buffers and assigns an energy penalty for buffer underflow. Our approach provides analytic formulas for calculating the optimal buffer sizes without subjective or heuristic decisions. We simulate four different hardware configurations with MPEG-1, MPEG-2, and MPEG-4 formats as a case study. Our results indicate that the optimal buffer size varies significantly for different data formats on different hardware. Simulation results indicate that 16 MB buffers are sufficient for MPEG-1, MPEG-2, and MPEG-4 video streams from a microdrive or a network card, but transfers from an IDE disk require buffer sizes ranging from 16 MB to 176 MB, depending on each video's statistical properties. Nathaniel Pettis, Le Cai, Yung-Hsiang Lu |
IEEE Trans. Computers | 3 |
| 2006 | Joint Power Management of Memory and Disk Under Performance ConstraintsabstractThis paper presents a method to combine memory resizing and disk shutdown to achieve better energy savings than can be achieved individually. The method periodically adjusts the size of physical memory and the timeout value to shut down a hard disk to reduce the average energy consumption. Pareto distributions are used to model the disk idle time. The parameters of the distributions are estimated at runtime and used to calculate the appropriate timeout value. The memory size is changed based on the predicted number of disk accesses at different memory sizes. The method also considers the delay caused by power management and limits the performance degradation. The method is simulated and compared with other power management methods. Simulation results show that the method consistently achieves better energy savings and less performance degradation across different workloads Le Cai, Nathaniel Pettis, Yung-Hsiang Lu |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2006 | Deployment of mobile robots with energy and timing constraintsabstractMobile robots can be used in many applications, such as carpet cleaning, search and rescue, and exploration. Many studies have been devoted to the control, sensing, and communication of robots. However, the deployment of robots has not been fully addressed. The deployment problem is to determine the number of groups unloaded by a carrier, the number of robots in each group, and the initial locations of those robots. This paper investigates robot deployment for coverage tasks. Both timing and energy constraints are considered; the robots carry limited energy and need to finish the tasks before deadlines. We build power models for mobile robots and calculate the robots' power consumption at different speeds. A speed-management method is proposed to decide the traveling speeds to maximize the traveling distance under both energy and timing constraints. Our method uses rectangle scanlines as the coverage routes, and solves the deployment problem using fewer robots. Finally, we provide an approach to consider areas with random obstacles. Compared with two simple heuristics, our solution uses 36% fewer robots for open areas and 32% fewer robots for areas with obstacles. Yongguo Mei, Yung-Hsiang Lu, Y. Charlie Hu, C. S. George Lee |
IEEE Trans. Robotics | 2 |
| 2005 | Joint Power Management of Memory and DiskabstractThe paper presents a scheme to combine memory and power management for achieving better energy reduction. Our method periodically adjusts the size of physical memory and the timeout value to shut down a hard disk for reducing the average power consumption. We use Pareto distributions to model the distributions of idle time. The parameters of the distributions are adjusted at run-time for calculating the corresponding timeout value of the disk power management. The memory size is changed based on the inclusion property to predict the number of disk accesses at different memory sizes. Experimental results show more than 50% energy savings compared to a 2-competitive fixed-timeout method. Le Cai, Yung-Hsiang Lu |
DATE | 2 |
| 2005 | An Efficient Group Communication Protocol for Mobile RobotsabstractMobile robot teams have many useful applications such as search and rescue, exploration and hazard detection and analysis. Communication between the robots of a team as well as between the robots and a human operator or controller are useful for many applications. Many applications of mobile robots involve scenarios in which no communication infrastructure such as base stations exist (e.g. demining in battlefields) or the existing infrastructure is damaged (e.g. search and rescue after an earthquake). In such scenarios, it is necessary for mobile robots to form an ad hoc network to enable communication by forwarding each other’s packets. In many applications, group communication can be used for flexible control, organization, and management of the mobile robots. Multicast provides a bandwidth efficient communication method between a source and a group of robots. In this paper, we propose an efficient multicast protocol MRMM (Mobile Robot Mesh Multicast) for deployment in mobile robot networks. MRMM exploits the fact that mobile robots know what velocity they are instructed to move at and for what distance in building a long lifetime sparse mesh for group communication that is more efficient. Our results show that MRMM provides an efficient group communication mechanism that can potentially be used in many mobile robot application scenarios. Saumitra M. Das, Y. Charlie Hu, C. S. George Lee, Yung-Hsiang Lu |
ICRA | 4 |
| 2005 | Efficient Unicast Messaging for Mobile RobotsabstractMobile multi-robot teams are useful in many critical applications such as search and rescue. Explicit communication among robots in such mobile multi-robot teams is useful for the coordination of such teams as well as exchanging data. Since many applications for mobile robots involve scenarios in which communication infrastructure may be damaged or unavailable, mobile robot teams frequently need to communicate with each other by using ad hoc networking. In such scenarios, energy efficient routing protocols to deliver messages among robots are a key requirement. In this paper, we propose and evaluate two routing protocols tailored for use in ad hoc networks formed by mobile multi-robot teams: Mobile Robot Distance Vector (MRDV) and Mobile Robot Source Routing (MRSR). Both protocols exploit the unique mobility characteristics of mobile robot networks to perform efficient routing. Our simulation study show that both MRDV and MRSR incur lower overhead while operating in mobile robot networks when compared to traditional mobile ad hoc network routing protocols such as DSR and AODV. Saumitra M. Das, Y. Charlie Hu, C. S. George Lee, Yung-Hsiang Lu |
ICRA | 4 |
| 2005 | Deployment Strategy for Mobile Robots with Energy and Timing ConstraintsabstractMobile robots usually carry limited energy and have to accomplish their tasks before deadlines. Examples of these tasks include search and rescue, landmine detection, and carpet cleaning. Many researchers have been studying control, sensing, and coordination for these tasks. However, one major problem has not been fully addressed: the initial deployment of mobile robots. The deployment problem considers the number of robots needed and their initial locations. In this paper, we present a solution for the deployment problem when robots have limited energy and time to collectively accomplish coverage tasks. Simulation results show that our method uses 26% fewer robots comparing with two heuristics for covering the same size of area. Yongguo Mei, Yung-Hsiang Lu, Y. Charlie Hu, C. S. George Lee |
ICRA | 2 |
| 2005 | Reducing the number of mobile sensors for coverage tasksabstractMobile robots provide new opportunities for research in environment sensing. By combining sensors and robots, we can develop applications to enhance security, search and rescue, and detect hazardous materials. One important problem is to use fewer mobile sensors to cover an area. This paper focuses on two problems in robot sensor deployment: speed management and detouring distance under energy and timing constraints. We determine the robots' traveling speed based on the remaining energy and remaining time before deadline. We use probabilistic models for environments with obstacles and propose an empirical analysis to estimate the extra traveling distance due to detouring. We compute the number of robots (fleet size) needed for cover an area and our approach reduces the fleet size. Compared with two simple heuristics, our method uses 21% fewer robots. Yongguo Mei, Yung-Hsiang Lu, Y. Charlie Hu, C. S. George Lee |
IROS | 2 |
| 2005 | Energy management using buffer memory for streaming dataabstractThis work presents a new approach for energy management by inserting data buffers. The buffers are managed using the concept of inventory control by treating energy as cost and computation as production. The data stored in the buffers are considered as merchandise. Our method provides a mathematical framework to develop general strategies managing the energy consumption in computers. The approach can solve a wide range of problems. For example, it can calculate the needed buffer sizes to achieve optimal energy savings. It can derive the conditions when a standby state saves energy. The method also shows the effect of load balancing on energy conservation and can handle the variations of data rates. We use sensor networks as case studies and demonstrate more than 20% energy savings. Le Cai, Yung-Hsiang Lu |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2004 | Dynamic Power Management Using Data BuffersabstractThis paper presents a method to reduce energy consumption by inserting data buffers. The method determines whether power can be reduced by inserting a buffer between two components and periodically turning off one of them. This method calculates the length of the period and the required buffer size to achieve the optimal energy savings. Our approach can be applied to any applications whose data arrival and departure rates are different and known in advance. Le Cai, Yung-Hsiang Lu |
DATE | 2 |
| 2004 | Improving FSM evolution with progressive fitness functionsabstractThis paper presents a method to reduce the total number of generations needed to evolve a nite state machine using genetic inferencing. Genetic inferencing is an evolution method that creates designs from their input-output relationship. We reduce the time required to evolve a design by only evolving a small partition of the input-output relationship. We construct the complete design by repeatedly evolving different input-output relationship partitions. Genetically inferring a design with our method can reduce evolution time by more than 80%. Jason W. Horihan, Yung-Hsiang Lu |
ACM Great Lakes Symposium on VLSI | 2 |
| 2004 | Program Counter Based Techniques for Dynamic Power ManagementabstractReducing energy consumption has become one of the major challenges in designing future computing systems. We propose a novel idea of using program counters to predict I/O activities in the operating system. We present a complete design of program-counter access predictor (PCAP) that dynamically learns the access patterns of applications and predicts when an I/O device can be shut down to save energy. PCAP uses path-based correlation to observe a particular sequence of program counters leading to each idle period, and predicts future occurrences of that idle period. PCAP differs from previously proposed shutdown predictors in its ability to: (1) correlate I/O operations to particular behavior of the applications and users, (2) carry prediction information across multiple executions of the applications, and (3) attain better energy savings while incurring low mispredictions. Chris Gniady, Y. Charlie Hu, Yung-Hsiang Lu |
HPCA | 3 |
| 2004 | Efficient Collection of Sensor Data in Remote Fields Using Mobile CollectorsabstractThis paper proposes using a mobile collector, such as an airplane or a vehicle, to collect sensor data from remote fields. We present three different schedules for the collector: round-robin, rate-based, and min movement. The data are not immediately transmitted to the base station after being sensed but buffered at a cluster head; hence, it is important to ensure the latency is within an acceptable range. We compare the latency and the energy expended of the three schedules. We use the ns-2 network simulator to study the scenarios and illustrate conditions under which rate-based outperforms round-robin in latency, and vice-versa. The benefit of min movement is in minimizing the energy expended. Yuldi Tirta, Zhiyuan Li 0001, Yung-Hsiang Lu, Saurabh Bagchi |
ICCCN | 3 |
| 2004 | Energy-time-efficient Adaptive Dispatching Algorithms for Ant-like Robot SystemsabstractIn this paper, we investigate energy-time-efficient dispatching methods for ant-like robots to cover an unmapped region effectively. These ant-like robots have very limited energy and sensor ability, making them practical and inexpensive to build. Our dispatching model was based on bio-inspired algorithms from an ant colony system. We assumed that all the ant-like robots start from their home starting point, the nest, and the region is composed of floor tiles that can be modelled as vertices in a graph. In this dispatching system, the ant-like robots leave pheromone on the tiles and use this information to cover the region. We developed and analyzed two different adaptive dispatching algorithms with different communication methods to the nest. We further compared these two adaptive dispatching algorithms with two non-adaptive methods. Extensive computer simulations validated the proposed adaptive algorithms, showing that they can dispatch ant-like mobile robots to cover an unmapped region with energy-time efficiency. H. Jacky Chang, C. S. George Lee, Yung-Hsiang Lu, Y. Charlie Hu |
ICRA | 3 |
| 2004 | Supporting many-to-one Communication in Mobile Multi-robot ad hoc Sensing NetworksabstractWe study the problem of supporting communication in mobile sensor networks formed by teams of mobile robots equipped with sensors. We assume that in addition to communicating with each other by implicit or environmental means, the robots are equipped with wireless communication capability and effectively form a mobile ad hoc network (MANET). However, unlike typical MANETs where the primary communication pattern is any-to-any, such mobile multi-robot sensing networks need to support sensing applications which exhibit a many-to-one communication pattern. We evaluate the ability of current ad hoc network routing protocols to support communication in such sensing networks using detailed simulations of 50 robots. We consider typical communication patterns that arise in sensing applications and their impact on success rates, routing overhead, and energy costs of the mobile sensor network. Saumitra M. Das, Y. Charlie Hu, C. S. George Lee, Yung-Hsiang Lu |
ICRA | 4 |
| 2004 | Energy-efficient Motion Planning for Mobile RobotsabstractThis paper presents a new approach to find energy-efficient motion plans for mobile robots. Motion planning has two goals: finding the routes and determining the velocities. We model the relationship of motors' speed and their power consumption with polynomials. The velocity of the robot is related to its wheels' velocities by performing a linear transformation. We compare the energy consumption of different routes at different velocities and consider the energy consumed for acceleration and turns. We use experiment-validated simulation to demonstrate up to 51% energy savings for searching an open area. Yongguo Mei, Yung-Hsiang Lu, Y. Charlie Hu, C. S. George Lee |
ICRA | 2 |
| 2004 | A computational efficient SLAM algorithm based on logarithmic-map partitioningabstractSimultaneous localization and map building (SLAM) is a fundamental and complex problem in mobile robot research. In SLAM, Kalman-filter-like implementations are widely adopted to localize a mobile robot and build a map simultaneously and incrementally. However, this approach requires extensive computations of order O(N/sup 2/), where N is the total number of landmarks. To make the computations more manageable, we propose a logarithmic map partitioning algorithm that partitions the global map into one local region and several sub-maps. The size of each sub-map is based on its distance from the mobile robot, and in each sub-map, a centroid landmark is selected to represent all the landmarks in the sub-map for SLAM computations. With this logarithmic-map partitioning, it maintains correlation updates with each sub-map and provides an efficient suboptimal solution to the SLAM problem. The number of landmarks reduces from N to a logarithm-based function of N, and the computational requirement reduces from O(N/sup 3/) to O(N/sup 2/), where N/sub L/ is the number of local landmarks. Furthermore, utilizing the compressed extended Kalman filter, the real-time computational complexity reduces to O(N/sub L//sup 2/). Computer simulation results showed that the proposed algorithm is consistent and efficient for a large number of landmarks. H. Jacky Chang, C. S. George Lee, Yung-Hsiang Lu, Y. Charlie Hu |
IROS | 3 |
| 2004 | Determining the fleet size of mobile robots with energy constraintsabstractAs robotics technologies improve, mobile robots can be used in many applications. A fundamental question is to decide the number of robots needed (i.e., the "fleet-size problem") to accomplish tasks. Previous studies did not consider the energy constraints of the fleet size problem. In this paper, we present a probabilistic method to decide the fleet size for serving random requests. A simplification method is provided for fast computation. Our method is computationally efficient, with an average of 4% errors validated by an event-driven simulator. Yongguo Mei, Yung-Hsiang Lu, Y. Charlie Hu, C. S. George Lee |
IROS | 2 |
| 2004 | Dynamic power management for streaming dataabstractThis paper presents a method that uses data buffers to smoothen request variations and to create long idleness for power management. This method considers the power consumed by the buffers and assigns an energy penalty for buffer underflow. Our approach provides analytic formulas for calculating the optimal buffer sizes and the amount of data to store in the buffers. We use video prefetching as a case study and obtain power savings of more than 74% for MPEG-1 and 34% for MPEG-2 videos. Nathaniel Pettis, Le Cai, Yung-Hsiang Lu |
ISLPED | 3 |
| 2004 | An overview of problems in image-based location awareness and navigationabstractIn this paper we describe some of the research issues and challenges in image-based location awareness and navigation. We will describe two systems being developed at Purdue University as testbeds for our ideas. The main system architecture combines image processing, mobility, wireless communication, and location awareness. We will describe two fundamental scenarios for using images to aid in mobile navigation problems. The first provides the ability to use a locally acquired image to determine the identity of an object, for example a building, as one roams in an area. The second problem is the use of images in a database to aid in vehicle navigation. The solution to these problems use location information, such as GPS signals, to compare and search location-annotated images in a database. We believe location information can improve the accuracy in image database search. Yung-Hsiang Lu, Edward J. Delp |
VCIP | 1 |
| 2002 | Dynamic Power Management for Nonstationary Service RequestsabstractDynamic power management (DPM) is a design methodology aimed at reducing power consumption of electronic systems by performing selective shutdown of idle system resources. The effectiveness of a power management scheme depends critically on accurate modeling of service requests and on computation of the control policy. In this work, we present an online adaptive DPM scheme for systems that can be modeled as finite-state Markov chains. Online adaptation is required to deal with initially unknown or nonstationary workloads, which are very common in real-life systems. Our approach moves from exact policy optimization techniques in a known and stationary stochastic environment and extends optimum stationary control policies to handle the unknown and nonstationary stochastic environment for practical applications. We introduce two workload learning techniques based on sliding windows and study their properties. Furthermore, a two-dimensional interpolation technique is introduced to obtain adaptive policies from a precomputed look-up table of optimum stationary policies. The effectiveness of our approach is demonstrated by a complete DPM implementation on a laptop computer with a power-manageable hard disk that compares very favorably with existing DPM schemes. Eui-Young Chung, Luca Benini, Alessandro Bogliolo, Yung-Hsiang Lu, Giovanni De Micheli |
IEEE Trans. Computers | 4 |
| 2002 | Dynamic frequency scaling with buffer insertion for mixed workloadsabstractThis paper presents a method to reduce the energy of interactive systems for mixed workloads: multimedia applications that require constant output rates and sporadic jobs that need prompt responses. The authors' method divides multimedia programs into stages and inserts data buffers between them. Data buffering has three purposes: (1) to support constant output rates; (2) to allow frequency scaling for energy reduction; and (3) to shorten the response times of sporadic jobs. The authors construct frequency-assignment graphs. Each vertex represents the current state of the buffers and the frequencies of the processor. The authors develop an efficient graph-walk algorithm that assigns frequencies to reduce energy. The same method. can be applied to perform voltage scaling and the combination of frequency and voltage scaling. The authors' experimental results on a Strong-ARM-based computer show that four discrete frequencies are sufficient to achieve nearly maximum energy saving. The method reduces the power consumption of an MPEG program by 46%. The authors also demonstrate a case that shortens the response time of a sporadic job by 55%. Yung-Hsiang Lu, Luca Benini, Giovanni De Micheli |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2002 | Power-aware operating systems for interactive systemsabstractMany portable systems deploy operating systems (OS) to support versatile functionality and to manage resources, including power. This paper presents a new approach for using OS to reduce the power consumption of I/O devices in interactive systems. Low-power OS observes the relationship between hardware devices and processes. The OS kernel estimates the utilization of a device from each process. If a device is not used by any running process, the OS puts it into a low-power state. This paper also explains how scheduling can facilitate power management. When processes are properly scheduled, power reduction can be achieved without degrading performance. We implemented a prototype on Linux to control two devices; experimental results showed nearly 70% power saving on a network card and a hard disk drive. Yung-Hsiang Lu, Luca Benini, Giovanni De Micheli |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2000 | Quantitative Comparison of Power Management AlgorithmsabstractDynamic power management saves power by shutting down idle devices. Several management algorithms have been proposed and demonstrated to be effective in certain applications. We quantitatively compare the power saving and performance impact of these algorithms on hard disks of a desktop and notebook computers. This paper has three contributions. First, we build a framework in Windows NT to implement power managers running realistic workloads and directly interacting with users. Second, we define performance degradation that reflects user perception. Finally, we compare power saving and performance of existing algorithms and analyze the difference. Yung-Hsiang Lu, Eui-Young Chung, Tajana Rosing, Giovanni De Micheli, Luca Benini |
DATE | 1 |
| 2000 | Operating-system directed power reductionabstractthis paper presents a new approach for power reduction by taking a global, software-centric view. It analyzes the sources of power consumption: tasks that require services from hardware components. When a component is not used by any task, it can enter a sleeping state to save power. Operating systems have detailed information about tasks; therefore, OS is the best place for identifying hardware idleness and shutting down unused components. We implement this technique in Linux and show that it can save more than 50% power compared to traditional hardware-centric shutdown techniques. Yung-Hsiang Lu, Luca Benini, Giovanni De Micheli |
ISLPED | 1 |
| 1999 | Adaptive Hard Disk Power Management on Personal ComputersabstractDynamic power management can be effective for designing low-power systems. In many systems, requests are clustered into sessions. This paper proposes an adaptive algorithm that can predict session lengths and shut down components between sessions to save power. Compared to other approaches, simulations show that this algorithm can reduce power consumption in hard disks with less impact on performance or reliability. Yung-Hsiang Lu, Giovanni De Micheli |
Great Lakes Symposium on VLSI | 1 |