Marilyn Wolf

dblp:376/6566 · also Marilyn Claire Wolf, Wayne H. Wolf, Wayne Hendrix Wolf · DBLP profile ↗
← Back
200ranked-venue papers
47as first author
12since 2021 · last 2026
0000-0002-4742-0841ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 120 · 28 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 48 · 9 first-author · 1 since 2021Software engineering, systems software and programming languages · 20 · 3 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 14 · 5 first-authorArtificial intelligence and machine learning · 8 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 6 · 2 since 2021Computer networks · 5Human-computer interaction and ubiquitous computing · 2 · 1 first-authorTheory of computation · 2 · 1 first-author
YearPublicationVenuePosition
2026 TCAD Editorial
Marilyn Wolf, Aviral Shrivastava
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2025 Law as a Design Consideration for Automated Vehicles Suitable to Transport Intoxicated Persons
abstract
This essay explains why an automated vehicle (AV) manufacturer should consider law during the design process for an AV intended as “fit-for-purpose” to transport intoxicated persons. It suggests that management, marketing, engineering and legal functions collaborate to develop product requirements and specifications that shield owner/occupants from criminal liability for DUI manslaughter, negligent homicide and similar charges, as well as guard against civil liability. This collaboration should occur for AV deployments in any state of the United States and in European countries.
William H. Widen, Marilyn Wolf
DATE2
2025 Network-on-Interposer Co-Design for Heterogeneous Chiplet-Based Integrated Systems
abstract
A co-design methodology is proposed for heterogeneous integrated systems with active interposers. Bus-based topologies are impractical for heterogeneous systems due to the latency of the chiplet interfaces which degrades the data transfer process. System-on-chip and Internet-of-Things applications present very different tradeoffs for interposers than compute-based systems. A network-on-interposer methodology is proposed which balances cost and system performance. Active interposers are shown to exhibit less than half the latency of passive interposers when interfacing between 7 nm and 45 nm chiplets. Design considerations for both passive and active interposer-based systems are explored in this paper, such as maximum bandwidth, chiplet placement, voltage conversion, and network area.
Andres Ayes, Eby G. Friedman, Marilyn Wolf
ISCAS3
2024 Adaptive Perception Control for Aerial Robots with Twin Delayed DDPG
abstract
Robotic perception is commonly assisted by convo-lutional neural networks. However, these networks are static in nature and do not adjust to changes in the environment. Additionally, these are computationally complex and impose latency in inference. We propose an adaptive perception system that changes in response to the robot's requirements. The perception controller has been designed using a recently proposed reinforcement learning technique called Twin Delayed DDPG (TD3). Our proposed method outperformed the baseline approaches.
Veera Venkata Ram Murali Krishna Rao Muvva, Kunjan Theodore Joseph, Kruttidipta Samal, Marilyn Wolf, Santosh Pitla
DATE4
2024 Corporate Governance and Management of AI-Driven Product Development: Vehicle Automation
abstract
This essay explores the interplay between proper corporate governance and engineering expertise in developing products that use artificial intelligence (with a focus on vehicle automation) and considers how an organization might balance maximizing earnings with the public interest in safe and secure AI-driven products. The essay recommends that for oversight directors use a management model with structures of communication and reporting which provides for more direct engagement with engineers and product managers.
William H. Widen, Marilyn Wolf
DATE2
2022 Attacks on Image Sensors
abstract
This paper provides a taxonomy of security vulnerabilities of smart image sensor systems. Image sensors form an important class of sensors. Many image sensors include computation units that can provide traditional algorithms such as image or video compression along with machine learning tasks such as classification. Some attacks rely on the physics and optics of imaging. Other attacks take advantage of the complex logic and software required to perform imaging systems.
Marilyn Wolf, Kruttidipta Samal
ICCAD1
2022 A Methodology for Understanding the Origins of False Negatives in DNN Based Object Detectors
abstract
In this paper we present two novel complimentary methods namely the gradient analysis and the activation discrepancy analysis to analyze the perception failures occurring inside the DNN based object detectors. The gradient analysis localizes the nodes within the network that fail consistently in a scenario, thus creating a ‘signature’ of False Negatives (FNs). This method traces a set of False Negatives through the network and finds sections of the network that contribute to this set. The signatures show the location of the faulty nodes is sensitive to input conditions (such as darkness, glare etc.), network architecture, training hyperparameters, object class etc. Certain nodes of the network fail consistently throughout the training process thus implying that some False Negatives occur due to the global optimization nature of Stochastic Gradient Descent (SGD) based training. This analysis requires the knowledge of False Negatives and therefore can be used for post-hoc diagnostic analysis. On the other hand, the activation discrepancy analysis analyzes the discrepancy in forward activations of a DNN. This method can be conducted online and shows that the pattern of the activation discrepancy is sensitive to input conditions and detection recall.
Kruttidipta Samal, Hemant Kumawat, Marilyn Wolf, Saibal Mukhopadhyay
IJCNN3
2022 The 5th Artificial Intelligence of Things (AIoT) Workshop
abstract
With advancement of recent network and chip technologies, IoT devices are becoming smarter with increasing compute power, bandwidth, and storage available on the device. This enables intelligent decision making and information transferring on the devices and unleashes the power of AIoT (Artificial Intelligence of Things) that supports applications such as smart city/agriculture/manufacturing/health care and self-driving scenarios.
Jian Tang 0008, Yiran Chen 0001, Jie Liu 0001, Jieping Ye, Marilyn Wolf, Narayanan Vijaykrishnan, Mani Srivastava 0001, Michael I. Jordan, Paramvir Bahl
KDD6
2022 MLCAD: A Survey of Research in Machine Learning for CAD Keynote Paper
abstract
Due to the increasing size of integrated circuits (ICs), their design and optimization phases (i.e., computer-aided design, CAD) grow increasingly complex. At design time, a large design space needs to be explored to find an implementation that fulfills all specifications and then optimizes metrics like energy, area, delay, reliability, etc. At run time, a large configuration space needs to be searched to find the best set of parameters (e.g., voltage/frequency) to further optimize the system. Both spaces are infeasible for exhaustive search typically leading to heuristic optimization algorithms that find some tradeoff between design quality and computational overhead. Machine learning (ML) can build powerful models that have successfully been employed in related domains. In this survey, we categorize how ML may be used and is used for design-time and run-time optimization and exploration strategies of ICs. A metastudy of published techniques unveils areas in CAD that are well explored and underexplored with ML, as well as trends in the employed ML algorithms. We present a comprehensive categorization and summary of the state of the art on ML for CAD. Finally, we summarize the remaining challenges and promising open research directions.
Martin Rapp, Hussam Amrouch, Yibo Lin, Bei Yu 0001, David Z. Pan, Marilyn Wolf, Jörg Henkel
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.6
2021 Introspective Closed-Loop Perception for Energy-efficient Sensors
abstract
Task-driven closed-loop perception-sensing systems have shown considerable energy savings over traditional open-loop systems. Prior works on such systems have used simple feedback signals such as object detections and tracking which led to poor perception quality. This paper proposes an improved approach based on perceptual risk. First, a method is proposed to estimate the risk of failure to detect a target of interest. The risk estimate is used as a signal in a feedback system to determine how sensor resources are utilized. Two feedback algorithms are proposed: one based on proportional/integral methods and the other based on 0/1 (bang-bang) methods. These feedback algorithms are compared based on the efficiency with which they use available sensor resources as well as their absolute detection rates. Experiments on two real-world autonomous driving datasets show that the proposed system has better object detection recall and lower marginal cost of prediction than prior work.
Kruttidipta Samal, Marilyn Wolf, Saibal Mukhopadhyay
AVSS2
2021 Closed-loop Approach to Perception in Autonomous System
abstract
Currently, functional tasks within Autonomous Systems are balkanized into several sub-systems such as object detection, tracking, motion planning, multi-sensor fusion etc. which are developed and tested in isolation. In recent times, deep learning is used in the perception systems for improved accuracy, but such algorithms are not adaptive to the transient real-world requirements of an Autonomous System such as latency and energy. These limitations are critical for resource constrained systems such as autonomous drones. Therefore, a holistic closed-loop system design is required for building reliable and efficient perception systems for autonomous drones. The closed-loop perception system creates a focus-of-attention based feedback from end-task such as motion planning to control computation within the deep neural networks (DNNs) used in early perception tasks such as object detection. We observe that this closed-loop perception system improves resource utilization of resource hungry DNNs within perception system with minimal impact on motion planning.
Kruttidipta Samal, Marilyn Wolf, Saibal Mukhopadhyay
DATE2
2021 The 4th Artificial Intelligence of Things (AIoT) Workshop
abstract
With advancement of recent network and chip technologies, IoT devices are becoming smarter with increasing compute power, bandwidth, and storage available on the device. This enables intelligent decision making and information transferring on the devices and unleashes the power of AIoT (Artificial Intelligence of Things) that supports scenarios such as smart city/agriculture/manufacturing/health care and self-driving scenarios. The AIoT Workshop is a forum for researchers, scientists, engineers, and practitioners to share and learn AI powered IoT solutions. The AIoT is a multi-disciplinary area, which include but not limited to IoT, AI/ML, embedded systems, and networking. The 4th AIoT workshop will be hosted virtually in conjunction with the 27th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining (KDD 2021). The workshop program consists of keynote(s), invited talks, accepted technical paper presentations, as well as an indoor location competition panel.
Jian Tang 0008, Yiran Chen 0001, Jie Liu 0001, Jieping Ye, Marilyn Wolf, Narayanan Vijaykrishnan, Mani Srivastava 0001, Michael I. Jordan, Paramvir Bahl
KDD6
2020 Neural Network-Based Side Channel Attacks and Countermeasures
abstract
This paper surveys results in the use of neural networks and deep learning in two areas of hardware security: power attacks and physically-unclonable functions (PUFs).
Dimitrios Serpanos, Shengqi Yang, Marilyn Wolf
DAC3
2020 Hybridization of Data and Model based Object Detection for Tracking in Flash Lidars
abstract
In recent times deep neural networks have become very successful in solving traditionally hard problems in Computer Vision such as Object Detection. This is due to their ability to find hidden patterns in high dimensional data such as images. But if there is a known structure within data that can be accurately represented by a pre-defined model, then by merging this model based algorithm and deep neural network, overall system accuracy can be increased. We apply this idea for solving the task of flash lidar object detection and tracking. Flash lidar is an emerging lidar sensing technology which is getting a lot of attention lately due to their lack of moving parts compared to prevalent scanning lidars. Samples from flash lidar suffer from both spatial and temporal noise which coupled with low angular resolution and FoV lead to low accuracy in object detection. In this paper we present a data driven deep learning based flash lidar object detector and tracker. To our knowledge, this is the first work to use deep learning for flash lidar object detection/tracking. Our tracker has two detectors- 1. supervised object detector and 2. unsupervised class agnostic foreground/moving object detector which are merged to achieve multi-object tracking accuracy of 47.9% on CAMEL dataset.
Kruttidipta Samal, Marilyn Wolf, Saibal Mukhopadhyay
IJCNN2
2020 Message from the IPSN 2020 Organizers
abstract
Welcome to the 19th ACM/IEEE International Conference on Information Processing in Sensor Networks (IPSN 2020), a premiere event on embedded sensing and networked systems that brings together researchers from academia, industry, and government. We are proud to see the continuing trend of increased interest in IPSN in recent years. A substantial 30% year-to-year increase in the number of submissions allowed us to share with you a particularly strong program this year. However, it is with mixed feelings that we write this message. We have all seen the world come to an abrupt stop in the last few weeks due to the spread of COVID-19 virus and will sadly not see you in the beautiful Darling Harbour in Sydney. Instead, the conference will happen in an entirely virtual format that will be split equally across three major world regions. This will inevitably harm the traditional cross-collaborative spirit of the Cyber- Physical Systems week, an event that brings together researchers across diverse fields, including Embedded, Hybrid, and Real-Time Systems. The inability to travel and exchange ideas face-to-face is a challenge, but also a wonderful opportunity to reach a wider global audience and we are determined to capitalize on the advantages and make the virtual IPSN a success. Testimony to the high quality of our program is the incredible work of our authors and organizers. Following the tradition in IPSN, our technical program committee brought together 27 distinguished experts covering the wide breadth of IPSN topics. The committee has collectively reviewed 124 submissions, providing a minimum of 3 high-quality reviews to each author. The top 50 of these submissions received additional 2 reviews and were discussed in person at the program committee meeting in St. Louis. We accepted 27 papers and followed a robust shepherding process to address the reviews prior to publication. We are proud of the authors and would like to thank them for striving to achieve the highest quality. We want to thank the program committee for their many hours of reviewing and discussion that made the program come together.
Branislav Kusy, Neal Patwari, Marilyn Wolf
IPSN3
2020 Efficient Model Solving for Markov Decision Processes
abstract
Markov decision processes provide powerful tools for adaptive management of computing and communication in cyber-physical systems. However, efficient solvers are required to provide these capabilities on embedded computing platforms. This paper describes two new MDP solvers for embedded applications: Sparse Value Iteration (SVI) uses sparse matrix methods and runs on small, single-threaded CPU platforms; Sparse Parallel Value Iteration (SPVI) extends this approach to leverage the parallelism of embedded graphics processing units (GPUs) to further improve performance on more sophisticated embedded platforms. Both solvers improve running time and reduce power consumption.
Adrian E. Sapio, Shuvra S. Bhattacharyya, Marilyn Wolf
ISCC3
2020 Achieving Resiliency and Behavior Assurance in Autonomous Navigation: An Industry Perspective
abstract
In this article, we present an industry perspective on key drivers for autonomous navigation, with a particular focus on resiliency and behavior assurance. We provide a brief survey of current deployed mobile autonomous systems and their capabilities (with a primary focus on the air domain but including other domains-underwater, ground, space, and surface-as well). We discuss techniques that are currently used for achieving resiliency and assurance in autonomous navigation, pointing out some of the shortcomings of these techniques. We describe techniques under development in the industry that aims to overcome these shortcomings by combining emerging approaches to resilient behavior with assured autonomous behavior constructs to yield reliable and mission-effective systems necessary to operate successfully in dynamic and adversarial environments. We briefly discuss ongoing efforts to develop multidomain standards that are designed to be applicable across these disparate vehicle domains.
Sanjoy Baruah, Prakash Sarathy, Marilyn Wolf
Proc. IEEE4
2020 Runtime Adaptation in Wireless Sensor Nodes Using Structured Learning
abstract
Markov Decision Processes (MDPs) provide important capabilities for facilitating the dynamic adaptation and self-optimization of cyber physical systems at runtime. In recent years, this has primarily taken the form of Reinforcement Learning (RL) techniques that eliminate some MDP components for the purpose of reducing computational requirements. In this work, we show that recent advancements in Compact MDP Models (CMMs) provide sufficient cause to question this trend when designing wireless sensor network nodes. In this work, a novel CMM-based approach to designing self-aware wireless sensor nodes is presented and compared to Q-Learning, a popular RL technique. We show that a certain class of CPS nodes is not well served by RL methods and contrast RL versus CMM methods in this context. Through both simulation and a prototype implementation, we demonstrate that CMM methods can provide significantly better runtime adaptation performance relative to Q-Learning, with comparable resource requirements.
Adrian E. Sapio, Shuvra S. Bhattacharyya, Marilyn Wolf
ACM Trans. Cyber Phys. Syst.3
2020 Introduction to the Special Issue on Machine Learning for CAD
abstract
No abstract available.
Jörg Henkel, Hussam Amrouch, Marilyn Wolf
ACM Trans. Design Autom. Electr. Syst.3
2019 A Camera with Brain - Embedding Machine Learning in 3D Sensors
abstract
The cameras today are designed to capture signals with highest possible accuracy to most faithfully represent what it sees. However, many mission-critical autonomous applications ranging from traffic monitoring to disaster recovery to defense requires quality of information, where useful information depends on the tasks and is defined using complex features, rather than only changes in captured signal. Such applications require cameras that capture useful information from a scene with highest quality while meeting system constraints such as power, performance, and bandwidth. This paper will discuss the feasibility of a camera that learns how to capture task-dependent information with highest quality, paving the pathway to design a camera with brain. 3D integration of digital pixel sensors with massively parallel computing platform for machine learning creates a hardware architecture for such a camera. The paper will discuss embedded machine learning algorithms that can run on such platform to enhance quality of useful information by real-time control of the sensor parameters. We conclude by identifying critical challenges as well as opportunities for hardware and algorithmic innovations to enable machine learning in the feedback loop of a 3D image sensor based camera.
Burhan Ahmad Mudassar, Priyabrata Saha, Mohammad Faisal Amir, Evan Gebhardt, Taesik Na, Jong Hwan Ko, Marilyn Wolf, Saibal Mukhopadhyay
DATE8
2019 Thoughts on Edge Intelligence
abstract
Machine learning methods have exploded in the past half-dozen years. Machine learning is being applied to a huge range of problems across the spectrum of applications. Initial results relied on server-oriented computations. But many applications will require deploying aspects of machine learning throughout the network hierarchy. Several factors motivate the development of Edge Intelligence architectures and algorithms: network bandwidth, power consumption, latency, privacy, etc. This talk will start the motivation for edge intelligence with several examples from manufacturing and health care and outline some important problems for VLSI systems.
Marilyn Wolf
ACM Great Lakes Symposium on VLSI1
2019 Machine Learning + Distributed IoT = Edge Intelligence
abstract
Internet-of-Things (IoT) systems provide large-scale sensor networks that can be used to monitor and analyze a wide range of physical systems. IoT systems can generate huge volumes of data that needs to be dealt with in a timely fashion. Machine learning (ML) provides compelling techniques for the analysis of large, complex data sets. Many existing machine learning systems operate in the cloud. However, bandwidth, power, latency, privacy, and other issues often require IoT systems to apply ML techniques at multiple levels of the network hierarchy. This paper describes important characteristics of edge intelligence and identifies several important research challenges related to machine learning + distributed IoT systems.
Marilyn Wolf
ICDCS1
2018 CAMEL Dataset for Visual and Thermal Infrared Multiple Object Detection and Tracking
abstract
We present a visual-infrared video sequence dataset for object detection and tracking, called the CAMEL dataset1. The dataset consists of 26 video sequences captured in the visible and thermal infrared domains. The sequences include multiple real world urban environments, as well as multiple targets. The goal is to provide a challenging benchmark similar to MOT challenge that includes sequences that have corresponding visual and infrared pairs. Our hope is that this dataset can be used to help improve work on visible-infrared fusion techniques, as well as object detection and tracking.
Evan Gebhardt, Marilyn Wolf
AVSS2
2018 The CAMEL approach to stacked sensor smart cameras
abstract
Stacked image sensor systems combine an image sensor, memory, and processors using 3D technology. Stacking camera components that have traditionally been packaged separately provides several benefits: very high bandwidth out of the image sensor, allowing for higher frame rates; very low latency, providing opportunities for image processing and computer vision algorithms which can adapt at very high rates; and lower power consumption. This paper will review the characteristics of stacked image sensor systems and discuss novel algorithmic and systems concepts that are made possible by these stacked sensors.
Saibal Mukhopadhyay, Marilyn Wolf, Mohammed Faisal Amir, Evan Gebhardt, Jong Hwan Ko, Jaeha Kung 0001, Burhan Ahmad Mudassar
DATE2
2018 Improving the Safety and Security of Wide-Area Cyber-Physical Systems Through a Resource-Aware, Service-Oriented Development Methodology
abstract
This paper presents a service-oriented development methodology for wide-area cyber-physical systems (CPS) such as smart grid and vehicular networks. Unlike the traditional task-based development approach from the domains of automotive and avionics, the proposed service-oriented development methodology inherently enables disruption-free incremental system deployment and reconfiguration that are fundamental requirements for handling the “always-online” nature of emerging wide-area CPS application domains such as smart grid and vehicular networks. The proposed service-oriented CPS development methodology extends the traditional service-oriented computing (SOC) paradigm for handling hard real-time CPS aspects by introducing resource-aware service deployment and quality-of-service (QoS)-aware service operation phases. The proposed CPS development methodology also supports a streamlined formal interface between the traditional computer-aided feedback controller design environments and SOC paradigm. The paper utilizes a simulation-based smart grid case study to illustrate the advantages of the proposed methodology for developing wide-area cyber-physical systems with improved safety and security characteristics. The paper also identifies a set of technological requirements for the proposed service-oriented CPS development methodology that should guide future research in this area.
Muhammad Umer Tariq, Jacques Florence, Marilyn Wolf
Proc. IEEE3
2018 Scanning The Issue
abstract
This special issue is devoted to the safety and security issues presented by cyber–physical systems (CPSs). CPSs use cyber software/hardware to perform real-time control on physical systems. Such systems are widely used in aerospace and automotive, medical, industrial, and critical infrastructure applications.
Marilyn Wolf, Dimitrios Serpanos
Proc. IEEE1
2018 Safety and Security in Cyber-Physical Systems and Internet-of-Things Systems
abstract
Safety and security have traditionally been distinct problems in engineering and computer science. The introduction of computing elements to create cyber-physical systems (CPSs) has opened up a vast new range of potential problems that do not always show up on the radar of traditional engineers. Security, in contrast, is traditionally viewed as a data or communications security problem to be handled by computer scientists and/or computer engineers. Advances in CPSs and the Internet-of-Things (IoT) requires us to take a unified view of safety and security. This paper defines a safety/security threat model for CPSs and IoT systems and surveys emerging techniques which improve the safety and security of CPSs and IoT systems.
Marilyn Wolf, Dimitrios Serpanos
Proc. IEEE1
2017 Design and implementation of adaptive signal processing systems using Markov decision processes
abstract
In this paper, we propose a novel framework, called Hierarchical MDP framework for Compact System-level Modeling (HMCSM), for design and implementation of adaptive embedded signal processing systems. The HMCSM framework applies Markov decision processes (MDPs) to enable autonomous adaptation of embedded signal processing under multidimensional constraints and optimization objectives. The framework integrates automated, MDP-based generation of optimal reconfiguration policies, dataflow-based application modeling, and implementation of embedded control software that carries out the generated reconfiguration policies. HMCSM systematically decomposes a complex, monolithic MDP into a set of separate MDPs that are connected hierarchically, and that operate more efficiently through such a modularized structure. We demonstrate the effectiveness of our new MDP-based system design framework through experiments with an adaptive wireless communications receiver.
Lin Li 0029, Adrian E. Sapio, Jiahao Wu 0001, Yanzhou Liu 0001, Kyunghun Lee, Marilyn Wolf, Shuvra S. Bhattacharyya
ASAP6
2017 Safety and Security of Cyber-Physical and Internet of Things Systems [Point of View]
abstract
Computer system security and engineering system safety have traditionally been very distinct topics pursued by people with very different expertise. The advent of cyber-physical systems and the Internet-of-Things (IoT) changes that dynamic. Safety and security are now inextricably linked through our linkage of computer hardware and software with complex physical plants. Computers have been added to traditional engineering systems to achieve goals that we cannot achieve using traditional mechanical control. The automobile provides an important early example of the benefits of cyber-physical systems: computer engine control allowed manufacturers to simultaneously meet stiff requirements on both fuel economy and emissions; features such as antilock brakes and traction control improved vehicle handling and safety; and a new generation of supercars use software to not only provide sophisticated vehicle capabilities but also to change the vehicle's handling characteristics at the push of a button.
Marilyn Wolf, Dimitrios Serpanos
Proc. IEEE1
2017 Guest Editorial: Special Issue on Embedded Computing for IoT
abstract
No abstract available.
Marilyn Wolf, Chun Jason Xue
ACM Trans. Embed. Comput. Syst.1
2015 What don't we know about CPS architectures?
abstract
This paper considers the challenges in the architectural design of cyber-physical systems (CPS). Cyber-physical systems are real-time control and coordination systems that rely on computational infrastructure. We help to elucidate these challenges by comparing cyber-physical system design to system-on-chip design. We then survey CPS architectures and identify several important challenges.
Marilyn Wolf, Eric Feron
DAC1
2013 Physics of computing as an introduction to computer engineering
abstract
This paper describes a new required course in the Georgia Tech computer engineering curriculum, ECE 3030, Physical Foundations of Computer Systems. Traditional introductory courses take a constructive approach to logic design and computer organization. 3030, in contrast, introduces the major physical concepts underlying computation. It shows how they determine basic properties of computers such as speed and energy consumption. It also explores design trade-offs by showing how changes that improve one type of property inevitably, due to physics, cause another useful property to degrade. The course emphasizes CMOS but many of its principles apply to other logic technologies as well. Students do not directly design logic or learn assembly language-for example, delay and energy consumption are studied for inverter chains. However, they have time in the course to study in detail the basic physical phenomena that underlie design choices in digital systems. Those principles help students absorb material in later classes such as VLSI design. 3030 introduces certain topics to students much earlier in the curriculum than is traditional. We believe that an early introduction to principles is important not just for students who become logic designers but for all computer engineers.
Marilyn Wolf, Saibal Mukhopadhyay
FIE1
2013 High-performance and low-energy buffer mapping method for multiprocessor DSP systems
abstract
When implementing digital signal processing (DSP) applications onto multiprocessor systems, one significant problem in the viewpoints of performance is the memory wall. In this paper, to help alleviate the memory wall problem, we propose a novel, high-performance buffer mapping policy for SDF-represented DSP applications on bus-based multiprocessor systems that support the shared-memory programming model. The proposed policy exploits the bank concurrency of the DRAM main memory system according to the analysis of hierarchical parallelism. Energy consumption is also a critical parameter, especially in battery-based embedded computing systems. In this paper, we apply a synchronization back-off scheme on the top of the proposed high-performance buffer mapping policy to reduce energy consumption. The energy saving is attained by minimizing the number of non-essential synchronization transactions. We measure throughput and energy consumption on both synthetic and real benchmarks. The simulation results show that the proposed buffer mapping policy is very useful in terms of performance, especially in memory-intensive applications where the total execution time of computational tasks is relatively small compared to that of memory operations. In addition, the proposed synchronization back-off scheme provides a reduction in the number of synchronization transactions without degrading performance, which results in system energy saving.
Dongwon Lee 0003, Marilyn Wolf, Shuvra S. Bhattacharyya
ACM Trans. Embed. Comput. Syst.2
2012 Power Analysis Attack Resistance Engineering by Dynamic Voltage and Frequency Scaling
abstract
This article proposes a novel approach to cryptosystem design to prevent power analysis attacks. Such attacks infer program behavior by continuously monitoring the power supply current going into the processor core. They form an important class of security attacks. Our approach is based on dynamic voltage and frequency scaling (DVFS), which hides processor state to make it harder for an attacker to gain access to a secure system. Three designs are studied to test the efficacy of the DVFS method against power analysis attacks. The advanced realization of our cryptosystem is presented which achieves enough high power and time trace entropies to block various kinds of power analysis attacks in the DES algorithm. We observed 27% energy reduction and 16% time overhead in these algorithms. Finally, DVFS hardness analysis is presented.
Shengqi Yang, Pallav Gupta, Marilyn Wolf, Dimitrios Serpanos, Narayanan Vijaykrishnan, Yuan Xie 0001
ACM Trans. Embed. Comput. Syst.3
2011 Modeling and Analysis of Image Dependence and Its Implications for Energy Savings in Error Tolerant Image Processing
abstract
We present an analysis of the relationship between input images and energy consumption in error tolerant image processing. Under aggressive voltage scaling, the output image quality of image processing depends on input images for two reasons: 1) error tolerance among images is naturally disparate in terms of perceptual image quality assessment, and 2) the error rate under aggressive voltage scaling varies by input image types. Based on both effects, the supply voltage can be optimized for a given quality requirement so as to achieve ultralow power/energy dissipation. Our analysis demonstrates the significance of the accurate delay estimation, which depends on not only combinational inputs but also the previous state of the logic. We present a new sequential model for accurate error estimation. Based on the model, our experimental results demonstrate that different input image types lead to very different output quality. We also present the effect of process variation on the relationship between input image and output quality. The dependence of energy consumption on input images provides a new perspective for low-power multimedia and image processing system design.
Se Hun Kim, Saibal Mukhopadhyay, Marilyn Wolf
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2011 Reconfigurable SRAM Architecture With Spatial Voltage Scaling for Low Power Mobile Multimedia Applications
abstract
This paper presents a dynamically reconfigurable SRAM array for low-power mobile multimedia application. The proposed structure use a lower voltage for cells storing low-order bits and a nominal voltage for cells storing higher order bits. The architecture allows reconfigure the number of bits in the low-voltage mode to change the error characteristics of the array in run-time. Simulations in predictive 70 nm nodes show that the proposed array can obtain 45% savings in memory power with a marginal (~10%) reduction in image quality.
Minki Cho, Jason Schlessman, Marilyn Wolf, Saibal Mukhopadhyay
IEEE Trans. Very Large Scale Integr. Syst.3
2010 Design space exploration of the turbo decoding algorithm on GPUs
abstract
In this paper, we explore the design space of the Turbo decoding algorithm on GPUs and find a performance bottleneck. We consider three axes for the design space exploration: a radix degree, a parallelization method, and the number of sub-frames per thread block. In Turbo decoding, a degree of radix affects computational complexity and memory access patterns in both algorithmic and implementation viewpoints. Second, computations of branch metrics (BMs) and state metrics (SMs) have a different degree of parallelism, which affects the mapping method of computational tasks to GPU threads. Finally, we can easily adjust the number of sub-frames per thread block to balance the occupancy and memory access traffic. Experimental results show that the radix-4 algorithm with the SM-centric mapping method shows the best performance at four sub-frames per thread block. According to our analysis, two factors -- the occupancy and shared memory bank conflicts -- differentiate the performance of different cases in the design space. We show further performance improvements by optimizing a kernel operation (max*) and applying the MAX-Log-Maximum A Posteriori (MAP) algorithm. A performance bottleneck at the finally optimized case is global memory access latency.
Dongwon Lee 0003, Marilyn Wolf, Hyesoon Kim
CASES2
2010 Detecting Moving Objects Using a Camera on a Moving Platform
abstract
This paper proposes a new ego-motion estimation and background/foreground classification method to effectively segment moving objects from videos captured by a moving camera on a moving platform. Existing methods for moving-camera detecting impose serious constraints. In our approach, ellipsoid scene shape is applied in the motion model and a complicated ego-motion estimation formula is derived. Genetic algorithm is introduced to accurately solve ego-motion parameters. After motion recovery, noisy result is refined by motion vector correlation and foreground is classified by pixel level probability model. Experiment results show that the method demonstrates significant detecting performance without further restrictions and performs effectively in complex detecting environment.
Chung-Ching Lin, Marilyn Wolf
ICPR2
2010 Hardware/Software Codesign of Aerospace and Automotive Systems
abstract
Electronics systems for modern vehicles must be designed to meet stringent requirements on real-time performance, safety, power consumption, and security. Hardware/software codesign techniques allow system designers to create platforms that can both meet those requirements and evolve as components and system requirements evolve. Design methodologies have evolved that allow systems-of-systems to be built from subsystems that are themselves embedded computing systems. Software performance is a key metric in the design of these systems. A number of methods-of-methods for the analysis of worst case execution time have been developed. More recently, we have developed new methods for software performance analysis based on design of experiments. Formal methods can be used to verify system properties. Systems must be architected to maintain their integrity in the face of attacks from the Internet. All of these techniques build upon generic hardware/software codesign techniques but with significant adaptations to the technical and economic context of vehicle design.
Ahmed Abdallah, Eric Feron, Graham R. Hellestrand, Philip Koopman, Marilyn Wolf
Proc. IEEE5
2010 System and software architectures of distributed smart cameras
abstract
In this article, we describe a distributed, peer-to-peer gesture recognition system along with a software architecture modeling technique and authority control protocol for ubiquitous cameras. This system performs gesture recognition in real time by combining imagery from multiple cameras without using a central server. We propose a system architecture that uses a network of inexpensive cameras to perform in-network video processing. A methodology for transforming well-designed single-node algorithm to distributed system is also proposed. Applications for ubiquitous cameras can be modeled as the composition of a finite-state machine of the system, functional services, and middleware. A service-oriented software architecture is proposed to dynamically reconfigure services when system state changes. By exchanging data and control messages between neighboring sensors, each node can maintain broader view of the environment with integrated video-processing results. Our prototype system is built on Windows machines, and uses standard video cameras as sensors and local network as a communication channel.
Chang Hong Lin, Marilyn Wolf, Xenofon Koutsoukos, Sandeep Neema, Janos Sztipanovits
ACM Trans. Embed. Comput. Syst.2
2010 Special Section on Distributed Camera Networks: Sensing, Processing, Communication, and Implementation
abstract
The eight papers in this special section span across theoretical and practical considerations of various aspects of distributed camera networks, including adaptive sensing, distributed processing, efficient communications, and versatile implementations.
Rama Chellappa, Wendi B. Heinzelman, Janusz Konrad, Dan Schonfeld, Marilyn Wolf
IEEE Trans. Image Process.5
2009 Accuracy-aware SRAM: a reconfigurable low power SRAM architecture for mobile multimedia applications
abstract
We propose a dynamically reconfigurable SRAM architecture for low-power mobile multimedia applications. Parametric failures due to manufacturing variations limit the opportunities for power saving in SRAM. We show that, using a lower voltage for cells storing low-order bits and a nominal voltage for cells storing higher order bits, ~45% savings in memory power can be achieved with a marginal (~10%) reduction in image quality. A reconfigurable array structure is developed to dynamically reconfigure the number of bits in different voltage domains.
Minki Cho, Jason Schlessman, Marilyn Wolf, Saibal Mukhopadhyay
ASP-DAC3
2009 Experimental analysis of sequence dependence on energy saving for error tolerant image processing
abstract
We present experimental analysis to exploit the sequence dependence on energy saving in error tolerant image processing. Our analysis shows that the error distributions depend not only on combinational inputs but also on the previous state of the logic. We present a new sequential model for low-power delay faults. Our experimental results demonstrate the importance of considering the state of logic when analyzing errors and its dependence on output quality and energy saving. The results show that different input image type leads to very different output quality. This means that more error tolerant image can be operated at lower voltage. The sequence dependence in output quality and energy saving provide a new perspective to design low power multimedia system.
Se Hun Kim, Saibal Mukhopadhyay, Marilyn Wolf
ISLPED3
2008 Multicore design is the challenge! what is the solution?
abstract
Multi Processor SoC (MPSoC) are being designed today. MPSoC design can help achieve aggressive performance and low power targets but it creates new design challenges: How to design the interconnect fabric and memory sub-system to allow the massive data movement required in a multi processor SoC environment? How to develop, debug and verify HW and SW functionality in a MPSoC design? Is MPSoC design an inflection point that will require new design methods including ESL methodologies?
Eshel Haritan, Toshihiro Hattori, Hiroyuki Yagi, Pierre G. Paulin, Marilyn Wolf, Achim Nohl, Drew Wingard, Mike Muller
DAC5
2008 An Optimized Message Passing Framework for Parallel Implementation of Signal Processing Applications
abstract
Novel reconfigurable computing platforms enable efficient realizations of complex signal processing applications by allowing exploitation of parallelization resulting in high throughput in a cost-efficient way. However, the design of such systems poses various challenges due to the complexities posed by the applications themselves as well as the heterogeneous nature of the targeted platforms. One of the most significant challenges is communication between the various computing elements for parallel implementation. In this paper, we present a communication interface, called the signal passing interface (SPI), that attempts to overcome this challenge by integrating relevant properties of two different yet important paradigms in this context - dataflow and the message passing interface (MPI). SPI is targeted towards signal processing applications and, due to its careful specialization, more performance-efficient for their embedded implementation. It is also more easier and intuitive to use. Earlier, a preliminary version of SPI was presented [12] which was restricted to static dataflow behavior. Here, we present a more complete version of SPI with new features to address both static and dynamic dataflow behavior, and to provide new optimization techniques. We develop a hardware description language (HDL) realization of the SPI library, and demonstrate its functionality on the Xilinx Virtex-4 FPGA. Details of the HDL-based SPI library along with experiments with two signal processing applications on the FPGA are also presented.
Sankalita Saha, Jason Schlessman, Sebastian Puthenpurayil, Shuvra S. Bhattacharyya, Marilyn Wolf
DATE5
2008 Using Empirical Science to Engineer Systems: Optimizing Cache for Power and Performance
abstract
The design process of modern embedded systems invariably places a large emphasis on power demands and system performance. Engineers seeking to optimize will inevitably look to adjust the microprocessor memory hierarchy. To provide adequate coverage for an extensive set of applications designers need to investigate as many cache parameter settings as possible. In this paper we describe the development of a methodology able to explore a vast design space. This methodology relies on the statistically based field of Design of Experiments (DOE) to efficiently navigate through these endless possibilities, and take on the chore of multiple objective optimization. We then also detail a tactic to determine an optimal set of configurations which will accommodate multiple applications on the same platform simultaneously. The intention here is not just to solve the problem of cache tuning, but to establish some of the structure necessary for the long overdue integration of the use of empirically based techniques, alongside other well-established methods, in the design and testing of systems.
Ahmed Abdallah, Marilyn Wolf, Graham R. Hellestrand
DSD2
2008 GLSVLSI 2008 invited/keynote talk
abstract
Distributed smart camera systems are physically distributed systems that perform real-time embedded computer vision. These systems perform large amounts of real-time computing at high rates, which has implications for the architectures of the nodes. In this talk, we will discuss the VLSI support that is required to make distributed smart cameras possible.
Marilyn Wolf
ACM Great Lakes Symposium on VLSI1
2008 Evaluation of functional architectures for cognitive radio systems
abstract
This paper proposes and compares two new physical layer architectures for cognitive radio. The idea of cognitive radio has been proposed for a while, but little research has been done on detailed architectural design and complexity analysis of physical layer functions and systems. Two cognitive radio functional architectures are proposed: one is based on the traditional communication components and the other is based on blind signal separation. Complexities of critical cognitive radio functions are evaluated and the two proposed architectures are compared.
Chia-han Lee, Marilyn Wolf
ICASSP2
2008 Middleware Architectures for Distributed Embedded Systems
abstract
A wide range of high-performance distributed embedded systems have been designed and deployed. Physically distributed embedded systems are used for manufacturing and control, traffic analysis, and other problems. Interestingly, today's systems-on-chips are sufficiently complex that they must be treated as distributed embedded systems. At all scales of physical extent, middleware is required to manage the computations. This paper looks at distributed embedded systems at several physical scales and considers the types of middleware that are needed to operate these systems.
Marilyn Wolf
ISORC1
2008 Task Scheduling for Control Oriented Requirements for Cyber-Physical Systems
abstract
The wide applications of cyber-physical systems (CPS) call for effective design strategies that optimize the performance of both computing units and physical plants.We study the task scheduling problem for a class of CPS whose behaviors are regulated by feedback control laws. We co-design the control law and the task scheduling algorithm for predictable performance and power consumption for both the computing and the physical systems. We use a typical example, multiple inverted pendulums controlled by one processor, to illustrate our method.
Fumin Zhang 0001, Klementyna Szwaykowska, Marilyn Wolf, Vincent John Mooney III
RTSS3
2008 Frame-level temporal calibration of video sequences from unsynchronized cameras
Senem Velipasalar, Marilyn Wolf
Mach. Vis. Appl.2
2008 Multiprocessor System-on-Chip (MPSoC) Technology
abstract
The multiprocessor system-on-chip (MPSoC) uses multiple CPUs along with other hardware subsystems to implement a system. A wide range of MPSoC architectures have been developed over the past decade. This paper surveys the history of MPSoCs to argue that they represent an important and distinct category of computer architecture. We consider some of the technological trends that have driven the design of MPSoCs. We also survey computer-aided design problems relevant to the design of MPSoCs.
Marilyn Wolf, Ahmed Amine Jerraya, Grant Martin
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2008 Case Study of Reliability-Aware and Low-Power Design
abstract
Based on the proposed reliability characterization model, reliability-aware and low-power design is illustrated for the first time as a design methodology to balance reliability enhancement and power reduction. Low-power and reliable SRAM cell design, reliable dynamic voltage scaling (DVS) algorithm design, and voltage island partitioning and floorplanning for reliable system-on-a-chip (SOC) design are demonstrated as case studies of this new design methodology.
Shengqi Yang, Wenping Wang 0004, Tiehan Lv, Marilyn Wolf, Narayanan Vijaykrishnan, Yuan Xie 0001
IEEE Trans. Very Large Scale Integr. Syst.4
2007 Real-Time Distributed Tracking
abstract
Distributed smart cameras use distributed computing architectures to analyze imagery from physically distributed cameras. Performing real-time distributed analysis of video introduces substantial new challenges, but also provides substantial benefits over server-based approaches. This work describes some of the algorithms and architectures we have developed for tracking using distributed smart camera systems, including fault-tolerance, synchronization, and multi-band fusion.
Marilyn Wolf, Senem Velipasalar, Jason Schlessman, Cheng-Yao Chen, Chang Hong Lin
ICASSP (4)1
2007 Register binding guided by the size of variables
abstract
An important problem is how to carry out register binding such that any register has to be bound to a set of variables such that the difference between their sizes is as small as possible. For the case of hardware implementations, satisfying this latter constraint will allow to reduce the complexity of the clock generation tree, and saving area occupied by the registers which is very important for the case of system-on-chip and some embedded systems. When registers are already built, satisfying this latter constraint would allow reducing power consumption due to useless switching activities that will happen into any register that is bound to variables with different sizes. Assuming that the size in bits of any variable is known, we propose in this paper exact algorithms to optimally solve this problem for the case of acyclic graphs. An extended version of this problem is how to solve it while controlling the number of variables to be assigned to a same register. We also propose exact algorithms to optimally solve this latter version of the problem. Experimental results are provided. We also test the impact of the proposed approach in the case of a hardware implementation using the design analyzer tool from Synopsys Inc.. Obtained results have shown that both area and power consumption have been reduced.
Noureddine Chabini, Marilyn Wolf
ICCD2
2007 Heterogeneous MPSoC Architectures for Embedded Computer Vision
abstract
In this paper, architectures for two distinct embedded computer vision operations are presented. Motivation is given for the utilization of heterogeneous processing cores on a single chip. In addition, a brief discussion of applicability of multi-processor system on a chip (MPSoC) design challenges and techniques to nascent multi-core development considerations is given. Furthermore, a composite architecture consisting of the two distinct operations is discussed, with relative merits of this approach provided. Finally, experimental analysis is given for the applicability and feasibility of these heterogeneous multiprocessor architectures. Area, power, and cycle times are provided for each of the aforementioned designs. The architectural mappings were implemented on a Xilinx Virtex-II Pro V2P30 FPGA, and are shown to operate without pipelining at 50 MHz, utilizing roughly 46% of FPGA resources, and consuming 565 mW of power.
Jason Schlessman, Mark Lodato, I. Burak Özer, Marilyn Wolf
ICME4
2007 VLSI models of network-on-chip interconnect
abstract
We use VLSI circuit models to analyze the relative delay of interconnect subsystems for networks-on-chips (NoCs). Most work in NoCs has selected a network topology based on higher-level performance models, such as packet delay. Our model parameterizes the interconnect subsystem size by N, the number of IP cores (processors, memories, etc.) to be connected. This paper analyzes busses, crossbars, and some multi-stage networks. We compare the delay required transfer a specific amount of information (bits) between two cores. Considering the data transfer parallelism in crossbars, we make 2 different comparisons: (i) transfer between 2 devices, and (ii) parallel transfers between all devices.
Dimitrios Serpanos, Marilyn Wolf
VLSI-SoC2
2007 Code Compression for VLIW Embedded Systems Using a Self-Generating Table
abstract
We propose a new class of methods for VLIW code compression using variable-sized branch blocks with self-generating tables. Code compression traditionally works on fixed-sized blocks with its efficiency limited by their small size. A branch block, a series of instructions between two consecutive possible branch targets, provides larger blocks for code compression. We compare three methods for compressing branch blocks: table-based, Lempel-Ziv-Welch (LZW)-based and selective code compression. Our approaches are fully adaptive and generate the coding table on-the-fly during compression and decompression. When encountering a branch target, the coding table is cleared to ensure correctness. Decompression requires a simple table lookup and updates the coding table when necessary. When decoding sequentially, the table-based method produces 4 bytes per iteration while the LZW-based methods provide 8 bytes peak and 1.82 bytes average decompression bandwidth. Compared to Huffman's 1 byte and variable-to-fixed (V2F)'s 13-bit peak performance, our methods have higher decoding bandwidth and a comparable compression ratio. Parallel decompression could also be applied to our methods, which is more suitable for VLIW architectures.
Chang Hong Lin, Yuan Xie 0001, Marilyn Wolf
IEEE Trans. Very Large Scale Integr. Syst.3
2007 Code Decompression Unit Design for VLIW Embedded Processors
abstract
Code size "bloating" in embedded very long instruction word (VLIW) processors is a major concern for embedded systems since memory is one of the most restricted resources. In this paper, we describe a code compression algorithm based on arithmetic coding, discuss how to design decompression architecture, and illustrate the tradeoffs between compression ratio and decompression overhead, by using different probability models. Experimental results for a VLIW embedded processor TMS320C6x show that compression ratios between 67% and 80% can be achieved, depending on the probability models used. A precache decompression unit design is implemented in TSMC 0.25 mum and a test chip is fabricated.
Yuan Xie 0001, Marilyn Wolf, Haris Lekatsas
IEEE Trans. Very Large Scale Integr. Syst.2
2006 Mapping Multimedia Applications Onto Configurable Hardware With Parameterized Cyclo-Static Dataflow Graphs
abstract
This paper develops methods for model-based design and implementation of image processing applications. We apply our previously developed meta-modeling technique of homogeneous parameterized dataflow (HPDF)[9] to the framework of cyclostatic dataflow (CSDF) [1], and demonstrate this integrated modeling methodology through hardware mapping of a gesture recognition application. We also provide a comparative study between HPDF/CSDF-based representation of the gesture recognition application, and a previously developed version based on applying HPDF in conjunction with conventional synchronous dataflow (SDF) semantics [9].
Fiorella Haim, Mainak Sen, Dong-Ik Ko, Shuvra S. Bhattacharyya, Marilyn Wolf
ICASSP (3)5
2006 Design and Verification of Communication Protocols for Peer-to-Peer Multimedia Systems
abstract
This paper addresses issues pertaining to the necessity of utilizing formal verification methods in the design of protocols for peer-to-peer multimedia systems. These systems require sophisticated communication protocols, and these protocols require verification. We discuss two sample protocols designed for two distinct peer-to-peer computer vision applications, namely multi-object multi-camera tracking and distributed gesture recognition. We present simulation and verification results for these protocols, obtained by using the SPIN verification tool, and discuss the importance of verifying the protocols used in peer-to-peer multimedia systems
Senem Velipasalar, Chang Hong Lin, Jason Schlessman, Marilyn Wolf
ICME4
2006 SCCS: A Scalable Clustered Camera System for Multiple Object Tracking Communicating Via Message Passing Interface
abstract
We introduce the scalable clustered camera system, a peer-to-peer multi-camera system for multi-object tracking, where different CPUs are used to process inputs from distinct cameras. Instead of transferring control of tracking jobs from one camera to another, each camera in our system performs its own tracking and keeps its own tracks for each target object, thus providing fault tolerance. A fast and robust tracking method is proposed to perform tracking on each camera view, while maintaining consistent labeling. In addition, we introduce a new communication protocol, where the decisions about when and with whom to communicate are made such that frequency and size of transmitted messages are minimized. This protocol incorporates variable synchronization capabilities, so as to allow flexibility with accuracy tradeoffs. We discuss our implementation, consisting of a parallel computing cluster, with communication between the cameras performed by MPI. We present experimental results which demonstrate the success of the proposed peer-to-peer multi-camera tracking system, with accuracy of 95% for a high frequency of synchronization, as well as a worst-case of 15 frames of latency in recovering correct labels at low synchronization frequencies
Senem Velipasalar, Jason Schlessman, Cheng-Yao Chen, Marilyn Wolf
ICME4
2006 Low power, low cost, wireless camera sensor nodes For human detection
abstract
Our demonstration consists of sensor nodes suitable for imageintensive network applications. We developed nodes for stationary and mobile deployment, for face recognition and human detection applications, respectively. Both designs consist of a visible spectrum camera sensor and a Zigbee-compliant wireless data transceiver. The stationary model receives input from a PIR motion sensor, while the mobile model incorporates distance measures from either an infrared or ultrasonic sensor and includes a robot for demonstration of mobility. Our node uses strictly commercial off the-shelf components, and not including the robotic accoutrement, is relatively inexpensive. The stationary node will be on display for face recognition of attendees. Additionally, we intend to demonstrate the mobile node for a human detection application, wherein the node detects objects, and transmits them to a host machine for face detection processing.
Jason Schlessman, Jae-Chang Shim, Ik-Dong Kim, Yun Cheol Baek, Marilyn Wolf
SenSys5
2006 Audiovisual Gunshot Event Recognition
abstract
In this paper, we introduce a gunshot event recognition system based on audio and visual feature analysis. We model the gunshot event by a hierarchical probabilistic system. By incorporating gunshot sounds, human emotion and human activity analysis, we developed an effective semantic gunshot scene description from consumer video sequences. Moreover, our system also detects possible threatening scenes and wounded victim scenes which are closely related to real world gunshot scenes of violence. In addition to event modeling, we also employ optimized hierarchical audiovisual models in feature state detection to determine additional details including different types of guns, human emotions and human gesture of weapon discharge. Experimental results indicate the precision of gunshot event video content recognition is encouraging while the rate of false alarms is low. These favorable results arise from effectively capturing not only the event features themselves but also human responses inside the event. The effectiveness and flexibility of our system can benefit applications in the field of content-based video indexing, multimedia surveillance and human-computer interaction.
Cheng-Yao Chen, Ahemd E. Abdallah, Marilyn Wolf
SMC3
2006 An efficient architecture for motion estimation and compensation in the transform domain
abstract
This paper describes a new architecture for discrete cosine transform (DCT)-based motion estimation and compensation. Previous methods do not take sufficient advantage of the sparseness of two-dimensional (2-D) DCT coefficients to reduce execution time. We first derive a recursion equation for transform domain motion estimation; we then use it to develop a wavefront array processor consisting of highly regular, parallel, and pipelined processing elements that more efficiently performs motion estimation. In addition, we show that the recursion equation enables motion predicted images with different frequency bands, for example, from the images with low-frequency components to the images with low- and high-frequency components. The wavefront array processor can reconfigure to different motion estimation algorithms, such as logarithmic search and three step search, without architectural modifications. These properties can be effectively used to reduce the energy required for video encoding and decoding. Simulation results on video sequences of different characteristics show that the proposed architecture achieves a significant reduction in computational complexity and processing time, with comparable performance to spatial domain approaches with respect to the peak signal to noise ratio (PSNR) and the compression ratio.
Jooheung Lee, Narayanan Vijaykrishnan, Mary Jane Irwin, Marilyn Wolf
IEEE Trans. Circuits Syst. Video Technol.4
2006 A design methodology for application-specific networks-on-chip
abstract
With the help of HW/SW codesign, system-on-chip (SoC) can effectively reduce cost, improve reliability, and produce versatile products. The growing complexity of SoC designs makes on-chip communication subsystem design as important as computation subsystem design. While a number of codesign methodologies have been proposed for on-chip computation subsystems, many works are needed for on-chip communication subsystems. This paper proposes application-specific networks-on-chip (ASNoC) and its design methodology. ASNoC is used for two high-performance SoC applications. The methodology (1) can automatically generate optimized ASNoC for different applications, (2) can generate a corresponding distributed shared memory along with an ASNoC, (3) can use both recorded and statistical communication traces for cycle-accurate performance analysis, (4) is based on standardized network component library and floorplan to estimate power and area, (5) adapts an industrial-grade network modeling and simulation environment, OPNET, which makes the methodology ready to use, and (6) can be easily integrated into current HW/SW codesign flow. Using the methodology, ASNoC is generated for a H.264 HDTV decoder SoC and Smart Camera SoC. ASNoC and 2D mesh networks-on-chip are compared in performance, power, and area in detail. The comparison results show that ASNoC provide substantial improvements in power, performance, and cost compared to 2D mesh networks-on-chip. In the H.264 HDTV decoder SoC, ASNoC uses 39% less power, 59% less silicon area, 74% less metal area, 63% less switch capacity, and 69% less interconnection capacity to achieve 2X performance compared to 2D mesh networks-on-chip.
Jiang Xu 0001, Marilyn Wolf, Jörg Henkel, Srimat T. Chakradhar
ACM Trans. Embed. Comput. Syst.2
2006 Code Compression for Embedded VLIW Processors Using Variable-to-Fixed Coding
abstract
In embedded system design, memory is one of the most restricted resources, posing serious constraints on program size. Code compression has been used as a solution to reduce the code size for embedded systems. Lossless data compression techniques are used to compress instructions, which are then decompressed on-the-fly during execution. Previous work used fixed-to-variable coding algorithms that translate fixed-length bit sequences into variable-length bit sequences. In this paper, we present a class of code compression techniques called variable-to-fixed code compression (V2FCC), which uses variable-to-fixed coding schemes based on either Tunstall coding or arithmetic coding. Though the techniques are suitable for both reduced instruction set computer (RISC) and very long instruction word (VLIW) architectures, they favor VLIW architectures which require a high-bandwidth instruction prefetch mechanism to supply multiple operations per cycle, and fast decompression is critical to overcome the communication bottleneck between memory and CPU. Experimental results for a VLIW embedded processor TMS320C6x show that the compression ratios using memoryless V2FCC and Markov V2FCC are around 82.5% and 70%, respectively. Decompression unit designs for memoryless V2FCC and Markov V2FCC are implemented in TSMC 0.25-/spl mu/m technology.
Yuan Xie 0001, Marilyn Wolf, Haris Lekatsas
IEEE Trans. Very Large Scale Integr. Syst.2
2005 Low-leakage robust SRAM cell design for sub-100nm technologies
abstract
A novel low-leakage robust SRAM design for sub-100nm technologies, Hybrid SRAM (HSRAM) cell, is presented in this paper. Leakage power, especially subthreshold leakage and gate leakage, and soft error are challenging the design of SRAM. While these important issues have been separately addressed in previous SRAM designs, there exists no design that simultaneously cuts down leakage power and enhances the resistance to soft error. In this work, we have built the first such SRAM cell, by hybrid of higlw,-. gate dielectric and dynamic threshold voltage which is realized in the form of jointly biased gate and substrate transistor. The HSRAM not only makes the gate leakage negligible, but lessens the severe increase of subthreshold leakage caused by Fringing/Field Induced Barrier Lowering (FIBL) effect accompanied with the introduction of high-ft gate dielectric, and in the same time reduces the susceptibility to soft error by increasing the node capacitance. Experiments were performed in both transistor level and circuit level for this novel HSRAM using ISE8.U and HSPICE. They indicate that up to 03% reduction in total leakage is possible by using HSRAM cell, with an up to 23% increase in reliability degree and and an up to 73% reduction in bitline delay, compared to standard 6T SRAM.
Shengqi Yang, Marilyn Wolf, Wenping Wang 0004, Narayanan Vijaykrishnan, Yuan Xie 0001
ASP-DAC2
2005 Real-time illumination compensation for face processing in video surveillance
abstract
Face processing under illumination variations from video has long been considered as an important and still challenging research issue in video surveillance. In this paper, we propose a real-time pre-processing system to compensate illumination for face processing by using scene lighting modeling. Our contribution lies in the system capability of accommodating multiple local light sources and efficient lighting matching mechanism. Furthermore it can enhance the performance of face processing under both non-standard global and local illumination condition while still maintaining reasonable computing burden for real-time video surveillance. By verifications from face detection results with customized test video clips and public face detection database, the performance of our system outperforms other illumination compensation techniques. This performance improvement in turn will benefit the following face recognition or tracking.
Cheng-Yao Chen, Marilyn Wolf
AVSS2
2005 Frame-level temporal calibration of video sequences from unsynchronized cameras by using projective invariants
abstract
This paper describes a new method for temporally calibrating multiple cameras by image processing operations. Existing multi-camera algorithms assume that the input sequences are synchronized either by genlock or by time stamp information and a centralized server. Yet, hardware-based synchronization increases installation cost. Hence, using image information is necessary to align frames from the cameras whose clocks are not synchronized. Our method uses image processing to find the frame offset between sequences so that they can be aligned. We track foreground objects, extract a point of interest for each object as its current location, and find the corresponding location of the object in the other sequence by using projective invariants in P/sup 2/. Our algorithm recovers the frame offset by matching the tracks in different views, and finding the most reliable match out of the possible track pairs. This method does not require information about intrinsic or extrinsic camera parameters, and thanks to information obtained from multiple tracks, is robust to possible errors in background subtraction or location extraction. We present results on different sequences from the PETS2001 database, which show the robustness of the algorithm in recovering the frame offset.
Senem Velipasalar, Marilyn Wolf
AVSS2
2005 Multimedia Applications of Multiprocessor Systems-on-Chips
abstract
The paper surveys the characteristics of multimedia systems. Multimedia applications today are dominated by compression and decompression, but multimedia devices must also implement many other functions, such as security and file management. We introduce some basic concepts of multimedia algorithms and the larger set of functions that multimedia systems-on-chips must implement.
Marilyn Wolf
DATE1
2005 Power Attack Resistant Cryptosystem Design: A Dynamic Voltage and Frequency Switching Approach
abstract
A novel power attack resistant cryptosystem is presented. Security in digital computing and communication is becoming increasingly important. Design techniques that can protect cryptosystems from leaking information have been studied by several groups. Power attacks, which infer program behavior from observing power supply current into a processor core, are important forms of attack. Various methods have been proposed to counter the popular and efficient power attacks. However, these methods do not adequately protect against power attacks and may introduce new vulnerabilities. We address a novel approach against power attacks, i.e., dynamic voltage and frequency switching (DVFS). Three designs, naive, improved and advanced implementations, have been studied to test the efficiency of DVFS against power attacks. A final advanced realization of our novel cryptosystem is presented; it achieves enough high power trace entropy and time trace entropy to block all kinds of power attacks, with 27% energy reduction and 16% time overhead for DES encryption and decryption algorithms.
Shengqi Yang, Marilyn Wolf, Narayanan Vijaykrishnan, Dimitrios Serpanos, Yuan Xie 0001
DATE2
2005 Modeling image processing systems with homogeneous parameterized dataflow graphs
abstract
We describe a new dataflow model called homogeneous parameterized dataflow (HPDF). This form of dynamic dataflow graph takes advantage of the fact that in a large number of image processing applications, data production and consumption rates, though dynamic, are equal across graph edges for any particular iteration, which leads to a homogeneous rate of actor execution, even though data production and consumption values are dynamic and vary across graph edges. We discuss existing dataflow models and formulate in detail the HPDF model. We develop examples of applications that are described naturally in terms of HPDF semantics and present experimental results that demonstrate the efficacy of the HPDF approach.
Mainak Sen, Shuvra S. Bhattacharyya, Tiehan Lv, Marilyn Wolf
ICASSP (5)4
2005 Multiple object tracking and occlusion handling by information exchange between uncalibrated cameras
abstract
We introduce a novel and robust method for multi-object tracking from multiple uncalibrated cameras. This method improves consistent labeling by incorporating the field of view lines and location information exchange between cameras by using the projective invariants in P/sup 2/. Each camera keeps its own tracks for each target object. This provides improved tracking as well as distributed processing, in which each camera is operated by a separate CPU that performs its own tracking and labeling. The tracking in each camera view is performed by using a two-level hierarchical structure. The main novelties of the proposed method include: a) the ability to communicate between the cameras at any time to improve and update the tracks of an object instead of tracking in each view independently, and to perform this without camera calibration; b) updating the track of an object without interruption and without any need for an estimation of the moving speed and direction, even if the object is totally invisible. The proposed method recovered 90% of the full occlusion cases. The hierarchical tracking structure makes the algorithm computationally efficient and, after background elimination, the first-level tracking runs at about 62 fps on a 2 GHz Celeron machine without code optimization. We present results obtained from the PETS2001 database, which show the success of the camera communication in partial and complete occlusions.
Senem Velipasalar, Marilyn Wolf
ICIP (2)2
2005 An Extended Motion-Estimation Architecture Applied to Shape Recognition
abstract
An architecture for shape recognition is presented, with emphasis on low-latency and power efficiency. This architecture is an extension of an existing architecture used for motion estimation. A number of algorithms were mapped to this architecture. Bounds related to power are given per frame for memory access rates. Face detection within CIPR CIF sequences was used as a target application, with feasible frame rates of 30 fps attained. Power results for this extended architecture correlate with power consumption of the existing architecture
Jason Schlessman, Sankalita Saha, Marilyn Wolf, Shuvra S. Bhattacharyya
ICME3
2005 H.264 HDTV Decoder Using Application-Specific Networks-On-Chip
abstract
This paper studied an H. 264 HDTV decoder on two multiprocessor system-on-chip architectures. Two types of networks-on-chip, the RAW network and the application specific networks-on-chip, were used. Regular-topology networks-on-chip (mesh, torus, and fat tree) have been proposed. However, we showed in this paper that the application-specific networks-on-chip provided substantial improvements in power, performance, and cost compared to regular-topology networks-on-chip. We measured the power, performance, area, total switch and link capacity, and switch and link utilization based on floorplans and circuit designs. Measurement results showed th at the application-specific networks-on-chip was both faster in absolute terms and more efficient. The application-specific networks-on-chip used 39% less power, 59% less silicon area, 74% less metal area, 63% less switch capacity, and 69% less link capacity to achieve 2X performance compared to the RAW network.
Jiang Xu 0001, Marilyn Wolf, Jörg Henkel, Srimat T. Chakradhar
ICME2
2005 Distributed Peer-to-Peer Smart Cameras: Algorithms and Architectures
abstract
Summary form only given. Advances in VLSI allow us to make cheap cameras and to supply them with powerful processors. To harness these capabilities, we need to move to peer-to-peer networks of smart cameras. Such systems perform distributed video analysis without a central server. Peer-to-peer systems save bandwidth and energy, are cheaper to install, and are more fault-tolerant. The Embedded Systems Group at Princeton University is developing smart cameras and peer-to-peer networks. After describing the application demands, we briefly describe architectures for embedded real-time video processing in smart cameras. We then describe video algorithms and network architectures for peer-to-peer gesture recognition and tracking.
Marilyn Wolf
ISM1
2005 Power and Performance Analysis of Motion Estimation Based on Hardware and Software Realizations
abstract
Motion estimation is the most computationally expensive task in MPEG-style video compression. Video compression is starting to be widely used in battery-powered terminals, but surprisingly little is known about the power consumption of modern motion estimation algorithms. This paper describes our effort to analyze the power and performance of realistic motion estimation algorithms in both hardware and software realizations. For custom hardware realizations, this paper presents a general model of VLSI motion estimation architectures. This model allows us to analyze in detail the power consumption of a large class of modern motion estimation engines that can execute the motion estimation algorithms of interest to us. We compare these algorithms in terms of their power consumption and performance. For software realizations, this paper provides the first detailed instruction-level simulation results on motion estimation based on a programmable CPU core. We analyzed various aspects of the selected motion estimation algorithms, such as search speed and power distribution. This paper provides a guideline to two types of machine designs for motion estimation: custom ASIC (application specific integrated circuit) design and custom ASIP (application specific instruction-set processor) designs.
Shengqi Yang, Marilyn Wolf, Narayanan Vijaykrishnan
IEEE Trans. Computers2
2005 Unification of scheduling, binding, and retiming to reduce power consumption under timings and resources constraints
abstract
Scheduling and binding are two tasks found in high-level synthesis of hardware as well as in compiling software. These tasks are realized on graphs that are models of the hardware or of the software to be compiled to run on a specific processor. Scheduling focuses on determining the start execution time of each node in the graph. Binding is the task of assigning each node in the graph to a specific computational element. Realize binding before or after scheduling can exclude generating high-quality designs (hardware or binary code). The latter statement is true in particular in the era of design for low power. Do not combine scheduling and binding can lead to designs with high switching activities and hence to high power consumption. To the best of our knowledge, there is no approach at this moment that addresses the problem of unifying scheduling and binding with an exact algorithm to produce designs with reduced power consumption. Known approaches to that problem are heuristics. That problem is NP-hard in general, since it is the composition of two NP-hard problems. Also, it has not yet been formulated in the literature. The problem becomes more complex when one has to deal with cyclic graphs and/or there are constraints to be met such as timings. For cyclic graphs, one has to integrate retiming in the unification of scheduling and binding. We propose a mathematical formulation to that problem. We extend this formulation to solve the problem of combining modulo scheduling, binding, and retiming under timings and resources constraints while reducing power consumption due to switching activities. The proposed approach is tested using known benchmarks. Based on obtained numerical results, this approach is able to reduce power consumption by 33.24% on average, with an average of 33.83 s as a run time.
Noureddine Chabini, Marilyn Wolf
IEEE Trans. Very Large Scale Integr. Syst.2
2005 Approximate arithmetic coding for bus transition reduction in low power designs
abstract
We present a method for reducing the power consumption of compressed-code systems by selectively inverting bits that are transmitted on the bus. By incorporating bus inversion into code compression/decompression, we reduce power consumption with no cost in hardware or power relative to code compression without inversion. Inverting has to be done carefully to ensure that the codes can still be decoded. As an additional challenge, compression will generally increase bit-toggling as it removes redundancies from the code transmitted. Therefore, we need to find the right balance between compression ratio and bit-toggling reduction. This paper presents a suitable algorithm that will combine approximate compression techniques with bit-toggling reduction and will explore the various tradeoffs. We take advantage of the approximations introduced to modify codes and reduce bit-toggling, while maintaining compression performance and decoding speed. An interesting result that is derived from our work is that high compression ratios do not necessarily result in the lowest power consumption. By using our method, bus-related power consumption has been reduced by as much as 35% compared to a system with no compression, and as much as 14% compared to a compressed-code system. Bit-toggling reduction does not impose any additional hardware costs other than the decompression engine. We also present a detailed analysis on how bus widths affect bit-toggling when transmitting compressed code, and we show experimental results on ARM, MIPS, and SPARC code. We finally compare our work with Bus Invert and show results that are superior except for the random data case where Bus Invert performs better.
Haris Lekatsas, Jörg Henkel, Marilyn Wolf
IEEE Trans. Very Large Scale Integr. Syst.3
2004 An approach for reducing dynamic power consumption in synchronous sequential digital designs
Noureddine Chabini, Marilyn Wolf
ASP-DAC2
2004 Synthesizing interconnect-efficient low density parity check codes
abstract
Error correcting codes are widely used in communication and storage applications. Codec complexity has usually been measured with a software implementation in mind. A recent hardware implementation of a Low Density Parity Check code (LDPC) indicates that interconnect complexity dominates the VLSI cost. We describe a heuristic interconnect-aware synthesis algorithm which generates LDPC codes that use an order of magnitude less wiring with little or no loss of coding efficiency.
Marghoob Mohiyuddin, Adnan Aziz, Marilyn Wolf
DAC4
2004 The future of multiprocessor systems-on-chips
abstract
This paper surveys the state-of-the-art and pending challenges in MPSoC design. Standards in communications, multimedia, networking, and other areas encourage the development of high-performance platforms that can support a range of implementations of the standard. A multiprocessor system-on-chip includes embedded processors, digital logic, and mixed-signal circuits combined into a heterogeneous multiprocessor. This mix of technologies creates a major challenge for MPSoC design teams. We will look at some existing MPSoC designs and then describe some hardware and software challenges for MPSoC designers.
Marilyn Wolf
DAC1
2004 LZW-Based Code Compression for VLIW Embedded Systems
abstract
We propose a new variable-sized-block method for VLIW code compression. Code compression traditionally works on fixed-sized blocks and its efficiency is limited by the small block size. Branch blocks-instructions between two consecutive possible branch targets-provide larger blocks for code compression. We propose LZW - based algorithms to compress branch blocks. Our approach is fully adaptive and generates coding table on-the-fly during compression and decompression. When encountering a branch target, the coding table is cleared to ensure correctness. Decompression requires only a simple lookup and update when necessary. Our method provides 8 bytes peak decompression bandwidth and 1.82 bytes in average. Compared to Huffman's 1 byte and V2F's 13-bit peak performance, our methods have higher decoding bandwidth and comparable compression ratio. Parallel decompression could also be applied to our methods, which is more suitable for VLIW architecture.
Chang Hong Lin, Yuan Xie 0001, Marilyn Wolf
DATE3
2004 A Case Study in Networks-on-Chip Design for Embedded Video
abstract
In this paper we study bus-based and switch-based on-chip networks for an embedded video application, the smart camera SoC (system on chip). We analyze network performance and overall system performance in detail. We explore system performance using crossbars with different sizes, fixed size but different numbers of ports, and different numbers of shared memories. We find that network is a performance bottleneck in our design, and the system using an optimized NoC can outperform one using a bus by 132%. Our simulations are based upon recorded real communication traces, which give more accurate system performance. Our study finds that for the Smart Camera system, a 16-bit/port 3/spl times/3 crossbar with two shared memories shows 85.7% performance improvement over the bus-based model and also has less maximum network throughput than the bus-based model. This design example illustrates a methodology to quickly and accurately estimate the performance of NoC's at architecture level.
Jiang Xu 0001, Marilyn Wolf, Jörg Henkel, Srimat T. Chakradhar, Tiehan Lv
DATE2
2004 An approach for integrating basic retiming and software pipelining
abstract
Basic retiming is an algorithm originally developed for hardware optimization. Software pipelining is a technique proposed to increase instruction-level parallelism for parallel processors. In this paper, we show that applying software pipelining alone for minimizing timings under resource constraints can lead to sub-optimal results, compared to the case if an unification of basic retiming and software pipelining is used. We propose an approach to realize this unification. The approach allows to minimize the code size of the optimized loop as well as minimizing the idleness of computational elements. We extend this approach to solve the problem of minimizing peak power consumption for time-constrained and resource-constrained software pipelined loops. Solving these problems is important for portable embedded systems as well as system-on-chip design. The approaches are tested using known benchmarks. On average, relative timing improvement is 60.19%, and relative reduction of peak power consumption is 13.17% without any trade-off in timings.
Noureddine Chabini, Marilyn Wolf
EMSOFT2
2004 Smart Cameras and Pervasive Information Systems
Marilyn Wolf
EUC1
2004 Challenges in System-Level Design
Marilyn Wolf
FMCAD1
2004 Recovering field of view lines by using projective invariants
abstract
Establishing correspondences between moving objects is an important problem in multiple camera tracking, and field of view (FOV) lines have been introduced in literature as an efficient tool to resolve the consistent labeling issue. We introduce a new and robust method to find the FOV lines, which uses projective invariants that does not rely on the object movement in the scene and does not require information about the camera parameters. As the labeling scheme suggested before is based on the distance of an object to an FOV line, accurate recovery of the FOV lines provides reliable labeling, and makes less errors in the process. We present results on different sequences, obtained from the PETS200I database, which show the robustness of the algorithm in recovering all visible FOV lines of another camera in the current camera view.
Senem Velipasalar, Marilyn Wolf
ICIP2
2004 A peer-to-peer architecture for distributed real-time gesture recognition
abstract
We describe a peer-to-peer multiple-camera architecture for a distributed real-time gesture recognition system. Previous work attaches multiple cameras to a server This simplifies many design problems but is impractical for real-world installations. Our architecture uses a network of relatively inexpensive cameras to gather images in order to provide high resolution at low cost. Computations are done at the embedded processors in each camera, without using a centralized server. We also propose a methodology for transforming well-defined single-camera algorithms to multiple cameras. We migrate our single-camera gesture recognition system into multiple cameras with slightly overlapped views. In order to minimize the communication bandwidth and power consumption, only selected contours or ellipses information is transmitted between the cameras.
Chang Hong Lin, Tiehan Lv, Marilyn Wolf, I. Burak Özer
ICME3
2004 A real-time background subtraction method with camera motion compensation
abstract
Background subtraction algorithms are critical to many video recognition/analysis systems and have been studied for decades. Most of the algorithms assume that the camera is fixed. We propose a background subtraction algorithm that works when a shaking camera is present. In this algorithm, the input frames are compensated and compared with the given reference frame to separate foreground objects from the background. Experimental results show that the proposed method outperforms the widely used Gaussian mixture model based method in both fixed camera and shaking camera scenarios with respect to accuracy, robustness, and efficiency.
Tiehan Lv, I. Burak Özer, Marilyn Wolf
ICME3
2004 Search speed and power driven integrated software and hardware optimizations for motion estimation algorithms
abstract
The motion estimation (ME) is the most time consuming and power demanding module in a video encoder. In this paper, we provide the first detailed instruction-level simulation results on motion estimation. Firstly, we analyze various aspects of the selected seven typical motion estimation algorithms, such as search speed, instruction frequencies, branch behavior, and power distribution. Then, we propose suitable compiler techniques, hardware techniques and optimal instruction set architecture to speed up and reduce the power consumption of the motion estimation. SimpleScalar and SimplePower are used to simulate the effects of the optimizations on the motion estimation algorithms. Experimental results show that the execution time spent on motion vector search for one frame are averagely reduced by 39% and the power consumption is reduced by 62% by applying these techniques
Shengqi Yang, Marilyn Wolf, Narayanan Vijaykrishnan
ICME2
2004 Reducing dynamic power consumption in synchronous sequential digital designs using retiming and supply voltage scaling
abstract
The problem of minimizing dynamic power consumption by scaling down the supply voltage of computational elements off critical paths is widely addressed in the literature for the case of combinational designs. The problem is NP-hard in general. To address the problem in the case of synchronous sequential digital designs, one needs to move some registers while applying voltage scaling. Moving these registers shifts some computational elements from critical paths, and can be done by basic retiming. Integrating basic retiming and supply voltage scaling to address this NP-hard problem cannot in general be done in polynomial run time. In this paper, we propose to first apply a guided retiming and then to apply supply voltage scaling on the retimed design. We devise new polynomial time algorithms to realize this guided retiming, and the supply voltage scaling on the retimed design. Also, we show that the problem in the case of combinational designs is not NP-hard for some combinational circuits with certain structure, and give a polynomial time algorithm to optimally solve it. Methods to determine lower bounds on the optimal reduction of dynamic power consumption are also provided. Experimental results on known benchmarks have shown that the proposed approach can reduce dynamic power consumption by factors as high as 61% for single-phase designs with minimal clock period. Also, they have shown that it can solve optimally the problem, and produce converter-free designs with reduced dynamic power consumption. For large size circuits from ISCAS'89 benchmark suite, the proposed algorithms run in 15 s-1 h.
Noureddine Chabini, Marilyn Wolf
IEEE Trans. Very Large Scale Integr. Syst.2
2003 Enhancing Signal Integrity through a Low-Overhead Encoding Scheme on Address Buses
abstract
Signal integrity is and will continue to be a major concern in deep sub-micron VLSI designs where the proximity of signal carrying lines leads to crosstalk, unpredictable signal delays and other parasitic side effects. Our scheme uses bus encoding that guarantees that at any time any two signal carrying lines will be separated by at least one grounded line and thus providing a high degree of signal integrity. This comes at a small overhead of only one additional bus line (the closest related work needs 14 additional lines for a 32-bit bus) and a small average performance decrease of 0.36%. By means of a large set of real-world applications, we compare our scheme to other state-of-the-art approaches and present comparisons in terms of degree of integrity, overhead (e.g. additional lines required) and a possible performance decrease.
Tiehan Lv, Jörg Henkel, Haris Lekatsas, Marilyn Wolf
DATE4
2003 Profile-Driven Selective Code Compression
Yuan Xie 0001, Marilyn Wolf, Haris Lekatsas
DATE2
2003 Code Compression Using Variable-to-fixed Coding Based on Arithmetic Coding
abstract
Embedded computing systems are space and cost sensitive. Memory is one of the most restricted resources that post serious constraints on program size. Code compression, which is a special case of data compression where the input source is in machine instructions, has been proposed as a solution to this problem. Previous work in code compression has focused on either fixed-to-variable coding or dictionary-based algorithms. Code compression schemes that use variable-to-fixed (V2F) length coding were proposed, based on arithmetic coding. Experiments have shown that the compression ratio, using memoryless V2F coding for the TMS320C6x processor, have an average of 82.5% and decompression can be parallelized. A Markov-based V2F coding based on arithmetic coding has achieved an average compression ratio of 72% for TMS320C6x while decompression cannot be parallelized. Furthermore, the given experiments have shown that arithmetic coding based V2F coding has similar compression performance to Tunstall coding. Finally, a power reduction scheme for the instruction bus using the V2F coding scheme was presented.
Yuan Xie 0001, Marilyn Wolf, Haris Lekatsas
DCC2
2003 Minimizing Variables' Lifetime in Loop-Intensive Applications
Noureddine Chabini, Marilyn Wolf
EMSOFT2
2003 Exploiting parallelism in media processing using VLIW processor
abstract
This paper presents the results of analyzing multimedia applications that have been implemented on a trimedia VLIW processor, TM1300. The TM1300 processor offers several media-specific features, such as multimedia instructions, that experience has shown to be very important for obtaining maximum performance from the VLlW machine. Other researchers have studied the effects of instructions on superscalar machines, but we know of no other work that measures the utilization of real VLIW instruction sets on significant applications. We will study static and dynamic instruction counts, interactions between instructions, and other effects. Our experiments show that a TM1300 processor provides a 5/spl times/ speedup over a single-issue RISC processor in equivalent technology on average and more than 10/spl times/ speedup in best case.
Marilyn Wolf, I. Burak Özer, Tiehan Lv
ICIP (3)1
2003 Architectures for distributed smart cameras
abstract
This paper describes our new multiple-camera architecture for real-time video analysis. This architecture uses an array of relatively inexpensive cameras to gather images in order to provide high resolution at low cost. The system also uses a hierarchy of cameras, including both wide-angle and telephoto views. Wide-angle cameras are responsible for camera coordination while telephoto cameras are primarily responsible for detailed processing of parts of the scene.
Marilyn Wolf, I. Burak Özer, Tiehan Lv
ICME1
2003 Memory system optimization of embedded software
abstract
The memory system often determines a great deal about the behavior of an embedded system: performance, power, and manufacturing cost. A great many software techniques have been developed over the past decade to optimize software to improve these characteristics. Embedded software design and compilation can take advantage of two important facts: the hardware target is known; and we can spend more time and computational effort to optimize the software. This paper surveys techniques for optimizing memory behavior of embedded software and points to some future trends in the field.
Marilyn Wolf, Mahmut T. Kandemir
Proc. IEEE1
2003 A dictionary-based en/decoding scheme for low-power data buses
abstract
As bus lengths on multihundred-million transistor systems-on-a-chip (SoC) grow, and as interwire capacitances of sub-0.10 /spl mu/m technologies advance, the resulting high-switching capacitances of buses (and interconnects in general) have a nonnegligible impact on the power consumption of a whole SoC. This trend has been recognized and recently addressed by various research groups. We address this problem by introducing our bus encoding technique, adaptive dictionary-encoding scheme "ADES" that minimizes the power consumption of data buses through a dictionary-based encoding technique. Based on exploration of data properties on buses, our technique saves on average more than 25% of bus energy compared to the nonencoded cases using a large set of real-world applications for both address and data buses. Furthermore, we compare our technique to the best-known data bus encoding techniques to date and we find that it exceeds all of them in terms of energy savings for the same set of applications.
Tiehan Lv, Jörg Henkel, Haris Lekatsas, Marilyn Wolf
IEEE Trans. Very Large Scale Integr. Syst.4
2002 Wave pipelining for application-specific networks-on-chips
abstract
This paper presents methods for optimizing application-specific networks-on-chips (NoCs). We show that wave pipelining provides more energy efficient data transport than non-wave pipelined communication. We observe 52% energy saving, 60% transistor area saving, and 1.7 times speedup by using wave pipelining in simulation. Wave pipelining is particularly well suited to networks-on-chips because the networkes structured interconnection provides better delay control. Our analysis shows how designers can tune their network to the requirements of the application by choosing a design point along area/performance or area/energy curves.
Jiang Xu 0001, Marilyn Wolf
CASES2
2002 Dynamic Runtime Re-Scheduling Allowing Multiple Implementations of a Task for Platform-Based Designs
abstract
This paper introduces an extension to the RMS scheduling technique that we call "hot swapping". Hot swapping enables a system to choose between various selected implementations of one task on-the-fly and thus to optimize the system's cost (e.g. power savings). The on-the-fly swapping between those implementations requires extra time to save and/or transform states of a certain task implementation. Even if the two steady-state schedules before and after the swapping are feasible, the transient schedule with the additional swapping computation time may exceed the system's capacity. Our technique is an extension to rate monotonic scheduling (RMS). While maintaining and meeting performance requirements, our technique shows an average reduction of 31% in power consumption compared to systems using a pure static scheduling approach (RMS) that cannot make use of task swapping. We have evaluated our algorithm through simulation of five real-world task sets and in addition by use of a large number of generated task sets.
Tin-Man Lee, Marilyn Wolf, Jörg Henkel
DATE2
2002 An Adaptive Dictionary Encoding Scheme for SOC Data Buses
abstract
As bus lengths on multi-hundred-million transistor SOCs (systems-on-a-chip) grow and as inter-wire capacitances of sub-0.1 /spl mu/m technologies increase, the resulting high switching capacitances of buses (and interconnects in general) have a non-negligible impact on the power consumption of a whole SOC. In this paper, we address this problem by introducing our bus encoding technique 'ADES' that minimizes the power consumption of data buses through a dictionary-based encoding technique. We show that our technique saves between 18% and 40% of bus energy compared to the non-encoded cases using a large set of (freely-accessible) real-world applications. Furthermore, we compare our technique to the best-known data bus encoding techniques to date and show that it exceeds all of them in energy savings for the same set of applications. The additional hardware effort for our bus en/decoder is thereby very small.
Tiehan Lv, Marilyn Wolf, Jörg Henkel, Haris Lekatsas
DATE2
2002 iSKIP: a fair and efficient scheduling algorithm for input-queued crossbar switches
abstract
Cell-based input-queued crossbars are widely used in networking equipments. A number of crossbar scheduling algorithms have been proposed to provide high performance scheduling. It is critical for a crossbar scheduling algorithm to be fair and efficient in real world conditions, which include the presence of over-subscribed ingress ports and the possible flow control from the egress ports. Several commonly deployed algorithms are not fair when there are ingress congestions due to over-subscription. Furthermore, these algorithms lose fairness and can even cause starvation when there is flow control or backpressure from the egress ports. We illustrate the problems in detail using the well-known iSLIP algorithm. In this paper, we present a new algorithm, iSKIP, which performs as efficiently as iSLIP in a benign environment but remains fair and starvation-free in the cases of ingress congestion and egress backpressure. The iSKIP algorithm can be implemented in fast and simple hardware. Simulation results are presented to illustrate the advantages of the iSKIP algorithm.
Walter Wang, Libin Dong, Marilyn Wolf
GLOBECOM3
2002 A bottom-up approach for activity recognition in smart rooms
abstract
We propose a smart camera system where the cameras detect the presence of a person and recognize activities of this person. A relational graph-based modeling of human body and a HMM-based activity recognition of the body parts are proposed for real-time video analysis. The results show that more than 86 percent of the body parts and 88 percent of the activities are correctly classified. We also describe the relationship between the activity detection algorithms and the architectures required to perform these tasks in real time. We achieve a processing rate of more than 20 frames per second for each TriMedia video capture board.
I. Burak Özer, Tiehan Lv, Marilyn Wolf
ICME (1)3
2002 A Distributed Switch Architecture with Dynamic Load-balancing and Parallel Input-Queued Crossbars for Terabit Switch Fabrics
abstract
Distributed switch architectures allow the partition of a switch fabric into smaller independent switches, which is a key advantage for building high capacity switching systems, such as terabit switches. The performance and the feasibility of a distributed switch architecture depend on two critical components: the design of the queueing structure and the load balancing algorithm. In this paper, we present a distributed switch architecture with a simple yet efficient queueing structure and load-balancing algorithm that can be easily implemented in a terabit switch fabric with OC-768 line rate. The queueing structure is based on distributed non-buffered input-queued crossbar switch elements with request-only virtual output queues. The distributed load-balancing algorithm dynamically balances workloads among the parallel switch elements by trying to equalize the length of request-only virtual output queues in each of the switch elements. We refer to our architecture as a distributed switch architecture (ADSA), which introduces little communication overhead and no throughput degradation. As a result, it enables non-blocking switching without the need for internal speed-up. The load-balancing algorithm can perform one load-balancing action in less than 10 ns, which is suitable for OC-768 line rates of 40 Gbps. We study the performance of the ADSA architecture by modeling it as discrete-time queues with uniform i.i.d. Bernoulli traffic. We use a combination of analytical and simulation approaches to show that the ADSA can be approximated as discrete-time Geom/G/P queues under both light and heavy loads. Then the Allen-Cunneen approximation formula is applied to derive the mean cell delay of the ADSA architecture as a function of the underlying crossbar scheduling algorithm, the number of parallel switch elements and the number of ports in the system.
Walter Wang, Libin Dong, Marilyn Wolf
INFOCOM3
2002 A Graph-Based Object Description for Information Retrieval in Digital Image and Video Libraries
I. Burak Özer, Marilyn Wolf, Ali N. Akansu
J. Vis. Commun. Image Represent.2
2002 Introduction to the inaugural issue
abstract
No abstract available.
Marilyn Wolf
ACM Trans. Embed. Comput. Syst.1
2002 A hierarchical human detection system in (un)compressed domains
abstract
We propose a hierarchical retrieval system where shape, color and motion characteristics of the human body are captured in compressed and uncompressed domains. The proposed retrieval method provides human detection and activity recognition at different resolution levels from low complexity to low false rates and connects low level features to high level semantics by developing relational object and activity presentations. The available information of standard video compression algorithms are used in order to reduce the amount of time and storage needed for the information retrieval. The principal component analysis is used for activity recognition using MPEG motion vectors and results are presented for walking, kicking, and running to demonstrate that the classification among activities is clearly visible. For low resolution and monochrome images it is demonstrated that the structural information of human silhouettes can be captured from AC-DCT coefficients.
I. Burak Özer, Marilyn Wolf
IEEE Trans. Multim.2
2001 Allocation and scheduling of conditional task graph in hardware/software co-synthesis
abstract
This paper introduces an allocation and scheduling algorithm that efficiently handles conditional execution in multi-rate embedded system. Control dependencies are introduced into the task graph model. We propose a mutual exclusion detection algorithm that helps the scheduling algorithm to exploit the resource sharing. Allocation and scheduling are performed simultaneously to take advantage of the resource sharing among those mutual exclusive tasks. The algorithm is fast and efficient, and so is suitable to be used in the inner loop of our hardware/software co-synthesis framework which must call the scheduling routine many times.
Yuan Xie 0001, Marilyn Wolf
DATE2
2001 Code Compression for VLIW Processors
Yuan Xie 0001, Haris Lekatsas, Marilyn Wolf
Data Compression Conference3
2001 Human detection in compressed domain
abstract
We propose an algorithm for human detection in JPEG compressed still images and MPEG I-frames. In this new algorithm, the overall shape of a standing or walking person is detected by using an eigenspace representation of human silhouettes obtained from AC-DCT coefficients. Our approach is invariant to changes in intensity, color and textures and has the advantage of using the available data in the standard compression algorithms. The algorithm achieves a correct detection rate of 80% for frontal and rear views of human bodies in cluttered scenes.
I. Burak Özer, Marilyn Wolf
ICIP (3)2
2001 Partial-Frame Transition Detection
abstract
This paper introduces a new video analysis problem, partial-frame transition detection, and describes an algorithm to solve the problem. Digital video allows the creation of new transitions between shots that were difficult to make with film—for example, wipes and dissolves can now be performed on only a part of the frame. These partial-frame transitions carry important semantic information but are often too small to be picked up by traditional transition detection algorithms. This paper describes the properties of partial-frame transitions and shows how to detect them.
Marilyn Wolf
ICME1
2001 Dynamic Parallel media processing using Speculative Broadcast Loop (SBL)
abstract
This paper presents the results of a study of dynamic parallel media processing using Speculative Broadcast Loop (SBL), a speculative run-time looplevel parallelization method. Due to processing regularity, multimedia applications typically contain extensive parallelism. Subword parallelism methods are commonly used to support data parallelism between independent loop iterations in inner loops, but much of the data parallelism in media processing resides in outer loops and cannot be supported with subword parallelism. Larger-scale parallel methods are needed to enable use of the full range of data parallelism in multimedia. Because static parallel compilation methods are often unable to recognize all parallelism at compile time, a run-time method is assumed for the speculative execution of potentially parallel loops. The SBL run-time method combines SIMD parallelism with large-scale speculation for supporting data parallelism in multimedia.
Jason Fritts, Marilyn Wolf
IPDPS2
2001 A code decompression architecture for VLIW processors
abstract
In embedded system design, memory has been one of the most restricted resources. Reducing program size has been an important goal when designing an embedded system. Most of the previous work on code compression has targeted RISC architectures. Recently VLIW processors became very popular, particularly for signal processing. Decompression speed is especially important for VLIW architectures given that the length of the instruction word is long. Furthermore, modem VLIW architectures use flexible instruction formats, which require new code compression approaches. Previous work has assumed that instruction positions within the long instruction word correspond to specific functional units. In contrast, our code compression algorithm is capable of compressing flexible instruction formats, where any functional unit can be used for any position in the instruction word. We demonstrate our methods by applying it to the TMS320C6x architecture. We also compare two techniques for decompressing the VLIW instruction packet to reduce the decompression time. A fast parallel decompression architecture is described, which is implemented in TSMC 0.25 technology.
Yuan Xie 0001, Marilyn Wolf, Haris Lekatsas
MICRO2
2001 Multi-Modal Dialog Scene Detection Using Hidden Markov Models for Content-Based Multimedia Indexing
A. Aydin Alatan, Ali N. Akansu, Marilyn Wolf
Multim. Tools Appl.3
2001 RAGS-real-analysis ALAP-guided synthesis
abstract
A new technique for single-bus heterogeneous system scheduling and hardware-software codesign is presented. This technique addresses challenging real-time problem domains (e.g., multirate periodic dependent tasks) for preemptive target systems including real-time operating system overheads and communication contention along with proper protocol handling. Such realistic system attributes are often times ignored in scheduling and codesign efforts, but are addressed here using a novel "real-analysis" approach. This technique foregoes the usual list- or cluster-oriented scheduling techniques for an as-late-as-possible (ALAP) guided iterative improvement procedure. A detailed system simulation is used at the scheduling level, which in turn is used as the core of the overall cosynthesis. Since an accurate simulation is the basis for the schedule feasibility check, separate verification steps become unnecessary. In addition to using real analysis as part of the scheduling, several other unique or unusual features are employed including the use of ALAP in heterogeneous scheduling, recursive search level at a time, allowing backtracking while maintaining polynomial execution time, as well as others. This framework appears to be unique in its ability to address the allocation/scheduling and codesign of heterogeneous systems, which, in particular, employ arbitrated buses for intertask communication.
David L. Rhodes, Marilyn Wolf
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2000 Co-synthesis with custom ASICs
abstract
Abstract- This paper introduces the first hardware/software co-synthesis algorithm that optimizes the implementations of ASICs that are used as processing elements for the embedded systems. Many real time embedded systems are composed of heterogeneous processing elements, such as general purpose CPUs, ASICs and FPGAs. Previous work has not considered how to select one of several possible ASIC implementations for a specific task. We have developed a heuristic iterative improvement algorithm for distributed embedded system cosynthesis. We use Monet, a behavioral level architectural exploration system, to generate multiple implementations of a behavioral description of an ASIC and to analyze their performance. To the best of our knowledge, this is the first cosynthesis algorithm that takes into account the impact of different ASIC implementations of tasks on system performance and cost in the co-synthesis process. I.
Yuan Xie 0001, Marilyn Wolf
ASP-DAC2
2000 Code compression for low power embedded system design
abstract
We propose instruction code compression as an efficient method for reducing power on an embedded system. Our approach is the first one to measure and optimize the power consumption of a complete SOC (System--On--a--Chip) comprising a CPU, instruction cache, data cache, main memory, data buses and address bus through code compression. We compare the pre-cache architecture (decompressor between main memory and cache) to a novel post-cache architecture (decompressor between cache and CPU). Our simulations and synthesis results show that our methodology results in large energy savings between 22% and 82% compared to the same system without code compression. Furthermore, we demonstrate that power savings come with reduced chip area and the same or even improved performance.
Haris Lekatsas, Jörg Henkel, Marilyn Wolf
DAC3
2000 Embedded systems education (panel abstract)
abstract
The design and design automation of embedded systems is rapidly emerging as a research area in its own right. It draws from several traditional areas of study such as system specification, modeling and analysis; computer architecture and micro-architecture; as well as compilers and operating systems. However, the embedded domain adds some interesting twists in terms of tighter problem constraints that demand a fresh look at even these traditional areas. In addition, there are several emerging EDA areas such as design reuse and integration of systems on a chip that are critical to the study of embedded systems. These aspects are not typically covered by computer engineering and EDA curricula. This panel addresses the challenges associated with the educational issues in embedded systems design and design automation. The panelists will examine issues in including embedded systems in university curricula, as well as in setting up research programs that are crucial for the education of graduate students.
Sharad Malik, D. K. Arvind 0001, Edward A. Lee, Philip Koopman, Alberto L. Sangiovanni-Vincentelli, Marilyn Wolf
DAC6
2000 Arithmetic Coding for Low Power Embedded System Design
abstract
We present a novel algorithm that assigns codes to instructions during instruction code compression in order to minimize bus-related bit-toggling and thus reduce power consumption. The target application area is embedded systems, where power consumption is increasingly becoming a dominant design constraint. Our algorithm is based on a variant of quasi-arithmetic coding where coding allows for random access and fast table-based decoding. We take advantage of the approximations introduced to modify codes and reduce bit-toggling, while maintaining compression performance and decoding speed. We present the first work to explore the trade-offs between compression ratios and bus-related power consumption and show that high compression ratios do not necessarily result in the lowest power consumption. By using our method, bus-related power consumption has been reduced by as much as 35% without imposing any additional hardware costs.
Haris Lekatsas, Marilyn Wolf, Jörg Henkel
Data Compression Conference2
2000 Comparative analysis of hidden Markov models for multi-modal dialogue scene indexing
abstract
A class of audio-visual content is segmented into dialogue scenes using the state transitions of a novel hidden Markov model (HMM). Each shot is classified using both the audio track and the visual content to determine the state/scene transitions of the model. After simulations with circular and left-to-right HMM topologies, it is observed that both performing very well with multi-modal inputs. Moreover, for the circular topology, the comparisons between different training and observation sets show that audio and face information together gives the most consistent results among different observation sets.
A. Aydin Alatan, Ali N. Akansu, Marilyn Wolf
ICASSP3
2000 A Decompression Architecture for Low Power Embedded Systems
abstract
We present an architecture for embedded systems that decompresses offline-compressed instructions during runtime. This is useful for compressed code systems where instructions are stored in a compressed format and decompressed on demand. The result is a significant reduction in power consumption, and in most cases a performance improvement. The stand-alone decompression engine is placed between the instruction cache and the CPU (post-cache architecture) as we have found this to be the most power-efficient architecture. This paper describes the design of this unit in detail and analyzes its power consumption and performance.
Haris Lekatsas, Jörg Henkel, Marilyn Wolf
ICCD3
2000 An Experimental Analysis of Digital Video Library Servers
Michael A. Kozuch, Marilyn Wolf, Andrew Wolfe
Multim. Syst.2
1999 Random Access Decompression Using Binary Arithmetic Coding
abstract
We present an algorithm based on arithmetic coding that allows decompression to start at any point in the compressed file. This random access requirement poses some restrictions on the implementation of arithmetic coding and on the model used. Our main application area is executable code compression for computer systems where machine instructions are decompressed on-the-fly before execution. We focus on the decompression side of arithmetic coding and we propose a fast decoding scheme based on finite state machines. Furthermore, we present a method to decode multiple bits per cycle, while keeping the size of the decoder small.
Haris Lekatsas, Marilyn Wolf
Data Compression Conference2
1999 Co-synthesis of heterogeneous multiprocessor systems using arbitrated communication
abstract
We describe the first co-design technique aimed at heterogeneous systems employing arbitrated communication. Arbitrated system design is especially difficult because communication scheduling is directly tied to task allocation. The method provides a complete co-design-i.e. generation of a hardware configuration along with an allocation and schedule for the execution of hard real-time data-dependent tasks. By using an actual scheduling analysis in the inner co-design loop, the method is readily able to address realistic system effects including various communication models like arbitration, as in PCI-based systems.
David L. Rhodes, Marilyn Wolf
ICCAD2
1999 CAD Techniques for Embedded Systems-on-Silicon
Marilyn Wolf
ICCD1
1999 Parallel Media Processors for the Billion-Transistor Era
abstract
This paper describes the challenges presented by single-chip parallel media processors (PMPs). These machines integrate multiple parallel function units, instruction execution, and memory hierarchies on a single chip. The combination of programmability and high performance on data parallelism is necessary to meet the demands of next-generation multimedia applications. Many research issues must be solved to realize the full potential of programmable media processors. This paper provides both a survey of research trends and issues in architecture and compiler design for programmable media processors, and an exploration of the potential performance of media processors over the next decade.
Jason Fritts, Marilyn Wolf
ICPP3
1999 Position Statement: Testing in a VLSI Design Course
abstract
I’ve believed for a long time that testing concepts are important to VLSI design. That’s why I included testing concepts in my book Modern VLSI Design [Wol98]. In our VLSI course at Princeton, I try to cover testing concepts at each level of the design hierarchy: basic stuck-at fault models and combinational test generation when we talk about combinational logic; sequential test generation when we talk about sequential logic; and highlevel design-for-testability when we discuss architecture design. I stress that testability is important and that it must be designed into the chip. I think that giving them some understanding of test fundamentals helps them understand both the importance of testing and some basic techniques for designing testable chips that they can apply when they go to industry.
Marilyn Wolf
ITC1
1999 A Hierarchical Multiresolution Video Shot Transition Detection Scheme
Hong Heather Yu, Marilyn Wolf
Comput. Vis. Image Underst.2
1999 Guest Editorial
Marilyn Wolf
Multim. Syst.1
1999 SAMC: a code compression algorithm for embedded processors
abstract
In this paper, we present a method for reducing the memory requirements of an embedded system by using code compression. We compress the instruction segment of the executable running on the embedded system, and we show how to design a run-time decompression unit to decompress code on the fly before execution. Our algorithm uses arithmetic coding in combination with a Markov model, which is adapted to the instruction set and the application. We provide experimental results on two architectures, Analog Devices' Share and ARM's ARM and Thumb instruction sets, and show that programs can often be reduced by more than 50%. Furthermore, we suggest a table-based design that allows multibit decoding to speed up decompression.
Haris Lekatsas, Marilyn Wolf
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
1999 Hardware/software co-synthesis with memory hierarchies
abstract
This paper introduces the first hardware/software co-synthesis algorithm of distributed real-time systems that optimizes the memory hierarchy along with the rest of the architecture. Memory hierarchies (caches) are essential for modern embedded cores to obtain high performance. They also represents a significant portion of the cost, size and power consumption of many embedded systems. Our algorithm synthesizes a set of real-time tasks with data dependencies onto a heterogeneous multiprocessor architecture that meets the performance constraints with minimized cost. Unlike previous work in co-synthesis, our algorithm not only synthesizes the hardware and software portions of the applications, but also the memory hierarchies. It chooses cache sizes and allocates tasks to caches as part of cosynthesis. The algorithm is built upon a task-level performance model for memory hierarchies. Experimental results, including examples from the literature and results on real-life examples such as an MPEG-2 encoder, show that our algorithm is efficient, and compared with existing algorithms, it can reduce the overall cost of the synthesized system.
Yanbing Li, Marilyn Wolf
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
1999 A methodology and algorithms for the design of hard real-time multitasking ASICs
abstract
Traditional high-level synthesis concentrates on the implementation of a single task (e.g. filter, linear controller, A/D converter). However, many applications—multifunctional embedded controllers intelligent wireless end-points, and DSP and multimedia servers—are defined as sets of several computational tasks. This paper describes new techniques for the synthesis of ASIC implementations that realize multiple computational processes under hard real-time constraints. Our synthesis methodology establishes connections between two important comengineering domains: operating systems and behavioral synthesis. Our hierarchical approach starts from an incompletely-specified preliminary solution and uses, interchangeably, operating system and behavioral synthesis techniques to derive increasingly more detailed and accurate design solutions. We have experimented with both optimal and heuristic algorithms to implement this methodology. The optimal algorithm uses several heuristics to speed up the average run time of an exhaustive branch-and-bound search. Force-directed optimization is the core of the heuristic synthesis method. Analysis of the proposed algorithms and the experiments shows that matching the number of bits and type of operational in taskes assigned to the same application-specific processor was the most important factor in obtaining area-efficient designs.
Miodrag Potkonjak, Marilyn Wolf
ACM Trans. Design Autom. Electr. Syst.2
1999 A circuit-driven design methodology for video signal-processing datapath elements
abstract
The programmable video signal processor (VSP) is an important category of processors for multimedia systems. Programmable video processors combine the flexibility of programmability with special architectural features that improve performance on video processing applications. VSPs are typically multiple processors with several processing elements (PEs) and a parallel memory system. This paper focuses on the architectural design of the PE's in a video processor and shows how technology and circuit parameters influence the structure of the datapath and, hence, the overall architecture of a programmable VSP. We emphasize the need to consider technological and circuit-level issues during the design of a system architecture and present a method whereby the conceptual organization of the PEs-the number of PEs, pipelining of the datapath, size of the register file, and number of register ports-can be evaluated in terms of a target set of applications before a detailed design is undertaken. We use motion-estimation and discrete cosine transform as example applications to illustrate how various technology parameters affect the architectural design choices. We show that the design of the register file and the datapath-pipeline depth can drastically affect PE utilization and, therefore, the number of PEs required for different applications. Our results demonstrate that pursuing the fastest cycle time can greatly increase the silicon area which must be devoted to PEs, due to both increased pipeline latency and reduced register file bandwidth.
Santanu Dutta, Marilyn Wolf
IEEE Trans. Very Large Scale Integr. Syst.2
1998 Code Compression for Embedded Systems
abstract
Memory is one of the most restricted resources in many modern embedded systems. Code compression can provide substantial savings in terms of size. In a compressed code CPU, a cache miss triggers the decompression of a main memory block, before it gets transferred to the cache. Because the code must be decompressible starting from any point (or at least at cache block boundaries), most file-oriented compression techniques cannot be used. We propose two algorithms to compress code in a space-efficient and simple to decompress way, one which is independent of the instruction set and another which depends on the instruction set. We perform experiments on two instruction sets, a typical RISC (MIPS) and a typical CISC (x86) and compare our results to existing file-oriented compression algorithms.
Haris Lekatsas, Marilyn Wolf
DAC2
1998 A multi-resolution video segmentation scheme for wipe transition identification
abstract
This paper presents a new methodology for wipe transition identification. Shot transition detection is an important technique for making videos easier to handle. Due to the wide variety, wipe transition appears to be the most difficult one to be detected among all types of shot transitions. We propose an approach that takes advantage of the production aspect of video. Each video frame is first decomposed into low-resolution and high-resolution components which are analyzed respectively and further recombined together to form a wipe transition detector. In our system, wavelet transformation is used for multi-resolution decomposition.
Hong Heather Yu, Marilyn Wolf
ICASSP2
1998 How will CAD handle billion-transistor systems? (panel)
abstract
No abstract available.
Robert C. Aitken, Jason Cong, Randy Harr, Kenneth L. Shepard, Marilyn Wolf
ICCAD5
1998 Real-time operating systems for embedded computing
abstract
No abstract available.
Serge Hustin, Miodrag Potkonjak, Eric Verhulst, Marilyn Wolf
ICCAD4
1998 Hardware/software co-synthesis with memory hierarchies
abstract
This paper introduces the first hardware/software co-synthesis algorithm of distributed real-time systems that op-timizes memory hierarchy along with the rest of the archi-tecture. Our algorithm synthesize a set of real-time tasks with data dependencies onto a heterogeneous multiproces-sor architecture that meets the performance constraints with minimized cost. Our algorithm chooses cache sizes and al-locates tasks to caches as part of co-synthesis. Experimental results, including examples from the literature and results on an MPEG-2 encoder, show that our algorithm is efficient and compared with existing algorithms, it can reduce the overall cost of the synthesized system. 1
Yanbing Li, Marilyn Wolf
ICCAD2
1998 An Algorithm for Wipe Detection
abstract
The detection of transitions between shots in video programs is an important first step in analyzing video content. The wipe is a frequently used transitional form between shots. Wipe detection is more involved than the detection of abrupt and other gradual transitions because a wipe may take various patterns and because of the difficulty in discriminating a wipe from object and camera motion. In this paper, we propose an algorithm for detecting wipes using both structural and statistical information. The algorithm can effectively detect most wipes used in current TV programs. It uses the DC sequence which can be easily extracted from the MPEG stream without full decompression.
Min Wu 0001, Marilyn Wolf, Bede Liu
ICIP (1)2
1998 Trace-Driven Studies of VLIW Video Signal Processors
abstract
This paper uses extensive traces to explore the parallel architectures of highly programmable Video Signal Processors (VSPs).First we briefly compare some existing architectures.Based on the parallel property of video applications, a general Very Long Instruction Word (VLIW) paradigm is proposed.Then we use a new technique, trace-driven architectural exploration, to evaluate architectures with different parameters such as number of registers, number and type of functional units, and number of pipeline stages, et al.Focusing on video applications, we have analyzed the traces of H.263, MPEG-2 and MPEG-4 video codecs.The results show us which architectural tradeoffs enhance the overall performance in the application domain and how we should balance processor resources among registers, functional units and memory units.
Marilyn Wolf
SPAA2
1998 Efficient Algorithms for Interface Timing Verification
Ti-Yen Yen, Alex Ishii, Albert E. Casavant, Marilyn Wolf
Formal Methods Syst. Des.4
1998 A design study of a 0.25-μm video signal processor
abstract
This paper presents a detailed design study of a high-speed, single-chip architecture for video signal processing (VSP), developed as part of the Princeton VSP Project. In order to define the architectural parameters by examining the area and delay tradeoffs, we start by designing parameterizable versions of key modules, and we perform VLSI modeling experiments in a 0.25 /spl mu/m process. Based on the properties of these modules, we propose a VLIW (very long instruction word) VSP architecture that features 32-64 operations per cycle at clock rates well in excess of 600 MHz, and that includes a significant amount of on-chip memory. VLIW architectures provide predictable, efficient, high performance, and benefit from mature compiler technology. As explained, a VLIW video processor design requires flexible, high-bandwidth interconnect at fast cycle times, and presents some unique VLSI tradeoffs and challenges in maintaining high clock rates while providing high parallelism and utilization.
Santanu Dutta, Kevin J. O'Connor, Marilyn Wolf, Andrew Wolfe
IEEE Trans. Circuits Syst. Video Technol.3
1998 A methodology to evaluate memory architecture design tradeoffs for video signal processors
abstract
Develops a methodology for the design of the memory and the memory-processor communication network in video signal processors. The memory subsystem is the bottleneck of most video computing systems and its design requires evaluating tradeoffs between area, cycle time, and utilization. We emphasize the need to consider technological and circuit-level issues during the design of a system architecture, particularly video signal processing (VSP) systems, and present a systematic method whereby the organization of the memory architecture can be analyzed and its cycle-time approximated before a detailed design is undertaken. We show how variations in sizes and circuit configurations help determine the variations in delay of both memory and network, and how the delay curves, thus determined, can be used to design, compare, and choose from different memory-system architectures; we also describe a technique that can be used to identify the on-chip-off-chip boundary with respect to a hierarchical memory-system design for a memory-intensive VSP module. All of our results are validated via layout and simulation of prototype circuits in two different process technologies. Motion estimation and discrete cosine transform (DCT) being two of the most important tasks in video processing, we use the design of a motion estimator and that of a DCT unit as examples to illustrate the high-level issues in designing the memory architecture for a VSP module. The analysis presented for the motion estimator and the DCT unit can also be applied to other processing blocks belonging to the system.
Santanu Dutta, Marilyn Wolf, Andrew Wolfe
IEEE Trans. Circuits Syst. Video Technol.2
1998 Performance Estimation for Real-Time Distributed Embedded Systems
abstract
Many embedded computing systems are distributed systems: communicating processes executing on several CPUs/ASICs. This paper describes a performance analysis algorithm for a set of tasks executing on a heterogeneous distributed system. Tight bounds are essential to the synthesis and verification of application-specific distributed systems, such as embedded computing systems. Our bounding algorithms are valid for a general problem model: The system can contain several tasks with hard real-time deadlines and different periods; each task is partitioned into a set of processes related by data dependencies. The periods of tasks and the computation times of processes are not necessarily constant and can be specified by a lower bound and an upper bound. Such a model requires a more sophisticated algorithm, but leads to more accurate results than previous work. Our algorithm both provides tighter bounds and is faster than previous methods.
Ti-Yen Yen, Marilyn Wolf
IEEE Trans. Parallel Distributed Syst.2
1997 A Task-Level Hierarchical Memory Model for System Synthesis of Multiprocessors
abstract
This paper introduces the first high-level (task-level)model of hierarchical memories and describes a scheduling andallocation algorithm for system-level synthesis of heterogeneousmultiprocessors. Caches are essential for modern RISC embeddedcores to obtain sustained high performance. However, caches havereceived limited use in priority-driven preemptive real-time systemsdue to the unpredictability of caches-average-case improvementsare of no use in systems with hard deadlines. Program-levelcache models do not take into account preemptions between multipletasks running at multiple rates on embedded cores. Our task-levelmodel of performance in the presence of memory hierarchiesprovides an efficient means to bound the guaranteed memory performanceof tasks running in a multi-rate, multi-tasking environment.Our system synthesis algorithm uses software-based cachepartitioning and reservation techniques to guarantee cache hitsfor some tasks and therefore improve task schedulability. Experimentalresults show that our model significantly improves schedulabilityof real-time tasks and can be evaluated efficiently duringsystem-level synthesis.
Yanbing Li, Marilyn Wolf
DAC2
1997 Hidden Markov model parsing of video programs
abstract
This paper introduces statistical parsing of video programs using hidden Markov models (HMMs). The fundamental units of a video program are shots and transitions (fades, dissolves, etc.). Those units are in turn used to create more complex structures, such as scenes. Parsing a video allows us to recognize higher-level story abstractions. These higher-level story elements can be used to create summarizations of the programs, to recognize the most important parts of a program, and many other purposes. The paper is of interest in cinematography for summarizing programs.
Marilyn Wolf
ICASSP1
1997 Critical technologies and methodologies for systems-on-chips (tutorial)
Wayne Wei-Ming Dai, Howard L. Kalter, Rob Roy, Marilyn Wolf
ICCAD4
1997 An Approach to Network Caching for Multimedia Objects
abstract
Caching is an important mechanism for improving both the performance and operational cost of multimedia networks. This paper presents a new approach to network caching for large multimedia objects, the Caching Tree Algorithm, which is based on previous research regarding the File Allocation Problem. This novel algorithm is both distributed and computationally tractable. We also describe some practical considerations in the application of the algorithm to modern networks.
Michael A. Kozuch, Marilyn Wolf, Andrew Wolfe
ICCD2
1997 Real-Time Operating Systems for Embedded Computing
abstract
The authors survey the state-of-the-art in real-time operating systems (RTOSs) from the system synthesis point of view. RTOSs have a very long research history which provides important theoretical results and useful industrial implementations. Convergence of applications, technology, and market trends of embedded systems implies a strong need for new generation of RTOS. Therefore, new system synthesis problem areas, notably hardware/software co-design and synthesis for systems-on-silicon (SOS), are opening up new avenues for RTOS research and development. The paper starts with a survey of classical academic and industrial RTOS work and continues with a survey of recent results related to co-design and design systems-on-silicon. They conclude by outlining future directions for the SOS RTOS.
Yanbing Li, Miodrag Potkonjak, Marilyn Wolf
ICCD3
1997 Allocation and Data Arrival Design of Hard Real-time Systems
abstract
The paper presents new models for process activation and process scheduling for real-time embedded systems. The authors introduce a realistic, yet high-level input data arrival model which includes both polled and interrupt-driven process activation. They consider the effect of combinations of these process activation styles on a static, priority-based, preemptive scheduler. Given a set of periodic tasks and a set of resources (e.g. processors), a configuration is defined as: i) a mapping of each process to a resource; ii) assignment of priority to each process; and iii) a mapping of each interprocess communication event to either a polled or interrupt-driven implementation. They present a new method which utilizes an exact schedule analysis to determine a configuration which can meet hard real time deadlines subject to a fixed limit on the number of interrupts available per resource. Task graph examples and comparisons are used to validate the method.
David L. Rhodes, Marilyn Wolf
ICCD2
1997 A Hierarchical, Multi-Resolution Method for Dictionary-Driven Content-Based Image Retrieval
abstract
A new methodology of image and keyframe retrieval for image and video databases is presented. In contrast to previous approaches, which only support similarity retrieval and direct image manipulation on-line, our system performs tagging off-line using neural network algorithm and answers query on-line using only the tags. This new method is more appealing in meeting user's needs than only emphasizing what current technology can offer. Experiments on our CAETI (Computer Assisted Education & Training Initiative) IML (Internet Multimedia Library) show that this model gives high quality query results with fast on-line performance. This visual search system is available at http:/www.videolib.princeton.edu/test/retrieve.
Hong Heather Yu, Marilyn Wolf
ICIP (2)2
1997 Redundancy Removal during High-Level Synthesis Using Scheduling Don't-Cares
Marilyn Wolf
J. Electron. Test.1
1997 Unifiable scheduling and allocation for minimizing system cycle time
abstract
This paper describes a new scheduling and allocation algorithm which optimizes a datapath-controller system for clock cycle time. The cycle time of a VLSI system depends not only on the characteristics of the datapath and controller in isolation but also on the interactions between them. A datapath may impose both arrival time constraints on controller inputs and departure time constraints on controller outputs. Late-arriving controller inputs may be generated by complex datapath functions, such as ALU carry-out, while early-departure controller outputs may be required to control slow datapath units. If the controller is not designed taking into account arrival and departure times, it may unnecessarily put control logic on the critical timing path. Our synthesis heuristic, which can be used in conjunction with other scheduling heuristics, identifies critical interactions between datapath and controller and reallocates/reschedules them to reduce system cycle time during high-level synthesis. Experimental results show that a unifiable scheduling and allocation (USA) can substantially improve system cycle time with only small area penalties.
Steve C.-Y. Huang, Marilyn Wolf
IEEE Trans. Very Large Scale Integr. Syst.2
1997 An architectural co-synthesis algorithm for distributed, embedded computing systems
abstract
Many embedded computers are distributed systems, composed of several heterogeneous processors and communication links of varying speeds and topologies. This paper describes a new, heuristic algorithm which simultaneously synthesizes the hardware and software architectures of a distributed system to meet a performance goal and minimize cost. The hardware architecture of the synthesized system consists of a network of processors of multiple types and arbitrary communication topology; the software architecture consists of an allocation of processes to processors and a schedule for the processes. Most previous work in co-synthesis targets an architectural template, whereas this algorithm can synthesize a distributed system of arbitrary topology. The algorithm works from a technology database which describes the available processors, communication links, I/O devices, and implementations of processes on processors. Previous work had proposed solving this problem by integer linear programming (ILP); our algorithm is much faster than ILP and produces high-quality results.
Marilyn Wolf
IEEE Trans. Very Large Scale Integr. Syst.1
1996 Heuristic techniques for synthesis of hard real-time DSP application specific systems
abstract
We introduce an approach for the design and optimization of ASIC implementations which realize multiple computational tasks under hard real-time constraints. The approach designs a multitask ASIC by combining techniques from hard real-time scheduling and behavioral synthesis. The key component of the methodology is the successive multiresolution synthesis technique. The technique starts from an incompletely specified preliminary solution and uses interchangeably operating systems and behavioral synthesis tools to derive increasingly more detailed and complete design solutions. The effectiveness of the optimization algorithms is demonstrated on several multiple task designs.
Miodrag Potkonjak, Marilyn Wolf
ICASSP2
1996 Key frame selection by motion analysis
abstract
This paper describes a new algorithm for identifying key frames in shots from video programs. We use optical flow computations to identify local minima of motion in a shot-stillness emphasizes the image for the viewer. This technique allows us to identify both gestures which are emphasized by momentary pauses and camera motion which links together several distinct images in a single shot. Results show that our algorithm can successfully select several key frames from a single complex shot which effectively summarize the shot.
Marilyn Wolf
ICASSP1
1996 New Challenges for Video Servers: Performance of Non-Linear Applications under User Choice
abstract
This paper outlines a framework for classifying video application types according to linearity and user-choice response constraint. It then presents the first analyses of video server performance under various non-linear video application loads.
Michael A. Kozuch, Marilyn Wolf, Andrew Wolfe
ICCD2
1996 A flexible parallel architecture adapted to block-matching motion-estimation algorithms
abstract
This paper describes a novel architecture that offers the flexibility of implementing widely varying motion-estimation algorithms. To achieve real-time performance, we employ multiple processing elements (PE's) which communicate with multiple memory banks via a multistage interconnection network. Three different block-matching algorithms-full search, three-step search, and conjugate-direction search-have been mapped onto this architecture to illustrate its programmability. We schedule the desired operations and design the required data-flow in such a way that processor utilization is high and memory bandwidth is at a feasible level. The details regarding the flow of the pixel data and the scheduling and allocation of the desired ALU operations (which pixels are processed on which processors in which clock cycles) are described in the paper. We analyze the performance of the proposed architecture for several different interconnection networks and data-memory organizations.
Santanu Dutta, Marilyn Wolf
IEEE Trans. Circuits Syst. Video Technol.2
1996 Object-oriented cosynthesis of distributed embedded systems
abstract
This article describes a new hardware-software cosynthesis algorithm that takes advantage of the structure inherent in an object-oriented specification. The algorithm creates a distributed system implementation with arbitrary topology, using the object-oriented structure to partition functionality in addition to scheduling and allocating processes. Process partitioning is an especially important optimization for such systems because the specification will not, in general, take into account the process structure required for efficient execution on the distributed engine. The object-oriented specification naturally provides both coarse-grained and fine-grained partitions of the system. Our algorithm uses that multilevel structure to guide synthesis. Experimental results show that our algorithm takes advantage of the object-oriented specification to quickly converge on high-quality implementations.
Marilyn Wolf
ACM Trans. Design Autom. Electr. Syst.1
1996 An efficient graph algorithm for FSM scheduling
abstract
This paper presents a new algorithm for scheduling control-dominated designs during high-level synthesis. Our algorithm can schedule systems with arbitrary control flow, including conditional branches and multiple loops. It can handle both upper bound and lower bound timing constraints. The timing constraints can cross basic block boundaries, span different iterations of a loop, and form interlocking cycles in the control flow. A scheduling problem is described by the behavior finite-state machine model, an automaton model for the behavioral specification and synthesis of control-dominated systems. We optimize the performance of the produced digital circuit implementation by minimizing the execution time of each state transition in the state transition graph. The finite-state machines (FSM) scheduling algorithm is based on previous work on cylindrical layout compaction; we extend that work to handle upper bound constraints, allow multiple loops, and not require an initial feasible solution. Experimental results for examples derived from real designs and benchmark descriptions demonstrate that the algorithm can handle complex combinations of constraints very efficiently.
Ti-Yen Yen, Marilyn Wolf
IEEE Trans. Very Large Scale Integr. Syst.2
1995 CAD challenges in multimedia computing
abstract
This tutorial surveys the present and future of multimedia computing systems and outlines new challenges for CAD presented by these systems. Multimedia computing is a challenging domain for several reasons: it requires both high computation rates and memory bandwidth; it is a multirate computing problem; and requires low-cost implementations for high-volume markets. As a result, the design of multimedia computing systems introduces new challenges for CAD at all levels of abstraction, ranging from layout to system design. After surveying the nature of the multimedia computing problem, we examine two experiences in multimedia computer design from a CAD perspective: the design of VLSI systems-on-chips for multimedia: and the successive refinement of an application from software to a high-volume chip using advanced CAD synthesis tools.
Paul E. R. Lippens, Vijay Nagasamy, Marilyn Wolf
ICCAD3
1995 Cost optimization in ASIC implementation of periodic hard-real time systems using behavioral synthesis techniques
abstract
Modern applications are often defined as sets of several computational tasks. This paper presents a synthesis algorithm for ASIC implementations which realize multiple computational tasks under hard real-time deadlines. The algorithm analyzes constraints imposed by task sharing as well as the traditional datapath synthesis criteria. In particular we demonstrated an efficient technique to combine rate-monotonic scheduling, a widely used hard real-time systems scheduling discipline, with estimations and scheduling and allocation algorithms. Matching the number of bits in tasks assigned to the same processor was the most important factor in obtaining good designs. We have demonstrated the effectiveness of our algorithms on several multiple-task examples.
Miodrag Potkonjak, Marilyn Wolf
ICCAD2
1995 Communication synthesis for distributed embedded systems
abstract
Communication synthesis is an essential step in hardware-software co-synthesis: many embedded systems use custom communication topologies and the communication links are often a significant part of the system cost. This paper describes new techniques for the analysis and synthesis of the communication requirements of embedded systems during co-synthesis. Our analysis algorithm derives delay bounds on communication in the system given an allocation of messages to links. This analysis algorithm is used by our synthesis algorithm to choose the required communication links in the system and assign interprocess communication to the links. Experimental results show that our algorithm finds good communication architectures in small amounts of CPU time.
Ti-Yen Yen, Marilyn Wolf
ICCAD2
1995 VLSI issues in memory-system design for video signal processors
abstract
This paper addresses the design of memory-system architectures for video signal processors. The memory subsystem is the bottleneck of most video computing systems and demands a careful analysis of the design tradeoffs related to area, cycle time, and utilization. We emphasize the need to consider technological and circuit-level issues during the design of a system architecture, particularly that of a video processor, and present a method whereby the conceptual organization of the memory architecture can be evaluated before a detailed design is undertaken. Our analysis suggests that the organization of an efficient memory hierarchy for video signal processors is different from the register-cache based hierarchy of general-purpose programmable microprocessors.
Santanu Dutta, Marilyn Wolf, Andrew Wolfe
ICCD2
1995 Performance estimation for real-time distributed embedded systems
abstract
Many embedded computing systems are distributed systems: communicating processes executing on several CPUs/ASICs connected by communication links. This paper describes a new, efficient analysis algorithm to derive tight bounds on the execution time required for an application task executing on a distributed system. Tight bounds are essential to cosynthesis algorithms. Our bounding algorithms are valid for a general problem model: the system can contain several tasks with different periods; each task is partitioned into a set of processes related by data dependencies; the periods and the computation times of processes are bounded but not necessarily constant. Experimental results show that our algorithm can find tight bounds in small amounts of CPU time.
Ti-Yen Yen, Marilyn Wolf
ICCD2
1995 An Automaton Model for Scheduling Constraints in Synchronous Machines
abstract
We present a finite-state model for scheduling constraints in digital system design. We define a two-level hierarchy of finite-state machines: a behavior FSM's input and output events are partially ordered in time; a register-transfer FSM is a traditional FSM whose inputs and outputs are totally ordered in time. Explicit modeling of scheduling constraints is useful for both high-level synthesis and verification-we can explicitly search the space of register-transfer FSM's which implement a desired schedule. State-based models for scheduling are particularly important in the design of control-dominated systems. This paper describes the BFSM I model, describes several important operations and algorithms on BFSM's and networks of communicating BFSM's, and illustrates the use of BFSM's in high-level synthesis.
Andrés Takach, Marilyn Wolf, Miriam Leeser
IEEE Trans. Computers2
1995 Asymptotic limits of video signal processing architectures
abstract
This paper analyzes the effects of technology scaling on video signal processing (VSP) architectures. We evaluate the processor, the memory, and the interconnect delays in terms of sophisticated delay models (that take into account deep-sub-micron device characteristics) and study how the response times of these logic components are affected when the feature sizes scale down. Equations for gate and interconnect delays, as functions of process scaling, are derived and the impact of these results examined in the context of heavily pipelined architectures, architectures featuring crossbar interconnection networks, and architectures whose performance is dominated by memory bandwidth. Architectural parameters such as clock skew, clock frequency, memory interleaving, memory efficiency, and average waiting times are analyzed in the light of the scaling behavior of the gate and the interconnect delays. In the context of scaling of interconnection lines and memory modules, we also highlight how the transmission-line characteristics of long lines are affected by technology scaling and how the delay associated with the memory subsystem-both the memory interleaving and the memory interconnect network-can be a potential bottleneck for the system's speed of operation. It is likely that sophisticated compilation and scheduling techniques must be employed along with architectural optimizations to achieve maximum system performance and ensure that the final hardware-software configuration does not overload the processor-memory communication.
Santanu Dutta, Marilyn Wolf
IEEE Trans. Circuits Syst. Video Technol.2
1995 Scheduling constraint generation for communicating processes
abstract
This paper describes a new algorithm for generation of scheduling constraints in networks of communicating processes. Our model of communication intertwines the schedules of the machines in the network: timing constraints of a machine may affect the schedules of machines communicating with it. This model of communication facilitates the modular specification of timing constraints. A feasible solution to the set of constraints generated gives a schedule for each machine in the network such that all internal constraints of each machine are satisfied and communication between machines is statically coordinated whenever possible. Static scheduling of communication saves on the cost of handshake associated with dynamic synchronization. Our algorithm can handle complex, state-dependent and cyclic timing constraints. Experimental results show that our algorithm is both effective and efficient.>
Andrés Takach, Marilyn Wolf
IEEE Trans. Very Large Scale Integr. Syst.2
1994 Asymptotic Limits of Video Signal Processing Architectures
abstract
This paper investigates the effects of technology scaling on video signal processing (VSP) architectures. We evaluate the processor, the memory, and the interconnect delays using RC models and study how the response times of these logic components scale with feature size. Architectural parameters such as clock skew, clock frequency, memory interleaving, memory efficiency, and average waiting times are analyzed in the light of the scaling behaviors of the above components.>
Santanu Dutta, Marilyn Wolf
ICCD2
1994 Hardware-software co-design of embedded systems
abstract
This paper surveys the design of embedded computer systems, which use software running on programmable computers to implement system functions. Creating an embedded computer system which meets its performance, cost, and design time goals is a hardware-software co-design problem-the design of the hardware and software components influence each other. This paper emphasizes a historical approach to show the relationships between well-understood design problems and the as-yet unsolved problems in co-design. We describe the relationship between hardware and software architecture in the early stages of embedded system design. We describe analysis techniques for hardware and software relevant to the architectural choices required for hardware-software co-design. We also describe design and synthesis techniques for co-design and related problems.>
Marilyn Wolf
Proc. IEEE1
1994 Performance-driven synthesis in controller-datapath systems
abstract
This paper describes new algorithms which combine state assignment and pipelining to perform timing-driven synthesis. A cycle time requirement for a controller-datapath system must be satisfied under fixed arrival times of datapath outputs and departure times of datapath inputs. As a result, the controller's design must take into account not only cycle time of the FSM in isolation, but also input arrival time specifications and output departure time requirements. Most state assignment methods minimize area; moreover, state assignment alone may not be sufficient to eliminate all delay bottlenecks. Performance-Driven Synthesis (PDS) applies both high-level and sequential optimizations to meet a cycle time requirement: we use new don't-care assignment algorithms to minimize the delays of FSM output signals on critical paths by reducing their dependencies on late-arriving FSM primary input signals; we also use new pipelining algorithms to break critical paths which cannot be fixed by state assignment. Experimental results show that PDS improves delays with little area overhead.>
Steve C.-Y. Huang, Marilyn Wolf
IEEE Trans. Very Large Scale Integr. Syst.2
1993 Behavioral Synthesis of Highly Testable Data Paths under the Non-Scan and Partial Scan Environments
abstract
Behavioral synthesis tools which only optimize area and performance can easily produce a hard-to-test architecture. In this paper, we propose a new behavioral synthesis algorithm for testability which reduces sequential loop size while minimizing area. The algorithm considers two levels of testability synthesis: synthesis for non-scan, which assumes no test strategy beforehand; and synthesis for partial scan, which uses the available scan information during resource allocation. Experimental results show that in almost all the cases our algorithm can synthesize benchmarks with a very high fault coverage in a small amount of test generation time, using the fewest registers and functional modules. Comparisons are also made with other behavioral synthesis algorithms which disregard testability in order to establish the efficacy of our approach.
Tien-Chien Lee, Niraj K. Jha, Marilyn Wolf
DAC3
1993 Embedded Systems and Hardware-Software Co-Design: Panacea or Pandora's Box? (Panel Abstract)
abstract
An embedded computer system (or simply embedded system) is a digital system which uses a microprocessor running software to implement some or all of its functions. Hardware-software co-design is any design technique for the creation of systems with both hardware and software components. Embedded systems are clearly important---microprocessors are used in products ranging from microwave ovens to laser printers to engine controllers. This panel considers what makes embedded system design difficult today and how CAD and CASE can help. Embedded system must be designed under sever constraints. On the one hand, they must often meet performance goals to be viable in the marketplace. On the otherhand, manufacturing cost is important---microprocessors are often used to save money, and the designer wastes money by using a larger processor than necessary. Time-to-market is often critical, since companies often believe (perhaps erroneously) that software implementations will be finished faster. Finally, embedded systems must often be expandable---companies expect to roll out new, improved versions of products by taking advantage of the existing hardware platform and software base. The design of an embedded system may require delicate balancing of hardware and software resource requirements. Performance requirements may force some operations to be done in custom hardware. In current design practice, hardware-software partitioning is often done too early, leading to suboptimal design times and delays in software development. Even if no custom hardware is required, the embedded system designer must choose a hardware architecture on which the software is to be executed. The designer must choose the size and number of CPUs, memory size, peripherals, and so forth. Hardware and software elements must be designed simultaneously, then integrated into a complete system. Concurrent design techniques help to keep track of the design and ensure that constraints are satisfied. Embedded system design practice today is about where VLSI design practice was in the late 1970's. Complex chips were designed before Mead and Conway, but with only the vaguest notion of design abstraction or methodology and with primitive tools. The VLSI revolution helped designers understand their chip designs more abstractly, which led to the development of methodologies and tools which vastly improved designer productivity. Similarly, embedded systems are designed today with only the crudest of tools and with very little learning from one system to the next. We need to understand the common characteristics of embedded systems and develop tools which automate specific steps in the embedded system design process. Only by rationalizing and automating embedded system design can we take advantage of the vast increases in programmable computing power delivered by VLSI.
Marilyn Wolf
DAC1
1993 Scheduling a minimum dependence in FSMs
abstract
We present a scheduling algorithm for controllers which reduces cycle time in datapath-controller systems. The algorithm chooses a schedule to reduce the primary output's dependence on late-arriving inputs - a minimum-dependence structure leads to a shorter system cycle time. Our algorithm accurately predicts from a behavioral description FSM properties which are correlated with fast logic. Experimental results show that the algorithm improves delay in critical outputs and control step-cycle time product.
Steve C.-Y. Huang, Marilyn Wolf
ICCAD2
1993 Optimal Scheduling of Finite-State Machines
abstract
The paper describes an algorithm for solving scheduling problems which contain multiple, interlocking cycles, such as scheduling constraints in state transition graphs. This algorithm is based on previous work on toroidal compaction but introduces three significant improvements: it allows the designer to use upper bound or equality constraints; it does not require an initial feasible solution; and it can handle multiple loops and conditional branches in the constraint system. Experimental results demonstrate the algorithm's effectiveness.>
Ti-Yen Yen, Marilyn Wolf
ICCD2
1993 A Conditional Resource-Sharing Method for Behavior Synthesis of Highly- Testable Data Paths
abstract
Existing conditional resource sharing methods using in behavioral synthesis focus on area and performance optimization and do not consider testability. This paper extends our previous work to handle conditional branches. A hierarchical control-data flow graph (HCDFG) is used to model the system behavior. A postorder traversal of the HCDFG is employed to reduce sequential depths and loops for testability synthesis. Experimental results for the benchmarks show that our method, with no a priori test strategy assumption, can achieve higher fault coverage in shorter test generation time than an algorithm which disregards testability, and, with partial scan test assumption, can have high testability with fewer scan registers than some design-for-test methods.>
Tien-Chien Lee, Niraj K. Jha, Marilyn Wolf
ITC3
1993 FSM decomposition for pipelined data
Marilyn Wolf
Integr.1
1992 The Princeton University Behavioral Synthesis System
Marilyn Wolf, Andrés Takach, Chun-Yao Huang, Richard Manno, Ephrem Wu
DAC1
1992 Behavioral synthesis for easy testability in data path scheduling
abstract
A data path scheduling algorithm to improve testability without assuming any particular test strategy is presented. A scheduling heuristic for easy testability, based on previous work on data path allocation for testability, is introduced. A mobility path scheduling algorithm to implement this heuristic while also minimizing area is developed. Experimental results on benchmark and example circuits show high fault coverage, short test generation time, and little or no area overhead.>
Tien-Chien Lee, Marilyn Wolf, Niraj K. Jha
ICCAD2
1992 The Future of Embedded System Design
abstract
As the scope of application of embedded CPUs changes, the design methodologies used to create embedded systems must keep pace. The authors explore what is understood of embedded system design, what problems need to be solved, and novel technologies that may help solve those problems.>
James H. Aylor, Raúl Camposano, Michael A. Schuette, Marilyn Wolf, Nam Sung Woo
ICCD4
1992 Behavioral Synthesis for Easy Testability in Data Path Allocation
abstract
The first behavioral synthesis scheme for improving testability in data path allocation independent of test strategy is presented. The authors propose two behavioral synthesis-for-test heuristics: improve observability and controllability of registers, and reduce sequential depth between registers. Also presented are algorithms that optimize a behavior-level design using these two criteria while minimizing area. Experimental results for benchmark circuits synthesized by the author's experimental system. PHITS, show that these methods give a high fault coverage in small amounts of CPU time at a low area overhead.>
Tien-Chien Lee, Marilyn Wolf, Niraj K. Jha, John M. Acken
ICCD2
1992 Tutorial on Embedded System Design
abstract
Typical embedded systems are described briefly. Design and debugging methodologies for such systems are reviewed. Programming tools are discussed.>
Marilyn Wolf, Ernest Frey
ICCD1
1992 Object-oriented Implementation Issues in an Experimental CAD System
abstract
Abstract This case study of object‐oriented program design illustrates two limitations of object‐oriented programming languages. Existing object‐oriented languages do not have good facilities to support two key program design problems: the definition of composite objects, or data structures that include sets of related subobjects; and the specification and run‐time management of temporary data structures required to implement efficient algorithms. Both composite objects and temporary data structures are important to the construction of a wide variety of programs. We use the design and implementation of an interactive computer‐aided design system to describe how the limitations of present object‐oriented languages complicate the design of composite objects and temporary data structures.
Marilyn Wolf
Softw. Pract. Exp.1
1991 A framework for industrial layout generators
abstract
The MACLOG family of layout generators creates parameterized custom layouts for many of the cells in the AT&T Cell Library. The authors describe not the MACLOG generators themselves, but the MACLOG framework that is the base of all the generators. The MACLOG generators are built on an object-oriented framework. Though object-oriented design techniques have been described in the literature, the MACLOG framework is one of the first such frameworks used to build an industrial-quality layout generator set. The attempt to build an object-oriented framework structure around the module types results in excessive code duplication and a hard-to-maintain structure. The authors found the service hierarchy-the types of information provided to the user by the generator-often to be a more effective axis for decomposition of functions in the framework. The authors describe the MACLOG framework and detail experiments with framework structure.>
Wayne Bower, Carl Seaquist, Marilyn Wolf
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
1990 A Framework for Industrial Layout Generators
abstract
The MACLOG layout generator set is built on an object-oriented framework. Most frameworks are built around the hierarchy of module types. We found the service hierarchy—the types of information provided to the user by the generator—made the framework and generators easier to design and modify. This paper describes the MACLOG framework and details our experience with framework structure.
Wayne Bower, Carl Seaquist, Marilyn Wolf
DAC3
1990 The FSM Network Model for Behavioral Synthesis of Control-Dominated Machines
abstract
Since many ASICs are dominated by control functions, control-dominated architectures form an important domain for behavioral synthesis. We propose modeling control-dominated architectures during behavioral synthesis as networks of communicating FSMs—the model more directly reflects behavior and allows more accurate cost estimation, especially for control, than do traditional data-directed representations for control-dominated machines. We show how to implement a number of important compiler optimizations on the FSM network model.
Marilyn Wolf
DAC1
1990 An Algorithm for Nearly-Minimal Collapsing of Finite-State Machine Networks
abstract
An algorithm is presented which simultaneously generates the Cartesian product of a network of finite-state machines and minimizes the resulting product machine. The algorithm can generate collapsed machines, removing a large set of redundant states on the fly, in CPU times comparable to the time required for simple Cartesian product collapsing. The algorithm makes it practical to generate and analyze a much larger class of collapsed FSM networks.>
Marilyn Wolf
ICCAD1
1990 Issues in synthesis of board-level systems
abstract
Board layout is fully or partially automated, thanks to placement and routing systems. The higher levels of the design process-component selection, architectural design, software development-are still done manually. The authors' research program at Princeton University concentrates on synthesis algorithms for high-level board design tasks. Based on case studies of system designs made both within the university and in industry, they have chosen a skeleton-based methodology for board specification and synthesis and are researching a variety of problems posed by this approach. A skeleton-based synthesis methodology assumes that a few key design decisions-major architectural choices and key component selections-are part of the board specification. The synthesis system's task is to complete the board design, based on a functional description of the board plus constraints on speed, size, power consumption, and interface behavior. A skeleton-based synthesis methodology is realistic and effective.>
Andrea S. LaPaugh, Marilyn Wolf
RSP2
1989 How to build a hardware description and measurement system on an object-oriented programming language
abstract
Techniques are described for applying the mechanisms of object-oriented programming languages to hardware description. Some object-oriented language mechanisms, like inheritance, directly simplify CAD (computer-aided design) programs; others, like data abstraction, allow more powerful CAD mechanisms based on them to be created. The author describes: how to extend class inheritance and to integrate it with procedural construction to simplify the description of hardware; how to create measurement methods than can measure a module whose components are described at different levels of abstraction; and how to implement a consistency-maintenance engine that ensures the consistency of the data kept for the design. The author has implemented these features in Fred, an object-oriented modeling system for VLSI modules. Fred is implemented in Flavors, an object-oriented extension of Lisp. The author also discusses how to implement its features in other languages.>
Marilyn Wolf
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
1989 Addendum to 'A kernel-finding state assignment algorithm for multi-level logic'
abstract
Two new sets of results are presented that extend work on state assignment for multilevel logic implementation previously reported by the authors (see Proc. IEEE/ACM 25th Design Autom Conf., p.433-8, 1988). Results are presented for several state assignment algorithms that were compared using an improved logic optimization method. The resulting logic implementations were smaller in absolute size and showed considerably more variation in random state assignment than in previous experiments. A new set of experiments that compare state assignment implementations in two-level and multilevel implementations using randomly generated state assignments is reported. These experiments show that, for the benchmark set used, a state assignment that gives a good two-level implementation also gives a good multilevel implementation.>
Marilyn Wolf, Kurt Keutzer, Janaki Akella
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
1988 What Is a Design Automation Framework, Anyway? (panel)
Marilyn Wolf
DAC1
1988 A Kernel-Finding State Assignment Algorithm for Multi-Level Logic
Marilyn Wolf, Kurt Keutzer, Janaki Akella
DAC1
1988 Experiments in logic optimization
abstract
A number of experiments have been conducted to answer several important questions about logic optimization and its application to high-level synthesis. The experiments are summarised and show that: logic optimization is competitive with manual design; stronger optimization methods give somewhat better average results (10%-30%) at much greater computational cost (8* and more); fast logic optimization methods can be used to estimate the average results of the more powerful, costly methods; and literal count is a good estimator of area before routing for standard cell designs.>
Michael R. Lightner, Marilyn Wolf
ICCAD2
1988 Anatomy of a Hardware Compiler
abstract
Programming-language compilers generate code targeted to machines with fixed architectures, either parallel or serial. Compiler techniques can also be used to generate the hardware on which these programming languages are executed. In this paper we demonstrate that many compilation techniques developed for programming languages are applicable to compilation of register-transfer hardware designs. Our approach uses a typical syntax-directed translation → global optimization → local optimization → code generation → peephole optimization method. In this paper we will describe ways in which we have both followed and diverged from traditional compiler approaches to these problems and compare our approach to other compiler oriented approaches to hardware compilation.
Kurt Keutzer, Marilyn Wolf
PLDI2
1988 Algorithms for optimizing, two-dimensional symbolic layout compaction
abstract
A set of algorithms that implement a technique called Supercompaction is described for two-dimensional compaction layouts. The algorithms minimize a one-dimensional objective function (pitch) by moving objects in the layout in two dimensions. The objective function can be monotonically reduced to a locally minimal value, greatly simplifying search. The algorithms can change the layout by simple motion of components or by automatic jog introduction. Experiments shows that: (1) the number of iterations required to reach a locally optimal layout is very small; (2) the time per iteration is competitive with other compaction techniques; (3) supercompacted layouts are smaller in pitch and area than those produced by existing one-dimensional algorithms; and (4) the results of supercompaction are more predictable than those of existing one-dimensional methods.>
Marilyn Wolf, Robert G. Mathews, John A. Newkirk, Robert W. Dutton
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
1986 Knowledge Engineering Issues in VLSI Synthesis
Marilyn Wolf, Thaddeus J. Kowalski, Michael C. McFarland
AAAI1
1986 An object-oriented, procedural database for VLSI chip planning
abstract
This paper describes Fred, a procedural database to support the architectural and floorplanning phases of VLSI design. The database holds hierarchical descriptions of modules that can be used to construct chips. Procedures provided by Fred can be used to compute the physical, electrical, timing, clocking, and functional properties of modules. The user can describe unimplemented modules with default values or approximation functions for these properties. The database can be searched by testing groups of modules with user-defined functions. Fred is implemented in Flavors, an object-oriented extension of Lisp; the object-oriented implementation aids both the designs of the database and the specification of modules for the database.
Marilyn Wolf
DAC1