EDBT 2026 Demo / reviewers in the wild / expert
Vivienne Sze
dblp:49/3209
· DBLP profile ↗
53ranked-venue papers
8as first author
17since 2021 · last 2024
0000-0003-4841-3990ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 26 · 6 first-author · 3 since 2021Systems, architecture and hardware · 15 · 1 first-author · 7 since 2021Artificial intelligence and machine learning · 10 · 4 since 2021Software engineering, systems software and programming languages · 8 · 6 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 2 first-author · 1 since 2021Computer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | ClearDepth: Addressing Depth Distortions Caused By Eyelashes For Accurate Geometric Gaze Estimation On Mobile DevicesabstractGaze estimation on mobile devices offers myriad opportunities such as human visual behavior tracking in natural environments, pervasive attentive user interfaces, and gaze-based wheelchair control. Geometric gaze estimation models are lightweight and thus suitable for mobile devices. However, they are sensitive to measurement noise and suffer from low accuracy in settings where such noise is present. Eyelashes are a critical source of noise for geometric models since they can obstruct the depth sensor’s view of the eyes, which distorts the pupil depth and hence 3D pupil coordinate that the model relies on. In the context of mobile devices, eyelashes are particularly problematic since the internal depth sensor’s location cannot readily be adjusted to minimize eyelash interference. We propose ClearDepth, a technique that derives undistorted pupil depth values using a single prerecorded eyelash-free depth map, while allowing free head movements of the user. ClearDepth accomplishes this by continuously transforming the eyelash-free depth map using head pose information to align it with the user’s current eye position in space. We demonstrate the significant accuracy benefits of ClearDepth using data recorded on iPads. Our approach is an important step towards enabling accurate geometric gaze estimation on mobile devices. Jamie Koerner, Vivienne Sze |
ICIP | 2 |
| 2024 | Architecture-Level Modeling of Photonic Deep Neural Network AcceleratorsabstractPhotonics is a promising technology to accelerate Deep Neural Networks as it can use optical interconnects to reduce data movement energy and it enables low-energy, high-throughput optical-analog computations. To realize these benefits in a full system (accelerator + DRAM), designers must ensure that the benefits of using the electrical, optical, analog, and digital domains exceed the costs of converting data between domains. Designers must also consider system-level energy costs such as data fetch from DRAM. Converting data and accessing DRAM can consume significant energy, so to evaluate and explore the photonic system space, there is a need for a tool that can model these full-system considerations. In this work, we show that similarities between Compute-in-Memory (CiM) and photonics let us use CiM system modeling tools to accurately model photonics systems. Bringing modeling tools to photonics enables evaluation of photonic research in a full-system context, rapid design space exploration, co-design, and comparison between systems. Using our open-source model, we show that cross-domain conversion and DRAM can consume a significant portion of photonic system energy. We then demonstrate optimizations that reduce conversions and DRAM accesses to improve photonic system energy efficiency by up to 3 x. Tanner Andrulis, Gohar Irfan Chaudhry, Vinith M. Suriyakumar, Joel S. Emer, Vivienne Sze |
ISPASS | 5 |
| 2024 | CiMLoop: A Flexible, Accurate, and Fast Compute-In-Memory Modeling ToolabstractCompute-In-Memory (CiM) is a promising solution to accelerate Deep Neural Networks (DNNs) as it can avoid energy-intensive DNN weight movement and use memory arrays to perform low-energy, high-density computations. These benefits have inspired research across the CiM stack, but CiM research often focuses on only one level of the stack (i.e., devices, circuits, architecture, workload, or mapping) or only one design point (e.g., one fabricated chip). There is a need for a full-stack modeling tool to evaluate design decisions in the context of full systems (e.g., see how a circuit impacts system energy) and to perform rapid early -stage exploration of the CiM co-design space. To address this need, we propose CiMLoop: an open-source tool to model diverse CiM systems and explore decisions across the CiM stack. CiMLoop introduces (1) a flexible specification that lets users describe, model, and map workloads to both circuits and architecture, (2) an accurate energy model that captures the interaction between DNN operand values, hardware data representations, and analog/digital values propagated by circuits, and (3) a fast statistical model that can explore the design space orders-of-magnitude more quickly than other high-accuracy models. Using CiMLoop, researchers can evaluate design choices at different levels of the CiM stack, co-design across all levels, fairly compare different implementations, and rapidly explore the design space. Tanner Andrulis, Joel S. Emer, Vivienne Sze |
ISPASS | 3 |
| 2024 | Gemino: Practical and Robust Neural Compression for Video Conferencing
Vibhaalakshmi Sivaraman, Pantea Karimi, Vedantha Venkatapathy, Mehrdad Khani Shirkoohi, Sadjad Fouladi, Mohammad Alizadeh, Frédo Durand, Vivienne Sze |
NSDI | 8 |
| 2024 | GMMap: Memory-Efficient Continuous Occupancy Map Using Gaussian Mixture ModelabstractEnergy consumption of memory accesses dominates the compute energy in energy-constrained robots, which require a compact 3-D map of the environment to achieve autonomy. Recent mapping frameworks only focused on reducing the map size while incurring significant memory usage during map construction due to the multipass processing of each depth image. In this work, we present a memory-efficient continuous occupancy map, named GMMap, that accurately models the 3-D environment using a Gaussian mixture model (GMM). Memory-efficient GMMap construction is enabled by the single-pass compression of depth images into local GMMs, which are directly fused together into a globally-consistent map. By extending Gaussian Mixture Regression (GMR) to model unexplored regions, occupancy probability is directly computed from Gaussians. Using a low-power ARM Cortex A57 CPU, GMMap can be constructed in real time at up to 60 images/s. Compared with prior works, GMMap maintains high accuracy while reducing the map size by at least 56%, memory overhead by at least 88%, dynamic random-access memory (DRAM) access by at least 78%, and energy consumption by at least 69%. Thus, GMMap enables real-time 3-D mapping on energy-constrained robots. Peter Zhi Xuan Li, Sertac Karaman, Vivienne Sze |
IEEE Trans. Robotics | 3 |
| 2023 | RAELLA: Reforming the Arithmetic for Efficient, Low-Resolution, and Low-Loss Analog PIM: No Retraining Required!abstractProcessing-In-Memory (PIM) accelerators have the potential to efficiently run Deep Neural Network (DNN) inference by reducing costly data movement and by using resistive RAM (ReRAM) for efficient analog compute. Unfortunately, overall PIM accelerator efficiency is limited by energy-intensive analog-to-digital converters (ADCs). Furthermore, existing accelerators that reduce ADC cost do so by changing DNN weights or by using low-resolution ADCs that reduce output fidelity. These strategies harm DNN accuracy and/or require costly DNN retraining to compensate. Tanner Andrulis, Joel S. Emer, Vivienne Sze |
ISCA | 3 |
| 2023 | LoopTree: Enabling Exploration of Fused-layer Dataflow AcceleratorsabstractMany accelerators today process deep neural networks layer by layer. As a consequence of this processing style, every intermediate feature map incurs expensive off-chip transfers. Layer fusion eliminates off-chip transfers of intermediate results, leading to better latency and energy efficiency. Prior works have explored only subsets of the fused-layer design space, looking only at a particular choice of tiling, scheduling, and buffering strategy. Their architectural models are also tailored for their proposed datatlow. The lack of a unified, systematic representation of designs and a versatile evaluation method has prevented thorough exploration of the design space. To enable systematic exploration of this design space, we present LoopTree, a framework for describing and evaluating any design in our expanded fused-layer datatlow design space. With a case study, we explore new designs to show that exploring our larger design space uncovers more efficient designs, especially for recent workloads with diverse layer types. Our design achieves 2.5× speedup and 2× lower energy compared to an optimized layer-by-layer design. Compared to a state-of-the-art fused-layer design, we match latency and energy while using 25% less onchip buffer space. Michael Gilbert, Yannan Nellie Wu, Angshuman Parashar, Vivienne Sze, Joel S. Emer |
ISPASS | 4 |
| 2023 | HighLight: Efficient and Flexible DNN Acceleration with Hierarchical Structured SparsityabstractDue to complex interactions among various deep neural network (DNN) optimization techniques, modern DNNs can have weights and activations that are dense or sparse with diverse sparsity degrees. To offer a good trade-off between accuracy and hardware performance, an ideal DNN accelerator should have high flexibility to efficiently translate DNN sparsity into reductions in energy and/or latency without incurring significant complexity overhead. Yannan Nellie Wu, Po-An Tsai, Saurav Muralidharan, Angshuman Parashar, Vivienne Sze, Joel S. Emer |
MICRO | 5 |
| 2023 | Tailors: Accelerating Sparse Tensor Algebra by Overbooking Buffer CapacityabstractSparse tensor algebra is a challenging class of workloads to accelerate due to low arithmetic intensity and varying sparsity patterns. Prior sparse tensor algebra accelerators have explored tiling sparse data to increase exploitable data reuse and improve throughput, but typically allocate tile size in a given buffer for the worst-case data occupancy. This severely limits the utilization of available memory resources and reduces data reuse. Other accelerators employ complex tiling during preprocessing or at runtime to determine the exact tile size based on its occupancy. Zi Yu Xue, Yannan Nellie Wu, Joel S. Emer, Vivienne Sze |
MICRO | 4 |
| 2022 | Memory-Efficient Gaussian Fitting for Depth Images in Real TimeabstractComputing consumes a significant portion of energy in many robotics applications, especially the ones involving energy-constrained robots. In addition, memory access accounts for a significant portion of the computing energy. For mapping a 3D environment, prior approaches reduce the map size while incurring a large memory overhead used for storing sensor measurements and temporary variables during computation. In this work, we present a memory-efficient algorithm, named Single-Pass Gaussian Fitting (SPGF), that accurately constructs a compact Gaussian Mixture Model (GMM) which approximates measurements from a depthmap generated from a depth camera. By incrementally constructing the GMM one pixel at a time in a single pass through the depthmap, SPGF achieves higher throughput and orders-of-magnitude lower memory overhead than prior multipass approaches. By processing the depthmap row-by-row, SPGF exploits intrinsic properties of the camera to efficiently and accurately infer surface geometries, which leads to higher precision than prior approaches while maintaining the same compactness of the GMM. Using a low-power ARM Cortex-A57 CPU on the NVIDIA Jetson TX2 platform, SPGF operates at 32fps, requires 43KB of memory overhead, and consumes only 0.11J per frame (depthmap). Thus, SPGF enables real-time mapping of large 3D environments on energy-constrained robots. Peter Zhi Xuan Li, Sertac Karaman, Vivienne Sze |
ICRA | 3 |
| 2022 | Uncertainty from Motion for DNN Monocular Depth EstimationabstractDeployment of deep neural networks (DNNs) for monocular depth estimation in safety-critical scenarios on resource-constrained platforms requires well-calibrated and efficient uncertainty estimates. However, many popular uncertainty estimation techniques, including state-of-the-art ensembles and popular sampling-based methods, require multiple inferences per input, making them difficult to deploy in latency-constrained or energy-constrained scenarios. We propose a new algorithm, called Uncertainty from Motion (UfM), that requires only one inference per input. UfM exploits the temporal redundancy in video inputs by merging incrementally the per-pixel depth prediction and per-pixel aleatoric uncertainty prediction of points that are seen in multiple views in the video sequence. When UfM is applied to ensembles, we show that UfM can retain the uncertainty quality of ensembles at a fraction of the energy by running only a single ensemble member at each frame and fusing the uncertainty over the sequence of frames. In a set of representative experiments using FCDenseNet and eight indistribution and out-of-distribution video sequences, UfM offers comparable uncertainty quality to an ensemble of size 10 while consuming only 11.3% of the ensemble's energy and running 6.4× faster on a single Nvidia RTX 2080 Ti GPU, enabling near ensemble uncertainty quality for resource-constrained, real-time scenarios. Soumya Sudhakar, Vivienne Sze, Sertac Karaman |
ICRA | 2 |
| 2022 | Sparseloop: An Analytical Approach To Sparse Tensor Accelerator ModelingabstractIn recent years, many accelerators have been proposed to efficiently process sparse tensor algebra applications (e.g., sparse neural networks). However, these proposals are single points in a large and diverse design space. The lack of systematic description and modeling support for these sparse tensor accelerators impedes hardware designers from efficient and effective design space exploration. This paper first presents a unified taxonomy to systematically describe the diverse sparse tensor accelerator design space. Based on the proposed taxonomy, it then introduces Sparseloop, the first fast, accurate, and flexible analytical modeling framework to enable early-stage evaluation and exploration of sparse tensor accelerators. Sparseloop comprehends a large set of architecture specifications, including various dataflows and sparse acceleration features (e.g., elimination of zero-based compute). Using these specifications, Sparseloop evaluates a design’s processing speed and energy efficiency while accounting for data movement and compute incurred by the employed dataflow, including the savings and overhead introduced by the sparse acceleration features using stochastic density models. Across representative accelerator designs and workloads, Sparseloop achieves over 2000× faster modeling speed than cycle-level simulations, maintains relative performance trends, and achieves 0.1% to 8% average error. The paper also presents example use cases of Sparseloop in different accelerator design flows to reveal important design insights. Yannan Nellie Wu, Po-An Tsai, Angshuman Parashar, Vivienne Sze, Joel S. Emer |
MICRO | 4 |
| 2021 | NetAdaptV2: Efficient Neural Architecture Search With Fast Super-Network Training and Architecture OptimizationabstractNeural architecture search (NAS) typically consists of three main steps: training a super-network, training and evaluating sampled deep neural networks (DNNs), and training the discovered DNN. Most of the existing efforts speed up some steps at the cost of a significant slowdown of other steps or sacrificing the support of non-differentiable search metrics. The unbalanced reduction in the time spent per step limits the total search time reduction, and the inability to support non-differentiable search metrics limits the performance of discovered DNNs.In this paper, we present NetAdaptV2 with three innovations to better balance the time spent for each step while supporting non-differentiable search metrics. First, we propose channel-level bypass connections that merge network depth and layer width into a single search dimension to reduce the time for training and evaluating sampled DNNs. Second, ordered dropout is proposed to train multiple DNNs in a single forward-backward pass to decrease the time for training a super-network. Third, we propose the multi-layer coordinate descent optimizer that considers the interplay of multiple layers in each iteration of optimization to improve the performance of discovered DNNs while supporting non-differentiable search metrics. With these innovations, NetAdaptV2 reduces the total search time by up to 5.8× on ImageNet and 2.4× on NYU Depth V2, respectively, and discovers DNNs with better accuracy-latency/accuracy-MAC trade-offs than state-of-the-art NAS works. Moreover, the discovered DNN outperforms NAS-discovered MobileNetV3 by 1.8% higher top-1 accuracy with the same latency.1 Tien-Ju Yang, Yi-Lun Liao, Vivienne Sze |
CVPR | 3 |
| 2021 | Domain-Specific Language Abstractions for CompressionabstractLittle attention has been given to language support for block-based compression algorithms, despite their high implementation complexity. Current implementations have to deal with both the intricacies of the algorithm itself, as well as the low-level optimizations necessary for generating fast code. However, many block-based compression algorithms share a common structure in terms of their data representations, data partitioning operations, and data traversals. In this work, we propose a set of high-level language abstractions that can succinctly capture this structure. These abstractions provide the building blocks for the development of a domain-specific language and an associated optimizing compiler. With compression-specific language support, researchers can focus on algorithm development rather than the low-level implementation details. Jessica Ray, Ajay Brahmakshatriya, Shoaib Kamil 0001, Albert Reuther, Vivienne Sze, Saman P. Amarasinghe |
DCC | 6 |
| 2021 | Efficient Computation of Map-scale Continuous Mutual Information on Chip in Real TimeabstractExploration tasks are essential to many emerging robotics applications, ranging from search and rescue to space exploration. The planning problem for exploration requires determining the best locations for future measurements that will enhance the fidelity of the map, for example, by reducing its total entropy. A widely-studied technique involves computing the Mutual Information (MI) between the current map and future measurements, and utilizing this MI metric to decide the locations for future measurements. However, computing MI for reasonably-sized maps is slow and power hungry, which has been a bottleneck towards fast and efficient robotic exploration. In this paper, we introduce a new hardware accelerator architecture for MI computation that features a low-latency, energy-efficient MI compute core and an optimized memory subsystem that provides sufficient bandwidth to keep the cores fully utilized. The core employs interleaving to counter the recursive algorithm, and workload balancing and numerical approximations to reduce latency and energy consumption. We demonstrate this optimized architecture with a Field-Programmable Gate Array (FPGA) implementation, which can compute MI for all cells in an entire 201-by-201 occupancy grid (e.g., representing a 20.1m-by-20.1m map at 0.1m resolution) in 1.55 ms while consuming 1.7 mJ of energy, thus finally rendering MI computation for the whole map real time and at a fraction of the energy cost of traditional compute platforms. For comparison, this particular FPGA implementation running on the Xilinx Zynq-7000 platform is two orders of magnitude faster and consumes three orders of magnitude less energy per MI map compute, when compared to a baseline GPU implementation running on an NVIDIA GeForce GTX 980 platform. The improvements are more pronounced when compared to CPU implementations of equivalent algorithms. Keshav Gupta 0004, Peter Zhi Xuan Li, Sertac Karaman, Vivienne Sze |
IROS | 4 |
| 2021 | Architecture-Level Energy Estimation for Heterogeneous Computing SystemsabstractDue to the data and computation intensive nature of many popular data processing applications, e.g., deep neural networks (DNNs), a variety of accelerators have been proposed to improve performance and energy efficiency. As a result, computing systems have become increasingly heterogeneous, with application-specific processing offloaded from the CPU to specialized accelerators. To understand the energy efficiency of such systems, it is desirable to characterize holistically the energy consumption of the CPU, the accelerator, and the data transfers in between. We present a modularized architecture-level energy estimation framework that captures the energy breakdown across the various CPU and accelerator components with a unified energy estimation back-end that allows easy integration of accelerator modeling frameworks for emerging designs. Using DNN workloads as examples, we show that CPU-end preprocessing and data transfers to and from the accelerator can account for up to 45-50% of total energy when assessing the system as a whole. Related open-source code is available at https://accelergy.mit.edu. Francis Wang, Yannan Nellie Wu, Matthew E. Woicik, Joel S. Emer, Vivienne Sze |
ISPASS | 5 |
| 2021 | Sparseloop: An Analytical, Energy-Focused Design Space Exploration Methodology for Sparse Tensor AcceleratorsabstractThis paper presents Sparseloop, the first infrastructure that implements an analytical design space exploration methodology for sparse tensor accelerators. Sparseloop comprehends a wide set of architecture specifications including various sparse optimization features such as compressed tensor storage. Using these specifications, Sparseloop can calculate a design's energy efficiency while accounting for both optimization savings and metadata overhead at each storage and compute level of the architecture using stochastic tensor density models. We validate Sparseloop on a well-known accelerator design and achieve ~99% accuracy in terms of runtime activities (e.g., compressed memory accesses). We also present a case study that highlights the key factors (e.g., uncompressed traffic, data density) that affect sparse optimization features' impact on energy efficiency. Tool available at: https://github.com/NVlabs/timeloop. Yannan Nellie Wu, Po-An Tsai, Angshuman Parashar, Vivienne Sze, Joel S. Emer |
ISPASS | 4 |
| 2020 | An Efficient and Continuous Approach to Information-Theoretic ExplorationabstractExploration of unknown environments is embedded and essential in many robotics applications. Traditional algorithms, that decide where to explore by computing the expected information gain of an incomplete map from future sensor measurements, are limited to very powerful computational platforms. In this paper, we describe a novel approach for computing this expected information gain efficiently, as principally derived via mutual information. The key idea behind the proposed approach is a continuous occupancy map framework and the recursive structure it reveals. This structure makes it possible to compute the expected information gain of sensor measurements across an entire map much faster than computing each measurements’ expected gain independently. Specifically, for an occupancy map composed of |M| cells and a range sensor that emits |Θ| measurement beams, the algorithm (titled FCMI) computes the information gain corresponding to measurements made at each cell in O(|Θ||M|) steps. To the best of our knowledge, this complexity bound is better than all existing methods for computing information gain. In our experiments, we observe that this novel, continuous approach is two orders of magnitude faster than the state-of-the-art FSMI algorithm. Theia Henderson, Vivienne Sze, Sertac Karaman |
ICRA | 2 |
| 2020 | Balancing Actuation and Computing Energy in Motion PlanningabstractWe study a novel class of motion planning problems, inspired by emerging low-energy robotic vehicles, such as insect-size flyers, chip-size satellites, and high-endurance autonomous blimps, for which the energy consumed by computing hardware during planning a path can be as large as the energy consumed by actuation hardware during the execution of the same path. We propose a new algorithm, called Compute Energy Included Motion Planning (CEIMP). CEIMP operates similarly to any other anytime planning algorithm, except it stops when it estimates further computing will require more computing energy than potential savings in actuation energy. We show that CEIMP has the same asymptotic computational complexity as existing sampling-based motion planning algorithms, such as PRM*. We also show that CEIMP outperforms the average baseline of using maximum computing resources in realistic computational experiments involving 10 floor plans from MIT buildings. In one representative experiment, CEIMP outperforms the average baseline 90.6% of the time when energy to compute one more second is equal to the energy to move one more meter, and 99.7% of the time when energy to compute one more second is equal to or greater than the energy to move 3 more meters. Soumya Sudhakar, Sertac Karaman, Vivienne Sze |
ICRA | 3 |
| 2020 | An Architecture-Level Energy and Area Estimator for Processing-In-Memory Accelerator DesignsabstractProcessing-in-memory (PIM) deep neural network (DNN) accelerators, which aim to improve energy/area efficiency of DNN processing by integrating computation into data storage, have gained popularity in recent years. Therefore, it is attractive to have a generally applicable framework that is able to quickly provide insights into the various trade-offs involved in PIM accelerator designs. We present an architecture-level design estimation framework for PIM accelerators that allows easy representations of the designs with provided architecture templates and component design templates, performs analytical runtime simulations, and produces technology-dependent area and energy estimations. We show that the framework can be easily used to evaluate state of the art PIM accelerator designs; it achieves 95% accurate total energy estimations and reproduces exact area breakdowns of the components in the design. Related open-source code is available at http://accelergy.mit.edu/. Yannan Nellie Wu, Vivienne Sze, Joel S. Emer |
ISPASS | 2 |
| 2020 | Low Power Depth Estimation of Rigid Objects for Time-of-Flight ImagingabstractDepth sensing is useful in a variety of applications that range from augmented reality to robotics. Time-of-flight (TOF) cameras are appealing because they obtain dense depth measurements with minimal latency. However, for many battery-powered devices, the illumination source of a TOF camera is power hungry and can limit the battery life of the device. To address this issue, we present an algorithm that lowers the power for depth sensing by reducing the usage of the TOF camera and estimating depth maps using concurrently collected images. Our technique also adaptively controls the TOF camera and enables it when an accurate depth map cannot be estimated. To ensure that the overall system power for depth sensing is reduced, we design our algorithm to run on a low power embedded platform, where it outputs 640 × 480 depth maps at 30 frames per second. We evaluate our approach on several RGB-D datasets, where it produces depth maps with an overall mean relative error of 0.96% and reduces the usage of the TOF camera by 85%. When used with commercial TOF cameras, we estimate that our algorithm can lower the total power for depth sensing by up to 73%. James Noraky, Vivienne Sze |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2020 | Measuring Saccade Latency Using Smartphone CamerasabstractOBJECTIVE: Accurate quantification of neurodegenerative disease progression is an ongoing challenge that complicates efforts to understand and treat these conditions. Clinical studies have shown that eye movement features may serve as objective biomarkers to support diagnosis and tracking of disease progression. Here, we demonstrate that saccade latency-an eye movement measure of reaction time-can be measured robustly outside of the clinical environment with a smartphone camera. METHODS: To enable tracking of saccade latency in large cohorts of patients and control subjects, we combined a deep convolutional neural network for gaze estimation with a model-based approach for saccade onset determination that provides automated signal-quality quantification and artifact rejection. RESULTS: Simultaneous recordings with a smartphone and a high-speed camera resulted in negligible differences in saccade latency distributions. Furthermore, we demonstrated that the constraint of chinrest support can be removed when recording healthy subjects. Repeat smartphone-based measurements of saccade latency in 11 self-reported healthy subjects resulted in an intraclass correlation coefficient of 0.76, showing our approach has good to excellent test-retest reliability. Additionally, we conducted more than 19 000 saccade latency measurements in 29 self-reported healthy subjects and observed significant intra- and inter-subject variability, which highlights the importance of individualized tracking. Lastly, we showed that with around 65 measurements we can estimate mean saccade latency to within less-than-10-ms precision, which takes within 4 min with our setup. CONCLUSION AND SIGNIFICANCE: By enabling repeat measurements of saccade latency and its distribution in individual subjects, our framework opens the possibility of quantifying patient state on a finer timescale in a broader population than previously possible. Hsin-Yu Lai, Gladynel Saavedra-Peña, Charles G. Sodini, Vivienne Sze, Thomas Heldt |
IEEE J. Biomed. Health Informatics | 4 |
| 2019 | Accelergy: An Architecture-Level Energy Estimation Methodology for Accelerator DesignsabstractWith Moore's law slowing down and Dennard scaling ended, energy-efficient domain-specific accelerators, such as deep neural network (DNN) processors for machine learning and programmable network switches for cloud applications, have become a promising way for hardware designers to continue bringing energy efficiency improvements to data and computation-intensive applications. To ensure the fast exploration of the accelerator design space, architecture-level energy estimators, which perform energy estimations without requiring complete hardware description of the designs, are critical to designers. However, it is difficult to use existing architecture-level energy estimators to obtain accurate estimates for accelerator designs, as accelerator designs are diverse and sensitive to data patterns. This paper presents Accelergy, a generally applicable energy estimation methodology for accelerators that allows design specifications comprised of user-defined high-level compound components and user-defined low-level primitive components, which can be characterized by third-party energy estimation plug-ins. An example with primitive and compound components for DNN accelerator designs is also provided as an application of the proposed methodology. Overall, Accelergy achieves 95% accuracy on Eyeriss, a well-known DNN accelerator design, and can correctly capture the energy breakdown of components at different granularities. The Accelergy code is available at http://accelergy.mit.edu. Yannan Nellie Wu, Joel S. Emer, Vivienne Sze |
ICCAD | 3 |
| 2019 | Low Power Adaptive Time-of-Flight Imaging for Multiple Rigid ObjectsabstractTime-of-flight (TOF) cameras are becoming increasingly popular for many mobile applications. To obtain accurate depth maps, TOF cameras must emit many pulses of light, which consumes a lot of power and lowers the battery life of mobile devices. However, lowering the number of emitted pulses results in noisy depth maps. To obtain accurate depth maps while reducing the overall number of emitted pulses, we propose an algorithm that adaptively varies the number of pulses to infrequently obtain high power depth maps and uses them to help estimate subsequent low power ones. To estimate these depth maps, our technique uses the previous frame by accounting for the 3D motion in the scene. We assume that the scene contains independently moving rigid objects and show that we can efficiently estimate the motions using just the data from a TOF camera. The resulting algorithm estimates 640 × 480 depth maps at 30 frames per second on an embedded processor. We evaluate our approach on data collected with a pulsed TOF camera and show that we can reduce the mean relative error of the low power depth maps by up to 64% and the number of emitted pulses by up to 81%. James Noraky, Charles Mathy, Alan Cheng, Vivienne Sze |
ICIP | 4 |
| 2019 | FastDepth: Fast Monocular Depth Estimation on Embedded SystemsabstractDepth sensing is a critical function for robotic tasks such as localization, mapping and obstacle detection. There has been a significant and growing interest in depth estimation from a single RGB image, due to the relatively low cost and size of monocular cameras. However, state-of-the-art single-view depth estimation algorithms are based on fairly complex deep neural networks that are too slow for real-time inference on an embedded platform, for instance, mounted on a micro aerial vehicle. In this paper, we address the problem of fast depth estimation on embedded systems. We propose an efficient and lightweight encoder-decoder network architecture and apply network pruning to further reduce computational complexity and latency. In particular, we focus on the design of a low-latency decoder. Our methodology demonstrates that it is possible to achieve similar accuracy as prior work on depth estimation, but at inference speeds that are an order of magnitude faster. Our proposed network, FastDepth, runs at 178 fps on an NVIDIA Jetson TX2 GPU and at 27 fps when using only the TX2 CPU, with active power consumption under 10 W. FastDepth achieves close to state-of-the-art accuracy on the NYU Depth v2 dataset. To the best of the authors' knowledge, this paper demonstrates real-time monocular depth estimation using a deep neural network with the lowest latency and highest throughput on an embedded platform that can be carried by a micro aerial vehicle. Diana Wofk, Fangchang Ma, Tien-Ju Yang, Sertac Karaman, Vivienne Sze |
ICRA | 5 |
| 2019 | FSMI: Fast Computation of Shannon Mutual Information for Information-Theoretic MappingabstractInformation-based mapping algorithms are critical to robot exploration tasks in several applications ranging from disaster response to space exploration. Unfortunately, most existing information-based mapping algorithms are plagued by the computational difficulty of evaluating the Shannon mutual information between potential future sensor measurements and the map. This has lead researchers to develop approximate methods, such as Cauchy-Schwarz Quadratic Mutual Information (CSQMI). In this paper, we propose a new algorithm, called Fast Shannon Mutual Information (FSMI), which is significantly faster than existing methods at computing the exact Shannon mutual information. The key insight behind FSMI is recognizing that the integral over the sensor beam can be evaluated analytically, removing an expensive numerical integration. In addition, we provide a number of approximation techniques for FSMI, which significantly improve computation time. Equipped with these approximation techniques, the FSMI algorithm is more than three orders of magnitude faster than the existing computation for Shannon mutual information; it also outperforms the CSQMI algorithm significantly, being roughly twice as fast, in our experiments. Zhengdong Zhang 0001, Trevor Henderson, Vivienne Sze, Sertac Karaman |
ICRA | 3 |
| 2018 | NetAdapt: Platform-Aware Neural Network Adaptation for Mobile Applications
Tien-Ju Yang, Andrew G. Howard, Bo Chen 0019, Alec Go, Mark Sandler 0002, Vivienne Sze, Hartwig Adam |
ECCV (10) | 7 |
| 2018 | Enabling Saccade Latency Measurements with Consumer-Grade CamerasabstractEye movements can be affected by a number of neurological, neuromuscular, and neurodegenerative disorders that are important to diagnose and track longitudinally. To enable unobtrusive tracking of disease progression, we tailored and evaluated a set of candidate eye-tracking algorithms to operate on video sequences obtained from an iPhone 6, for accurate and robust determination of the time between the presentation of a visual stimulus and the beginning of the eye movement toward the stimulus (saccade latency). Additionally, we proposed a model-based method to determine the onset of the eye movement and demonstrate that the associated residual normalized root-mean-squared error can be used to automatically flag saccade tracings that should not be included in further analysis. A variant of the iTracker algorithm performs most robustly and results in mean saccade latencies and associated standard deviations on iPhone recordings that are essentially the same as those obtained from simultaneous recordings using a high-end, high-speed camera. Our results suggest that accurate and robust saccade latency determination is feasible using consumer-grade cameras and therefore might enable unobtrusive tracking of neurodegenerative disease progression. Hsin-Yu Lai, Gladynel Saavedra-Peña, Charles G. Sodini, Thomas Heldt, Vivienne Sze |
ICIP | 5 |
| 2018 | Depth Estimation of Non-Rigid Objects for Time-Of-Flight ImagingabstractDepth sensing is useful for a variety of applications that range from augmented reality to robotics. Time-of- flight (TOF) cameras are appealing because they obtain dense depth measurements with low latency. However, for reasons ranging from power constraints to multi -camera interference, the frequency at which accurate depth measurements can be obtained is reduced. To address this, we propose an algorithm that uses concurrently collected images to estimate the depth of non-rigid objects without using the TOF camera. Our technique models non-rigid objects as locally rigid and uses previous depth measurements along with the optical flow of the images to estimate depth. In particular, we show how we exploit the previous depth measurements to directly estimate pose and how we integrate this with our model to estimate the depth of non-rigid objects by finding the solution to a sparse linear system. We evaluate our technique on a RGB-D dataset of deformable objects, where we estimate depth with a mean relative error of 0.37% and outperform other adapted techniques. James Noraky, Vivienne Sze |
ICIP | 2 |
| 2017 | Designing Energy-Efficient Convolutional Neural Networks Using Energy-Aware PruningabstractDeep convolutional neural networks (CNNs) are indispensable to state-of-the-art computer vision algorithms. However, they are still rarely deployed on battery-powered mobile devices, such as smartphones and wearable gadgets, where vision algorithms can enable many revolutionary real-world applications. The key limiting factor is the high energy consumption of CNN processing due to its high computational complexity. While there are many previous efforts that try to reduce the CNN model size or the amount of computation, we find that they do not necessarily result in lower energy consumption. Therefore, these targets do not serve as a good metric for energy cost estimation. To close the gap between CNN design and energy consumption optimization, we propose an energy-aware pruning algorithm for CNNs that directly uses the energy consumption of a CNN to guide the pruning process. The energy estimation methodology uses parameters extrapolated from actual hardware measurements. The proposed layer-by-layer pruning algorithm also prunes more aggressively than previously proposed pruning methods by minimizing the error in the output feature maps instead of the filter weights. For each layer, the weights are first pruned and then locally fine-tuned with a closed-form least-square solution to quickly restore the accuracy. After all layers are pruned, the entire network is globally fine-tuned using back-propagation. With the proposed pruning method, the energy consumption of AlexNet and GoogLeNet is reduced by 3.7x and 1.6x, respectively, with less than 1% top-5 accuracy loss. We also show that reducing the number of target classes in AlexNet greatly decreases the number of weights, but has a limited impact on energy consumption. Tien-Ju Yang, Vivienne Sze |
CVPR | 3 |
| 2017 | Low power depth estimation for time-of-flight imagingabstractDepth sensing is used in a variety of applications that range from augmented reality to robotics. One way to measure depth is with a time-of-flight (TOF) camera, which obtains depth by emitting light and measuring its round trip time. However, for many battery powered devices, the illumination source of the TOF camera requires a significant amount of power and further limits its battery life. To minimize the power required for depth sensing, we present an algorithm that exploits the apparent motion across images collected alongside the TOF camera to obtain a new depth map without illuminating the scene. Our technique is best suited for estimating the depth of rigid objects and obtains low latency, 640 × 480 depth maps at 30 frames per second on a low power embedded platform by using block matching at a sparse set of points and least squares minimization. We evaluated our technique on an RGB-D dataset where it produced depth maps with a mean relative error of 0.85% while reducing the total power required for depth sensing by 3×. James Noraky, Vivienne Sze |
ICIP | 2 |
| 2017 | Towards closing the energy gap between HOG and CNN features for embedded vision (Invited paper)abstractComputer vision enables a wide range of applications in robotics/drones, self-driving cars, smart Internet of Things, and portable/wearable electronics. For many of these applications, local embedded processing is preferred due to privacy and/or latency concerns. Accordingly, energy-efficient embedded vision hardware delivering real-time and robust performance is crucial. While deep learning is gaining popularity in several computer vision algorithms, a significant energy consumption difference exists compared to traditional hand-crafted approaches. In this paper, we provide an in-depth analysis of the computation, energy and accuracy trade-offs between learned features such as deep Convolutional Neural Networks (CNN) and hand-crafted features such as Histogram of Oriented Gradients (HOG). This analysis is supported by measurements from two chips that implement these algorithms. Our goal is to understand the source of the energy discrepancy between the two approaches and to provide insight about the potential areas where CNNs can be improved and eventually approach the energy-efficiency of HOG while maintaining its outstanding performance accuracy. Amr Suleiman, Joel S. Emer, Vivienne Sze |
ISCAS | 4 |
| 2017 | Efficient Processing of Deep Neural Networks: A Tutorial and SurveyabstractDeep neural networks (DNNs) are currently widely used for many artificial intelligence (AI) applications including computer vision, speech recognition, and robotics. While DNNs deliver state-of-the-art accuracy on many AI tasks, it comes at the cost of high computational complexity. Accordingly, techniques that enable efficient processing of DNNs to improve energy efficiency and throughput without sacrificing application accuracy or increasing hardware cost are critical to the wide deployment of DNNs in AI systems. This article aims to provide a comprehensive tutorial and survey about the recent advances toward the goal of enabling efficient processing of DNNs. Specifically, it will provide an overview of DNNs, discuss various hardware platforms and architectures that support DNNs, and highlight key trends in reducing the computation cost of DNNs either solely via hardware design changes or via joint hardware design and DNN algorithm changes. It will also summarize various development resources that enable researchers and practitioners to quickly get started in this field, and highlight important benchmarking metrics and design considerations that should be used for evaluating the rapidly growing number of DNN hardware designs, optionally including algorithmic codesigns, being proposed in academia and industry. The reader will take away the following concepts from this article: understand the key design considerations for DNNs; be able to evaluate different DNN hardware implementations with benchmarks and comparison metrics; understand the tradeoffs between various hardware architectures and platforms; be able to evaluate the utility of various DNN design techniques for efficient processing; and understand recent implementation trends and opportunities. Vivienne Sze, Tien-Ju Yang, Joel S. Emer |
Proc. IEEE | 1 |
| 2016 | Eyeriss: A Spatial Architecture for Energy-Efficient Dataflow for Convolutional Neural NetworksabstractDeep convolutional neural networks (CNNs) are widely used in modern AI systems for their superior accuracy but at the cost of high computational complexity. The complexity comes from the need to simultaneously process hundreds of filters and channels in the high-dimensional convolutions, which involve a significant amount of data movement. Although highly-parallel compute paradigms, such as SIMD/SIMT, effectively address the computation requirement to achieve high throughput, energy consumption still remains high as data movement can be more expensive than computation. Accordingly, finding a dataflow that supports parallel processing with minimal data movement cost is crucial to achieving energy-efficient CNN processing without compromising accuracy. In this paper, we present a novel dataflow, called row-stationary (RS), that minimizes data movement energy consumption on a spatial architecture. This is realized by exploiting local data reuse of filter weights and feature map pixels, i.e., activations, in the high-dimensional convolutions, and minimizing data movement of partial sum accumulations. Unlike dataflows used in existing designs, which only reduce certain types of data movement, the proposed RS dataflow can adapt to different CNN shape configurations and reduces all types of data movement through maximally utilizing the processing engine (PE) local storage, direct inter-PE communication and spatial parallelism. To evaluate the energy efficiency of the different dataflows, we propose an analysis framework that compares energy cost under the same hardware area and processing parallelism constraints. Experiments using the CNN configurations of AlexNet show that the proposed RS dataflow is more energy efficient than existing dataflows in both convolutional (1.4× to 2.5×) and fully-connected layers (at least 1.3× for batch size larger than 16). The RS dataflow has also been demonstrated on a fabricated chip, which verifies our energy analysis. Joel S. Emer, Vivienne Sze |
ISCA | 3 |
| 2016 | Introduction to the Special Issue on HEVC Extensions and Efficient HEVC ImplementationsabstractHigh Efficiency Video Coding (HEVC) is the most recent standard in the series of major video coding standards jointly produced by the ITU-T Video Coding Experts Group (VCEG) and the ISO/IEC Moving Picture Experts Group (MPEG). HEVC was first approved in 2013 in the ITU-T as Recommendation H.265 and in ISO/IEC as International Standard 23008-2, and it offers an unprecedented degree of compression capability for a very wide variety of applications. In the three years since its initial completion, it has been extended in several important ways to further broaden its scope. This special issue on HEVC features two sections: 1) HEVC extensions and 2) efficient HEVC implementations. Jens-Rainer Ohm, Gary J. Sullivan, Vivienne Sze, Thomas Wiegand 0001, Madhukar Budagavi |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2015 | Rotate intra block copy for still image codingabstractThis paper proposes a method called rotate intra block copy, which extends the intra block copy technique by making the block matching process invariant to rotation. HEVC intra prediction plus rotate intra block copy gives an average of 20% reduction in residual energy (i.e. prediction error) compared to HEVC intra prediction plus intra block copy. As the motion vector correlation in rotate intra block copy is different from the intra block copy, a new method of motion vector coding is presented. The impact of angular resolution on residual energy reduction is also evaluated. In a full codec pipeline, this reduction in residual energy translates into a coding gain in BD-rate of 3.4% over HEVC intra prediction plus intra block copy for both screen content and camera-captured gray scale images. Zhengdong Zhang 0001, Vivienne Sze |
ICIP | 2 |
| 2015 | A Deeply Pipelined CABAC Decoder for HEVC Supporting Level 6.2 High-Tier ApplicationsabstractHigh Efficiency Video Coding (HEVC) is the latest video coding standard that specifies video resolutions up to 8K ultra-high definition (UHD) at 120 frames/s to support the next decade of video applications. This results in high-throughput requirements for the context-adaptive binary arithmetic coding (CABAC) entropy decoder, which was already a well-known bottleneck in H.264/AVC. To address the throughput challenges, several modifications were made to CABAC during the standardization of HEVC. This paper leverages these improvements in the design of a high-throughput HEVC CABAC decoder. It also supports the high-level parallel processing tools introduced by HEVC, including tile and wavefront parallel processing. The proposed design uses a deeply pipelined architecture to achieve a high clock rate. Additional techniques such as the state prefetch logic, latched-based context memory, and separate finite state machines are applied to minimize stall cycles, while multibypass-bin decoding is used to further increase the throughput. The design is implemented in an International Business Machines 45-nm silicon on insulator process. After place and route, its operating frequency reaches 1.6 GHz. The corresponding throughputs achieve up to 1696 and 2314 Mbin/s under common and theoretical worst-case test conditions, respectively. The results show that the design is sufficient to decode in real-time high-tier video bitstreams at level 6.2 (8K UHD at 120 frames/s), or main-tier bitstreams at level 5.1 (4K UHD at 60 frames/s) for applications requiring subframe latency, such as video conferencing. Vivienne Sze |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2014 | A 2014 Mbin/s deeply pipelined CABAC decoder for HEVCabstractHigh Efficiency Video Coding (HEVC) is the latest video standard that specifies video resolutions up to 8K Ultra-HD (UHD) at 120 fps to support the next decade of video applications. This results in high throughput requirements for the Context Adaptive Binary Arithmetic Coding (CABAC) entropy decoder, which was already a well-known bottleneck in H.264/AVC. Several modifications were made to the HEVC CABAC to address the throughput challenges. This work leverages these improvements in the design of a high throughput HEVC CABAC decoder. The proposed design uses a deeply pipelined architecture to achieve a high clock rate. Additional techniques such as state prefetch logic, latched-based context memory, and separate finite state machines are applied to minimize stall cycles, while multi-bypass bins decoding is used to further increase the throughput. The design is synthesized in a IBM 45nm SOI process, and achieves throughputs up to 2014 and 2748 Mbin/s under common and worst-case test conditions, respectively, at 1.9 GHz operating frequency. The results show that the design is sufficient to decode video bitstreams in real-time at Level 6.2, or at Level 6.0 for applications requiring sub-frame latency. Vivienne Sze |
ICIP | 2 |
| 2014 | Energy and area-efficient hardware implementation of HEVC inverse transform and dequantizationabstractHigh Efficiency Video Coding (HEVC) inverse transform for residual coding uses 2-D 4×4 to 32×32 transforms with higher precision as compared to H.264/AVC's 4×4 and 8×8 transforms resulting in an increased hardware complexity. In this paper, an energy and area-efficient VLSI architecture of an HEVC-compliant inverse transform and dequantization engine is presented. We implement a pipelining scheme to process all transform sizes at a minimum throughput of 2 pixel/cycle with zero-column skipping for improved throughput. We use data-gating in the 1-D Inverse Discrete Cosine Transform engine to improve energy-efficiency for smaller transform sizes. A high-density SRAM-based transpose memory is used for an area-efficient design. This design supports decoding of 4K Ultra-HD (3840×2160) video at 30 frame/sec. The inverse transform engine takes 98.1 kgate logic, 16.4 kbit SRAM and 10.82 pJ/pixel while the dequantization engine takes 27.7 kgate logic, 8.2 kbit SRAM and 1.10 pJ/pixel in 40 nm CMOS technology. Although larger transforms require more computation per coefficient, they typically contain a smaller proportion of non-zero coefficients. Due to this trade-off, larger transforms can be more energy-efficient. Mehul Tikekar, Chao-Tsung Huang, Vivienne Sze, Anantha P. Chandrakasan |
ICIP | 3 |
| 2012 | Unified forward+inverse transform architecture for HEVCabstractThe upcoming HEVC video coding standard supports many transform sizes ranging from 4-point to 32-point in square and rectangular form. Multiple transform sizes improve coding efficiency, but also increase the implementation complexity. Furthermore both forward and inverse transforms need to be supported in various consumer devices. This paper presents a unified forward+inverse transform architecture for HEVC. The unified architecture makes use of symmetry properties that exist in the HEVC forward and inverse transform matrices to achieve hardware sharing across different transform sizes and also between forward and inverse transforms. It uses 43-45% less area than separate forward and inverse core transform implementations. Madhukar Budagavi, Vivienne Sze |
ICIP | 2 |
| 2012 | Hardware-aware motion estimation search algorithm development for high-efficiency video coding (HEVC) standardabstractThis work presents a hardware-aware search algorithm for HEVC motion estimation. Implications of several decisions in search algorithm are considered with respect to their hardware implementation costs (in terms of area and bandwidth). Proposed algorithm provides 3X logic area in integer motion estimation, 16% on-chip reference buffer area and 47X maximum off-chip bandwidth savings when compared to HM-3.0 fast search algorithm. Mahmut E. Sinangil, Anantha P. Chandrakasan, Vivienne Sze, Minhua Zhou |
ICIP | 3 |
| 2012 | Memory cost vs. coding efficiency trade-offs for HEVC motion estimation engineabstractThis paper presents a comparison between various High Efficiency Video Coding (HEVC) motion estimation configurations in terms of coding efficiency and memory cost in hardware. An HEVC motion estimation hardware model that is suitable to implement HEVC reference software (HM) search algorithm is created and memory area and data bandwidth requirements are calculated based on this model. 11 different motion estimation configurations are considered. Supporting smaller block sizes is shown to impose significant memory cost in hardware although the coding gain achieved through supporting them is relatively smaller. Hence, depending on target encoder specifications, the decision can be made not to support certain block sizes. Specifically, supporting only 64x64, 32x32 and 16x16 block sizes provide 3.2X on-chip memory area, 26X on-chip bandwidth and 12.5X off-chip bandwidth savings at the expense of 12% bit-rate increase when compared to the anchor configuration supporting all block sizes. Mahmut E. Sinangil, Anantha P. Chandrakasan, Vivienne Sze, Minhua Zhou |
ICIP | 3 |
| 2012 | Parallelization of CABAC transform coefficient coding for HEVCabstractData dependencies in CABAC make it difficult to parallelize and thus limits its throughput. The majority of the bins processed by the CABAC are used to represent the prediction error/residual in terms of quantized transform coefficients. This paper provides an overview of the various improvements to context selection and scans in transform coefficient coding that enable HEVC to potentially achieve higher throughput relative to AVC/H.264. Specifically, it describes changes that remove data dependencies in significance map and coefficient level coding. Proposed and adopted techniques up to HM-4.0 are discussed. This work illustrates that accounting for implementation cost when designing video coding algorithms can result in a design that can enable higher processing speed and reduce hardware cost, while still delivering high coding efficiency. Vivienne Sze, Madhukar Budagavi |
PCS | 1 |
| 2012 | High Throughput CABAC Entropy Coding in HEVCabstractContext-adaptive binary arithmetic coding (CAB-AC) is a method of entropy coding first introduced in H.264/AVC and now used in the newest standard High Efficiency Video Coding (HEVC). While it provides high coding efficiency, the data dependencies in H.264/AVC CABAC make it challenging to parallelize and thus, limit its throughput. Accordingly, during the standardization of entropy coding for HEVC, both coding efficiency and throughput were considered. This paper highlights the key techniques that were used to enable HEVC to potentially achieve higher throughput while delivering coding gains relative to H.264/AVC. These techniques include reducing context coded bins, grouping bypass bins, grouping bins with the same context, reducing context selection dependencies, reducing total bins, and reducing parsing dependencies. It also describes reductions to memory requirements that benefit both throughput and implementation costs. Proposed and adopted techniques up to draft international standard (test model HM-8.0) are discussed. In addition, analysis and simulation results are provided to quantify the throughput improvements and memory reduction compared with H.264/AVC. In HEVC, the maximum number of context-coded bins is reduced by 8×, and the context memory and line buffer are reduced by 3× and 20×, respectively. This paper illustrates that accounting for implementation cost when designing video coding algorithms can result in a design that enables higher processing speed and lowers hardware costs, while still delivering high coding efficiency. Vivienne Sze, Madhukar Budagavi |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2011 | Joint algorithm-architecture optimization of CABAC to increase speed and reduce area costabstractTo address the increasing demand for higher resolution and frame rates, processing speed (i.e. performance) and area cost need to be considered in the development of next generation video coding. Accordingly, both algorithm and architecture should be taken into account during video codec design. This paper proposes joint optimization of both the algorithm and architecture to ensure that high coding efficiency can be achieved in conjunction with high processing speed and low area cost. Specifically, it presents two optimizations that can be performed on Context-based Adaptive Binary Arithmetic Coding (CABAC), a form of entropy coding in H.264/AVC. First, subinterval reordering is proposed for the arithmetic de coder to increase the processing speed by 14 to 22% with no cost to coding efficiency. Second, modification of the motion vector difference (mvd) context selection is proposed to reduce memory requirements (i.e. area cost) by 50% with negligible coding efficiency impact (≤0.02%). These joint algorithm and architecture optimizations are non-standard com pliant and thus are well suited to be used in High Efficiency Video Coding (HEVC), the successor to H.264/AVC. Vivienne Sze, Anantha P. Chandrakasan |
ICASSP | 1 |
| 2011 | HEVC ALF decode complexity analysis and reductionabstractThis paper analyzes the decoder implementation complexity of a new tool called Adaptive Loop Filtering (ALF) being considered for the ITU-T/ISO/IEC High Efficiency Video Coding (HEVC) standard, and proposes new luma filters (Nx7 and Nx5) for ALF that reduce memory bandwidth, memory size requirements, and number of computations. The luma filters in ALF of the initial version HEVC Test Model (HM-1.0) have a maximum vertical size of 9. The vertical size of the ALF filters determines the memory size (line buffers) and memory bandwidth requirements. Accordingly, this paper proposes reducing the vertical size of ALF filters to 7 and 5, which are referred to as Nx7 and Nx5 filter sets respectively. These filters reduce memory bandwidth and size requirements by 25% and 50% respectively with minimal impact on coding efficiency. In addition, the worst case computational complexity is reduced by -10% and -20% respectively. Reduced vertical size luma ALF filters are under consideration for inclusion in HEVC standard with Nx7 being been adopted into HM- 2.0 and Nx5 being under consideration for HM-4.0. Madhukar Budagavi, Vivienne Sze, Minhua Zhou |
ICIP | 2 |
| 2010 | Technologies for Ultradynamic Voltage ScalingabstractEnergy efficiency of electronic circuits is a critical concern in a wide range of applications from mobile multi-media to biomedical monitoring. An added challenge is that many of these applications have dynamic workloads. To reduce the energy consumption under these variable computation requirements, the underlying circuits must function efficiently over a wide range of supply voltages. This paper presents voltage-scalable circuits such as logic cells, SRAMs, ADCs, and dc-dc converters. Using these circuits as building blocks, two different applications are highlighted. First, we describe an H.264/AVC video decoder that efficiently scales between QCIF and 1080p resolutions, using a supply voltage varying from 0.5 V to 0.85 V. Second, we describe a 0.3 V 16-bit micro-controller with on-chip SRAM, where the supply voltage is generated efficiently by an integrated dc-dc converter. Anantha P. Chandrakasan, Denis C. Daly, Daniel F. Finchelstein, Joyce Kwong, Yogesh K. Ramadass, Mahmut E. Sinangil, Vivienne Sze, Naveen Verma |
Proc. IEEE | 7 |
| 2009 | A high throughput CABAC algorithm using syntax element partitioningabstractEnabling parallel processing is becoming increasingly necessary for video decoding as performance requirements continue to rise due to growing resolution and frame rate demands. It is important to address known bottlenecks in the video decoder such as entropy decoding, specifically the highly serial Context-based Adaptive Binary Arithmetic Coding (CABAC) algorithm. Concurrency must be enabled with minimal cost to coding efficiency, power, area and delay. This work proposes a new CABAC algorithm for the next generation standard in which binary symbols are grouped by syntax elements and assigned to different partitions which can be decoded in parallel. Furthermore, since the distribution of binary symbols changes with quantization, an adaptive binary symbol allocation scheme is proposed to maximize throughput. Application of this next generation CABAC algorithm on five 720p sequences shows a throughput increase of up to 3x can be achieved with negligible impact on coding efficiency (0.06% to 0.37%), which is a 2 to 4x reduction in coding penalty compared with H.264/AVC and entropy slices. Area cost is also reduced by 2x. This increased throughput can be traded-off for low power consumption in mobile applications. Vivienne Sze, Anantha P. Chandrakasan |
ICIP | 1 |
| 2009 | Low-Power Impulse UWB Architectures and CircuitsabstractUltra-wide-band (UWB) communication has a variety of applications ranging from wireless USB to radio frequency (RF) identification tags. For many of these applications, energy is critical due to the fact that the radios are situated on battery-operated or even batteryless devices. Two custom low-power impulse UWB systems are presented in this paper that address high- and low-data-rate applications. Both systems utilize energy-efficient architectures and circuits. The high-rate system leverages parallelism to enable the use of energy-efficient architectures and aggressive voltage scaling down to 0.4 V while maintaining a rate of 100 Mb/s. The low-rate system has an all digital transmitter architecture, 0.65 and 0.5 V radio-frequency (RF) and analog circuits in the receiver, and no RF local oscillators, allowing the chipset to power on in 2 ns for highly duty-cycled operation. Anantha P. Chandrakasan, Fred S. Lee, David D. Wentzloff, Vivienne Sze, Brian P. Ginsburg, Patrick P. Mercier, Denis C. Daly, Raúl Blázquez |
Proc. IEEE | 4 |
| 2009 | Multicore Processing and Efficient On-Chip Caching for H.264 and Future Video DecodersabstractPerformance requirements for video decoding will continue to rise in the future due to the adoption of higher resolutions and faster frame rates. Multicore processing is an effective way to handle the resulting increase in computation. For power-constrained applications such as mobile devices, extra performance can be traded-off for lower power consumption via voltage scaling. As memory power is a significant part of system power, it is also important to reduce unnecessary on-chip and off-chip memory accesses. This paper proposes several techniques that enable multiple parallel decoders to process a single video sequence; the paper also demonstrates several on-chip caching schemes. First, we describe techniques that can be applied to the existing H.264 standard, such as multiframe processing. Second, with an eye toward future video standards, we propose replacing the traditional raster-scan processing with an interleaved macroblock ordering; this can increase parallelism with minimal impact on coding efficiency and latency. The proposed architectures allowNparallel hardware decoders to achieve a speedup of up to a factor ofN. For example, ifN=3, the proposed multiple frame and interleaved entropy slice multicore processing techniques can achieve performance improvements of 2.64times and 2.91times, respectively. This extra hardware performance can be used to decode higher definition videos. Alternatively, it can be traded-off for dynamic power savings of 60% relative to a single nominal-voltage decoder. Finally, on-chip caching methods are presented that significantly reduce off-chip memory bandwidth, leading to a further increase in performance and energy efficiency. Data-forwarding caches can reduce off-chip memory reads by 53%, while using a last-frame cache can eliminate 80% of the off-chip reads. The proposed techniques were validated and benchmarked using full-system Verilog hardware simulations based on an existing decoder; they should also be applicable to most other decoder architectures. The metrics used to evaluate the ideas in this paper are performance, power, area, memory efficiency, coding efficiency, and input latency. Daniel F. Finchelstein, Vivienne Sze, Anantha P. Chandrakasan |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2008 | Parallel CABAC for low power video codingabstractWith the growing presence of high definition video content on battery-operated handheld devices such as camera phones, digital still cameras, digital camcorders, and personal media players, it is becoming ever more important that video compression be power efficient. A popular form of entropy coding called Context-Based Adaptive Binary Arithmetic Coding (CABAC) provides high coding efficiency but has limited throughput. This can lead to high operating frequencies resulting in high power dissipation. This paper presents a novel parallel CABAC scheme which enables a throughput increase of N-fold (depending on the degree parallelism), reducing the frequency requirement and expected power consumption of the coding engine. Experiments show that this new scheme (with N=2) can deliver ∼2x throughput improvement at a cost of 0.76% average increase in bit-rate or equivalently a decrease in average PSNR of 0.025dB on five 720p resolution video clips when compared with H.264/AVC. Vivienne Sze, Anantha P. Chandrakasan, Madhukar Budagavi, Minhua Zhou |
ICIP | 1 |
| 2007 | A 0.4-V UWB baseband processorabstractA 0.4-V UWB digital baseband processor has been fabricated in a standard-VT 90-nm CMOS technology. The base-band processor operates at an ultra-low supply voltage to reduce energy consumption and utilizes a highly parallelized architecture to meet throughput constraints. While ultra-low voltage operation is usually limited to low energy, low performance applications, this work examines how it can be applied to low energy, high performance applications. Measured results for a 20-pJ/bit 100-Mbps UWB baseband processor are presented. Architectural techniques and design methodologies for reducing additional complexity due to parallelism are discussed. Vivienne Sze, Anantha P. Chandrakasan |
ISLPED | 1 |
| 2006 | An Energy Efficient Sub-Threshold Baseband Processor Architecture for Pulsed Ultra-Wideband CommunicationsabstractThis paper describes how parallelism in the digital baseband processor can reduce the energy required to receive ultra-wideband (UWB) packets. The supply voltage of the digital baseband is lowered so that the correlator operates near its minimum energy point resulting in a 68% energy reduction across the entire baseband. This optimum supply voltage occurs below the threshold voltage, placing the circuit in the sub-threshold region. The correlator and the rest of the baseband must be parallelized to maintain throughput at this reduced voltage. While sub-threshold operation is traditionally used for low energy, low frequency applications such as wrist-watches, this paper examines how sub-threshold operation can be applied to low energy, high performance applications. The correlators are further parallelized for a 31x reduction in the synchronization time, which along with duty-cycling, lowers the energy per packet by 43% for a 500 byte packet. Simulation results for a 100 Mbps UWB baseband processor are described Vivienne Sze, Raúl Blázquez, Manish Bhardwaj, Anantha P. Chandrakasan |
ICASSP (3) | 1 |