EDBT 2026 Demo / reviewers in the wild / expert
Bevan M. Baas
dblp:22/1108
· DBLP profile ↗
38ranked-venue papers
2as first author
5since 2021 · last 2023
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 34 · 1 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-authorComputer networks · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | Energy-efficient canonical Huffman decoders on many-core processor arrays and FPGAs
Satyabrata Sarangi, Bevan M. Baas |
Integr. | 2 |
| 2022 | A Low-Overhead Method for the Accurate Estimation of the Maximum Operating Clock FrequencyabstractThe maximum operating clock frequency of a digital system depends on static factors such as process variations and dynamic factors such as temperature and instantaneous supply voltage. The more accurately this maximum operating frequency can be known, the higher the clock frequency can be set thereby maximizing performance and minimizing the performance loss margin. We designed, fabricated, and measured circuits for the accurate estimation of the maximum operating frequency of a digital system. The circuits require: only 67.6 μm^2 die area in 32 nm CMOS, only a few clock cycles per measurement, minor modifications to the base system’s hardware, and small increases, if any, to the delays of the system’s critical paths. The technique also enables great potential flexibility in varying which circuits are tested. Measurements from fabricated silicon show an average 44%–54% reduction in frequency margin below the maximum operating clock frequency compared to the commonly-used critical path replica approach. Brent Bohnenstiehl, Aaron Stillmaker, Timothy Andreas, Bevan M. Baas |
VLSI-SoC | 4 |
| 2022 | Architecture and 28 nm CMOS Design of a 1886 MBin/sec Context-Adaptive Binary Arithmetic Coder (CABAC) EncoderabstractEntropy encoding is a key element in H.264/MPEG-4 AVC video coding responsible for the efficient and lossless compression of transformed and quantized data. Context-Adaptive Binary Arithmetic Coding (CABAC) is one of the two entropy coding methods available in the H.264 standard, and achieves a 15%–19% greater bit-rate reduction than the optional CAVLC method. However, the computational complexity of CABAC is much greater due to its combination of arithmetic coding and adaptive context modeling, so hardware acceleration for CABAC is necessary for real-time high resolution video coding. We present a hardware CABAC encoder for the H.264 main profile that supports all CABAC functions including context initialization, binarization, context modeling and binary arithmetic coding in hardware. The architecture of the CABAC encoder contains a six-stage pipeline and is capable of encoding at a rate of 1 bin per cycle where a bin is a binarized encoded symbol. The CABAC encoder was implemented in a 28 nm FD-SOI CMOS technology occupying a chip area of 33,411 µm2which is equivalent to 23.4K minimum-sized NAND logic gates. The synthesized CABAC encoder achieves a clock rate of 1.886 GHz which is 3.04× greater than the previously fastest known design, and it achieves a throughput of 1,886 MBin/s which is 1.10× greater than the previously fastest known design. The design was also laid out in a standard cell chip design and achieves a clock rate of 1.495 GHz which is 4.53× greater than the previously fastest known design, and it achieves a throughput of 1,495 MBin/s which is 1.06× greater than the previously fastest known design. At a supply voltage of 0.8 V, the laid out chip design can process 769 million bins per second which is sufficient to encode real-time 4K UHD (3840×2160) video at 60 frames per second while dissipating an average power of 11.48 mW. Aaron Stillmaker, Bevan M. Baas |
VLSI-SoC | 3 |
| 2021 | Canonical Huffman Decoder on Fine-grain Many-core Processor ArraysabstractCanonical Huffman codecs have been used in a wide variety of platforms ranging from mobile devices to data centers which all demand high energy efficiency and high throughput. This work presents bit-parallel canonical Huffman decoder implementations on a fine-grain many-core array built using simple RISC-style programmable processors. We develop multiple energy-efficient and area-efficient decoder implementations and the results are compared with an Intel i7-4850HQ and a massively parallel GT 750M GPU executing the corpus benchmarks: Calgary, Canterbury, Artificial, and Large. The many-core implementations achieve a scaled throughput per chip area that is 324x and 2.7x greater on average than the i7 and GT 750M respectively. In addition, the many-core implementations yield a scaled energy efficiency (bytes decoded per energy) that is 24.1x and 4.6x greater than the i7 and GT 750M respectively. Satyabrata Sarangi, Bevan M. Baas |
ASP-DAC | 2 |
| 2021 | DeepScaleTool: A Tool for the Accurate Estimation of Technology Scaling in the Deep-Submicron EraabstractThe estimation of classical CMOS "constant-field" or "Dennard" scaling methods that define scaling factors for various dimensional and electrical parameters have become less accurate in the deep-submicron regime, which drives the need for better estimation approaches especially in the educational and research domains. We present DeepScaleTool, a tool for the accurate estimation of deep-submicron technology scaling by modeling and curve fitting published data by a leading commercial fabrication company for silicon fabrication technology generations from 130 nm to 7 nm for the key parameters of area, delay, and energy. Compared to 10 nm-7 nm scaling data published by a leading foundry, the DeepScaleTool achieves an error of 1.7% in area, 2.5% in delay, and 5% in power. This compares favorably with another leading academic estimation method that achieves an error of 24% in area, 9.1% in delay, and 24.9% in power. Satyabrata Sarangi, Bevan M. Baas |
ISCAS | 2 |
| 2020 | Scalable energy-efficient parallel sorting on a fine-grained many-core processor array
Aaron Stillmaker, Brent Bohnenstiehl, Lucas Stillmaker, Bevan M. Baas |
J. Parallel Distributed Comput. | 4 |
| 2019 | Corrigendum to "Scaling equations for the accurate prediction of CMOS device performance from 180 nm to 7 nm" [Integr. VLSI J. 58. (2017) 74-81]
Aaron Stillmaker, Bevan M. Baas |
Integr. | 2 |
| 2017 | Scaling equations for the accurate prediction of CMOS device performance from 180 nm to 7 nm
Aaron Stillmaker, Bevan M. Baas |
Integr. | 2 |
| 2017 | EditorialabstractAs I start my second two-year term (2017–2018) as the Editor-in-Chief (EIC) of the IEEE Transactions on Very Large Scale Integration Systems (TVLSI), I wish the TVLSI readership a very happy new year and continued professional success. It gives me great pleasure to report on the state of the journal and our performance metrics. Over the past two years, TVLSI has seen a healthy increase in the number of submissions—from 687 in 2014 to 770 in 2015, and at the time of writing of this editorial, we are at 760 submissions for 2016. We expect the number of submissions for 2016 to cross 800 before the end of the year. TVLSI, therefore, continues to be the premier archival journal for university researchers and industry practitioners in the broad area of VLSI system design. Krishnendu Chakrabarty, Massimo Alioto, Bevan M. Baas, Chirn Chye Boon, Meng-Fan Chang, Naehyuck Chang, Yao-Wen Chang, Chip-Hong Chang, Shih-Chieh Chang 0001, Poki Chen, Masud H. Chowdhury, Pasquale Corsonello, Ibrahim M. Elfadel, Said Hamdioui, Masanori Hashimoto, Tsung-Yi Ho, Houman Homayoun, Yuh-Shyan Hwang, Rajiv V. Joshi, Tanay Karnik, Mehran Mozaffari Kermani, Chulwoo Kim, Jaydeep P. Kulkarni, Eren Kursun, Erik Larsson, Hai Li 0001, Huawei Li 0001, Patrick P. Mercier, Prabhat Mishra 0001, Makoto Nagata, Arun Natarajan 0001, Koji Nii, Partha Pratim Pande, Ioannis Savidis, Mingoo Seok, Sheldon X.-D. Tan, Mark Tehranipoor, Aida Todri, Miroslav N. Velev, Xiaoqing Wen, Jiang Xu 0001, Wei Zhang 0012, Zhengya Zhang, Stacey Weber |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2017 | Hybrid Hardware/Software Floating-Point Implementations for Optimized Area and Throughput TradeoffsabstractHybrid floating-point (FP) implementations improve software FP performance without incurring the area overhead of full hardware FP units. The proposed implementations are synthesized in 65-nm CMOS and integrated into small fixed-point processors with a RISC-like architecture. Unsigned, shift carry, and leading zero detection (USL) support is added to a processor to augment an existing instruction set architecture and increase FP throughput with little area overhead. The hybrid implementations with USL support increase software FP throughput per core by 2.18× for addition/subtraction, 1.29× for multiplication, 3.07-4.05× for division, and 3.11-3.81× for square root, and use 90.7-94.6% less area than dedicated fused multiply- add (FMA) hardware. Hybrid implementations with custom FP-specific hardware increase throughput per core over a fixed-point software kernel by 3.69-7.28× for addition/subtraction, 1.22-2.03× for multiplication, 14.4× for division, and 31.9× for square root, and use 77.3-97.0% less area than dedicated FMA hardware. The circuit area and throughput are found for 38 multiply-add, 8 addition/subtraction, 6 multiplication, 45 division, and 45 square root designs. Thirty-three multiply- add implementations are presented, which improve throughput per core versus a fixed-point software implementation by 1.11-15.9× and use 38.2-95.3% less area than dedicated FMA hardware. Jon J. Pimentel, Brent Bohnenstiehl, Bevan M. Baas |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2016 | KiloCore: A 32 nm 1000-processor arrayabstractPresents a collection of slides covering the following topics: KiloCore; globally asynchronous-locally synchronous clocking; dynamic packet routing network; and processor data memory. Brent Bohnenstiehl, Aaron Stillmaker, Jon J. Pimentel, Timothy Andreas, Bin Liu 0037, Emmanuel Adeagbo, Bevan M. Baas |
Hot Chips Symposium | 8 |
| 2014 | Time-Scalable Mapping for Circuit-Switched GALS Chip Multiprocessor PlatformsabstractWe study the problem of mapping concurrent tasks of an application to cores of a chip multiprocessor that utilize circuit-switched interconnect and global asynchronous local synchronous (GALS) clocking domains. We develop a configurable algorithm that naturally handles a number of practical requirements, such as architectural features of the target platform, core failures, and hardware accelerators, and in addition, is scalable to a large number of tasks and cores. Experiments with several real life applications show that our algorithm outperforms manual mapping, integer linear programming-based mapping after ten days of solver run time, and a recent packet-switched network on chip-based task mapper through which, we underscore the unique requirements of task mapping for circuit-switched GALS architectures. Mohammad H. Foroozannejad, Matin Hashemi, Alireza Mahini, Bevan M. Baas, Soheil Ghiasi |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2014 | Achieving High-Performance On-Chip Networks With Shared-Buffer RoutersabstractOn-chip routers typically have buffers dedicated to their input or output ports for temporarily storing packets in case contention occurs on output physical channels. Buffers, unfortunately, consume significant portions of router area and power budgets. While running a traffic trace, however, not all input ports of routers have incoming packets needed to be transferred simultaneously. Therefore, a large number of buffer queues in the network are empty and other queues are mostly busy. This observation motivates us to design router architecture with shared queues (RoShaQ), router architecture that maximizes buffer utilization by allowing the sharing multiple buffer queues among input ports. Sharing queues, in fact, makes using buffers more efficient hence is able to achieve higher throughput when the network load becomes heavy. On the other side, at light traffic load, our router achieves low latency by allowing packets to effectively bypass these shared queues. Experimental results on a 65-nm CMOS standard-cell process show that over synthetic traffics RoShaQ has 17% less latency and 18% higher saturation throughput than a typical virtualchannel (VC) router. Because of its higher performance, RoShaQ consumes 9% less energy per transferred packet than VC router given the same buffer space capacity. Over real multitask applications and E3S embedded benchmarks using near-optimal NMAP mapping algorithm, RoShaQ has 32% lower latency than VC router and targeting the same application throughput with 30% lower energy per packet. Anh Thien Tran, Bevan M. Baas |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2014 | Processor Tile Shapes and Interconnect Topologies for Dense On-Chip NetworksabstractWe propose two eight-neighbor, two five-nearest-neighbor, and three six-nearest-neighbor interconnection topologies for many-core processor arrays-three of which use five-sided or hexagonal processor tiles-which typically reduce application communication distance and result in an overall application processor that requires fewer cores and lower power consumption. A 16-bit processor with the appropriate number of input and output ports is implemented in all topologies and tile shapes. The hexagonal and five-sided processor tiles and arrays of tiles are laid out with industry standard automatic place and route design flow and Manhattan-style wires without full-custom layout. A 1080p H.264/AVC residual video encoder and a 54 Mb/s 802.11a/g OFDM wireless local area network baseband receiver are mapped onto all topologies. The six-neighbor hexagonal tile incurs a 2.9% area increase per tile compared with the four-neighbor 2-D mesh, but its much more effective interprocessor interconnect yields an average total application area reduction of 22% and an average application power savings of 17%. Bevan M. Baas |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2013 | Parallel AES Encryption Engines for Many-Core Processor ArraysabstractBy exploring different granularities of data-level and task-level parallelism, we map 16 implementations of an Advanced Encryption Standard (AES) cipher with both online and offline key expansion on a fine-grained many-core system. The smallest design utilizes only six cores for offline key expansion and eight cores for online key expansion, while the largest requires 107 and 137 cores, respectively. In comparison with published AES cipher implementations on general purpose processors, our design has 3.5-15.6 times higher throughput per unit of chip area and 8.2-18.1 times higher energy efficiency. Moreover, the design shows 2.0 times higher throughput than the TI DSP C6201, and 3.3 times higher throughput per unit of chip area and 2.9 times higher energy efficiency than the GeForce 8800 GTX. Bin Liu 0037, Bevan M. Baas |
IEEE Trans. Computers | 2 |
| 2012 | Fine-Grained Energy-Efficient Sorting on a Many-Core Processor ArrayabstractData centers require significant and growing amounts of power to operate, and with increasing numbers of data centers worldwide, power consumption for enterprise workloads is a significant concern. Sorting is a key computational kernel in large database systems, and the development of energy efficient sorting capabilities would therefore significantly reduce data center power usage. We propose highly parallel sorting algorithms and mappings using a modular design for a fine-grained many-core system that greatly decreases the amount of energy consumed to perform sorts of arbitrarily large data sets. The memory, computational, and nearest-neighbor inter-processor communication hardware of the many-core processor array require relatively small die area. We present the design and implementation of several sorting variants that perform the first phase of an external sort. They are built using program kernels operating on independent processors in a many-core array with 256 bytes of data memory and fewer than 128 instructions per processor. The algorithms employed are simple and the vast majority of processors contain identical programs. Compared to a quicksort implementation on an Intel Core 2 Duo T9600 the highest throughput design achieves up to 27× higher throughput per chip area, and the most energy efficient sort yields a 330× reduction in energy dissipated per sorted block. Compared to a radix sort implementation on a GPU, the highest throughput design achieves up to 22× higher throughput per chip area, and the most energy efficient sort yields a 750× reduction in energy dissipated per sorted block. Aaron Stillmaker, Lucas Stillmaker, Bevan M. Baas |
ICPADS | 3 |
| 2012 | A hexagonal shaped processor and interconnect topology for tightly-tiled many-core architecture
Bevan M. Baas |
VLSI-SoC | 2 |
| 2011 | RoShaQ: High-performance on-chip router with shared queuesabstractOn-chip router typically has buffers dedicated to its input or output ports for temporarily storing packets in case contention occurs on output physical channels. Buffers, unfortunately, consume significant portions of router area and power. While running a traffic trace, however, not all input ports of routers have incoming packets needed to be transferred at the same time. As a result, a large number of buffer queues in the network are empty while other queues are mostly busy. This observation motivates us to design RoShaQ, a router architecture that maximizes buffer utilization by allowing to share multiple buffer queues among input ports. Sharing queues, in fact, makes using buffers more efficient hence is able to achieve higher throughput when the network load becomes heavy. On the other side, at light traffic load, our router achieves low latency by allowing packets to effectively bypass these shared queues. Experimental results show that RoShaQ is 21% less latency and 14% higher saturation throughput than a typical virtual-channel (VC) router with 4% higher power and 16% larger area. Due to its higher performance, RoShaQ consumes 7% less energy per a transferred packet than a VC router given the same buffer space capacity. Anh Thien Tran, Bevan M. Baas |
ICCD | 2 |
| 2011 | Low power LDPC decoder with efficient stopping scheme for undecodable blocksabstractAn efficient technique for early detection of undecodable blocks during LDPC decoding is introduced. The proposed method avoids unnecessary decoding iterations by predicting decoding failure and therefore results in significant improvement in power and latency in low SNR values. The proposed method which has a low hardware overhead compares the parity checksum against predefined threshold values for three iterations and terminates decoding if a condition is met. A 5.25 mm210GBASE-T Split-Row Threshold decoder is implemented using the proposed technique in 65 nm CMOS. The postlayout results show that at low SNR value of 3.0 dB, the decoder requires 2.3 times fewer decoding iterations which results in 23 pJ/bit energy dissipation. This is 2.4 times lower than the energy dissipation of Split-Row Threshold decoder without the proposed early stopping technique. Tinoosh Mohsenin, Houshmand Shirani-mehr, Bevan M. Baas |
ISCAS | 3 |
| 2011 | A 1080p H.264/AVC Baseline Residual Encoder for a Fine-Grained Many-Core SystemabstractThis paper presents a baseline residual encoder for H.264/AVC on a programmable fine-grained many-core processing array that utilizes no application-specific hardware. The software encoder contains integer transform, quantization, and context-based adaptive variable length coding functions. By exploiting fine-grained data and task-level parallelism, the residual encoder is partitioned and mapped to an array of 25 small processors. The proposed encoder encodes video sequences with variable frame sizes and can encode 1080p high-definition television at 30 f/s with 293 mW average power consumption by adjusting each processor to workload-based optimal clock frequencies and dual supply voltages-a 38.4% power reduction compared to operation with only one clock frequency and supply voltage. In comparison to published implementations on the TI C642 digital signal processing platform, the design has approximately 2.9-3.7 times higher scaled throughput, 11.2-15.0 times higher throughput per chip area, and 4.5-5.8 times lower energy per pixel. Compared to a heterogeneous single instruction, multiple data architecture customized for H.264, the presented design has 2.8-3.6 times greater throughput, 4.5-5.9 times higher area efficiency, and similar energy efficiency. The proposed fine-grained parallelization methodology provides a new approach to program a large number of simple processors allowing for a higher level of parallelization and energy-efficiency for video encoding than conventional processors while avoiding the cost and design time of implementing an application specific integrated circuit or other application-specific hardware. Bevan M. Baas |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2010 | Circuit modeling for practical many-core architecture design explorationabstractCurrent tools for computer architecture design lack standard support for multi- and many-core development. We propose using circuit models to describe the multiple processor architecture to motivate a circuit based design approach applied to a computer architecture problem. A test chip with DVFS capable processors is used to explore our ideas through software implementation. Dean Nguyen Truong, Bevan M. Baas |
DAC | 2 |
| 2010 | A Reconfigurable Source-Synchronous On-Chip Network for GALS Many-Core PlatformsabstractThis paper presents a globally-asynchronous locally-synchronous (GALS)-compatible circuit-switched on-chip network that is well suited for use in many-core platforms targeting streaming digital signal processing and embedded applications which typically have a high degree of task-level parallelism among computational kernels. Inter-processor communication is achieved through a simple yet effective reconfigurable source-synchronous network. Interconnect paths between processors can sustain a peak throughput of one word per cycle. A theoretical model is developed for analyzing the performance of the network. A 65 nm complementary metal-oxide–semiconductor GALS chip utilizing this network was fabricated which contains 164 programmable processors, three accelerators and three shared memory modules. For evaluating the efficiency of this platform, a complete 802.11a wireless local area network baseband receiver was implemented. It has a real-time throughput of 54 Mb/s with all processors running at 594 MHz and 0.95-V, and consumes an average of 174.8 mW with 12.2 mW (or 7.0%) dissipated by its interconnect links and switches. With the chip's dual supply voltages set at 0.95-V and 0.75-V, and individual processors' oscillators operating at workload-based optimal frequencies, the receiver consumes 123.2 mW, which is a 29.5% reduction in power. Measured power consumption values from the chip are within 2–5% of the estimated values. Anh Thien Tran, Dean Nguyen Truong, Bevan M. Baas |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2010 | A Low-Area Multi-Link Interconnect Architecture for GALS Chip MultiprocessorsabstractA new inter-processor communication architecture for chip multiprocessors is proposed which has a low area cost, flexible routing capability, and supports globally asynchronous locally synchronous (GALS) clocking styles. To achieve a low area cost, the proposed statically-configurable asymmetric architecture assigns large buffer resources to only the nearest neighbor interconnect and much smaller buffer resources for long distance interconnect. To maintain flexible routing capability, each neighboring processor pair has multiple connecting links. The architecture supports long distance communication in GALS systems by transferring the source clock with the data signals along the entire path for write synchronization. Compared to a traditional dynamically-configurable interconnect architecture with symmetric buffer allocation and single-links between neighboring processor pairs, this implementation has approximately two times smaller communication circuitry area with a similar routing capability. Area and speed estimates are obtained with the physical design of seven chips in 0.18-¿m CMOS. Zhiyi Yu, Bevan M. Baas |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2009 | An Improved Split-Row Threshold Decoding Algorithm for LDPC CodesabstractWe present an improved thresholding LDPC decoding algorithm which outperforms the split-row and original split-row threshold decoders with a small increase in hardware. Simulation results show that the algorithm provides 0.27- 0.50 dB coding gain over split-row, 0.10-0.20 dB over split-row threshold, and is within 0.08-0.13 dB of SPA. Compared with the original threshold algorithm the check node processor's gate count is increased by 3% while total chip area is kept the same. Tinoosh Mohsenin, Dean Nguyen Truong, Bevan M. Baas |
ICC | 3 |
| 2009 | The Design of a Reconfigurable Continuous-flow Mixed-radix FFT ProcessorabstractThe design of a highly configurable continuous flow mixed-radix (CFMR) Fast Fourier Transform (FFT) processor is presented. It computes fixed-point complex FFTs and inverse FFTs (IFFTs), and utilizes a flexible addressing scheme to enable runtime configuration of the FFT length from 16-points to 4096-points. A configurable block floating point (BFP) unit increases numerical performance. Compared to a floating point Matlab FFT function, the accuracy of the proposed architecture is 80 dB for a 64-point FFT and 74 dB for a 1024-point FFT with random complex input data. Anthony T. Jacobson, Dean Nguyen Truong, Bevan M. Baas |
ISCAS | 3 |
| 2009 | Multi-Split-Row Threshold Decoding Implementations for LDPC CodesabstractThe recently introduced Split-Row Threshold algorithm significantly improves the error performance when compared to the non- threshold Split-Row algorithm while requiring a very small increase in hardware complexity. The Multi-Split-Row Threshold decoding algorithm presented in this paper enables further reductions in routing complexity for greater throughput and smaller circuit area implementations. Several Multi-Split-Row Threshold decoder designs have been implemented in 65 nm CMOS and the impact of the different levels of partitioning on error performance, wire interconnect complexity, decoder area, and speed are investigated. The Split-Row-16 Threshold decoder occupies 3.8 mm2, runs at 100 MHz, delivers a throughput of 13.8 Gbps at 15 iterations and is only 0.28 dB and 0.22 dB away from SPA and MinSum Normalized. Tinoosh Mohsenin, Dean Nguyen Truong, Bevan M. Baas |
ISCAS | 3 |
| 2009 | A Low-cost High-speed Source-synchronous Interconnection Technique for GALS Chip MultiprocessorsabstractThe globally asynchronous locally synchronous (GALS) design style for a large area chip has become increasingly attractive due to the difficulty of designing global clocking circuits at high clock frequencies in the GHz range. In this paper, we present a high-speed interconnect network for a GALS multiprocessing system composed of a 2-D mesh array of processors. Processors are locally clocked by their own oscillators and communicate together using a static circuit-switched technique combined with a source-synchronous communication scheme. A technique to maximize the timing reliability on long-distance interconnects at high clock rates is proposed that is area and power efficient with low latency and allows a sustained ideal peak throughput of one word per cycle. Anh Thien Tran, Dean Nguyen Truong, Bevan M. Baas |
ISCAS | 3 |
| 2009 | A GALS many-core heterogeneous DSP platform with source-synchronous on-chip interconnection networkabstractThis paper presents a many-core heterogeneous computational platform that employs a GALS compatible circuit-switched on-chip network. The platform targets streaming DSP and embedded applications that have a high degree of task-level parallelism among computational kernels. The test chip was fabricated in 65nm CMOS consisting of 164 simple small programmable cores, three dedicated-purpose accelerators and three shared memory modules. All processors are clocked by their own local oscillators and communication is achieved through a simple yet effective source-synchronous communication technique that allows each interconnection link between any two processors to sustain a peak throughput of one data word per cycle. A complete 802.11a WLAN baseband receiver was implemented on this platform. It has a real-time throughput of 54 Mbps with all processors running at 594 MHz and 0.95 V, and consumes an average 174.76 mW with 12.18 mW (or 7.0%) dissipated by its interconnection links. We can fully utilize the benefit of the GALS architecture and by adjusting each processor's oscillator to run at a workload-based optimal clock frequency with the chip's dual supply voltages set at 0.95 V and 0.75 V, the receiver consumes only 123.18 mW, a 29.5% in power reduction. Measured results of its power consumption on the real chip come within the difference of only 2-5% compared with the estimated results showing our design to be highly reliable and efficient. Anh Thien Tran, Dean Nguyen Truong, Bevan M. Baas |
NOCS | 3 |
| 2009 | High Performance, Energy Efficiency, and Scalability With GALS Chip MultiprocessorsabstractChip multiprocessors with globally asynchronous locally synchronous (GALS) clocking styles are promising candidates for processing computationally-intensive and energy-constrained workloads. The GALS methodology simplifies clock tree design, provides opportunities to use clock and voltage scaling jointly in system submodules to achieve high energy efficiencies, and can also result in easily scalable clocking systems. However, its use typically also introduces performance penalties due to additional communication latency between clock domains. We show that GALS chip multiprocessors (CMPs) with large inter-processor first-inputs-first-outputs (FIFOs) buffers can inherently hide much of the GALS performance penalty while executing applications that have been mapped with few communication loops. In fact, the penalty can be driven tozerowith sufficiently large FIFOs and the removal of multiple-loop communication links. We present an example mesh-connected GALS chip multiprocessor and show it has a less than 1% performance (throughput) reduction on average compared to the corresponding synchronous system for many DSP workloads. Furthermore, adaptive clock and voltage scaling for each processor provides an approximately 40% power savings without any performance reduction. These results compare favorably with the GALS uniprocessor, which compared to the corresponding synchronous uniprocessor, has a reported greater than 10% performance (throughput) reduction and an energy savings of approximately 25% using dynamic clock and voltage scaling for many general purpose applications. Zhiyi Yu, Bevan M. Baas |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2008 | A high-performance parallel CAVLC encoder on a fine-grained many-core systemabstractThis paper presents a high-performance parallel context-based adaptive length coding (CAVLC) encoder implemented on a fine-grained many-core system. The software encoder is designed for a H.264/AVC baseline profile encoder. By utilizing arithmetic table elimination and compression techniques, the data-flow of the CAVLC encoder has been partitioned and mapped to an array of 15 small processors. The parallel workload of each processor is characterized and balanced for further throughput optimization. The proposed parallel CAVLC encoder achieves the real-time processing requirement of 30 frames per second for 720 p HDTV. Our experiments show that the presented CAVLC encoder has 4.86 to 6.83 times higher throughput and requires far smaller chip area than the identical encoder implemented on state-of-art general-purpose processors. In comparison to published implementations on common DSP processors, the design has approximately 1.0 to 6.15 times higher throughput while requiring less than 6 times smaller area. Bevan M. Baas |
ICCD | 2 |
| 2008 | Dynamic voltage and frequency scaling circuits with two supply voltagesabstractThis paper presents circuits that enable dynamic voltage and frequency scaling (DVFS) for fine-grained chip multi-processors to reduce both dynamic and leakage power dissipation. Each processor can run on either a high voltage or low voltage power supply, or disconnect from both. Switching between power supplies is performed dynamically, where scaling decisions are based on each processor's workload, allowing for reduced power consumption without a significant impact on performance. Tradeoffs in performance versus circuit area and supply noise are examined. The DVFS circuits are designed in a wrapper around each individual processor, resulting in a 12% area overhead. DVFS operation utilizing supply voltages of 1.3 V and 0.8 V on a nine-processor JPEG application reduces average energy consumption by 48% while reducing performance by only 8%. Wayne H. Cheng, Bevan M. Baas |
ISCAS | 2 |
| 2008 | A low-area interconnect architecture for chip multiprocessorsabstractA new inter-processor communication architecture for chip multiprocessors is proposed which has a low area cost and flexible routing capability. To achieve a low area cost, the proposed statically-configurable asymmetric architecture assigns large buffer resources only to the nearest neighbor interconnect and much smaller buffer resources for long distance interconnect. To maintain flexible routing capability, each neighboring processor pair has two connecting links. Compared to a traditional dynamically-configurable interconnect architecture with symmetric buffer allocation and single-links between neighboring processor pairs, this implementation has approximately 2 times smaller communication circuitry area with a similar routing capability. Area and speed estimates are obtained with the physical design of seven chips in 0.18 μm CMOS. Zhiyi Yu, Bevan M. Baas |
ISCAS | 2 |
| 2007 | High-Throughput LDPC Decoders Using A Multiple Split-Row MethodabstractWe propose the "multi-split-row'" LDPC decoding method which allows further reductions in routing complexity, greater throughput, and smaller circuit area implementations compared to the previously proposed split-row decoding method. Multi-split-row is especially useful for regular high row weight LDPC codes. A 2048-bit full parallel decoder is implemented in a 0.18 μm CMOS technology using standard MinSum, split-row-2 and split-row-4 methods. The split-row-4 decoder delivers 7.1 Gbps throughput with 15 decoding iterations, and has 3.2 times smaller circuit area and 5.2 times higher throughput than the standard MinSum decoder. Tinoosh Mohsenin, Bevan M. Baas |
ICASSP (2) | 2 |
| 2007 | A Scalable Dual-Clock FIFO for Data Transfers Between Arbitrary and Haltable Clock DomainsabstractA robust, scalable, and power efficient dual-clock first-input first-out (FIFO) architecture which is useful for transferring data between modules operating in different clock domains is presented. The architecture supports correct operation in applications where multiple clock cycles of latency exist between the data producer, FIFO, and the data consumer; and with arbitrary clock frequency changes, halting, and restarting in either or both clock domains. The architecture is demonstrated in both a 0.18- mum CMOS full-custom design and a 0.18-mum CMOS standard cell design used in a globally asynchronous locally synchronous array processor. It achieves 580-MHz operation and 10.3-mW power dissipation while performing simultaneous FIFO read and write operations at 1.8 V. Ryan W. Apperson, Zhiyi Yu, Michael J. Meeuwsen, Tinoosh Mohsenin, Bevan M. Baas |
IEEE Trans. Very Large Scale Integr. Syst. | 5 |
| 2006 | Hardware and applications of AsAP: An asynchronous array of simple processors
Bevan M. Baas, Zhiyi Yu, Michael J. Meeuwsen, Omar Sattari, Ryan W. Apperson, Eric W. Work, Jeremy W. Webb, Michael A. Lai, Daniel Gurman, Jason Cheung, Dean Nguyen Truong, Tinoosh Mohsenin |
Hot Chips Symposium | 1 |
| 2006 | Split-Row: A Reduced Complexity, High Throughput LDPC Decoder ArchitectureabstractA reduced complexity LDPC decoding method is presented that dramatically reduces wire interconnect complexity, which is a major issue in LDPC decoders. The proposed split-row method makes column processing parallelism easier to exploit, doubles available row processor parallelism, and significantly simplifies row processors - which results in smaller area, higher speeds, and lower energy dissipation. Simulation results over an additive white Gaussian channel show that the error performance of high row-weight codes with split-row decoding is within 0.3-0.6 dB of the min-sum and sum-product decoding algorithms. A full parallel decoder for a (3,6) LDPC code with a code length of 1536 bits is implemented in a 0.18 mum CMOS technology twice: once using the split-row method, and once using the min-sum algorithm for comparison. The split-row decoder operates at 53 MHz and delivers a throughput of 5.4 Gbps with 15 decoding iterations per block. The split-row decoder is about 1.3 times smaller, has an average wire length 1.5 times shorter, and has a throughput 1.6 times higher than the min-sum decoder. Tinoosh Mohsenin, Bevan M. Baas |
ICCD | 2 |
| 2006 | Implementing Tile-based Chip Multiprocessors with GALS Clocking StylesabstractThis paper investigates implementation techniques for tile-based chip multiprocessors with Globally Asynchronous Locally Synchronous (GALS) clocking styles. These architectures can simplify the physical design flow since they allow focusing on a single processor when designing an entire chip. However, they also introduce challenges to maintain system robustness and scalability. We propose a physical design flow for these architectures, investigate timing issues for robust implementations, and propose methods to take full advantage of their potential scalability. As a design example, we present data from a recently implemented single-chip 6 x 6 tile-based GALS processing array. Zhiyi Yu, Bevan M. Baas |
ICCD | 2 |
| 2005 | A generalized cached-FFT algorithmabstractFast Fourier transform (FFT) algorithms are typically designed to minimize the number of multiplications and additions while maintaining a simple form. Few FFT algorithms are designed to take advantage of hierarchical memory systems, which are easy to include in special-purpose processors, and nearly universal in modern programmable processors. We present a new generalized algorithm, called the cached-FFT, which is designed explicitly to operate on a processor with a hierarchical memory system. By taking advantage of a small and fast cache memory, the algorithm enables higher clock frequencies (for special-purpose processor applications), reduced data communication energy, and increased energy-efficiency - since smaller memories require lower energy per access and can be positioned closer to the processor. Bevan M. Baas |
ICASSP (5) | 1 |