EDBT 2026 Demo / reviewers in the wild / expert
Mau-Chung Frank Chang
dblp:c/MCFrankChang · also M. Frank Chang, M.-C. Frank Chang
· DBLP profile ↗
38ranked-venue papers
6as first author
6since 2021 · last 2025
0000-0002-2934-9359ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 32 · 5 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-authorComputer networks · 2Software engineering, systems software and programming languages · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | SAT-Accel: A Modern SAT Solver on a FPGAabstractBoolean satisfiability (SAT) solving is the first known NP-complete problem and is widely used in many application domains. Over the years, there have been so many consistent improvements in this area such that larger instances can be solved relatively quickly. Although these improvements have found their way onto CPU implementations, there has been limited progress adopting this on hardware accelerators mainly because it is difficult to implement the dynamic data structures needed to support a modern SAT solving algorithm. Michael Lo, Mau-Chung Frank Chang, Jason Cong |
FPGA | 2 |
| 2023 | HMLib: Efficient Data Transfer for HLS Using Host MemoryabstractStreaming applications compose an important portion of the workloads that FPGAs may accelerate but suffer from inefficient data movement. The inefficiency stems from copying data indirectly into the FPGA DRAM rather than directly into its on-chip memory, substantially diminishing the end-to-end speedup, especially for small workloads (hundreds of kilobytes). AMD Xilinx's Host Memory IP (HMI) aims to address the data movement problem by exposing to the developer an High-Level Synthesis (HLS) interface that moves the data from the host directly to the FPGA's on-chip memory. However, using HMI purely for its interface without additional code changes incurred a 3.3x slowdown in comparison with the current programming model. The slowdown mainly originates from OpenCL call overhead and the kernel control logic unnecessarily switching states. To overcome these issues, we propose Host Memory Library (HMLib), an efficient HLS-based library that facilitates data transfer on behalf of the user. HMLib not only optimizes the runtime stack for efficient data transfer, but also provides HLS compatible and user-friendly interfaces. We demonstrate HMLib's effectiveness for streaming applications (Deflate compression and CRC32) with improvements of up to up to 36.2X over OpenCL-DDR and up to 79.5X over raw HMI for small-scale data while maintaining little-to-no performance loss for large scale inputs. We plan to open source our work in the future. Michael Lo, Weikang Qiao, Mau-Chung Frank Chang, Jason Cong |
FPGA | 4 |
| 2022 | TopSort: A High-Performance Two-Phase Sorting Accelerator Optimized on HBM-based FPGAsabstractThe emergence of high-bandwidth memory (HBM) brings new opportunities to boost the performance of sorting acceleration on FPGAs, which was conventionally bounded by the available off-chip memory bandwidth. However, it is nontrivial for designers to fully utilize this immense bandwidth. First, the existing sorter designs cannot be directly scaled at the increasing rate of available off-chip bandwidth, as the required on-chip resource usage grows at a much faster rate and would bound the sorting performance in turn. Second, designers need an in-depth understanding of HBM’s characteristics to effectively utilize the HBM bandwidth. To tackle these challenges, we present TopSort, a novel two-phase sorting solution optimized for HBMbased FPGAs. TopSort can sort up to 4 GB data using all 32 HBM channels, with an overall sorting performance of 15.6 GB/s. TopSort is 6.7× and 2.2× faster than state-of-the-art CPU and FPGA sorters. Weikang Qiao, Licheng Guo, Zhenman Fang, Mau-Chung Frank Chang, Jason Cong |
FCCM | 4 |
| 2022 | A 14-bit 1-GS/s SiGe Bootstrap Sampler for High Resolution ADC with 250-MHz InputabstractAn 86.6-dB SFDR, 1-GS/s differential bootstrap sampler in a 0.18-um SiGe BiCMOS technology is presented. The performance is achieved using an amplitude-modulated bootstrap circuit. The results show 14-bit linearity over nearly 500-MHz bandwidth, while consuming less power compared to a conventional MOSFET switched-capacitor bootstrap circuit due to less parasitic capacitance and the use of high ftHBT. Jiazhang Song, Li-Yang Chen, Mau-Chung Frank Chang, Sudhakar Pamarti, Chih-Kong Ken Yang |
ISCAS | 3 |
| 2021 | FANS: FPGA-Accelerated Near-Storage SortingabstractLarge-scale sorting is always an important yet demanding task for data center applications. In addition to powerful processing capability, high-performance sorting system requires efficient utilization of the available bandwidth of various levels in the memory hierarchy. Nowadays, with the explosive data size, the frequent data transfers between the host and the storage device are becoming increasingly a performance bottleneck. Fortunately, the emergence of near-storage computing devices gives us the opportunity to accelerate large-scale sorting by avoiding the back and forth data transfer. Near-storage sorting is promising for extra performance improvement and power reduction. However, it is still an open question of how to achieve the optimal sorting performance on the existing near-storage computing device.In this work, we first perform an in-depth analysis of the sorting performance on the newly released Samsung SmartSSD platform. Contrary to the previous belief, our analysis shows that the end-to-end sorting performance is bound by not only the bandwidth of the flash, but also the main memory bandwidth, the configuration of the sorting kernel and the intermediate sorting status. Based on our modeling, we propose FANS, an FPGA accelerated near-storage sorting system which selects the optimized design configuration and achieves the theoretically maximum end-to-end performance when using a single Samsung SmartSSD device. The experiments demonstrate more than 3× performance speedup over the state-of-art FPGA-accelerated flash storage. Weikang Qiao, Jihun Oh, Licheng Guo, Mau-Chung Frank Chang, Jason Cong |
FCCM | 4 |
| 2021 | Self-Synchronized DS/SS With High Spread Factors for Robust Millimeter-Wave DatalinksabstractThis paper presents direct sequence spread spectrum (DS/SS) datalinks operating at 92-100 GHz with spreading factors up to 100K. Unlike traditional DS/SS datalinks which rely on preamble and cyclic dispreading for synchronization, the reported datalink uses a self-synchronous demodulation scheme, avoiding the large hardware complexity associated with exceptionally large spreading factors, and avoiding the difficulties of synchronizing the long de-spreading codes that are applicable to mm-wave systems. A demonstration link chipset is presented operating at 92-100 GHz with a spreading factor of 104K and a spread bandwidth of 6.0 GHz (baseband bandwidth of 56Kb/s) while consuming a total of 425mW. Adrian Tang 0002, Rulin Huang, Gabriel Virbila, Mau-Chung Frank Chang |
IEEE Trans. Circuits Syst. I Regul. Pap. | 4 |
| 2020 | Algorithm-Hardware Co-design for BQSR Acceleration in Genome Analysis ToolKitabstractGenome sequencing is one of the key applications in healthcare and has a great potential to realize precision medicine and personalized healthcare. However, its computing process is very time consuming. Even pre-processing the raw sequence data of a whole genome for a single person to the analysis ready data can take several days on a single-core CPU.In this paper, we propose to accelerate the performance of the widely used Genome Analysis ToolKit (GATK) using FPGAs. More specifically, we focus on the algorithm and hardware co-design for the Base Quality Score Re-calibration (BQSR) step in GATK, which is an important and time-consuming step to correct systematic errors made by a sequencing machine. Prior studies did not consider hardware acceleration for BQSR because it requires a large amount of memory with random access and has a lot of control flow. To address these challenges, we first adapt the algorithm to resolve the random memory access conflicts to achieve a fully pipelined accelerator design and reduce its dataset size. Second, we leverage the newly introduced large-capacity UltraRAM (URAM) in Xilinx UltraScale+ FPGAs to butter BQSR’s large dataset on chip, and further optimize its operating frequency. Finally, we also explore the coarse-grained pipeline and parallelism to improve the overall performance of the BQSR accelerator. Compared to the latest software implementation of BQSR on GATK 4.1, running on single-thread and 56-thread CPUs (14nm Xeon E5-2680 v4), our FPGA accelerator running on Xilinx 16nmUltraScale+VCUl525 board achieves up to 40. 7x and 8. 5x speedups, respectively. Michael Lo, Zhenman Fang, Jie Wang 0022, Peipei Zhou 0001, Mau-Chung Frank Chang, Jason Cong |
FCCM | 5 |
| 2020 | Bonsai: High-Performance Adaptive Merge Tree SortingabstractSorting is a key computational kernel in many big data applications. Most sorting implementations focus on a specific input size, record width, and hardware configuration. This has created a wide array of sorters that are optimized only to a narrow application domain.In this work we show that merge trees can be implemented on FPGAs to offer state-of-the-art performance over many problem sizes. We introduce a novel merge tree architecture and develop Bonsai, an adaptive sorting solution that takes into consideration the off-chip memory bandwidth and the amount of on-chip resources to optimize sorting time. FPGA programmability allows us to leverage Bonsai to quickly implement the optimal merge tree configuration for any problem size and memory hierarchy.Using Bonsai, we develop a state-of-the-art sorter which specifically targets DRAM-scale sorting on AWS EC2 F1 instances. For 4-32 GB array size, our implementation has a minimum of 2.3x, 1.3x, 1.2x and up to 2.5x, 3.7x, 1.3x speedup over the best designs on CPUs, FPGAs, and GPUs, respectively. Our design exhibits 3.3x better bandwidth-efficiency compared to the best previous sorting implementations. Finally, we demonstrate that Bonsai can tune our design over a wide range of problem sizes(megabyte to terabyte) and memory hierarchies including DDR DRAMs, high-bandwidth memories (HBMs) and solid-state disks (SSDs). Nikola Samardzic, Weikang Qiao, Vaibhav Aggarwal, Mau-Chung Frank Chang, Jason Cong |
ISCA | 4 |
| 2019 | An FPGA-Based BWT Accelerator for Bzip2 Data CompressionabstractThe Burrows-Wheeler Transform (BWT) has played an important role in lossless data compression algorithms. To achieve a good compression ratio, the BWT block size needs to be several hundreds of kilobytes, which requires a large amount of on-chip memory resources and limits effective hardware implementations. In this paper, we analyze the bottleneck of the BWT acceleration and present a novel design to map the anti-sequential suffix sorting algorithm to FPGAs. Our design can perform BWT with a block size of up to 500KB (i.e., bzip2 level 5 compression) on the Xilinx Virtex UltraScale+ VCU1525 board, while the state-of-art FPGA implementation can only support 4KB block size. Experiments show our FPGA design can achieve ~2x speedup compared to the best CPU implementation using standard large Corpus benchmarks. Weikang Qiao, Zhenman Fang, Mau-Chung Frank Chang, Jason Cong |
FCCM | 3 |
| 2019 | A 7.5-mW 10-Gb/s 16-QAM wireline transceiver with carrier synchronization and threshold calibration for mobile inter-chip communications in 16-nm FinFETabstractA compact energy-efficient 16-QAM wireline transceiver with carrier synchronization and threshold calibration is proposed to leverage high-density fine-pitch interconnects. Utilizing frequency-division multiplexing, the transceiver transfers four-bit data through one RF band to reduce intersymbol interferences. A forwarded clock is also transmitted through the same interconnect with the data simultaneously to enable low-power PVT-insensitive symbol clock recovery. A carrier synchronization algorithm is proposed to overcome nontrivial current and phase mismatches by including DC offset calibration and dedicated I/Q phase adjustments. Along with this carrier synchronization, a threshold calibration process is used for the transceiver to tolerate channel and circuit variations. The transceiver implemented in 16-nm FinFET occupies only 0.006-mm2 and achieves 10 Gb/s with 0.75-pJ/bit efficiency and <2.5-ns latency. Jieqiong Du, Chien-Heng Wong, Yo-Hao Tu, Wei-Han Cho, Yilei Li, Yuan Du, Po-Tsang Huang, Sheau Jiung Lee, Mau-Chung Frank Chang |
NOCS | 9 |
| 2019 | An Analog Neural Network Computing Engine Using CMOS-Compatible Charge-Trap-Transistor (CTT)abstractAn analog neural network computing engine based on CMOS-compatible charge-trap transistor (CTT) is proposed in this paper. CTT devices are used as analog multipliers. Compared to digital multipliers, CTT-based analog multiplier shows significant area and power reduction. The proposed computing engine is composed of a scalable CTT multiplier array and energy efficient analog-digital interfaces. By implementing the sequential analog fabric, the engine's mixed-signal interfaces are simplified and hardware overhead remains constant regardless of the size of the array. A proof-of-concept 784 by 784 CTT computing engine is implemented using TSMC 28-nm CMOS technology and occupies 0.68 mm2. The simulated performance achieves 76.8 TOPS (8-bit) with 500 MHz clock frequency and consumes 14.8 mW. As an example, we utilize this computing engine to address a classic pattern recognition problem-classifying handwritten digits on MNIST database and obtained a performance comparable to state-of-the-art fully connected neural networks using 8-bit fixed-point resolution. Yuan Du, Xuefeng Gu, Jieqiong Du, X. Shawn Wang, Boyu Hu, Mingzhe Jiang, Xiaoliang Chen 0001, Subramanian S. Iyer, Mau-Chung Frank Chang |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 10 |
| 2018 | High-Throughput Lossless Compression on Tightly Coupled CPU-FPGA PlatformsabstractData compression techniques have been widely used to reduce data storage and movement overhead, especially in the big data era. While FPGAs are well suited to accelerate the computation-intensive lossless compression algorithms, big data compression with parallel requests intrinsically poses two challenges to the overall system throughput. First, scaling existing single-engine FPGA compression accelerator designs already encounters bottlenecks which will result in lower clock frequency, saturated throughput and lower area efficiency. Second, when such FPGA compression accelerators are integrated with the processors, the overall system throughput is typically limited by the communication between a CPU and an FPGA. We propose a novel multi-way parallel and fully pipelined architecture to achieve high-throughput lossless compression on modern Intel-Altera HARPv2 platforms. To compensate for the compression ratio loss in a multi-way design, we implement novel techniques, such as a better data feeding method and a hash chain to increase the hash dictionary history. Our accelerator kernel itself can achieve a compression throughput of 12.8 GB/s (2.3x better than the current record throughput) and a comparable compression ratio of 2.03 for standard benchmark data. Our approach enables design scalability without a reduction in clock frequency and also improves the performance per area efficiency (up to 1.5x). Moreover, we exploit the high CPU-FPGA communication bandwidth of HARPv2 platforms to improve the compression throughput of the overall system, which can achieve an average practical end-to-end throughput of 10.0 GB/s (up to 12 GB/s for larger input files) on HARPv2. Weikang Qiao, Jieqiong Du, Zhenman Fang, Michael Lo, Mau-Chung Frank Chang, Jason Cong |
FCCM | 5 |
| 2018 | High-Throughput Lossless Compression on Tightly Coupled CPU-FPGA Platforms: (Abstract Only)abstractData compression techniques have been widely used to reduce the data storage and movement overhead, especially in the big data era. Recent studies demonstrate the great promise of FPGAs to improve the throughput of lossless compression algorithms that are very computation-intensive. However, when such FPGA-based compression accelerators are integrated with the processors, the overall system throughput is typically limited by the communication between a CPU and an FPGA. This study proposes a novel scheme to achieve high-throughput lossless compression on modern Intel-Altera HARPv2 platforms, where a Xeon CPU and an Altera FPGA are tightly coupled to improve the CPU-FPGA communication. First, it implements a multi-way parallel and fully pipelined compression accelerator based on Deflate algorithm. The accelerator itself can achieve a maximum throughput of 12.8 GB/s and a compression ratio of 2.03 over standard benchmarks. In addition, various trade-offs among compression throughput, compression ratio, FPGA resource utilization and scalability are explored to optimize the accelerator design based on different application requirements. Moreover, this study exploits the high CPU-FPGA communication bandwidth of HARPv2 platforms to improve the compression throughput of the overall system, which can achieve an average practical end-to-end throughput of 10.0 GB/s (up to 12 GB/s for larger input files) on HARPv2. Weikang Qiao, Jieqiong Du, Zhenman Fang, Michael Lo, Mau-Chung Frank Chang, Jason Cong |
FPGA | 6 |
| 2018 | A 2.6GS/s Spectrometer System in 65nm CMOS for Spaceborne Telescopic SensingabstractA fully integrated spectrometer system-on-a-chip (SoC) is demonstrated for the first time to support the back-end processing of spaceborne telescopic sensing. Like with all space-borne instruments, payload size, weight, and power are critically restricted by the launch vehicle capacity and available solar power. A custom integrated circuit (IC) approach naturally prevails over FPGA-based discrete solutions in these areas. Rather than concocting a system out of circuitries intended for different applications, as in [2], each component in this work is optimized along the operating principle of radio-frequency (RF) spectroscopy. Running at 2.6 GHz, the presented spectrometer features a three-bit flash ADC, an 8192-point polyphase filter bank (PFB), a 2048-point FFT processor, and a billion-count accumulator (ACC), all running off clocks derived from an integrated phase locked loop (PLL). The entire system consumes a peak power of 650 mW. It achieves the highest level of integration and best efficiency among current spectrometer solutions, and serves as the baseline component for upcoming NASA astrophysics spectroscopy missions. Yan Zhang 0050, Yanghyo Kim, Adrian Tang 0002, Jon Kawamura, Theodore Reck, Mau-Chung Frank Chang |
ISCAS | 6 |
| 2018 | A Single Layer 3-D Touch Sensing System for Mobile Devices ApplicationabstractTouch sensing has been widely implemented as a main methodology to bridge human and machine interactions. The traditional touch sensing range is 2-D and therefore limits the user experience. To overcome these limitations, we propose a novel 3-D contactless touch sensing called Airtouch system, which improves user experience by remotely detecting single/multi-finger position. A single layer touch panel with triangle-shaped electrodes is proposed to achieve multitouch detection capability as well as manufacturing cost reduction. Moreover, an oscillator-based-capacitive touch sensing circuit is implemented as the sensing hardware with the bootstrapping technique to eliminate the interchannel coupling effects. To further improve the system accuracy, a grouping algorithm is proposed to group the useful channels' data and filter out hardware noise impact. Finally, improved algorithms are proposed to eliminate the fringing capacitance effect and achieve accurate finger position estimation. EM simulation proved that the proposed algorithm reduced the maximum systematic error by 11 dB in the horizontal position detection. The proposed system consumes 2.3 mW and is fully compatible with existing mobile device environments. A prototype is built to demonstrate that the system can successfully detect finger movement in a vertical direction up to 6 cm and achieve a horizontal resolution up to 0.6 cm at 1 cm finger-height. As a new interface for human and machine interactions, this system offers great potential in finger movement detection and gesture recognition for small-sized electronics and advanced human interactive games for mobile device. Yan Zhang 0050, Yilei Li, Yuan Du, Yen-Cheng Kuan, Mau-Chung Frank Chang |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 7 |
| 2018 | A Novel Fully Synthesizable All-Digital RF Transmitter for IoT ApplicationsabstractIn this paper, a fully synthesizable all-digital transmitter (ADTX) is first proposed. This transmitter (TX) uses Cartesian architecture and supports wide-band quadratic-amplitude modulation with wide carrier frequency range. Furthermore, the design methodology for ADTX and corresponding bandpass filter is discussed. This TX is synthesized with digital register transfer level-graphic database system flow, and can be easily implemented in any standard CMOS technology. An exemplary TX is synthesized by TSMC 28-nm standard cell library with extremely small area (0.0009 mm2) and supports carrier frequency as high as 6 GHz with excellent error vector magnitude (<;-30 dB). To the best of the authors' knowledge, this is the first work on a fully synthesizable design of RF transistors, allowing easy technology migration and portability. Yilei Li, Kirti Dhwaj, Chien-Heng Wong, Yuan Du, Yiwu Tang, Yiyu Shi 0001, Tatsuo Itoh, Mau-Chung Frank Chang |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 9 |
| 2017 | An R2R-DAC-Based Architecture for Equalization-Equipped Voltage-Mode PAM-4 Wireline Transmitter DesignabstractThis brief presents a wireline transmitter architecture, enabling multilevel signaling with feedforward equalization (FFE) in voltage-mode. A compact R2R-DAC-based front end is proposed and analyzed in terms of its speed, power consumption, and linearity. A voltage-mode PAM-4 transmitter with 2-tap FFE utilizing the proposed architecture is implemented in the 65-nm CMOS technology. It achieves a data rate of 34 Gb/s and an energy efficiency of 2.7 mW/Gb/s. Boyu Hu, Yuan Du, Rulin Huang, Jeffrey Lee, Young-Kai Chen, Mau-Chung Frank Chang |
IEEE Trans. Very Large Scale Integr. Syst. | 6 |
| 2016 | Invited - Airtouch: a novel single layer 3D touch sensing system for human/mobile devices interactionsabstractTouchscreen technology plays an important role in the booming mobile devices market. Traditional touchscreen only provides 2D interactions with limited user experience. To overcome these limitations, we propose a novel 3D touch sensing system called the Airtouch system, which can recognize the movement of the finger in a 3-dimensional space. Half of the manufacturing cost is reduced by applying only single layer electrodes in the touch panel design. Moreover, an oscillator based correlated double sampling circuit is implemented as the self-capacitive sensor with bootstrapping technique to reduce inter-channel-coupling effect. Additionally, new algorithm for finger positioning is created with grouping filter invented to reduce system background noise. The demonstrated setup can successfully detect finger movement within a vertical range of 6cm and achieve a horizontal resolution up to 1cm. This system offers great potential in both gesture recognition for small-sized electronics, and advanced human interactive games for TV and mobile device. Adrian Tang 0002, Yan Zhang 0050, Yilei Li, Kye Cheung, Mau-Chung Frank Chang |
DAC | 7 |
| 2016 | Invited - A 2.2 GHz SRAM with high temperature variation immunity for deep learning application under 28nmabstractWith the coming era of Big Data, hardware implementation of machine learning has become attractive for many applications, such as real-time object recognition and face recognition. The implementation of machine learning algorithms needs intensive memory access, and SRAM is critical for the overall performance. This paper proposes a new design of high speed SRAM for machine learning purposes. With fast access time (cycle time: 650 ps, access time: 350 ps), low sensitivity to temperature variation and high configurability (less than 10% performance difference between 125_rcw_tt vs 0_rcw_tt), the proposed SRAM is a better candidate for hardware machine learning system than the conventional SRAM. Compared with Samsung HL 152, our design has smaller size (121×43 um2 vs 127×44 um2) with half the number of pins ports (12 vs 25) and higher speed (2.2GHz vs 0.8GHz). Yen-Hsiang Wang, Yilei Li, Chien-Heng Wong, Tien Pei Chou, Young-Kai Chen, Mau-Chung Frank Chang |
DAC | 7 |
| 2016 | The SMEM Seeding Acceleration for DNA Sequence AlignmentabstractThe advance of next-generation sequencing technology has dramatically reduced the cost of genome sequencing. However, processing and analyzing huge amounts of data collected from sequencers introduces significant computation challenges, these have become the bottleneck in many research and clinical applications. For such applications, read alignment is usually one of the most compute-intensive steps. Billions of reads generated from the sequencer need to be aligned to the long reference genome. Recent state-of-the-art software read aligners follow the seed-andextend model. In this paper we focus on accelerating the first seeding stage, which generates the seeds using the supermaximal exact match (SMEM) seeding algorithm. The two main challenges for accelerating this process are 1) how to process a huge number of short reads with high throughput, and 2) how to hide the frequent and long random memory access when we try to fetch the value of the reference genome. In this paper, we propose a scalable array-based architecture, which is composed by many processing engines (PEs) to process large amounts of data simultaneously for the demand of high throughput. Furthermore, we provide a tight software/hardware integration that realizes the proposed architecture on the Intel-Altera HARP system. With a 16-PE accelerator engine, we accelerate the SMEM algorithm by 4x, and the overall SMEM seeding stage by 26% when compared with 16-thread CPU execution. We further analyze the performance bottleneck of the design due to extensive DRAM accesses and discuss the possible improvements that are worthwhile to be explored in the future. Mau-Chung Frank Chang, Yuting Chen 0003, Jason Cong, Po-Tsang Huang, Chun-Liang Kuo, Cody Hao Yu |
FCCM | 1 |
| 2015 | Wireless Gigabit Data Telemetry for Large-Scale Neural RecordingabstractImplantable wireless neural recording from a large ensemble of simultaneously acting neurons is a critical component to thoroughly investigate neural interactions and brain dynamics from freely moving animals. Recent researches have shown the feasibility of simultaneously recording from hundreds of neurons and suggested that the ability of recording a larger number of neurons results in better signal quality. This massive recording inevitably demands a large amount of data transfer. For example, recording 2000 neurons while keeping the signal fidelity ( > 12 bit, > 40 KS/s per neuron) needs approximately a 1-Gb/s data link. Designing a wireless data telemetry system to support such (or higher) data rate while aiming to lower the power consumption of an implantable device imposes a grand challenge on neuroscience community. In this paper, we present a wireless gigabit data telemetry for future large-scale neural recording interface. This telemetry comprises of a pair of low-power gigabit transmitter and receiver operating at 60 GHz, and establishes a short-distance wireless link to transfer the massive amount of neural signals outward from the implanted device. The transmission distance of the received neural signal can be further extended by an externally rendezvous wireless transceiver, which is less power/heat-constraint since it is not at the immediate proximity of the cortex and its radiated signal is not seriously attenuated by the lossy tissue. The gigabit data link has been demonstrated to achieve a high data rate of 6 Gb/s with a bit-error-rate of 10(-12) at a transmission distance of 6 mm, an applicable separation between transmitter and receiver. This high data rate is able to support thousands of recording channels while ensuring a low energy cost per bit of 2.08 pJ/b. Yen-Cheng Kuan, Yi-Kai Lo, Yanghyo Kim, Mau-Chung Frank Chang, Wentai Liu |
IEEE J. Biomed. Health Informatics | 4 |
| 2013 | A 100Gb/s quad-rate transformer-coupled injection-locking CDR circuit in 65nm CMOSabstractThis paper presents an injection-locking clock and data recovery circuit (CDR) for serial link receivers. A transformer-coupled injection-locking scheme with all passive components is proposed to lock the quadrature voltage controlled oscillator (QVCO) to align the received data. The quad-rate CDR successfully regenerates the serial 100 Gb/s PRBS 231-1 data into 4 parallel data streams at 25 Gb/s. The fabricated chip occupies 1.92 mm2in 65 nm standard CMOS process with recovered data peak-to-peak jitter of 0.84ps and consumes 130 mW power with 1.0-V supply. Fanta Chen, Jen-Ming Wu, Jenny Yi-Chun Liu, Mau-Chung Frank Chang |
ISCAS | 4 |
| 2013 | Stream arbitration: Towards efficient bandwidth utilization for emerging on-chip interconnectsabstractAlternative interconnects are attractive for scaling on-chip communication bandwidth in a power-efficient manner. However, efficient utilization of the bandwidth provided by these emerging interconnects still remains an open problem due to the spatial and temporal communication heterogeneity. In this article, a Stream Arbitration scheme is proposed, where at runtime any source can compete for any communication channel of the interconnect to talk to any destination. We apply stream arbitration to radio frequency interconnect (RF-I). Experimental results show that compared to the representative token arbitration scheme, stream arbitration can provide an average 20% performance improvement and 12% power reduction. Chunhua Xiao, Mau-Chung Frank Chang, Jason Cong, Michael Gill, Zhangqin Huang, Chunyue Liu, Glenn Reinman, Hao Wu 0026 |
ACM Trans. Archit. Code Optim. | 2 |
| 2012 | Utilizing RF-I and intelligent scheduling for better throughput/watt in a mobile GPU memory systemabstractSmartphones and tablets are becoming more and more powerful, replacing desktops and laptops as the users' main computing system. As these systems support higher and higher resolutions with more complex 3D graphics, a high-throughput and low-power memory system is essential for the mobile GPU. In this article, we propose to improve throughput/watt in a mobile GPU memory system by using intelligent scheduling to reduce power and multi-band radio frequency interconnect (MRF-I) to offset any throughput degradation caused by our intelligent scheduling. Overall, we are able to improve throughput 17% up to 66% while increasing throughput per watt by an average of 18% up to 26%. Kanit Therdsteerasukdi, Gyungsu Byun, Jason Cong, Mau-Chung Frank Chang, Glenn Reinman |
ACM Trans. Archit. Code Optim. | 4 |
| 2011 | The DIMM tree architecture: A high bandwidth and scalable memory systemabstractThe demand for capacity and off-chip bandwidth to DRAM will continue to grow as we integrate more cores onto a die. However, as the data rate of DRAM has increased, the number of DIMMs supported on a multi-drop bus has decreased. Therefore, traditional memory systems are not sufficient to meet both these demands. We propose the DIMM tree architecture for better scalability by connecting the DIMMs as a tree. The DIMM tree architecture is able to grow the number of DIMMs exponentially with each level of latency in the tree. We also propose application of Multiband Radio Frequency Interconnect (MRF-I) to the DIMM tree architecture for even greater scalability and higher throughput. The DIMM tree architecture without MRF-I was able to scale up to 64 DIMMs with only an 8% degradation in throughput over an ideal system. The DIMM tree architecture with MRF-I was able to increase throughput by 68% (up to 200%) on a 64-DIMM system over a 4-DIMM system. Kanit Therdsteerasukdi, Gyungsu Byun, Jeremy Ir, Glenn Reinman, Jason Cong, Mau-Chung Frank Chang |
ICCD | 6 |
| 2011 | A progammable baseband anti-alias filter for a passive-mixer-based, SAW-less, multi-band, multi-mode WEDGE transmitterabstractA programmable baseband anti-alias filter (AAF) for a passive-mixer-based, 1.8V, SAW-less, multi-band, multi-mode WEDGE (WCDMA/HSUPA/EGPRS) cellular transmitter (TX) is described. This paper presents an AAF which results in ultra-low, -170dBc/Hz, receive-band noise, and enables the first single-mixer-based, multi-mode, SAW-less TX. By providing the noise and linearity performance only on-demand from system requirements, the presented AAF enables power savings of 14mW, or 34% of the total 41mW TX power. The AAF has been fabricated as part of a single-chip 0.13μm CMOS transceiver. Sandeep D'Souza, Mau-Chung Frank Chang, Sudhakar Pamarti, Bipul Agarwal, Hossein Zarei, Tirdad Sowlati, Roc Berenguer |
ISCAS | 2 |
| 2011 | RF/wireless-interconnect: The next wave of connectivity
Sai-Wang Tam, Mau-Chung Frank Chang |
Sci. China Inf. Sci. | 2 |
| 2009 | A scalable micro wireless interconnect structure for CMPsabstractThis paper describes an unconventional way to apply wireless networking in emerging technologies. It makes the case for using a two-tier hybrid wireless/wired architecture to interconnect hundreds to thousands of cores in chip multiprocessors (CMPs), where current interconnect technologies face severe scaling limitations in excessive latency, long wiring, and complex layout. We propose a recursive wireless interconnect structure called the WCube that features a single transmit antenna and multiple receive antennas at each micro wireless router and offers scalable performance in terms of latency and connectivity. We show the feasibility to build miniature on-chip antennas, and simple transmitters and receivers that operate at 100 − 500 GHz sub-terahertz frequency bands. We also devise new two-tier wormhole based routing algorithms that are deadlock free and ensure a minimum-latency route on a 1000core on-chip interconnect network. Our simulations show that our protocol suite can reduce the observed latency by 20 % to 45%, and consumes power that is comparable to or less than current 2-D wiredmeshdesigns. Suk-Bok Lee, Sai-Wang Tam, Ioannis Pefkianakis, Songwu Lu, Mau-Chung Frank Chang, Chuanxiong Guo, Glenn Reinman, Chunyi Peng 0001, Mishali Naik, Lixia Zhang 0001, Jason Cong |
MobiCom | 5 |
| 2008 | CMP network-on-chip overlaid with multi-band RF-interconnectabstractIn this paper, we explore the use of multi-band radio frequency interconnect (or RF-I) with signal propagation at the speed of light to provide shortcuts in a many core network-on-chip (NoC) mesh topology. We investigate the costs associated with this technology, and examine the latency and bandwidth benefits that it can provide. Assuming a 400mm2die, we demonstrate that in exchange for 0.13% of area overhead on the active layer, RF-I can provide an average 13% (max 18%) boost in application performance, corresponding to an average 22% (max 24%) reduction in packet latency. We observe that RF access points may become traffic bottlenecks when many packets try to use the RF at once, and conclude by proposing strategies that adapt RF-I utilization at runtime to actively combat this congestion. Mau-Chung Frank Chang, Jason Cong, Adam Kaplan, Mishali Naik, Glenn Reinman, Eran Socher, Sai-Wang Tam |
HPCA | 1 |
| 2008 | Lower-Complexity Layered Belief-Propagation Decoding of LDPC CodesabstractThe design of LDPC decoders with low complexity, high throughput, and good performance is a critical task. A well-known strategy is to design structured codes such as quasi- cyclic LDPC (QC-LDPC) that allow partially-parallel decoders. Sequential schedules, such as Layered Belief-Propagation (LBP), converge faster than the traditional flooding schedule while allowing parallel decoding of QC-LDPC codes. In this paper, we propose a novel low-complexity sequential schedule called Zigzag LBP (Z-LBP). Current LBP schedules do not allow partially- parallel architectures in the regime of high-rate codes with small- to-medium blocklengths. Our proposed algorithm can still be implemented in a partially-parallel manner in this regime. Z-LBP provides the same benefits as LBP including faster convergence speed and lower frame error rates than flooding. Yuan-Mao Chang, Andres I. Vila Casado, Mau-Chung Frank Chang, Richard D. Wesel |
ICC | 3 |
| 2008 | RF interconnects for communications on-chipabstractIn this paper, we propose a new way of implementing on-chip global interconnect that would meet stringent challenges of core-to-core communications in latency, data rate, and re-configurability for future chip-microprocessors (CMP) with efficient area and energy overheads. We discuss the limitation of traditional RC-limited interconnects and possible benefits of multi-band RF-interconnect (RF-I) through on-chip differential transmission lines. The physical implementation of RF-I and its projected performance versus overhead as the function of CMOS technology scaling are discussed as well Mau-Chung Frank Chang, Eran Socher, Sai-Wang Tam, Jason Cong, Glenn Reinman |
ISPD | 1 |
| 2008 | Power reduction of CMP communication networks via RF-interconnectsabstractAs chip multiprocessors scale to a greater number of processing cores, on-chip interconnection networks will experience dramatic increases in both bandwidth demand and power dissipation. Fortunately, promising gains can be realized via integration of radio frequency interconnect (RF-I) through on-chip transmission lines with traditional interconnects implemented with RC wires. While prior work has considered the latency advantage of RF-I, we demonstrate three further advantages of RF-I: (1) RF-I bandwidth can be flexibly allocated to provide an adaptive NoC, (2) RF-I can enable a dramatic power and area reduction by simplification of NoC topology, and (3) RF-I provides natural and efficient support for multicast. In this paper, we propose a novel interconnect design, exploiting dynamic RF-I bandwidth allocation to realize a reconfigurable network-on-chip architecture. We find that our adaptive RF-I architecture on top of a mesh with 4B links can even outperform the baseline with 16B mesh links by about 1%, and reduces NoC power by approximately 65% including the overhead incurred for supporting RF-I. Mau-Chung Frank Chang, Jason Cong, Adam Kaplan, Chunyue Liu, Mishali Naik, Jagannath Premkumar, Glenn Reinman, Eran Socher, Sai-Wang Tam |
MICRO | 1 |
| 2008 | A Cost-Effective Latency-Aware Memory Bus for Symmetric Multiprocessor SystemsabstractThis paper presents how a multi-core system can benefit from the use of a latency-aware memory bus capable of dual-concurrent data transfers on a single wire line: Source synchronous CDMA interconnect (SSCDMA-I) has been adopted to implement the memory bus of a shared-memory multi-core system. Two types of bus-based homogeneous and heterogeneous multi-core systems are modeled and simulated by a cycle-accurate simulation platform. Unlike the conventional time-division multiplexing (TDM) bus-based multi-core system that shows degradation in performance as the number of processing cores increases, the proposed SSCDMA bus-based multi-core shows higher performance up to 23.1% for 4 cores. The maximum latency of a heterogeneous multi-core system with a mix of traffic loads has been reduced up to 78%. These results demonstrate that the performance of multi-core systems can be improved with less cost and network complexity by reducing the bus contention interferences and by supporting higher concurrency in memory accesses that brings shorter critical word access latency. Jongsun Kim, Bo-Cheng Lai, Mau-Chung Frank Chang, Ingrid Verbauwhede |
IEEE Trans. Computers | 3 |
| 2007 | Design of an Interconnect Architecture and Signaling Technology for Parallelism in CommunicationabstractThe need for efficient interconnect architectures beyond the conventional time-division multiplexing (TDM) protocol-based interconnects has been brought on by the continued increase of required communication bandwidth and concurrency of small-scale digital systems. To improve the overall system performance without increasing communication resources and complexity, this paper presents a cost-effective interconnect architecture, communication protocol, and signaling technology that exploits parallelism in board-level communication, resulting in shorter latency and higher concurrency on a shared bus or link: the proposed source synchronous CDMA interconnect (SSCDMA-I) enables dual concurrent transactions on a single wire line as well as flexible input/output (I/O) reconfiguration. The SSCDMA-I utilizes 2-bit orthogonal CDMA coding and a variation of source synchronous clocking for multilevel superposition; a single 3-level SSCDMA-I line operates as if it consists of dual virtual time-multiplexed interconnects, which exploits communication parallelism with a reduced number of pins, wires, and complexity. The unique multiple access capability of the SSCDMA-I improves real-time communication between multiple semiconductor intellectual property (IP) blocks on a shared link or bus by reducing the bus contention interference from simultaneous traffic requests and by taking advantage of shorter request latency. The prototype transceiver chip is implemented in 0.18-m CMOS and the 10-cm test PC board system achieves an aggregate data rate of 2.5 Gb/s/pin between four off-chip (2Tx-to-2Rx) I/Os. Jongsun Kim, Ingrid Verbauwhede, Mau-Chung Frank Chang |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2005 | CDMA/FDMA-interconnects for future ULSI communicationsabstractFuture inter- and intra-ULSI interconnect systems demand extremely high data rates as well as bidirectional multi-I/O concurrent service, reconfigurable computing/processing architecture, and total compatibility with mainstream silicon SOC (system-on-chip) and SIP (system-in-package) technologies. In this talk, we review recent advances in CDMA and FDMA interconnect schemes that promise to meet all of the above system requirements. The physical transmission line is no longer limited to a direct-coupled metal wire. Rather, it can be accomplished via either wired or wireless mediums through capacitor couplers that reduce the baseband noise and DC power consumption. These new advances in interconnect schemes would fundamentally alter the paradigm of ULSI data communications and enable the design of next generation computing/processing systems. Mau-Chung Frank Chang |
ICCAD | 1 |
| 2004 | A simple DDS architecture with highly efficient sine function lookup tableabstractA simple architecture for direct digital frequency synthesis (DDS) is presented. The proposed architecture uses a sampling-only-algorithm (SOA) to achieve a high compression ratio of 558 in realizing the sine function look up table, which is higher than the prior arts. Mau-Chung Frank Chang, Jessica Chiatai Chou |
ACM Great Lakes Symposium on VLSI | 2 |
| 2001 | RF/wireless interconnect for inter- and intra-chip communicationsabstractRecent studies showed that conventional approaches being used to solve problems imposed by hard-wired metal interconnects will eventually encounter fundamental limits and may impede the advance of future ultralarge-scale integrated circuits (ULSls). To surpass these fundamental limits, we introduce a novel RF/wireless interconnect concept for future inter- and intra-ULSI communications. Unlike the traditional "passive" metal interconnect, the "active" RF/wireless interconnect is based on low loss and dispersion-free microwave signal transmission, near-field capacitive coupling, and modem multiple-access algorithms. In this paper we address issues relevant to the signal channeling of the RF/wireless interconnect and discuss its advantages in speed, signal integrity, and channel reconfiguration. The electronic overhead required in the RF/wireless-interconnect system and its compatibility with the future ULSI and MCM (multi-chip-module) will be discussed as well. Mau-Chung Frank Chang, Vwani P. Roychowdhury, Hyunchol Shin, Yongxi Qian |
Proc. IEEE | 1 |
| 1993 | GaAs-based heterojunction bipolar transistors for very high performance electronic circuitsabstractThis paper reviews the principles and status of AlGaAs/GaAs heterojunction bipolar transistor technology. Comparisons of this technology with Si bipolar transistor and GaAs field-effect transistor technologies are made. Epitaxial materials, fabrication processes, transistor DC and RF characteristics, and modeling of AlGaAs/GaAs HBT's are described. Key areas of HBT application are also highlighted.> Peter M. Asbeck, Mau-Chung Frank Chang, Keh-Chung Wang, Gerard J. Sullivan, Derek T. Cheung |
Proc. IEEE | 2 |