EDBT 2026 Demo / reviewers in the wild / expert
Bingjun Xiao
dblp:30/8314
· DBLP profile ↗
18ranked-venue papers
2as first author
0since 2021 · last 2016
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 16 · 2 first-authorArtificial intelligence and machine learning · 1Software engineering, systems software and programming languages · 1Applied, interdisciplinary, general and emerging computing · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
8 papers |
Electronic design automation · 47% Hardware accelerators and domain-specific architectures · 16% Reconfigurable computing and FPGAs · 12% | |
| Software engineering, system software, and programming languages
1 paper |
Compilers and program optimization · 100% |
Topics — the 26 heaviest of 27, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Electronic design automation
high-level synthesis |
0.7 | 3 | 2016 | An Optimal Microarchitecture for Stencil Computation Acceleration Based on Nonuniform Partitioning of Data Reuse Buffers · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2016 CMOST: a system-level FPGA compilation framework · DAC 2015 An Optimal Microarchitecture for Stencil Computation Acceleration Based on Non-Uniform Partitioning of Data Reuse Buffers · DAC 2014 |
Electronic design automation › high-level synthesis › memory synthesis
memory partitioning |
0.4 | 2 | 2016 | An Optimal Microarchitecture for Stencil Computation Acceleration Based on Nonuniform Partitioning of Data Reuse Buffers · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2016 An Optimal Microarchitecture for Stencil Computation Acceleration Based on Non-Uniform Partitioning of Data Reuse Buffers · DAC 2014 |
Hardware accelerators and domain-specific architectures › scientific computing accelerator
stencil computation accelerator |
0.4 | 2 | 2016 | An Optimal Microarchitecture for Stencil Computation Acceleration Based on Nonuniform Partitioning of Data Reuse Buffers · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2016 An Optimal Microarchitecture for Stencil Computation Acceleration Based on Non-Uniform Partitioning of Data Reuse Buffers · DAC 2014 |
Hardware reliability and fault tolerance
defect tolerance |
0.3 | 2 | 2013 | Defect recovery in nanodevice-based programmable interconnects (abstract only) · FPGA 2013 Defect tolerance in nanodevice-based programmable interconnects: utilization beyond avoidance · DAC 2013 |
Electronic design automation
physical design |
0.3 | 2 | 2013 | Defect recovery in nanodevice-based programmable interconnects (abstract only) · FPGA 2013 Defect tolerance in nanodevice-based programmable interconnects: utilization beyond avoidance · DAC 2013 |
Electronic design automation › physical design › routing
timing-driven routing |
0.3 | 2 | 2013 | Defect recovery in nanodevice-based programmable interconnects (abstract only) · FPGA 2013 Defect tolerance in nanodevice-based programmable interconnects: utilization beyond avoidance · DAC 2013 |
Memory systems
on-chip memory |
0.3 | 2 | 2016 | An Optimal Microarchitecture for Stencil Computation Acceleration Based on Nonuniform Partitioning of Data Reuse Buffers · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2016 An Optimal Microarchitecture for Stencil Computation Acceleration Based on Non-Uniform Partitioning of Data Reuse Buffers · DAC 2014 |
Compilers and program optimization
hardware compilation |
0.2 | 1 | 2015 | CMOST: a system-level FPGA compilation framework · DAC 2015 |
Hardware accelerators and domain-specific architectures › machine learning accelerator
CNN accelerator |
0.2 | 1 | 2015 | Optimizing FPGA-based Accelerator Design for Deep Convolutional Neural Networks · FPGA 2015 |
Electronic design automation › high-level synthesis › hardware compilation
c-to-hardware compilation |
0.2 | 1 | 2015 | CMOST: a system-level FPGA compilation framework · DAC 2015 |
Electronic design automation
design space exploration |
0.2 | 1 | 2015 | Optimizing FPGA-based Accelerator Design for Deep Convolutional Neural Networks · FPGA 2015 |
Hardware accelerators and domain-specific architectures › machine learning accelerator › CNN accelerator
FPGA-based CNN accelerator |
0.2 | 1 | 2015 | Optimizing FPGA-based Accelerator Design for Deep Convolutional Neural Networks · FPGA 2015 |
Reconfigurable computing and FPGAs
FPGA compilation |
0.2 | 1 | 2015 | CMOST: a system-level FPGA compilation framework · DAC 2015 |
Performance modeling and evaluation › analytical modeling
roofline model |
0.2 | 1 | 2015 | Optimizing FPGA-based Accelerator Design for Deep Convolutional Neural Networks · FPGA 2015 |
Interconnection networks and networks-on-chip › interconnect architecture
reconfigurable interconnect |
0.2 | 2 | 2013 | FPGA-RR: an enhanced FPGA architecture with RRAM-based reconfigurable interconnects (abstract only) · FPGA 2012 Defect recovery in nanodevice-based programmable interconnects (abstract only) · FPGA 2013 |
Electronic design automation › physical design › routing
FPGA routing |
0.2 | 1 | 2013 | Defect recovery in nanodevice-based programmable interconnects (abstract only) · FPGA 2013 |
Reconfigurable computing and FPGAs
FPGA routing architecture |
0.2 | 1 | 2013 | Defect tolerance in nanodevice-based programmable interconnects: utilization beyond avoidance · DAC 2013 |
Electronic design automation › physical design
routing |
0.2 | 1 | 2013 | Defect tolerance in nanodevice-based programmable interconnects: utilization beyond avoidance · DAC 2013 |
Reconfigurable computing and FPGAs
FPGA architecture |
0.1 | 1 | 2012 | FPGA-RR: an enhanced FPGA architecture with RRAM-based reconfigurable interconnects (abstract only) · FPGA 2012 |
Memory systems › non-volatile memory
resistive memory |
0.1 | 1 | 2012 | FPGA-RR: an enhanced FPGA architecture with RRAM-based reconfigurable interconnects (abstract only) · FPGA 2012 |
Energy systems and smart grids › energy storage
battery management |
0.1 | 1 | 2010 | A universal state-of-charge algorithm for batteries · DAC 2010 |
Integrated circuit design
analog and mixed-signal circuits |
0.1 | 1 | 2010 | A universal state-of-charge algorithm for batteries · DAC 2010 |
Energy-efficient computing › battery management
battery modeling |
0.1 | 1 | 2010 | A universal state-of-charge algorithm for batteries · DAC 2010 |
Reconfigurable computing and FPGAs
FPGA accelerator |
0.1 | 1 | 2015 | CMOST: a system-level FPGA compilation framework · DAC 2015 |
Electronic design automation › physical design
placement |
0.0 | 1 | 2013 | Defect tolerance in nanodevice-based programmable interconnects: utilization beyond avoidance · DAC 2013 |
Reconfigurable computing and FPGAs
FPGA design flow |
0.0 | 1 | 2012 | FPGA-RR: an enhanced FPGA architecture with RRAM-based reconfigurable interconnects (abstract only) · FPGA 2012 |
Methods — techniques the papers use, named apart from their topics
non-uniform memory partitioning · 0.4task-level dependence analysis · 0.4block-based data streaming · 0.4automated SDF generation · 0.4HLS · 0.2roofline model · 0.2loop transformation · 0.2loop tiling · 0.2automated design flow · 0.2shorting constraints · 0.2open-circuit voltage estimation · 0.1linear system analysis · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2016 | Novel applications of deep learning hidden features for adaptive testingabstractAdaptive test of integrated circuits (IC) promises to increase the quality and yield of products with reduced manufacturing test cost compared to traditional static test flows. Two mostly widely used techniques are Statistical Process Control (SPC) and Part Average Testing (PAT), whose capabilities to capture complex correlation between test measurements and the underlying IC's physical and electrical properties are, however, limited. Based on recent progress on machine learning, this paper proposes a novel deep learning based method for adaptive test. Compared to most machine learning techniques, deep learning has the distinctive advantage of being able to capture the underlying key features automatically from data without manual intervention. In this paper, we start from a trained deep neuron network (DNN) with a much higher accuracy than the conventional test flow for the pass and fail prediction. We further develop two novel applications by leveraging the features learned from DNN: one to enable partial testing, i.e., make decisions on pass and fail without finishing the entire test flow, and two to enable dynamic test ordering, i.e., changing the sequence of tests adaptively. Experiment results show significant improvement on the accuracy and effectiveness of our proposed method. Bingjun Xiao, Jinjun Xiong, Yiyu Shi 0001 |
ASP-DAC | 1 |
| 2016 | An Optimal Microarchitecture for Stencil Computation Acceleration Based on Nonuniform Partitioning of Data Reuse BuffersabstractHigh-level synthesis (HLS) tools have made significant progress in compiling high-level descriptions of computation into highly pipelined register-transfer level specifications. The high-throughput computation raises a high data demand. To prevent data accesses from being the bottleneck, on-chip memories are used as data reuse buffers to reduce off-chip accesses. Also memory partitioning is explored to increase the memory bandwidth by scheduling multiple simultaneous memory accesses to different memory banks. Prior work on memory partitioning of data reuse buffers is limited to uniform partitioning. In this paper, we perform an early-stage exploration of nonuniform memory partitioning. We use the stencil computation, a popular communication-intensive application domain, as a case study to show the potential benefits of nonuniform memory partitioning. Our novel method can always achieve the minimum memory size and the minimum number of memory banks, which cannot be guaranteed in any prior work. We develop a generalized microarchitecture to decouple stencil accesses from computation, and an automated design flow to integrate our microarchitecture with the HLS-generated computation kernel for a complete accelerator. Jason Cong, Peng Li 0031, Bingjun Xiao, Peng Zhang 0007 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2015 | CMOST: a system-level FPGA compilation frameworkabstractProgramming difficulty is a key challenge to the adoption of FPGAs as a general high-performance computing platform. In this paper we present CMOST, an open-source automated compilation flow that maps C-code to FPGAs for acceleration. CMOST establishes a unified framework for the integration of various system-level optimizations and for different hardware platforms. We also present several novel techniques on integrating optimizations in CMOST, including task-level dependence analysis, block-based data streaming, and automated SDF generation. Experimental results show that automatically generated FPGA accelerators can achieve over 8x speedup and 120x energy gain on average compared to the multi-core CPU results from similar input C programs. CMOST results are comparable to those obtained after extensive manual source-code transformations followed by high-level synthesis. Peng Zhang 0007, Muhuan Huang, Bingjun Xiao, Hui Huang 0001, Jason Cong |
DAC | 3 |
| 2015 | Optimizing FPGA-based Accelerator Design for Deep Convolutional Neural NetworksabstractConvolutional neural network (CNN) has been widely employed for image recognition because it can achieve high accuracy by emulating behavior of optic nerves in living creatures. Recently, rapid growth of modern applications based on deep learning algorithms has further improved research and implementations. Especially, various accelerators for deep CNN have been proposed based on FPGA platform because it has advantages of high performance, reconfigurability, and fast development round, etc. Although current FPGA accelerators have demonstrated better performance over generic processors, the accelerator design space has not been well exploited. One critical problem is that the computation throughput may not well match the memory bandwidth provided an FPGA platform. Consequently, existing approaches cannot achieve best performance due to under-utilization of either logic resource or memory bandwidth. At the same time, the increasing complexity and scalability of deep learning applications aggravate this problem. In order to overcome this problem, we propose an analytical design scheme using the roofline model. For any solution of a CNN design, we quantitatively analyze its computing throughput and required memory bandwidth using various optimization techniques, such as loop tiling and transformation. Then, with the help of rooine model, we can identify the solution with best performance and lowest FPGA resource requirement. As a case study, we implement a CNN accelerator on a VC707 FPGA board and compare it to previous approaches. Our implementation achieves a peak performance of 61.62 GFLOPS under 100MHz working frequency, which outperform previous approaches significantly. Chen Zhang 0001, Peng Li 0031, Guangyu Sun 0003, Yijin Guan, Bingjun Xiao, Jason Cong |
FPGA | 5 |
| 2015 | ARACompiler: a prototyping flow and evaluation framework for accelerator-rich architecturesabstractAccelerator-rich architectures (ARAs) provide energy-efficient solutions for domain-specific computing in the age of dark silicon. However, due to the complex interaction between the general-purpose cores, accelerators, customized onchip interconnects, customized memory systems, and operating systems, it has been difficult to get detailed and accurate evaluations and analyses of ARAs on complex real-life benchmarks using the existing full-system simulators. In this paper we develop the ARACompiler, which is a highly automated design flow for prototyping ARAs and performing evaluation on FPGAs. An efficient system software stack is generated automatically to handle resource management and TLB misses.We further provide application programming interfaces (APIs) for users to develop their applications using accelerators. The flow can provide 2.9x to 42.6x evaluation time saving over the full-system simulations. Yuting Chen 0003, Jason Cong, Bingjun Xiao |
ISPASS | 3 |
| 2014 | An Optimal Microarchitecture for Stencil Computation Acceleration Based on Non-Uniform Partitioning of Data Reuse BuffersabstractHigh-level synthesis (HLS) tools have made significant progress in compiling high-level descriptions of computation into highly pipelined register-transfer level (RTL) specifications. The high-throughput computation raises a high data demand. To prevent data accesses from being the bottleneck, on-chip memories are used as data reuse buffers to reduce off-chip accesses. Also memory partitioning is explored to increase the memory bandwidth by scheduling multiple simultaneous memory accesses to different memory banks. Prior work on memory partitioning of data reuse buffers is limited to uniform partitioning. In this paper, we perform an early-stage exploration of non-uniform memory partitioning. We use the stencil computation, a popular communication-intensive application domain, as a case study to show the potential benefits of non-uniform memory partitioning. Our novel method can always achieve the minimum memory size and the minimum number of memory banks, which cannot be guaranteed in any prior work. We develop a generalized microarchitecture to decouple stencil accesses from computation, and an automated design flow to integrate our microarchitecture with the HLS-generated computation kernel for a complete accelerator. Jason Cong, Peng Li 0031, Bingjun Xiao, Peng Zhang 0007 |
DAC | 3 |
| 2014 | A Fully Pipelined and Dynamically Composable Architecture of CGRAabstractFuture processor chips will not be limited by the transistor resources, but will be mainly constrained by energy efficiency. Reconfigurable fabrics bring higher energy efficiency than CPUs via customized hardware that adapts to user applications. Among different reconfigurable fabrics, coarse-grained reconfigurable arrays (CGRAs) can be even more efficient than fine-grained FPGAs when bit-level customization is not necessary in target applications. CGRAs were originally developed in the era when transistor resources were more critical than energy efficiency. Previous work shares hardware among different operations via modulo scheduling and time multiplexing of processing elements. In this work, we focus on an emerging scenario where transistor resources are rich. We develop a novel CGRA architecture that enables full pipelining and dynamic composition to improve energy efficiency by taking full advantage of abundant transistors. Several new design challenges are solved. We implement a prototype of the proposed architecture in a commodity FPGA chip for verification. Experiments show that our architecture can fully exploit the energy benefits of customization for user applications in the scenario of rich transistor resources. Jason Cong, Hui Huang 0001, Chiyuan Ma, Bingjun Xiao, Peipei Zhou 0001 |
FCCM | 4 |
| 2014 | Minimizing Computation in Convolutional Neural Networks
Jason Cong, Bingjun Xiao |
ICANN | 2 |
| 2014 | FPGA-RPI: A Novel FPGA Architecture With RRAM-Based Programmable InterconnectsabstractIn this paper we introduce a novel field programmable gate array (FPGA) architecture with resistive random access memory (RRAM)-based programmable interconnects (FPGA-RPI). Programmable interconnects are the dominant part of FPGA. We use RRAMs to build programmable interconnects, and optimize their structures by exploiting opportunities that emerge in RRAM-based circuits. FPGA-RPI can be fabricated by the existing CMOS-compatible RRAM process. Using an advanced placement and routing tool named VPR-RPI which was developed to deal with the novel architecture, a customized CAD flow is provided for FPGA-RPI. Results show that the programmable interconnects of FPGA-RR have a 96% smaller footprint, 55% higher performance, and 79% lower power consumptions compared to other FPGA counterparts. Jason Cong, Bingjun Xiao |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2013 | Optimizing routability in large-scale mixed-size placementabstractOne of the necessary requirements for the placement process is that it should be capable of generating routable solutions. This paper describes a simple but effective method leading to the reduction of the routing congestion and the final routed wirelength for large-scale mixed-size designs. In order to reduce routing congestion and improve routability, we propose blocking narrow regions on the chip. We also propose dummy-cell insertion inside regions characterized by reduced fixed-macro density. Our placer consists of three major components: (i) narrow channel reduction by performing neighbor-based fixed-macro inflation; (ii) dummy-cell insertion inside large regions with reduced fixed-macro density; and (iii) pre-placement inflation by detecting tangled logic structures in the netlist and minimizing the maximum pin density. We evaluated the quality of our placer using the newly released DAC 2012 routability-driven placement contest designs and we compared our results to the top four teams that participated in the placement contest. The experimental results reveal that our placer improves the routability of the DAC 2012 placement contest designs and effectively reduces the routing congestion. Jason Cong, Guojie Luo, Kalliopi Tsota, Bingjun Xiao |
ASP-DAC | 4 |
| 2013 | Defect tolerance in nanodevice-based programmable interconnects: utilization beyond avoidanceabstractThis work focuses on defect tolerance for nanodevice-based programmable interconnects of FPGAs. First, we show that the stuck-closed defects of nanodevices have a much higher impact than the stuck-open defects. Instead of simply avoiding the stuck-closed defects, we use them by treating them as shorting constraints in the routing. We develop a scalable algorithm to perform timing-driven routing under these extra constraints. We also enhance the placement algorithm to recover logic blocks which become virtually unusable due to shorted pins. Simulation results show that at the up-to-date level of nanodevice defects (108--1011x higher than CMOS), compared to the simple avoidance method, our approach reduces the degradation of resource usage by 87%, improves the routability by 37%, and reduce the degradation of circuit performance by 36%, at a negligible overhead of tool runtime. Jason Cong, Bingjun Xiao |
DAC | 2 |
| 2013 | Defect recovery in nanodevice-based programmable interconnects (abstract only)abstractThis work focuses on defect tolerance for nanodevice-based programmable interconnects of FPGAs. A single nanodevice can function as a routing switch in place of a pass transistor and its six-transistor SRAM cell in conventional FPGAs. Defects of nanodevices in programmable interconnects are manifested as losses of configurability and can be categorized into stuck- open defect and stuck- closed defect. First, we show that the stuck-closed defects of nanodevices have a much higher impact than the stuck-open defects. Instead of simply avoiding the stuck-closed defects, we recover them by treating them as shorting constraints in the routing. We develop a scalable algorithm to perform timing-driven routing under these extra constraints. We extend the idea of the resource negotiation to balance the goals of timing and routability under shorting constraints. We also develop several techniques to guide the router to map the shorting clusters to those nets with more shared paths for better utilization of routing resources while automatically balancing it with circuit performance. We also enhance the placement algorithm to recover logic blocks which become virtually unusable due to shorted pins. Simulation results show that at the up-to-date level of nanodevice defects (108-1011x higher than CMOS), compared to the simple avoidance method, our approach reduces the degradation of resource usage by 87%, improves the routability by 37%, and reduce the degradation of circuit performance by 36%, at a negligible overhead of tool runtime. Jason Cong, Bingjun Xiao |
FPGA | 2 |
| 2013 | Optimization of interconnects between accelerators and shared memories in dark siliconabstractApplication-specific accelerators provide orders-of-magnitude improvement in energy-efficiency over CPUs, and accelerator-rich computing platforms are showing promise in the dark silicon age. Memory sharing among accelerators leads to huge transistor savings, but needs novel designs of interconnects between accelerators and shared memories. Accelerators run 100x faster than CPUs and post a high demand on data. This leads to resource-consuming interconnects if we follow the same design rules as those for interconnects between CPUs and shared memories, and simply duplicate the interconnect hardware to meet the accelerator data demand. In this work we develop a novel design of interconnects between accelerators and shared memories and exploit three optimization opportunities that emerge in accelerator-rich computing platforms: 1) The multiple data ports of the same accelerators are powered on/off together, and the competition for shared resources among these ports can be eliminated to save interconnect transistor cost; 2) In dark silicon, the number of active accelerators in an accelerator-rich platform is usually limited, and the interconnects can be partially populated to just fit the data access demand limited by the power budget; 3) The heterogeneity of accelerators leads to execution patterns among accelerators and, based on the probability analysis to identify these patterns, interconnects can be optimized for the expected utilization. Experiments show that our interconnect design outperforms prior work that was optimized for CPU cores or signal routing. Jason Cong, Bingjun Xiao |
ICCAD | 2 |
| 2013 | Accelerator-rich CMPs: From concept to real hardwareabstractApplication-specific accelerators provide 10-100× improvement in power efficiency over general-purpose processors. The accelerator-rich architectures are especially promising. This work discusses a prototype of accelerator-rich CMPs (PARC). During our development of PARC in real hardware, we encountered a set of technical challenges and proposed corresponding solutions. First, we provided system IPs that serve a sea of accelerators to transfer data between userspace and accelerator memories without cache overhead. Second, we designed a dedicated interconnect between accelerators and memories to enable memory sharing. Third, we implemented an accelerator manager to virtualize accelerator resources for users. Finally, we developed an automated flow with a number of IP templates and customizable interfaces to a C-based synthesis flow to enable rapid design and update of PARC. We implemented PARC in a Virtex-6 FPGA chip with integration of platform-specific peripherals and booting of unmodified Linux. Experimental results show that PARC can fully exploit the energy benefits of accelerators at little system overhead. Yuting Chen 0003, Jason Cong, Mohammad Ali Ghodrat, Muhuan Huang, Chunyue Liu, Bingjun Xiao, Yi Zou 0001 |
ICCD | 6 |
| 2013 | Energy-efficient computing using adaptive table lookup based on nonvolatile memoriesabstractTable lookup based function computation can significantly save energy consumption. However existing table lookup methods are mostly used in ASIC designs for some fixed functions. The goal of this paper is to enable table lookup computation in general-purpose processors, which requires adaptive lookup tables for different applications. We provide a complete design flow to support this requirement. We propose a novel approach to build the reconfigurable lookup tables based on emerging nonvolatile memories (NVMs), which takes full advantages of NVMs over conventional SRAMs and avoids the limitation of NVMs. We provide compiler support to optimize table resource allocation among functions within a program. We also develop a runtime table manager that can learn from history and improve its arbitration of the limited on-chip table resources among programs. Jason Cong, Milos D. Ercegovac, Muhuan Huang, Bingjun Xiao |
ISLPED | 5 |
| 2012 | FPGA-RR: an enhanced FPGA architecture with RRAM-based reconfigurable interconnects (abstract only)abstractIn this study, we explore the use of Resistive RAMs (RRAMs) as candidates for programmable interconnects in FPGAs. An RRAM cell can be programmed between high resistance state and low resistance state, with an on/off ratio close to MOSFET. It provides an opportunity to use an RRAM as a routing switch at a much smaller area cost than its CMOS counterpart. RRAMs can be fabricated over CMOS circuits using CMOS-compatible processes to have a more compact gate array. Our recent work (presented in NanoArch'2011) demonstrated significant potential of area, delay, and power reduction from using RRAMs in FPGAs. But some design problems remain open. The programming of RRAM switches integrated in interconnects is one important problem. We show that the high-level architecture of programming circuits for RRAM switches should be modified to avoid potential logic hazard. Also the programming cells used in previous works have an area overhead even larger than RRAM itself. We manage to reduce this overhead significantly with utilization of the non-arbitrary pattern of RRAM integration in FPGA interconnects. In addition we suggest a novel buffering solution for FPGA interconnects in light of the low area cost of RRAM-based routing switch. We propose on-demand buffer insertion, where buffers can be connected to interconnects via RRAMs to dynamically reflect the demand of the netlist to map onto FPGA. Compared to conventional buffering solution which are pre-determined during fabrication and can only be optimized for general case, our solution shows further area savings and performance improvement. The resulting FPGA architecture using RRAM for programmable interconnects is named FPGA-RR. We provide a complete CAD flow for FPGA-RR. Jason Cong, Bingjun Xiao |
FPGA | 2 |
| 2011 | Domain-specific processor with 3D integration for medical image processingabstractThe growth of 3D technology had led to opportunities for stacked multiprocessor-accelerator computing platforms with high-bandwidth and low-latency TSV connections between them, resulting in high computing performance and better energy efficiency. This work evaluates the performance and energy benefits of such an advanced architecture and addresses associated design problems. To better utilize the reconfigurable hardware resource and to explore the opportunity of kernel sharing across applications, we propose to use a dedicated domain-specific computing platform. In particular, we have chosen medical image processing as the domain in this work to accelerate due to its growing for real-time processing demand yet inadequete performance on conventional computing architectures. A design flow is proposed in this work for the 3D multiprocessor-accelerator platform and a number of methods are applied to optimize the average performance of all the applications in the targeted domain under area and bandwidth constraints. Experiments show that the applications in this domain can gain a 7.4× speed-up and 18.8× energy savings on average running on our platform using CMP cores and domain-specific accelerators as compared to their counterparts coded in CPU only. Jason Cong, Karthik Gururaj, Muhuan Huang, Bingjun Xiao, Yi Zou 0001 |
ASAP | 5 |
| 2010 | A universal state-of-charge algorithm for batteriesabstractState-of-charge (SOC) measures energy left in a battery, and it is critical for modeling and managing batteries. Developing efficient yet accurate SOC algorithms remains a challenging task. Most existing work uses regression based on a time-variant circuit model, which may be hard to converge and often does not apply to different types of batteries. Knowing open-circuit voltage (OCV) leads to SOC due to the well known mapping between OCV and SOC. In this paper, we propose an efficient yet accurate OCV algorithm that applies to all types of batteries. Using linear system analysis but without a circuit model, we calculate OCV based on the sampled terminal voltage and discharge current of the battery. Experiments show that our algorithm is numerically stable, robust to history dependent error, and obtains SOC with less than 4% error compared to a detailed battery simulation for a variety of batteries. Our OCV algorithm is also efficient, and can be used as a real-time electro-analytical tool revealing what is going on inside the battery. Bingjun Xiao, Yiyu Shi 0001, Lei He 0001 |
DAC | 1 |