EDBT 2026 Demo / reviewers in the wild / expert
Xin Fan 0002
dblp:87/3021-2
· DBLP profile ↗
5ranked-venue papers
4as first author
3since 2021 · last 2025
0000-0001-6954-5074ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 5 · 4 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | A Compact SHA256 Accelerator in 22nm for Energy Bounded Use-Cases with 8.2GHash/JabstractA 6.8•103μm2SHA256 hardware accelerator, achieving 8.2GHash/J and 3MHash/s throughput at 460mV, is fabricated in 22nm CMOS. Round-based dataflow with added pipelining and hybrid shift-FIFO structure provides 31% increase in clock frequency and 21% reduction in energy consumption. Replacing DFFs with pulsed D-latch can further increase clock frequency by 15% and decrease energy cost by 13%. Combination of high energy efficiency, high clock frequency and compact area makes the proposed SHA256 engines suitable for a wide range of applications spanning from IoT to bitcoin mining. Xin Fan 0002, Michael Gansen, Tobias Gemmeke |
ISCAS | 2 |
| 2022 | Compiling All-Digital-Embedded Content Addressable Memories on Chip for Edge ApplicationabstractA spectrum of emerging applications, including edge artificial intelligence, advocates the precompute-and-search scheme with embedded small-size content addressable memory (CAM) for its hardware efficiency instead of repetitive arithmetic operations. However, the integration of the CAM macros that are conventionally implemented with full custom analog circuits renders design-space exploration and optimization to be difficult at system level. As an alternative, a complete design flow for compiling ternary CAM on chip using foundry-supplied digital standard cells is introduced in this article. Based on the novel CAM architecture and logic design, we leverage guided placement and routing with mainstream EDA tools for exploiting the inherent structure regularity of CAM-cell arrays in the layout. An analytical model is also presented, which allows us for a systematic investigation on the energy reduction by adapting our design to various presearch structures. Validated on a postlayout$32{\times }64$ternary CAM in 28 nm, our parallel-matching scheme performs at 2.6 GHz with 0.42 fJ/bit/search, and the (8-bit) presearch scheme achieves 0.19 fJ/bit/search at 1.1 GHz, both under the 0.9-V supply voltage. In addition to flexibility for tradeoffs between the search throughput and energy, our all-digital CAM design enables voltage scaling aggressively down to 0.45 V with a minimum energy consumption of 0.06 fJ/bit/search at 50 MHz. Xin Fan 0002, Niklas Meyer, Tobias Gemmeke |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2021 | Plesiochronous Spread Spectrum Clocking With Guaranteed QoS for In-Band Switching Noise ReductionabstractSpread spectrum clocking (SSC) conventionally uses frequency modulation (FM) to suppress digital switching noise in the frequency domain. While clock-FM effectively reduces spectral noise peaks, it maintains the synchronous operation per cycle with total noise unchanged. In this paper, we introduce plesiochronous design as a general applicable de-synchronization solution for the spectral switching noise optimization with guaranteed quality-of-service. By modeling on-chip aperiodic supply current as a poly-cyclostationary random process, we theoretically prove that digital plesiochronous design contributes to reducing both, total and peak switching noise, in a harmonic frequency band of interest logarithmically proportional to the number of adopted clock domains over the synchronous baseline. A complete framework is also developed to implement plesiochronous design with the optimal clock domain partitioning and FIFO-based synchronization that features a minimum depth of six by employing Johnson encoding fully compatible with mainstream design flow. Validated on a 130nm pipelined FFT test chip across 25 dies thus taking process variations into account, our plesiochronous SSC achieves on average 5.1dB total power reductions in addition to 12.8dB peak power reductions of substrate noise at the clock fundamental frequency, which match our predictions, with only marginal hardware overhead in terms of cell area and power consumption. Xin Fan 0002, Milan Babic, Eckhard Grass, Milos Krstic |
IEEE Trans. Circuits Syst. I Regul. Pap. | 1 |
| 2020 | Approximation of Transcendental Functions With Guaranteed Algorithmic QoS by Multilayer Pareto OptimizationabstractDesign space exploration of approximate computing is deemed tough. While restricting the optimization to reduce approximate errors at the individual-function level might simplify the problem to be tackled, the increased hardware cost should be well justified by the improvement of the quality of service (QoS) at the algorithm level. In light of the loose correlation between atomic errors and algorithmic QoS, it is imperative but computationally expensive to explore approximate design space incorporating a variety of alternatives for evaluation. Despite being addressed extensively in the literature, we consider the transcendental sigmoid (σ) and hyperbolic tangent (tanh) functions as typical examples manifesting the dilemma of approximate computing. In this work, we leverage Pareto-front optimization at three hierarchies, from the parameter layer up to the structure and algorithm layers, to effectively tailor the approximate design space of the σ and tanh functions for hardware-efficient as well as algorithm-feasible implementations. Our investigations are performed based on a comprehensive design library consisting of representative approximate schemes in radically different hardware structures featuring both linear and nonlinear approximations. As tested on MNIST with 99% accuracy and Wisconsin Breast Cancer data set with 96.6% accuracy, we identify a novel and compact shift-based approximation that directly applies to the two's complement numbers achieving a 1.5× area reduction and a 3.4× energy reduction compared with the prior art. Provided the flexibility in approximate functions is of concern, we also present a uniform yet concise structure for implementing the Chebyshev-polynomials-based approximation adaptive in silicon to arbitrary nonlinear functions and error constraints. Xin Fan 0002, Tobias Gemmeke |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2018 | Physical modeling of bitcell stability in subthreshold SRAMs for leakage-area optimization under PVT variationsabstractSubthreshold SRAM design is crucial for addressing the memory bottleneck in energy constrained applications. While statistical optimization can be applied based on Monte-Carlo (MC) simulation, exploration of bitcell design space is time consuming. This paper presents a framework for model-based design and optimization of subthreshold SRAM bitcells under random PVT variations. By incorporating key design and process features, a physical model of bitcell static noise margin (SNM) has been derived analytically. It captures intra-die SNM variations by the combination of a folded-normal distribution and a non-central chi-squared distribution. Validations with MC simulation show its accuracy of modeling SNM distributions down to 25mV beyond 6-sigma for typical bitcells in 28nm. Model-based tuning of subthreshold SRAM bitcells is investigated for design tradeoff between leakage, area and stability. When targeting a specific SNM constraint, we show that an optimal standby voltage exists which offers minimum bitcell leakage power – any deviation above or below increases the power consumption. When targeting a specific standby voltage, our design flow identifies bitcell instances of 12× less leakage power or 3× reductions in area as compared to the minimum-length design. Xin Fan 0002, Tobias Gemmeke |
ICCAD | 1 |