EDBT 2026 Demo / reviewers in the wild / expert
Patricia Gonzalez-Guerrero
dblp:204/6026 · also Luisa Patricia Gonzalez-Guerrero
· DBLP profile ↗
9ranked-venue papers
3as first author
7since 2021 · last 2025
0000-0003-4377-7496ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 9 · 3 first-author · 7 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Enabling Classical-Quantum Interface Using Digital SFQ for Pulse-Phase Driven Control for Superconducting QubitsabstractIn the interest of alleviating qubit coherence constraints and making quantum control circuits more efficient, this paper explores qubit control using superconducting Single-Flux Quantum (SFQ) digital circuits. By using SFQ pulse trains and SFQ switches, we represent qubit state transformations through their equivalent phase changes, where each SFQ pulse represents a phase change of certain degrees. We propose classical-based unitary quantum gates represented by SFQ pulses. We create a classical equivalent model for quantum gates based on an SFQ-based parametric pulse sequence equivalent to quantum Pauli gates. To generate pulse sequences of configurable parameters, we implement a versatile multi-frequency pulse generator that seamlessly integrates with the qubit-resonator cavity. Meriam Gay Bautista, Patricia Gonzalez-Guerrero, George Michelogiannakis, Anastasiia Butko |
ISCAS | 2 |
| 2024 | Triangle Counting in the Temporal DomainabstractTriangle counting is a graph kernel that reveals information about the communities in real-world networks. Recently, circuits that employ time-domain as opposed to traditional binary computing have been developed to try and improve the energy efficiency and throughput of multiple graph processing problems. Here, we explore the tradeoffs of using race logic (RL), where inputs are encoded as timing delays, to count triangles in a graph through neighbor-set intersection. Using theoretical analysis, we investigate the scaling efficiency when using a circuit of fixed dimension to find the intersection between sets of different sizes. We investigate three different circuit array sizes: the dimension that minimizes latency, the dimension that is ideal for processing sets of the mean length, and the dimension 3 X 3. We evaluate the energy efficiency and throughput of the proposed circuit and find that our approach demonstrates the most improvement over current state-of-the-art (SOTA) with unimodal, right-skewed, and low range set-length distributions such as road networks. Caroline Hammond, Patricia Gonzalez-Guerrero, Meriam Gay Bautista, Nirmalendu Bikash Patra |
ISLPED | 2 |
| 2024 | Toward Practical Superconducting Accelerators for Machine Learning Using U-SFQabstractMost popular superconducting circuits operate on information carried by ps-wide, μV-tall, single flux quantum (SFQ) pulses. These circuits can operate at frequencies of hundreds of GHz with orders of magnitude lower switching energy than complementary-metal-oxide-semiconductors (CMOS). However, under the stringent area constraints of modern superconductor technologies, fully-fledged, CMOS-inspired superconducting architectures cannot be fabricated at large scales. Unary SFQ (U-SFQ) is an alternative computing paradigm that can address these area constraints. In U-SFQ, information is mapped to a combination of streams of SFQ pulses and in the temporal domain. In this work, we extend U-SFQ to introduce novel building blocks such as a multiplier and an accumulator. These blocks reduce area and power consumption by 2 \(\times\) and 4 \(\times\) compared with previously proposed U-SFQ building blocks and yield at least 97% area savings compared with binary approaches. Using these multiplier and adder, we propose a U-SFQ Convolutional Neural Network (CNN) hardware accelerator capable of comparable peak performance with state-of-the-art superconducting binary approach (B-SFQ) in 32 \(\times\) less area. CNNs can operate with 5–8 bits of resolution with no significant degradation in classification accuracy. For 5 bits of resolution, our proposed accelerator yields 5 \(\times\) to 63 \(\times\) better performance than CMOS and 15 \(\times\) to 173 \(\times\) better area efficiency than B-SFQ. Patricia Gonzalez-Guerrero, Kylie Huch, Nirmalendu Bikash Patra, Doru-Thom Popovici, George Michelogiannakis |
ACM J. Emerg. Technol. Comput. Syst. | 1 |
| 2024 | Area Efficient Asynchronous SFQ Pulse Round-Robin Distribution NetworkabstractWe present an area-efficient, asynchronous, single-input, and multiple-output Rapid Single Flux Quantum (RSFQ) pulse round-robin distribution network. We adopt the structure of a two-output toggle flip flop (TFF) where incoming pulses are temporarily stored in a SQUID loop in the form of magnetic flux quanta and then directed to outputs in a round-robin fashion. To support additional toggle outputs, we design a new circuit based on TFFs to distribute pulses in a round-robin mechanism to more than two outputs. We also elaborate on our design methodology that can support a different number of outputs while minimizing the number of JJs and power consumption. We then demonstrate a three- and four-output round-robin distribution network constructed with only 14-JJs and 18-JJs with$14.85 ~\mu W$and 18.86-$\mu W$power dissipation, respectively. Our four-output design has 40%–70% fewer JJ compared to a similar-functioning four-output network composed of TFFs, DFFs, splitters, and mergers. Our design also consumes 38%–70% less power and has a 16%–64% reduced delay compared to the same four-output TFF network. Finally, we demonstrate the usability of our four-output design in the context of a periodic counting network and pulse generator. Meriam Gay Bautista, Darren Lyles, Kylie Huch, Patricia Gonzalez-Guerrero, George Michelogiannakis |
IEEE Trans. Circuits Syst. I Regul. Pap. | 4 |
| 2023 | Superconducting Shuttle-Flux Shift Register for Race Logic and Its ApplicationsabstractThis paper presents a superconducting, magnetically-coupled, shuttle-flux shift register (SF-SR) that stores single flux quantum (SFQ) pulses. This shift register has a DC bias operating margin of ±34% at 10 GHz, with a power dissipation of$3.6~\mu W$and 38% fewer Josephson junctions (JJs) when scaled up to multiple stages compared to a data flip-flop (DFF) based shift register. The clock input is inductively coupled and is independent from the data input. We then present three applications for our SF-SR. In the first application, we add two non-destructive readout (NDRO) cells to construct a buffer that temporarily stores the temporal information of a series of race logic (RL) pulses. The second application is a pseudo-random number generator based on a linear function shift register (LFSR). The third application is N parallel SF-SRs that can act similar to a deserializer or instead can emulate a single SF-SR of N times higher clock frequency. These three applications motivate deep shift registers with many shifting intervals, which our SF-SR can implement with fewer JJs and lower power consumption compared to DFF-based shift registers. Meriam Gay Bautista, Patricia Gonzalez-Guerrero, Darren Lyles, George Michelogiannakis |
IEEE Trans. Circuits Syst. I Regul. Pap. | 2 |
| 2022 | Temporal and SFQ pulse-streams encoding for area-efficient superconducting acceleratorsabstractSuperconducting technology is a prime candidate for the future of computing. However, current superconducting prototypes are limited to small-scale examples due to stringent area constraints and complex architectures inspired from voltage-level encoding in CMOS; this is at odds with the ps-wide Single Quantum Flux (SFQ) pulses used in superconductors to carry information. In this work, we propose a wave-pipelined Unary SFQ (U-SFQ) architecture that leverages the advantages of two data representations: pulse-streams and Race Logic (RL). We introduce novel building blocks such as multipliers, adders, and memory cells, which leverage the natural properties of SFQ pulses to mitigate area constraints. We then design and simulate three popular hardware accelerators: i) a Processing Element (PE), typically used in spatial architectures; ii) A dot-product-unit (DPU), one of the most popular accelerators in artificial neural networks and digital signal processing (DSP); and iii) A Finite Impulse Response (FIR) filter, a popular and computationally demanding DSP accelerator. The proposed U-SFQ building blocks require up to 200× fewer JJs compared to their SFQ binary counterparts, exposing an area-delay trade-off. This work mitigates the stringent area constraints of superconducting technology. Patricia Gonzalez-Guerrero, Meriam Gay Bautista, Darren Lyles, George Michelogiannakis |
ASPLOS | 1 |
| 2021 | SRNoC: A Statically-Scheduled Circuit-Switched Superconducting Race Logic NoCabstractTemporal encoding has been shown to be a natural fit for single flux quantum (SFQ) superconducting computing since SFQ already encodes information with the presence or absence of voltage pulses. However, past work in SFQ has focused on binary-encoded networks on chip (NoCs). In this paper, we propose superconducting rotary NoC (SRNoC), a NoC where both data and control paths operate in the temporal domain following the race logic (RL) convention. Therefore, SFQ chips with temporal compute or memory can use SRNoC to avoid converting between the temporal and binary domains that would result from using a binary-encoded NoC. Using RL also enables SRNoC to be area-efficient, mitigating SFQ technology's low device density. SRNoC treats pulses as independent packets and delivers them to outputs without changing their value, i.e. preserving the RL convention. SRNoC operates on a fixed, rotating connection schedule between inputs and outputs. In each connection window, multiple pulses (packets) can be transmitted sequentially. SRNoC provides 13.1x higher throughput per port per Josephson junction (JJ) compared to the best-performing of three demonstrated NoCs. George Michelogiannakis, Darren Lyles, Patricia Gonzalez-Guerrero, Meriam Gay Bautista, Dilip P. Vasudevan, Anastasiia Butko |
IPDPS | 3 |
| 2020 | Fulcrum: A Simplified Control and Access Mechanism Toward Flexible and Practical In-Situ AcceleratorsabstractIn-situ approaches process data very close to the memory cells, in the row buffer of each subarray. This minimizes data movement costs and affords parallelism across subarrays. However, current in-situ approaches are limited to only row-wide bitwise (or few-bit) operations applied uniformly across the row buffer. They impose a significant overhead of multiple row activations for emulating 32-bit addition and multiplications using bitwise operations and cannot support operations with data dependencies or based on predicates. Moreover, with current peripheral logic, communication among subarrays is inefficient, and with typical data layouts, bits in a word are not physically adjacent. The key insight of this work is that in-situ, single-word ALUs outperform in-situ, parallel, row-wide, bitwise ALUs by reducing the number of row activations and enabling new operations and optimizations. Our proposed lightweight access and control mechanism, Fulcrum, sequentially feeds data into the single-word ALU and enables operations with data dependencies and operations based on a predicate. For algorithms that require communication among subarrays, we augment the peripheral logic with broadcasting capabilities and a previously-proposed method for low-cost inter-subarray data movement. The sequential processor also enables overlapping of broadcasting and computation, and reuniting bits that are physically adjacent. In order to realize true subarray-level parallelism, we introduce a lightweight column-selection mechanism through shifting one-hot encoded values. This technique enables independent column selection in each subarray. We integrate Fulcrum with Compress Express Link (CXL), a new interconnect standard. Fulcrum with one memory stack delivers on average (up to) 23.4 (76) speedup over a server-class GPU, NVIDIA P100, with three stacks of HBM2 memory, (ii) 70 (228) times speedup per memory stack over the GPU, and (iii) 19 (178.9) times speedup per memory stack over an ideal model of the GPU, which only accounts for the overhead of data movement. Marzieh Lenjani, Patricia Gonzalez-Guerrero, Elaheh Sadredini, Shuangchen Li, Yuan Xie 0001, Ameen Akel, Sean Eilert, Mircea R. Stan, Kevin Skadron |
HPCA | 2 |
| 2020 | Towards on-node Machine Learning for Ultra-low-power Sensors Using Asynchronous Σ Δ StreamsabstractWe propose a novel architecture to enable low-power, complex on-node data processing, for the next generation of sensors for the internet of things (IoT), smartdust, or edge intelligence. Our architecture combines near-analog-memory-computing (NAM) and asynchronous-computing-with-streams (ACS), eliminating the need for ADCs. ACS enables ultra-low power, massive computational resources required to execute on-node complex Machine Learning (ML) algorithms; while NAM addresses the memory-wall that represents a common bottleneck for ML and other complex functions. In ACS an analog value is mapped to an asynchronous stream that can take one of two logic levels ( v h , v l ). This stream-based data representation enables area/power-efficient computing units such as a multiplier implemented as an AND gate yielding savings in power of ∼90% compared to digital approaches. The generation of streams for NAM and ACS in a brute force manner, using analog-to-digital-converters (ADCs) and digital-to-streams-converters, would sky-rocket the power-latency-energy cost making the approach impractical. Our NAM-ACS architecture eliminates expensive conversions, enabling an end-to-end processing on asynchronous streams data-path. We tailor the NAM-ACS architecture for random forest (RaF), an ML algorithm, chosen for its ability to classify using a reduced number of features. Simulations show that our NAM-ACS architecture enables 75% of savings in power compared with a single ADC, obtaining a classification accuracy of 85% using an RaF-inspired algorithm. Patricia Gonzalez-Guerrero, Tommy Tracy II, Xinfei Guo, Rahul Sreekumar, Marzieh Lenjani, Kevin Skadron, Mircea R. Stan |
ACM J. Emerg. Technol. Comput. Syst. | 1 |