Anastasiia Butko

dblp:120/2710 · DBLP profile ↗
← Back
7ranked-venue papers
3as first author
3since 2021 · last 2026
0000-0002-3265-2885ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 7 · 3 first-author · 3 since 2021
YearPublicationVenuePosition
2026 C-3PQ: A Closeness Centrality-based Circuit Partitioner for Quantum Simulations
abstract
Simulating quantum circuits (QC) on high-performance computing (HPC) systems is essential for validating quantum algorithms and probing the potential of large-scale quantum computation before the advent of scalable quantum hardware. However, state vector simulations are resource-intensive, often requiring large clusters with thousands of compute nodes and large amounts of memory. To address this, we introduce C-3PQ, an end-to-end framework that minimizes inter-node data movement through an efficient circuit partitioning scheme alongside a flexible code generator to offer a portable simulation solution. By formulating the distribution of the state vector and circuit over multiple nodes as a graph problem, we utilize closeness centrality to assess gate importance and design a fast, scalable partitioning method. C-3PQ compiles the resulting partitions into highly optimized kernels targeting a wide range of hardware platforms from Intel, AMD, NVIDIA and ARM. Moreover, compared to a state-of-the-art implementation specific for NVIDIA platforms, C-3PQ achieves a speedup of up to \(40\%\) at a fraction of the cost when partitioning the state vector.
Doru-Thom Popovici, Harlin Lee, Naoki Yoshioka, Mauro Del Ben, Nobuyasu Ito, Katherine Klymko, Daan Camps, Anastasiia Butko
ICS8
2025 Enabling Classical-Quantum Interface Using Digital SFQ for Pulse-Phase Driven Control for Superconducting Qubits
abstract
In the interest of alleviating qubit coherence constraints and making quantum control circuits more efficient, this paper explores qubit control using superconducting Single-Flux Quantum (SFQ) digital circuits. By using SFQ pulse trains and SFQ switches, we represent qubit state transformations through their equivalent phase changes, where each SFQ pulse represents a phase change of certain degrees. We propose classical-based unitary quantum gates represented by SFQ pulses. We create a classical equivalent model for quantum gates based on an SFQ-based parametric pulse sequence equivalent to quantum Pauli gates. To generate pulse sequences of configurable parameters, we implement a versatile multi-frequency pulse generator that seamlessly integrates with the qubit-resonator cavity.
Meriam Gay Bautista, Patricia Gonzalez-Guerrero, George Michelogiannakis, Anastasiia Butko
ISCAS4
2021 SRNoC: A Statically-Scheduled Circuit-Switched Superconducting Race Logic NoC
abstract
Temporal encoding has been shown to be a natural fit for single flux quantum (SFQ) superconducting computing since SFQ already encodes information with the presence or absence of voltage pulses. However, past work in SFQ has focused on binary-encoded networks on chip (NoCs). In this paper, we propose superconducting rotary NoC (SRNoC), a NoC where both data and control paths operate in the temporal domain following the race logic (RL) convention. Therefore, SFQ chips with temporal compute or memory can use SRNoC to avoid converting between the temporal and binary domains that would result from using a binary-encoded NoC. Using RL also enables SRNoC to be area-efficient, mitigating SFQ technology's low device density. SRNoC treats pulses as independent packets and delivers them to outputs without changing their value, i.e. preserving the RL convention. SRNoC operates on a fixed, rotating connection schedule between inputs and outputs. In each connection window, multiple pulses (packets) can be transmitted sequentially. SRNoC provides 13.1x higher throughput per port per Josephson junction (JJ) compared to the best-performing of three demonstrated NoCs.
George Michelogiannakis, Darren Lyles, Patricia Gonzalez-Guerrero, Meriam Gay Bautista, Dilip P. Vasudevan, Anastasiia Butko
IPDPS6
2019 TIGER: topology-aware task assignment approach using ising machines
abstract
Optimal mapping of a parallel code's communication graph is increasingly important as both system size and heterogeneity increase. However, the topology-aware task assignment problem is an NP-complete graph isomorphism problem. Existing task scheduling approaches are either heuristic or based on physical optimization algorithms, providing different speed and solution quality tradeoffs. Ising machines such as quantum and digital annealers have recently become available offering an alternative hardware solution to solve certain types of optimization problems. We propose an algorithm that allows expressing the problem for such machines and a domain specific partition strategy that enables to solve larger scale problems. TIGER - topology-aware task assignment mapper tool - implements the proposed algorithm and automatically integrates task - communication graph and an architecture graph into the quantum software environment. We use D-Wave's quantum annealer to demonstrate the solving algorithm and evaluate the proposed tool flow in terms of performance, partition efficiency and solution quality. Results show significant speed-up of the tool flow and reliable solution quality while using TIGER together with the proposed partition.
Anastasiia Butko, George Michelogiannakis, David Donofrio, John Shalf
CF1
2019 Extending classical processors to support future large scale quantum accelerators
abstract
Extensive research in material science together with outstanding engineering efforts allowed quantum technology to be significantly improved hence enabling continuing scaling of quantum circuit size. In around 10 years, quantum annealing circuits have reached 103 qubits and trailing by several years, universal quantum circuits now demonstrate similar trends. From the current trends we can expect that quantum computers will reach thousands of qubits in the next 5--10 years.
Anastasiia Butko, George Michelogiannakis, David Donofrio, John Shalf
CF1
2016 Efficient Embedded Software Migration towards Clusterized Distributed-Memory Architectures
abstract
A large portion of existing multithreaded embedded software has been programmed according to symmetric shared memory platforms where a monolithic memory block is shared by all cores. Such platforms accommodate popular parallel programming models such as POSIX threads and OpenMP. However with the growing number of cores in modern manycore embedded architectures, they present a bottleneck related to their centralized memory accesses. This paper proposes a solution tailored for an efficient execution of applications defined with shared-memory programming models onto on-chip distributed-memory multicore architectures. It shows how performance, area and energy consumption are significantly improved thanks to the scalability of these architectures. This is illustrated in an open-source realistic design framework, including tools from ASIC to microkernel.
Rafael Garibotti, Anastasiia Butko, Luciano Ost, Abdoulaye Gamatié, Gilles Sassatelli, Chris Adeniyi-Jones
IEEE Trans. Computers2
2015 A trace-driven approach for fast and accurate simulation of manycore architectures
abstract
International audience
Anastasiia Butko, Rafael Garibotti, Luciano Ost, Vianney Lapotre, Abdoulaye Gamatié, Gilles Sassatelli, Chris Adeniyi-Jones
ASP-DAC1