EDBT 2026 Demo / reviewers in the wild / expert
Mingfei Yu
dblp:280/0391
· DBLP profile ↗
13ranked-venue papers
7as first author
13since 2021 · last 2026
0009-0009-6816-8903ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 10 · 4 first-author · 10 since 2021Software engineering, systems software and programming languages · 4 · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 first-author · 2 since 2021Security and privacy · 1 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Faster Homomorphic Operations and Beyond: Expediting Homomorphic Computation via Boolean Circuit OptimizationabstractAbstract Fully homomorphic encryption (FHE) enables secure data processing without compromising data access. However, its computational cost and slower execution compared to plaintext operations present significant challenges. The increasing interest in FHE-based secure computation underscores the need to accelerate homomorphic computations. Existing research predominantly focuses on reducing the multiplicative depth (MD) of FHE circuits, as a lower MD enhances the execution efficiency of each homomorphic operation. However, this often comes at the expense of increased multiplicative complexity (MC), leading to more homomorphic multiplications — a computationally intensive task. Currently, there is a lack of approaches that effectively balance the trade-off between MD reduction and MC increase, potentially resulting in sub-optimal outcomes. This paper addresses this critical gap with three main contributions: (a) an exact synthesis paradigm for generating optimal FHE circuit implementations, (b) a heuristic circuit optimization algorithm, named MC-aware MD minimization, that leverages the exact synthesis paradigm to optimize FHE circuits efficiently, and (c) an FHE circuit optimization flow that integrates MC-aware MD minimization with existing MD reduction techniques. Experimental results demonstrate a 21.32% average reduction in homomorphic computation time and highlight significantly improved efficiency in circuit optimization. Mingfei Yu, Giovanni De Micheli |
J. Cryptol. | 1 |
| 2025 | Back-end-aware Fault-tolerant Quantum Oracle SynthesisabstractQuantum oracle synthesis involves compiling arbitrary Boolean functions into quantum circuits using specific quantum gates supported by the target quantum computer. The Clifford+T gate library is particularly common in fault-tolerant quantum computing systems. Utilizing XOR-AND-inverter graphs (XAGs) as the logic representation for the target Boolean functions has received extensive attention due to the observed direct correlation between the number of AND nodes in an XAG and the T count and the helper qubit count of the quantum oracle optimally compiled from it. However, to be deployed onto fault-tolerant quantum hardware, quantum gates must be further re-expressed by logical quantum error correction (QEC) code operations, a process known as back-end compilation. This paper enhances the current XAG-based oracle synthesis techniques by establishing a link between the properties of XAGs and quality measures of back-end-compiled quantum oracles. This link unlocks more optimization opportunities---experimental results demonstrate average reductions of 4.49% in T count, 7.00% in logical time steps, and 14.89% in helper qubit count, respectively, on benchmarks optimized by the proposed back-end-aware XAG optimization approaches. Mingfei Yu, Alessandro Tempia Calvino, Mathias Soeken, Giovanni De Micheli |
ASP-DAC | 1 |
| 2025 | Making the Best Switch: Encoding Strategy Management for Efficient TFHE Circuit EvaluationabstractThis work addresses the synthesis of efficient torus fully homomorphic encryption (TFHE) circuits for private Boolean function evaluation through encoding strategy management. Modern TFHE implementations support multiple plaintext encoding spaces, each offering distinct trade-offs between computational cost and the expressiveness enabled by larger plaintext domains. Smartly switching between encoding strategies to maximize evaluation efficiency remains challenging due to the lack of algorithmic support for determining when and where such transitions should occur. To address this, we propose a synthesis framework that enables encoding-switch-aware TFHE circuit generation. Our approach leverages the structural properties of the exclusive-or sum of products (ESOP) representation to partition Boolean functions into encoding-aligned regions, enabling cost-effective evaluation while minimizing switch overhead. Experimental results demonstrate that our encoding-aware synthesis technique significantly accelerates homomorphic Boolean function evaluation – achieving up to 53.46% and 23.34% average evaluation time reduction on general-purpose Boolean benchmarks – compared to advanced synthesis baselines lacking explicit encoding-switch management. This work lays the groundwork for systematic encoding strategy management in TFHE circuits and highlights the role of logic-level design automation in advancing efficient homomorphic evaluation. Mingfei Yu, Gabrielle De Micheli, Giovanni De Micheli |
ICCAD | 1 |
| 2024 | Unleashing the Power of T1-cells in SFQ Arithmetic CircuitsabstractRapid single-flux quantum (RSFQ) is one of the most advanced cryogenic superconductive electronics technologies. With orders of magnitude smaller power dissipation, RSFQ is an attractive technology for cloud computing, aerospace electronics, and high-speed interfacing with quantum computing systems. Technological challenges however greatly complicate the realization of VLSI-complexity RSFQ systems. For example, gate-level pipelining in SFQ systems incurs a significant area overhead due to the need for path balancing. This issue is particularly detrimental to SFQ systems due to the limited layout density of RSFQ systems. Rassul Bairamkulov, Mingfei Yu, Giovanni De Micheli |
DAC | 2 |
| 2024 | Technology-Aware Logic Synthesis for Superconducting ElectronicsabstractSuperconducting electronics provide us with cryogenic digital circuits that can rival established technologies in performance and energy consumption. Today, the lack of tools for the design of large-scale integrated superconducting circuits is a major obstacle to their deployment. Few research institutions and companies have contributed to making such tools available. This review focuses on methods, algorithms, and open-source design tools for logic synthesis of superconducting circuits in two major families: single-flux quantum (SFQ) circuits and adiabatic quantum flux parametron (AQFP). Rassul Bairamkulov, Siang-Yun Lee, Alessandro Tempia Calvino, Dewmini Sudara Marakkalage, Mingfei Yu, Giovanni De Micheli |
DATE | 5 |
| 2024 | Unleashing the Power of T1-Cells in SFQ Arithmetic CircuitsabstractRapid single-flux quantum (RSFQ), a leading cryogenic superconductive electronics (SCE) technology, offers extremely low power dissipation and high speed. However, implementing RSFQ systems at VLSI complexity faces challenges, such as substantial area overhead from gate-level pipelining and path balancing, exacerbated by RSFQ's limited layout density. T1 flip-flop (T1-FF) is an RSFQ logic cell operating as a pulse counter. Using T1-FF the full adder function can be realized with only 40% of the area required by the conventional realization. This cell however imposes complex constraints on input signal timing, complicating its use. Multiphase clocking has been recently proposed to alleviate gate-level pipelining overhead. The fanin signals can be efficiently controlled using multiphase clocking. We present the novel two-stage SFQ technology mapping methodology supporting the T1-FF. Compatible parts of the SFQ network are first replaced by the efficient T1-FFs. Multiphase retiming is next applied to assign clock phases to each logic gate and insert DFFs to satisfy the input timing. Using our flow, the area of the SFQ networks is reduced, on average, by 6% with up to 25% reduction in optimizing the 128-bit adder. Rassul Bairamkulov, Mingfei Yu, Giovanni De Micheli |
DATE | 2 |
| 2024 | RareLS: Rarity-Reducing Logic Synthesis for Mitigating Hardware Trojan ThreatsabstractHardware Trojan (HT) poses a critical security threat to integrated circuits, which can change circuit functionality or leak sensitive data. HTs are typically activated under low-probability conditions by exploiting "rare signals" in logic circuits. In this paper, we propose RareLS, rarity-reducing logic synthesis for mitigating HT threats. Specifically, RareLS reduces the number of rare signals through rarity-oriented technology-independent optimization and technology mapping. Experimental results show that RareLS reduces rare signals by 63.4% on average, with a small overhead of 4.0% in area, 1.9% in delay, and 6.2% in power. Moreover, RareLS complicates HT insertion for attackers by reducing HT trigger logic by 92.94%, and aids defenders in detecting HTs by shortening the test length by more than 80.83%. Chang Meng, Mingfei Yu, Wayne P. Burleson, Giovanni De Micheli |
ICCAD | 2 |
| 2023 | Striving for Both Quality and Speed: Logic Synthesis for Practical Garbled CircuitsabstractGarbled circuit (GC) is one of the few promising protocols to realize general-purpose secure computation. The target computation is represented by a Boolean circuit that is subsequently transformed into a network of encrypted tables for execution. The need of distributing GCs among parties however requires excessive data communication, called garbling cost, which bottlenecks system performance. Due to the zero garbling cost of XOR operations, existing works reduce garbling cost by representing the target computation as the XOR-AND graph (XAG) with minimum multiplicative complexity. Recently, an XOR-OneHot graph (X1G) has been proposed as an efficient GC representation. However, there is a lack of formal proof of X1G performance in the literature. In this paper, we prove that starting from any XAG, there exists an X1G implementation with equal or lower garbling cost. Based on our findings, we propose (a) an affine function classification-based database generation method, which decouples time-consuming on-the-fly exact syn-thesis from Boolean rewriting; (b) a novel optimal X1G synthesis approach to accelerate the database generation procedure. The proposals jointly facilitate a performant Boolean rewriting-based X1G optimization method. Experimental evaluations show significant improvement in both garbling cost and runtime: with a 2273.63× speed-up achieved on average, the proposed method realized an up-to 8.20% improvement in garbling cost reduction compared to the state-of-the-art. Mingfei Yu, Giovanni De Micheli |
ICCAD | 1 |
| 2023 | SCP-SLAM: Accelerating DynaSLAM With Static Confidence PropagationabstractDynaSLAM is the state-of-the-art visual simultaneous localization and mapping (SLAM) in dynamic environments. It adopts a convolutional neural network (CNN) for moving object detection, but usually incurs a very high computational cost because it performs semantic segmentation using the CNN model on every frame. This paper proposes SCP-SLAM, which accelerates DynaSLAM by running the CNN only on keyframes and propagating static confidence through other frames in parallel. The proposed static confidence characterizes the moving object features by the residual defined by inter-frame geometry transformation, which can be computed quickly. Our method combines the effectiveness of a CNN with the efficiency of static confidence in a tightly coupled manner. Extensive experiments on the publicly available TUM and Bonn RGB-D dynamic benchmark datasets demonstrate the efficacy of the method. Compared with DynaSLAM, it enables acceleration by a factor of ten on average, but retains comparable localization accuracy. Mingfei Yu, Lei Zhang 0021, Wu-Fan Wang |
VR | 1 |
| 2021 | A Decomposition-Based Synthesis Algorithm for Sparse Matrix-Vector Multiplication in Parallel Communication StructureabstractThere is an obvious trend that hardware including many-core CPU, GPU and FPGA are always made use of to conduct computationally intensive tasks of deep learning implementations, a large proportion of which can be formulated into the format of sparse matrix-vector multiplication(SpMV). In contrast with dense matrix-vector multi-plication(DMV), scheduling solutions for SpMV targeting parallel processing turn out to be irregular, leading to the dilemma that the optimum synthesis problems are time-consuming or even infeasible, when the size of the involved matrix increases. In this paper, the minimum scheduling problem of 4x4 SpMV on ring-connected architecture is first studied, with two concepts named Multi-Input Vector and Multi-Output Vector introduced. The classification of 4x4 sparse matrices has been conducted, on account of which a decomposition-based synthesis algorithm for larger matrices is put forward. As the proposed method is guided by known sub-scheduling solutions, search space of the synthesis problem is considerably reduced. Through comparison with an exhaustive search method and a brute force-based parallel scheduling method, the proposed method is proved to be able to offer scheduling solutions of high-equality: averagely utilize 65.27% of the sparseness of the involved matrices and achieve 91.39% of the performance of the solutions generated by exhaustive search, with a remarkable saving of compilation time and the best scalability among the above-mentioned approaches. Mingfei Yu, Ruitao Gao |
ASP-DAC | 1 |
| 2021 | Logic Synthesis for Generalization and Learning AdditionabstractLogic synthesis generates a logic circuit of a given Boolean function, where the size and depth of the circuit are optimized for small area and low delay. On the other hand, machine learning has been extensively studied and used for many applications these days. Its general approach of training a model from a set of input-output samples is similar to logic synthesis with external don't-cares, except that in the case of machine learning the goal is to come up with a general understanding from the given samples. Seeing this resemblance from another perspective, we can think of logic synthesis targeting a generalization of the care-set. In this paper, we try such logic synthesis that generates a logic circuit where the given incomplete relation between input and output is generalized. We compared popular logic synthesis methods and machine learning models and analyzed their characteristics. We found that there were some arithmetic functions that these conventional models cannot effectively learn. Out of them, we further experimented on addition operations using tree models and found a heuristic minimization method of BDD achieves the highest accuracy. Yukio Miyasaka, Xinpei Zhang, Mingfei Yu, Qingyang Yi |
DATE | 3 |
| 2021 | Logic Synthesis Meets Machine Learning: Trading Exactness for GeneralizationabstractLogic synthesis is a fundamental step in hardware design whose goal is to find structural representations of Boolean functions while minimizing delay and area. If the function is completely-specified, the implementation accurately represents the function. If the function is incompletely-specified, the implementation has to be true only on the care set. While most of the algorithms in logic synthesis rely on SAT and Boolean methods to exactly implement the care set, we investigate learning in logic synthesis, attempting to trade exactness for generalization. This work is directly related to machine learning where the care set is the training set and the implementation is expected to generalize on a validation set. We present learning incompletely-specified functions based on the results of a competition conducted at IWLS 2020. The goal of the competition was to implement 100 functions given by a set of care minterms for training, while testing the implementation using a set of validation minterms sampled from the same function. We make this benchmark suite available and offer a detailed comparative analysis of the different approaches to learning. Shubham Rai, Walter Lau Neto, Yukio Miyasaka, Xinpei Zhang, Mingfei Yu, Qingyang Yi, Masahiro Fujita 0004, Guilherme B. Manske, Matheus F. Pontes, Leomar S. da Rosa Jr., Marilton S. de Aguiar, Paulo F. Butzen, Po-Chun Chien, Yu-Shan Huang, Hoa-Ren Wang, Jie-Hong Roland Jiang, Jiaqi Gu 0002, Zheng Zhao 0003, Zixuan Jiang, David Z. Pan, Brunno Abreu, Isac de Souza Campos, Augusto Andre Souza Berndt, Cristina Meinhardt, Jônata Tyska Carvalho, Mateus Grellert, Sergio Bampi, Aditya Lohana, Akash Kumar 0001, Wei Zeng 0015, Azadeh Davoodi, Rasit Onur Topaloglu, Jordan Dotzel, Yichi Zhang 0006, Hanyu Wang 0005, Zhiru Zhang, Valerio Tenace, Pierre-Emmanuel Gaillardon, Alan Mishchenko, Satrajit Chatterjee |
DATE | 5 |
| 2021 | Loop Closure Detection by Using Global and Local Features With Photometric and Viewpoint InvarianceabstractLoop closure detection plays an important role in many Simultaneous Localization and Mapping (SLAM) systems, while the main challenge lies in the photometric and viewpoint variance. This paper presents a novel loop closure detection algorithm that is more robust to the variance by using both global and local features. Specifically, the global feature with the consolidation of photometric and viewpoint invariance is learned by a Siamese Network from the intensity, depth, gradient and normal vectors distribution. The local feature with rotation invariance is based on the histogram of relative pixel intensity and geometric information like curvature and coplanarity. Then, these two types of features are jointly leveraged for the robust detection of loop closures. The extensive experiments have been conducted on the publicly available RGB-D benchmark datasets like TUM and KITTI. The results demonstrate that our algorithm can effectively address challenging scenarios with large photometric and viewpoint variance, which outperforms other state-of-the-art methods. Mingfei Yu, Lei Zhang 0021, Wufan Wang, Hua Huang 0001 |
IEEE Trans. Image Process. | 1 |