VLDB 2026 Research / reviewers in the wild / expert
Jin Miao
dblp:123/7053
· DBLP profile ↗
14ranked-venue papers
5as first author
2since 2021 · last 2022
0000-0002-0150-4599ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 14 · 5 first-author · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
5 papers |
Electronic design automation · 71% Integrated circuit design · 25% Interconnection networks and networks-on-chip · 4% | |
| Network and information security
2 papers |
Hardware security and side channels · 100% |
Topics — the 14 heaviest of 15, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Electronic design automation
design space exploration |
1.0 | 2 | 2022 | High-Speed Adder Design Space Exploration via Graph Neural Processes · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2022 Cross-Layer Optimization for High Speed Adders: A Pareto Driven Machine Learning Approach · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2019 |
Electronic design automation
logic synthesis |
1.0 | 2 | 2022 | High-Speed Adder Design Space Exploration via Graph Neural Processes · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2022 Cross-Layer Optimization for High Speed Adders: A Pareto Driven Machine Learning Approach · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2019 |
Hardware security and side channels › hardware security primitives
physical unclonable function |
0.6 | 2 | 2018 | SD-PUF: Spliced Digital Physical Unclonable Function · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2018 Practical public PUF enabled by solving max-flow problem on chip · DAC 2016 |
Integrated circuit design › digital circuit design › arithmetic circuit design
adder design |
0.6 | 1 | 2022 | High-Speed Adder Design Space Exploration via Graph Neural Processes · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2022 |
Integrated circuit design › digital arithmetic circuits › parallel adder
parallel prefix adder |
0.6 | 1 | 2022 | High-Speed Adder Design Space Exploration via Graph Neural Processes · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2022 |
Electronic design automation › design space exploration
pareto optimization |
0.6 | 1 | 2022 | High-Speed Adder Design Space Exploration via Graph Neural Processes · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2022 |
Electronic design automation
physical design |
0.5 | 2 | 2020 | A Unified Framework for Simultaneous Layout Decomposition and Mask Optimization · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2020 Cross-Layer Optimization for High Speed Adders: A Pareto Driven Machine Learning Approach · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2019 |
Electronic design automation
design technology co-optimization |
0.4 | 1 | 2020 | A Unified Framework for Simultaneous Layout Decomposition and Mask Optimization · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2020 |
Electronic design automation › physical design › lithography
layout decomposition |
0.4 | 1 | 2020 | A Unified Framework for Simultaneous Layout Decomposition and Mask Optimization · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2020 |
Electronic design automation › physical design
mask optimization |
0.4 | 1 | 2020 | A Unified Framework for Simultaneous Layout Decomposition and Mask Optimization · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2020 |
Hardware security and side channels › hardware security primitives › physical unclonable function
machine-learning-resistant PUF |
0.3 | 1 | 2018 | SD-PUF: Spliced Digital Physical Unclonable Function · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2018 |
Hardware security and side channels › hardware security primitives › physical unclonable function
public PUF |
0.2 | 1 | 2016 | Practical public PUF enabled by solving max-flow problem on chip · DAC 2016 |
Integrated circuit design
analog and mixed-signal circuits |
0.2 | 1 | 2016 | Practical public PUF enabled by solving max-flow problem on chip · DAC 2016 |
Interconnection networks and networks-on-chip › switch architecture
crossbar array |
0.2 | 1 | 2016 | Practical public PUF enabled by solving max-flow problem on chip · DAC 2016 |
Methods — techniques the papers use, named apart from their topics
shuffle-splice mechanism · 0.7variational graph autoencoder · 0.6neural process · 0.6graph neural process · 0.6gaussian process surrogate · 0.6max-flow problem · 0.5gradient-based optimization · 0.4discrete optimization · 0.4machine learning · 0.4active learning · 0.4source degeneration · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2022 | High-Speed Adder Design Space Exploration via Graph Neural ProcessesabstractAdders are the primary components in the data-path logic of a microprocessor, and thus, adder design has been always a critical issue in the very large-scale integration (VLSI) industry. However, it is infeasible for designers to obtain optimal adder architecture by exhaustively running EDA flow due to the extremely large design space. Previous arts have proposed the machine learning-based framework to explore the design space. Nevertheless, they fall into suboptimality due to a two-stage flow of the learning process and less efficient nor effective feature representations of prefix adder structures. In this article, we first integrate a variational graph autoencoder and a neural process (NP) into an end-to-end, multibranch framework, which is termed thegraph neural process. The former performs automatic feature learning of prefix adder structures, whilst the latter one is designed as an alternative to the Gaussian process. Then, we propose a sequential optimization framework with the graph NP as the surrogate model to explore the Pareto-optimal prefix adder structures with tradeoff among Quality-of-Result (QoR) metrics, such as power, area, and delay. The experimental results show that compared with state-of-the-art methodologies, our framework can achieve a much better Pareto frontier in multiple QoR metric spaces with fewer design-flow evaluations. Hao Geng, Yuzhe Ma, Qi Xu 0004, Jin Miao, Subhendu Roy, Bei Yu 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2021 | Correlated Multi-objective Multi-fidelity Optimization for HLS Directives DesignabstractHigh-level synthesis (HLS) tools have gained great attention in recent years because it emancipates engineers from the complicated and heavy hardware description language writing, by using high-level languages and HLS directives. However, previous works seem powerless, due to the time-consuming design processes, the contradictions among design objectives, and the accuracy difference between the three stages (fidelities). To find good HLS directives, in this paper, a novel correlated multi-objective non-linear optimization algorithm is proposed to explore the Pareto solutions while making full use of data from different fidelities. A non-linear Gaussian process is proposed to model relationships among the analysis reports from different fidelities for the same objective. For the first time, correlated multivariate Gaussian process models are introduced into this domain to characterize the complex relationships of multiple objectives in each design fidelity. A tree-based method is proposed to erase invalid solutions and obviously non-optimal solutions. Experimental results show that our non-linear and pioneering correlated models can approximate the Pareto-frontier of the directive design space in a shorter time with much better performance and good stability, compared with the state-of-the-art. Qi Sun 0002, Tinghuan Chen, Siting Liu 0002, Jin Miao, Jianli Chen, Hao Yu 0001, Bei Yu 0001 |
DATE | 4 |
| 2020 | Hotspot Detection via Attention-based Deep Layout Metric LearningabstractWith the aggressive and amazing scaling of the feature size of semiconductors, hotspot detection has become a crucial and challenging problem in the generation of optimized mask design for better printability. Machine learning techniques, especially deep learning, have attained notable success on hotspot detection tasks. However, most existing hotspot detectors suffer from suboptimal performance due to two-stage flow and less efficient representations of layout features. What is more, most works can only solve simple benchmarks with apparent hotspot patterns like ICCAD 2012 Contest benchmarks. In this paper, we firstly develop a new end-to-end hotspot detection flow where layout feature embedding and hotspot detection are jointly performed. An attention mechanism-based deep convolutional neural network is exploited as the backbone to learn embeddings for layout features and classify the hotspots simultaneously. Experimental results demonstrate that our framework achieves accuracy improvement over prior arts with fewer false alarms and faster inference speed on much more challenging benchmarks. Hao Geng, Jin Miao, Fan Yang 0001, Xuan Zeng 0001, Bei Yu 0001 |
ICCAD | 4 |
| 2020 | A Unified Framework for Simultaneous Layout Decomposition and Mask OptimizationabstractIn advanced technology nodes, layout decomposition (LD) and mask optimization (MO) are two key stages in integrated circuit design. Due to the inconsistency of the objectives of these two stages, the performance of conventional layout and MO may be suboptimal. To tackle this problem, in this article, we propose a unified framework, which seamlessly integrates LD and MO. We propose a gradient-based approach to solve the unified mathematical formulation, as well as a set of discrete optimization techniques to avoid being stuck in local optimum. The conventional optimization process can be accelerated as some inferior decomposition results can be smartly pruned in early stages. The experimental results show that the proposed unified framework can achieve more than 34× speed-up compared with the conventional two-stage flow, meanwhile, it can dramatically reduce EPE violations by more than 8×, and thus maintain better design quality. Yuzhe Ma, Shuxiang Hu, Jhih-Rong Gao, Jian Kuang 0001, Jin Miao, Bei Yu 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 6 |
| 2019 | Power-Driven DNN Dataflow Optimization on FPGAabstractDeep neural networks (DNNs) have been proven to achieve unprecedented success on modern artificial intelligence (AI) tasks, which have also greatly motivated the rapid developments of novel DNN models and hardware accelerators. Many challenges still remain towards the design of power efficient DNN accelerator due to the intrinsically intensive data computation and transmission in DNN algorithms. However, most existing efforts in the domain have taken latency as the sole optimization objective, which may often result in sub-optimality in power consumption. In this paper, we propose a framework to optimize the power efficiency of DNN dataflow on FPGA while maximally minimizing the impact on latency. We first propose power and latency models that are built upon different dataflow configurations. Then a power-driven dataflow formulation is proposed, which enables a hierarchical exploration strategy on the dataflow configurations, leading to efficient power consumption at limited latency loss. Experimental results have demonstrated the effectiveness of our proposed models and exploration strategies, where power improvement has shown up to 31% with latency degradation of no worse than 6.5%. Qi Sun 0002, Tinghuan Chen, Jin Miao, Bei Yu 0001 |
ICCAD | 3 |
| 2019 | Cross-Layer Optimization for High Speed Adders: A Pareto Driven Machine Learning ApproachabstractIn spite of maturity to the modern electronic design automation (EDA) tools, optimized designs at architectural stage may become suboptimal after going through physical design flow. Adder design has been such a long studied fundamental problem in very large-scale integration industry yet designers cannot achieve optimal solutions by running EDA tools on the set of available prefix adder architectures. In this paper, we enhance a state-of-the-art prefix adder synthesis algorithm to obtain a much wider solution space in architectural domain. On top of that, a machine learning-based design space exploration methodology is applied to predict the Pareto frontier of the adders in physical domain, which is infeasible by exhaustively running EDA tools for innumerable architectural solutions. Considering the high cost of obtaining the true values for learning, an active learning algorithm is proposed to select the representative data during learning process, which uses less labeled data while achieving better quality of Pareto frontier. Experimental results demonstrate that our framework can achieve Pareto frontier of high quality over a wide design space, bridging the gap between architectural and physical designs. Source code and data are available athttps://github.com/yuzhe630/adder-DSE. Yuzhe Ma, Subhendu Roy, Jin Miao, Bei Yu 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2018 | SD-PUF: Spliced Digital Physical Unclonable FunctionabstractDigital circuit physical unclonable function (PUF) has been attracting attentions for the merits of resilience to the environmental and operational variations that analog PUFs suffer from. Existing state-of-the-art digital circuit PUFs, however, are either hybrid of analog-digital circuits which are still under the shadow of vulnerability, or impractical for real-world applications. In this paper, we propose a novel highly nonlinear and secure digital PUF (D-PUF) and the spliced version SD-PUF. The fingerprints are extracted from intentionally induced very large-scale integration interconnect randomness during lithography process, as well as a post-silicon shuffling process. Strongly skewed CMOS latches are used to ensure the immunity against environmental and operational variations. Crucially, a highly nonlinear logic network is proposed to effectively spread and augment any subtle interconnect randomness, which also enables strong resilience against machine learning attacks. On top of it, the expandable architecture of the proposed logic network empowers a novel post-silicon shuffle-splice mechanism, where multiple randomly selected D-PUFs are spliced to be one SD-PUF, pushing the statistical security to a much higher level, while significantly reducing the mask cost per PUF device. It also decouples the trustworthy demands enforced to the foundries or other third party manufacturers. Our proposed PUFs demonstrate close to ideal performance in terms of statistical metrics, including 0 intra-Hamming distance. Various state-of-the-art machine learning models show prediction accuracies almost no better than random guesses when attacking to the proposed PUFs. We also mathematically prove the probability of existence of identical SD-PUF pair is significantly lower than that of D-PUF pair, e.g., such probability of an SD-PUF spliced by 30 D-PUFs is 2.3 × 10-22, which is 19 order magnitude lower than that of D-PUF. Benefited from the proposed shuffle-splice mechanism, the mask cost per SD-PUF is also reduced by 300× than that of D-PUF. Jin Miao, Meng Li 0004, Subhendu Roy, Yuzhe Ma, Bei Yu 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2017 | A unified framework for simultaneous layout decomposition and mask optimizationabstractIn advanced technology nodes, layout decomposition and mask optimization are two key stages in integrated circuit design. Due to the inconsistency of the objectives of these two stages, the performance of conventional layout and mask optimization may be suboptimal. To tackle this problem, in this paper we propose a unified framework, which seamlessly integrates layout decomposition and mask optimization. We propose a gradient based approach to solve the unified mathematical formulation, as well as a set of discrete optimization techniques to avoid being stuck in local optimum. The conventional optimization process can be accelerated as some inferior decompositions can be smartly pruned in early stages. The experimental results show that the proposed unified framework can achieve more than 17 x speed-up compared with the conventional two-stage flow, meanwhile it can reduce EPE violations by 18%, and thus maintain better design quality. Yuzhe Ma, Jhih-Rong Gao, Jian Kuang 0001, Jin Miao, Bei Yu 0001 |
ICCAD | 4 |
| 2017 | A learning bridge from architectural synthesis to physical design for exploring power efficient high-performance addersabstractIn spite of maturity to the modern electronic design automation (EDA) tools, optimized designs at architectural stage may become sub-optimal after going through physical design flow. Adder design has been such a long studied fundamental problem in VLSI industry yet designers cannot achieve optimal solutions by running EDA tools on the set of available prefix adder architectures. In this paper, we enhance a state-of-the-art prefix adder synthesis algorithm to obtain a much wider solution space in architectural domain. On top of that, a machine learning based design space exploration methodology is applied to predict the Pareto frontier of the adders in physical domain, which is infeasible by exhaustively running EDA tools for innumerable architectural solutions. Experimental results demonstrate that our framework can achieve near-optimal delay vs. power/area Pareto frontier over a wide design space, bridging the gap between architeon the set of available prefix adder architectures. In this paper, we enhance a state-of-the-art prefix adder synthesis algorithm to obtain a much wider solution space in architectural domain. On top of that, a machine learning based design space exploration methodology is applied to predict the Pareto frontier of the adders in physical domain, which is infeasible by exhaustively running EDA tools for innumerable architectural solutions. Experimental results demonstrate that our framework can achieve near-optimal delay vs. power/area Pareto frontier over a wide design space, bridging the gap between architectural andctural and physical designs. Subhendu Roy, Yuzhe Ma, Jin Miao, Bei Yu 0001 |
ISLPED | 3 |
| 2016 | Practical public PUF enabled by solving max-flow problem on chipabstractThe execution-simulation gap (ESG) is a fundamental property of public physical unclonable function (PPUF), which exploits the time gap between direct IC execution and computer simulation. ESG needs to consider both advanced computing scheme, including parallel and approximate computing scheme, and IC physical realization. In this paper, we propose a novel PPUF design, whose execution is equivalent to solving the hard-to-parallel and hard-to-approximate max-flow problem in a complete graph on chip. Thus, max-flow problem can be used as the simulation model to bound the ESG rigorously. To enable an efficient physical realization, we propose a crossbar structure and adopt source degeneration technique to map the graph topology on chip. The difference on asymptotic scaling between execution delay and simulation time is examined in the experimental results. The measurability of output difference is also verified to prove the physical practicality. Meng Li 0004, Jin Miao, Kai Zhong 0007, David Z. Pan |
DAC | 2 |
| 2016 | LRR-DPUF: learning resilient and reliable digital physical unclonable functionabstractConventional silicon physical unclonable function (PUF) extracts fingerprints from transistor's analog attributes, which are vulnerable to environmental and operational variations. Recently, digitalized PUF prototypes have emerged to overcome the vulnerability issues, however, the existing prototypes are either hybrid of analog-digital PUFs which are still under the shadow of vulnerability, or impractical for real-world implementation. To address the above limitations, we propose a learning resilient and reliable digital PUF (LRR-DPUF). The fingerprints are extracted from VLSI interconnect geometrical randomness induced by lithography variations. Crucially, we use strongly skewed latches to ensure the immunity against environmental and operational variations. Further, a cross-coupled, highly non-linear logic network is proposed to effectively spread and augment even subtle interconnect randomness, as well as to achieve strong resilience to machine learning attacks. We demonstrate that a 64-bit LRR-DPUF exhibits close to ideal statistical performances, including 0 intra Hamming Distance. We also mathematically prove that each output of the LRR-DPUF follows uniform distribution. Various state-of-the-art machine learning models show almost no better than random prediction accuracies when applied to LRR-DPUF. Jin Miao, Meng Li 0004, Subhendu Roy, Bei Yu 0001 |
ICCAD | 1 |
| 2014 | Multi-level approximate logic synthesis under general error constraintsabstractWe address the problem of multi-level approximate logic synthesis. Our strategy assumes existence of an optimized exact Boolean network, which is critical in practice since arithmetic blocks are rarely synthesized from 2-level representation automatically. The goal is to produce minimum cost circuits whose logic function deviates in a controlled manner from the exact function with deviations quantified by the magnitude and frequency of errors. We rely on network simplifications allowed by external don't cares (EXDCs). We formulate the error-magnitude constrained problem by using Boolean relations to capture the allowed error behavior in a more general manner compared to incompletely specified functions. Our key contribution is in finding sets of external don't cares that maximally approach the Boolean relation. The algorithm starts with an EXDC set that is overly relaxed and iteratively, and in a greedy fashion, identifies a feasible EXDC set by solving a series of conventional EXDC-based network optimizations. The algorithm then ensures compliance to error frequency constraints by recovering the correct outputs on the sought number of error-producing inputs while aiming to minimize the network cost increase. We applied the algorithm to several well-known adder and multiplier designs of varying bit-width. Even for small error magnitudes, the algorithm produces networks with gate count reduced by 30-50%, when the error frequency constraint is loose. This is up to 20% fewer gates than a naive EXDC-based approach. Jin Miao, Andreas Gerstlauer, Michael Orshansky |
ICCAD | 1 |
| 2013 | Approximate logic synthesis under general error magnitude and frequency constraintsabstractRecent interest in approximate circuit design is driven by its potential for large energy savings. In this paper, we address the problem of approximate logic synthesis (ALS). ALS is concerned with formally synthesizing a minimum-cost approximate Boolean network whose behavior deviates in a well-defined manner from a specified exact Boolean function, where in this work, we allow the deviation to be constrained by both the magnitude and frequency of the error. We make two contributions in solving this general ALS problem: The first contribution is to establish that the approximate synthesis problem un-constrained by the frequency of errors is isomorphic with the Boolean relations (BR) minimization problem. That equivalence allows us to exploit recently developed fast algorithms for BR problems to solve the error magnitude-only constrained ALS problem. The second contribution is an efficient heuristic algorithm for iteratively refining the magnitude-constrained solution to arrive at a solution also satisfying the error frequency constraint. Our combined greedy approximate logic synthesis (GALS) approach is able to operate on any Boolean network for which the deviation measures can be specified and is most immediately applicable to arithmetic blocks. Experiments on adder and multiplier blocks demonstrate literal count reductions of up to 60% under tight error frequency and magnitude constraints. Jin Miao, Andreas Gerstlauer, Michael Orshansky |
ICCAD | 1 |
| 2012 | Modeling and synthesis of quality-energy optimal approximate addersabstractRecent interest in approximate computation is driven by its potential to achieve large energy savings. This paper formally demonstrates an optimal way to reduce energy via voltage over-scaling at the cost of errors due to timing starvation in addition. We identify a fundamental trade-off between error frequency and error magnitude in a timing-starved adder. We introduce a formal model to prove that for signal processing applications using a quadratic signal-to-noise ratio error measure, reducing bit-wise error frequency is sub-optimal. Instead, energy-optimal approximate addition requires limiting maximum error magnitude. Intriguingly, due to possible error patterns, this is achieved by reducing carry chains significantly below what is allowed by the timing budget for a large fraction of sum bits, using an aligned, fixed internal-carry structure for higher significance bits. Jin Miao, Ku He, Andreas Gerstlauer, Michael Orshansky |
ICCAD | 1 |