Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Jin Miao

dblp:123/7053 · DBLP profile ↗
← Back
14ranked-venue papers
5as first author
2since 2021 · last 2022
0000-0002-0150-4599ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 14 · 5 first-author · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
5 papers
Electronic design automation · 71% Integrated circuit design · 25% Interconnection networks and networks-on-chip · 4%
Network and information security
2 papers
Hardware security and side channels · 100%

Topics — the 14 heaviest of 15, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Electronic design automation
design space exploration
1.022022
High-Speed Adder Design Space Exploration via Graph Neural Processes · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2022
Cross-Layer Optimization for High Speed Adders: A Pareto Driven Machine Learning Approach · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2019
Electronic design automation
logic synthesis
1.022022
High-Speed Adder Design Space Exploration via Graph Neural Processes · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2022
Cross-Layer Optimization for High Speed Adders: A Pareto Driven Machine Learning Approach · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2019
Hardware security and side channels › hardware security primitives
physical unclonable function
0.622018
SD-PUF: Spliced Digital Physical Unclonable Function · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2018
Practical public PUF enabled by solving max-flow problem on chip · DAC 2016
Integrated circuit design › digital circuit design › arithmetic circuit design
adder design
0.612022
High-Speed Adder Design Space Exploration via Graph Neural Processes · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2022
Integrated circuit design › digital arithmetic circuits › parallel adder
parallel prefix adder
0.612022
High-Speed Adder Design Space Exploration via Graph Neural Processes · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2022
Electronic design automation › design space exploration
pareto optimization
0.612022
High-Speed Adder Design Space Exploration via Graph Neural Processes · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2022
Electronic design automation
physical design
0.522020
A Unified Framework for Simultaneous Layout Decomposition and Mask Optimization · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2020
Cross-Layer Optimization for High Speed Adders: A Pareto Driven Machine Learning Approach · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2019
Electronic design automation
design technology co-optimization
0.412020
A Unified Framework for Simultaneous Layout Decomposition and Mask Optimization · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2020
Electronic design automation › physical design › lithography
layout decomposition
0.412020
A Unified Framework for Simultaneous Layout Decomposition and Mask Optimization · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2020
Electronic design automation › physical design
mask optimization
0.412020
A Unified Framework for Simultaneous Layout Decomposition and Mask Optimization · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2020
Hardware security and side channels › hardware security primitives › physical unclonable function
machine-learning-resistant PUF
0.312018
SD-PUF: Spliced Digital Physical Unclonable Function · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2018
Hardware security and side channels › hardware security primitives › physical unclonable function
public PUF
0.212016
Practical public PUF enabled by solving max-flow problem on chip · DAC 2016
Integrated circuit design
analog and mixed-signal circuits
0.212016
Practical public PUF enabled by solving max-flow problem on chip · DAC 2016
Interconnection networks and networks-on-chip › switch architecture
crossbar array
0.212016
Practical public PUF enabled by solving max-flow problem on chip · DAC 2016

Methods — techniques the papers use, named apart from their topics

shuffle-splice mechanism · 0.7variational graph autoencoder · 0.6neural process · 0.6graph neural process · 0.6gaussian process surrogate · 0.6max-flow problem · 0.5gradient-based optimization · 0.4discrete optimization · 0.4machine learning · 0.4active learning · 0.4source degeneration · 0.2
YearPublicationVenuePosition
2022 High-Speed Adder Design Space Exploration via Graph Neural Processes
abstract
Adders are the primary components in the data-path logic of a microprocessor, and thus, adder design has been always a critical issue in the very large-scale integration (VLSI) industry. However, it is infeasible for designers to obtain optimal adder architecture by exhaustively running EDA flow due to the extremely large design space. Previous arts have proposed the machine learning-based framework to explore the design space. Nevertheless, they fall into suboptimality due to a two-stage flow of the learning process and less efficient nor effective feature representations of prefix adder structures. In this article, we first integrate a variational graph autoencoder and a neural process (NP) into an end-to-end, multibranch framework, which is termed thegraph neural process. The former performs automatic feature learning of prefix adder structures, whilst the latter one is designed as an alternative to the Gaussian process. Then, we propose a sequential optimization framework with the graph NP as the surrogate model to explore the Pareto-optimal prefix adder structures with tradeoff among Quality-of-Result (QoR) metrics, such as power, area, and delay. The experimental results show that compared with state-of-the-art methodologies, our framework can achieve a much better Pareto frontier in multiple QoR metric spaces with fewer design-flow evaluations.
Hao Geng, Yuzhe Ma, Qi Xu 0004, Jin Miao, Subhendu Roy, Bei Yu 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2021 Correlated Multi-objective Multi-fidelity Optimization for HLS Directives Design
abstract
High-level synthesis (HLS) tools have gained great attention in recent years because it emancipates engineers from the complicated and heavy hardware description language writing, by using high-level languages and HLS directives. However, previous works seem powerless, due to the time-consuming design processes, the contradictions among design objectives, and the accuracy difference between the three stages (fidelities). To find good HLS directives, in this paper, a novel correlated multi-objective non-linear optimization algorithm is proposed to explore the Pareto solutions while making full use of data from different fidelities. A non-linear Gaussian process is proposed to model relationships among the analysis reports from different fidelities for the same objective. For the first time, correlated multivariate Gaussian process models are introduced into this domain to characterize the complex relationships of multiple objectives in each design fidelity. A tree-based method is proposed to erase invalid solutions and obviously non-optimal solutions. Experimental results show that our non-linear and pioneering correlated models can approximate the Pareto-frontier of the directive design space in a shorter time with much better performance and good stability, compared with the state-of-the-art.
Qi Sun 0002, Tinghuan Chen, Siting Liu 0002, Jin Miao, Jianli Chen, Hao Yu 0001, Bei Yu 0001
DATE4
2020 Hotspot Detection via Attention-based Deep Layout Metric Learning
abstract
With the aggressive and amazing scaling of the feature size of semiconductors, hotspot detection has become a crucial and challenging problem in the generation of optimized mask design for better printability. Machine learning techniques, especially deep learning, have attained notable success on hotspot detection tasks. However, most existing hotspot detectors suffer from suboptimal performance due to two-stage flow and less efficient representations of layout features. What is more, most works can only solve simple benchmarks with apparent hotspot patterns like ICCAD 2012 Contest benchmarks. In this paper, we firstly develop a new end-to-end hotspot detection flow where layout feature embedding and hotspot detection are jointly performed. An attention mechanism-based deep convolutional neural network is exploited as the backbone to learn embeddings for layout features and classify the hotspots simultaneously. Experimental results demonstrate that our framework achieves accuracy improvement over prior arts with fewer false alarms and faster inference speed on much more challenging benchmarks.
Hao Geng, Jin Miao, Fan Yang 0001, Xuan Zeng 0001, Bei Yu 0001
ICCAD4
2020 A Unified Framework for Simultaneous Layout Decomposition and Mask Optimization
abstract
In advanced technology nodes, layout decomposition (LD) and mask optimization (MO) are two key stages in integrated circuit design. Due to the inconsistency of the objectives of these two stages, the performance of conventional layout and MO may be suboptimal. To tackle this problem, in this article, we propose a unified framework, which seamlessly integrates LD and MO. We propose a gradient-based approach to solve the unified mathematical formulation, as well as a set of discrete optimization techniques to avoid being stuck in local optimum. The conventional optimization process can be accelerated as some inferior decomposition results can be smartly pruned in early stages. The experimental results show that the proposed unified framework can achieve more than 34× speed-up compared with the conventional two-stage flow, meanwhile, it can dramatically reduce EPE violations by more than 8×, and thus maintain better design quality.
Yuzhe Ma, Shuxiang Hu, Jhih-Rong Gao, Jian Kuang 0001, Jin Miao, Bei Yu 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.6
2019 Power-Driven DNN Dataflow Optimization on FPGA
abstract
Deep neural networks (DNNs) have been proven to achieve unprecedented success on modern artificial intelligence (AI) tasks, which have also greatly motivated the rapid developments of novel DNN models and hardware accelerators. Many challenges still remain towards the design of power efficient DNN accelerator due to the intrinsically intensive data computation and transmission in DNN algorithms. However, most existing efforts in the domain have taken latency as the sole optimization objective, which may often result in sub-optimality in power consumption. In this paper, we propose a framework to optimize the power efficiency of DNN dataflow on FPGA while maximally minimizing the impact on latency. We first propose power and latency models that are built upon different dataflow configurations. Then a power-driven dataflow formulation is proposed, which enables a hierarchical exploration strategy on the dataflow configurations, leading to efficient power consumption at limited latency loss. Experimental results have demonstrated the effectiveness of our proposed models and exploration strategies, where power improvement has shown up to 31% with latency degradation of no worse than 6.5%.
Qi Sun 0002, Tinghuan Chen, Jin Miao, Bei Yu 0001
ICCAD3
2019 Cross-Layer Optimization for High Speed Adders: A Pareto Driven Machine Learning Approach
abstract
In spite of maturity to the modern electronic design automation (EDA) tools, optimized designs at architectural stage may become suboptimal after going through physical design flow. Adder design has been such a long studied fundamental problem in very large-scale integration industry yet designers cannot achieve optimal solutions by running EDA tools on the set of available prefix adder architectures. In this paper, we enhance a state-of-the-art prefix adder synthesis algorithm to obtain a much wider solution space in architectural domain. On top of that, a machine learning-based design space exploration methodology is applied to predict the Pareto frontier of the adders in physical domain, which is infeasible by exhaustively running EDA tools for innumerable architectural solutions. Considering the high cost of obtaining the true values for learning, an active learning algorithm is proposed to select the representative data during learning process, which uses less labeled data while achieving better quality of Pareto frontier. Experimental results demonstrate that our framework can achieve Pareto frontier of high quality over a wide design space, bridging the gap between architectural and physical designs. Source code and data are available athttps://github.com/yuzhe630/adder-DSE.
Yuzhe Ma, Subhendu Roy, Jin Miao, Bei Yu 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2018 SD-PUF: Spliced Digital Physical Unclonable Function
abstract
Digital circuit physical unclonable function (PUF) has been attracting attentions for the merits of resilience to the environmental and operational variations that analog PUFs suffer from. Existing state-of-the-art digital circuit PUFs, however, are either hybrid of analog-digital circuits which are still under the shadow of vulnerability, or impractical for real-world applications. In this paper, we propose a novel highly nonlinear and secure digital PUF (D-PUF) and the spliced version SD-PUF. The fingerprints are extracted from intentionally induced very large-scale integration interconnect randomness during lithography process, as well as a post-silicon shuffling process. Strongly skewed CMOS latches are used to ensure the immunity against environmental and operational variations. Crucially, a highly nonlinear logic network is proposed to effectively spread and augment any subtle interconnect randomness, which also enables strong resilience against machine learning attacks. On top of it, the expandable architecture of the proposed logic network empowers a novel post-silicon shuffle-splice mechanism, where multiple randomly selected D-PUFs are spliced to be one SD-PUF, pushing the statistical security to a much higher level, while significantly reducing the mask cost per PUF device. It also decouples the trustworthy demands enforced to the foundries or other third party manufacturers. Our proposed PUFs demonstrate close to ideal performance in terms of statistical metrics, including 0 intra-Hamming distance. Various state-of-the-art machine learning models show prediction accuracies almost no better than random guesses when attacking to the proposed PUFs. We also mathematically prove the probability of existence of identical SD-PUF pair is significantly lower than that of D-PUF pair, e.g., such probability of an SD-PUF spliced by 30 D-PUFs is 2.3 × 10-22, which is 19 order magnitude lower than that of D-PUF. Benefited from the proposed shuffle-splice mechanism, the mask cost per SD-PUF is also reduced by 300× than that of D-PUF.
Jin Miao, Meng Li 0004, Subhendu Roy, Yuzhe Ma, Bei Yu 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2017 A unified framework for simultaneous layout decomposition and mask optimization
abstract
In advanced technology nodes, layout decomposition and mask optimization are two key stages in integrated circuit design. Due to the inconsistency of the objectives of these two stages, the performance of conventional layout and mask optimization may be suboptimal. To tackle this problem, in this paper we propose a unified framework, which seamlessly integrates layout decomposition and mask optimization. We propose a gradient based approach to solve the unified mathematical formulation, as well as a set of discrete optimization techniques to avoid being stuck in local optimum. The conventional optimization process can be accelerated as some inferior decompositions can be smartly pruned in early stages. The experimental results show that the proposed unified framework can achieve more than 17 x speed-up compared with the conventional two-stage flow, meanwhile it can reduce EPE violations by 18%, and thus maintain better design quality.
Yuzhe Ma, Jhih-Rong Gao, Jian Kuang 0001, Jin Miao, Bei Yu 0001
ICCAD4
2017 A learning bridge from architectural synthesis to physical design for exploring power efficient high-performance adders
abstract
In spite of maturity to the modern electronic design automation (EDA) tools, optimized designs at architectural stage may become sub-optimal after going through physical design flow. Adder design has been such a long studied fundamental problem in VLSI industry yet designers cannot achieve optimal solutions by running EDA tools on the set of available prefix adder architectures. In this paper, we enhance a state-of-the-art prefix adder synthesis algorithm to obtain a much wider solution space in architectural domain. On top of that, a machine learning based design space exploration methodology is applied to predict the Pareto frontier of the adders in physical domain, which is infeasible by exhaustively running EDA tools for innumerable architectural solutions. Experimental results demonstrate that our framework can achieve near-optimal delay vs. power/area Pareto frontier over a wide design space, bridging the gap between architeon the set of available prefix adder architectures. In this paper, we enhance a state-of-the-art prefix adder synthesis algorithm to obtain a much wider solution space in architectural domain. On top of that, a machine learning based design space exploration methodology is applied to predict the Pareto frontier of the adders in physical domain, which is infeasible by exhaustively running EDA tools for innumerable architectural solutions. Experimental results demonstrate that our framework can achieve near-optimal delay vs. power/area Pareto frontier over a wide design space, bridging the gap between architectural andctural and physical designs.
Subhendu Roy, Yuzhe Ma, Jin Miao, Bei Yu 0001
ISLPED3
2016 Practical public PUF enabled by solving max-flow problem on chip
abstract
The execution-simulation gap (ESG) is a fundamental property of public physical unclonable function (PPUF), which exploits the time gap between direct IC execution and computer simulation. ESG needs to consider both advanced computing scheme, including parallel and approximate computing scheme, and IC physical realization. In this paper, we propose a novel PPUF design, whose execution is equivalent to solving the hard-to-parallel and hard-to-approximate max-flow problem in a complete graph on chip. Thus, max-flow problem can be used as the simulation model to bound the ESG rigorously. To enable an efficient physical realization, we propose a crossbar structure and adopt source degeneration technique to map the graph topology on chip. The difference on asymptotic scaling between execution delay and simulation time is examined in the experimental results. The measurability of output difference is also verified to prove the physical practicality.
Meng Li 0004, Jin Miao, Kai Zhong 0007, David Z. Pan
DAC2
2016 LRR-DPUF: learning resilient and reliable digital physical unclonable function
abstract
Conventional silicon physical unclonable function (PUF) extracts fingerprints from transistor's analog attributes, which are vulnerable to environmental and operational variations. Recently, digitalized PUF prototypes have emerged to overcome the vulnerability issues, however, the existing prototypes are either hybrid of analog-digital PUFs which are still under the shadow of vulnerability, or impractical for real-world implementation. To address the above limitations, we propose a learning resilient and reliable digital PUF (LRR-DPUF). The fingerprints are extracted from VLSI interconnect geometrical randomness induced by lithography variations. Crucially, we use strongly skewed latches to ensure the immunity against environmental and operational variations. Further, a cross-coupled, highly non-linear logic network is proposed to effectively spread and augment even subtle interconnect randomness, as well as to achieve strong resilience to machine learning attacks. We demonstrate that a 64-bit LRR-DPUF exhibits close to ideal statistical performances, including 0 intra Hamming Distance. We also mathematically prove that each output of the LRR-DPUF follows uniform distribution. Various state-of-the-art machine learning models show almost no better than random prediction accuracies when applied to LRR-DPUF.
Jin Miao, Meng Li 0004, Subhendu Roy, Bei Yu 0001
ICCAD1
2014 Multi-level approximate logic synthesis under general error constraints
abstract
We address the problem of multi-level approximate logic synthesis. Our strategy assumes existence of an optimized exact Boolean network, which is critical in practice since arithmetic blocks are rarely synthesized from 2-level representation automatically. The goal is to produce minimum cost circuits whose logic function deviates in a controlled manner from the exact function with deviations quantified by the magnitude and frequency of errors. We rely on network simplifications allowed by external don't cares (EXDCs). We formulate the error-magnitude constrained problem by using Boolean relations to capture the allowed error behavior in a more general manner compared to incompletely specified functions. Our key contribution is in finding sets of external don't cares that maximally approach the Boolean relation. The algorithm starts with an EXDC set that is overly relaxed and iteratively, and in a greedy fashion, identifies a feasible EXDC set by solving a series of conventional EXDC-based network optimizations. The algorithm then ensures compliance to error frequency constraints by recovering the correct outputs on the sought number of error-producing inputs while aiming to minimize the network cost increase. We applied the algorithm to several well-known adder and multiplier designs of varying bit-width. Even for small error magnitudes, the algorithm produces networks with gate count reduced by 30-50%, when the error frequency constraint is loose. This is up to 20% fewer gates than a naive EXDC-based approach.
Jin Miao, Andreas Gerstlauer, Michael Orshansky
ICCAD1
2013 Approximate logic synthesis under general error magnitude and frequency constraints
abstract
Recent interest in approximate circuit design is driven by its potential for large energy savings. In this paper, we address the problem of approximate logic synthesis (ALS). ALS is concerned with formally synthesizing a minimum-cost approximate Boolean network whose behavior deviates in a well-defined manner from a specified exact Boolean function, where in this work, we allow the deviation to be constrained by both the magnitude and frequency of the error. We make two contributions in solving this general ALS problem: The first contribution is to establish that the approximate synthesis problem un-constrained by the frequency of errors is isomorphic with the Boolean relations (BR) minimization problem. That equivalence allows us to exploit recently developed fast algorithms for BR problems to solve the error magnitude-only constrained ALS problem. The second contribution is an efficient heuristic algorithm for iteratively refining the magnitude-constrained solution to arrive at a solution also satisfying the error frequency constraint. Our combined greedy approximate logic synthesis (GALS) approach is able to operate on any Boolean network for which the deviation measures can be specified and is most immediately applicable to arithmetic blocks. Experiments on adder and multiplier blocks demonstrate literal count reductions of up to 60% under tight error frequency and magnitude constraints.
Jin Miao, Andreas Gerstlauer, Michael Orshansky
ICCAD1
2012 Modeling and synthesis of quality-energy optimal approximate adders
abstract
Recent interest in approximate computation is driven by its potential to achieve large energy savings. This paper formally demonstrates an optimal way to reduce energy via voltage over-scaling at the cost of errors due to timing starvation in addition. We identify a fundamental trade-off between error frequency and error magnitude in a timing-starved adder. We introduce a formal model to prove that for signal processing applications using a quadratic signal-to-noise ratio error measure, reducing bit-wise error frequency is sub-optimal. Instead, energy-optimal approximate addition requires limiting maximum error magnitude. Intriguingly, due to possible error patterns, this is achieved by reducing carry chains significantly below what is allowed by the timing budget for a large fraction of sum bits, using an aligned, fixed internal-carry structure for higher significance bits.
Jin Miao, Ku He, Andreas Gerstlauer, Michael Orshansky
ICCAD1