EDBT 2026 Demo / reviewers in the wild / expert
Lang Feng 0001
dblp:211/0071-1
· DBLP profile ↗
28ranked-venue papers
10as first author
23since 2021 · last 2026
0000-0001-9943-0550ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 26 · 8 first-author · 23 since 2021Software engineering, systems software and programming languages · 4 · 2 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | N for One: Reticle-Reuse-Driven Routing for Silicon InterposersabstractAs ultra-large-scale integrated circuits continue to evolve, advanced packaging has become a promising approach to boost system performance, especially in high-performance computing applications. To support heterogeneous integration, multiple chiplets are integrated via large-area silicon interposers. However, due to reticle size limitations, interposer lithography requires multi-reticle stitching, which introduces high manufacturing costs and alignment-induced yield degradation. In this paper, we propose CIT-R3, a reticle-reuse-driven router that formulates routing and reticle reuse as a differentiable optimization problem. Additionally, a graph-patching algorithm is applied to enforce layout consistency across reused reticle regions. Experimental results on multi-chiplet benchmarks demonstrate significant reticle reuse improvements with minimal routing cost overhead. Xiaokun Lin, Lang Feng 0001, Jixiang Zhu, Xupengkai Lu, Ying Wang 0001, Fengwei Dai, Yinhe Han 0001 |
ACM Great Lakes Symposium on VLSI | 3 |
| 2026 | A Potential-Guided Efficient Timing-Driven Obstacle-Avoiding Routing Tree Algorithm
Changhao Sun, Hongxin Kong, Lang Feng 0001 |
ISCAS | 3 |
| 2026 | A High Efficient and Scalable Obstacle-Avoiding VLSI Global Routing FlowabstractRouting is a crucial step in the VLSI design flow. With advancements in manufacturing technology, more constraints have emerged in design rules, particularly regarding obstacles during routing, leading to increased routing complexity. Unfortunately, many global routers struggle to generate efficient obstacle-free solutions due to the lack of scalable obstacle-avoiding tree generation methods and the capability to handle modern designs with complex obstacles and nets. In this work, we propose an efficient obstacle-aware global routing flow for VLSI designs with obstacles. The flow includes a rule-based obstacle-avoiding rectilinear Steiner minimal tree (OARSMT) algorithm during the tree generation phase. This algorithm is both scalable and fast, providing tree topologies avoiding obstacles in the early stage globally. With its guidance, in the later stages, the OARSMT-guided and obstacle-aware sparse maze routing are proposed to further minimize obstacle violations and reduce overflow costs. Compared to previously advanced methods on the benchmark with obstacles, our approach successfully eliminates obstacle violations and reduces wirelength and overflow cost, while sacrificing only a limited number of via counts and runtime overhead. Junhao Guo, Hongxin Kong, Lang Feng 0001 |
ACM Trans. Design Autom. Electr. Syst. | 3 |
| 2026 | ETOF: An Efficient Transistor-Level Optimization Flow for Large-Scale CMOS CircuitsabstractThe computing efficiency of digital VLSIs has increased with advancements in manufacturing, but this progress is slowing due to post-Moore physical limits. To improve efficiency, better standard cell-level synthesis can reduce transistor counts for low-power design. However, the design space at this level is limited, leaving room for transistor-level optimization. While previous research has explored transistor-level optimization, most focus on small-scale circuits, and few large-scale approaches are coarse-grained and lack a global perspective. In this article, we propose an efficient transistor-level optimization flow for CMOS VLSIs. It includes (1) a partition algorithm with a fast quality estimation method based on a metric named weighted cell sharing rate, (2) a neural network model with dedicated feature selection to provide an accurate optimization potential evaluation, and (3) an effective iterative partition selection method with global consideration of the partitions’ dependencies, for obtaining partitions suitable for transistor-level synthesis tools. This flow can optimize a given digital circuit’s netlist for reducing the transistor count. The experimental results demonstrate that the proposed flow achieves an average reduction of 11.04% and 7.94% in transistor counts compared to standard cell logic synthesis and the advanced large-scale transistor-level optimization work, respectively. Runquan Lei, Lang Feng 0001, Zetao Zhang |
ACM Trans. Design Autom. Electr. Syst. | 2 |
| 2025 | An MIP-based Force-directed Large Scale Placement Refinement AlgorithmabstractPlacement is an important part in the flow of physical design, which can affect the performance of a circuit significantly. Many algorithms have been proposed to refine the placement in the past years and Mixed Integer Programming (MIP) is one of the directions that can further improve the placement quality, since MIP is able to perform a finer-grained placement with precise MIP formulations. Many previous MIP works try to prune the search space for efficiency, but the strategies for selecting valuable search space do not contain enough analysis of the initial placement before refinement. In this work, we propose an MIP-based algorithm that can refine large scale placement by considering more global factors from initial placement, while achieving the trade-off between efficiency and quality. A force-directed displacement technique is proposed, which quantifies multiple metrics in each orientation to assign a potential region for each cell's movement. Meanwhile, we also propose an accurate wirelength prediction method for high-degree nets by introducing the concept of centroid for net breaking. Experiments on benchmarks of ISPD18 and ISPD19 show that our algorithm is able to reduce the wirelength and vias by 1.02% and 0.58% on average and our work outperforms the state-of-the-art related work in wirelength optimization under both its comprehensive mode and wirelength-only mode. Ke Tang 0006, Lang Feng 0001, Zhongfeng Wang 0001 |
ASP-DAC | 3 |
| 2025 | SeDA: Secure and Efficient DNN Accelerators with Hardware/Software SynergyabstractEnsuring the confidentiality and integrity of DNN accelerators is paramount across various scenarios spanning autonomous driving, healthcare, and finance. However, current security approaches typically require extensive hardware resources, and incur significant off-chip memory access overheads. This paper introduces SeDA, which utilizes 1) a bandwidth-aware encryption mechanism to improve hardware resource efficiency, 2) optimal block granularity through intra-layer and inter-layer tiling patterns, and 3) a multi-level integrity verification mechanism that minimizes, or even eliminates, memory access overheads. Experimental results show that SeDA decreases performance overhead by over 12% for both server and edge neural processing units (NPUs), while ensuring robust scalability.11SeDA source code:https://github.com/wayne4s/seda.git Lang Feng 0001, Ning Lin, Zihao Xuan, Rongliang Fu, Tsung-Yi Ho, Yuzhong Jiao, Luhong Liang |
DAC | 3 |
| 2025 | Hybrid Exact and Heuristic Efficient Transistor Network Optimization for Multi-Output LogicabstractWith the approaching post-Moore era, it is becoming increasingly impractical to decrease the transistor size in digital VLSI for better performance. To address this issue, one approach is to optimize the digital circuit at the transistor level to reduce the transistor count. Although previous works have explored ways to conduct transistor network optimization, most of these efforts have focused on single-output networks or applied heuristics only, limiting their scope or optimization quality. In this paper, we propose an exact transistor network optimization algorithm that supports multi-output logic and is formulated as a SAT problem. Our approach maintains a high optimization level by employing the exact algorithm, while also incorporating a hybrid process that uses a heuristic algorithm to predict the solution range as a guidance for better efficiency. Experimental results show that the proposed algorithm has a 5.32% better optimization level given 54% less runtime compared with the state-of-the-art work. Lang Feng 0001, Rongjian Liang, Hongxin Kong |
DATE | 1 |
| 2025 | CIT-CTPlacer: An Analytical RDL Chiplet-Terminal Co-Placement Algorithm for Large-Scale 2.5D ICabstractAs the number of chiplets in 2.5D IC continues to increase, existing chiplet placement method faces two main challenges: (1) the combinatorial explosion in the search space, and (2) the difficulty of achieving global optimization through iterativing chiplet and terminal placement. To tackle these challenges, we develop an efficient analytical RDL chiplet-terminal co-placement algorithm, to ensure simultaneous placement of chiplets and terminals. Our algorithm employs RDL chiplet-terminal co-placement in three stages: analytical global placement, legalization, and bump-terminal assignment, to achieve high-quality placement results that comply with design rules. Experimental results demonstrate that our algorithm reduces average wirelength by 31% compared to prior work for the common testcases, with a maximum speedup of up to 6500× in testcases with more than 10 chiplets. Xihao Liang, Xupengkai Lu, Lang Feng 0001, Jixiang Zhu, Ying Wang 0001, Yinhe Han 0001 |
ISCAS | 4 |
| 2025 | PreSIT: Predict Cryptography Computations in SGX-Style Integrity TreesabstractIn recent years, SGX-style integrity trees (SITs) have been applied in trusted execution environments (TEEs) to protect the off-chip memory security from physical attacks, replay attacks, etc. As a tradeoff, SIT implementations incur a huge performance overhead due to its extra computations and memory accesses. Recent works reduced the memory accesses and significantly increased the speed, but the performance overhead is still high. To tackle this challenge, this article performs an insight analysis under current SIT implementations, then identifies two critical remaining performance overhead causes: 1) hash and 2) decryption computations. Next, a novel design named PreSIT based on a proposed parallel operation flow is introduced. It performs prefetching and precomputes partial results of hash and decryption processes. By designing the prediction algorithm flow and predict-assisted algorithms, along with the proposed dedicated prediction data structure to effectively store and access the precomputed results, the performance overhead is further reduced. According to the evaluations of GEM5 on SPEC 2017, GAP, PARSEC, and SPLASH2x, PreSIT improves the average instruction per cycle (IPC) by 5.1% (at most 21.6%) based on the SIT in VAULT without any security degradation. Lang Feng 0001, Zhongfeng Wang 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2025 | Resister: A Resilient Interposer Architecture for Chiplet to Mitigate Timing Side-Channel AttacksabstractChiplet technology has been a hot topic due to its potential for more efficient implementation of large-scale integrated circuits. In chiplet manufacturing, the general-purpose active interposer usually integrates chiplets from different vendors with a typical mesh network. This method of manufacturing is broadly recognized for its cost-efficiency. However, untrusted vendors make the chiplet system vulnerable to security threats such as timing side-channel attacks (TSA) based on network contention information. Even worse, the reliability of each chiplet is usually unknown beforehand to a general-purpose interposer’s manufacturer, so that TSAs can be on arbitrary chiplets at arbitrary time in the manufacturer’s view. To address this challenge, this work first quantitatively analyzes the attack patterns including reinforced styles, based on which, a resilient interposer architecture named Resister is proposed. A hardware defender is designed in every router to globally detect the malicious transaction patterns at runtime, and adaptively detour the transaction packets accordingly for security while maintaining the performance. According to the evaluation of GEM5 on SPEC 2017 and PARSEC benchmarks, Resister can effectively mitigate TSA with only a 1.7% performance overhead. Lang Feng 0001, Taotao Xu, Yinhe Han 0001, Zhongfeng Wang 0001 |
ACM Trans. Design Autom. Electr. Syst. | 2 |
| 2025 | A CPU+FPGA OpenCL Heterogeneous Computing Platform for Multi-Kernel PipelineabstractOver the past decades, Field-Programmable Gate Arrays (FPGAs) have become a choice for heterogeneous computing due to their flexibility, energy efficiency, and processing speed. OpenCL is used in FPGA heterogeneous computing for its high-level abstraction and cross-platform compatibility. Previous works have introduced optimization techniques in OpenCL for FPGAs to leverage FPGA-specific advantages. However, the multi-kernel pipeline technique, which can raise throughput and resource utilization, has not performed well. This article presents a CPU+FPGA heterogeneous platform with a novel execution model to optimize multi-kernel pipeline. Firstly, we extend OpenCL by introducing new APIs and additional functions to represent the execution model. Secondly, a hardware-software co-scheduling scheme is employed to manage execution. Thirdly, we design a holistic development flow and toolkit to facilitate the deployment of algorithms on the platform or the integration of RTL IP cores to the OpenCL environment. We validate the platform using a Range Doppler algorithm. The proposed development flow and integrated toolchain enhance the efficiency of integrating traditional RTL IP cores into the OpenCL environment. Experimental results demonstrate that, with a comparable processing speed (averaging 95%) to traditional RTL implementations, the platform successfully establishes the multi-kernel pipelines. Leveraging the multi-kernel pipeline, the platform achieves a significant improvement in multi-frame processing speed compared to traditional OpenCL. Yuefei Wang, Wendong Mao, Lang Feng 0001, Jin Sha 0001, Zhongfeng Wang 0001 |
ACM Trans. Design Autom. Electr. Syst. | 3 |
| 2024 | Hardware-Assisted Control-Flow Integrity Enhancement for IoT DevicesabstractInternet of Things (IoT) devices face an escalating threat from code reuse attacks (CRAs) as they can reuse existing code for malicious purpose. Thus a practical cost-effective Control-Flow Integrity (CFI) mechanism for IoT devices is urgently needed. However, existing CFI solutions suffer from impractical-ities, including high performance overhead and a heavy reliance on offline perfect Control-Flow Graph (CFG) generation. To tackle these challenges, we propose a fine-grained dependable CFI scheme for IoT devices that real-time updates the CFG of devices. We evaluate the implementation on RISC-V architectures and the results show that our CFI scheme provides both backward- and forward-edge protection with almost no performance overhead in the case of fixed CFG, negligible power overhead, and low hardware overhead. Compared to the current hardware-assisted CFI designs, our design eliminates the dependence on the offline perfect CFG generation and performs real-time CFG updating for better practicality. Lang Feng 0001, Zhiguo Shi 0001, Cheng Zhuo, Jiming Chen 0001 |
DATE | 2 |
| 2024 | A Rule-Based High Efficient Obstacle-Avoiding RSMT Algorithm for VLSI RoutingabstractFor VLSI physical design, the routing problem has attracted attention in recent years due to the emerging manufacturing technologies. Tree generation is one key routing step directly affecting the routing quality, which is to find the rectilinear steiner minimal tree (RSMT) of each net. Ordinary RSMT algorithms such as FLUTE fail to generate valid trees avoiding obstacles. In contrast, current obstacle-avoiding RSMT (OARSMT) algorithms can incur a large runtime overhead compared with FLUTE. To reduce the runtime cost while maintaining the quality, a novel OARSMT algorithm is proposed in this work. By proposing multiple rule-based routing schemes, which are fast while maintaining the awareness of global conditions from mature RSMT solutions, OARSMT solutions with reasonable qualities can be quickly obtained, even for large and complicated cases. Compared with the state-of-the-art works, traded with limited wirelength overhead, the proposed algorithm has ∼10x-2700x and ∼150x-5800x runtime speedup under randomized testcases and standard benchmarks, respectively. Junhao Guo, Hongxin Kong, Lang Feng 0001 |
ISCAS | 3 |
| 2024 | RISC-V Custom Instructions of Elementary Functions for IoT Endpoint DevicesabstractThe computation of elementary functions is required in many tasks of Internet of Things (IoT) endpoint devices, for example, communications, image processing, and biomedical signal processing. IoT endpoint devices generally adopt software approaches to compute elementary functions, which take many cycles. To improve efficiency, this work proposes custom instructions for elementary functions to the open-source RISC-V instruction set architecture (ISA). In particular, several variants of the custom instructions (fast, intermediate, and tiny variants) are developed to satisfy the needs of various types of IoT devices. Microarchitecture design and VLSI circuit design are then proposed to efficiently support the extended ISA. Both software emulation and on-board evaluation of the new architecture are carried out with testbenches covering typical communication and computation tasks for IoT devices. The custom instructions gain speedups ranging from 3.3 to 18.0 compared to a baseline RV32IM design. ASIC synthesis results under TSMC 28nm technology demonstrate that the power overhead is$ \lt $5% with the tiny variant,$ \lt $17% with the intermediate variant, and$ \lt $26% with the fast variant, which is not significant considering the achieved speedup. The experimental results further confirm that the proposed custom instructions are computation-efficient and versatile to adapt to different IoT devices for various applications. Yuxing Chen 0001, Suwen Song, Lang Feng 0001, Zhongfeng Wang 0001 |
IEEE Trans. Computers | 4 |
| 2024 | Prefender: A Prefetching Defender Against Cache Side Channel Attacks as a PretenderabstractCache side channel attacks are increasingly alarming in modern processors due to the recent emergence of Spectre and Meltdown attacks. A typical attack performs intentional cache access and manipulates cache states to leak secrets by observing the victim’s cache access patterns. Different countermeasures have been proposed to defend against both general and transient execution based attacks. Despite their effectiveness, they mostly trade some level of performance for security, or have restricted security scope. In this paper, we seek an approach to enforcing security while maintaining performance. We leverage the insight that attackers need to access cache in order to manipulate and observe cache state changes for information leakage. Specifically, we propose Prefender, a secure prefetcher that learns and predicts attack-related accesses for prefetching the cachelines to simultaneously help security and performance. Our results show that Prefenderis effective against several cache side channel attacks while maintaining or even improving performance for SPEC CPU 2006 and 2017 benchmarks. Jiayi Huang 0001, Lang Feng 0001, Zhongfeng Wang 0001 |
IEEE Trans. Computers | 3 |
| 2024 | Mixed Integer Programming based Placement Refinement by RSMT Model with Movable PinsabstractPlacement is a critical step in the physical design for digital application specific integrated circuits (ASICs), as it can directly affect the design qualities such as wirelength and timing. For many domain specific designs, the demands for high performance parallel computing result in repetitive hardware instances, such as the processing elements in the neural network accelerators. As these instances can dominate the area of the designs, the runtime of the complete design’s placement can be traded for optimizing and reusing one instance’s placement to achieve higher quality. Therefore, this work proposes a mixed integer programming (MIP)-based placement refinement algorithm for the repetitive instances. By efficiently modeling the rectilinear steiner tree wirelength, the placement can be precisely refined for better quality. Besides, the MIP formulations for timing-driven placement are proposed. A theoretical proof is then provided to show the correctness of the proposed wirelength model. For the instances in various popular fields, the experiments show that given the placement from the commercial placers, the proposed algorithm can perform further placement refinement to reduce 3.76%/3.64% detailed routing wirelength and 1.68%/2.42% critical path delay under wirelength/timing-driven mode, respectively, and also outperforms the state-of-the-art previous work. Ke Tang 0006, Lang Feng 0001, Zhongfeng Wang 0001 |
ACM Trans. Design Autom. Electr. Syst. | 2 |
| 2023 | 1+1 <2: Efficient Automatic Standard Cell Sharing Between Digital VLSI Designs for Area SavingabstractIn the field of digital VLSI design, multimode circuits are the designs where the modes can be switched according to different application scenarios, and are commonly used in communication systems. In a multimode circuit, different modes are usually implemented by different circuits, which can lead to large circuit area consumption. For different modes, sharing their isomorphic circuit regions in the standard cell level can save the area. This goal is similar to that in the subgraph isomorphism problem, which is to check if a given graph is a subgraph of another one. However, subgraph isomorphism needs unacceptable runtime to solve as it is NP-complete. Even worse, finding the largest isomorphic regions of different circuits is a problem harder than subgraph isomorphism. In this article, we propose a novel algorithm for efficiently finding enough isomorphic circuit regions of different digital circuits in polynomial time, and give the theoretical proof of the correctness. The experiments show that the proposed approach can save 20%–25% area on average by sharing the standard cells between 2 and 4 circuits, while keeping the functional correctness. The proposed algorithm also has reasonable runtime and mostly incurs negligible timing overhead. Lang Feng 0001, Jin Sha 0001, Zhongfeng Wang 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2023 | ProMiSE: A High-Performance Programmable Hardware Monitor for High Security Enforcement of Software ExecutionabstractIn recent years, to prevent computer systems from software attacks, hardware monitors are proposed as a type of efficient security enforcement scheme, which can detect software attacks at runtime. However, due to the limited flexibility of dedicated hardware monitors, one monitor can be only applied to a few targeted application scenarios, and is hard to defend against unconsidered attacks. This leads to high cost for redesigning monitors for new scenarios. Although recent studies propose flexible hardware monitors, the scope and security of the reconfigurable monitoring policies are still limited. To further improve the flexibility and security, this work proposes a monitor instruction set and multiple security-assisting designs for supporting general operations needed by various attack detection schemes. Based on the above efforts, an efficient programmable hardware monitor named ProMiSE is designed. After implemented on the RocketChip RISC-V processor, ProMiSE can be programmed to realize a wider range of monitoring policies with higher security and similar hardware resource overhead, compared with stateof-the-art flexible hardware monitors. With these advantages, ProMiSE still has the detection latency as low as 18-59 CPU cycles. The performance overhead ranges from 0%-23.4%, which is also reasonable compared with the dedicated hardware monitors of corresponding policies. Lang Feng 0001, Zhongfeng Wang 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2023 | GANDSE: Generative Adversarial Network-based Design Space Exploration for Neural Network Accelerator DesignabstractWith the popularity of deep learning, the hardware implementation platform of deep learning has received increasing interest. Unlike the general purpose devices, e.g., CPU or GPU, where the deep learning algorithms are executed at the software level, neural network hardware accelerators directly execute the algorithms to achieve higher energy efficiency and performance improvements. However, as the deep learning algorithms evolve frequently, the engineering effort and cost of designing the hardware accelerators are greatly increased. To improve the design quality while saving the cost, design automation for neural network accelerators was proposed, where design space exploration algorithms are used to automatically search the optimized accelerator design within a design space. Nevertheless, the increasing complexity of the neural network accelerators brings the increasing dimensions to the design space. As a result, the previous design space exploration algorithms are no longer effective enough to find an optimized design. In this work, we propose a neural network accelerator design automation framework named GANDSE, where we rethink the problem of design space exploration, and propose a novel approach based on the generative adversarial network (GAN) to support an optimized exploration for high-dimension large design space. The experiments show that GANDSE is able to find the more optimized designs in negligible time compared with approaches including multilayer perceptron and deep reinforcement learning. Lang Feng 0001, Chuliang Guo, Ke Tang 0006, Cheng Zhuo, Zhongfeng Wang 0001 |
ACM Trans. Design Autom. Electr. Syst. | 1 |
| 2022 | PREFENDER: A Prefetching Defender against Cache Side Channel Attacks as A PretenderabstractCache side channel attacks are increasingly alarming in modern processors due to the recent emergence of Spectre and Meltdown attacks. A typical attack performs intentional cache access and manipulates cache states to leak secrets by observing the victim's cache access patterns. Different countermeasures have been proposed to defend against both general and transient execution based attacks. Despite their effectiveness, they all trade some level of performance for security. In this paper, we seek an approach to enforcing security while maintaining performance. We leverage the insight that attackers need to access cache in order to manipulate and observe cache state changes for information leakage. Specifically, we propose PREFENDER,a secure prefetcher that learns and predicts attack-related accesses for prefetching the cachelines to simultaneously help security and performance. Our results show that PREFENDER is effective against several cache side channel attacks while maintaining or even improving performance for SPEC CPU2006 benchmarks. Jiayi Huang 0001, Lang Feng 0001, Zhongfeng Wang 0001 |
DATE | 3 |
| 2022 | RvDfi: A RISC-V Architecture With Security Enforcement by High Performance Complete Data-Flow IntegrityabstractWith the rapid revolution of open-source hardware, RISC-V architecture has been prevalent in both academic research and industrial developments. Due to the increasing threats of information leakage, it is imperative to provide a secure RISC-V ecosystem to defend against malicious software exploits. Toward this goal, data-flow integrity (DFI) is employed as a strict security policy for enforcing the legitimacy of each data access, thereby filtering out most of the attack exploits. However, due to the intensive computations needed by DFI, there are only limited proposals successfully implementing partial DFI with low performance overhead. Moreover, all the previous studies failed to enforce thecompleteDFI policy in a real hardware platform, while trading off security strength for performance efficiency. To provide RISC-V architecture with high security enforcement and low performance overhead, we leverage the open-source Rocket Chip and proposeRvDfi, the first complete DFI implementation based on RISC-V architecture with only 17.8% performance overhead on average and 3.9% in minimum, incurring much less performance loss compared to the 166.3% overhead caused by previous complete DFI implementation. Lang Feng 0001, Jiayi Huang 0001, Zhongfeng Wang 0001 |
IEEE Trans. Computers | 1 |
| 2022 | Toward Taming the Overhead Monster for Data-flow IntegrityabstractData-Flow Integrity (DFI) is a well-known approach to effectively detecting a wide range of software attacks. However, its real-world application has been quite limited so far because of the prohibitive performance overhead it incurs. Moreover, the overhead is enormously difficult to overcome without substantially lowering the DFI criterion. In this work, an analysis is performed to understand the main factors contributing to the overhead. Accordingly, a hardware-assisted parallel approach is proposed to tackle the overhead challenge. Simulations on SPEC CPU 2006 benchmark show that the proposed approach can completely enforce the DFI defined in the original seminal work while reducing performance overhead by 4×, on average. Lang Feng 0001, Jiayi Huang 0001, Jeff Huang 0001, Jiang Hu 0001 |
ACM Trans. Design Autom. Electr. Syst. | 1 |
| 2021 | FastCFI: Real-time Control-Flow Integrity Using FPGA without Code InstrumentationabstractControl-Flow Integrity (CFI) is an effective defense technique against a variety of memory-based cyber attacks. CFI is usually enforced through software methods, which entail considerable performance overhead. Hardware-based CFI techniques can largely avoid performance overhead, but typically rely on code instrumentation, forming a non-trivial hurdle to the application of CFI. Taking advantage of the tradeoff between computing efficiency and flexibility of FPGA, we develop FastCFI, an FPGA-based CFI system that can perform fine-grained and stateful checking without code instrumentation. We also propose an automated Verilog generation technique that facilitates fast deployment of FastCFI, and a compression algorithm for reducing the hardware expense. Experiments on popular benchmarks confirm that FastCFI can detect fine-grained CFI violations over unmodified binaries. When using FastCFI on prevalent benchmarks, we demonstrate its capability to detect fine-grained CFI violations in unmodified binaries, while incurring an average of 0.36% overhead and a maximum of 2.93% overhead. Lang Feng 0001, Jeff Huang 0001, Jiang Hu 0001, Abhijith Reddy |
ACM Trans. Design Autom. Electr. Syst. | 1 |
| 2019 | Layout recognition attacks on split manufacturingabstractOne technique to prevent attacks from an untrusted foundry is split manufacturing, where only a part of the layout is sent to the untrusted high-end foundry, and the rest is manufactured at a trusted low-end foundry. The untrusted foundry has front-end-of-line (FEOL) layout and the original circuit netlist and attempts to identify critical components on the layout for Trojan insertion. Although defense methods for this scenario have been developed, the corresponding attack technique is not well explored. For instance, Boolean satisfiability (SAT) based bijective mapping attack is mentioned without detailed research. Hence, the defense methods are mostly evaluated with the k-security metric without actual attacks. We provide the first systematic study, to the best of our knowledge, on attack techniques in this scenario. Besides of implementing SAT-based bijective mapping attack, we develop a new attack technique based on structural pattern matching. Experimental comparison with bijective mapping attack shows that the new attack technique achieves about the same success rate with much faster speed for cases without the k-security defense, and has a much better success rate at the same runtime for cases with k-security defense. The results offer an alternative and practical interpretation for k-security in split manufacturing. Lang Feng 0001, Jeyavijayan Rajendran, Jiang Hu 0001 |
ASP-DAC | 2 |
| 2019 | FastCFI: Real-Time Control Flow Integrity Using FPGA Without Code Instrumentation
Lang Feng 0001, Jeff Huang 0001, Jiang Hu 0001, Abhijith Reddy |
RV | 1 |
| 2018 | Exploring Serverless Computing for Neural Network TrainingabstractServerless or functions as a service runtimes have shown significant benefits to efficiency and cost for event-driven cloud applications. Although serverless runtimes are limited to applications requiring lightweight computation and memory, such as machine learning prediction and inference, they have shown improvements on these applications beyond other cloud runtimes. Training deep learning can be both compute and memory intensive. We investigate the use of serverless runtimes while leveraging data parallelism for large models, show the challenges and limitations due to the tightly coupled nature of such models, and propose modifications to the underlying runtime implementations that would mitigate them. For hyperparameter optimization of smaller deep learning models, we show that serverless runtimes can provide significant benefit. Lang Feng 0001, Prabhakar Kudva, Dilma Da Silva, Jiang Hu 0001 |
IEEE CLOUD | 1 |
| 2017 | Making split fabrication synergistically secure and manufacturableabstractSplit fabrication is a promising approach to security against attacks by untrusted foundries. While existing split fabrication methods consider the overhead of conventional objectives such as wirelength and timing, they mostly neglect manufacturability - an unavoidable challenge in nanometer technologies. Observing that security and manufacturability can be addressed in a synergistic manner, this work introduces routing techniques that can simultaneously improve both security and manufacturability in terms of either Chemical Mechanical Planarization (CMP) uniformity or Self-Aligned Double Patterning (SADP) compliance. The effectiveness of these techniques is confirmed by experiments on benchmark circuits. Lang Feng 0001, Jiang Hu 0001, Wai-Kei Mak, Jeyavijayan Rajendran |
ICCAD | 1 |
| 2017 | Making split fabrication synergistically secure and manufacturableabstractSplit fabrication is a promising approach to security against attacks by untrusted foundries. While existing split fabrication methods consider the overhead of conventional objectives such as wirelength and timing, they mostly neglect manufacturability - an unavoidable challenge in nanometer technologies. Observing that security and manufacturability can be addressed in a synergistic manner, this work introduces routing techniques that can simultaneously improve both security and manufacturability in terms of either Chemical Mechanical Planarization (CMP) uniformity or Self-Aligned Double Patterning (SADP) compliance. The effectiveness of these techniques is confirmed by experiments on benchmark circuits. Lang Feng 0001, Jiang Hu 0001, Wai-Kei Mak, Jeyavijayan Rajendran |
ICCAD | 1 |