VLDB 2026 Research / reviewers in the wild / expert
Zhifeng Lin
dblp:02/7753
· DBLP profile ↗
43ranked-venue papers
7as first author
35since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 32 · 5 first-author · 26 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 1 first-author · 7 since 2021Software engineering, systems software and programming languages · 4 · 2 first-author · 3 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | An Adaptive Cost-based Via and Congestion Co-optimization Framework for VLSI Global RoutingabstractGlobal routing is a critical stage in VLSI physical design, directly affecting the final Power, Performance, and Area (PPA) metrics. In this paper, we propose a high-performance global router that optimizes via count and routing congestion simultaneously. We first generate a 2D via-aware spine tree, which incorporates bend cost to reduce via usage while minimizing wire length. Then, a fast maze routing algorithm is employed to efficiently find a blockage-free path, followed by a congestion-aware layer assignment method to generate the 3D routing solution. Finally, we present an iterative rip-up and reroute strategy to resolve remaining congestion using the 3D bidirectional A* search. The A* search is guided by an adaptive cost function that dynamically adjusts via costs based on congestion, facilitating the co-optimization of congestion and via count. Compared to an advanced commercial tool and the leading academic engine OpenROAD, our algorithm achieves the best results in both overflow and via count, while preserving almost the same wire length. Zhaoyi Wu, Haishan Huang, Jianli Chen, Zhifeng Lin |
DATE | 4 |
| 2026 | Layer-Aware Timing-Driven Global Routing for Advanced Technology Nodes
Zhaoyi Wu, Haishan Huang, Benchao Zhu, Jianli Chen, Zhifeng Lin |
ISCAS | 5 |
| 2026 | A Co-optimization Framework for Multi-layer Design Rule ConstraintsabstractCompliance with design rule constraints constitutes a fundamental prerequisite for successful fabrication in advanced integrated circuit design. As foundries progressively introduce process-specific customization for better performance, new design rule challenges emerge across the device layers, such as implant layer constraints in the designs with multiple threshold voltages, and the trim poly layer constraints in the self-aligned double patterning (SADP). Conventional methodologies typically address such topological constraints in the legalization stage. In addition, filler insertion during chip finishing serves to improve manufacturability, such as a more uniform chip surface and a more robust power integrity. However, improper filler insertion might undermine the previous legalized layout. This work presents a co-optimization framework in the legalization and filler insertion stage, adaptable to multi-layer constraint scenarios. We model the filler insertion problem as a multi-branch tree and develop a dynamic programming-based pre-pruning algorithm, which is also able to detect violations in the legalization stage. To reduce runtime, two violation detectors are introduced for legalization, including a look-up table (LUT) inference method and a greedy scanning algorithm. These components are systematically integrated into a co-optimization framework, with configurable parameterization to ensure scalability across diverse constraints. Experimental results show that our algorithm can significantly reduce the number of violations compared with state-of-the-art work and the commercial tool. Guohao Chen 0001, Chang Liu 0131, Xingyu Tong 0001, Jianli Chen, Zhifeng Lin |
ACM Trans. Design Autom. Electr. Syst. | 6 |
| 2026 | G2: A customizable web-based framework for authoring interactive visualizationsabstractSpecifying visual encodings and interactions is exhausting but essential for authoring interactive visualizations. In this paper, we present G2, a customizable framework designed to support rapid generation of interactive visualizations with uniform specifications. G2 employs a data-driven grammar of graphics and defines a uniform set of interaction specifications. We discuss the design and implementation of G2 with rich examples, and a user interview to demonstrate its effectiveness. Since its first release in March 2016, G2 has undergone 360 iterations, received 12,000 stars, and supported over 22,600 related projects on Github. Bairui Su, Zhifeng Lin, Xiaojuan Liao, Zihan Zhou 0009, Minfeng Zhu 0001, Wei Chen 0001 |
Vis. Informatics | 3 |
| 2025 | Do Multimodal Language Models Really Understand Direction? A Benchmark for Compass Direction ReasoningabstractDirection reasoning is essential for intelligent systems to understand the real world. While existing work focuses primarily on spatial reasoning, compass direction reasoning remains underexplored. To address this, we propose the Compass Direction Reasoning (CDR) benchmark, designed to evaluate the direction reasoning capabilities of multimodal language models (MLMs). CDR includes three types images to test spatial (up, down, left, right) and compass (north, south, east, west) directions. Our evaluation reveals that most MLMs struggle with direction reasoning, often performing at random guessing levels. Experiments show that training directly with CDR data yields limited improvements, as it requires an understanding of real-world physical rules. We explore the impact of mixdata and CoT fine-tuning methods, which significantly enhance MLM performance in compass direction reasoning by incorporating diverse data and step-by-step reasoning, improving the model’s ability to understand direction relationships. Hang Yin 0007, Zhifeng Lin, Xin Liu 0132, Bin Sun 0004, Kan Li 0001 |
ICASSP | 2 |
| 2025 | Analytical Layer Assignment with Simulated Annealing RefinementabstractRouting is a critical and time-consuming stage in circuit physical design. The typical approach involves 2D routing followed by 3D layer assignment, with most state-of-the-art methods using sequential assignments, which limits the solution space due to the fixed order in which nets are processed. This paper proposes a two-stage layer assignment paradigm inspired by the placement process. First, we apply an analytical method to simultaneously assign layers for all nets, leveraging GPU acceleration to enhance computational efficiency. Then, a simulated annealing algorithm further optimizes the segment assignments. Experimental results show that, compared to state-of-the-art sequential and concurrent layer assignment algorithms, our method reduces via count by 16.9% and 1.5% in global routing and by 5.3% and 3.5% in detailed routing, respectively, with minimal wirelength increases. Additionally, our algorithm achieves the fewest DRC violations across all benchmarks. Zhijie Cai, Xiqiong Bai, Zhifeng Lin, Jianli Chen |
ISCAS | 5 |
| 2025 | Multi-Bit Flip-Flop Based Timing and Power Optimization under Advanced Technology NodesabstractMulti-bit flip-flops (MBFFs) are widely employed in modern digital design due to their reduced power and area consumption compared to single-bit flip-flops (SBFFs). In this paper, we present an MBFF-based framework that simultaneously optimizes the crucial power, area, and timing metrics. First, we present a mean shift-based clustering algorithm to generate power and area-friendly clusters while considering multiple clocks. Then, a feasible-region-based declustering method is developed to produce the desired timing solution. Finally, we propose a timing-aware refinement strategy to further improve the solution quality. Compared with the competitive works, the experimental results show that our proposed algorithm achieves the best performance within the shortest runtime. Tingxuan Gong, Wenxu Ruan, Dongwei Tan, Zhendong He, Zhifeng Lin, Jianli Chen |
ISCAS | 5 |
| 2025 | Legalization Framework with Design Rule Constraints Enhanced by Monte-Carlo-Based Cell Priority OptimizationabstractLegalization holds significant importance in VLSI physical design, as it significantly influences the manufacturability and reliability of circuits. Recently, advanced foundry nodes introduced complex constraints in standard-cell legalization, which makes legalization even harder. In this paper, we develop a legalization framework with design rule constraints enhanced by Monte-Carlo-Based cell priority optimization. We first handle abnormal density distribution to reduce the subsequent legalization’s hardness. Then, we propose an interval-assisted sequential legalization algorithm considering multiple design rule constraints with a Monte-Carlo-Based cell priority decision technique. Besides, based on the characteristics and complexity of different design rule constraints, we present a refinement phase to handle the remaining design rule constraints with corresponding detectors. Compared with a leading commercial tool, experiments on industrial benchmarks show that our legalization framework achieves 11% smaller average displacement, 15% smaller maximum displacement, 1.63× speedup, and 13% fewer remaining design rule violations on average. Benchao Zhu, Guohao Chen 0001, Zhifeng Lin, Jianli Chen |
ISCAS | 6 |
| 2025 | Two stage Ordered Escape Routing combined with LP and heuristic algorithm for large scaled PCB
Disi Lin, Chuandong Chen, Rongshan Wei, Qinghai Liu, Ziran Zhu, Zhifeng Lin, Jianli Chen |
Integr. | 7 |
| 2025 | O.O: Optimized one-die placement for face-to-face bonded 3D ICsabstractAs the miniaturization of integrated circuits (ICs) reaches its physical limits, the industry is entering a “more-than-Moore” era, demanding new Electronic Design Automation (EDA) tools. Existing TSV-based 3D placers focus on minimizing cuts while burgeoning F2F-bonded ICs feature dense interconnection between two planar die. Towards this novel structure, we proposed an integrated adaptation methodology upon mature one-die-based placement strategies. First, we instructively utilized a one-die placer to provide a statistical looking-ahead net diagnosis. The netlist henceforth shall be coarsened topologically and geometrically using a multi-level framework. Our multi-objective gain formulation guides a level-by-level refinement of the partition. This formulation considers factors like cut expectation, heterogeneous row heights, and balanced cell distribution, enabling efficient incremental calculations at each level. Given the partition, we synchronized the behavior of analytical planar placers by balancing the density and wirelength objective function among asymmetric layers. Finally, the result will be further improved by heuristic detail placement of bonding terminals and a post-place partition adjustment. Experimental results demonstrate that our fine-grained fusion of partitioning and placement techniques are competitive compared with the top three winners of the 2022 ICCAD CAD Contest, achieving the best normalized average wirelength with competitive runtime under various 3D architectural constraints . Xingyu Tong 0001, Yuhao Ren, Zhijie Cai, Yuan Wen, Zhifeng Lin, Jianli Chen |
Integr. | 7 |
| 2025 | An analytical placement algorithm with looking-ahead routing topology optimization
Xingyu Tong 0001, Zhijie Cai, Zhifeng Lin, Jianli Chen |
Integr. | 5 |
| 2025 | A Matching-Based Escape Routing Algorithm With Variable Design Rules and Multiple ConstraintsabstractEscape routing is a critical problem in PCB routing, and its quality dramatically affects the cost of the PCB design. Unlike the traditional escape routing that works mainly for the BGA with unique line width and space, this paper presents a high-performance escape routing algorithm to handle problems with variable design rules and multiple constraints. We first propose a novel obstacle-avoiding method to project pins to the boundary and construct a channel projection graph combined with a channel merging technique to handle complex irregular packages. We then construct a bi-projection graph and propose a matching-based hierarchical sequencing algorithm to consider manual constraints. We perform global routing for each pin/differential pair by congestion-avoiding path initializing and rip-up and reroute path optimizing. Finally, a length-aware detail routing algorithm is developed to optimize the line length while ensuring the differential pair constraints. The experimental results on industrial PCB instances show that our algorithm can achieve 100% routability without violating the design rules and constraints, while two state-of-the-art PCB routers, FreeRouting and Allegro, cannot complete escape routing. Chuandong Chen, Disi Lin, Qinghai Liu, Zhifeng Lin, Genggeng Liu, Jianli Chen, Yao-Wen Chang |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2025 | An analytical timing-driven placer for modern heterogeneous FPGAs
Zhifeng Lin, Yilu Chen, Yanyue Xie, Chuandong Chen, Jianli Chen |
J. Supercomput. | 1 |
| 2025 | Obstacle-Avoiding X-Architecture Bounded-Skew Tree Algorithm Under Timing Slack ConstraintsabstractAs interconnect delay increasingly becomes the primary source of chip delay, timing analysis in the very large-scale integration (VLSI) routing process is becoming more crucial. Concurrently, to maintain computational synchronization in the chip, the bounded-skew constraint must be introduced. Additionally, the issue of obstacle-avoiding has gained attention due to the presence of routing obstacles on the chip. Furthermore, the introduction of X-architecture enables more efficient utilization of routing resources. In this article, we propose an obstacle-avoiding X-architecture bounded-skew tree (BST) algorithm under timing slack constraints, which, for the first time, simultaneously considers timing slack, bounded-skew, obstacle-avoidance, and X-architecture in a unified framework. First, an effective preprocessing strategy is presented to support fast information retrieval for the subsequent strategies. Second, a BST construction strategy is developed to ensure compliance with skew constraints by consulting and updating a dedicated skew table. Third, a local worst negative slack (WNS) optimization strategy is designed to improve the WNS of critical paths by balancing wirelength (WL) and radius. Fourth, an obstacle-avoiding strategy is implemented to navigate around routing obstacles while minimizing unnecessary WL overhead. Finally, a path refinement strategy is designed to select routing structures with maximal edge sharing to replace the initial structure, thereby further optimizing WL. Experimental results demonstrate that the proposed algorithm significantly improves both WL and the key timing metric WNS, while satisfying obstacle-avoidance and bounded-skew constraints. Genggeng Liu, Ren Lu, Zhifeng Lin, Chuandong Chen, Min Gan, Jianli Chen, Wenzhong Guo |
IEEE Trans. Syst. Man Cybern. Syst. | 4 |
| 2024 | O.O: Optimized One-die Placement for Face-to-face Bonded 3D ICsabstractThe expansion of the IC dimension is ushering in a more-than-Moore era, necessitating corresponding EDA tools. Existing TSV-based 3D placers focus on minimizing cuts, while burgeoning F2F-bonded ICs features dense interconnection between two planar die. Towards this novel structure, we proposed an integrated adaptation methodology upon mature one-die-based placement strategies. First, we instructively utilized a one-die placer to provide a statistical looking-ahead net diagnosis. The netlist henceforth shall be coarsened topologically and geometrically with a multi-level framework. Level by level, the partition will be refined according to a multi-objective gain formulation, including cut expectation, heterogeneous row height, and balanced cell distribution. Given the partition, we synchronized the behavior of analytical planar placers by balancing the density and wirelength objective function among asymmetric layers. Finally, the result will be further improved by heuristic bonding terminals’ detail placement and a post-place partition adjustment. Compared to the top three winners of the 2022 CAD Contest at ICCAD, experiment results show that our fine-grained fusion upon partitioning and placement gets the best normalized average wirelength with a fairly reasonable runtime under all 3D architectural constraints. Xingyu Tong 0001, Zhijie Cai, Yuan Wen, Zhifeng Lin, Jianli Chen |
ASPDAC | 6 |
| 2024 | An Analytical Placement Algorithm with Routing topology OptimizationabstractPlacement is a critical step in the modern VLSI design flow, as it dramatically determines the performance of circuit designs. Most placement algorithms estimate the design performance with a half-perimeter wirelength (HPWL) and target it as their optimization objective. The wirelength model used by these algorithms limits their ability to optimize the internal routing topology, which can lead to discrepancies between estimates and the actual routing wirelength. This paper proposes an analytical placement algorithm to optimize the internal routing topology. We first introduce a differential wirelength model in the global placement stage based on an ideal routing topology RSMT. Through screening and tracing various segments, this model can generate meaningful gradients for interior points during gradient computation. Then, after global placement, we propose a cell refinement algorithm and further optimize the routing wirelength with swift density control. Experiments on ICCAD2015 benchmarks show that our algorithm can achieve a 3% improvement in routing wirelength, 0.8% in HPWL, and 23.8% in TNS compared with the state-of-the-art analytical placer. Xingyu Tong 0001, Zhijie Cai, Zhifeng Lin, Jianli Chen |
ASPDAC | 5 |
| 2024 | Effective Analytical Placement for Advanced Hybrid-Row-Height Circuit DesignsabstractRecently, hybrid-row-height designs have been introduced to achieve performance and area co-optimization in advanced nodes. Hybrid-row-height designs incur challenging issues to layout due to the heterogeneous cell and row structures. In this paper, we present an effective algorithm to address the hybrid-row-height placement problem in two major stages: (1) global placement, and (2) legalization. Inspired by the multi-channel processing method in convolutional neural networks (CNN), we use the feature extraction technique to equivalently transform the hybrid-row-height global placement problem into two sub-problems that can be solved effectively. We propose a multi-layer nonlinear framework with alignment guidance and a self-adaptive parameter adjustment scheme, which can obtain a high-quality solution to the hybrid-row-height global placement problem. In the legalization stage, we formulate the hybrid-row-height legalization problem into a convex quadratic programming (QP) problem, then apply the robust modulus-based matrix splitting iteration method (RMMSIM) to solve the QP efficiently. After RMMSIM-based global legalization, Tetris-like allocation is used to resolve remaining physical violations. Compared with the state-of-the-art work, experiments on the 2015 ISPD Contest benchmarks show that our algorithm can achieve 7%; shorter final total wirelength and $2.23 \times $ speedup. Yuan Wen, Benchao Zhu, Zhifeng Lin, Jianli Chen |
ASPDAC | 3 |
| 2024 | Late Breaking Results: Coulomb Force-Based Routability-Driven Placement Considering Global and Local CongestionabstractPlacement is a critical stage for VLSI routability optimization. A placement engine without considering the layout congestion might lead to poor solutions with routing failures. This paper introduces a Coulomb force-based global placement framework that addresses global and local routing congestions. We first present a routing path-based cell padding strategy for local congestion mitigation. Then, we construct a routability-aware placement model that utilizes virtual Coulomb forces to eliminate crucial global congestion. Compared with a leading academic placer, RePlAce, and the advanced commercial tool, Innovus, the experimental results on industrial benchmark suites show that our proposed algorithm achieves the best routability within the shortest runtime. Jihai Meng, Shaohong Weng, Zhijie Cai, Yilu Chen, Zhifeng Lin, Jianli Chen |
DAC | 5 |
| 2024 | Electrostatics-Based Analytical Global Placement for Timing OptimizationabstractPlacement is a critical stage for VLSI timing closure. A global placer without considering timing delay might lead to inferior solutions with timing violations. This paper proposes an electrostatics-based timing optimization method for VLSI global placement. Simulating the optimal buffering behavior, we first present an analytical delay model to calculate each connection delay accurately. Then, a timing-driven block distribution scheme is developed to optimize the critical path delay while considering the path-sharing effect. Finally, we develop a timing-aware precondition technique to speed up placement convergence without degrading timing quality. Experimental results on industrial benchmark suites show that our timing-driven placement algorithm outperforms a leading commercial tool by 6.7% worst negative slack (WNS) and 21.6% total negative slack (TNS). Zhifeng Lin, Yilu Chen, Jianli Chen, Yao-Wen Chang |
DATE | 1 |
| 2024 | Layout-level Hardware Trojan Prevention in the Context of Physical DesignabstractA growing recognition of potential vulnerabilities to layout-level Hardware Trojan (HT) attacks has spurred significant research efforts aimed at enhancing the resilience of ICs against such threats. However, traditional hardware security has been predominantly concerned with defensive measures, often overlooking the original key metrics in physical design evaluation: power, performance, and area (PPA). This study introduces an automated methodology incorporating HT considerations into the practical physical design process. Utilizing a Bayesian optimization framework, it effectively navigates the operation of commercial physical implementation tools in the solution space of hyper-parameter settings. Innovative strategies inspired by mosaic techniques, such as cell shifting and buffer insertion, realize additional improvements in layout-level trojan prevention. Comparative evaluations have shown that our approach outperforms leading entries from the ISPD 2023 Contest in terms of PPA and HT prevention metrics, thereby providing significant insights into the synergy between these critical factors. Xingyu Tong 0001, Guohao Chen 0001, Zhijie Cai, Zhifeng Lin, Jianli Chen |
ICCAD | 6 |
| 2024 | Global and Local Attention-Based Inception U-Net for Static IR Drop PredictionabstractStatic IR drop analysis is a fundamental and critical task in chip design since the IR drop will significantly affect the design's functionality, performance, and reliability. However, the process of IR drop analysis can be time-consuming, potentially taking several hours. Therefore, a fast and accurate IR drop prediction is paramount for reducing the overall time invested in chip design. In this paper, we propose a global and local attention-based Inception U-Net for static IR drop prediction. Our U-Net incorporates the Transformer, CBAM, and Inception architectures to enhance its feature capture capability at different scales and improve the accuracy of predicted IR drop. Moreover, we propose 4 new features, which enhance our model with richer information. Finally, to balance the sampling probabilities across different regions in one design, we propose a series of novel data spatial adjustment techniques, with each batch randomly selecting one of them during training. Experimental results demonstrate that our proposed algorithm can achieve the best results among the winning teams of the ICCAD 2023 contest and the state-of-the-art algorithms. Yilu Chen, Zhijie Cai, Zhifeng Lin, Jianli Chen |
ICCD | 4 |
| 2024 | Decoupling Heterogeneous Features for Robust 3D Interacting Hand Poses EstimationabstractEstimating the 3D poses of interacting hands from a monocular image is challenging due to the similarity in appearance between hand parts. Therefore, utilizing the appearance features alone tends to result in unreliable pose estimation. Existing approaches directly fuse the appearance features with position features, ignoring that the two types of features are heterogeneous. Here, the appearance features are derived from the RGB values of pixels, while the position features are mapped from the coordinates of pixels or joints. To address this problem, we present a novel framework called Decoupled Feature Learning (DFL ) for 3D pose estimation of interacting hands. By decoupling the appearance and position features, we facilitate the interactions within each feature type and those between both types of features. First, we compute the appearance relationships between the joint queries and the image feature maps; we utilize these relationships to aggregate each joint's appearance and position features. Second, we compute the 3D spatial relationships between hand joints using their position features; we utilize these relationships to guide the feature enhancement of joints. Third, we calculate appearance relationships and spatial relationships between the joints and image using the appearance and position features, respectively; we utilize these complementary relationships to promote the joints' location in the image. The two processes mentioned above are conducted iteratively. Finally, only the refined position features are used for hand pose estimation. This strategy avoids the step of mapping heterogeneous appearance features to hand-joint positions. Our method significantly outperforms state-of-the-art methods on the large-scale InterHand2.6M dataset. More impressively, our method exhibits strong generalization ability on in-the-wild images. Huan Yao, Changxing Ding, Xuanda Xu, Zhifeng Lin |
ACM Multimedia | 4 |
| 2024 | A fast and high-performance global router with enhanced congestion control
Xiqiong Bai, Yilu Chen, Zhifeng Lin, Zhijie Cai, Ziran Zhu, Jianli Chen |
Integr. | 3 |
| 2024 | High-correlation 3D routability estimation for congestion-guided global routing
Yilu Chen, Miaodi Su, Hongzhi Ding, Shaohong Weng, Zhifeng Lin, Xiqiong Bai |
J. Supercomput. | 5 |
| 2024 | AVA: An automated and AI-driven intelligent visual analytics frameworkabstractWith the incredible growth of the scale and complexity of datasets, creating proper visualizations for users becomes more and more challenging in large datasets. Though several visualization recommendation systems have been proposed, so far, the lack of practical engineering inputs is still a major concern regarding the usage of visualization recommendations in the industry. In this paper, we proposed AVA, an open-sourced web-based framework for Automated Visual Analytics. AVA contains both empiric-driven and insight-driven visualization recommendation methods to meet the demands of creating aesthetic visualizations and understanding expressible insights respectively. The code is available at https://github.com/antvis/AVA. Jiazhe Wang, Chenlu Li, Zeyu Wang 0005, Yuhui Gu, Xingui Lai, Xiaoqing Dong, Zhifeng Lin, Jiehui Zhou, Xingyu Liu 0003, Wei Chen 0001 |
Vis. Informatics | 11 |
| 2023 | Harmonious Feature Learning for Interactive Hand-Object Pose EstimationabstractJoint hand and object pose estimation from a single image is extremely challenging as serious occlusion often occurs when the hand and object interact. Existing approaches typically first extract coarse hand and object features from a single backbone, then further enhance them with reference to each other via interaction modules. However, these works usually ignore that the hand and object are competitive in feature learning, since the backbone takes both of them as foreground and they are usually mutually occluded. In this paper, we propose a novel Harmonious Feature Learning Network (HFL-Net). HFL-Net introduces a new framework that combines the advantages of single-and double-stream backbones: it shares the parameters of the low-and high-level convolutional layers of a common ResNet-50 model for the hand and object, leaving the middle-level layers unshared. This strategy enables the hand and the object to be extracted as the sole targets by the middle-level layers, avoiding their competition in feature learning. The shared high-level layers also force their features to be harmonious, thereby facilitating their mutual feature enhancement. In particular, we propose to enhance the feature of the hand via concatenation with the feature in the same location from the object stream. A subsequent self-attention layer is adopted to deeply fuse the concatenated feature. Experimental results show that our proposed approach consistently outperforms state-of-the-art methods on the popular HO3D and Dex-Ycb databases. Notably, the performance of our model on hand pose estimation even surpasses that of existing works that only perform the single-hand pose estimation task. Code is available at https://github.com/lzjJf12/HFL-Net. Zhifeng Lin, Changxing Ding, Huan Yao, Zengsheng Kuang, Shaoli Huang |
CVPR | 1 |
| 2023 | Toward Optimal Filler Cell Insertion with Complex Implant Layer ConstraintsabstractModern circuits often contain standard cells of different threshold voltages (multi-VTs) to achieve a better trade-off between timing and power consumption. Due to the heterogeneous cell structures, the multi-VTs cells impose various implant layer constraints, further complicating the already time-consuming filler cell insertion process. In this paper, we present a fast and near-optimal algorithm to solve the filler insertion problem with complex implant layer rules and minimum filler width constraints. We first propose an inference-driven detecting algorithm to identify each design rule violation accurately. Then, a dynamic-programming-based insertion method is developed to reduce the implant layer violations. Finally, we design a contour-driven violation refinement strategy to further improve manufacturability. Experimental results show that our algorithm can reduce the number of violations significantly compared with state-of-the-art works. Besides, with our identifier in the legalization stage, we can avoid conflicts in advance and solve almost all violations after filler insertion in industrial cases. Guohao Chen 0001, Zhifeng Lin, Jun Yu 0010, Jianli Chen |
DAC | 3 |
| 2023 | Incremental 3-D Global Routing Considering Cell Movement and Complex Routing ConstraintsabstractPlacement and routing are two critical problems in very large-scale integration physical design. However, there may be out-of-sync between the two problems considering congestion and wirelength. Therefore, it is desirable to design an efficient and highly coupled placement and routing engine to narrow the gap and minimize the mismatch between placement and routing. This article proposes an incremental 3-D global routing engine considering cell movement and complex routing constraints to relocate cells and reroute nets. We first apply a queue-based congestion-aware 3-D maze routing with routing height restriction to improve the initial routing solution. Efficient multinet-based location estimation is then presented to find the best location for each cell in multiple cell movement rounds. In each step of cell movement, we reroute nets for all candidate cell locations in parallel using a guided stack-based 3-D routing algorithm while considering the routing constraints. Finally, we adopt an edge-adjusting technique to improve the routed wirelength further. Compared with the champion of the 2020 CAD Contest at ICCAD (Hu et al., 2020) and the state-of-the-art works, experiment results based on the contest benchmarks show that our proposed algorithm achieves the best routing wirelength and competitive runtime without maximum cell movement constraint. Zhijie Cai, Zhifeng Lin, Chenyue Ma, Jun Yu 0010, Jianli Chen |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2022 | Learning Granularity-Unified Representations for Text-to-Image Person Re-identificationabstractText-to-image person re-identification (ReID) aims to search for pedestrian images of an interested identity via textual descriptions. It is challenging due to both rich intra-modal variations and significant inter-modal gaps. Existing works usually ignore the difference in feature granularity between the two modalities, i.e., the visual features are usually fine-grained while textual features are coarse, which is mainly responsible for the large inter-modal gaps. In this paper, we propose an end-to-end framework based on transformers to learn granularity-unified representations for both modalities, denoted as LGUR. LGUR framework contains two modules: a Dictionary-based Granularity Alignment (DGA) module and a Prototype-based Granularity Unification (PGU) module. In DGA, in order to align the granularities of two modalities, we introduce a Multi-modality Shared Dictionary (MSD) to reconstruct both visual and textual features. Besides, DGA has two important factors, i.e., the cross-modality guidance and the foreground-centric reconstruction, to facilitate the optimization of MSD. In PGU, we adopt a set of shared and learnable prototypes as the queries to extract diverse and semantically aligned features for both modalities in the granularity-unified feature space, which further promotes the ReID performance. Comprehensive experiments show that our LGUR consistently outperforms state-of-the-arts by large margins on both CUHK-PEDES and ICFG-PEDES datasets. Code will be released at https://github.com/ZhiyinShao-H/LGUR. Zhiyin Shao, Xinyu Zhang 0015, Zhifeng Lin, Jian Wang 0066, Changxing Ding |
ACM Multimedia | 4 |
| 2022 | Mixed-Cell-Height Placement With Complex Minimum-Implant-Area ConstraintsabstractMixed-cell-height standard cells are prevailingly used in advanced technologies to achieve better design tradeoffs among timing, power, and routability. As feature size decreases, the placement of cells with multiple threshold voltages may violate the complex minimum-implant-area (MIA) layer rule arising from the limitations of patterning technologies. Existing works consider the mixed-cell-height placement problem only during legalization or handle the MIA constraints during detailed placement. In this article, we address the mixed-cell-height placement problem with MIA constraints in two major stages: 1) post-global placement (Post-GP) and 2) MIA-aware legalization. In the Post-GP stage, we first present a continuous and differentiable cost function to address the Vdd/Vss alignment constraints and add weighted pseudonets to MIA-violation cells dynamically. Then, we propose a proximal optimization method based on the given global placement result to simultaneously consider Vdd/Vss alignment constraints, MIA constraints, cell distribution, cell displacement, and total wirelength. In the MIA-aware legalization stage, we develop a graph-based method to cluster cells of specific threshold voltages and apply a strip-packing-based binary linear programming to reshape cells. Then, we propose a matching-based technique to resolve intrarow MIA violations and reduce filler insertion. Furthermore, we formulate inter-row MIA-aware legalization as a quadratic programming problem, which is efficiently solved by a modulus-based matrix splitting iteration method. Finally, MIA-aware cell allocation and refinement are performed to further improve the result. Experimental results show that without any extra area overhead, our algorithm still can achieve 5.4% shorter final total wirelength than the state-of-the-art work. Jianli Chen, Zhifeng Lin, Yanyue Xie, Wenxing Zhu, Yao-Wen Chang |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2022 | An Incremental Placement Flow for Advanced FPGAs With Timing AwarenessabstractAs interconnects dominate circuit performance in modern field programmable gate arrays (FPGAs), placement becomes a crucial stage for timing closure. Traditional FPGA placers seldom consider the timing constraints and, thus, may lead to illegal routing solutions. In this article, we present an incremental timing-driven placement flow for advanced FPGAs. First, a timing-based global placement strategy is designed to guide heterogeneous blocks to desired locations with satisfied timing constraints. Then, a timing-aware packing algorithm is developed to mitigate the design complexity while improving the timing results. Finally, we propose a critical path-based optimization method to generate optimized layout without timing violations. We evaluate our algorithm based on industrial circuits using an advanced FPGA device. The experimental results show that our placer achieves a 5.1% improvement in worst slack and produce placements that require 16.7% less time to route when compared with the leading commercial tool Xilinx Vivado. Zhifeng Lin, Yanyue Xie, Sifei Wang, Jun Yu 0010, Jianli Chen |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2021 | Origami Inference: Private Inference Using Hardware EnclavesabstractThis work presents Origami, a framework which provides privacy-preserving inference for large deep neural network (DNN) models through a combination of enclave execution, cryptographic blinding, interspersed with accelerator-based computation. Origami partitions the ML model into multiple partitions. The first partition receives the encrypted user input within an SGX enclave. The enclave decrypts the input and then applies cryptographic blinding to the input data and the model parameters. The layer computation is offloaded to a GPU/CPU and the computed output is returned to the enclave, which decodes the computation on noisy data using the unblinding factors privately stored within SGX. This process may be repeated for each DNN layer, as has been done in prior work Slalom. However, the overhead of blinding and unblinding the data is a limiting factor to scalability. Origami relies on the empirical observation that the feature maps after the first several layers can not be used, even by a powerful conditional GAN adversary to reconstruct input. Hence, Origami dynamically switches to executing the rest of the DNN layers directly on an accelerator. We empirically demonstrate that using Origami, a conditional GAN adversary, even with an unlimited inference budget, cannot reconstruct the input. Compared to running the entire VGG-19 model within SGX, Origami inference improves the performance of private inference from 11x while using Slalom to 15. 1x. Krishna Narra, Zhifeng Lin, Yongqin Wang, Keshav Balasubramanian, Murali Annavaram |
CLOUD | 2 |
| 2021 | Late Breaking Results: Incremental 3D Global Routing Considering Cell MovementabstractPlacement and routing are two key problems in VLSI physical design. However, there may be out of sync between the two problems with congestion and routing resources. Therefore, it is desirable to design an efficient and highly coupled placement and routing engine. This paper proposes an incremental 3D global routing engine considering cell movement and complex routing constraints to relocate cells and reroute nets. We develop an efficient movement evaluation method to find desired locations and estimated routing resources for each cell. Then, we adopt an iterative approach to move cells to reduce routing resources. To reduce the time consumption of rerouting, we propose two technologies (searching space reduction and data structure optimization) to speed up the rerouting process. Compared with the participating teams at the 2020 CAD Contest at ICCAD based on the contest benchmarks, experiment results show that our proposed algorithm achieves the best runtime and routing resources while satisfying all the routing constraints. Zhifeng Lin, Chenyue Ma, Jun Yu 0010, Jianli Chen |
DAC | 2 |
| 2021 | Timing-Driven Placement for FPGAs with Heterogeneous Architectures and Clock ConstraintsabstractModern FPGAs often contain heterogeneous architectures and clocking resources which must be considered to achieve desired solutions. As the design complexity keeps growing, placement has become critical for FPGA timing closure. In this paper, we present an analytical placement algorithm for heterogeneous FPGAs to optimize its worst slack and clock constraints simultaneously. First, a heterogeneity-aware and memory-friendly delay model is developed to accurately and rapidly assess each connection delay. Then, a two-stage clock region refinement method is presented to effectively resolve the clock and resource violations. Finally, we develop a novel timing-based co-optimization method to generate optimized placement without any clocking violations. Compared with the state-of-the-art placer based on the advanced commercial tool Xilinx Vivado 2019.1 with the Xilinx 7 Series FPGA architecture, our algorithm achieves the best worst slack and routed wirelength while satisfying all clock constraints. Zhifeng Lin, Yanyue Xie, Gang Qian, Jianli Chen, Sifei Wang, Jun Yu 0010, Yao-Wen Chang |
DATE | 1 |
| 2021 | G6: A web-based library for graph visualizationabstractAuthoring graph visualization poses great challenges to developers due to its high requirements on both domain knowledge and development skills. Although existing libraries and tools reduce the difficulty of generating graph visualization, there are still many challenges. We work closely with developers and formulate several design goals, then design and implement G6, a web-based library for graph visualization. It combines template-based configuration for high usability and flexible customization for high expressiveness. To enhance development efficiency, G6 proposes a range of optimizations, including state management and interaction modes. We demonstrate its capabilities through an extensive gallery, a quantitative performance evaluation, and an expert interview. G6 was first released in 2017 and has been iterated for 317 versions. It has served as a web-based library for thousands of applications and received 8312 stars on GitHub. Zhanning Bai, Zhifeng Lin, Xiaoqing Dong, Yingchaojie Feng, Jiacheng Pan, Wei Chen 0001 |
Vis. Informatics | 3 |
| 2020 | Late Breaking Results: An Analytical Timing-Driven Placer for Heterogeneous FPGAs*abstractAs the feature sizes keep shrinking, interconnect delays have become a major limiting factor for FPGA timing closure. Traditional placement algorithms that address wirelength alone are no longer sufficient to close timing, especially for the large-scale heterogeneous FPGAs. In this paper, we resolve the crucial FPGA placement problem by optimizing wirelength and timing simultaneously. First, a smoothed routing-architecture-aware timing model is proposed to accurately estimate each interconnect delay. Then, a timing-driven delay look-up table is constructed to further speed up delay access. Finally, we present an effective wirelength and timing co-optimization strategy to produce high-quality placements without timing violations. Compared with Vivado 2019.1 on Xilinx benchmark suites for xc7k325t device, experimental results show that our algorithm achieves not only a 6.6% improvement in worst slack but also a 3.2% reduction for routed wirelength. Zhifeng Lin, Yanyue Xie, Gang Qian, Sifei Wang, Jun Yu 0010, Jianli Chen |
DAC | 1 |
| 2020 | Time-Division Multiplexing Based System-Level FPGA Routing for Logic VerificationabstractMulti-FPGA prototyping is widely used for modern VLSI verification, but the limited number of inter-FPGA connections in a multi-FPGA system may cause routing failures. As a result, the time-division multiplexing (TDM) technique is adopted to increase its resource utilization by transmitting multiple signals through the same routing channel. Due to the large signal delay between FPGA pairs, however, the performance of such a system greatly depends on the inter-FPGA routing quality. In this paper, we propose a TDM-based system-level routing algorithm to simultaneously minimize the maximum TDM (signal multiplexing) ratio and runtime, considering the crucial ratio constraints. By weighting the routing edges, we first model the net routing as a Steiner minimum tree (SMT) problem and solve it with an approximation algorithm with the performance bound 2(1 - 1/1), where l is the number of leaves in an optimal SMT. Then, a timing-driven assignment method is presented to evenly distribute the TDM ratio to routing signals, followed by a novel reassignment algorithm to efficiently handle unbalanced net groups. Finally, a ratio-aware refinement technique is employed to further improve the solution quality. Compared with the top-3 winners at the 2019 CAD Contest at ICCAD based on the contest benchmarks, experiment results show that our proposed algorithm achieves the best runtime and TDM ratio while satisfying all TDM constraints. Zhifeng Lin, Xiao Shi 0001, Jianli Chen, Jun Yu 0010, Yao-Wen Chang |
DAC | 2 |
| 2020 | Collage Inference: Using Coded Redundancy for Lowering Latency Variation in Distributed Image Classification SystemsabstractMLaaS (ML-as-a-Service) offerings by cloud computing platforms are becoming increasingly popular. Hosting pre-trained machine learning models in the cloud enables elastic scalability as the demand grows. But providing low latency and reducing the latency variance is a key requirement. Variance is harder to control in a cloud deployment due to uncertain-ties in resource allocations across many virtual instances. We propose the collage inference technique, which uses a novel convolutional neural network model, collage-cnn, to provide low-cost redundancy. A collage-cnn model takes a collage image formed by combining multiple images and performs multi-image classification in one shot, albeit at slightly lower accuracy. We augment a collection of traditional single image classifier models with a single collage-cnn classifier, which acts as their low-cost redundant backup. Collage-cnn provides backup classification results if any single image classification requests experience a slowdown. Deploying the collage-cnn models in the cloud, we demonstrate that the 99th percentile tail latency of inference can be reduced by 1.2x to 2x compared to replication-based approaches while providing high accuracy. Variation in inference latency can be reduced by 1.8x to 15x. Krishna Narra, Zhifeng Lin, Ganesh Ananthanarayanan, Amir Salman Avestimehr, Murali Annavaram |
ICDCS | 2 |
| 2020 | Clock-Aware Placement for Large-Scale Heterogeneous FPGAsabstractA modern field-programmable gate array (FPGA) often contains an ASIC-like clocking architecture which is crucial to achieve better skew and performance. Existing conventional FPGA placement algorithms seldom consider clocking resources, and thus may lead to clock routing failures. To address the special FPGA clocking architecture, this article presents an effective clock-aware placement algorithm for large-scale heterogeneous FPGAs. Our algorithm consists of four major technologies: 1) a combinatorial clock fence region method to effectively reduce the overuse of clocking resources; 2) a smoothed heterogeneous density function to lead heterogeneous blocks to desired sites and a coordinate transformation technique to facilitate CLB cell spreading; 3) a heterogeneous force modulation algorithm to stabilize placement movement and a hierarchical contraction technique to remedy an insufficiency of the multilevel placement framework; and 4) a two-level clock-aware packing and legalization scheme to generate an optimized, clocking-violation-free placement. We evaluate our results based on the ISPD 2017 Clock-Aware Placement Contest benchmark suite. Compared with the state-of-the-art placers, the experimental results show that our algorithm achieves the best-routed wirelength. Jianli Chen, Zhifeng Lin, Yun-Chih Kuo, Chau-Chin Huang, Yao-Wen Chang, Shih-Chun Chen, Chun-Han Chiang, Sy-Yen Kuo |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2019 | Slack squeeze coded computing for adaptive straggler mitigationabstractWhile performing distributed computations in today's cloud-based platforms, execution speed variations among compute nodes can significantly reduce the performance and create bottlenecks like stragglers. Coded computation techniques leverage coding theory to inject computational redundancy and mitigate stragglers in distributed computations. In this paper, we propose a dynamic workload distribution strategy for coded computation called Slack Squeeze Coded Computation (S2C2). S2C2 squeezes the compute slack (i.e., overhead) that is built into the coded computing frameworks by efficiently assigning work for all fast and slow nodes according to their speeds and without needing to re-distribute data. We implement an LSTM-based speed prediction algorithm to predict speeds of compute nodes. We evaluate S2C2 on linear algebraic algorithms, gradient descent, graph ranking, and graph filtering algorithms. We demonstrate 19% to 39% reduction in total computation latency using S2C2 compared to job replication and coded computation. We further show how S2C2 can be applied beyond matrix-vector multiplication. Krishna Narra, Zhifeng Lin, Mehrdad Kiamari, Amir Salman Avestimehr, Murali Annavaram |
SC | 2 |
| 2018 | Modular Routing Design for Chiplet-Based SystemsabstractSystem-on-Chip (SoC) complexity and the increasing costs of silicon motivate the breaking of an SoC into smaller "chiplets." A chiplet-based SoC design process has the promise to enable fast SoC construction by using advanced packaging technologies to tightly integrate multiple disparate chips (e.g., CPU, GPU, memory, FPGA). However, when assembling chiplets into a single SoC, correctness validation becomes a significant challenge. In particular, the network-on-chip (NoC) used within the individual chiplets and across chiplets to tie them together can easily have deadlocks, especially if each chip is designed in isolation. We introduce a simple, modular, yet elegant methodology for ensuring deadlock-free routing in multi-chiplet systems. As an example, we focus on future systems combining chiplets on an active silicon interposer. To maximize modularity, each individual chiplet is free to implement its own NoC topology and local routing algorithm, and the interposer can implement its own independent topology and routing. Our methodology imposes a few simple turn restrictions applied only to traffic as it flows into or out of the chiplets from the interposer, and we provide a way to determine these restrictions. The end result is an overall approach that enables highly-modular, chiplet-based SoC construction while eliminating deadlocks with high performance. Jieming Yin, Zhifeng Lin, Onur Kayiran, Matthew Poremba, Muhammad Shoaib Bin Altaf, Natalie D. Enright Jerger, Gabriel H. Loh |
ISCA | 2 |
| 2018 | GradiVeQ: Vector Quantization for Bandwidth-Efficient Gradient Aggregation in Distributed CNN TrainingabstractData parallelism can boost the training speed of convolutional neural networks (CNN), but could suffer from significant communication costs caused by gradient aggregation. To alleviate this problem, several scalar quantization techniques have been developed to compress the gradients. But these techniques could perform poorly when used together with decentralized aggregation protocols like ring all-reduce (RAR), mainly due to their inability to directly aggregate compressed gradients. In this paper, we empirically demonstrate the strong linear correlations between CNN gradients, and propose a gradient vector quantization technique, named GradiVeQ, to exploit these correlations through principal component analysis (PCA) for substantial gradient dimension reduction. GradiveQ enables direct aggregation of compressed gradients, hence allows us to build a distributed learning system that parallelizes GradiveQ gradient compression and RAR communications. Extensive experiments on popular CNNs demonstrate that applying GradiveQ slashes the wall-clock gradient aggregation time of the original RAR by more than 5x without noticeable accuracy loss, and reduce the end-to-end training time by almost 50%. The results also show that \GradiveQ is compatible with scalar quantization techniques such as QSGD (Quantized SGD), and achieves a much higher speed-up gain under the same compression ratio. Mingchao Yu, Zhifeng Lin, Krishna Narra, Youjie Li, Nam Sung Kim, Alexander G. Schwing, Murali Annavaram, Amir Salman Avestimehr |
NeurIPS | 2 |
| 2009 | Thresholded Range Aggregation in Sensor NetworksabstractThe recent advances in wireless sensor technologies (e.g., Mica, Telos motes) enable the economic deployment of lightweight sensors for capturing data from their surrounding environment, serving various monitoring tasks, like forest wildfire alarming and volcano activity. We propose a novel query called thresholded range aggregate query (TRA), which retrieves the IDs of the sensors for which the average measurement in their neighborhood exceeds a user-given threshold. This query provides results that they are robust against individual sensor abnormality, and yet precisely summarize the sensors' status in each local region. In order to process the (snapshot) TRA query, we develop energy-efficient protocols based on appropriate operators and filters in sensor nodes. The design of these operators and filters is non-trivial, due to the fact that each sensor measurement influences the actual results of other nodes in its neighborhood region. Furthermore, we extend our protocols for continuous evaluation of the TRA query. Experimental results show that our proposed solutions indeed offer substantial energy savings for both real and synthetic sensor networks. Zhifeng Lin, Man Lung Yiu, Nikos Mamoulis |
Mobile Data Management | 1 |