Qiang Zhou 0001

dblp:43/3182-1 · DBLP profile ↗
← Back
135ranked-venue papers
9as first author
31since 2021 · last 2025
0000-0003-3697-9348ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 84 · 2 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 26 · 4 first-author · 14 since 2021Applied, interdisciplinary, general and emerging computing · 16 · 1 first-author · 2 since 2021Artificial intelligence and machine learning · 15 · 2 first-author · 12 since 2021Software engineering, systems software and programming languages · 4 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 2
YearPublicationVenuePosition
2025 CR2PQ: Continuous Relative Rotary Positional Query for Dense Visual Representation Learning
abstract
Dense visual contrastive learning (DRL) shows promise for learning localized information in dense prediction tasks, but struggles with establishing pixel/patch correspondence across different views (cross-contrasting). Existing methods primarily rely on self-contrasting the same view with variations, limiting input variance and hindering downstream performance. This paper delves into the mechanisms of self-contrasting and cross-contrasting, identifying the crux of the issue: transforming discrete positional embeddings to continuous representations. To address the correspondence problem, we propose a Continuous Relative Rotary Positional Query ({\mname}), enabling patch-level representation learning. Our extensive experiments on standard datasets demonstrate state-of-the-art (SOTA) results. Compared to the previous SOTA method (PQCL), our approach achieves significant improvements on COCO: with 300 epochs of pretraining, {\mname} obtains \textbf{3.4\%} mAP$^{bb}$ and \textbf{2.1\%} mAP$^{mk}$ improvements for detection and segmentation tasks, respectively. Furthermore, {\mname} exhibits faster convergence, achieving \textbf{10.4\%} mAP$^{bb}$ and \textbf{7.9\%} mAP$^{mk}$ improvements over SOTA with just 40 epochs of pretraining.
Shaofeng Zhang, Qiang Zhou 0001, Sitong Wu, Haoru Tan, Zhibin Wang 0004, Jinfa Huang, Junchi Yan
ICLR2
2025 EasyOutPainter: One Step Image Outpainting With Both Continuous Multiple and Resolution
abstract
Image outpainting aims to generate the content of an input sub-image outside its boundaries, which remains open for existing generative models. This paper explores image outpainting in three directions that have not been achieved in literature to our knowledge: outpainting 1) with continuous multiples (in contrast to the discrete ones by existing methods); 2) with arbitrary resolutions; and 3) in a single step (for any multiples and resolutions). The arbitrary multiple outpainting is achieved by utilizing randomly cropped views from the same image during training to capture arbitrary relative positional information. Specifically, by feeding one view and relative positional embeddings as queries, we can reconstruct another view. At inference, we generate images with arbitrary expansion multiples by inputting an anchor image and its corresponding positional embeddings. The continuous-resolution outpainting is achieved by introducing the multi-scale training strategy into generative models. Specifically, by disentangling the image resolution and the number of patches, it can generate images with arbitrary resolutions without post-processing. Meanwhile, we propose a query-based contrastive objective to make our method not rely on a pre-trained backbone network which is otherwise often required in peer methods. The comprehensive experimental results on public benchmarks show its superior performance over state-of-the-art approaches.
Shaofeng Zhang, Qiang Zhou 0001, Zhibin Wang 0004, Hao Li 0030, Junchi Yan
IEEE Trans. Pattern Anal. Mach. Intell.2
2025 Few-Shot Semantic Segmentation on Remote Sensing Images With Learnable Prototype
abstract
Deep learning-based semantic segmentation has been the dominant solution to quickly capture regions of interest (ROIs) in remote sensing images. However, the annotation and training cost of a fully-supervised segmentation model is often too high due to the requirement for elaborate masks. Additionally, trained models are limited to recognize only those classes defined in the training set. This has led to increased interest in how to cheaply adapt learned knowledge to new unseen objects. In this paper, we propose a meta-learning-based few-shot method called Learnable Prototype Few-Shot Segmentation (LPFS) to quickly adapt models to previously unseen geographic categories with only a few support examples of remote sensing images. Specifically, we first build a learnable prototype module based on variational auto-encoder (VAE) to eliminate inter-class ambiguity and extract high-level semantic prototypes from the support set effectively. We then design a global-attention correlation map to achieve low-level structural feature alignment between the support and query images. Additionally, we introduce a base learner to alleviate the bias caused by the meta-learning network on base classes. The extensive experiments on the public few-shot segmentation benchmark iSAID-5idemonstrate that our method sets a new strong baseline for few-shot semantic segmentation on remote sensing images.
Jing Wang 0224, Yuang Liu, Qiang Zhou 0001, Zhibin Wang 0004, Fan Wang 0019
IEEE Trans. Geosci. Remote. Sens.3
2024 An Empirical Study and Analysis of Text-to-Image Generation Using Large Language Model-Powered Textual Representation
Zhiyu Tan, Mengping Yang, Luozheng Qin, Ye Qian, Qiang Zhou 0001, Cheng Zhang 0014, Hao Li 0030
ECCV (80)6
2024 DMT: Comprehensive Distillation with Multiple Self-Supervised Teachers
abstract
Numerous self-supervised learning paradigms, such as contrastive learning and masked image modeling, have been proposed to acquire powerful and general representations from unlabeled data. However, these models are commonly pretrained within their specific framework alone, failing to consider the complementary nature of visual representations. To tackle this issue, we introduce Comprehensive Distillation with Multiple Self-supervised Teachers (DMT) for pretrained model compression, which leverages the strengths of multiple off-the-shelf self-supervised models. Our experimental results on prominent benchmark datasets exhibit that the proposed method significantly surpasses state-of-the-art competitors while retaining favorable efficiency metrics. On classification tasks, our DMT framework utilizing three different self-supervised ViT-Base teachers enhances the performance of both small/tiny models and the base model itself. For dense tasks, DMT elevates the AP/mIoU of standard SSL models on MS-COCO and ADE20K datasets by 4.0%.
Yuang Liu, Jing Wang 0224, Qiang Zhou 0001, Fan Wang 0019, Jun Wang 0006, Wei Zhang 0056
ICASSP3
2024 Language-Guided Few-Shot Semantic Segmentation
abstract
Few-shot learning is a promising way for reducing the label cost in new categories adaptation with the guidance of a small, well labeled support set. But for few-shot semantic segmentation, the pixel-level annotations of support images are still expensive. In this paper, we propose an innovative solution to tackle the challenge of few-shot semantic segmentation using only language information, i.e.image-level text labels. Our approach involves a vision-language-driven mask distillation scheme, which contains a vision-language pretraining (VLP) model and a mask refiner, to generate high quality pseudo-semantic masks from text prompts. We additionally introduce a distributed prototype supervision method and complementary correlation matching module to guide the model in digging precise semantic relations among support and query images. The experiments on two benchmark datasets demonstrate that our method establishes a new baseline for language-guided few-shot semantic segmentation and achieves competitive results to recent vision-guided methods.
Jing Wang 0224, Yuang Liu, Qiang Zhou 0001, Fan Wang 0019
ICASSP3
2024 Continuous-Multiple Image Outpainting in One-Step via Positional Query and A Diffusion-based Approach
abstract
Image outpainting aims to generate the content of an input sub-image beyond its original boundaries. It is an important task in content generation yet remains an open problem for generative models. This paper pushes the technical frontier of image outpainting in two directions that have not been resolved in literature: 1) outpainting with arbitrary and continuous multiples (without restriction), and 2) outpainting in a single step (even for large expansion multiples). Moreover, we develop a method that does not depend on a pre-trained backbone network, which is in contrast commonly required by the previous SOTA outpainting methods. The arbitrary multiple outpainting is achieved by utilizing randomly cropped views from the same image during training to capture arbitrary relative positional information. Specifically, by feeding one view and positional embeddings as queries, we can reconstruct another view. At inference, we generate images with arbitrary expansion multiples by inputting an anchor image and its corresponding positional embeddings. The one-step outpainting ability here is particularly noteworthy in contrast to previous methods that need to be performed for $N$ times to obtain a final multiple which is $N$ times of its basic and fixed multiple. We evaluate the proposed approach (called PQDiff as we adopt a diffusion-based generator as our embodiment, under our proposed \textbf{P}ositional \textbf{Q}uery scheme) on public benchmarks, demonstrating its superior performance over state-of-the-art approaches. Specifically, PQDiff achieves state-of-the-art FID scores on the Scenery (\textbf{21.512}), Building Facades (\textbf{25.310}), and WikiArts (\textbf{36.212}) datasets. Furthermore, under the 2.25x, 5x and 11.7x outpainting settings, PQDiff only takes \textbf{40.6\%}, \textbf{20.3\%} and \textbf{10.2\%} of the time of the benchmark state-of-the-art (SOTA) method.
Shaofeng Zhang, Jinfa Huang, Qiang Zhou 0001, Zhibin Wang 0004, Fan Wang 0019, Jiebo Luo 0001, Junchi Yan
ICLR3
2024 Multimodal LLM Enhanced Cross-lingual Cross-modal Retrieval
abstract
Cross-lingual cross-modal retrieval (CCR) aims to retrieve visually relevant content based on non-English queries, without relying on human-labeled cross-modal data pairs during training. One popular approach involves utilizing machine translation (MT) to create pseudo-parallel data pairs, establishing correspondence between visual and non-English textual data. However, aligning their representations poses challenges due to the significant semantic gap between vision and text, as well as the lower quality of non-English representations caused by pre-trained encoders and data noise. To overcome these challenges, we propose LECCR, a novel solution that incorporates the multi-modal large language model (MLLM) to improve the alignment between visual and non-English representations. Specifically, we first employ MLLM to generate detailed visual content descriptions and aggregate them into multi-view semantic slots that encapsulate different semantics. Then, we take these semantic slots as internal features and leverage them to interact with the visual features. By doing so, we enhance the semantic information within the visual features, narrowing the semantic gap between modalities and generating local visual semantics for subsequent multi-level matching. Additionally, to further enhance the alignment between visual and non-English features, we introduce softened matching under English guidance. This approach provides more comprehensive and reliable inter-modal correspondences between visual and non-English features. Extensive experiments on four CCR benchmarks, i.e., Multi30K, MSCOCO, VATEX, and MSR-VTT-CN, demonstrate the effectiveness of our proposed method. Code: https://github.com/LiJiaBei-7/leccr.
Le Wang 0003, Qiang Zhou 0001, Zhibin Wang 0004, Hao Li 0030, Gang Hua 0001, Wei Tang 0016
ACM Multimedia3
2024 I2EBench: A Comprehensive Benchmark for Instruction-based Image Editing
abstract
Significant progress has been made in the field of Instruction-based Image Editing (IIE). However, evaluating these models poses a significant challenge. A crucial requirement in this field is the establishment of a comprehensive evaluation benchmark for accurately assessing editing results and providing valuable insights for its further development. In response to this need, we propose I2EBench, a comprehensive benchmark designed to automatically evaluate the quality of edited images produced by IIE models from multiple dimensions. I2EBench consists of 2,000+ images for editing, along with 4,000+ corresponding original and diverse instructions. It offers three distinctive characteristics: 1) Comprehensive Evaluation Dimensions: I2EBench comprises 16 evaluation dimensions that cover both high-level and low-level aspects, providing a comprehensive assessment of each IIE model. 2) Human Perception Alignment: To ensure the alignment of our benchmark with human perception, we conducted an extensive user study for each evaluation dimension. 3) Valuable Research Insights: By analyzing the advantages and disadvantages of existing IIE models across the 16 dimensions, we offer valuable research insights to guide future development in the field. We will open-source I2EBench, including all instructions, input images, human annotations, edited images from all evaluated methods, and a simple script for evaluating the results from new IIE models. The code, dataset, and generated images from all IIE models are provided in GitHub: https://github.com/cocoshe/I2EBench.
Jiayi Ji, Ke Ye, Weihuang Lin, Zhibin Wang 0004, Yonghan Zheng, Qiang Zhou 0001, Xiaoshuai Sun, Rongrong Ji
NeurIPS7
2024 Dynamic Token-Pass Transformers for Semantic Segmentation
abstract
Vision transformers (ViT) usually extract features via forwarding all the tokens in the self-attention layers from top to toe. In this paper, we introduce dynamic token-pass vision transformers (DoViT) for semantic segmentation, which can adaptively reduce the inference cost for images with different complexity. DoViT gradually stops partial easy tokens from self-attention calculation and keeps the hard tokens forwarding until meeting the stopping criteria. We employ lightweight auxiliary heads to make the token-pass decision and divide the tokens into keeping/stopping parts. With a token separate calculation, the self-attention layers are speeded up with sparse tokens and still work friendly with hardware. A token reconstruction module is built to collect and reset the grouped tokens to their original position in the sequence, which is necessary to predict correct semantic masks. We conduct extensive experiments on two common semantic segmentation tasks, and demonstrate that our method greatly reduces about 40% ∼ 60% FLOPs and the drop of mIoU is within 0.8% for various segmentation transformers. The throughput and inference speed of ViT-L/B are increased to more than 2× on Cityscapes. Code is available at https://github.com/FLHonker/DoViT-code.
Yuang Liu, Qiang Zhou 0001, Jing Wang 0224, Zhibin Wang 0004, Fan Wang 0019, Jun Wang 0006, Wei Zhang 0056
WACV2
2023 Point-Teaching: Weakly Semi-supervised Object Detection with Point Annotations
abstract
Point annotations are considerably more time-efficient than bounding box annotations. However, how to use cheap point annotations to boost the performance of semi-supervised object detection is still an open question. In this work, we present Point-Teaching, a weakly- and semi-supervised object detection framework to fully utilize the point annotations. Specifically, we propose a Hungarian-based point-matching method to generate pseudo labels for point-annotated images. We further propose multiple instance learning (MIL) approaches at the level of images and points to supervise the object detector with point annotations. Finally, we propose a simple data augmentation, named Point-Guided Copy-Paste, to reduce the impact of those unmatched points. Experiments demonstrate the effectiveness of our method on a few datasets and various data regimes. In particular, Point-Teaching outperforms the previous best method Group R-CNN by 3.1 AP with 5% fully labeled data and 2.3 AP with 30% fully labeled data on the MS COCO dataset. We believe that our proposed framework can largely lower the bar of learning accurate object detectors and pave the way for its broader applications. The code is available at https://github.com/YongtaoGe/Point-Teaching.
Yongtao Ge, Qiang Zhou 0001, Chunhua Shen, Zhibin Wang 0004, Hao Li 0030
AAAI2
2023 Static Probability Analysis Guided RTL Hardware Trojan Test Generation
abstract
Directed test generation is an effective method to detect potential hardware Trojan (HT) in RTL. While the existing works are able to activate hard-to-cover Trojans by covering security targets, the effectiveness and efficiency of identifying the targets to cover are ignored. We propose a static probability analysis method for identifying the hard-to-active data channel targets and generating the corresponding assertions for the HT test generation. Our method could generate test vectors to trigger Trojans from Trusthub, DeTrust, and OpenCores in 1 minute and get 104.33X time improvement on average compared with the existing method.
Haoyi Wang, Qiang Zhou 0001, Yici Cai
ASP-DAC2
2023 Foundation Model Drives Weakly Incremental Learning for Semantic Segmentation
abstract
Modern incremental learning for semantic segmentation methods usually learn new categories based on dense annotations. Although achieve promising results, pixel-by-pixel labeling is costly and time-consuming. Weakly incremental learning for semantic segmentation (WILSS) is a novel and attractive task, which aims at learning to segment new classes from cheap and widely available image-level labels. Despite the comparable results, the image-level labels can not provide details to locate each segment, which limits the performance of WILSS. This inspires us to think how to improve and effectively utilize the supervision of new classes given image-level labels while avoiding forgetting old ones. In this work, we propose a novel and data-efficient frame-work for WILSS, named FMWISS. Specifically, we propose pre-training based co-segmentation to distill the knowledge of complementary foundation models for generating dense pseudo labels. We further optimize the noisy pseudo masks with a teacher-student architecture, where a plug-in teacher is optimized with a proposed dense contrastive loss. Moreover, we introduce memory-based copy-paste augmentation to improve the catastrophic forgetting problem of old classes. Extensive experiments on Pascal VOC and COCO datasets demonstrate the superior performance of our framework, e.g., FMWISS achieves 70.7% and 73.3% in the 15–5 VOC setting, outperforming the state-of-the-art method by 3.4% and 6.1%, respectively.
Chaohui Yu, Qiang Zhou 0001, Jingliang Li, Jianlong Yuan, Zhibin Wang 0004, Fan Wang 0019
CVPR2
2023 D2Q-DETR: Decoupling and Dynamic Queries for Oriented Object Detection with Transformers
abstract
Despite the promising results, existing oriented object detection methods usually involve heuristically designed rules, e.g., RRoI generation, rotated NMS. In this paper, we propose an end-to-end framework for oriented object detection, which simplifies the model pipeline and obtains superior performance. Our framework is based on DETR, with the box regression head replaced with a points prediction head. The learning of points is more flexible, and the distribution of points can reflect the angle and size of the target rotated box. We further propose to decouple the query features into classification and regression features, which significantly improves the model precision. Aerial images usually contain thousands of instances. To better balance model precision and efficiency, we propose a novel dynamic query design, which reduces the number of object queries in stacked decoder layers without sacrificing model performance. Finally, we rethink the label assignment strategy of existing DETR-like detectors and propose an effective label re-assignment strategy for improved performance. We name our method D2Q-DETR. Experiments on the largest and challenging DOTA-v1.0 and DOTA-v1.5 datasets show that D2Q-DETR outperforms existing NMS-based and NMS-free oriented object detection methods and achieves the new state-of-the-art.
Qiang Zhou 0001, Chaohui Yu, Zhibin Wang 0004, Fan Wang 0019
ICASSP1
2023 LMSeg: Language-guided Multi-dataset Segmentation
Qiang Zhou 0001, Yuang Liu, Chaohui Yu, Jingliang Li, Zhibin Wang 0004, Fan Wang 0019
ICLR1
2023 Patch-level Contrastive Learning via Positional Query for Visual Pre-training
abstract
Dense contrastive learning (DCL) has been recently explored for learning localized information for dense prediction tasks (e.g., detection and segmentation). It still suffers the difficulty of mining pixels/patches correspondence between two views. A simple way is inputting the same view twice and aligning the pixel/patch representation. However, it would reduce the variance of inputs, and hurts the performance. We propose a plug-in method PQCL (Positional Query for patch-level Contrastive Learning), which allows performing patch-level contrasts between two views with exact patch correspondence. Besides, by using positional queries, PQCL increases the variance of inputs, to enhance training. We apply PQCL to popular transformer-based CL frameworks (DINO and iBOT, and evaluate them on classification, detection and segmentation tasks, where our method obtains stable improvements, especially for dense tasks. It achieves new state-of-the-art in most settings. Code is available at https://github.com/Sherrylone/Query_Contrastive.
Shaofeng Zhang, Qiang Zhou 0001, Zhibin Wang 0004, Fan Wang 0019, Junchi Yan
ICML2
2023 Points-to-3D: Bridging the Gap between Sparse Points and Shape-Controllable Text-to-3D Generation
abstract
Text-to-3D generation has recently garnered significant attention, fueled by 2D diffusion models trained on billions of image-text pairs. Existing methods primarily rely on score distillation to leverage the 2D diffusion priors to supervise the generation of 3D models, e.g., NeRF. However, score distillation is prone to suffer the view inconsistency problem, and implicit NeRF modeling can also lead to an arbitrary shape, thus leading to less realistic and uncontrollable 3D generation. In this work, we propose a flexible framework of Points-to-3D to bridge the gap between sparse yet freely available 3D points and realistic shape-controllable 3D generation by distilling the knowledge from both 2D and 3D diffusion models. The core idea of Points-to-3D is to introduce controllable sparse 3D points to guide the text-to-3D generation. Specifically, we use the sparse point cloud generated from the 3D diffusion model, Point-E, as the geometric prior, conditioned on a single reference image. To better utilize the sparse 3D points, we propose an efficient point cloud guidance loss to adaptively drive the NeRF's geometry to align with the shape of the sparse 3D points. In addition to controlling the geometry, we propose to optimize the NeRF for a more view-consistent appearance. To be specific, we perform score distillation to the publicly available 2D image diffusion model ControlNet, conditioned on text as well as depth map of the learned compact geometry. Qualitative and quantitative comparisons demonstrate that Points-to-3D improves view consistency and achieves good shape controllability for text-to-3D generation. Points-to-3D provides users with a new way to improve and control text-to-3D generation.
Chaohui Yu, Qiang Zhou 0001, Jingliang Li, Zhe Zhang 0049, Zhibin Wang 0004, Fan Wang 0019
ACM Multimedia2
2023 McPAT-Calib: A RISC-V BOOM Microarchitecture Power Modeling Framework
abstract
Power efficiency has become a nonneglected issue of modern CPUs. Therefore, accurate and robust power models are highly demanded in academia and industry. However, it is hard for existing power models to balance modeling speed, generality, and accuracy well. This article introduces McPAT-Calib, a microarchitecture power modeling framework, which combines McPAT with machine learning (ML) calibration and active learning (AL) sampling. McPAT-Calib can quickly and accurately estimate the power of different benchmarks executed on different CPU configurations, and provide an effective evaluation tool for the early design stage. First, McPAT-7nm is introduced to support the preliminary analytical power modeling for the 7-nm technology node. Then, a wide range of modeling features are identified, and automatic feature selection and advanced nonlinear regression are used to calibrate the McPAT-7nm modeling results, greatly improving the accuracy. Moreover, a novel AL approach termed power greedy sampling (PowerGS) embedded with domain knowledge is leveraged to reduce the modeling cost effectively. We use up to 15 configurations of the RISC-V Berkeley out-of-order machine (BOOM) along with 80 benchmarks, targeting 7-nm technology, to extensively evaluate McPAT-Calib. Compared with state-of-the-art (SOTA) microarchitecture power models, McPAT-Calib can reduce the mean absolute percentage error (MAPE) under different cross-validation (CV) strategies by 3.64%–6.14% (absolute reduction). Meanwhile, PowerGS is superior to the existing AL approaches, which can significantly reduce the demand for labeled samples to speed up model construction. The effectiveness of the overall modeling and estimation flow with AL sampling has also been verified.
Jianwang Zhai, Binwu Zhu, Yici Cai, Qiang Zhou 0001, Bei Yu 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.5
2023 Object Detection Made Simpler by Eliminating Heuristic NMS
abstract
It is valuable and promising to remove post-processing non-maximum suppression (NMS) for object detectors, making detectors simpler and purely end-to-end. Removing NMS is possible if the object detector can identify only one positive sample for prediction for each ground-truth object instance in an image. In this work, we propose a compact and plug-in head, named PSS head, which can be attached to any one-stage detectors to make them NMS-free. Specifically, the PSS head works by automatically selecting a positive sample for each instance to be detected, so that the detectors with our PSS head can directly remove NMS. The success of our PSS head lies in three aspects, namely one-to-one label assignment, stop-gradient operation for eliminating optimization conflicts, and the pss loss and ranking loss specifically designed for the PSS head. Experiments on the COCO dataset demonstrate the effectiveness of our method. In particular, when compared with stage-of-the-art NMS-free methods, our (attaching PSS head to VFNET) achieves 44.0% mAP, which exceeds the 41.5% mAP of DeFCN with a large margin. When taking Res2Net-101-DCN as backbone network, our achieves 50.3% mAP on the COCO test set, which is a promising performance even among NMS-based methods.
Qiang Zhou 0001, Chaohui Yu
IEEE Trans. Multim.1
2022 TransMarker: A Pure Vision Transformer for Facial Landmark Detection
abstract
Recent years, Convolution Neural Networks (CNNs) have achieved impressive results in facial landmark detection task. Especially, the u-shaped architecture, also known as U-net, has become the de-facto standard and achieved tremendous success. However, due to the locality property of convolution operation, it has a limitation in modeling global and long-range semantic information interaction, which is essential in localization tasks. In this work, we propose a Unet-like pure transformer method TransMarker, in which we give a new perspective to tackle facial landmark detection task in a sequence-to-sequence manner. We first split the input image into non-overlapping patches, which are seen as tokens in NLP tasks. Then, we feed the image patches into a symmetric u-shaped Encoder-Decoder architecture for local-global semantic feature learning. In addition, we introduce a Dense Skip-Connection schema to leverage the multi-level information within different resolutions. Note that, unlike conventional U-net architecture, we design the network with pure Transformer blocks, without any conventional operations. Extensive experiments demonstrate the state-of-the-art performance of our method on several standard datasets, i.e., WFLW, COFW and 300W, which remarkably outperform previous convolutional-based methods.
Wenyan Wu 0005, Yici Cai, Qiang Zhou 0001
ICPR3
2022 MimCo: Masked Image Modeling Pre-training with Contrastive Teacher
abstract
Recent masked image modeling (MIM) has received much attention in self-supervised learning (SSL), which requires the target model to recover the masked part of the input image. Although MIM-based pre-training methods achieve new state-of-the-art performance when transferred to many downstream tasks, the visualizations show that the learned representations are less separable, especially compared to those based on contrastive learning pre-training. This inspires us to think whether the linear separability of MIM pre-trained representation can be further improved, thereby improving the pre-training performance. Since MIM and contrastive learning tend to utilize different data augmentations and training strategies, combining these two pretext tasks is not trivial. In this work, we propose a novel and flexible pre-training framework, named MimCo, which combines MIM and contrastive learning through two-stage pre-training. Specifically, MimCo takes a pre-trained contrastive learning model as the teacher model and is pre-trained with two types of learning targets: patch-level and image-level reconstruction losses.
Qiang Zhou 0001, Chaohui Yu, Hao Luo 0004, Zhibin Wang 0004, Hao Li 0030
ACM Multimedia1
2022 Intelligent and kernelized placement: A survey
Yici Cai, Qiang Zhou 0001
Integr.3
2022 A survey on machine learning-based routing for VLSI physical design
Lin Li 0072, Yici Cai, Qiang Zhou 0001
Integr.3
2021 Instant-Teaching: An End-to-End Semi-Supervised Object Detection Framework
abstract
Supervised learning based object detection frameworks demand plenty of laborious manual annotations, which may not be practical in real applications. Semi-supervised object detection (SSOD) can effectively leverage unlabeled data to improve the model performance, which is of great significance for the application of object detection models. In this paper, we revisit SSOD and propose Instant-Teaching, a completely end-to-end and effective SSOD framework, which uses instant pseudo labeling with extended weak-strong data augmentations for teaching during each training iteration. To alleviate the confirmation bias problem and improve the quality of pseudo annotations, we further propose a co-rectify scheme based on Instant-Teaching, denoted as Instant-Teaching∗. Extensive experiments on both MS-COCO and PASCAL VOC datasets substantiate the superiority of our framework. Specifically, our method surpasses state-of-the-art methods by 4.2 mAP on MS-COCO when using 2% labeled data. Even with full supervised information of MS-COCO, the proposed method still outperforms state-of-the-art methods by about 1.0 mAP. On PASCAL VOC, we can achieve more than 5 mAP improvement by applying VOC07 as labeled data and VOC12 as unlabeled data.
Qiang Zhou 0001, Chaohui Yu, Zhibin Wang 0004, Qi Qian 0001, Hao Li 0030
CVPR1
2021 SRL: Separation-and-Recombination Learning for Video Facial Landmark Detection with Limited Data
abstract
Recent video facial landmark detection methods heavily rely on the supervised learning with large amount of annotated data. Nevertheless, the annotation of data on the video is very labor-intensive and time-consuming. Also, the supervised learning with massive parameters is easy to make the network suffer from overfitting and generalization-losing. In this work, we propose the Separation-and-Recombination Learning (SRL) framework to tackle this problem, in which the crucial idea is to adequately mine the inherent information of the limited labeled data in a semi-supervised manner. Specifically, we split the SRL framework into two stages, a separation stage and a recombination stage. Firstly, in the separation stage, we propose to train an Auto-Encoder network, disentangling-net, taking multi-frame temporal cues as input and with reconstruction and KL-divergence loss as constraints. In this stage, we successfully disentangle the face into two weak-coupling latent spaces, i.e., structure and appearance space. Then, in the recombination stage, with the trained disentangling-net, the limited labeled data can be greatly expended as pseudo paired data, with the recombination of structure and appearance code. Finally, we train a replaceable landmark detection network, predicting-net, with the supervision of both labeled and pseudo-labeled data. In the experiment, we demonstrate state-of-the-art performance on several well-known benchmarks, i.e., 300VW [56], blurred-300VW [60] and RWMB [60] dataset. Most importantly, our method is able to maintain impressive accuracy on extremely small training sets down to as few as 50% samples.
Wenyan Wu 0005, Yici Cai, Qiang Zhou 0001
FG3
2021 McPAT-Calib: A Microarchitecture Power Modeling Framework for Modern CPUs
abstract
Energy efficiency has become the core issue of modern CPUs, and it is difficult for existing power models to balance speed, generality, and accuracy. This paper introduces McPAT-Calib, a microarchitecture power modeling framework, which combines McPAT with machine learning (ML) calibration methods. McPAT-Calib can quickly and accurately estimate the power of different benchmarks running on different CPU configurations, and provide an effective evaluation tool for the design of modern CPUs. First, McPAT-7nm is introduced to support the analytical power modeling for the 7nm technology node. Then, a wide range of modeling features are identified, and automatic feature selection and advanced regression methods are used to calibrate the McPAT-7nm modeling results, which greatly improves the generality and accuracy. Moreover, a sampling algorithm based on active learning (AL) is leveraged to effectively reduce the labeling cost. We use up to 15 configurations of 7nm RISC-V Berkeley Out-of-Order Machine (BOOM) along with 80 benchmarks to extensively evaluate the proposed framework. Compared with state-of-the-art microarchitecture power models, McPAT-Calib can reduce the mean absolute percentage error (MAPE) of shuffle-split cross-validation by 5.95%. More importantly, the MAPE is reduced by 6.14% and 3.64% for the evaluations of unknown CPU configurations and benchmarks, respectively. The AL sampling algorithm can reduce the demand of labeled samples by 50 %, while the accuracy loss is only 0.44 %.
Jianwang Zhai, Binwu Zhu, Yici Cai, Qiang Zhou 0001, Bei Yu 0001
ICCAD5
2021 Multiple Features Driven Author Name Disambiguation
abstract
Author Name Disambiguation (AND) has received more attention recently, accompanied by the increase of academic publications. To tackle the AND problem, existing studies have proposed many approaches based on different types of information, such as raw document feature (e.g., co-author, title, and keywords), fusion feature (e.g., a hybrid publication embedding based on raw document feature), local structural information (e.g., a publication's neighborhood information on a graph), and global structural information (e.g., the interactive information between a node and others on a graph). However, there has been no work taking all the above-mentioned information into account for the AND problem so far. To fill the gap, we propose a novel framework namely MFAND (Multiple Features Driven Author Name Disambiguation). Specifically, we first employ the raw document and fusion feature to construct six similarity graphs for each author name to be disambiguated. Next, the global and local structural information extracted from these graphs is fed into a novel encoder called R3JG, which integrates and reconstructs the above-mentioned four types of information associated with an author, with the goal of learning the latent information to enhance the generalization ability of the MFAND. Then, the integrated and reconstructed information is fed into a binary classification model for disambiguation. Note that, several pruning strategies are applied before the information extraction to remove noise effectively. Finally, our proposed framework is investigated on two real-world datasets, and the experimental results show that MFAND performs better than all state-of-the-art methods.
Qiang Zhou 0001, Wei Chen 0070, Weiqing Wang 0001, Jiajie Xu 0001, Lei Zhao 0001
ICWS1
2021 An Efficient Approach for DRC Hotspot Prediction with Convolutional Neural Network
abstract
Predicting the design rule check (DRC) violation hotspots in an early stage plays an essential role in the efficiency of the physical design. Multiple factors that affect the performance of a DRC hotspot predictor, among them, the efficacy of the extracted features plays a substantial role. In this paper, we propose a connectivity-based DRC hotspot prediction method using a convolutional neural network. We show that the proposed method is efficient in both training and prediction. The relation between pin features and predictor performance is further investigated and two weighted connectivity-based route map features are introduced. Experimental results demonstrate that the proposed algorithm can predict on average 73% of the DRC hotspots with only 2.7% false alarms.
Lin Li 0072, Yici Cai, Qiang Zhou 0001
ISCAS3
2021 A Power Grids Electromigration Analysis with Via Array Using Current-Tracing Model
abstract
Electromigration (EM) has been considered to be a severe reliability issue in power grid networks of large integrated circuits (IC). The via array possesses special EM characteristics that have been observed to be distinct from a single via. In this study, a compact analytical model for the fast estimation of EM for via array was proposed by calculating the current distribution in the via arrays. The proposed model was then analyzed in a multi-layer power grid, which, for the first time, considered the impacts of the current propagation that exists in the vertical via array connected within the multi-level interconnection to improve the accuracy of the analytical model further. According to the model, a novel methodology for full- chip EM checking for multi-layered power grids was proposed. This method factored in the routing structure of the multi-layer power grid network, ensuring the EM assessment analysis's efficiency for large-scale power grid networks without sacrificing accuracy.
Jing Wang 0224, Yici Cai, Qiang Zhou 0001
ISCAS3
2021 A game theory approach for RTL security verification resources allocation
Haoyi Wang, Yici Cai, Qiang Zhou 0001
CCF Trans. High Perform. Comput.3
2021 Temperature-Aware Electromigration Analysis with Current-Tracking in Power Grid Networks
Jing Wang 0224, Yici Cai, Qiang Zhou 0001
J. Comput. Sci. Technol.3
2020 TransMoMo: Invariance-Driven Unsupervised Video Motion Retargeting
abstract
We present a lightweight video motion retargeting approach TransMoMo that is capable of transferring motion of a person in a source video realistically to another video of a target person. Without using any paired data for supervision, the proposed method can be trained in an unsupervised manner by exploiting invariance properties of three orthogonal factors of variation including motion, structure, and view-angle. Specifically, with loss functions carefully derived based on invariance, we train an auto-encoder to disentangle the latent representations of such factors given the source and target video clips. This allows us to selectively transfer motion extracted from the source video seamlessly to the target video in spite of structural and view-angle disparities between the source and the target. The relaxed assumption of paired data allows our method to be trained on a vast amount of videos needless of manual annotation of source-target pairing, leading to improved robustness against large structural variations and extreme motion in videos. We demonstrate the effectiveness of our method over the state-of-the-art methods. Code, model and data are publicly available on our project page (https://yzhq97.github.io/transmomo).
Zhuoqian Yang, Wentao Zhu 0004, Wayne Wu, Chen Qian 0006, Qiang Zhou 0001, Bolei Zhou, Chen Change Loy
CVPR5
2019 FAB: A Robust Facial Landmark Detection Framework for Motion-Blurred Videos
abstract
Recently, facial landmark detection algorithms have achieved remarkable performance on static images. However, these algorithms are neither accurate nor stable in motion-blurred videos. The missing of structure information makes it difficult for state-of-the-art facial landmark detection algorithms to yield good results. In this paper, we propose a framework named FAB that takes advantage of structure consistency in the temporal dimension for facial landmark detection in motion-blurred videos. A structure predictor is proposed to predict the missing face structural information temporally, which serves as a geometry prior. This allows our framework to work as a virtuous circle. On one hand, the geometry prior helps our structure-aware deblurring network generates high quality deblurred images which lead to better landmark detection results. On the other hand, better landmark detection results help structure predictor generate better geometry prior for the next frame. Moreover, it is a flexible video-based framework that can incorporate any static image-based methods to provide a performance boost on video datasets. Extensive experiments on Blurred-300VW, the proposed Realworld Motion Blur (RWMB) datasets and 300VW demonstrate the superior performance to the state-of-the-art methods. Datasets and models will be publicly available at https://keqiangsun.github.io/projects/FAB/FAB.html.
Keqiang Sun, Wayne Wu, Tinghao Liu, Shuo Yang 0003, Qiang Zhou 0001, Zuochang Ye, Chen Qian 0006
ICCV6
2019 Composite Optimization for Electromigration Reliability and Noise in Power Grid Networks
abstract
Electromigration(EM) and power supply noise has been considered serious reliability issue in the power grid networks. Several performance goals in EM reliability optimization and power supply noise optimization are typically conflict with each other. In this paper, we propose a composite optimization method trading off EM and power noise optimization process. In the method, we expand a temperature-aware EM model, which takes EM transient effect into account. Experimental results show that composite reliability optimization method can lengthen the lifetime of an entire circuit by approximately 10% compared with previous respective optimization strategy and no power noise violations exists after the composite optimization.
Jing Wang 0224, Yici Cai, Qiang Zhou 0001
ISCAS4
2019 Deep coupling neural network for robust facial landmark detection
Wenyan Wu 0005, Xingzhe Wu, Yici Cai, Qiang Zhou 0001
Comput. Graph.4
2019 A high-level information flow tracking method for detecting information leakage
abstract
In this paper, we note that the hardware Trojans that leak information through the unspecified output pins are difficult to detect by functional testing or side-channel signal analysis. Especially, the Trojans that leak the information through the side channel has proven stealthy to be detected. To solve this problem, we propose a feature matching method based on information flow tracking at high abstraction level. In this paper, the Trojans features are summarized with the format of high-level information flow tracking, which can be used to detect the Trojans. Experimental results show that our method can successfully identify the above-mentioned Trojans from Trust-hub, DeTrust, and OpenCores in less than 20 ms, showing significantly lower time complexity compared with the existing works.
Haoyi Wang, Chenguang Wang 0003, Yici Cai, Qiang Zhou 0001
Integr.4
2019 Parallelizing SAT-based de-camouflaging attacks by circuit partitioning and conflict avoiding
Qiang Zhou 0001, Yici Cai, Gang Qu 0001
Integr.2
2019 Toward a Formal and Quantitative Evaluation Framework for Circuit Obfuscation Methods
abstract
Since the first circuit obfuscation technique was proposed to thwart reverse engineering (RE) attacks to integrated circuits (ICs), there have been active research in de-obfuscation attacks and new obfuscation countermeasures. Although it is crucial for an obfuscation method to be secure against known de-obfuscation attacks, it is equally important to keep the cost of circuit obfuscation low. Most importantly, obfuscation methods need to be formally analyzed for their effectiveness and efficiency. In this paper, we propose a set of quantitatively evaluable metrics for this purpose, particularly facilitated by a recently proposed circuit partition attack (CPA) and the powerful SAT-based attack (SATA). Moreover, we find that CPA can be applied prior to any de-obfuscation attacks to reduce RE efforts exponentially. We then propose a new equivalent class guided obfuscation scheme (ECG-Obfus) to defeat CPA which leverages specially designed camouflaged cells to replace judiciously selected logic gates. Specifically, we select candidate gates for obfuscation from one certain equivalent class, in which the underlying equivalent relation is defined based on IC topological structure information. We evaluate ECG-Obfus using the proposed metrics and conduct experiments on ISCAS 85/89 standard benchmark suites and OpenSparc T1 microprocessor. The results show that ECG-Obfus gains good resilience against known de-obfuscation attacks (including CPA and SATA), with low design complexity and performance overhead.
Qiang Zhou 0001, Yici Cai, Gang Qu 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2019 Exploiting Spin-Orbit Torque Devices As Reconfigurable Logic for Circuit Obfuscation
abstract
Circuit obfuscation is a frequently used approach to conceal logic functionalities in order to prevent reverse engineering attacks on fabricated chips. Efficient obfuscation implementations are expected with lower design complexity and overhead but higher attack difficulties. In this paper, an emerging obfuscation approach is proposed by leveraging spin-orbit torque (SOT) devices-based look-up-tables as reconfigurable logic to replace the carefully selected gates. It is essentially impossible to identify the obfuscated gate with SOTs inside according to the physical geometry characteristics because the configured functionalities are represented by magnetization states. Such an obfuscation approach makes the circuit security further improved with high exponential attack complexities. Experiments on MCNC and ISCAS 85/89 benchmark suits show that the proposed approach could reduce the area overheads due to obfuscation by 10% averagely.
Jianlei Yang 0001, Qiang Zhou 0001, Zhaohao Wang, Hai Li 0001, Yiran Chen 0001, Weisheng Zhao 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2018 HLIFT: A high-level information flow tracking method for detecting hardware Trojans
abstract
In this paper, we note that the hardware Trojans that leak information through the unspecified output pins are difficult to detect by functional testing or side-channel signal analysis. To solve this problem, we propose a feature matching method based on information flow tracking at high abstraction level. Experimental results show that our method can successfully identify the above-mentioned Trojans from Trust-hub, DeTrust, and OpenCores in less than 20 ms, showing significantly lower time complexity compared with the existing works.
Chenguang Wang 0003, Yici Cai, Qiang Zhou 0001
ASP-DAC3
2018 ASAX: Automatic security assertion extraction for detecting Hardware Trojans
abstract
Hardware Trojans (HT) has been one of the major concerns of IC designers, and formal methods have been applied to the HT detection. In general, the assertions for detecting HT are manually defined, which is time-consuming and error-prone even for an expert engineer. However, there is a lack of studies on the automatic definition for security assertions. To fill in this gap, we propose an automatic security assertion extraction (ASAX) tool. ASAX labels the candidate signals and infers the proposed register transfer level (RTL) invariants from simulation traces. Next, the security assertions are mined from the inferred RTL invariants. By adopting a two-step invariants inferring technique, ASAX can extract high-coverage assertions with a low runtime. We validate the effectiveness and efficiency of ASAX through experiments on the benchmarks from Trust-hub, DeTrust and OpenCores. The results show that the HT can be 100% detected by model checking with the extracted security assertions.
Chenguang Wang 0003, Yici Cai, Qiang Zhou 0001, Haoyi Wang
ASP-DAC3
2018 A conflict-free approach for parallelizing SAT-based de-camouflaging attacks
abstract
As one of the most effective proactive countermeasures against reverse engineering, circuit camouflaging has emerged to be a hot research topic and it is becoming a mature technology with the development of various de-camouflaging attacks. Among them, the SAT-based method is the most powerful one to defeat circuit camouflaging. However, SAT-based attacks have scalability problem due to the complexity of the underlying SAT solvers, and straightforward approach to parallelize SAT-based attacks will fail. In this paper, we propose a two-level partition method (independent module partitioning and k-medoids clustering), together with a novel conflict avoidance strategy to solve the problem. Experimental results on OpenSparc T1 microprocessor controller demonstrate that our approach can on average reduce the scales of the SAT formulas by more than 50% and achieve 3.6× speedup on the best-known SAT-based de-camouflaging tool.
Qiang Zhou 0001, Yici Cai, Gang Qu 0001
ASP-DAC2
2018 Look at Boundary: A Boundary-Aware Face Alignment Algorithm
abstract
We present a novel boundary-aware face alignment algorithm by utilising boundary lines as the geometric structure of a human face to help facial landmark localisation. Unlike the conventional heatmap based method and regression based method, our approach derives face landmarks from boundary lines which remove the ambiguities in the landmark definition. Three questions are explored and answered by this work: 1. Why using boundary? 2. How to use boundary? 3. What is the relationship between boundary estimation and landmarks localisation? Our boundary-aware face alignment algorithm achieves 3.49% mean error on 300-W Fullset, which outperforms state-of-the-art methods by a large margin. Our method can also easily integrate information from other datasets. By utilising boundary information of 300-W dataset, our method achieves 3.92% mean error with 0.39% failure rate on COFW dataset, and 1.25% mean error on AFLW-Full dataset. Moreover, we propose a new dataset WFLW to unify training and testing across different factors, including poses, expressions, illuminations, makeups, occlusions, and blurriness. Dataset and model are publicly available at https://wywu.github.io/projects/LAB/LAB.html
Wayne Wu, Chen Qian 0006, Shuo Yang 0003, Yici Cai, Qiang Zhou 0001
CVPR6
2018 Electromigration Design Rule aware Global and Detailed Routing Algorithm
abstract
Electromigration (EM) in interconnects is becoming a major concern as the scaling of technology nodes. Electromigration affects chip performance and signal integrity seriously by generating shorts or opens, and then shortens the life-time of integrated circuits. In this paper, we propose an EM-aware routing algorithm in both global and detailed routing stages. Based on physics-based EM modeling and analysis, EM issue is modeled as physical design rule. In global routing stage, an efficient EM-aware Mazerouting algorithm is implemented. An concurrent EM-aware detailed router is then proposed based on multi-commodity flow method. Experimental results show that comparing with general routing algorithm, the proposed EM-aware algorithm could effectively reduce EM risk of signal wires by 92% with slight increasing of wire length and via count.
Xiaotao Jia, Jing Wang 0224, Yici Cai, Qiang Zhou 0001
ACM Great Lakes Symposium on VLSI4
2018 Electromagnetic equalizer: an active countermeasure against EM side-channel attack
abstract
Electromagnetic (EM) analysis is to reveal the secret information by analyzing the EM emission from a cryptographic device. EM analysis (EMA) attack is emerging as a serious threat to hardware security. It has been noted that the on-chip power grid (PG) has a security implication on EMA attack by affecting the fluctuations of supply current. However, there is little study on exploiting this intrinsic property as an active countermeasure against EMA. In this paper, we investigate the effect of PG on EM emission and propose an active countermeasure against EMA, i.e. EM Equalizer (EME). By adjusting the PG impedance, the current waveform can be flattened, equalizing the EM profile. Therefore, the correlation between secret data and EM emission is significantly reduced. As a first attempt to the co-optimization for power and EM security, we extend the EME method by fixing the vulnerability of power analysis. To verify the EME method, several cryptographic designs are implemented. The measurement to disclose (MTD) is improved by 1138x with area and power overheads of 0.62% and 1.36%, respectively.
Chenguang Wang 0003, Yici Cai, Haoyi Wang, Qiang Zhou 0001
ICCAD4
2018 An Efficient Technique to Reverse Engineer Minterm Protection Based Camouflaged Circuit
Ning Xu 0006, Qiang Zhou 0001
J. Comput. Sci. Technol.4
2018 Spear and Shield: Evolution of Integrated Circuit Camouflaging
Qiang Zhou 0001, Yici Cai, Gang Qu 0001
J. Comput. Sci. Technol.2
2018 A Multicommodity Flow-Based Detailed Router With Efficient Acceleration Techniques
abstract
Detailed routing is an important stage in very large scale integrated physical design. Due to the extreme scaling of transistor feature size and the complicated design rules, ensuring routing completion without design rule checking (DRC) violations becomes more and more difficult. Studies have shown that the low routing quality partly results from nonoptimal net-ordering nature of traditional sequential methods. The concurrent routing strategy is always based on an NP-hard model, thus is at a disadvantage in runtime. In this paper, we present a novel concurrent detailed routing algorithm that routes all nets simultaneously. Based on the multicommodity flow model, detailed routing problem with complex design rule constraints is formulated as an integer linear programming. Some model simplification heuristics and efficient model solving algorithms are proposed to improve the runtime. Experimental results show that, the proposed algorithms can reduce the DRC violations by 80%, meanwhile can reduce wirelength and via count by 5% and 8% compared with an industry tool. In addition, the proposed algorithm is general that it can be adopted as an incremental detailed router to refine a routing solution, so the number of DRC violations that industry tool cannot fix are further reduced by 27%.
Xiaotao Jia, Yici Cai, Qiang Zhou 0001, Bei Yu 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2017 An Empirical Study on Gate Camouflaging Methods Against Circuit Partition Attack
abstract
Gate camouflaging has emerged as a leading proactive countermeasure for reverse engineering (RE) attacks. However, a recently proposed circuit partition attack (CPA) can significantly reduce the complexity of revealing the original design from a camouflaged circuit. In this paper, we first conduct an empirical study on how CPA can facilitate the state-of-the-art de-camouflaging methods to perform more efficient attacks. We then study how an equivalent class guided camouflaging approach may thwart these de-camouflaging attempts and re-establish the defense against RE. Experimental results demonstrate that (1) CPA is an effective pre-processing technique to boost de-camouflaging methods, and (2) Equivalent class guided camouflaging technique is resilient against the union of CPA and existing de-camouflaging methods.
Qiang Zhou 0001, Yici Cai, Gang Qu 0001
ACM Great Lakes Symposium on VLSI2
2017 Automatic Security Property Generation for Detecting Information-Leaking Hardware Trojans
abstract
In recent years, formal methods have been adopted to detect the hardware Trojans (HT). However, they generally suffer from the time-consuming and error-prone development for property, lack of self-learning system to counter with the future HT types, and high computational complexity due to the growth of design scales. To overcome the above limitations, we propose an automatic security property generation method (ASPG) by feature analysis and property matching techniques. Machine learning is applied to systematically training the property library from the suspicious behaviors in unknown designs, which is expected to counter with the future HT. To reduce the computational complexity, we transform the register-transfer level (RTL) code into an introduced succinct abstract format to remove the redundant information which is unnecessary for depicting HT features. Experimental results show that the properties are generated in less than 50 ms with low memory consumption and the benchmarks from Trust-hub and DeTrust can be successfully detected with 0 false negatives and positives.
Chenguang Wang 0003, Yici Cai, Qiang Zhou 0001
ICCD3
2017 Power Profile Equalizer: A Lightweight Countermeasure against Side-Channel Attack
abstract
Power attack is an important side-channel attack (SCA) method based on the correlation between measured power profile and internal switching activities. Various techniques have been proposed to prevent power attack. It has been noted that the on-chip power grid (PG) has a vital effect on the effectiveness of power attack by inducing a noise in the power profile. However, there is a lack of study on this intrinsic effect of PG. In this paper, we explore the methods of exploiting the PG-induced noise to counter with power attack. We note that the PG-induced noise strongly depends on the PG impedance and it can be regulated by adjusting the PG capacitor to control the power profile to fixed values, which contributes to reducing the power leakage. Further, we propose a novel adjustment technique for PG capacitor, i.e. power profile equalizer (PPE), as a lightweight (low-overhead) countermeasure against power attack. PPE exploits the regulated noise to equalize the power profile without violating the layout and supply noise constraints. To reduce the overheads, random walk is adopted to utilize the utmost on-chip resources. Moreover, PPE is implemented by optimizing PG which is an essential IC component rather than producing new circuits. As a result, PPE incurs low overheads. Experimental results show that PPE is able to improve the measurements to disclose (MTD) by 1800x while the area and power increase respectively by 0.12% and 0.91%.
Chenguang Wang 0003, Yici Cai, Qiang Zhou 0001, Jianlei Yang 0001
ICCD4
2017 Cell spreading optimization for force-directed global placers
abstract
Wirelength is a traditional optimization objective in global placement algorithms. To eliminate cell overlaps, spreading forces need to be added to pull cells away from highly congested areas. At the same time, to optimize wirelength, the quadratic nature should be maintained. In this paper, several techniques are proposed to optimize spreading force orientation and modulation. Specifically, a percentage-driven method is proposed to cluster overfilled bins, followed by a center-uniformization algorithm to demarcate the expand region for the cluster. Finally, cells are distributed evenly within each expand region while maintaining relative cell positions and minimizing cell displacements. Experimental results show that the global placer that integrated with the proposed strategies achieves 13.0% and 2.1% less wirelength compared with Capo10.5 and Aplace3, respectively.
Yici Cai, Qiang Zhou 0001
ISCAS3
2016 MCFRoute 2.0: A Redundant Via Insertion Enhanced Concurrent Detailed Router
abstract
In modern VLSI design, manufacturing yield and chip performance are seriously affected by via failure. Redundant via insertion is an effective technique recommended by foundries to deal with the via failure. However, due to the extreme scaling of feature size, it is more and more difficult to resolve redundant via insertion (RVI) with limited routing resource while obeying complicated design rules. In this paper, we propose an RVI enhanced concurrent detailed router, MCFRoute 2.0, which effectively avoids design rule violations through a compact integer linear programming (ILP) model. The proposed router can not only route all nets simultaneously but also search for redundant via positions for all via simultaneously during routing stage. In addition, it proposes an RVI aware pin access allocation to further improve the routing performance. Experimental results show that our detailed router outperforms an industry EDA tool that it improves the redundant via insertion rate by 21%, while reducing design rule checking violation count, total wire length and via count by 47%, 4% and 14%, respectively.
Xiaotao Jia, Yici Cai, Qiang Zhou 0001, Bei Yu 0001
ACM Great Lakes Symposium on VLSI3
2016 Secure and Low-Overhead Circuit Obfuscation Technique with Multiplexers
abstract
Circuit obfuscation techniques have been proposed to conceal circuit's functionality in order to thwart reverse engineering (RE) attacks to integrated circuits (IC). We believe that a good obfuscation method should have low design complexity and low performance overhead, yet, causing high RE attack complexity. However, existing obfuscation techniques do not meet all these requirements. In this paper, we propose a polynomial obfuscation scheme which leverages special designed multiplexers (MUXs) to replace judiciously selected logic gates. Candidate to-be-obfuscated logic gates are selected based on a novel gate classification method which utilizes IC topological structure information. We show that this scheme is resilient to all the known attacks, hence it is secure. Experiments are conducted on ISCAS 85/89 and MCNC benchmark suites to evaluate the performance overhead due to obfuscation.
Xiaotao Jia, Qiang Zhou 0001, Yici Cai, Jianlei Yang 0001, Gang Qu 0001
ACM Great Lakes Symposium on VLSI3
2016 An efficient framework for configurable RO PUF
abstract
Physical Unclonable Function (PUF) is one of the most efficient technique to generate unique and random identification for chip authentication. Ring oscillator (RO) PUF takes advantage of delay variations of a pair of ROs, which is easy to implement on FPGAs. An important consideration for FPGA based RO PUF is how to eliminate systematic variation without reducing the number of output bits. To address this problem, we introduce high performance RO organization and comparison framework. Moreover, an enhanced configurable RO, which has up to 512 different configurations but only occupies one FPGA slice, is proposed to improve the reliability and output bits number. Experimental results demonstrate that our PUF achieves best value on bit-aliasing rate (50.37%) compared with other existing configurable RO PUFs. The output bits number also increases by the factors of 2.1-9.2.
Zhuwei Chen, Yici Cai, Qiang Zhou 0001, Gang Qu 0001
ISCAS3
2016 Is the Secure IC camouflaging really secure?
abstract
Circuit camouflaging techniques have been proposed to thwart reverse engineering (RE) attacks to integrated circuits (IC). In one of the most well-known camouflaging methods, selective XOR, NAND, and NOR gates are replaced by configurable logic units which have the same appearance to the RE attackers. It is argued that a successful attack has to brute force search all the camouflaged gates' possible {XOR, NAND, NOR} combinations, resulting in the attack complexity exponential to the number of camouflaged gates. In this paper, we have reported an attack to significantly reduce this complexity by partitioning the IC to many subcircuits to attack individually. We validate the power of the proposed circuit partition based attack on IS CA S benchmark suite and OpenSparc T1 microprocessor, and propose a potential countermeasure to re-secure IC camouflaging.
Qiang Zhou 0001, Yici Cai, Gang Qu 0001
ISCAS2
2016 Techniques for Design and Implementation of an FPGA-Specific Physical Unclonable Function
Jiliang Zhang 0002, Qiang Wu 0015, Yipeng Ding, Yongqiang Lyu 0001, Qiang Zhou 0001, Zhihua Xia, Xingming Sun, Xingwei Wang 0001
J. Comput. Sci. Technol.5
2015 Reliable and Anti-cloning PUFs Based on Configurable Ring Oscillators
abstract
Ring oscillator Physical Unclonable Function (RO PUF) is a popular silicon PUF due to its ease of implementation on both ASIC and FPGA. However, RO PUFs have severe reliability issues when the operating environment deviates from the nominal condition and security issues as cloning attacks have been reported. In this work, we propose to build configurable RO PUFs based on the notions of configurable RO PUF [6, 16] and highly flexible RO PUF [22] to address these concerns. First, we demonstrate how to build RO PUF from single flexible ROs, which improves both the reliability and hardware efficiency of RO PUFs. Then we propose a novel dual voltage based configurable RO PUF to mitigate the cloning attacks. Our experimental results show that our configurable RO PUFs are more reliable and hardware efficient than the existing RO PUF designs. Using the flexible RO PUF [22] as baseline, we have reduced the bit flip rate by 69% and improve the hardware utilization by 136%. In addition, the anti-cloning approach generates PUF data significantly different from the original PUF secret (average 47.5% Hamming distance) which makes potential cloning attacks very difficult.
Khai Lai, Jiliang Zhang 0002, Gang Qu 0001, Aijiao Cui, Qiang Zhou 0001
CAD/Graphics6
2015 SIAR: Customized real-time interactive router for analog circuits
Hailong Yao 0002, Yici Cai, Qiang Zhou 0001, Chiu-Wing Sham
Integr.4
2015 Register Clustering Methodology for Low Power Clock Tree Synthesis
Yici Cai, Qiang Zhou 0001
J. Comput. Sci. Technol.3
2015 Design-Rule-Aware Congestion Model with Explicit Modeling of Vias and Local Pin Access Paths
Zhongdong Qi, Yici Cai, Qiang Zhou 0001
J. Comput. Sci. Technol.3
2015 Obstacle-Avoiding and Slew-Constrained Clock Tree Synthesis With Efficient Buffer Insertion
abstract
As VLSI technology continuously scales down, buffered clock tree synthesis (CTS) has become increasingly critical in an attempt to generate a high-performance synchronous chip design. This paper presents a novel obstacle-avoiding CTS approach with slew constraints satisfied and signal polarity corrected. We build a look-up table through NGSPICE simulation to achieve accurate buffer delay and slew, which guarantees that the final skew after NGSPICE simulation is as satisfactory as expected. Aiming at skew optimization under constraints of slew and obstacles, our CTS approach features the clock tree construction stage with the obstacle-aware topology generation algorithm called OBB, balanced insertion of candidate buffer positions and a fast heuristic buffer insertion algorithm. With an overall view on obstacles to explore the global optimization space, our CTS approach effectively overcomes the negative influence on skew brought by the obstacles. Experimental results show the effectiveness of our CTS approach with significantly improved skew and latency by 69.0% and 72.0% on average. In addition, the accuracy of the look-up table is demonstrated through the huge skew reduction by 87.3% on average. Moreover, our OBB heuristic algorithm obtains 53.2% improvement in skew than the classic balanced bipartition algorithm.
Yici Cai, Qiang Zhou 0001, Hailong Yao 0002, Feifei Niu, Cliff C. N. Sze
IEEE Trans. Very Large Scale Integr. Syst.3
2015 A Selected Inversion Approach for Locality Driven Vectorless Power Grid Verification
abstract
Vectorless power grid verification is a practical approach for early stage safety check without input current patterns. The power grid is usually formulated as a linear system and requires intensive matrix inversion and numerous linear programming (LP), which is extremely time-consuming for large-scale power grid verification. In this paper, the power grid is represented in the manner of domain-decomposition approach, and we propose a selected inversion technique to reduce the computation cost of matrix inversion for vectorless verification. The locality existence among power grids is exploited to decide which blocks of matrix inversion should be computed while remaining blocks are not necessary. The vectorless verification could be purposefully performed by this manner of selected inversion, while previous direct approaches are required to perform full matrix inversion and then discard small entries to reduce the complexity of LP. Meanwhile, constraint locality is proposed to improve the verification accuracy. In addition, a concept of quasi-Poisson block is introduced to exploit grid locality among realistic power grids and a scheme of pad-aware partitioning is proposed to enable the selected inversion approach available for practical use. Experimental results show that the proposed approach could achieve significant speedups compared with previous approaches while still guaranteeing the quality of solution accuracy.
Jianlei Yang 0001, Yici Cai, Qiang Zhou 0001
IEEE Trans. Very Large Scale Integr. Syst.3
2014 VFGR: A very fast parallel global router with accurate congestion modeling
abstract
With the rapid growth of design size and complexity, global routing has always been a hard problem. Several new factors contribute to global routing congestion and can only be measured and optimized in 3-D global routing rather than 2-D routing. We propose an enhanced congestion model in global routing to capture local congestion and more accurately reflect modern design rule requirements. To achieve better global and detailed routing solution quality, we propose a 3-D global router VFGR with parallel computing using this congestion model. Experimental results show that VFGR can achieve comparable or better global routing solution quality with two start-of-the-art global routers in shorter runtime. It is also demonstrated that adopting proposed congestion model in global routing, higher solution quality and much shorter runtime can be achieved in detailed routing stage.
Zhongdong Qi, Yici Cai, Qiang Zhou 0001, Zhuoyuan Li 0003
ASP-DAC3
2014 Power supply noise aware evaluation framework for side channel attacks and countermeasures
abstract
Side Channel Attack (SCA) aims to extract the secret information from cryptography chips by analyzing the leakage of physical parameters. Power analysis based SCA is a popular approach to obtain secret keys by monitoring the power consumption of cryptography chips. However, most SCA evaluation methods are performed on FPGA platforms while many parasitic physical effects cannot be revealed before the cryptography chips are taped out. Roughly ignoring these effects will significantly increase the attack difficulties due to the corresponding measurement noise. Power supply noise has been observed to be critical for power analysis based SCA. This paper demonstrates a power supply noise aware evaluation framework for practical side channel attack from cryptography system design to physical design. On-chip power delivery network is implemented among physical design stage. Consequently the supply noise of power network can be explored according to the post-layout implementation. Additionally, the countermeasures of cryptography chips could be enhanced by on-chip decapacitors placement due to its influences on the characteristics of power delivery network.
Jianlei Yang 0001, Chenguang Wang 0003, Yici Cai, Qiang Zhou 0001
FPT4
2014 MCFRoute: a detailed router based on multi-commodity flow method
abstract
Detailed routing is an important stage in VLSI physical design. Due to the high routing complexity, it is difficult for existing routing methods to guarantee total completion without design rule checking violations (DRCs) and it generally takes several days for designers to fix remaining DRC-s. Studies has shown that the low routing quality partly results from non-optimal net-ordering nature of traditional sequential methods. In this paper, a novel concurrent detailed routing algorithm is presented that overcomes the net-order problem. Based on the multi-commodity flow (M-CF) method, detailed routing problem with complex design rule constraints is formulated as an integer linear programming (ILP) problem. Experiments show that the proposed algorithm is capable of reducing design rule violations while introducing no negative effects on wirelength and via count. Implemented as a detailed router following track assignment, the algorithm can reduce the DRCs by 38%, meantime, wirelength and via count are reduced by 3% and 2.7% respectively comparing with an industry tool. Additionally, the algorithm is adopted as an incremental detailed router to refine a routing solution, and experimental results show that the number of DRCs that industry tool can't fix are further reduce by half. Utilizing the independency between subregions, an efficient parallelization algorithm is implemented that can get a close to linear speedup.
Xiaotao Jia, Yici Cai, Qiang Zhou 0001, Zhuoyuan Li 0003, Zuowei Li
ICCAD3
2014 Accurate prediction of detailed routing congestion using supervised data learning
abstract
Routing congestion model is of great importance in design stages of modern physical synthesis, e.g. global routing and routability estimation during placement. As the technology node becomes smaller, routing congestion is more difficult to estimate during design stages ahead of detailed routing. In this paper, we propose a framework using nonparametric regression technique in machine learning to construct routing congestion model. The constructed model can capture multiple factors and enables direct prediction of detailed routing congestion with high accuracy. By using this model in global routing, significant reduction of design rule violations and detailed routing runtime can be achieved compared with the model in previous work, with small overhead in global routing runtime and memory usage.
Zhongdong Qi, Yici Cai, Qiang Zhou 0001
ICCD3
2014 A register clustering algorithm for low power clock tree synthesis
abstract
Clock networks dissipate a significant fraction of the entire chip power budget. Therefore, the optimization for power consumption of clock networks has become one of the most important objectives in high performance IC designs. In contrast to most of the traditional works that handle this problem with clock routing or buffer sizing, this paper proposes a novel register clustering algorithm in generating the leaf level topology of the clock tree to reduce the power consumption. Aiming to guarantee the stability of our register clustering algorithm, an effective initialization algorithm called “K-Splitting” and a “Pseudo Center” technology are developed. Meanwhile, a buffer allocation algorithm is proposed to satisfy the slew constraints within the clusters at a minimum cost of power consumption. We implement the clock tree synthesis (CTS) flow in [2] to test our approach on ISPD'10 benchmark circuits. Experimental results show that our register clustering algorithm achieves a 29.0% reduction in power consumption as well as a 5.7% reduction in max latency without affecting the clock skew. Moreover, the total runtime of the CTS flow with our register clustering algorithm is significantly reduced by 87.3%.
Yici Cai, Qiang Zhou 0001
ISCAS3
2014 Trusted Integrated Circuits: The Problem and Challenges
Yongqiang Lyu 0001, Qiang Zhou 0001, Yici Cai, Gang Qu 0001
J. Comput. Sci. Technol.2
2014 A Survey on Silicon PUFs and Recent Advances in Ring Oscillator PUFs
Jiliang Zhang 0002, Gang Qu 0001, Yongqiang Lyu 0001, Qiang Zhou 0001
J. Comput. Sci. Technol.4
2014 Friendly Fast Poisson Solver Preconditioning Technique for Power Grid Analysis
abstract
Robust and efficient algorithms for power grid analysis are crucial for both VLSI design and optimization. Due to the increasing size of power grids, IR drop analysis has become more computationally challenging both in runtime and memory consumption. This paper presents a Fast Poisson Solver (FPS) preconditioned method for unstructured power grids with unideal boundary conditions. Unstructured power grids are transformed to structured grids, which can be modeled as Poisson blocks by analytic formulation. The analytic formulation of transformed structured grids is adopted as an analytic preconditioner for original unstructured grids, in which the analytic preconditioner can be considered as a sparse approximate inverse technique. By combining this analytic preconditioner with robust conjugate gradient method, we demonstrate that this approach is totally robust for extremely large scale power grid simulations. Theoretical proof and experimental results show that iterations of our proposed method will hardly increase with the increasing of grid size as long as the pads density and the distribution range of metal conductance value have been decided. We demonstrate that the run efficiency of our approach is much higher than classical incomplete Cholesky factorization preconditioned conjugate gradient solver and random walk-based hybrid solver.
Jianlei Yang 0001, Yici Cai, Qiang Zhou 0001
IEEE Trans. Very Large Scale Integr. Syst.3
2014 PowerRush: An Efficient Simulator for Static Power Grid Analysis
abstract
Efficient power grid analysis is critical for modern very large scale integration design but is computationally challenging in runtime and memory consumption because of the increasing size of power grids. PowerRush is proposed as an efficient IR-drop simulator, which includes an efficient SPICE parser, a robust circuit builder, and a linear solver Algebraic MultiGrid Preconditioned Conjugate Gradient. The proposed AMG-PCG solver is a pure algebraic method, which can provide stable convergence without geometric information. Aggregation-based AMG with K-cycle acceleration is adopted as a preconditioner to improve the scalability of iterative method. In multigrid scheme, double pairwise aggregation technique is applied to matrix graph in coarsening to ensure low setup cost and memory requirement. Furthermore, K-cycle multigrid scheme is adopted to provide Krylov subspace acceleration at each level to guarantee enhanced robustness and scalability. The experimental results for large-scale power grids have shown that PowerRush has remarkable scalability both in runtime and memory consumption. DC analysis of power grid with 60-million nodes can be solved by PowerRush for 0.01 $mV$ accuracy within 150 s and 21.99 GB total memory used. Moreover, the proposed AMG-PCG solver can perform much better than widely used direct solver Cholmod and well-developed Hybrid solver both on runtime and memory consumption.
Jianlei Yang 0001, Zuowei Li, Yici Cai, Qiang Zhou 0001
IEEE Trans. Very Large Scale Integr. Syst.4
2013 Bridging the Gap between Global Routing and Detailed Routing: A Practical Congestion Model
abstract
To capture detailed routing congestion factors in sub-90nm technology nodes, we propose a practical congestion model embedded in 3-D global routing grid graph. Using a concept of pass-through capacity and demand, intra-gcell congestion contributed by fat vias, stacked vias, local nets and related design rules can be measured and optimized. Proposed congestion model is compatible with existing widely-used path search algorithms in global routing. Experimental results validate proposed model, and demonstrate that 42% less design rule violations and 46% shorter full-flow routing runtime, as well as 3% shorter wire length and 4% less via count in detailed routing results can be achieved using proposed congestion model in global routing stage.
Zhongdong Qi, Yici Cai, Qiang Zhou 0001
CAD/Graphics3
2013 Design and Implementation of a Delay-Based PUF for FPGA IP Protection
abstract
Physical Unclonable Function (PUF) makes use of the uncontrollable process variations during the production of IC to generate a unique signature for each IC. It has a wide application in security such as FPGA Intellectual Property (IP) protection, key generation and digital rights management. Ring Oscillator (RO) based PUF and Arbiter-based PUF are the most popular PUFs, but they are not specially designed for FPGA. RO-based PUF incurs high resource overhead while obtaining less challenge-response pairs, and requires ``hard macros'' to implement on FPGA. The arbiter-based PUF brings low resource overhead, but its structure is hard to be mapped on FPGA. Anderson'PUF can address these weaknesses of current Arbiter-based and RO-based PUFs. However, it cannot be directly implemented on the new generation FPGAs, and therefore it has the scalability issue. In order to address these problems, this paper presents a delay-based PUF using the intrinsic structure of FPGA (look-up table and multiplexer). The proposed delay-based PUF is completely realized on 28nm FPGAs. The experimental results show its high uniqueness and reliability. Moreover, we test the proposed PUF in the high temperature, and the results show its availability. Finally, the prospect of the proposed PUF in the FPGA IP protection is discussed.
Jiliang Zhang 0002, Qiang Wu 0015, Yongqiang Lyu 0001, Qiang Zhou 0001, Yici Cai, Yaping Lin, Gang Qu 0001
CAD/Graphics4
2013 Design and implementation of a group-based RO PUF
abstract
The silicon physical unclonable functions (PUF) utilize the uncontrollable variations during integrated circuit (IC) fabrication process to facilitate security related applications such as IC authentication. In this paper, we describe a new framework to generate secure PUF secret from ring oscillator (RO) PUF with improved hardware efficiency. Our work is based on the recently proposed group-based RO PUF with the following novel concepts: an entropy distiller to filter the systematic variation; a simplified grouping algorithm to partition the ROs into groups; a new syndrome coding scheme to facilitate error correction; and an entropy packing method to enhance coding efficiency and security. Using RO PUF dataset available in the public domain, we demonstrate these concepts can create PUF secret that can pass the NIST randomness and stability tests. Compared to other state-of-the-art RO PUF design, our approach can generate an average of 72% more PUF secret with the same amount of hardware.
Chi-En Daniel Yin, Gang Qu 0001, Qiang Zhou 0001
DATE3
2013 Binding Hardware IPs to Specific FPGA Device via Inter-twining the PUF Response with the FSM of Sequential Circuits
abstract
The continuous growth in both capability and capacity for FPGA now requires significant resources invested in the hardware design, which results in two classes of main security issues: 1) the unauthorized use and piracy attacks including cloning, reverse engineering, tampering etc. 2) the licensing issue. Binding hardware IPs (HW-IPs) to specific FPGA devices can efficiently resolve these problems. However, previous binding techniques are all based on encryption and hence have three main drawbacks: 1) encryption-based proposals in commercial are limited to protect the single large FPGA configuration, 2) many encryption-based proposals depend on a trusted third party to involve the licensing protocol, and 3) the encryption-based binding methods use costly mechanisms such as secure ROM or flash memory to store FPGA specific cryptographic keys, which is not only expensive but also vulnerable to side-channel attacks, and the management and transport of secret keys became a practical issue. In this work, we propose a PUF-FSM binding technique completely different from the traditional encryption-based methods to address these shortcomings.
Jiliang Zhang 0002, Yaping Lin, Yongqiang Lyu 0001, Ray C. C. Cheung, Wenjie Che, Qiang Zhou 0001, Jinian Bian
FCCM6
2013 FPGA IP protection by binding Finite State Machine to Physical Unclonable Function
abstract
In this paper we propose a novel binding mechanism that can protect FPGA IP from being cloned, tampered, or misused; and facilitate the pay-per-use licensing to limit the FPGA IP's execution to specific FPGA devices only. In this mechanism, the FPGA vendors will provide each enrolled device with a Physical Unclonable Function (PUF) that can be deployed securely during fabrication process. The core vendor will embed an augmented Finite State Machine (FSM) into the original FSM structure of the hardware IP (HW-IP) to react on the PUF response to a given challenge. The proposed binding method does not need any Trusted Third Party (TTP) or block cipher for key management and exchange. We analyze several known attacks to hardware IP and show that our method is secure against these attacks. Experimental results on MCNC benchmarks show that the proposed method incurs small design overhead in terms of area, power and delay.
Jiliang Zhang 0002, Yaping Lin, Yongqiang Lyu 0001, Gang Qu 0001, Ray C. C. Cheung, Wenjie Che, Qiang Zhou 0001, Jinian Bian
FPL7
2013 Selected inversion for vectorless power grid verification by exploiting locality
abstract
Vectorless power grid verification is a practical approach for early stage safety check without input current patterns. The power grid is usually formulated as a linear system and requires intensive matrix inversion and numerous linear programming, which is extremely time-consuming for large scale power grid verification. In this paper, the power grid is represented in the manner of domain-decomposition approach, and we propose a selected inversion technique to reduce the computation cost of matrix inversion for vectorless verification. The locality existence among power grids is exploited to decide which blocks of matrix inversion should be computed while remaining blocks are not necessary. The vectorless verification could be purposefully performed by this manner of selected inversion while previous direct approaches are required to perform full matrix inversion and then discard small entries to reduce the complexity of linear programming. Meanwhile, constraint locality is proposed to improve the verification accuracy. Experimental results show that the proposed approach could achieve significant speedups compared to previous approaches while still guaranteeing the quality of solution accuracy.
Jianlei Yang 0001, Yici Cai, Qiang Zhou 0001
ICCD3
2013 Thermal-aware P/G TSV planning for IR drop reduction in 3D ICs
Zuowei Li, Yuchun Ma, Qiang Zhou 0001, Yici Cai, Yuan Xie 0001
Integr.3
2012 Thermal-aware power network design for IR drop reduction in 3D ICs
abstract
Due to the high integration on vertical stacked layers, power/ground network design becomes one of the critical challenges in 3D IC design. With the leakage-thermal dependency, the increasing on-chip temperature in 3D designs has serious impact on IR drop due to the increased wire resistance and increased leakage current. Power/ground (P/G) TSVs can help to relieve the IR drop violation by vertically connecting the on-chip P/G networks on different layers. However, most previous work only fulfills a margin of the full potential of PG TSVs planning since the P/G grids are restricted in a uniform topology. Besides, the overlook of resistance variation and leakage current will make the results less accurate. In this paper, we present an efficient thermal-aware P/G TSVs planning algorithm based on a sensitivity model with temperature-dependent leakage current considered. The proposed method can overcome the limitation of uniform P/G grid topology and make full use of P/G TSVs planning for the optimization of P/G network by allowing short wires to connect the P/G TSVs to P/G grids in non-uniform topology. Moreover, with resistance variation and increased leakage current caused by high temperature in 3D ICs, more accurate result can be obtained. Both the theoretical analysis and experimental results show the efficiency of our approach. Results show that neglecting thermal impacts on power delivery can underestimate IR drop by about 11%. To relieve the severe IR drop violation, 51.8% more P/G TSVs are needed than the cases without thermal impacts considered. Results also show that our P/G TSV planning based on the sensitivity model can reduce max IR drop by 42.3% and reduce the number of violated nodes by 82.4%.
Zuowei Li, Yuchun Ma, Qiang Zhou 0001, Yici Cai, Yu Wang 0002, Yuan Xie 0001
ASP-DAC3
2012 PowerRush : Efficient transient simulation for power grid analysis
abstract
Transient analysis is the most practical and effective approach for power grid validation, but which is very challengeable for large scale VLSI chips because it is really time consuming and requires large memory resources. In this paper we proposed a parallel transient simulation approach for efficient power grid analysis. Firstly we adopt symmetric formulation for NA equation of RLC power grid to reduce memory usage. Meanwhile, fast Cholesky factorization solver can be used to improve simulation efficiency. Secondly, we perform partition-based parallel transient simulation for naturally independent subnets without accuracy lost. Thirdly, we propose a composite simulation flow for efficient and practical transient analysis for industrial power grid. Finally, several industrial power grid benchmarks are evaluated on our approaches for high accurate transient simulation with extremely low memory consumption.
Jianlei Yang 0001, Zuowei Li, Yici Cai, Qiang Zhou 0001
ICCAD4
2011 A fast recursive detailed routing algorithm for hierarchical FPGAs
abstract
Traditional sequence based routing algorithms for FPGAs usually route only one net at a time, so as to simplify the routing problems. However, with the number of logic blocks in the FPGAs becomes larger and larger, the time need to route each net can increase significantly. A new recursive detailed routing algorithm is proposed to address this problem. As decided by its recursive nature, this algorithm can only be applied for hierarchical FPGAs, which of the architectural features with its connection patterns is also presented in detail in this paper. The overall algorithm begins its routing from the topmost cluster and continues to route for each cluster from top down recursively, where the routing clusters map to the architectural cluster exactly. At each cluster level, a new heuristic is proposed to solve the specific routing problem. The scale of the problem is so small that the heuristic can be considered deterministic and quickly to solve. The proposed algorithm also takes advantages of the architectural features such as the connection patterns of switch box. As a result, the proposed algorithm is very fast in runtime due to all these facts. The experimental results show that detailed routing for a very large circuit can be done very quickly in just a few seconds.
Jinian Bian, Qiang Zhou 0001, Yici Cai
CSCWD3
2011 Obstacle-avoiding and slew-constrained buffered clock tree synthesis for skew optimization
abstract
Buered clock tree synthesis (CTS) is increasingly critical as VLSI technology continually scales down. Many researches have been done on this topic due to its key role in CTS, but current approaches either lack the obstacle-avoiding functionality or lead to large clock latency and/or skew. This paper presents a new obstacle-avoiding CTS approach with separate clock tree construction and buer insertion stages based on an integral view to explore the global optimization space. Aiming at skew optimization under constraints of slew and obstacles, our CTS approach features the clock tree construction stage with the obstacle-aware topology generation algorithm called OBB, balanced insertion of candidate buer positions, and a fast heuristic buer insertion algorithm. Experimental results show the eectiveness of our CTS approach with significantly improved skew and latency than [6] by 46% and 63% on average, and 15.3% reduction in skew than [5]. Our OBB heuristic obtains 36% improvement in skew than the classic balanced bipartition algorithm (BB) in [10].
Feifei Niu, Qiang Zhou 0001, Hailong Yao 0002, Yici Cai, Jianlei Yang 0001, Cliff C. N. Sze
ACM Great Lakes Symposium on VLSI2
2011 SIAR: splitting-graph-based interactive analog router
abstract
As analog and mixed-signal (AMS) circuitry gains increasing portions in modern SoCs, automotive analog routing is becoming more and more important. This paper presents a fast real-time interactive analog router called SIAR based on a splitting graph. A key feature is that SIAR allows real-time interactions between the router and the designer. The designer can try different guiding points by moving the cursor in the user window and the router will show the corresponding routing solutions in real-time for the designer to select the most satisfactory one. To enable real-time interactions, we present a new splitting graph to represent the routing area, which greatly enhances the routing efficiency. Different design rules such as variable wire and via width/spacing are supported by the router. Moreover, SIAR supports different routing modes such as point-to-point, point-to-module and module-to-module. Experimental results show that SIAR obtains promising routing efficiency with upto 28.6x speedup and better routing solutions compared with the commercial router Laker as well as upto 108x speedup compared with a modified implication-graph-based gridless routing approach [13].
Hailong Yao 0002, Qiang Zhou 0001, Yici Cai
ACM Great Lakes Symposium on VLSI3
2011 Fast poisson solver preconditioned method for robust power grid analysis
abstract
Robust and efficient algorithms for power grid analysis are crucial for both VLSI design and optimization. Due to the increasing size of power grids IR drop analysis has become more computationally challenging both in runtime and memory consumption. This work presents a fast Poisson solver preconditioned method for unstructured power grid with unideal boundary conditions. In fact, by taking the advantage of analytical formulation of power grids this analytical preconditioner can be considered as sparse approximate inverse technique. By combining this analytical preconditioner with robust conjugate gradient method, we demonstrate that this approach is totally robust for extremely large scale power grid simulations. Experimental results have shown that iterations of our proposed method will hardly increase with grid size increasing once the pads density and the range of metal resistances value distribution have been decided. We demonstrated that this approach solves an unstructured power grid with 2.56M nodes in only 1/3 iterations of classical ICCG solver, and achieves almost 20X speedups over the classical ICCG solver on runtime.
Jianlei Yang 0001, Yici Cai, Qiang Zhou 0001
ICCAD3
2011 PowerRush: A linear simulator for power grid
abstract
As the increasing size of power grids, IR drop analysis has become more computationally challenging both in runtime and memory consumption. In this paper, we propose a linear complexity simulator named PowerRush, which consists of an efficient SPICE Parser, a robust circuit Builder and a linear solver. The proposed solver is a pure algebraic method which can provide an optimal convergence without geometric information. It is implemented by Algebraic Multigrid Preconditioned Conjugate Gradient method, in which an aggregation based algebraic multigrid with K-Cycle acceleration is adopted as a preconditioner to improve the robustness of conjugate gradient iterative method. In multigrid scheme, double pairwise aggregation technique is applied to the matrix graph in coarsening procedure to ensure low setup cost and memory requirement. Further, a K-Cycle multigrid scheme is adopted to provide Krylov subspace acceleration at each level to guarantee optimal or near optimal convergence. Experimental results on real power grids have shown that PowerRush has a linear complexity in runtime cost and memory consumption. The DC analysis of a 60 Million nodes power grid can be solved by PowerRush for 0.01mV accuracy in 170 seconds with 21.89GB memory used.
Jianlei Yang 0001, Zuowei Li, Yici Cai, Qiang Zhou 0001
ICCAD4
2011 Floorplanning Considering IR Drop in Multiple Supply Voltages Island Designs
abstract
Voltage island has become a very effective design style for power saving in low-power design. However, the new design style also brings forward new challenges, especially to the designers of power/ground (P/G) networks. In this paper, we study the power delivery problem in voltage island designs, and propose to consider voltage drop during the floorplanning process to reduce design iterations. Our analysis shows that it is unnecessary to consider the pitch of the P/G network in the floorplan stage. By using the simplified searching strategy in floorplanning, we can obtain more robust low power design within reasonable runtime. Experimental results have demonstrated the effectiveness of our approach.
Qiang Zhou 0001, Bin Liu 0007, Yici Cai
IEEE Trans. Very Large Scale Integr. Syst.1
2010 An architecture-aware routing optimization via satisfiabilty for hierarchical FPGA
abstract
Boolean Satisfiability (SAT) has successfully been applied to the FPGA routing. It has many advantages over the conventional one-net-a-time routing algorithm such as routing all nets concurrently, higher flexibility and unroutability provable. However it also has the limits of scalability and is time-consuming. This paper presents some optimizations to the SAT-based routing approach by applying some architecture related features to the generated the Boolean constraints function. Specifically, Switch Box based connectivity optimization to reduce the variable number for each net, Logic Block pins rearrangement to improve the flexibility for each net and Exclusivity constraints optimization based on net-to-track distribution. Each of the optimizations is discussed in detail in this paper. Some heuristics and algorithms are also presented to implement the optimizations. We implement the SAT-based routing strategy as well as the optimizations on a general hierarchical FPGA architecture. The experimental results show that we can greatly reduce the variable and constraint number of the generated Boolean SAT functions. Hence, the generated SAT functions can be solved much more quickly. It also shows that high routing flexibility is also achieved due to the pins rearrangement.
Qiang Zhou 0001, Yici Cai, Jinian Bian
CSCWD2
2010 SAT based multi-net rip-up-and-reroute for manufacturing hotspot removal
abstract
Manufacturing hotspots are the layout patterns which cause excessive difficulties to manufacturing process. Design rules are effective at handling sizing/spacing induced hotspots, but are inadequate at dealing with topological hotspots. In wire routings, existing approaches often remove the hotspots through iteratively ripping up and rerouting one net at a time guided by litho-simulations. This procedure can be very time-consuming because litho-simulation is typically very slow and the rerouting may result in new hotspots due to its heuristic nature. In this paper, we propose a new approach for improving the efficiency of hotspot removal. In our approach, multiple nets in each hotspot region are simultaneously ripped up and rerouted based on Boolean satisfiability (SAT). The hotspot patterns, which are described and stored in a pre-built library, are forbidden to appear in the reroute through SAT constraints. Since multiple nets are simultaneously processed and SAT can guarantee to find a feasible solution if it exists, our approach can greatly accelerate the convergence on manufacturability. Experimental results on benchmark circuits show that our approach can remove over 90% of the hotspots in less than one minute on circuits with more than 20K nets and hundreds of hotspots.
Yici Cai, Qiang Zhou 0001, Jiang Hu 0001
DATE3
2010 Behavioral level dual-vth design for reduced leakage power with thermal awareness
abstract
Dual-Vthdesign is an effective leakage power reduction technique at behavioral synthesis level. It allows designers to replace modules on non-critical path with the high-Vthimplementation. However, the existing constructive algorithms fail to find the optimal solution due to the complexity of the problem and do not consider the on-chip temperature variation. In this paper, we propose a two-stage thermal-dependent leakage power minimization algorithm by using dual-Vthlibrary during behavioral synthesis. In the first stage, we quantitatively evaluate the timing impact on other modules caused by replacing certain modules with high Vth. Based on this analysis and the characteristics of the dual-Vthmodule library, we generate a small set of candidate solutions for the module replacement. Then in the second stage, we obtain the on-chip thermal information from thermal-aware floorplanning and thermal analysis to select the final solution from the candidate set. Experimental results show an average of 17.8% saving in leakage power consumption and a slightly shorter runtime compared to the best known work. In most cases, our algorithm can actually find the optimal solutions obtained from a complete solution space exploration.
Junbo Yu, Qiang Zhou 0001, Gang Qu 0001, Jinian Bian
DATE2
2010 Peak current reduction by simultaneous state replication and re-encoding
abstract
Peak current is one of the important considerations for circuit design and testing in the deep sub-micron technology. In a synchronous finite state machine (FSM), it is observed that the peak current happens at the moment of state transitions and it has a strong correlation with the maximum number of state registers switching in the same direction simultaneously [2], which we refer to as the peak switching value (PSV). We propose a FSM synthesis method to reduce P SV by seamlessly combining state replication and state re-encoding techniques. Our experiments show that out of 52 FSM benchmarks encoded by a state-of-the-art power-driven encoding algorithm POW3 [1], 36 of them are not optimal in terms of PSV. Our approach can improve on 34 of them with an average 39.2% PSV reduction, while the only comparable PSV-driven FSM synthesis technique [2] can improve on 27 benchmarks with an average 24.5% reduction. Furthermore, we compare our approach with [2] after the FSMs are implemented using an industry EDA tool. The results show that our approach reduces the peak current in the circuits by 13% on average and the total power by 3% with a mere 2% overhead in area.
Junjun Gu, Gang Qu 0001, Qiang Zhou 0001
ICCAD4
2010 Multilevel Optimization for Large-Scale Hierarchical FPGA Placement
Hui Dai, Qiang Zhou 0001, Jinian Bian
J. Comput. Sci. Technol.2
2010 ECP- and CMP-Aware Detailed Routing Algorithm for DFM
abstract
In this paper, a novel design-for-manufacture-aware detailed routing algorithm that seeks to minimize the thickness range of the chip surface after copper damascene process is proposed. The paper is based on an electroplating (ECP) and chemical mechanical polishing (CMP) model and predictors for final thickness range are abstracted. The proposed detailed routing is implemented in a W-shape multilevel full-chip routing framework using depth first search and branch-and-bound techniques in maze backtracking. Experimental results show that compared to maze routing (MR) (that does not consider CMP), the improvements in the average metal density standard and the average amount of dummy fill are 12.0% and 6.99% respectively. Compared to density-driven maze routing (DMR) that considers only CMP but does not consider ECP, the improvements in the average metal density standard and the average amount of dummy fill are 0.53% and 0.72%, respectively. So, the proposed algorithm can obtain improvement in optimizing CMP while the wire length and vias are not increased clearly and the completion rate is guaranteed. Therefore, the yield of chips is improved.
Yin Shen, Qiang Zhou 0001, Yici Cai, Xianlong Hong
IEEE Trans. Very Large Scale Integr. Syst.2
2009 Peak temperature control in thermal-aware behavioral synthesis through allocating the number of resources
abstract
High temperature adversely impacts on reliability, performance, and leakage power of ICs. In behavioral synthesis, both resource usage allocation and resource binding influence the final thermal profile. Previous thermal-aware behavioral syntheses only focused on binding, ignoring allocation. This paper proposes thermal-aware behavioral synthesis with resource usage allocation. According to power density and feedbacks from thermal simulation, we allocate the number of resources under area constraint. Our flow effectively controls peak temperature and creates even power densities among resources of “different” and “same” types. Compared to classic behavioral synthesis of peak temperature control, our technique reduces peak temperature by 11.1 °C on average with no area overhead and only 1.2 more steps latency overhead.
Junbo Yu, Qiang Zhou 0001, Jinian Bian
ASP-DAC2
2009 Fast placement for large-scale hierarchical FPGAs
abstract
In this paper, we propose a fast placer for FPGA placement on a new commercial hierarchical FPGA device. The novelty of this research lies in the application of a multilevel V-shape optimization flow including an architecture related cluster process and a constructive placement. The new placer can handle large-scale FPGA placement problem quickly. Experimental results show that the proposed placer can further reduced the wirelength average 28.3% compared with simulated annealing based tool while achieving near 5X speedup in runtime for the five largest MCNC benchmarks.
Hui Dai, Qiang Zhou 0001, Yici Cai, Jinian Bian, Xianlong Hong
CAD/Graphics2
2009 A thermal-driven force-directed floorplanning algorithm for 3D ICs
abstract
The three-dimensional (3D) integration circuit is a new technology with higher integration density. To solve the critical thermal issue in 3D layout, we propose a thermal-driven force-directed floorplanning algorithm. Based on the characteristic of the different stages of floorplanning, this algorithm applies different methods to calculate the thermal distribution to reach a tradeoff between time efficiency and accuracy. And a new effective strategy of the layer assignment is used in which we consider the area, the overlaps and the power densities simultaneously. Experimental results show that, compared with the recent thermal-driven force-directed 3D floorplanner, it averagely decreases the temperature by 8% and runtime by 10.7% while only increases the area and wirelength by 3% at most.
Qiang Zhou 0001, Yici Cai, Haixia Yan
CAD/Graphics2
2009 Global density smoothing technique for analytical placement algorithm
abstract
Cell migration has been widely used in global placement for the highly efficiency in smoothing cells overlap. The current cell migration methods only locally or globally smooth the density without considering the relation between the local and the global density. Furthermore, the cells generally are treated as points with area. In this paper we present a new cell migration technique called CSAGO to even the overlap. CSAGO obtains the movement distance while considering the local and the global density distribution simultaneously. By separating the standard cells and the macro blocks crossing multi-bin, a more smooth migration speed can be obtained. CSAGO has been embedded into the global placement process, experimental result shows that average HPWL have reduced 3% and the total runtime is 1.78 times faster than DPlace.
Qiang Zhou 0001, Jinian Bian, Yanming Jia
CAD/Graphics2
2009 Information hiding for trusted system design
abstract
For a computing system to be trusted, it is equally important to verify that the system performs no more and no less functionalities than desired. Traditional testing and verification methods are developed to validate whether the system meets all the requirements. They cannot detect the existence or show the non-existence of the unknown undesired functionalities. In this paper, we propose a novel approach that converts this problem to a less challenging design quality measuring problem. Our approach is based on information hiding and constraint manipulation of the original system design specification. We lay out the basic requirements for our approach and demonstrate it through the popular graph coloring problem. Results show that information can be embedded into the original graph without significant impact to the solution quality. However, when the same information is added to the graph modified based on our approach, there will be noticeable drop in the solution quality.
Junjun Gu, Gang Qu 0001, Qiang Zhou 0001
DAC3
2009 Improve clock gating through power-optimal enable function selection
abstract
Clock gating technology can reduce the consumption of clock signals' switching power of flip-flops. The clock gate enable functions can be identified by Boolean analysis of the logic inputs for all flip flops. However, the enable functions of clock gate can be further simplified, and the average number of flip flops driven by enable functions can be improved. In this way, the circuit area can be reduced; therefore, the clock gating can be improved and power saving can be achieved. This paper presents a technique for improving clock gating by optimizing the enable functions. The problem of improving clock gating is formulated as finding the optimal set of enable functions in the shared logic cone that leads to best power reduction on flip flops. First, enable functions are identified by random simulation and SAT. Then the optimal set of enable functions is found with partition method. This paper demonstrates the effectiveness of the approach through testing on MCNC benchmarks and industrial circuits. The experimental results show that the algorithm will get as much power saving as 3 times of that of the original clock gating circuits, and all benchmarks can run in tens of seconds.
Yunjian Jiang, Qiang Zhou 0001
DDECS4
2009 Fast congestion-aware timing-driven placement for island FPGA
abstract
A new fast timing-driven placement is presented in this paper, which is partitioning-based method, explicitly considering the congestion for island style FPGAs. The most distinct feature of this approach is that it not only reduces the circuit critical path delay efficiently, but also takes congestion into account. The harmony between partitioning objective and timing improvement goal is kept; moreover, the congestion constraint is added to cost function to improve routability in the meantime. As a result, it avoids the excessive usage of local routing resources while remaining circuit performance much better. The experimental results show our method, FCTP, is very fast. It is able to produce solutions with equal or better routability and up to average 8.19% improvement on performance but only less 1/3 average runtime compared to TVPR [1]. It also achieves much better results than PPFF [7] in terms of timing and congestion with negligible runtime penalty.
Jinpeng Zhao, Qiang Zhou 0001, Yici Cai
DDECS2
2009 Decoupling capacitance efficient placement for reducing transient power supply noise
abstract
Decoupling capacitance (decap) is an efficient way to reduce transient noise in on-chip power supply networks. However, excessive decap may cause more leakage power, chip resource waste, and even lead to more design iterations. In this paper, we present a novel decap-efficient placement algorithm for transient power supply noise reduction. In contrast to traditional design flow, our approach considers decap impacts at the placement stage to seek the placement minimizing decap requirements while still satisfying the traditional placement objectives. In the new method, we first devise a fast procedure to assess the decap requirement for the force-based placement framework, in which the required decap is modeled as a density function over the chip. Then, we build a corresponding supply and demand system to adjust the placement in favor of minimizing decap. Finally, we develop a decap efficient placement algorithm with a new force induced by imbalance between power supply and power demands. Experimental results show that the new combined placement and decap optimization flow could reduce the minimum decap area by 35% with a wire length increase of only 0.5% at nearly the same computational cost, which is efficient for practical problems.
Yici Cai, Qiang Zhou 0001, Sheldon X.-D. Tan, Thom Jefferson A. Eguia
ICCAD3
2009 Thermal aware placement in 3D ICs using quadratic uniformity modeling approach
Haixia Yan, Qiang Zhou 0001, Xianlong Hong
Integr.2
2009 An MTCMOS technology for low-power physical design
Qiang Zhou 0001, Yici Cai, Xianlong Hong
Integr.1
2008 Low power clock buffer planning methodology in F-D placement for large scale circuit design
abstract
Traditionally, clock network layout is performed after cell placement. Such methodology is facing a serious problem in nanometer IC designs where people tend to use huge clock buffers for robustness against variations. That is, clock buffers are often placed far from ideal locations to avoid overlap with logic cells. As a result, both power dissipation and timing are degraded. In order to solve this problem, we propose a low power clock buffer planning methodology which is integrated with cell placement. A Bin- Divided Grouping algorithm is developed to construct virtual buffer tree, which can explicitly model the clock buffers in placement. The virtual buffer tree is dynamically updated during the placement to reflect the changes of latch locations. To reduce power dissipation, latch clumping is incorporated with the clock buffer planning. The experimental results show that our method can reduce clock power significantly by 21% on average.
Qiang Zhou 0001, Yici Cai, Jiang Hu 0001, Xianlong Hong, Jinian Bian
ASP-DAC2
2008 MacroMap: A technology mapping algorithm for heterogeneous FPGAs with effective area estimation
abstract
Recent generation of FPGA devices takes advantage of speed and density benefits resulted from heterogeneous FPGA architecture, in which several basic LUTs can be combined to form one larger size LUT called Macro. Large Macros not only decrease network depth efficiently but also reduce area. In this paper, a new technology mapping algorithm, named MacroMap is proposed for the heterogeneous FPGAs with effective area estimation to overcome the main disadvantage that traditional technology mapping algorithms only generate one kind of typical K-LUT and cannot make full use of LUTs with different sizes (basic LUTs and Macros). Experimental results show that MacroMap can obtain 19% gain on area while keeping the network depth optimal compared with the existing heterogeneous FPGA mapping algorithm heteromap[8].
Qiang Zhou 0001, Yici Cai, Jinian Bian, Xianlong Hong
FPL3
2008 A novel performance driven power gating based on distributed sleep transistor network
abstract
Power Gating is an effective method to reduce leakage power. One of the most important issues in power gating design is the decision on the size of sleep transistor, which is mostly determined by the maximum instantaneous current (MIC) and the maximum tolerable voltage drop. In order to reduce the sleep transistor area, the distributed sleep transistor network (DSTN) was proposed to reduce MIC by connecting all the virtual ground nets together. Most of the following works focused on estimating the MICs through sleep transistors accurately. But the previous works use a pre-defined global voltage drop constraint on circuit, which leads to a uniform gate slowdown. In this paper, we propose a performance driven methodology for DSTN design, which exploits the maximum tolerable voltage drops of gates, particularly the non-critical ones, to reduce the total sleep transistor area without additional performance loss. Moreover, a clustering strategy in placement is proposed to help further reduce the total sleep transistor area. Experimental results show that the proposed approach can reduce the total sleep transistor area by about 36% on average.
Liangpeng Guo, Yici Cai, Qiang Zhou 0001, Xianlong Hong
ACM Great Lakes Symposium on VLSI3
2008 Application of optical proximity correction technology
Yici Cai, Qiang Zhou 0001, Xianlong Hong
Sci. China Ser. F Inf. Sci.2
2007 Logic and Layout Aware Voltage Island Generation for Low Power Design
abstract
Multiple supply voltage (MSV) is one of the most effective schemes to achieve low power, but most works are based on logic level. A few recent works are based on physical level but all of them do not consider level converters which have an important effect in dual-vdd design. In this work we propose a logic and layout aware approach for voltage assignment and voltage island generation in placement process to minimize the number of level converters and to implement voltage islands with minimal overheads. Experimental results show that our approach uses much less level converters than the approach in (Bin Liu, 2006) (reduced by 59.50% on average) when achieving the same power savings. The approach is able to produce feasible placement with a small impact to traditional placement goals.
Liangpeng Guo, Yici Cai, Qiang Zhou 0001, Xianlong Hong
ASP-DAC3
2007 Micro-architecture Pipelining Optimization with Throughput-Aware Floorplanning
abstract
For modern processor designs in nanometer technologies, both block and interconnect pipelining are needed to achieve multi-gigahertz clock frequency, but previous approaches consider block pipelining and interconnect pipelining separately. For example, all recent works on wire pipelining assume pre-pipelined components and consider only inserting pipeline stages on point-to-point wire or bus connections. To the best of our knowledge, this paper is the first that considers block pipelining and interconnect pipelining simultaneously. We optimize multiple critical paths or loops in the micro-architecture and insert the pipelines stages optimally in the blocks and wires of these loops to meet the clock frequency requirement. We propose two approaches to this problem. The first approach is based on mixed integer linear programming (MILP) which is theoretically guaranteed to produce the optimal solution, and the second one is an efficient graph-based algorithm that produces near-optimal solutions. Experimental results show that simultaneous block and interconnect pipelining leads to more than 20% improvement over wire-pipelining alone on the overall processor performance. Moreover, the graph-based approach gives solutions very close to the MILP results ( 2% more than MILP results on average) but in a much shorter runtime.
Yuchun Ma, Zhuoyuan Li 0003, Jason Cong, Xianlong Hong, Glenn Reinman, Sheqin Dong, Qiang Zhou 0001
ASP-DAC7
2007 Practical Implementation of Stochastic Parameterized Model Order Reduction via Hermite Polynomial Chaos
abstract
This paper describes the stochastic model order reduction algorithm via stochastic Hermite polynomials from the practical implementation perspective. Comparing with existing work on stochastic interconnect analysis and parameterized model order reduction, we generalized the input variation representation using polynomial chaos (PC) to allow for accurate modeling of non-Gaussian input variations. We also explore the implicit system representation using sub-matrices and improved the efficiency for solving the linear equations utilizing block matrix structure of the augmented system. Experiments show that our algorithm matches with Monte Carlo methods very well while keeping the algorithm effective. And the PC representation of non-Gaussian variables gains more accuracy than Taylor representation used in previous work (Wang et al., 2004).
Yici Cai, Qiang Zhou 0001, Xianlong Hong, Sheldon X.-D. Tan
ASP-DAC3
2007 VPH: Versatile Routability-Driven Place Algorithm for Hierarchical FPGAs Based on VPR
abstract
VPH (versatile placer for hierarchical FPGAs, HFPGAs) is a place tool aiming at a routability-driven placement process for HFPGAs. It improves the placement algorithm of VPR (versatile place and route) by taking into consideration of specific constraints of hierarchical architectures for HFPGAs, and updates the place process based on it. Further more, thanks to the interconnect predictability, VPH can take into account routing constraints in early placement stage and evaluate efficiently the routability of the circuit. In this paper, we introduce the VPH design framework and its related placement algorithms. Its effectiveness is validated by the experimental results of MCNC benchmark.
Qiang Zhou 0001, Jinian Bian, Junhua Qu
CAD/Graphics2
2007 Thermal Effects with Leakage Power Considered in 2D/3D Floorplanning
abstract
Leakage power is becoming a key design challenge in current and future CMOS designs. Due to technology scaling, the leakage power is rising so quickly that it largely elevates the die temperature. In this paper, we deeply investigate the impact of leakage power on thermal profile in 2D and 3D floorplanning. Our results show that chip temperature can increase by about 11 V in 2D design and 68 V for 3D case with leakage power considered. Then we propose a thermal-driven floorplanning flow integrated with an iterative leakage-aware thermal analysis process to optimize chip temperature and save leakage power consumption. Experimental results show that for 2D design, the max chip temperature can be reduced by about 8 "C and the proportion of leakage power to total power can be reduced from 19.17% to 11.12%. The corresponding results for 3D are 60 degC temperature reduction and 16.3% less leakage power proportion.
Pingqiang Zhou, Yuchun Ma, Qiang Zhou 0001, Xianlong Hong
CAD/Graphics3
2007 New timing and routability driven placement algorithms for FPGA synthesis
abstract
We present new timing and congestion driven FPGA placement algorithms with minimal runtime overhead. By predicting the post-routing critical edges and estimating congestion accurately, our algorithms simultaneously reduce the critical path delay and the minimum number of routing tracks. The core of our algorithm consists of a criticality history record of connection edges and a congestion map. This approach is applied to the 20 largest MCNC benchmark circuits. Experimental results show that compared with VPR [1], our algorithms yield an average of 8.1% reduction (maximum 30.5%) in the critical path delay and 5% reduction in channel width. Meanwhile, the average runtime of our algorithms is only 2.3X as of VPR's.
Hao Li 0030, Qiang Zhou 0001, Yici Cai, Xianlong Hong
ACM Great Lakes Symposium on VLSI3
2007 3D-STAF: scalable temperature and leakage aware floorplanning for three-dimensional integrated circuits
abstract
Thermal issues are a primary concern in the threedimensional (3D) integrated circuit (IC) design. Temperature, area, and wire length must be simultaneously optimized during 3D floorplanning, significantly increasing optimization complexity. Most existing floorplanners use combinatorial stochastic optimization techniques, hampering performance and scalability when used for 3D floorplanning. In this work, we propose and evaluate a scalable, temperature-aware, force-directed floorplanner called 3D-STAF. Force-directed techniques, although efficient at reacting to physical information such as temperature gradients, must eventually eliminate overlap. This can cause significant displacement when used for heterogeneous blocks. To smooth the transition from an unconstrained 3D placement to a legalized, layer-assigned floorplan, we propose a three-stage force-directed optimization flow combined with new legalization techniques that eliminate white spaces and block overlapping during multi-layer floorplanning. A temperature-dependent leakage model is used within 3D-STAF to permit optimization based on the feedback loop connecting thermal profile and leakage power consumption. 3D-STAF has good performance that scales well for large problem instances. Compared to recently published 3D floorplanning work, 3D-STAF improves the area by 6%, wire length by 16%, via count by 22%, peak temperature by 6% while running nearly 4× faster on average.
Pingqiang Zhou, Yuchun Ma, Zhuoyuan Li 0003, Robert P. Dick, Hai Zhou 0001, Xianlong Hong, Qiang Zhou 0001
ICCAD8
2007 Clock-Tree Aware Placement Based on Dynamic Clock-Tree Building
abstract
Minimization of clock network is traditionally achieved by clock routing, which may be helpless for a poor placement result. In this paper, a novel Dynamic Clock-Tree Building technique integrated into placement for zero-skew design is proposed. This method combines a pre-designed clock-tree with the Force-Directed Placement procedure to navigate the register placement for minimizing the clock network. Meanwhile, a new model of Multi-Level Bounding Box and technique of Multi-Level Attractive Force are proposed to give a better local distribution of registers. Experiments on several standard-cell benchmarks indicate an average 26.1% clock network reduction with the logic cell placement preserved well.
Qiang Zhou 0001, Xianlong Hong, Yici Cai
ISCAS2
2007 Unified Quadratic Programming Approach For 3-D Mixed Mode Placement
abstract
An efficient analytical 3D placement algorithm for mixed-mode placement is presented, which consists of 3D global placement and detailed placement. In global placement, wire length and cell division are unified into a quadratic objective function. It takes advantage of quadratic programming to optimize the unified objective efficiently. 3D discrete cosine transformation (DCT) is introduced to help divide cells into different layers. The number of vertical vias gets better controlled during global placement and a new method to optimize cell division after global placement is presented. For detailed placement, we traverse 3D to 2D by net decomposition and finish detailed placement by network flow algorithm. Experimental results show that the 3D placement algorithm is very promising.
Haixia Yan, Zhuoyuan Li 0003, Xianlong Hong, Qiang Zhou 0001
ISCAS4
2007 An efficient quadratic placement based on search space traversing technology
Yongqiang Lyu 0001, Xianlong Hong, Qiang Zhou 0001, Yici Cai
Integr.3
2007 A Yield-Driven Gridless Router
Qiang Zhou 0001, Yici Cai, Xianlong Hong
J. Comput. Sci. Technol.1
2007 Efficient Thermal via Planning Approach and Its Application in 3-D Floorplanning
abstract
In this paper, we investigate thermal via (T-via) planning during three-dimensional (3-D) floorplanning. First, we consider the temperature constrained T-via planning (TVP) problem on a given 3-D floorplan. Second, we integrate dynamic TVP into 3-D floorplanning process. Our main contribution and results can be summarized as follows. We solve the temperature constrained TVP problem by solving a sequence of simplified interlayer and intralayer TVP subproblems. Each subproblem is formulated as convex programming problem and we derive nearly optimal solution for detailed T-via distribution. Based on the TVP solution, we implement the integrated TVP and 3-D floorplanning algorithm in a two-stage approach. Before floorplanning, blocks are assigned into different layers by solving a sequence of knapsack problems. During floorplanning, T-vias are allocated with white space redistribution to optimize T-via insertion. Experimental results show that our TVP approach can reduce T-vias by 12% compared with a recent published work (J. Cong and Y. Zhang, "Thermal via planning for 3-D ICs," in Proc. Int. Conf. Comput.-Aided Des., Nov. 2005, pp.745-752). Compared with the postfloorplanning optimization approach, integrating TVP into floorplanning process can reduce T-vias by 16% with 21% runtime overhead
Zhuoyuan Li 0003, Xianlong Hong, Qiang Zhou 0001, Shan Zeng, Jinian Bian, Wenjian Yu, Hannah Honghua Yang, Vijay Pitchumani, Chung-Kuan Cheng
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2006 Power driven placement with layout aware supply voltage assignment for voltage island generation in Dual-Vdd designs
abstract
In this paper we propose a method for standard cell placement with support for dual supply voltages, aiming to reduce total power under timing constraints and to implement voltage islands with minimal overheads. The method begins with timing and power driven coarse placement, followed by a few iterations between voltage assignment and placement refinement to generate voltage islands. Several techniques, including timing and power driven net weighting, seed growth based voltage assignment, and soft clustering strategy for placement refinements are employed in our implementation. Experimental results on a set of MCNC benchmarks show that our approach is able to produce feasible placement for dual-Vdd designs and significantly reduce total power with a wirelength increase within 14% compared to a power and timing driven placer without voltage islands.
Bin Liu 0007, Yici Cai, Qiang Zhou 0001, Xianlong Hong
ASP-DAC3
2006 A novel technique integrating buffer insertion into timing driven placement
abstract
Increasing buffer number for future technology makes traditional one-pass-flow (timing driven placement is followed by buffer insertion and legalization) failed, since accommodation for buffers significantly disturbs original design. This paper exploits the delicate relationship between buffer insertion and timing driven placement, and proposes a novel method to incorporate buffer insertion during timing driven placement. Experimental results show that this incorporation not only ensures design convergence, but also benefits timing behavior and alleviates buffer explosion
Lijuan Luo, Qiang Zhou 0001, Yici Cai, Xianlong Hong, Yibo Wang 0009
ISCAS2
2006 A novel low-power physical design methodology for MTCMOS
abstract
The optimization of virtual supply network plays an important role in MTCMOS low power design. Existing low power works are mainly on gate-level without any optimization on physical design level, which can lead to large amount of virtual supply networks. This paper presents (1) a low power driven physical design flow; (2) a novel low power placement to simultaneously place standard cells and sleep transistors and (3) sleep transistor relocation technique to further reduce the virtual supply networks. Experiment results are promising for both achieving up to 28.15% savings for virtual supply networks and well controlling the increase of signal nets
Yici Cai, Qiang Zhou 0001, Xianlong Hong
ISCAS3
2006 Integrating dynamic thermal via planning with 3D floorplanning algorithm
abstract
Incorporating thermal vias into 3D ICs is a promising way to reduce circuit temperature by lowering down the thermal resistances between device layers. In this paper, we integrate dynamic thermal via planning into 3D floorplanning process. Our 3D floorplanning and thermal via planning approaches are implemented in a two-stage approach. Before floorplanning, the temperature-constrained vertical thermal via planning is formulated as a convex programming problem. Based on the analytical solution, blocks are assigned into different layers by solving a sequence of knapsack problems. Then a SA engine is used to generate floorplans of all these layers simultaneously. During floorplanning, thermal vias are distributed horizontally in each layer with white space redistribution to optimize thermal via insertion. Experimental results show that compared to a recent published result from [14], our method can reduce thermal vias by 15% with 38% runtime overhead.
Zhuoyuan Li 0003, Xianlong Hong, Qiang Zhou 0001, Shan Zeng, Jinian Bian, Hannah Honghua Yang, Vijay Pitchumani, Chung-Kuan Cheng
ISPD3
2006 Priority-Based Routing Resource Assignment Considering Crosstalk
Yici Cai, Bin Liu 0007, Yan Xiong 0001, Qiang Zhou 0001, Xianlong Hong
J. Comput. Sci. Technol.4
2006 Efficient thermal-oriented 3D floorplanning and thermal via planning for two-stacked-die integration
abstract
New three-dimensional (3D) floorplanning and thermal via planning algorithms are proposed for thermal optimization in two-stacked die integration. Our contributions include (1) a two-stage design flow for 3D floorplanning, which scales down the enlarged solution space due to multidevice layer structure; (2) an efficient thermal-driven 3D floorplanning algorithm with power distribution constraints; (3) a thermal via planning algorithm considering congestion minimization. Experiments results show that our approach is nine times faster with better solution quality compared to a recent published result. In addition, the thermal via planning approach is proven to be very efficient to eliminate localized hot spots directly.
Zhuoyuan Li 0003, Xianlong Hong, Qiang Zhou 0001, Jinian Bian, Hannah Honghua Yang, Vijay Pitchumani
ACM Trans. Design Autom. Electr. Syst.3
2005 Clock network minimization methodology based on incremental placement
abstract
In ultra-deep submicron VLSI circuits, clock network is a major source of power consumption and power supply noise. Therefore, it is very important to minimize clock network size. Traditional design methodologies usually let the clock router to undertake the task of clock network minimization independently. Since a clock routing is carried out based on register locations, register placement actually has fundamental influence to a clock network size. In this paper, we propose a new clock network design methodology that Incorporates register placement optimization. Given a cell placement result, incremental modifications are performed according to clock skew specifications. The incremental placement change moves registers toward preferred locations that may enable a small clock network size. At the same time, the side-effect to logic cell placement and wire connections is controlled. Experimental results on benchmark circuits show that the proposed methodology can reduce clock network size considerably with limited impact on signal net wirelength and critical path delay.
Yici Cai, Qiang Zhou 0001, Xianlong Hong, Jiang Hu 0001, Yongqiang Lyu 0001
ASP-DAC3
2005 Register placement for low power clock network
abstract
In modern VLSI designs, the increasingly severe power problem requests to minimize clock routing wirelength so that both power consumption and power supply noise can be alleviated. In contrast to most of traditional works that handle this problem only in clock routing, we propose to navigate standard cell register placement to locations that enable further less clock routing wirelength and power. To minimize adverse impacts to conventional cell placement goals such as signal net wirelength and critical path delay, the register placement is carried out in the context of a quadratic placement. The proposed technique is particularly effective for the recently popular prescribed skew clock routing. Experiments on benchmark circuits show encouraging results.
Yongqiang Lyu 0001, Cliff C. N. Sze, Xianlong Hong, Qiang Zhou 0001, Yici Cai, Jiang Hu 0001
ASP-DAC4
2005 Analysis of buffered hybrid structured clock networks
abstract
This paper presents a novel approach for fast transient analysis of buffered hybrid structured clock networks. The new method applies structure reduction and relaxed hierarchical analysis methods to reduce the circuit complexity and speedup the simulation. A simple controlled sources model is used for modeling clock buffers to deal with nonlinearity in the buffered clock trees. Our experiment results show that the proposed algorithm is about two orders of magnitude faster than HSPICE without loss on accuracy and stability. The relatively errors on delay times are within a few percent of the exact ones.
Qiang Zhou 0001, Yici Cai, Xianlong Hong, Sheldon X.-D. Tan
ASP-DAC2
2005 Navigating registers in placement for clock network minimization
abstract
The progress of VLSI technology is facing two limiting factors: power and variation. Minimizing clock network size can lead to reduced power consumption, less power supply noise, less number of clock buffers and therefore less vulnerability to variations. Previous works on clock network minimization are mostly focused on clock routing and the improvements are often limited by the input register placement. In this work, we propose to navigate registers in cell placement for further clock network size reduction. To solve the conflict between clock network minimization and traditional placement goals, we suggest the following techniques in a quadratic placement framework: (1) Manhattan ring based register guidance; (2) center of gravity constraints for registers; (3) pseudo pin and net; (4) register cluster contraction. These techniques work for both zero skew and prescribed skew designs in both wirelength driven and timing driven placement. Experimental results show that our method can reduce clock net wirelength by 16%~33% with no more than 0.5% increase on signal net wirelength compared with conventional approaches.
Yongqiang Lyu 0001, Cliff C. N. Sze, Xianlong Hong, Qiang Zhou 0001, Yici Cai, Jiang Hu 0001
DAC4
2005 A new algorithm for layout of dark field alternating phase shifting masks
abstract
A new methodology is proposed to accelerate AltPSM design flow for dark field AltPSM for large-scale layouts. When scaling to large-scale layouts, designing AltPSM may be much time consuming. Our new algorithm solves this problem by splitting a layout into smaller, easier-to-solve parts, solving the sub-layouts independently and simultaneously, and then recombining the sub-layouts. The experimental results on industry layouts indicate that parallel algorithm has potential to provide more significant improvement in speed and achieve a better quality of solutions.
Qinglang Luo, Xianlong Hong, Qiang Zhou 0001, Yici Cai
ACM Great Lakes Symposium on VLSI3
2005 Improved multilevel routing with redundant via placement for yield and reliability
abstract
This paper presents an improved multilevel Full-chip routing system which integrates global routing and detailed routing algorithms to achieve great enhancement in yield and reliability considering the redundant via placement. The system features a pre-coarsening stage which is equipped with a fast congestion-driven L-pattern global routing followed by the rvia-driven detailed routing. The L-pattern global routing benefits a lot to the reduction of vias and thus relieves the burden of redundant via addition. Then the rvia-driven maze routing algorithm considers the addition of redundant vias during routing. Finally the redundant via placement heuristic also contributes to improve the completion rate. We have tested the system on a set of commonly used benchmark circuits and compared the results with a previous multilevel routing framework. The experimental results are promising.
Hailong Yao 0002, Yici Cai, Xianlong Hong, Qiang Zhou 0001
ACM Great Lakes Symposium on VLSI4
2005 Multi-stage Detailed Placement Algorithm for Large-Scale Mixed-Mode Layout Design
Lijuan Luo, Qiang Zhou 0001, Xianlong Hong, Hanbin Zhou
ICCSA (4)2
2005 Shielding Area Optimization Under the Solution of Interconnect Crosstalk
Yici Cai, Qiang Zhou 0001, Xianlong Hong
J. Comput. Sci. Technol.3
2005 Crosstalk-Aware Routing Resource Assignment
Hailong Yao 0002, Yici Cai, Qiang Zhou 0001, Xianlong Hong
J. Comput. Sci. Technol.3
2004 A Fast Delay Analysis Algorithm for The Hybrid Structured Clock Network
abstract
This paper presents a novel approach to reducing the complexity of the transient linear circuit analysis for a hybrid structured clock network. Topology reduction is first used to reduce the complexity of the circuits and a preconditioned Krylov-subspace iterative method is then used to perform the nodal analysis on the reduced circuits. By proper choice of the simulation time step based on Elmore delay model, the delay of the clock signal between the clock source and the sink node and the skews between the sink nodes can be obtained efficiently and accurately. Our experimental results show that the proposed algorithm is two orders of magnitude faster than HSPICE without loss of accuracy and stability and the maximum error is within 0.4% of the exact delay time.
Yici Cai, Qiang Zhou 0001, Xianlong Hong, Sheldon X.-D. Tan
ICCD3