EDBT 2026 Demo / reviewers in the wild / expert
Yu Hu 0001
dblp:08/6001-1
· DBLP profile ↗
117ranked-venue papers
5as first author
36since 2021 · last 2026
0000-0001-8818-4075ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 84 · 4 first-author · 15 since 2021Artificial intelligence and machine learning · 39 · 34 since 2021Software engineering, systems software and programming languages · 18 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 7 · 5 since 2021Security and privacy · 5 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 5 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1Theory of computation · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SEA: Hierarchically searching efficient adapters for pre-trained models
Shun Lu 0001, Fangyuan Mao, Junkun Chen, Jilin Mei, Yu Hu 0001 |
Neural Networks | 7 |
| 2026 | PID: Physics-Informed Diffusion Model for Infrared Image Generation
Fangyuan Mao, Jilin Mei, Shun Lu 0001, Fuyang Liu, Fangzhou Zhao, Yu Hu 0001 |
Pattern Recognit. | 7 |
| 2025 | ROD: RGB-Only Fast and Efficient Off-Road Freespace DetectionabstractOff-road freespace detection is more challenging than on-road scenarios because of the blurred boundaries of traversable areas. Previous state-of-the-art (SOTA) methods employ multi-modal fusion of RGB images and LiDAR data. However, due to the significant increase in inference time when calculating surface normal maps from LiDAR data, multimodal methods are not suitable for real-time applications, particularly in real-world scenarios where higher FPS is required compared to slow navigation. This paper presents a novel RGB-only approach for off-road freespace detection, named ROD, eliminating the reliance on LiDAR data and its computational demands. Specifically, we utilize a pre-trained Vision Transformer (ViT) to extract rich features from RGB images. Additionally, we design a lightweight yet efficient decoder, which together improve both precision and inference speed. ROD establishes a new SOTA on ORFD and RELLIS-3D datasets, as well as an inference speed of 50 FPS, significantly outperforming prior models. Our code will be available at https://github.com/STLIFE97/offroad_roadseg. Hongliang Ye, Jilin Mei, Fangzhou Zhao, Leiqiang Zong, Yu Hu 0001 |
ICRA | 7 |
| 2025 | Adaptive Sliding Window Optimization for Multi-Modal LiDAR Inertial Odometry and MappingabstractFixed-Lag smoothing is widely employed as a backend in localization tasks. Generally, increasing the window length leads to better accuracy, but demands more computational resources. Therefore, determining an appropriate window length and whether a fixed length should be maintained throughout the localization process are worth studying. Assuming independent and identically distributed noise based on the distance-independent characteristic of LiDAR ranging errors, we propose an uncertainty-based adaptive sliding window (ASW) strategy. Through mathematical derivation, the reference uncertainty is affected by the LiDAR feature distribution of each frame. Consequently, we develop a multimodal LiDAR inertial odometry and mapping framework based on ASW, which integrates mechanical and solid-state LiDAR to enhance odometry accuracy and mapping density. By designing a joint matching module, our approach leverages the strengths of distinct scanning patterns. Additionally, we incorporate loop closure detection in the mapping process to minimize cumulative drift. Extensive experiments conducted on both public and self-collected datasets demonstrate the effectiveness of our method. Compared to the state-of-the-art method, our approach improves the average accuracy by 10.3%. We also provide an open-source implementation for further studies. https://github.com/wowhhhhgd/ASW-LIOM. Wei Li 0235, Yu Hu 0001 |
IROS | 3 |
| 2025 | CORENet: Cross-Modal 4D Radar Denoising Network with LiDAR Supervision for Autonomous Drivingabstract4D radar-based object detection has garnered great attention for its robustness in adverse weather conditions and capacity to deliver rich spatial information across diverse driving scenarios. Nevertheless, the sparse and noisy nature of 4D radar point clouds poses substantial challenges for effective perception. To address the limitation, we present CORENet, a novel cross-modal denoising framework that leverages LiDAR supervision to identify noise patterns and extract discriminative features from raw 4D radar data. Designed as a plug-and-play architecture, our solution enables seamless integration into voxel-based detection frameworks without modifying existing pipelines. Notably, the proposed method only utilizes LiDAR data for cross-modal supervision during training while maintaining full radar-only operation during inference. Extensive evaluation on the challenging Dual-Radar dataset, which is characterized by elevated noise level, demonstrates the effectiveness of our framework in enhancing detection robustness. Comprehensive experiments validate that CORENet achieves superior performance compared to existing mainstream approaches. The code is available at https://github.com/charlesuv/corenet.git. Fuyang Liu, Jilin Mei, Fangyuan Mao, Yu Hu 0001 |
IROS | 6 |
| 2025 | CA2Point: Learning Keypoint Detection and Description with Context Aggregation and Cross AugmentationabstractKeypoint detection and description are fundamental tasks for a variety of computer vision applications. Due to the limited receptive field of convolutional neural networks, most existing methods based on deep learning mainly focus on the local features, instead of taking into account the global context from entire image. The purpose of this work is to enhance the detection and description process of keypoints by leveraging global information obtained from Transformer, and to boost the consistence between keypoints and descriptors through their interaction. Specifically, the above two improvements are respectively implemented through the Local & Global Context Aggregation (LGCA) Module and Point & Descriptor Cross Augmentation (PDCA) Module proposed in this article. The LGCA module, which can model the long-range context, is inserted a Feature Pyramid Network (FPN) to extract features which contain diverse scales and different receptive fields. Moreover, the PDCA module enhances descriptors by the geometry information of keypoints detected, while enhancing the keypoint detection process by the position coordinates of correctly matched descriptors. Finally, we design a lightweight model to improve the running efficiency. Extensive experiments on various tasks demonstrate that our method achieves a substantial performance improvement over the current feature extraction methods. Code is available at: https://github.com/meng152634/CA2Point. Xuebin Meng, Wei Li 0235, Yu Hu 0001, Yinhe Han 0001 |
IROS | 3 |
| 2025 | Vibration-Aware Trajectory Optimization for Mobile Robots in Wild Environments via Physics-Informed Neural NetworkabstractThe suspension system, through effective damping of vibrations and shocks, can enhance the stability of wheeled robots traversing challenging terrain. Because the suspension system decouples the rigid correspondence between terrain changes and robot vibrations, considering suspension modeling in trajectory planning offers the advantage of more accurate prediction of the robot’s response to terrain. This improved predictive capability facilitates the planning of safer trajectories and may reduce tracking errors in the subsequent control process. In this work, inspired by the structure of Physics-Informed Neural Network (PINN), we propose a physics-informed planning method that considers the vibrational effects of complex nonlinear suspension systems. In addition, we design a two-stage process to accelerate training. By incorporating PINN, our method can better guarantee the physical feasibility of the planned trajectories. The proposed approach has been evaluated on a real robot platform. Compared to state-of-the-art baseline methods, our proposed approach achieves a 15.38% reduction in hazardous planning for mobile robots in wild environments. Aochun Xu, Andong Yang, Wei Li 0235, Yu Hu 0001 |
IROS | 4 |
| 2025 | Generating Synthetic Deviation Maps for Prior-Enhanced Vectorized HD Map ConstructionabstractHigh-definition (HD) maps are essential for autonomous driving, providing detailed and accurate environmental information. Recent advancements in online vectorized HD map construction have shown great promise, particularly methods that leverage existing maps as prior knowledge to improve performance. However, the robustness of these prior-enhanced methods under varying deviations between the priors and the real world remains a critical concern. This paper introduces a novel framework for generating synthetic maps, which allows for the controllable magnitude of diverse deviations, including geometric distortions, topological errors, and semantic inconsistencies, simulating real-world scenarios where prior maps may be outdated or inaccurate. Furthermore, lane group constraints are designed to avoid positional conflicts when map elements are modified. The synthesis method can overcome the time-consuming challenge of collecting real road changes. We demonstrate the utility of the synthetic deviation maps by incorporating them into the state-of-the-art prior-enhanced construction methods. The results reveal how different types and degrees of deviations affect the prediction accuracy, providing worthwhile insights into their robustness. Overall, this work contributes to a data augmentation and provides a valuable tool for developing more robust and reliable autonomous driving systems. The code is opensource and available at https://github.com/healenrens/Syn-D-maps. Yiyang Xiao, Wei Li 0235, Yu Hu 0001 |
IV | 4 |
| 2024 | F3DMP: Foresighted 3D Motion Planning of Mobile Robots in Wild EnvironmentsabstractIn wild environments, motion planning for mobile robots faces the challenge of local optimal path traps due to limited sensor perception range and lack of spatial awareness. Existing approaches that avoid local optimum by designing heuristic functions or high-quality global paths in wild environments are time-consuming and unstable. This work proposes F3DMP, which consists of two parts to alleviate the local optimum solution and better utilize distant terrain information. First, the entire planning framework is adapted to the three-dimensional space so that the planning result conforms to the geometric characteristics of the terrain. Second, a time allocation function based on offline reinforcement learning is proposed. This function can anticipate potential challenges or opportunities based on semantic information for the image and proactively determine a time allocation. Our planner is integrated into a complete mobile robot system and deployed to a real robot. Experiments in simulation and the real world demonstrate that our method can improve the success rate by 28% and the trajectory smoothness by 27% compared with traditional methods. Andong Yang, Wei Li 0235, Yu Hu 0001 |
ICRA | 3 |
| 2024 | A Safe and Efficient Timed-Elastic-Band Planner for Unstructured EnvironmentsabstractIn unstructured environments with complex obstacles and obscure road boundaries, the local planner faces more severe challenges in terms of safety and real-time performance. In order to fulfill these emerging requirements, we propose a novel Timed-Elastic-Band approach for unstructured environments, abbreviated as TEB-U. This approach incorporates a free space extraction optimization module for 2D occupancy grid maps, which efficiently transforms irregular free space boundaries into polygons and restrains robots within the boundaries. Moreover, a dynamic global point adjustment module is designed to adaptively correct the trajectory points obtained from the global planner, thereby enabling robots to travel along the centerline of free space and providing a better initial trajectory for subsequent modules. To reduce the computational cost, we replace the obstacle constraint of TEB with the boundary constraint in hyper-graph optimization. We evaluate our planner in three distinct scenarios, and the results show that TEB-U improves the average success rate by 21% and reduces the planning time by 23% compared to TEB in unstructured road, which demonstrates its safety and efficiency. Haoyu Xi, Wei Li 0235, Fangzhou Zhao, Yu Hu 0001 |
IROS | 5 |
| 2024 | SCOML: Trajectory Planning Based on Self-Correcting Meta-Reinforcement Learning in Hybrid Terrain for Mobile RobotabstractTrajectory planning is important for ground robots to achieve safe and efficient autonomous navigation in unstructured off-road environments. Most existing methods treat each terrain as a single type. However, in the real world, a ground usually consists of hybrid terrains. In this work, we propose a novel trajectory planning network that handles hybrid terrain. To further enhance safety, we have designed a self-correcting structure based on historical planning data. This structure can correct the trajectory when an inappropriate one is planned. To train the network, we introduce a two-stage training scheme based on Offline Meta-Reinforcement Learning, which can train the network with pre-collected non-optimal datasets and reduce the occurrence of hazardous planning. The proposed approach has been evaluated on both simulated datasets and a real robot platform. Compared to state-of-the-art baseline methods, the proposed approach reduces hazardous planning by 59.3% in hybrid terrains. Andong Yang, Wei Li 0235, Yu Hu 0001 |
IROS | 3 |
| 2024 | TeFF: Tracking-enhanced Forgetting-free Few-shot 3D LiDAR Semantic SegmentationabstractIn autonomous driving, 3D LiDAR plays a crucial role in understanding the vehicle’s surroundings. However, the newly emerged, unannotated objects presents few-shot learning problem for semantic segmentation. This paper addresses the limitations of current few-shot semantic segmentation by exploiting the temporal continuity of LiDAR data. Employing a tracking model to generate pseudo-ground-truths from a sequence of LiDAR frames, our method significantly augments the dataset, enhancing the model’s ability to learn on novel classes. However, this approach introduces a data imbalance biased to novel data that presents a new challenge of catastrophic forgetting. To mitigate this, we incorporate LoRA, a technique that reduces the number of trainable parameters, thereby preserving the model’s performance on base classes while improving its adaptability to novel classes. This work represents a significant step forward in few-shot 3D LiDAR semantic segmentation for autonomous driving. Our code is available at https://github.com/BowmanChow/Track-no-forgetting. Junbao Zhou, Jilin Mei, Pengze Wu, Fangzhou Zhao, Xijun Zhao, Yu Hu 0001 |
IROS | 7 |
| 2024 | SAM-PS: Zero-shot Parking-slot Detection based on Large Visual ModelabstractLarge visual models have recently demonstrated their promising performance on zero-shot transfer. However, so far, none of the existing methods explicitly possess the ability to perform zero-shot transfer on parking-slot detection, which results in current deep-learning based methods relying on training datasets, and methods based on traditional computer vision exhibiting poor robustness. In this paper, we propose a large visual model-based parking-slot detection method, which utilizes a large visual model (segment anything) to segment an around-view image and infer parking-slots by analyzing the relationship of marking-points in masks. In addition, we classify real-world parking-slots into two categories, line-based and area-based. The proposed method employs a two-stage approach which has a manually designed post-processing step without training. Multiple experiments have been carried out on public benchmarks, and our method demonstrates the capability for zero-shot transfer. The code will be released at https://github.com/Zhai0123/SAM-PS. Heng Zhai, Jilin Mei, Fangzhou Zhao, Xijun Zhao, Yu Hu 0001 |
IV | 6 |
| 2024 | FusionOcc: Multi-Modal Fusion for 3D Occupancy Predictionabstract3D occupancy prediction (OCC) aims to estimate and predict the semantic occupancy state of the surrounding environment, which is crucial for scene understanding and reconstruction in the real world. However, existing methods for 3D OCC mainly rely on surround-view camera images, whose performance is still insufficient in some challenging scenarios, such as low-light conditions. To this end, we propose a new multi-modal fusion network for 3D occupancy prediction by fusing features of LiDAR point clouds and surround-view images, called FusionOcc. Our model fuses features of these two modals in 2D and 3D space, respectively. By integrating the depth information from point clouds, a cross-modal fusion module is designed to predict a 2D dense depth map, enabling an accurate depth estimation and a better transition of 2D image features into 3D space. In addition, features of voxelized point clouds are aligned and merged with image features converted by a view-transformer in 3D space. Experiments show that FusionOcc establishes the new state of the art on Occ3D-nuScenes dataset, achieving a mIoU score of 35.94% (without visibility mask) and 56.62% (with visibility mask), showing an average improvement of 3.42% compared to the best previous method. Our work provides a new baseline for further research in multi-modal fusion for 3D occupancy prediction. Codes will be made publicly at https://github.com/ShuoZhang-code/FusionOcc. Yupeng Zhai, Jilin Mei, Yu Hu 0001 |
ACM Multimedia | 4 |
| 2024 | PHD-NAS: Preserving helpful data to promote Neural Architecture Search
Shun Lu 0001, Yu Hu 0001, Longxing Yang, Jilin Mei, Jianchao Tan, Chengru Song |
Neurocomputing | 2 |
| 2024 | Trajectory Planning for Autonomous Driving Featuring Time-Varying Road Curvature and Adhesion ConstraintsabstractAmong the various driving situations, there are challenging road conditions where both the texture and curvature are variables over time (e.g., mountainous area). However, it is found that the characteristics of road texture and curvature have been respectively considered in some of the existing studies to determine the vehicle speed for trajectory planning, but the complementary effect of these two factors is still yet to be incorporated. This could lead to unsafe vehicle behaviour. This limitation has led us to develop a trajectory planning method that gives a systematic consideration of road conditions and leverages the complementary effect of road curvature and adhesion on the vehicle speed. It prioritises the trajectory safety through a preview of road constraints (i.e., waypoints, curvature and adhesion) in a look-ahead distance and the real-time computation of the vehicle speed that satisfies the constraints. In the experiment, our method was compared with the state-of-the-art techniques in a simulated mountainous driving environment, namely Model Predictive Control (MPC), Deep Reinforcement Learning (DRL) and Hybrid A*. The environment was built with abundant variation in road curvature and adhesion. The results showed that our approach was able to generate safe and comfort trajectories in both sharp turn and ice-covered driving scenarios, in which the vehicle successfully passed through the whole length of the global path without producing large deviations and exceeding lane boundaries. Whereas, the MPC, DRL and Hybrid A* approaches resulted in the vehicle exceeding lanes at some point with completeness levels of 77.72%, 75.31% and 79.53%, respectively. Yifan Gao 0002, Wei Li 0235, Yu Hu 0001 |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2023 | PINAT: A Permutation INvariance Augmented Transformer for NAS PredictorabstractTime-consuming performance evaluation is the bottleneck of traditional Neural Architecture Search (NAS) methods. Predictor-based NAS can speed up performance evaluation by directly predicting performance, rather than training a large number of sub-models and then validating their performance. Most predictor-based NAS approaches use a proxy dataset to train model-based predictors efficiently but suffer from performance degradation and generalization problems. We attribute these problems to the poor abilities of existing predictors to character the sub-models' structure, specifically the topology information extraction and the node feature representation of the input graph data. To address these problems, we propose a Transformer-like NAS predictor PINAT, consisting of a Permutation INvariance Augmentation module serving as both token embedding layer and self-attention head, as well as a Laplacian matrix to be the positional encoding. Our design produces more representative features of the encoded architecture and outperforms state-of-the-art NAS predictors on six search spaces: NAS-Bench-101, NAS-Bench-201, DARTS, ProxylessNAS, PPI, and ModelNet. The code is available at https://github.com/ShunLu91/PINAT. Shun Lu 0001, Yu Hu 0001, Peihao Wang, Yan Han 0001, Jianchao Tan, Sen Yang 0004, Ji Liu 0002 |
AAAI | 2 |
| 2023 | PA&DA: Jointly Sampling PAth and DAta for Consistent NASabstractBased on the weight-sharing mechanism, one-shot NAS methods train a supernet and then inherit the pre-trained weights to evaluate sub-models, largely reducing the search cost. However, several works have pointed out that the shared weights suffer from different gradient descent directions during training. And we further find that large gradient variance occurs during supernet training, which degrades the supernet ranking consistency. To mitigate this issue, we propose to explicitly minimize the gradient variance of the supernet training by jointly optimizing the sampling distributions of PAth and DAta (PA&DA). We theoretically derive the relationship between the gradient variance and the sampling distributions, and reveal that the optimal sampling probability is proportional to the normalized gradient norm of path and training data. Hence, we use the normalized gradient norm as the importance indicator for path and training data, and adopt an importance sampling strategy for the supernet training. Our method only requires negligible computation cost for optimizing the sampling distributions of path and data, but achieves lower gradient variance during supernet training and better generalization performance for the supernet, resulting in a more consistent NAS. We conduct comprehensive comparisons with other improved approaches in various search spaces. Results show that our method surpasses others with more reliable ranking performance and higher accuracy of searched architectures, showing the effectiveness of our method. Code is available at https://github.com/ShunLu91/PA-DA. Shun Lu 0001, Yu Hu 0001, Longxing Yang, Jilin Mei, Jianchao Tan, Chengru Song |
CVPR | 2 |
| 2023 | Unleashing the Power of Gradient Signal-to-Noise Ratio for Zero-Shot NASabstractNeural Architecture Search (NAS) aims to automatically find optimal neural network architectures in an efficient way. Zero-Shot NAS is a promising technique that leverages proxies to predict the accuracy of candidate architectures without any training. However, we have observed that most existing proxies do not consistently perform well across different search spaces, and are less concerned with generalization. Recently, the gradient signal-to-noise ratio (GSNR) was shown to be correlated with neural network generalization performance. In this paper, we not only explicitly give the probability that larger GSNR at network initialization can ensure better generalization, but also theoretically prove that GSNR can ensure better convergence. Then we design the ξ-based gradient signal-to-noise ratio (ξ-GSNR) as a Zero-Shot NAS proxy to predict the network accuracy at initialization. Extensive experiments in different search spaces demonstrate that ξ-GSNR provides superior ranking consistency compared to previous proxies. Moreover, ξ-GSNR-based Zero-Shot NAS also achieves outstanding performance when directly searching for the optimal architecture in various search spaces and datasets. The source code is available at https://github.com/Sunzh1996/Xi-GSNR. Longxing Yang, Shun Lu 0001, Jilin Mei, Wen-Xiao Zhao, Yu Hu 0001 |
ICCV | 7 |
| 2023 | Zero-shot Object Detection Based on Dynamic Semantic VectorsabstractZero-shot object detection has shown its ability to overcome the problems of data scarcity and novel classes. Existing methods generally utilize static semantic vectors to classify objects and guide the network to map visual features to semantic vectors. However, the distribution of semantic vectors cannot adequately represent visual features, which makes migration from seen to unseen classes difficult. This work explores the dynamic semantic vector method to align the distributions of semantic vectors and visual features. The main challenge is to get a more reasonable distribution of semantic vectors. To address this issue, we proposed a two-way classification branch network and introduce N-pair loss into the dynamic semantic vector optimization process. Experiments on the MS-COCO dataset and SiTi (a real-world autonomous driving dataset collected by us) demonstrate the effectiveness and generalization of our method. Our code is available at https://github.com/HaoyuLizju/ZSD_tcb Jilin Mei, Jiancong Zhou, Yu Hu 0001 |
ICRA | 4 |
| 2023 | Few-shot 3D LiDAR Semantic Segmentation for Autonomous DrivingabstractIn autonomous driving, the novel objects and lack of annotations challenge the traditional 3D LiDAR semantic segmentation based on deep learning. Few-shot learning is a feasible way to solve these issues. However, currently few-shot semantic segmentation methods focus on camera data, and most of them only predict the novel classes without considering the base classes. This setting cannot be directly applied to autonomous driving due to safety concerns. Thus, we propose a few-shot 3D LiDAR semantic segmentation method that predicts both novel and base classes simultaneously. Our method tries to solve the background ambiguity problem in generalized few-shot semantic segmentation. We first review the original cross-entropy and knowledge distillation losses, then propose a new loss function that incorporates the background information to achieve 3D LiDAR few-shot semantic segmentation. Extensive experiments on SemanticKITTI demonstrate the effectiveness of our method. Jilin Mei, Junbao Zhou, Yu Hu 0001 |
ICRA | 3 |
| 2023 | Active Visual SLAM Based on Hierarchical Reinforcement LearningabstractWe present AVS-HRL, a modular Active visual SLAM system based on hierarchical reinforcement learning. The reward function explicitly considers the efficiency of exploration and the accuracy of mapping by utilizing the internal variables of SLAM, such as feature points distribution and loop-closure signal. Compared to end-to-end active SLAM methods, we designed a map reconstruction module that can correct the cumulative error in the incremental mapping process. Furthermore, the inputs of all neural network modules use more abstract and general information, such as grid maps, rather than raw sensor observations. We conducted experiments in two different simulators and real-world environments. In the noisy setting of Habitat environments, our method improves the accuracy of the mapped areas by 68.48% as an average of Gibson and MP3D datasets. Moreover, our method's generalization performance was demonstrated through direct transfer across different simulators and real-world environments. Wensong Chen, Wei Li 0235, Andong Yang, Yu Hu 0001 |
IROS | 4 |
| 2023 | Generalized Few-shot Semantic Segmentation for LiDAR Point CloudsabstractSemantic segmentation of LiDAR point clouds can provide assistance for precise perception in autonomous driving, but traditional segmentation methods face challenges such as unbalanced class distribution and insufficient labeling. Generalized few-shot learning has been researched on image data, but these methods are difficult to apply directly to LiDAR point clouds. To tackle these challenges, we propose a generalized few-shot semantic segmentation method based on LiDAR point cloud data, enabling us to predict base and novel classes simultaneously. To improve the performance with limited novel class samples, we integrate semantic vectors and leverage the intrinsic relationship between base and novel class vectors to facilitate learning. We conduct comprehensive comparisons with other methods on the SemanticKITTI and constantly surpass them with higher mIoU, demonstrating the effectiveness of our method. Pengze Wu, Jilin Mei, Xijun Zhao, Yu Hu 0001 |
IROS | 4 |
| 2023 | A Tightly-Coupled GNSS RTK/INS Positioning Algorithm Based on Adaptive Lag SmootherabstractHow to take into account both the calculation cost and positioning accuracy when driving over a long distance in the scene of changing satellite visibility, such as urban areas and mountain roads, is a research topic worth of attention for intelligent vehicles. In this paper, a tightly-coupled RTK/INS positioning algorithm base on adaptive lag smoother is proposed. By combining the uncertainty of the state to be estimated at the current time and the quantitative score of satellite visibility, the lag-length in smoother can be adjusted adaptively, so as to ensure positioning accuracy while reducing the calculation cost of marginal benefit consumption as much as possible. The proposed algorithm is demonstrated in both the simulator and real-world urban roads. From the experimental results, it is found that the estimation accuracy achieved by the adaptive lag smoother is similar to that of smoothers with long lag length, but the time consumption is reduced by about 30%. Under the same condition that the positioning can be completed in real time, the accuracy of the algorithm in this paper is 27% higher than that of the tightly-coupled RTK/INS system based on extended Kalman filter. Wei Li 0235, Yu Hu 0001 |
IV | 3 |
| 2023 | M2F2-Net: Multi-Modal Feature Fusion for Unstructured Off-Road Freespace DetectionabstractFreespace detection is an important part of autonomous driving technology. Compared with structured on-road scenes, unstructured off-road scenes face more challenges. Multi-modal fusion method is a viable solution to these challenges. But existing fusion methods do not fully utilize the multi-modal features. In this paper, we propose an effective multi-modal network named M2F2-Net for freespace detection in unstructured off-road scenes. We propose a multi-modal feature fusion strategy named Multi-modal Cross Fusion (MCF). MCF module is simple but effective in fusing the features of RGB images and surface normal maps. Meanwhile, a multi-modal segmentation decoder module is designed to decouple the segmentation of two modalities, and it further helps the features of both modalities to be fully utilized. In order to solve the problem that the road edge is difficult to extract in the unstructured scenes, we also propose an edge segmentation decoder module. Extensive experiments show that our approach can lead to significant improvements, which brings 6.1% F1 and 10.8% IoU improvements. Our code will be available at https://github.com/yhl1010/M2F2-Net. Hongliang Ye, Jilin Mei, Yu Hu 0001 |
IV | 3 |
| 2023 | PMR-CNN: Prototype Mixture R-CNN for Few-Shot Object DetectionabstractFew-shot object detection is a challenging task because of the limited annotation data. Under the limitation of few-shot samples, images from the same class may differ significantly in appearance and pose. Although the research has progressed considerably since adding the prototype vector to few-shot object detection, the previous paradigm is still constrained by several factors: (1) using a single prototype to represent the support image tends to cause semantic ambiguity; (2) the way of extracting prototypes is too simple, like global average pooling, which makes prototypes not representative enough. In this work, we design PMR-CNN to address the above limitations. PMR-CNN proposes a new method of prototype generation and enhances the representative information by using multiple prototypes to represent support images. For experiments, we not only evaluate our method on general image dataset MS COCO, but also evaluate on SiTi (a real-world autonomous driving dataset collected by us). Experiment on the few-shot object detection benchmark shows that we have a significant advantage over the previous methods. Code is available at: https://github.com/Chientsung-Chou/PMR-CNN. Jiancong Zhou, Jilin Mei, Yu Hu 0001 |
IV | 4 |
| 2023 | Sweet Gradient matters: Designing consistent and efficient estimator for Zero-shot Architecture Search
Longxing Yang, Yanxin Fu, Shun Lu 0001, Jilin Mei, Wen-Xiao Zhao, Yu Hu 0001 |
Neural Networks | 7 |
| 2022 | Repeatable Pattern Mining for Accurate Subtraction of Backgrounds with Waving Objects in Underwater VideosabstractThe success of advanced Background Subtraction (BGS) algorithms for dynamic backgrounds is mostly in land scenes such as those in CDNet benchmarks; few handle underwater scenes, since existing underwater video datasets are either in low resolution or with only static backgrounds. Consequently, the lack of reliable BGS support makes supervised Moving-Objects Segmentation (MOS) algorithms much harder to adapt to unknown underwater scenes because of the diversities of the aquatic environments. For example, those trained by the latest underwater image dataset, SUIM, are ineffective in the underwater videos of our experiments.The underwater waving objects (e.g., plants) often render existing BGS algorithms inaccurate due to three types of errors: (a) incompletely identified MOs (Moving Objects), (b) missing MOs, and (c) falsely identified MOs. In this paper, we propose a novel Clustering-Based Multi-State Background Representation (CBMSBR) model to learn and represent the repeatable patterns of waving movements in k background states (i.e., color ranges) per pixel, and thus accurately subtract the background waving objects to reduce these errors. In addition, we further develop a CBMSBR+ model to remove the more challenging background objects in unusually large magnitudes of wavings. Both models come from a basic observation: the video pixels in the waving zones repeatedly switch among multiple background states; e.g., a pixel switches among water state, plant 1 state, and plant 2 state. To test our proposed models, we create experiments using three types of challenging scenarios that each often covers at least two error types, i.e., the scattered MOs scenario covering (b) and (c), the crowded MOs scenario covering (a) - (c), and the slow MOs scenario covering (a) and (c). Experiments on these scenarios demonstrate the accuracy, effectiveness, and efficiency of our models and their applications in MOS improvements. Junfeng Wu 0010, Guangyan Huang, Hui Zheng 0001, Guang-Li Huang, Yu Hu 0001, Jing He 0004 |
DSAA | 5 |
| 2022 | AGNAS: Attention-Guided Micro and Macro-Architecture SearchabstractMicro- and macro-architecture search have emerged as two popular NAS paradigms recently. Existing methods leverage different search strategies for searching micro- and macro- architectures. When using architecture parameters to search for micro-structure such as normal cell and reduction cell, the architecture parameters can not fully reflect the corresponding operation importance. When searching for the macro-structure chained by pre-defined blocks, many sub-networks need to be sampled for evaluation, which is very time-consuming. To address the two issues, we propose a new search paradigm, that is, leverage the attention mechanism to guide the micro- and macro-architecture search, namely AGNAS. Specifically, we introduce an attention module and plug it behind each candidate operation or each candidate block. We utilize the attention weights to represent the importance of the relevant operations for the micro search or the importance of the relevant blocks for the macro search. Experimental results show that AGNAS can achieve 2.46% test error on CIFAR-10 in the DARTS search space, and 23.4% test error when directly searching on ImageNet in the ProxylessNAS search space. AGNAS also achieves optimal performance on NAS-Bench-201, outperforming state-of-the-art approaches. The source code can be available at https://github.com/Sunzh1996/AGNAS. Yu Hu 0001, Shun Lu 0001, Longxing Yang, Jilin Mei, Yinhe Han 0001, Xiaowei Li 0001 |
ICML | 2 |
| 2022 | Searching for BurgerFormer with Micro-Meso-Macro Space DesignabstractWith the success of Transformers in the computer vision field, the automated design of vision Transformers has attracted significant attention. Recently, MetaFormer found that simple average pooling can achieve impressive performance, which naturally raises the question of how to design a search space to search diverse and high-performance Transformer-like architectures. By revisiting typical search spaces, we design micro-meso-macro space to search for Transformer-like architectures, namely BurgerFormer. Micro, meso, and macro correspond to the granularity levels of operation, block and stage, respectively. At the microscopic level, we enrich the atomic operations to include various normalizations, activation functions, and basic operations (e.g., multi-head self attention, average pooling). At the mesoscopic level, a hamburger structure is searched out as the basic BurgerFormer block. At the macroscopic level, we search for the depth, width, and expansion ratio of the network based on the multi-stage architecture. Meanwhile, we propose a hybrid sampling method for effectively training the supernet. Experimental results demonstrate that the searched BurgerFormer architectures achieve comparable even superior performance compared with current state-of-the-art Transformers on the ImageNet and COCO datasets. The codes can be available at https://github.com/xingxing-123/BurgerFormer. Longxing Yang, Yu Hu 0001, Shun Lu 0001, Jilin Mei, Yinhe Han 0001, Xiaowei Li 0001 |
ICML | 2 |
| 2022 | Closing the Dynamics Gap via Adversarial and Reinforcement Learning for High-Speed RacingabstractAutonomous racing has lately gained popularity because of its entertainment value and potential of advancing autonomous driving in high-speed situations. These high-speed racing efforts usually focus on a road domain with fixed dynamics. They cannot meet the challenge of policy adaptation between domains with large dynamics gaps. Meanwhile, existing policy adaptation methods either rely on experts to build new environments for policy training, or only handle a small dynamics gap for low-speed control tasks due to limited dynamics modeling and rigorous data collection assumptions. To overcome these drawbacks, we introduce DAARL, a novel policy adaptation algorithm that uses adversarial and reinforcement learning to bridge the large dynamics gap between different domains. It has two training stages. In the first training stage, a domain transfer function is learned by adversarial learning to better capture the dynamics gap. The single domain transfer function integrates with the source domain to implement the dynamics of different target domains virtually without the help of experts. We name these virtual domains the imaginary target domains. In the second training stage, the knowledge of the source-domain policy guides the reinforcement learning of a target-domain policy on an imaginary target domain. It improves the convergence of the target-domain policy. Five experiments have been conducted on a racing simulator with different road domains. All results show that DAARL outperforms baselines in terms of driving speed, stability, success rate, and domain scalability. Jingyu Niu, Yu Hu 0001, Wei Li 0235, Guangyan Huang, Yinhe Han 0001, Xiaowei Li 0001 |
IJCNN | 2 |
| 2022 | SMS-MPC: Adversarial Learning-based Simultaneous Prediction Control with Single Model for Mobile RobotsabstractModel predictive control is a promising method in robot control tasks. How to design an effective model structure and efficient prediction framework for model predictive control is still an open challenge. To reduce the time consumption and avoid compounding-error of the multi-step prediction process in model predictive control, we propose a single-model simultaneous framework, which uses single dynamics model to predict the entire prediction horizon simultaneously by taking all control actions with the current state as inputs. Based on this framework, we further propose an adversarial dynamics model that contains two parts. The generator provides a dynamics model for the prediction process, while the discriminator provides constraints that are hard to describe by manually defined loss. This adversarial dynamics model can accelerate training and improve model accuracy in unstructured environments. Experiments conducted in Gazebo simulator and on a real mobile robot demonstrate the efficiency and accuracy of the single-model simultaneous framework with an adversarial dynamics model. Andong Yang, Wei Li 0235, Yu Hu 0001 |
IROS | 3 |
| 2022 | STC-NAS: Fast neural architecture search with source-target consistency
Yu Hu 0001, Longxing Yang, Shun Lu 0001, Jilin Mei, Yinhe Han 0001, Xiaowei Li 0001 |
Neurocomputing | 2 |
| 2021 | DDSAS: Dynamic and Differentiable Space-Architecture SearchabstractNeural Architecture Search (NAS) has made remarkable progress in automatically designing neural networks. However, existing differentiable NAS and stochastic NAS methods are either biased towards exploitation and thus may converge to a local minimum, or biased towards exploration and thus converge slowly. In this work, we propose a Dynamic and Differentiable Space-Architecture Search (DDSAS) method to address the exploration-exploitation dilemma. DDSAS dynamically samples space, searches architectures in the sampled subspace with gradient descent, and leverages the Upper Confidence Bound (UCB) to balance exploitation and exploration. The whole search space is elastic, offering flexibility to evolve and to consider resource constraints. Experiments on image classification datasets demonstrate that with only 4GB memory and 3 hours for searching, DDSAS achieves 2.39% test error on CIFAR10, 16.26% test error on CIFAR100, and 23.9% test error when transferring to ImageNet. When directly searching on ImageNet, DDSAS achieves comparable accuracy with more than 6.5 times speedup over state-of-the-art methods. The source codes are available at https://github.com/xingxing-123/DDSAS. Longxing Yang, Yu Hu 0001, Shun Lu 0001, Jilin Mei, Yiming Zeng 0003, Zhi-Ping Shi 0002, Yinhe Han 0001, Xiaowei Li 0001 |
ACML | 2 |
| 2021 | DU-DARTS: Decreasing the Uncertainty of Differentiable Architecture Search
Shun Lu 0001, Yu Hu 0001, Longxing Yang, Jilin Mei, Yiming Zeng 0003, Xiaowei Li 0001 |
BMVC | 2 |
| 2021 | KFS-LIO: Key-Feature Selection for Lightweight Lidar Inertial OdometryabstractFeature-based lidar odometry methods have attracted increasing attention due to their low computational cost. However, theoretically analysis of the effect of extracted features on pose estimation is still lacked. In this paper, we propose a method of key-feature selection for lightweight lidar inertial odometry, KFS-LIO, to further enhance the real-time performance by selecting the most effective subset of lidar feature constraints. Aiming at explaining the correlation between the feature distribution and state errors, a quantitative evaluation method of lidar constraints is introduced. In addition, to avoid recalculating the reprojection matrices in de-skewing step, we use the intermediate variables in IMU preintegration to compensate for lidar motion distortion. The experimental results demonstrate that KFS-LIO can reduce half of the LOAM features and provide comparable accuracy with the state-of-the-art odometry. Wei Li 0235, Yu Hu 0001, Yinhe Han 0001, Xiaowei Li 0001 |
ICRA | 2 |
| 2020 | Survey: Hardware Trojan Detection for NetlistabstractThe development of integrated circuit technology is accompanied by potential threats. Malicious modifications to circuits, known as hardware Trojans, are major security concerns. This paper gives a survey of hardware Trojan detection methods towards gate-level netlists. The detection methods are divided into search-based, threshold-based, and machine learning-based ones. This paper compares and analyzes existing works from aspects of feature selection, data balancing techniques, classification criterion, detection range. The experimental results are also selected for comparison. Yipei Yang, Jing Ye 0001, Yuan Cao 0003, Jiliang Zhang 0002, Xiaowei Li 0001, Huawei Li 0001, Yu Hu 0001 |
ATS | 7 |
| 2020 | Exploring Spatial-Temporal Multi-Frequency Analysis for High-Fidelity and Temporal-Consistency Video PredictionabstractVideo prediction is a pixel-wise dense prediction task to infer future frames based on past frames. Missing appearance details and motion blur are still two major problems for current models, leading to image distortion and temporal inconsistency. We point out the necessity of exploring multi-frequency analysis to deal with the two problems. Inspired by the frequency band decomposition characteristic of Human Vision System (HVS), we propose a video prediction network based on multi-level wavelet analysis to uniformly deal with spatial and temporal information. Specifically, multi-level spatial discrete wavelet transform decomposes each video frame into anisotropic sub-bands with multiple frequencies, helping to enrich structural information and reserve fine details. On the other hand, multilevel temporal discrete wavelet transform which operates on time axis decomposes the frame sequence into sub-band groups of different frequencies to accurately capture multifrequency motions under a fixed frame rate. Extensive experiments on diverse datasets demonstrate that our model shows significant improvements on fidelity and temporal consistency over the state-of-the-art works. Source code and videos are available at https://github.com/Bei-Jin/STMFANet. Beibei Jin, Yu Hu 0001, Qiankun Tang, Jingyu Niu, Zhi-Ping Shi 0002, Yinhe Han 0001, Xiaowei Li 0001 |
CVPR | 2 |
| 2020 | Prediction Stability: A New Metric for Quantitatively Evaluating DNN OutputsabstractIn many realistic applications, the collected inputs of DNN face a big challenge: perturbations. Although the perturbations are imperceptible, they may cause incorrect prediction results. This paper proposes prediction stability to quantitatively evaluate whether the prediction result of an input is instable and easy to be perturbed. Prediction stability can guide the DNN system to cope with the situation where the prediction result has a high confidence but with a low stability. Experimental result shows that, using the proposed metrics to evaluate the stability of prediction results, over 99.8 cases are consistent with the real stable/instable conditions. Qingli Guo, Jing Ye 0001, Jiliang Zhang 0002, Yu Hu 0001, Xiaowei Li 0001, Huawei Li 0001 |
ACM Great Lakes Symposium on VLSI | 4 |
| 2020 | Lightdet: A Lightweight and Accurate Object Detection NetworkabstractThe extensive computational burden limits the usage of accurate but complex object detectors in resource-bounded scenarios. In this paper, we present a lightweight object detector, named LightDet, to address this dilemma. We design a lightweight backbone that is able to capture rich low-level features by the proposed Detail-Preserving Module. To effectively aggregate bottom and top-down features, we introduce an efficient Feature-Preserving and Refinement Module. A lightweight prediction head is employed to further reduce the entire network complexity. Experimental results show that our LightDet achieves 75.5% mAP on PASCAL VOC 2007 at the speed of 250 FPS and 24.0% mAP on MS COCO dataset. Qiankun Tang, Zhi-Ping Shi 0002, Yu Hu 0001 |
ICASSP | 4 |
| 2020 | Two-Stage Safe Reinforcement Learning for High-Speed Autonomous RacingabstractDecision making for autonomous driving is a safety-critical control problem. Prior works of safe reinforcement learning either tackle the problem with reward shaping or with modifying the reinforcement learning exploration process. However, the former cannot guarantee the safety during the learning process, while the latter relies heavily on expertise to design exquisite exploration policy. Currently, only short-term decision makings for low-speed driving were achieved in road scenes with basic geometries. In this paper, we propose a two-stage safe reinforcement learning algorithm to automatically learn a long-term policy for high-speed driving that guarantees safety during the entire training. In the first learning stage, model-free reinforcement learning is followed by a rule-based safeguard module to avoid danger at low speed without expert ne-tuning. In the second learning stage, the rule-based module is replaced with a data-driven counterpart to develop a closed-form analytical safety solution for high-speed driving. Moreover, an adaptive reward function is designed to match the different objectives of the two learning stages for faster convergence to an optimal policy. Experiments are conducted on a racing simulator TORCS which has complex racing tracks (e.g. sharp turns, hills). Compared with the state-of-the-art baselines, the results show that our method achieves zero safety violation and quickly converges to a more efficient and stable policy with an average speed of 127 km/h (3.3% higher than the best result of baselines) and an average swing of 3.96 degrees. Jingyu Niu, Yu Hu 0001, Beibei Jin, Yinhe Han 0001, Xiaowei Li 0001 |
SMC | 2 |
| 2020 | Sequence Triggered Hardware Trojan in Neural Network AcceleratorabstractWith the rapid development of deep learning techniques, the security issue for Neural Network (NN) systems has emerged as an urgent and severe problem. Hardware Trojan attack is one of the threatens, which provides attackers backdoors to control the prediction results of NN systems. This paper proposes a sequence triggered hardware Trojan. Normal images but with specific sequence are used to trigger the hardware Trojan and let attackers fully control the prediction results. This kind of trigger is not only robust to image pre-processing, but also unrecognizable by human beings. In comparison with existing hardware Trojan design, it is more practical and less hardware overhead. The experiments on MNIST, CIFAR100, and ISLVRC show that the proposed hardware Trojan is rarely triggered in normal working status while the hardware cost is reduced by 19X. Zizhen Liu, Jing Ye 0001, Xing Hu 0001, Huawei Li 0001, Xiaowei Li 0001, Yu Hu 0001 |
VTS | 6 |
| 2020 | INOR - An Intelligent noise reduction method to defend against adversarial audio examples
Qingli Guo, Jing Ye 0001, Yiran Chen 0001, Yu Hu 0001, Yazhu Lan, Guohe Zhang, Xiaowei Li 0001 |
Neurocomputing | 4 |
| 2019 | Implementation of Parametric Hardware Trojan in FPGAabstractThe reconfigurability of FPGA makes it flexible for different applications. However, an FPGA may be delivered, designed, and deployed by different persons during its lifecycle, so anyone who can access the FPGA may bring in security issues. This paper proposes an implementation method of a parametric hardware Trojan in the FPGA. This hardware Trojan does not add any extra circuits, so many existing detection methods based on analyzing the design files are invalid. Yipei Yang, Jing Ye 0001, Xiaowei Li 0001, Yinhe Han 0001, Huawei Li 0001, Yu Hu 0001 |
ITC-Asia | 6 |
| 2019 | PUFPass: A password management mechanism based on software/hardware codesign
Qingli Guo, Jing Ye 0001, Bing Li 0017, Yu Hu 0001, Xiaowei Li 0001, Yazhu Lan, Guohe Zhang |
Integr. | 4 |
| 2018 | PUF Based Pay-Per-Device Scheme for IP Protection of CNN ModelabstractWith great success of Convolutional Neural Network (CNN) in many applications, it is not surprising that the CNN models will become commercial IPs. This paper proposes a Physical Unclonable Function (PUF) based pay-per-device scheme for protecting IPs of CNN models. PUFs are embedded into the FPGA based CNN accelerator. The original CNN model trained by the IP vendor is obfuscated based on the PUFs before being distributed to the end users. The PUF challenges come from obfuscated CNN model parameters, and the PUF responses determine outputs of convolutional layers. In this way, the obfuscated CNN model is limited to be correctly executed in one specific FPGA. Experiments on AlexNet show that performance and hardware overhead of the CNN accelerator are negligible. For authorized end users, the prediction accuracy of the obfuscated CNN model is the same as that of the original one, while for adversaries, prediction accuracies of guessed ones are nearly 0. Qingli Guo, Jing Ye 0001, Yu Hu 0001, Xiaowei Li 0001 |
ATS | 4 |
| 2018 | Hardware Trojan in FPGA CNN AcceleratorabstractMaliciously manipulating prediction results of Convolutional Neural Network (CNN) is a severe security threat. Previous works studied this threat from the aspects of dataset and model. However, with the increasing developments of CNN accelerators nowadays, the role of hardware in this threat lacks attentions. This paper inserts a hardware Trojan into the convolutional operations of a FPGA CNN accelerator. The experiments on ImageNet show that, with only 0.0051% hardware overhead to the accelerator and 0.000356% modification to an image, the hardware Trojan can be triggered to 100% precisely control the CNN classification result of the image. Jing Ye 0001, Yu Hu 0001, Xiaowei Li 0001 |
ATS | 2 |
| 2018 | VarNet: Exploring Variations for Unsupervised Video PredictionabstractUnsupervised video prediction is a very challenging task due to the complexity and diversity in natural scenes. Prior works directly predicting pixels or optical flows either have the blurring problem or require additional assumptions. We highlight that the crux for video frame prediction lies in precisely capturing the inter-frame variations which encompass the movement of objects and the evolution of the surrounding environment. We then present an unsupervised video prediction framework - Variation Network (VarNet) to directly predict the variations between adjacent frames which are then fused with current frame to generate the future frame. In addition, we propose an adaptively re-weighting mechanism for loss function to offer each pixel a fair weight according to the amplitude of its variation. Extensive experiments for both short-term and long-term video prediction are implemented on two advanced datasets - KTH and KITTI with two evaluating metrics - PSNR and SSIM. For the KTH dataset, the VarNet outperforms the state-of-the-art works up to 11.9% on PSNR and 9.5% on SSIM. As for the KITTI dataset, the performance boosts are up to 55.1% on PSNR and 15.9% on SSIM. Moreover, we verify that the generalization ability of our model excels other state-of-the-art methods by testing on the unseen CalTech Pedestrian dataset after being trained on the KITTI dataset. Source code and video are available at https://github.com/jinbeibei/VarNet. Beibei Jin, Yu Hu 0001, Yiming Zeng 0003, Qiankun Tang, Shice Liu, Jing Ye 0001 |
IROS | 2 |
| 2018 | Grey Zone in Pre-Silicon Hardware Trojan DetectionabstractPre-Silicon hardware Trojan detection has been studied for years. The most popular benchmark circuits are from the Trust-Hub. Their common feature is that the probability of activating hardware Trojans is very low. This leads to a series of machine learning based hardware Trojan detection methods which try to find the nets with low signal probability of 0 or 1. On the other hand, it is considered that, if the probability of activating hardware Trojans is high, these hardware Trojans can be easily found through behaviour simulations or during functional test. This paper explores the "grey zone" between these two opposite scenarios: if the activation probability of a hardware Trojan is not low enough for machine learning to detect it and is not high enough for behaviour simulation or functional test to find it, it can escape from detection. Experiments show the existence of such hardware Trojans, and this paper suggests a new set of hardware Trojan benchmark circuits for future study. Jing Ye 0001, Yipei Yang, Yu Hu 0001, Xiaowei Li 0001 |
ITC-Asia | 4 |
| 2018 | See and Think: Disentangling Semantic Scene CompletionabstractSemantic scene completion predicts volumetric occupancy and object category of a 3D scene, which helps intelligent agents to understand and interact with the surroundings. In this work, we propose a disentangled framework, sequentially carrying out 2D semantic segmentation, 2D-3D reprojection and 3D semantic scene completion. This three-stage framework has three advantages: (1) explicit semantic segmentation significantly boosts performance; (2) flexible fusion ways of sensor data bring good extensibility; (3) progress in any subtask will promote the holistic performance. Experimental results show that regardless of inputing a single depth or RGB-D, our framework can generate high-quality semantic scene completion, and outperforms state-of-the-art approaches on both synthetic and real datasets. Shice Liu, Yu Hu 0001, Yiming Zeng 0003, Qiankun Tang, Beibei Jin, Yinhe Han 0001, Xiaowei Li 0001 |
NeurIPS | 2 |
| 2018 | Modeling attacks on strong physical unclonable functions strengthened by random number and weak PUFabstractPhysical Unclonable Function (PUF) is a promising hardware security primitive. One important category of PUFs is the strong PUF with numerous Challenge-Response Pairs (CRPs). Since the typical strong PUFs, the arbiter PUF and several its variants, were broken by modeling attacks, many new designs for resisting modeling attacks have been proposed. Do they really achieve their promise, or are they only another pipe dream? This paper targets two PUF designs: the randomized PUF and the obfuscation PUF, which strengthen the arbiter PUF by leveraging the random number and the weak PUF, respectively. A heuristic algorithm is proposed for attacking these PUFs. The algorithm is implemented in CUDA. Some PUFs that cannot be broken in several months by CPU show their vulnerabilities in days by leveraging the GPU acceleration. The experimental results show that, for certain scales of objective PUFs, the prediction accuracy is beyond the reliability of CRPs, indicating successful attacks. Jing Ye 0001, Qingli Guo, Yu Hu 0001, Huawei Li 0001, Xiaowei Li 0001 |
VTS | 3 |
| 2018 | Deterministic and Probabilistic Diagnostic Challenge Generation for Arbiter Physical Unclonable FunctionabstractPhysical unclonable functions (PUFs) have broad application prospects in the field of hardware security. Like faults in general-purpose circuits, faults may also occur in PUFs. Fault diagnosis plays an important role in the yield learning process. Traditional fault diagnosis methods are based on comparing the fault-free responses of a design and the failing responses of chips. However, different manufactured, fault-free PUFs with the same design have different challenge-response pairs, so PUFs do not have deterministic, fault-free responses. Hence, traditional fault diagnosis methods are unsuitable for PUFs. To effectively diagnose PUFs, this paper proposes a diagnostic challenge generation method for the typical PUF: arbiter PUF. The diagnostic challenges that can deterministically or probabilistically distinguish the suspect faults of arbiter PUFs are generated. Simulation experiments on diagnosing failing arbiter PUF instances show that all the actual fault locations are accurately included in the candidate sets, and the average number of candidate locations (i.e., diagnostic resolution) is 1.585. FPGA experiments on diagnosing real PUFs show that the diagnostic accuracy is also 1, and the average diagnostic resolution is 1.602. Jing Ye 0001, Qingli Guo, Yu Hu 0001, Xiaowei Li 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2017 | Fault diagnosis of arbiter physical unclonable functionabstractPhysical Unclonable Function (PUF) has broad application prospects in the field of hardware security. If faults happen in PUF during manufacturing, the security of whole chip will be threatened. Fault diagnosis plays an important role in the yield learning process. However, since different manufactured PUFs with the same design have different Challenge-Response Pairs (CRPs), which cannot be predicted, the traditional fault diagnosis method based on comparing the fault-free responses of a design and the failing responses of chips is no longer suitable for diagnosing PUF. Therefore, this paper proposes a fault diagnosis method toward classic arbiter PUF. The stuck-at faults and the delay faults are considered. Based on the expected uniformity of arbiter PUF, a diagnostic challenge generation method and a corresponding CRP analysis method are proposed to distinguish faults within the arbiter PUF. Experimental results show that the diagnostic accuracy achieves 100.0% with good diagnostic resolution. Jing Ye 0001, Qingli Quo, Yu Hu 0001, Xiaowei Li 0001 |
DATE | 3 |
| 2017 | Leveraging FVT-margins in design space exploration for FFGA-based CNN acceleratorsabstractThe performance of an FPGA based CNN accelerator is determined by both parallelism and frequency, however, most prior works optimize the parallelism in the RTL design and resolve the frequency after the synthesis. This paper presents a design space exploration method for the pipeline implementation of the deep CNN models, which concurrently optimizes parallelism and frequency to achieve a comprehensive optimization on throughput. In addition to the quantitative modeling on parallelism, the maximum achievable system frequency under various parallelism is explored to leverage the PVT-margins in real-life scenarios and is adopted to guide the design space exploration for further performance boost. A case study of the AlexNet model is implemented using the proposed method on the Altera DE5a-Net board. The experimental results demonstrate that our method can achieve the throughput up to 906.25GOP/s, which gains 1.39× improvement compared to state-of-the-art RTL optimization methods. Weina Lu, Wenyan Lu, Jing Ye 0001, Yu Hu 0001, Xiaowei Li 0001 |
FPL | 4 |
| 2017 | Polymorphic PUF: Exploiting reconfigurability of CPU+FPGA SoC to resist modeling attackabstractPhysical Unclonable Function (PUF) is severely threatened by modeling attacks. This paper proposes a novel Polymorphic PUF for CPU+FPGA SoC. We fully exploit the dynamic reconfigurability of the SoC to minimize the Challenge Response Pair (CRP) correlation so as to resist modeling attacks. An asymmetric RO pair is proposed to produce the response. Experiments on real CPU+FPGA SoCs show the high resistance of Polymorphic PUF against modeling attacks, with good uniformity and uniqueness. Jing Ye 0001, Yu Hu 0001, Xiaowei Li 0001 |
IOLTS | 3 |
| 2017 | VPUF: Voter based physical unclonable function with high reliability and modeling attack resistanceabstractPhysical Unclonable Function (PUF) has broad application prospects in the field of hardware security. Arbiter PUF is a typical PUF, but is threatened by modeling attacks. To resist attack, XOR arbiter PUF employs multiple basic arbiter PUFs and XOR their response bits to generate the final response bit. However, its low reliability not only limits its applications, but also leaks information to enhance modeling attacks. To improve both the reliability and the modeling attack resistance, we propose the Voter based PUF (VPUF), which also employs multiple basic arbiter PUFs. It has two key components: (1) an on-line reliability checker to evaluate the reliability level of each internal response bit produced by each basic arbiter PUF; (2) a weighted voter, instead of XOR gates, to produce the final response bit. Experiments in FPGAs show 7.6%~23.4% reliability improvement of the VPUF than the XOR arbiter PUF, and prove the VPUF can resist modeling attacks. Jing Ye 0001, Yu Hu 0001, Xiaowei Li 0001 |
IOLTS | 2 |
| 2017 | GeoCueDepth: Exploiting geometric structure cues to estimate depth from a single imageabstractDepth estimation from a single image is very challenging due to the inherent ambiguity of mapping a color image to a depth map. Previous work tackles this problem by exploiting various levels of features with multi-scale deep convolutional neural networks. However, most of the local geometric structure related monocular depth cues are lost when being propagated through convolutional neural network. Moreover, the error of depth cues related to local geometric structures is not considered in the loss function. In this work, we propose the GeoCueDepth convolutional neural network to exploit local geometric structure cues and propose a training loss that takes the geometric error into consideration, which significantly improve the performance of depth prediction in both accuracy and sharpness. Experiments show that the proposed method achieves 0.122 average relative error and 0.078 square relative error on the NYU Depth v2 data set, which outperforms state-of-the-art monocular depth estimation approaches. Yiming Zeng 0003, Yu Hu 0001, Shice Liu, Qiankun Tang, Jing Ye 0001, Xiaowei Li 0001 |
IROS | 2 |
| 2017 | Power-Utility-Driven Write Management for MLC PCMabstractPhase change memory (PCM) is a promising alternative to Dynamic Random Access Memory (DRAM) as main memory due to its merits of high density and low leakage power. Multi-level Cell (MLC) PCM is more attractive than Single-level Cell (SLC) PCM, because it can store multiple bits per cell to achieve higher density and lower per-bit cost. With the iterative program-verify write technique, MLC PCM writes demand at much higher power than DRAM writes, while the power supply system of MLC memory system is similar to that of DRAM, and the power capability is limited. The incompatibility of high write power and limited power budget results in the degradation of the write throughput and performance in MLC PCM. In this work, we investigate both write scheduling policy and power management to improve the MLC power utility and alleviate the negative impacts induced by high write power. We identify the power-utility-driven write scheduling as an online bin-packing problem and then derive a power-utility-driven scheduling (PUDS) policy from the First Fit algorithm to improve the write power usage. Based on the ramp-down characteristic of the SET pulse (the pulse changes the PCM to high resistance), we propose the SET Power Amortization (SPA) policy, which proactively reclaims the power tokens at the intra-SET level to promote the power utilization. Our experimental results demonstrate that the PUDS and SPA respectively achieve 24% and 27% performance improvement over the state-of-the-art power management technique, and the PUDS8SPA has an overall 31% improvement of the power utility and 50% increase of performance compared to the baseline system. Bing Li 0017, Yu Hu 0001, Ying Wang 0001, Jing Ye 0001, Xiaowei Li 0001 |
ACM J. Emerg. Technol. Comput. Syst. | 2 |
| 2017 | Going Cooler With Timing-Constrained TeSHoP: A Temperature Sensing-Based Hotspot-Driven Placement Technique for FPGAsabstractThe continuous shrinking of the feature size in CMOS technology has significantly increased the power densities of integrated circuits, leading to severe temperature issues. However, the previous offline simulation-based thermal optimization works cast large deviations with the reality, while online sensing-based thermal managements usually incur significant performance overhead. Therefore, it is crucial to propose a method that could achieve fine-grained optimization with accurate temperature profiles. In this paper, we propose a timing-constraint temperature sensing-based hotspot-driven placement technique for field-programmable gate arrays (FPGAs). The hotspot optimization issue is modeled as a hyper minimum bipartite matching problem and is solved by a place adjustment with the input of an online sensed temperature profile. We propose an open-source/commercial hybrid design flow to implement the whole optimization in Xilinx Virtex-6 FPGA. Experimental results demonstrate a significant reduction in peak temperature and a great improvement on thermal uniformity, with slight performance overhead under timing constraints. Weina Lu, Yu Hu 0001, Jing Ye 0001, Xiaowei Li 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2016 | Efficient Attack on Non-linear Current Mirror PUF with Genetic AlgorithmabstractPhysical Unclonable Function (PUF) is a new hardware security primitive that exploits the manufacturing variations of integrated circuits. Traditional arbiter PUF is vulnerable to machine learning based modeling attacks due to its linearity. Current mirror PUF uses non-linear current mirror to bring non-linearity into the challenge-response relationship and is claimed resistant to modeling attacks. This paper further tests its security, and proves that the current mirror PUF is not as secure as claimed. A genetic algorithm based method is proposed to attack the current mirror PUF. By modeling the relationship between the output current and the input current of each current mirror, and fitting the model using genetic algorithm, we are able to predict the responses of current mirror PUF. Experiments prove that the prediction accuracy towards current mirror PUF is up to 99.27%. Qingli Guo, Jing Ye 0001, Yu Hu 0001, Xiaowei Li 0001 |
ATS | 4 |
| 2016 | POSTER: Attack on Non-Linear Physical Unclonable FunctionabstractPhysical Unclonable Function (PUF) is a promising hardware security primitive with broad application prospect. However, the strong PUF with numerous Challenge and Response Pairs (CRPs), e.g. the arbiter PUF, is vulnerable to modeling attacks. There are two major kinds of countermeasures. One is restricting CRP access interface, such as controlled PUF and XOR arbiter PUF, which unfortunately has been broken with the help of side-channels. The other is using non-linear electronic characteristics to produce CRPs, such as the current mirror PUF and the voltage transfer PUF. They are only proved to be resistant to SVM based attack, while no more analysis is further explored so far. In this paper, we propose an attack method based on compound heuristic algorithms of evolution strategy, simulated annealing, and ant colony to efficiently attack these two non-linear PUFs. This paper reveals that current mirror and voltage transfer are still not able to help strong PUF resist attacks. Our experimental results show that the average CRP prediction accuracy is as high as 99%. Jing Ye 0001, Yu Hu 0001, Xiaowei Li 0001 |
CCS | 2 |
| 2016 | DCPUF: Placement and Routing Constraint based Dynamically Configured Physical Unclonable Function on FPGA (Abstact Only)abstractWith the development of Integrated Circuit (IC), it is a growing trend that the CPU and the FPGA are integrated into one chip. To improve the security of CPU+FPGA IC, we explore the reconfigurable feature of FPGA to implement a novel Dynamically Configured Physical Unclonable Function (DCPUF). PUF is a hardware security primitive that utilizes unpredictable process variations to produce particular challenge-response pairs, so even the chips with the same design would produce different responses for the same challenge. In the DCPUF, the FPGA configuration bits, which are specifically designed with dedicated placement and routing constraint, constitute the challenge. When a challenge is input to a CPU+FPGA IC, the CPU uses it to configure or partially configure the FPGA, and then waits for the FPGA to reply a response. In comparison with existing PUFs, the DCPUF has three major advantages: (1) different from existing PUFs with fixed designs, the logic of DCPUF is dynamically configured for each challenge, i.e. the circuits for producing different responses are different, leading to higher security; (2) much more electronic parameters affected by process variation are leveraged to make DCPUF more robust against attacks; (3) for CPU+FPGA IC, no extra hardware is needed. The experiments on real CPU+FPGA ICs show the proposed DCPUF keeps good randomness and stability. Jing Ye 0001, Yu Hu 0001, Xiaowei Li 0001 |
FPGA | 2 |
| 2016 | TeSHoP: A Temperature Sensing based Hotspot-Driven Placement technique for FPGAsabstractThe rapid shrinking of the feature size in CMOS technology has significantly increased the power density of integrated circuits, leading to excessive temperature. Though online thermal management techniques such as DVFS and task migration can mitigate the temperature issue, but usually incur significant performance penalty. Therefore, it is crucial to optimize temperature at the design stage. In this work, we propose the TeSHoP, a Temperature Sensing based Hotspot-Driven Placement technique for FPGAs. Firstly, the un-optimized circuit along with a sensor network is run in FPGA to obtain the real temperature profile of the circuit. Then, based on the temperature profile, we proceed a one-off adjustment of the circuit placement for hotspot optimization. The optimization is modeled as a Hyper Minimum Bipartite Matching problem for solving. We implement the whole optimization flow in a real FPGA, with extension of the VTR-to-Bitstream tool. Experimental results on Xilinx Virtex-6 FPGA show that the reduction of peak temperature and the improvement of thermal uniformity can be up to 7.5°C and 13.9% respectively. Weina Lu, Yu Hu 0001, Jing Ye 0001, Xiaowei Li 0001 |
FPL | 2 |
| 2016 | TSocket: Thermal Sustainable Power BudgetingabstractAs technology scales, thermal management for multicore architectures becomes a critical challenge due to increasing power density. Existing power budgeting techniques focus on maximizing performance under a given power budget by optimizing the core configurations. In multicore era, a chip-wide power budget, however, is not sufficient to ensure thermal constraints because the thermal sustainable power capacity varies with different threading strategies and core configurations. In this article, we propose two models to dynamically estimate the thermal sustainable power capacity in homogeneous multicore systems: uniform power model and nonuniform power model . These two models convert the thermal effect of threading strategies and core configurations into power capacity, which provide a context-based core power capacity for power budgeting. Based on these models, we introduce a power budgeting framework aiming to improve the performance within thermal constraints, named as TSocket. Compared to the chip-wide power budgeting solution, TSocket shows 19% average performance improvement for the PARSEC benchmarks in single program scenario and up to 11% performance improvement in multiprogram scenario. The performance improvement is achieved by reducing thermal violations and exploring thermal headrooms. Yi Xu 0010, Xing Hu 0001, Xiangyang Guo, Yu Hu 0001, Yuan Xie 0001 |
ACM Trans. Design Autom. Electr. Syst. | 6 |
| 2015 | Impact assessment of net metering on smart home cyberattack detectionabstractDespite the increasing popularity of the smart home concept, such a technology is vulnerable to various security threats such as pricing cyberattacks. There are some technical advances in developing detection and defense frameworks against those pricing cyberattacks. However, none of them considers the impact of net metering, which allows the customers to sell the excessively generated renewable energy back to the grid. At a superficial glance, net metering seems to be irrelevant to the cybersecurity, while this paper demonstrates that its implication is actually profound. Yang Liu 0064, Shiyan Hu 0001, Jie Wu 0023, Yiyu Shi 0001, Yier Jin, Yu Hu 0001, Xiaowei Li 0001 |
DAC | 6 |
| 2015 | OPUF: Obfuscation logic based physical unclonable functionabstractThe Physical Unclonable Function (PUF) has broad application prospects in the field of hardware security. The arbiter PUF is a typical kind of strong PUF. However, due to its deterministic logic, attackers can use modeling techniques to break it in short time. Therefore, this paper proposes an Obfuscation logic based PUF (OPUF) design. A Boolean obfuscation module is proposed to obfuscate the logic which is employed to select the path segments in the arbiter PUF. In this way, the nondeterminacy of PUF is improved, and the computation complexities of modeling attacks are significantly increased, making the OPUF much safer against modeling attack. Both the theoretical analysis and the experimental results show the proposed OPUF design has good stability and randomness. Jing Ye 0001, Yu Hu 0001, Xiaowei Li 0001 |
IOLTS | 2 |
| 2015 | Diagnosis and Layout Aware (DLA) Scan Chain StitchingabstractWithout appropriate stitching of scan chains, even with good diagnosis algorithm and diagnostic pattern generation, the chain diagnostic resolution may still be bad. In this paper, we propose a novel pattern-independent diagnosis and layout aware (DLA) scan chain stitching method: 1) the resolution is improved by increasing and properly distributing the sensitive scan cells, which can capture useful diagnostic information under both single- and multiple-fault situations; and 2) the scan cell layout placement is taken into account to reduce routing overhead and hence preserve the chip performance. Experiments using two different techniques to diagnose ISCAS'89/ITC'99 benchmark circuits with/without embedded scan compaction show the effectiveness of the proposed method in improving the diagnostic resolution. Impacts on chip performance, embedded scan compaction, transition fault coverage, and test power dissipation are negligible. The proposed method is also successfully applied to an industry circuit manufactured with 20-nm technology. The silicon results show 7× average resolution improvement comparing to without using the DLA scan chain stitching. Jing Ye 0001, Yu Huang 0005, Yu Hu 0001, Wu-Tung Cheng, Ruifeng Guo, Liyang Lai, Ting-Pu Tai, Xiaowei Li 0001, Wei-pin Changchien, Daw-Ming Lee, Ji-Jan Chen, Sandeep C. Eruvathi, Kartik K. Kumara, Charles C. C. Liu, Sam Pan |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2014 | SwimmingLane: A composite approach to mitigate voltage droop effects in 3D power delivery networkabstractOne of the design challenges for the emerging 3D ICs is the power integrity. With multiple dies stacked vertically, the voltage droop may result in severe power integrity issues. In this paper, we first analyze the impact of application behaviors on voltage droop in a 3D power supply network (PDN) and observe that voltage droop is extremely imbalanced either across different layers or among the cores in the same layer. Based on the observation, we propose Swimming Lane, a hardware/software co-design method with two key schemes: (1) Mitigating the interference among different dies via a layer-independent scheme, and (2) balancing the intra-layer voltage droop and reducing the worst-case margin via OS scheduling. Compared to conventional designs, our method can reduce power consumption by 18%, worst-case voltage droops by 13%, and the number of voltage violations by 40%. Xing Hu 0001, Yi Xu 0010, Yu Hu 0001, Yuan Xie 0001 |
ASP-DAC | 3 |
| 2014 | Thermal-Sustainable Power Budgeting for Dynamic ThreadingabstractAs technology scales, thermal management for multi-core architectures becomes a critical challenge due to increased power density and higher integration density. Existing power budgeting techniques focus on maximizing performance under a given power budget by optimizing the core dynamics. However, in multi-core era, a chip-wide power budget is not sufficient to ensure thermal constraints because the thermal sustainable power capacity varies with different threading strategies and core configurations. In this paper, we propose a model which estimates the thermal sustainable power capacity considering these two run-time factors. The model converts the thermal effect of threading strategies and core configurations into power capacity, which provides a context-based power budget for the power budgeting. Based on this model, we introduce a power budgeting framework aiming to optimize the performance within thermal constraints, named as TSocket. Compared to the chip-wide power budgeting solution, TSocket shows 19% of performance improvement for the PARSEC benchmarks by reducing thermal violations and providing extra power budget for performance improvement. Xing Hu 0001, Yi Xu 0010, Yu Hu 0001, Yuan Xie 0001 |
DAC | 5 |
| 2014 | Partial-SET: Write speedup of PCM main memoryabstractPhase change memory (PCM) is a promising nonvolatile memory technology developed as a possible DRAM replacement. Although it offers the read latency close to that of DRAM, PCM generally suffers from the long write latency. Long write request may block the read requests on the critical path of cache/memory access, incurring adverse impact on the system performance. Besides, the write performance of PCM is very asymmetric, i.e, the SET operation (writing `1') is much slower than that of the RESET operation (writing `0'). In this work, we re-examine the resistance transform process during the SET operation of PCM and propose a novel Partial-SET scheme to alleviate the long write latency issue of PCM. During a write access to a memory line, a short Partial-SET pulse is applied first to program the PCM cells to a pre-stable state, achieving the same write latency as RESET. The partially-SET cells are then fully programmed within the retention window to preserve the data integrity. Experimental results show that our Partial-SET scheme can improve the memory access performance of PCM by more than 45% averagely with very marginal storage overhead. Bing Li 0017, Shuchang Shan, Yu Hu 0001, Xiaowei Li 0001 |
DATE | 3 |
| 2014 | Orchestrator: Guarding Against Voltage Emergencies in Multithreaded ApplicationsabstractVoltage emergency (VE) has become a critical challenge with decreasing feature size and increasing power capacity. Destructive core interference is one main source of VE in multicore processors. We observed that the applications following single program and multiple data programming model tend to spark domain-wide destructive core interference because multiple threads exhibit similar power activity. We analyze and quantify this effect and propose one low-cost solution, Orchestrator, to avoid voltage droop synergy among cores. Orchestrator leverages the thread diversity to smooth voltage droops in multicore architectures based on thread scheduling. The thread migration impact on performance is also considered. Experimental results show that Orchestrator can significantly reduce VEs, thereby improving performance. Xing Hu 0001, Guihai Yan, Yu Hu 0001, Xiaowei Li 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2014 | Reliability-Oriented Placement and Routing Algorithm for SRAM-Based FPGAsabstractAs the feature size shrinks to the nanometer scale, SRAM-based FPGAs will become increasingly vulnerable to soft errors. Existing reliability-oriented placement and routing approaches primarily focus on reducing the fault occurrence probability (node error rate) of soft errors. However, our analysis shows that, besides the fault occurrence probability, the propagation probability (error propagation probability) plays an important role and should be taken into consideration. In this paper, we first propose a cube-based analysis algorithm to efficiently and accurately estimate the error propagation probability. Based on such a model, we propose a novel reliability-oriented placement and routing algorithm that combines both the fault occurrence probability and the error propagation probability together to enhance system-level robustness against soft errors. Experimental results show that, compared with the baseline versatile place and route technique, the proposed scheme can reduce the failure rate by 20.73%, and increase the mean time between failures by 39.44%. Keheng Huang, Yu Hu 0001, Xiaowei Li 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2014 | Diagnose Failures Caused by Multiple Locations at a TimeabstractFault diagnosis plays an important role in physical failure analysis and yield learning process. With tens of billions of transistors being integrated in one chip, multiple faults may exist. With multiple faults, fault masking and reinforcing effects may appear. They may cause the conventional single-fault-based diagnosis methods such as the single location at a time (SLAT) to be invalid. The popular SLAT approach fails if there are not enough SLAT patterns that can be explained by a single stuck-at fault. Moreover, a real silicon defect may behave as different fault models (DM) under different failing patterns, which may invalidate the SLAT approach that uses a single-fault model across all failing patterns. In this paper, we introduce the concept of fault element to support multiple fault models, and use a fault-element graph (FEG) to consider fault masking and reinforcing effects among multiple faults. Based on the FEGs of all failing patterns, the most likely fault locations and their fault elements are iteratively identified. Meanwhile, the FEGs are iteratively pruned to keep track of the remaining multiple fault effects until all the fault locations are identified and all the FEGs are reduced to null. Experiments demonstrate that the proposed diagnosis method can identify the locations of multiple faults even under DM with high diagnostic accuracy and resolution. Jing Ye 0001, Yu Hu 0001, Xiaowei Li 0001, Wu-Tung Cheng, Yu Huang 0005, Huaxing Tang |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2013 | Orchestrator: a low-cost solution to reduce voltage emergencies for multi-threaded applicationsabstractVoltage emergencies have become a major challenge to multi-core processors because core-to-core resonance may put all cores into danger which jeopardizes system reliability. We observed that the applications following SPMD (Single Program and Multiple Data) programming model tend to spark domain-wide voltage resonance because multiple threads sharing the same function body exhibit similar power activity. When threads are judiciously relocated among the cores, the voltage droops can be greatly reduced. We propose “Orchestrator”, a sensor-free non-intrusive scheme for multi-core architectures to smooth the voltage droops. Orchestrator focuses on the inter-core voltage interactions, and maximally leverages the thread diversity to avoid voltage droops synergy among cores. Experimental results show that Orchestrator can reduce up to 64% voltage emergencies on average, meanwhile improving performance. Xing Hu 0001, Guihai Yan, Yu Hu 0001, Xiaowei Li 0001 |
DATE | 3 |
| 2013 | Capturing post-silicon variation by layout-aware path-delay testingabstractWith aggressive device scaling, the impact of parameter variation is becoming more prominent, which results in the uncertainty of a chip's performance. Techniques that capture post-silicon variation by deploying on-chip monitors suffer from serious area overhead and low testing reliability, while techniques using non-invasion test are limited in small scale circuits. In this paper, a novel layout-aware post-silicon variation extraction method which is based on non-invasive path-delay test is proposed. The key technique of the proposed method is a novel layout-aware heuristic path selection algorithm which takes the spatial correlation and linear dependence between paths into consideration. Experimental results show that the proposed technique can obtain an accurate timing variation distribution with zero area overhead. Moreover, the test cost is much smaller than the existing non-invasion method. Jing Ye 0001, Yu Hu 0001, Xiaowei Li 0001 |
DATE | 3 |
| 2013 | HHC: Hierarchical hardware checkpointing to accelerate fault recovery for SRAM-based FPGAsabstractAs the feature size shrinks to the nanometer scale, SRAM-based FPGAs are increasingly vulnerable to soft errors. Checkpointing is an effective fault recovery technique that can restore the faulty system to its previous fault free state. Since the function of the system needs to be suspended during checkpoint saving and checkpoint restoring, so the Mean Time to Repair (MTTR) of the system is critical to the system performance. In this work, we propose a hierarchical hardware checkpointing (HHC) technique that contains a high-speed on-chip checkpoint and a low-speed off-chip checkpoint to accelerate fault recovery for SRAM-based FPGAs. Most of single event effect (SEE) faults can be recovered by the high-speed on-chip checkpoint, which significantly reduces the MTTR of the system. The memory resource occupation of the on-chip checkpoint is low because HHC only stores the logic states of user bits and check information for configuration bits. Experimental results show that, compared with traditional off-chip checkpoint strategies, the proposed technique can reduce the MTTR of the system by 94.30%. In addition, the memory resource occupation is 11.11% of FPGAs, a little high but can be further optimized. Enshan Yang, Keheng Huang, Yu Hu 0001, Xiaowei Li 0001, Hongjin Liu, Bo Liu 0018 |
IOLTS | 3 |
| 2013 | Diagnosis and Layout Aware (DLA) scan chain stitchingabstractWithout appropriate stitching of scan chains, even with good diagnosis algorithm and diagnostic pattern generation, it may still result in bad scan chain diagnostic resolution. To improve the diagnostic resolution, we propose a novel Diagnosis and Layout Aware (DLA) scan chain stitching method, which is pattern independent and supports embedded scan compaction. It is based on three ideas: (1) increasing the number of sensitive scan cells, which can capture useful diagnostic information; (2) properly distributing the sensitive scan cells along the scan chains to enhance the overall resolution; (3) stitching scan cells based on their placement at layout to preserve the chip performance. Experiments on ISCAS'89/ITC'99 benchmark circuits and a real industry circuit based on 20nm technology with silicon results show that, the proposed DLA scan chain stitching method effectively improves the resolution, with negligible impact on chip performance, embedded scan compaction, transition fault coverage, and test power dissipation. The silicon results even show 7X average resolution improvement comparing to without using the proposed method. Jing Ye 0001, Yu Huang 0005, Yu Hu 0001, Wu-Tung Cheng, Ruifeng Guo, Liyang Lai, Ting-Pu Tai, Xiaowei Li 0001, Wei-pin Changchien, Daw-Ming Lee, Ji-Jan Chen, Sandeep C. Eruvathi, Kartik K. Kumara, Charles C. C. Liu, Sam Pan |
ITC | 3 |
| 2013 | Tolerating Noise in MLC PCM with Multi-Bit Error Correction CodeabstractPhase change memory (PCM) has emerged as a mostly promising non-volatile memory. Multi-level Cell (MLC) PCM that stores multiple bits in a single cell, has the benefits of increasing capacity and lower cost-per-bit. However, as feature size scales down, prior work reports that low frequency noise and random telegraph noise would greatly jeopardize the reliability of MLC PCM. In this paper, we firstly analyze the multi-bit error rate induced by noise and then propose a multi-bit ECC (Error Correction Code) to alleviate the deleterious noise effects in MLC PCM. As far as we know, this is the first paper to utilize of error correction method to mitigate the impact of noise at architectural level. However, a strong multi-bit ECC requires additional storage and latency. Thus, we propose a 6EC-7ED BCH scheme which achieves a tradeoff between correction capability and overhead. Compared to conventional DRAM ECC, this scheme effectively improves the reliability of MLC PCM system, while has the comparable storage overhead. Moreover, the experimental results show this scheme incurs negligible latency cost with merely 1% performance degradation. Bing Li 0017, Shuchang Shan, Yu Hu 0001, Xiaowei Li 0001 |
PRDC | 3 |
| 2012 | In-Field Testing of NAND Flash Storage: Why and How?abstractNAND Flash memories have rapidly emerged as a storage class memory such as SSD (Solid State Disk), CF (Compact Flash) Card, SD (Secure Digital Memory) Card. Due to its distinct operation mechanisms, NAND Flash memory suffers from erase/program endurance, data retention and program/read disturbance problems. Specifically, erase and program operation keeps in developing bad blocks during the lifetime of memory chips. Bad blocks are blocks that contain faulty bits but the ECC (Error Correction Code) algorithm cannot correct them. Although wear leveling tries to balance the erase/program operations on different blocks so that all blocks can wear out at a similar pace, new bad blocks still inevitably occur.We propose an in-field testing technique which takes some pages in a block as predictors. Due to wear out faster than the other pages, the predictors will become bad before the other pages in the block become bad. The further questions are (1) how to detect those wearing fast pages so as to use them as predictors, (2) how many predictors are needed to achieve a satisfactory prediction accuracy, (3) misprediction will result in what negative impact on performance and endurance. Yu Hu 0001, Xinli Gu, Xiaowei Li 0001 |
Asian Test Symposium | 1 |
| 2012 | Off-path leakage power aware routing for SRAM-based FPGAsabstractAs the feature size and threshold voltage reduce, leakage power dissipation becomes an important concern in SRAM-based FPGAs. This work focuses on reducing the leakage power in routing resources, and more specifically, the leakage power dissipated in the used part of FPGA device, which is known as the active leakage power. We observe that the leakage power in off-path transistors takes up most of the active leakage power in multiplexers that control routing, and strongly depends on Hamming distance between the state of the on-path input and the states of the off-path inputs. Hence, an off-path leakage power aware routing algorithm is proposed to minimize Hamming distance between the state of on-path input and the states of off-path inputs for each multiplexer. Experimental results on MCNC benchmark circuits show that, compared with the baseline VPR technique, the proposed off-path leakage aware routing algorithm can reduce active leakage power in routing resources by 16.79%, and the increment of critical-path delay is only 1.06%. Keheng Huang, Yu Hu 0001, Xiaowei Li 0001, Bo Liu 0018, Hongjin Liu |
DATE | 2 |
| 2012 | IVF: Characterizing the Vulnerability of Microprocessor Structures to Intermittent FaultsabstractAs CMOS technology scales into the nanometer era, future shipped microprocessors will be increasingly vulnerable to intermittent faults. Quantitatively characterizing the vulnerability of microprocessor structures to intermittent faults at an early design stage is significantly helpful in balancing system reliability and performance. Prior researches have proposed several metrics to analyze the vulnerability of microprocessor structures to soft errors and hard faults, however, the vulnerability of these structures to intermittent faults is rarely considered yet. In this work, we propose a metric intermittent vulnerability factor (IVF) to characterize the vulnerability of microprocessor structures to intermittent faults. A structure's IVF is the probability an intermittent fault in that structure causes an external visible error (failure). We compute IVFs for reorder buffer and register file considering three intermittent fault models: intermittent stuck-at-1 and stuck-at-0 fault model, intermittent open and short fault model, and intermittent timing fault model. Experimental results show that, among the three types of intermittent faults, intermittent stuck-at-1 faults have the most serious impact on program execution. Besides, IVF varies significantly across individual structures and programs, which implies partial protection to the most vulnerable structures and program phases for minimizing performance and/or energy overheads. Songjun Pan, Yu Hu 0001, Xiaowei Li 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2011 | Exploiting Free LUT Entries to Mitigate Soft Errors in SRAM-based FPGAsabstractAs the feature size of FPGA shrinks to nanometers, SRAM-based FPGAs are more vulnerable to soft errors. During logic synthesis, reliability of the design can be improved by introducing logic masking effect. In this work, we observe that there are a lot of not-fully occupied look-up tables (LUTs) after logic synthesis. Hence, we propose a functional equivalent class based soft error mitigation scheme to exploit free LUT entries in the circuit. The proposed technique replaces not fully-occupied LUTs with corresponding functional equivalent classes, which can improve the reliability while preserve the functionality of the design. Experimental results show that, compared with the baseline ABC mapper, the proposed technique can reduce the soft error rate by 21%, and the critical-path delay increase is only 4.25%. Keheng Huang, Yu Hu 0001, Xiaowei Li 0001, Gengxin Hua, Hongjin Liu, Bo Liu 0018 |
Asian Test Symposium | 2 |
| 2011 | Cross-layer optimized placement and routing for FPGA soft error mitigationabstractAs the feature size of FPGA shrinks to nanometers, soft errors increasingly become an important concern for SRAM-based FPGAs. Without consideration of the application level impact, existing reliability-oriented placement and routing approaches analyze soft error rate (SER) only at the physical level, consequently completing the design with suboptimal soft error mitigation. Our analysis shows that the statistical variation of the application level factor is significant. Hence in this work, we first propose a cube-based analysis to efficiently and accurately evaluate the application level factor. And then we propose a cross-layer optimized placement and routing algorithm to reduce the SER by incorporating the application level and the physical level factor together. Experimental results show that, the average difference of the application level factor between our cube-based method and Monte Carlo golden simulation is less than 0.01. Moreover, compared with the baseline VPR placement and routing technique, the cross-layer optimized placement and routing algorithm can reduce the SER by 14% with no area and performance overhead. Keheng Huang, Yu Hu 0001, Xiaowei Li 0001 |
DATE | 2 |
| 2011 | A cost-effective substantial-impact-filter based method to tolerate voltage emergenciesabstractSupply voltage fluctuation caused by inductive noises has become a critical problem in microprocessor design. A voltage emergency occurs when supply voltage variation exceeds the acceptable voltage margin, jeopardizing the microprocessor reliability. Existing techniques assume all voltage emergencies would definitely lead to incorrect program execution and prudently activate rollbacks or flushes to recover, and consequently incur high performance overhead. We observe that not all voltage emergencies result in external visible errors, which can be exploited to avoid unnecessary protection. In this paper, we propose a substantial-impact-filter based method to tolerate voltage emergencies, including three key techniques: 1) Analyze the architecture-level masking of voltage emergencies during program execution; 2) Propose a metric intermittent vulnerability factor for intermittent timing faults (IV Fitf) to quantitatively estimate the vulnerability of microprocessor structures (load/store queue and register file) to voltage emergencies; 3) Propose a substantial-impact-filter based method to handle voltage emergencies. Experimental results demonstrate our approach gains back nearly 57% of the performance loss compared with the once-occur-then-rollback approach. Songjun Pan, Yu Hu 0001, Xing Hu 0001, Xiaowei Li 0001 |
DATE | 2 |
| 2011 | On diagnosis of multiple faults using compacted responsesabstractWith the exponential growth in the number of transistors, not only test data volume and test application time may increase, but also multiple faults may exist in one chip. Test compaction has been a de-facto design-for-testability technique to reduce the test cost. However, the compacted test responses make multiple-fault diagnosis rather difficult. When there is no space compactor, the most likely suspect fault is considered producing the failing responses most similar to the failing responses observed from the automatic test equipment. But when compactor exists, those suspect faults may no longer have the same high possibility of being the actual faults. To address this problem, we introduce a novel metric explanation necessity. By using both of the new metric and the traditional metric explanation capability, we evaluate the possibility of a suspect fault to be the actual fault. For ISCAS'89 and ITC'99 benchmark circuits equipped with extreme space compactors, experimental results show that 98.8% of the top-ranked suspect faults hit the actual faults, outperforming a previous work by 11.3%. Jing Ye 0001, Yu Hu 0001, Xiaowei Li 0001 |
DATE | 2 |
| 2011 | Transparent dynamic binding with fault-tolerant cache coherence protocol for chip multiprocessorsabstractAggressive technology scaling causes chip multiprocessors increasingly error-prone. Core-level fault-tolerant approaches bind two cores to implement redundant execution and error detection. However, along with more cores integrated into one chip, existing static and dynamic binding schemes suffer from the scalability problem when considering the violation effects caused by external write operations. In this paper, we present a transparent dynamic binding (TDB) mechanism to address the issue. Learning from static binding schemes, we involve the private caches to hold identical data blocks, thus we reduce the global masters-lave consistency maintenance to the scale of the private caches. With our fault-tolerant cache coherence protocol, TDB satisfies the objective of private cache consistency, therefore provides excellent scalability and flexibility. Experimental results show that, for a set of parallel workloads, the overall performance of our TDB scheme is very close to that of baseline fault-tolerant systems, outperforming dynamic core coupling by 9.2%, 10.4%, 18% and 37.1% when considering 4, 8, 16 and 32 cores respectively. Shuchang Shan, Yu Hu 0001, Xiaowei Li 0001 |
DSN | 2 |
| 2011 | Scan chain design for shift power reduction in scan-based testing
Jia Li 0022, Yu Hu 0001, Xiaowei Li 0001 |
Sci. China Inf. Sci. | 2 |
| 2011 | Capture-power-aware test data compression using selective encoding
Jia Li 0022, Xiao Liu 0011, Yubin Zhang, Yu Hu 0001, Xiaowei Li 0001, Qiang Xu 0001 |
Integr. | 4 |
| 2011 | Using Launch-on-Capture for Testing Scan Designs Containing Synchronous and Asynchronous Clock DomainsabstractThis paper presents a hybrid automatic test pattern generation (ATPG) technique using the staggered launch-on capture (LOC) scheme followed by the one-hot LOC scheme for testing delay faults in a scan design containing asynchronous clock domains. Typically, the staggered scheme produces small test sets but needs long ATPG runtime, whereas the one-hot scheme takes short ATPG runtime but yields large test sets. The proposed hybrid technique is intended to reduce test pattern count with acceptable ATPG runtime for multi-million-gate scan designs. In case the scan design contains multiple synchronous clock domains, each group of synchronous clock domains is treated as a clock group and tested using a launch aligned or a capture aligned LOC scheme. By combining these schemes together, we found the pattern counts for two large industrial designs were reduced by approximately 1.1X to 2.1X, while the ATPG runtime was increased by 10% to 50%, when compared to the one-hot clocking scheme alone. Shianling Wu, Laung-Terng Wang, Xiaoqing Wen, Lang Tan, Yu Hu 0001, Wen-Ben Jone, Michael S. Hsiao, Chien-Mo James Li, Jiun-Lang Huang, Lizhen Yu |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 7 |
| 2010 | Substantial Fault Pair At-a-Time (SFPAT): An Automatic Diagnostic Pattern Generation MethodabstractVolume diagnosis plays an important role in the yield learning process. To get a high quality diagnosis result, patterns with high distinguish ability are essential. However, the test patterns used by volume diagnosis commonly have low distinguish ability to specific faults. In our experiments, we observe that on average, under automatic generated test patterns, faults in the same fan out free region (FFR) account for only 6% of all possible fault pairs, but their share in total indistinguishable faults is 70%, faults in different FFRs but with the same observation points account for 4% of all fault pairs, but their share in total indistinguishable faults is 22%. Exploiting this fact that faults in the same FFRs are harder to be distinguished, we propose an Automatic Diagnostic Pattern Generation (ADPG) method named Substantial Fault Pairs at-A-Time (SFPAT)-ADPG. By applying a transformed circuit and a new fault list to an existing Automatic Test Pattern Generation (ATPG) tool, we generate the compressed test patterns which are also the diagnostic patterns with high distinguish ability for the original circuit. Experiments on ISCAS'89 and ITC'99 benchmark circuits show the effectiveness of the proposed SFPAT-ADPG method. Jing Ye 0001, Yu Hu 0001, Xiaowei Li 0001 |
Asian Test Symposium | 3 |
| 2010 | IVF: Characterizing the vulnerability of microprocessor structures to intermittent faultsabstractWith the advancement of CMOS manufacturing process to nano-scale, future shipped microprocessors will be increasingly vulnerable to intermittent faults. Quantitatively characterizing the vulnerability of microprocessor structures to intermittent faults at early design stage is significantly helpful to balance system performance and reliability. Prior researches have proposed several metrics to characterize the vulnerability of microprocessor structures to soft errors and permanent faults, however, the vulnerability of these structures to intermittent faults are still rarely considered. In this work, we propose a metric intermittent vulnerability factor (IVF) to characterize the vulnerability of microprocessor structures to intermittent faults. A structure's IVF is the probability an intermittent fault in that structure causes an external visible error. We instrument a cycle-accurate execution-driven simulator Sim-Alpha to compute IVFs for reorder buffer and register file. Experimental results show that the IVF of reorder buffer is much higher than that of register file. Besides, IVF varies significantly across different structures and workloads, which implies partial protection to the most vulnerable structures to improve system reliability with less overhead. Songjun Pan, Yu Hu 0001, Xiaowei Li 0001 |
DATE | 2 |
| 2010 | Diagnosis of multiple arbitrary faults with mask and reinforcement effectabstractWe propose a multiple-fault diagnosis method with high diagnosability, resolution, first-hit and short run time. The method has no assumption on fault models, thus can diagnose arbitrary faults. To cope with the multiple-fault mask and reinforcement effect, two key techniques of construction and scoring of fault-tuple equivalence trees are introduced to choose and rank the final candidate locations. Experimental results show that, when the circuits have 2 arbitrary faults, the average diagnosability and resolution are 98% and 0.95, respectively, with the best case 100% and 1.00. Moreover, in average, even when 21 arbitrary faults exist, our method can still identify 93% of them with the resolution 0.78, increased by 41% and 39% in comparison with the latest work where the diagnosability and resolution are 66% and 0.56. Finally, 96% of our top-ranked candidate locations are actual fault locations. Jing Ye 0001, Yu Hu 0001, Xiaowei Li 0001 |
DATE | 2 |
| 2010 | X-Filling for Simultaneous Shift- and Capture-Power Reduction in At-Speed Scan-Based TestingabstractPower consumption during at-speed scan-based testing can be significantly higher than that during normal functional mode in both shift and capture phases, which can cause circuits' reliability concerns during manufacturing test. This paper proposes a novel X-filling technique, namely “iFill”, to address the above issue, by analyzing the impact ofX-bitson switching activities of the circuit nodes in the two different phases. In addition, different from prior$X$-filling methods for shift-power reduction that can only reduce shift-in power, our method is able to cut down power consumptions in both shift-in and shift-out processes. Experimental results on benchmark circuits show that the proposed technique can guarantee the power safety in both shift and capture phases during at-speed scan-based testing. Jia Li 0022, Qiang Xu 0001, Yu Hu 0001, Xiaowei Li 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2009 | Online Computing and Predicting Architectural Vulnerability Factor of Microprocessor StructuresabstractSoft Errors have emerged as a key challenge to microprocessor design. Traditional soft error tolerance techniques (such as redundant multithreading and instruction duplication) can achieve high fault coverage but at the cost of significant performance degradation. Prior research reports that soft errors can be masked at the architecture level, and the degree of such masking, named as architecture vulnerability factor (AVF), can vary significantly across workloads and individual structures, hence strict redundant execution may not be necessary for soft error tolerance. In this work, we exploit the AVF varying feature to adaptively tune reliability and performance. We present an infrastructure to online compute and predict AVF for three microprocessor structures (IQ, ROB, and LSQ), guiding when the protection scheme should be activated to improve reliability. Experimental results show that our method can efficiently compute the AVF for different structures independent of hardware configurations. The average differences between our method and a prior offline AVF computing method are 0.10, 0.01, and 0.039 for IQ, ROB, and LSQ, respectively. Songjun Pan, Yu Hu 0001, Xiaowei Li 0001 |
PRDC | 2 |
| 2008 | Robust test generation for power supply noise induced path delay faultsabstractIn deep sub-micron designs, the delay caused by power supply noise (PSN) can no longer be ignored. A PSN-induced path delay fault (PSNPDF) model is proposed in this paper, and should be tested to enhance chip quality. Based on precise timing analysis, we also propose a robust test generation technique for PSNPDF. Concept of timing window is introduced into the PSNPDF model. If two devices in the same feed region simultaneously switch in the same direction, the current waveform of the two devices will have an overlap and excessive PSN will be produced. Experimental results on ISCAS’89 circuits showed test generation can be finished in a few seconds. Xiang Fu 0007, Huawei Li 0001, Yu Hu 0001, Xiaowei Li 0001 |
ASP-DAC | 3 |
| 2008 | Localized random access scan: Towards low area and routing overheadabstractConventional random access scan (RAS) designs, although economic in test power dissipation, test application time and test data volume, are expensive in area and routing overhead. In this paper, we present a localized RAS architecture (LRAS) to address this issue. A novel scan cell structure, which has fewer transistors than the multiplexer-type scan cell, is proposed to eliminate the global test enable signal and to localize the row enable and the column enable signals. Experimental results on ISCAS'89 and ITC'99 benchmark circuits demonstrate that LRAS has 54% less area overhead than multiplexer-type scan chain based designs, while significantly outperforms the state-of-the-art RAS scheme in routing overhead. Yu Hu 0001, Xiang Fu 0007, Xiaoxin Fan, Hideo Fujiwara |
ASP-DAC | 1 |
| 2008 | On reducing both shift and capture power for scan-based testingabstractPower consumption in scan-based testing is a major concern nowadays. In this paper, we present a new X-filling technique to reduce both shift power and capture power during scan tests, namely LSC-filling. The basic idea is to use as few as possible X-bits to keep the capture power under the peak power limit of the circuit under test (CUT), while using the remaining X-bits to reduce the shift power to cut down the CUT’s average power consumption during scan tests as much as possible. In addition, by carefully selecting the X-filling order, our X-filling technique is able to achieve lower capture power when compared to existing methods. Experimental results on ISCAS’89 benchmark circuits show the effectiveness of the proposed methodology. Jia Li 0022, Qiang Xu 0001, Yu Hu 0001, Xiaowei Li 0001 |
ASP-DAC | 3 |
| 2008 | A design- for-diagnosis technique for diagnosing both scan chain faults and combinational circuit faultsabstractThe amount of die area consumed by scan chains and scan control circuit can range from 15%~30%, and scan chain failures account for almost 50% of chip failures. As the conventional diagnosis process usually runs on the faulty free scan chain, scan chain faults may disable the diagnostic process, leaving large failure area to time-consuming failure analysis. In this paper, a design-for-diagnosis (DFD) technique is proposed to diagnose faulty scan chains precisely and efficiently, moreover, with the assistant of the proposed technique, the conventional logic diagnostic process can be carried on with faulty scan chains. The proposed approach is entirely compatible with conventional scan-based design. Previously proposed software-based diagnostic methods for conventional scan designs can still be applied to our design. Experiments on ISCAS'89 benchmark circuits are conducted to demonstrate the efficiency of the proposed DFD technique. Yu Hu 0001, Huawei Li 0001, Xiaowei Li 0001 |
ASP-DAC | 2 |
| 2008 | Observation Point Oriented Deterministic Diagnosis Pattern Generation (DDPG) for Chain DiagnosisabstractScan is a widely used Design-for-Testability technique to improve test and diagnosis quality. Many defects may cause scan chains to fail. In this paper, an observation point oriented Deterministic Diagnostic Pattern Generation (DDPG) method was proposed for compound defects, which tolerates the system defects during scan chain diagnosis. Instead of sensitizing multiple paths proposed in our prior work, the proposed new DDPG method directly targets as many observation points as possible to observe the loading error occurred on the targeted scan cell. Experimental results on ISCASpsila89 benchmark circuits show that the proposed DDPG method improves the effectiveness and efficiency of diagnosing compound defects, compared to our prior research. Yu Hu 0001, Yu Huang 0005, Jing Ye 0001, Xiaowei Li 0001 |
ATS | 2 |
| 2008 | iFill: An Impact-Oriented X-Filling Method for Shift- and Capture-Power Reduction in At-Speed Scan-Based TestingabstractIn scan-based tests, power consumptions in both shift and capture phases may be significantly higher than that in normal mode, which threatens circuits' reliability during manufacturing test. In this paper, by analyzing the impact of X-bits on circuit switching activities, we present an X-filling technique that can decrease both shift- and capture-power to guarantees the reliability of scan tests, called iFill. Moreover, different from prior work on X-filling for shift-power reduction which can only reduce shift-in power, iFill is able to decrease power consumptions during both shift-in and shift-out. Experimental results on ISCAS' 89 benchmark circuits show the effectiveness of the proposed technique. Jia Li 0022, Qiang Xu 0001, Yu Hu 0001, Xiaowei Li 0001 |
DATE | 3 |
| 2008 | On capture power-aware test data compression for scan-based testingabstractLarge test data volume and high test power are two of the major concerns for the industry when testing large integrated circuits. With given test cubes in scan-based testing, the ldquodonpsilat-carerdquo bits can be exploited for test data compression and/or test power reduction. Prior work either targets only one of these two issues or considers to reduce test data volume and scan shift power together. In this paper, we propose a novel capture power-aware test compression scheme that is able to keep scan capture power under a safe limit with little loss in test compression ratio. Experimental results on benchmark circuits demonstrate the efficacy of the proposed approach. Jia Li 0022, Xiao Liu 0011, Yubin Zhang, Yu Hu 0001, Xiaowei Li 0001, Qiang Xu 0001 |
ICCAD | 4 |
| 2008 | Deterministic Diagnostic Pattern Generation (DDPG) for Compound DefectsabstractScan chain failure diagnosis has become an important means for silicon debug and yield improvement. Although plenty of prior work discussed how to perform scan chain diagnosis, most of the previously proposed techniques made an assumption that the system logic is fault-free, which could be an impractical assumption leading to incorrect diagnostic results. In this paper, we propose a scan chain deterministic diagnostic pattern generation (DDPG) method that can tolerate the faults in the system logic without degradation of chain diagnostic resolution and precision. The entire flow includes three steps. In the first step, patterns are created to propagate the state of a targeted scan cell to as many reliable observation points as possible. In the second step, the load error probability of each targeted scan cell is calculated based on the hamming distances between the observed responses and the expected good or faulty responses. In the last step, a suspect profile is plotted, which can be used to identify the suspect scan cell(s) based on ranking scores. Experimental results show that the diagnostic resolution and precision are not degraded even with dozens of faults injected into the system logic. Yu Hu 0001, Huawei Li 0001, Xiaowei Li 0001, Jing Ye 0001, Yu Huang 0005 |
ITC | 2 |
| 2008 | Diagnosis of Mask-Effect Multiple Timing Faults in Scan ChainsabstractA deterministic diagnosis method for multiple timing faults in scan chains is proposed. Compared to prior work, our approach can diagnose mask-effect multiple timing faults as well as conventional mixed multiple timing faults. Experimental results on ISCAS'89 benchmark circuits demonstrate that the average diagnosis resolution of two faults is less than 3. Jing Ye 0001, Yu Hu 0001, Xiaowei Li 0001 |
ITC | 3 |
| 2008 | Codeword Selection for Crosstalk Avoidance and Error Correction on InterconnectsabstractCrosstalk effects and soft errors on interconnects have been increasingly serious, which affects normal communication among cores. Therefore, it is desirable to design a reliable bus system without causing unacceptable performance reduction. In this paper, a new bus encoding method based on codeword selection is presented for enduring crosstalk-induced effects, which can avoid crosstalk and provide error correction as well. This method finds a subset from crosstalk avoidance code (CAC) to provide error correction. It can avoid crosstalk induced by late signal transition on checking bits in the previous methods. Extra wires for checking bus are never required in the proposed method. Experiment shows that the method reduces 6% wire overhead compared to the former methods. And it can also improve bus performance and reduce power dissipation. Ying Zhang 0040, Huawei Li 0001, Xiaowei Li 0001, Yu Hu 0001 |
VTS | 4 |
| 2008 | Design-for-Testability Features and Test Implementation of a Giga Hertz General Purpose Microprocessor
Yu Hu 0001, Huawei Li 0001, Xiaowei Li 0001 |
J. Comput. Sci. Technol. | 2 |
| 2007 | An On-Chip Test Clock Control Scheme for Multi-Clock At-Speed TestingabstractTo test timing-related faults between synchronous clocks, an at-speed test clock and an automatic test pattern generation scheme are needed. However, previous work on designing on-chip at-speed test clock controllers for multi-clock has quadratic increasing area overhead along with linearly increasing clocks. This paper presents a clock-chain based test clock control scheme using an internal phase-locked-loop (PLL) as the at-speed test clock generator, which supports at-speed testing for inter-clock domain and intra-clock domain logic. Experimental results demonstrate that the proposed design has low area overhead when increasing the number of clocks. Xiaoxin Fan, Yu Hu 0001, Laung-Terng Wang |
ATS | 2 |
| 2007 | The design-for-testability features of a general purpose microprocessorabstractThis paper describes the design-for-testability (DFT) features and test challenges in a general purpose microprocessor design. An optimized DFT architecture with its implementation strategies are presented in detail. Major DFT solutions are implemented which can meet high-volume manufacturing (HVM) and high quality test goals. Xiaoxin Fan, Xiang Fu 0007, Huawei Li 0001, Yu Hu 0001, Xiaowei Li 0001 |
ITC | 8 |
| 2007 | Leakage Current Optimization Techniques During Test Based on Don't Care Bits Assignment
Yu Hu 0001, Yinhe Han 0001, Xiaowei Li 0001, You-Sheng Zhang |
J. Comput. Sci. Technol. | 2 |
| 2007 | Embedded Test Decompressor to Reduce the Required Channels and Vector Memory of Tester for Complex Processor CircuitabstractAn embedded test stimulus decompressor is presented for the test patterns decompression, which can reduce the required channels and vector memory of automatic test equipment (ATE) for complex processor circuit. The proposed decompressor mainly consists of a periodically alterable MUX network which has multiple configurations to decode the input information flexibly and efficiently. In order to reduce the number of test patterns and configurations, a test patterns compaction algorithm, using CI-Graph merging, is proposed. With the proposed periodically alterable MUX network and the patterns compaction algorithm, smaller test data volume and required external pins can be achieved as compared to previous techniques Yinhe Han 0001, Yu Hu 0001, Xiaowei Li 0001, Huawei Li 0001, Anshuman Chandra |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2006 | Test data compression based on clustered random access scanabstractWe proposed clustered random access scan (CRAS) architecture to reduce test data volume. CRAS makes use of the compatibility of the test stimuli to cluster the scan cells, and assigns every cluster a unique address. The compression ratio upper bound of CRAS is analyzed based on the random graph theory. Experimental results on ISCAS'89 benchmarks and two industry designs show that the proposed CRAS architecture can yield on average 67.3% reduction in test data volume, with reasonable area and routing overhead than scan design Yu Hu 0001, Jia Li 0022, Yinhe Han 0001, Xiaowei Li 0001, Huawei Li 0001, Laung-Terng Wang, Xiaoqing Wen |
ATS | 1 |
| 2006 | A Scan Chain Adjustment Technology for Test Power ReductionabstractRecently test power dissipation has become a more and more challenging issue. This paper proposes a technique to solve this problem through scan chain adjustment to eliminate unnecessary transitions in scan chains. An extended WTM (EWTM) metric is proposed to estimate dynamic power dissipation in circuit under test caused by transitions in test stimulus and response vectors. And the routing overhead of this methodology can be reduced through scan chain adjustment guided with our Distance of EWTM (DEWTM) metric. Experimental results on ISCAS'89 benchmarks circuits show that the proposed approach can reduce average power dissipation during scan test by 72.2% on average, with negligible routing overhead. Jia Li 0022, Yu Hu 0001, Xiaowei Li 0001 |
ATS | 2 |
| 2006 | An on-chip combinational decompressor for reducing test data volumeabstractUtilizing an on-chip decompressor is an efficient method to reduce test data volume in multiple-scan-chain designs. This paper investigates a new technique to implement the decompressor by combinational circuits. The proposed architecture drives a large number of internal scan chains with far fewer external input pins, thus delivering significant reductions in test data volume. Based on the analysis of compatible relationships among scan slices, the number of external scan inputs can be minimized. The effectiveness and applicability of the proposed scheme are demonstrated by experimental results. Jie Don, Yu Hu 0001, Yinhe Han 0001, Xiaowei Li 0001 |
ISCAS | 2 |
| 2005 | Theoretic analysis and enhanced X-tolerance of test response compact based on convolutional codeabstractThis paper addresses the problem of test response compaction. In order to maximize compaction ratio, a single-output encoder based on check matrix of a (n, n-1, m, 3) convolutional code is proposed. Theoretic analysis for this encoder is presented to avoid two and any odd erroneous bit cancellations, handle one unknown bit(X bit) and diagnose one erroneous bit. The X-bits tolerance capacity can be enhanced by choosing a proper memory size and weight of check matrix, which can also be obtained by an optimized input assignment algorithm. The theoretic analysis and experimental results on aliasing shows the efficiency of the proposed encoder. Yinhe Han 0001, Yu Hu 0001, Huawei Li 0001, Xiaowei Li 0001 |
ASP-DAC | 2 |
| 2005 | Scan Data Volume Reduction Using Periodically Alterable MUXs DecompressorabstractThis paper presents a decompression architecture using a periodically alterable MUXs decompressor for scan data volume reduction. Compared to static XOR network, the periodically alterable MUXs decompressor has multiple configurations to decode the input information more efficiently. Three different DFT techniques are proposed to handle hard, firm and soft cores, respectively. With the proposed pattern decompression algorithms and scan decompression architecture, smaller test data volume and test application time can be achieved as compared to previous techniques. Yinhe Han 0001, Xiaowei Li 0001, Shivakumar Swaminathan, Yu Hu 0001, Anshuman Chandra |
Asian Test Symposium | 4 |
| 2005 | Compression/Scan Co-Design for Reducing Test Data Volume, Scan-in Power Dissipation and Test Application TimeabstractTesting chips is very critical to guarantee chips are fault-free before they are integrated in a system, so as to increase the reliability of the system. Although full-scan is a widely adopted design-for-test technique for LSI design and testing, the need for reducing the test data volume, scan-in power dissipation and test application time (VPT) of the full-scan designed chip is imperative. Based on the analysis of the characteristics of the variable-to-fixed run-length coding technique and the random access scan architecture, this paper presents a novel design scheme tackling all VPT issues simultaneously. Experimental results on ISCAS'89 benchmarks have shown on average 51.2%, 99.5%, 99.3% and 85.5% reduction in test data volume, average scan-in power dissipation, peak scan-in power dissipation and test application time, respectively. Yu Hu 0001, Xiaowei Li 0001, Huawei Li 0001, Xiaoqing Wen |
PRDC | 1 |
| 2004 | Rapid and Energy-Efficient Testing for Embedded CoresabstractConventional serial connection of internal scan chains brings the power and time penalty. A parallel core wrapper design (pCWD) approach is presented in this paper for reducing test power and test application time. The pCWD utilizes overlapping scan slices to reduce the number of scan slices loading. Experimental results on d695 of ITC2002 benchmark demonstrated that, about 2/spl times/ shift time and 20/spl times/ test power reduction can be achieved. Yinhe Han 0001, Yu Hu 0001, Huawei Li 0001, Xiaowei Li 0001, Anshuman Chandra |
Asian Test Symposium | 2 |
| 2004 | Pair Balance-Based Test Scheduling for SOCsabstractAlong with more pre-designed and pre-verified cores are integrated into a single chip to construct an entire system, the test application time increases significantly. This paper presents a novel test scheduling solution, unlike previous techniques that take advantage of balanced scan chains of every single core, utilizing the balance of pairwise combined cores. Experimental results for two ITC '02 SOC benchmarks show that the pair balance-based test scheduling technique achieves less test time compared to the previous approaches. Yu Hu 0001, Yinhe Han 0001, Huawei Li 0001, Tao Lv 0001, Xiaowei Li 0001 |
Asian Test Symposium | 1 |