Danella Zhao

dblp:11/6855 · also Dan Zhao 0001 · DBLP profile ↗
← Back
32ranked-venue papers
8as first author
11since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 23 · 8 first-author · 5 since 2021Computer networks · 2 · 2 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021Security and privacy · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Silentflow: Leveraging Trusted Execution for Resource-Limited MPC via Hardware-Algorithm Co-design
abstract
Secure Multi-Party Computation (MPC) offers a practical foundation for privacy-preserving machine learning at the edge, with MPC commonly employed to support nonlinear operations. These MPC protocols fundamentally rely on Oblivious Transfer (OT), particularly Correlated OT (COT), to generate correlated randomness essential for secure computation. Although COT generation is efficient in conventional two-party settings with resource-rich participants, it becomes a critical bottleneck in real-world inference on resource-constrained devices (e.g., IoT sensors and wearables), due to both communication latency and limited computational capacity. To enable realtime secure inference, we introduce Silentflow, a highly efficient Trusted Execution Environment (TEE)-assisted protocol that eliminates communication in COT generation. We tackle the core performance bottleneck—low computational intensity—through structured algorithmic decomposition: kernel fusion for parallelism, Blocked On-chip eXpansion (BOX) to improve memory access patterns, and vectorized batch operations to maximize memory bandwidth utilization. Through design space exploration, we balance end-to-end latency and resource demands, achieving up to $39.51 \times$ speedup over state-of-the-art protocols. By offloading COT computations to a Zynq-7000 SoC, SilentFlow accelerates PPMLaaS inference on the ImageNet dataset under resource constraints, achieving a $4.62 \times$ and $3.95 \times$ speedup over Cryptflow2 and Cheetah, respectively.
Hanieh Totonchi Asl, Ebrahim Nouri, Danella Zhao
ASP-DAC5
2026 SecDTD: Dynamic Token Drop for Secure Transformers Inference
Yizhou Feng, Qiao Zhang 0002, Hongyi Wu, Danella Zhao, Chunsheng Xin
EuroS&P6
2026 Bridging Trust and Efficiency: TEE-Accelerated Nonlinear Evaluation in Multiparty Computing
Hanieh Totonchi Asl, Ebrahim Nouri, Danella Zhao
ISCAS4
2026 DeepRevelio: Multivariate Time Series Detection for Stealthy IoT Malware via Variable-length Entropy Analysis
Danella Zhao
ISCAS2
2025 Lynx-Net: Privacy-Preserving Neural Network Training and Malware Detection at IoT-Edge
abstract
Cyberattacks on IoT devices are accelerating at an unprecedented rate, largely driven by IoT malware activities. The IoT malware attacks are typically composed of three stages: intrusion, infection, and execution. It is essential to instantaneously detect malware at the early stage of intrusion on IoT devices before massive attacks. To build an efficient and scalable instruction detection system, a multi-client single-server privacy-preserved neural network model is proposed. Further, it is important to train the neural network model on the fly to cope with the fast-evolving malware variations. For online training and detection, we propose Lynx-Net, a distributed deep learning online training model for inferring malicious activities by analyzing power side-channel signals. To protect user private data and model parameters, and reduce prediction latency, we implement a novel privacy-preserved protocol via secret sharing and packed hybrid homomorphic encryption. Through theoretical analysis and empirical experiments, we demonstrate that Lynx-Net can detect infection activities of different IoT malware with high accuracy. Our extensive experiments demonstrate not only stable training but also a 1.13 to 2.67 times speedup compared to the state-of-the-art training model and an 8 to 500 times improvement in Lynx-Net’s inference prediction latency compared to the state-of-the-art inference model.
Sabbir Ahmed Khan, Danella Zhao, Ravi Mukkamala, Woosub Jung
SERA3
2024 SEEK+: Securing vehicle GPS via a sequential dashcam-based vehicle localization framework
Peng Jiang 0027, Hongyi Wu, Yanxiao Zhao, Danella Zhao, Gang Zhou 0002, Chunsheng Xin
Pervasive Mob. Comput.4
2024 ZeroD-fender: A Resource-aware IoT Malware Detection Engine via Fine-grained Side-channel Analysis
abstract
In early 2023, cyberattacks experienced a significant rise due to unknown (zero-day) malware targeting Internet of Things (IoT) devices. To tackle the challenge of zero-day detection within a highly resource-constrained IoT environment, we propose a novel design that utilizes fine-grained power side-channel analysis with deep learning techniques. Our approach introduces an innovative concept called multiscale feature extraction to identify the most representative malware features across diverse architectures, thereby enhancing deep learning based detection performance against zero-day malware. Specifically, we employ a fine-grained power side-channel analysis of more than 120,000 honeypot-collected malware files across a hierarchy of commands , functions , and modules to identify the unique zero-day malware behaviors. With these identified features to train our model, ZeroD-fender’s performance in detecting zero-day malware has significantly improved. In pursuit of on-device detection, we present a resource-aware online inference customization framework. This framework features our lightweight network, ThingNetV2, which uses specialized 1-D depthwise separable convolution paired with h-swish activation, leading to significant resource savings. By applying the fine-grained power analysis, ZeroD-fender demonstrates a detection rate of 95.88% across various architecture zero-day malware, achieving detection speeds ranging between 16.083 ms and 23.961 ms , depending on the specific scenario.
Danella Zhao
ACM Trans. Design Autom. Electr. Syst.2
2023 SEEK: Detecting GPS Spoofing via a Sequential Dashcam-Based Vehicle Localization Framework
abstract
GPS spoofing is a great threat to the safety of transportation systems as well as other systems that rely on GPS for navigation. This paper proposes a novel computer vision based approach for GPS spoofing detection, termed SEquential dashcam-based vEhicle localization frameworK (SEEK). SEEK utilizes vehicle dashcam images to identify a vehicle's true location and detects possible GPS spoofing attacks through verifying if the reported GPS locations of the vehicle are correct. However, it is nontrivial to use dashcam images for vehicle localization due to multiple challenges caused by real-world driving, including the complicated lighting/weather conditions, season/timing variations of the images, large blockage ratio in the images, and varying driving speeds. SEEK features a unique design with novel schemes to address complicated lighting/weather conditions, transform images to align with season changes, reduce blockage, and adopt a sequential image matching scheme. The performance evaluation shows that SEEK significantly outperforms the previous GPS spoofing detection scheme, and achieves a detection accuracy of up to 94%.
Peng Jiang 0027, Hongyi Wu, Yanxiao Zhao, Danella Zhao, Chunsheng Xin
PERCOM4
2022 ThingNet: A Lightweight Real-time Mirai IoT Variants Hunter through CPU Power Fingerprinting
abstract
Internet of Things (IoT) devices have become attractive targets of cyber criminals, whereas attackers have been leveraging these vulnerable devices most notably via the infamous Mirai-based botnets, accounting for nearly 90% of IoT malware attacks in 2020. In this work, we propose a robust, universal and non-invasive Mirai-based malware detection engine employing a compact deep neural network architecture. Our design allows programmatic collection of CPU power footprints with integrated current sensors under various device states, such as idle, service and attack. A lightweight online inference model is deployed in the CPU for on-the-fly classification. Our model is robust against noisy environment with a lucid design of noise reduction function. This work appears to be the first step towards a viable CPU malware detection engine based on power fingerprinting. The extensive simulation study under ARM architecture that is widely used in IoT devices, demonstrates a high detection accuracy of 99.1% at a speed less than 1ms. By analyzing Mirai-based infection under distinguishable phases for power feature extraction, our model has further demonstrated an accuracy of 96.3% on model-unknown variants detection.
Danella Zhao
DATE2
2022 DeepAuditor: Distributed Online Intrusion Detection System for IoT Devices via Power Side-channel Auditing
abstract
As the number of IoT devices has increased rapidly, IoT botnets have exploited the vulnerabilities of IoT devices. However, it is still challenging to detect the initial intrusion on IoT devices prior to massive attacks. Recent studies have utilized power side-channel in-formation to identify this intrusion behavior on IoT devices but still lack accurate models in real-time for ubiquitous botnet detection. We propose the first online intrusion detection system called DeepAuditor for multiple IoT devices via power auditing. To de-velop the real-time system, we propose a lightweight power auditing device called Power Auditor. We also design a distributed CNN classifier for online inference in a laboratory setting. In order to protect data leakage and reduce networking redundancy, we then propose a privacy-preserved inference protocol via Packed Homo-morphic Encryption and a sliding window protocol in our system. The classification accuracy and processing time are measured, and the proposed classifier outperforms a baseline classifier, especially against unseen patterns. We also demonstrate that the distributed CNN design is secure against any distributed components. Over-all, the measurements are shown to the feasibility of our real-time distributed system for intrusion detection on IoT devices.
Woosub Jung, Yizhou Feng, Sabbir Ahmed Khan, Chunsheng Xin, Danella Zhao, Gang Zhou 0002
IPSN5
2022 Demo Abstract: A Distributed Power Side-channel Auditing System for Online loT Intrusion Detection
abstract
As the number of IoT devices has increased rapidly, IoT botnets have exploited the vulnerabilities of IoT devices. However, it is still challenging to detect the initial intrusion on IoT devices prior to massive attacks. Thus, a new approach that monitors these ini-tial intrusions is needed. Power side-channel information can be used because it does not require any modification in programming languages or operating systems on diverse IoT devices. We propose a distributed power side-channel auditing system for online IoT intrusion detection. To meet the real-time requirement, we develop a lightweight power auditing device. We then design a distributed CNN classifier for online inference in a laboratory setting. Two distributed protocols are also proposed in order to protect data leakage and reduce networking redundancy. In this work, we demonstrate the feasibility of our real-time distributed system for intrusion detection on IoT devices.
Woosub Jung, Yizhou Feng, Sabbir Ahmed Khan, Chunsheng Xin, Danella Zhao, Gang Zhou 0002
IPSN5
2018 Neuro-NoC: Energy Optimization in Heterogeneous Many-Core NoC using Neural Networks in Dark Silicon Era
abstract
Due to the end of Dennard Scaling and the rise of dark silicon, it is essential to design energy-efficient heterogeneous NoC under critical power and thermal constraints. The challenge is to determine and configure NoC resources while meeting the application(s) requirements. Because of the large and complex many-core NoC design space (voltage/frequency scaling, link bandwidth, power-gating, etc.), design space becomes difficult to explore within a reasonable time for optimal decision at run-time. Furthermore, reactive resource management is not effective in preventing problems, such as creating thermal hotspots and exceeding power budget, from happening. Therefore, we propose a Neuro-NoC model, which utilizes neural networks learning algorithm to dynamically monitor, predict, and configure NoC resources based on online learning of the system status. Distributed cluster-wise neural network and a global neural network model for resource monitoring and configuration in many-core NoC has been proposed. Simulations demonstrate that Neuro-NoC can predict the global optimal NoC configuration with high accuracy (88%), sensitivity (97% true positive), and specificity (88% true negative).
Md Farhadur Reza, Tung Thanh Le, Bappaditya Dey, Magdy A. Bayoumi, Danella Zhao
ISCAS5
2017 Dark silicon-power-thermal aware runtime mapping and configuration in heterogeneous many-core NoC
abstract
To address power-thermal-dark silicon issues in many-core chip, run time task-resource and voltage co-allocation with reconfigurable network-on-chip (NoC) framework for energy and hotspots minimization is proposed in this work. At runtime, the global manager with the help of proposed MinEnergy mapping algorithm reconfigures the NoC links bandwidth and nodes voltage-level and power-gated the resources depending on the traffic demand and resource statistics collected from the distributed cluster-managers. MinEnergy mapping algorithm minimizes overall chip power and thermal hotspots in heterogeneous large-scale NoC. We have formulated the mapping and configuration problem into a linear optimization model and implemented a traditional minimum-path contiguous mapping for comparisons. Simulations show that MinEnergy dynamic mapping solution is 80-90% close to the optimal solution, and significantly better than the minimum-path mapping solution.
Md Farhadur Reza, Danella Zhao, Magdy A. Bayoumi
ISCAS2
2017 Multi-objective Task Mapping Approach for Wireless NoC in Dark Silicon Age
abstract
Hybrid Wireless Network-on-Chip (HWNoC) provides high bandwidth, low latency and flexible topology configurations, making this emerging technology a scalable communication fabric for future Many-Core System-on-Chips (MCSoCs). On the other hand, dark silicon is dominating the chip footage of upcoming MCSoCs since Dennard scaling fails due to the voltage scaling problem that results in higher power densities. Moreover, congestion avoidance and hot-spot prevention are two important challenges of HWNoC-based MCSoCs in dark silicon age, Therefore, in this paper, a novel task mapping approach for HWNoC is introduced in order to first balance the usage of wireless links by avoiding congestion over wireless routers and second spread temperature across the whole chip by utilizing dark silicon. Simulation results show significant improvement in both congestion and temperature control of the system, compared to state-of-the-art works.
Amin Rezaei 0001, Danella Zhao, Masoud Daneshtalab, Hai Zhou 0001
PDP2
2016 Shift sprinting: fine-grained temperature-aware NoC-based MCSoC architecture in dark silicon age
abstract
Reliability is a critical feature of chip integration and unreliability can lead to performance, cost, and time-to-market penalties. Moreover, upcoming Many-Core System-on-Chips (MCSoCs), notably future generations of mobile devices, will suffer from high power densities due to the dark silicon problem. Thus, in this paper, a novel NoC-based MCSoC architecture, called Shift Sprinting, is introduced in order to reliably utilize dark silicon under the power budget constraint. By employing the concept of distributional sprinting, our proposed architecture provides Quality of Service (QoS) to efficiently run real-time streaming applications in mobile devices. Simulation results show meaningful gain in performance and reliability of the system compared to state-of-the-art works.
Amin Rezaei 0001, Danella Zhao, Masoud Daneshtalab, Hongyi Wu
DAC2
2016 Task-Resource Co-Allocation for Hotspot Minimization in Heterogeneous Many-Core NoCs
abstract
To fully exploit the massive parallelism of many cores, this work tackles the problem of mapping large-scale applications onto heterogeneous on-chip networks (NoCs) to minimize the peak workload for energy hotspot avoidance. A task-resource co-optimization framework is proposed which configures the on-chip communication infrastructure and maps the applications simultaneously and coherently, aiming to minimize the peak load under the constraints of computation power and communication capacity and a total cost budget of on-chip resources. The problem is first formulated into a linear programming model to search for optimal solution. A heuristic algorithm is further developed for fast design space exploration in extremely large-scale many-core NoCs. Extensive simulations are carried out under real-world benchmarks and randomly generated task graphs to demonstrate the effectiveness and efficiency of the proposed schemes.
Md Farhadur Reza, Danella Zhao, Hongyi Wu
ACM Great Lakes Symposium on VLSI2
2016 Efficient Congestion-Aware Scheme for Wireless on-Chip Networks
abstract
Wireless NoC is becoming popular to be a promising future on-chip interconnection network as a result of high bandwidth, low latency and flexible topology configurations provided by this emerging technology. Nonetheless, congestion occurrence in wireless routers negatively affects the usability of high speed wireless links and considerably increases the network latency, therefore, in this paper, a congestion-aware platform (CAP-W) is introduced for wireless NoCs in order to reduce both internal and external congestions. The whole platform of CAP-W consists of an adaptive routing algorithm that balances utilization of wired and wireless networks, a dynamic task mapping approach that tries to minimize congestion probability, and a task migration strategy that considers dynamic variation of application behaviors. Simulation results show significant gain in congestion control over PEs of wireless NoC, compared to state-of-the-art works.
Amin Rezaei 0001, Masoud Daneshtalab, Maurizio Palesi, Danella Zhao
PDP4
2015 Dynamic Application Mapping Algorithm for Wireless Network-on-Chip
abstract
Because of high bandwidth, low latency and flexible topology configurations provided by wireless NoC, this emerging technology is gaining momentum to be a promising future on-chip interconnection paradigm. However, congestion occurrence in wireless routers reduces the benefit of high speed wireless links and significantly increases the network latency, therefore, in this paper, a Dynamic Application Mapping Algorithm (DAMA) is introduced for wireless NoCs in order to reduce both internal and external congestion. DAMA has three key steps: finding the first node to map, choosing the first task to be mapped onto the first node, and allocation of the remaining tasks to the remaining nodes. Simulation results show significant gain in the mapping cost functions compared to state-of-the-art works.
Amin Rezaei 0001, Masoud Daneshtalab, Danella Zhao, Farshad Safaei, Xiaohang Wang 0001, Masoumeh Ebrahimi
PDP3
2012 DuSCA: A multi-channeling strategy for doubling communication capacity in wireless NoC
abstract
To bridge the widening gap between computation requirements and communication efficiency faced by many-core chips, Wireless Network-on-Chip (WiNoC) has been proposed by using ultra-wideband interconnect. While prior research has demonstrated the salient features of WiNoC as high perlink data rate, high accumulated bandwidth, high flexibility, low overhead and low power consumption, this research aims to develop a multi-access WiNoC to substantially improve the end-to-end performance of on-chip communication. Enabled by time hopping PPM multi-channel capability, we propose an efficient multi-channel distribution and arbitration scheme for improving communication concurrency and resolving channel competition among multiple users to achieve the desired network performance. Our simulation studies based on synthetic traffics demonstrate the efficiency, cost effectiveness and scalability of the channel arbitration scheme and the promising network performance of WiNoC.
Yi Wang 0007, Danella Zhao, Jian Li 0059
ICCD2
2011 Design of multi-channel wireless NoC to improve on-chip communication capacity!
abstract
Many-core chip design has become a popular means to sustain the exponential growth of chip-level computing performance. The main advantage lies in the exploitation of parallelism, distributively and massively. Consequently, the on-chip communication fabric becomes the performance determinant. In the meantime, the introduction of Ultra-Wideband (UWB) interconnect brings in the new opportunity for giga-bps communication bandwidth, milliwatts communication power, and low cost implementation for millimeter range on-chip communication for future chip generations. In this paper, we study multi-channel wireless Network-on-Chip (McWiNoC) with ultra-short RF/wireless links for multi-hop communication. We first present the benefit of high bandwidth, low latency and flexible topology configurations provided by this new on-chip inter-connection network. We then propose a distributed and deadlock-free location based routing scheme. We further design an efficient channel arbitration scheme to grant multi-channel access. With a few representative synthetic traffic patterns and SPLASH-II benchmarks, we demonstrate that McWiNoC can achieve 23.3% average performance improvement and 65.3% average end-to-end latency reduction over a baseline NoC of 8 × 8 metal wired mesh.
Danella Zhao, Yi Wang 0007, Jian Li 0059, Takamaro Kikkawa
NOCS1
2010 A Low-Cost Deadlock-Free Design of Minimal-Table Rerouted XY-Routing for Irregular Wireless NoCs
abstract
To bridge the widening gap between computation requirements of terascale application and communication efficiency faced by gigascale multi-processor system-on-chip devices, a new on-chip communication system, dubbed Wireless Network-on-Chip (WNoC), has been proposed. This work centers on the design of a high-efficient, low-cost, deadlock-free routing scheme for domain-specific irregular mesh WNoCs. A distributed minimal table based routing scheme is designed to facilitate segmented XY-routing. Deadlock-free data transmission is achieved by implementing a new turn classes based buffer ordering scheme. The simulation study demonstrates high routing efficiency, low cost and scalability of the routing scheme and the promising network performance of WNoC.
Ruizhe Wu, Yi Wang 0007, Danella Zhao
NOCS3
2009 Distributed Flow Control and Buffer Management for Wireless Network-on-Chip
abstract
To bridge the widening gap between computation requirements of ubiquitous application and communication efficiency faced by gigascale multi-processor system-on-chip (MP-SoC) devices, a new on-chip communication system, dubbed wireless network-on-chip (WNoC), is proposed by using the recently developed CMOS ultra wideband (UWB) intrachip wireless communication. In this work, we propose a distributed flow control and buffer management strategy to improve WNoC end-to-end performance due to the close coupling between shared medium contention and network congestion. Such a strategy involves multiple mechanisms such as fast forwarding by prioritized contention, congestion aware traffic admission control, dynamic virtual output queuing with a shared buffer. The simulation results demonstrate the promising performance of the approach.
Yi Wang 0007, Danella Zhao
ISCAS2
2008 SD-MAC: Design and Synthesis of a Hardware-Efficient Collision-Free QoS-Aware MAC Protocol for Wireless Network-on-Chip
abstract
To bridge the widening gap between computation requirements and communication efficiency faced by gigascale heterogeneous SoCs in the upcoming ubiquitous era, a new on-chip communication system, dubbed Wireless Network-on-Chip (WNoC), is introduced by using the recently developed CMOS UWB wireless interconnection technology. In this paper, a synchronous and distributed medium access control (SD-MAC) protocol is designed and implemented. Tailored for WNoC, SD-MAC employs a binary countdown approach to resolve channel contention between RF nodes. The receiver_select_sender mechanism and hidden terminal elimination scheme are proposed to increase the throughput and channel utilization of the system. Our simulation study shows the promising performance of SD-MAC in terms of throughput, latency, and network utilization. We further propose a QoS-aware SD-MAC to ensure the serviceability of the entire system and to improve the bandwidth utilization. As a major component of simple and compact RF node design, a MAC unit implements the proposed SD-MAC that guarantees correct operation of synchronized frames while keeping overhead low. The synthesis results demonstrate several attractive features such as high speed, low power consumption, nice scalability and low area cost.
Danella Zhao, Yi Wang 0007
IEEE Trans. Computers1
2008 MTNet: Design of a Wireless Test Framework for Heterogeneous Nanometer Systems-on-Chip
abstract
The rapid migration to nanometer design processes has brought an unprecedented level of integration by allowing system designers to pack a wide variety of functionalities on-chip, namely, systems-on-a-chip (SoCs). In the meantime, electronic testing becomes an enabling technology for this SoC paradigm, since the integration of various core tests is a big challenge, and has revealed a widening gap between design and manufacturing. In particular, the increasing complexity and density of nanometer SoCs have led to the problem of visibility and accessibility in testing. In this paper, we propose an integrated wireless test framework to resolve the acerbated core accessibility problem and to eliminate the incompatibility between the existing SoC test strategies and the next generation billion-transistor SoC specification. Under such a test strategy, the intra-chip wireless links form the wireless test access mechanism (TAM) to transport test data chip-wide. We present a self-configurable multi-hop wireless test micronetwork, dubbed MTNet, with simple and efficient data transmission protocols, and develop a system level design-for-testability structure. Consequently, we propose a geographic routing algorithm to find the test access paths for the deeply embedded cores and a path driven test scheduling algorithm to design and integrate the MTNet-based SoC test access architecture. Extensive simulation study show the feasibility and applicability of MTNet.
Danella Zhao, Yi Wang 0007
IEEE Trans. Very Large Scale Integr. Syst.1
2007 The design and synthesis of a synchronous and distributed MAC protocol for wireless network-on-chip
abstract
To bridge the widening gap between computation requirements and communication efficiency faced by gigascale heterogeneous SoCs in the upcoming ubiquitous era, a new on-chip communication system, dubbed Wireless Network-on-Chip (WNoC), is introduced by using the recently developed CMOS proximity wireless interconnection technology. In this paper, a synchronous and distributed medium access control (SD-MAC) protocol is designed and implemented. Tailored for WNoC, SD-MAC employs a binary countdown approach to resolve channel contention between RF nodes. The receiver select sender mechanism and hidden terminal elimination scheme are proposed to increase the throughput and channel utilization of the system. Our simulation study shows the promising performance of SD-MAC in terms of throughput, latency, and network utilization. As a major component of simple and compact RF node design, a MAC unit implements the proposed SDMAC that guarantees correct operation of synchronized frames while keeping overhead low. The synthesis results demonstrate several attractive features such as high speed, low power consumption, nice scalability and low area cost.
Yi Wang 0007, Danella Zhao
ICCAD2
2007 Design and Implementation of Routing Scheme for Wireless Network-on-Chip
abstract
Modern embedded SoC design uses a rapidly increasing number of processing units for ubiquitous computing. When moving towards a billion-transistor era, ever increasing complexity, density and heterogeneity of embedded components exacerbate on-chip communication, which serves as the fabric to integrate these components and provide a communication mechanism among them. In order to bridge the widening gap between computation requirements and communication efficiency, the authors develop a new on-chip communication system, dubbed wireless network-on-chip (WNoC), by using the RoC technology. This paper mainly focused on the design and hardware implementation of a region-aided routing scheme that realizes loop-free, minimum path cost and high scalability for irregular WNoC infrastructure.
Yi Wang 0007, Danella Zhao
ISCAS2
2007 Using Domain Partitioning in Wrapper Design for IP Cores Under Power Constraints
abstract
This paper presents a novel design method for power-aware test wrappers targeting embedded cores with multiple clock domains. We show that effective partitioning of clock domains combined with bandwidth conversion and gated-clocks would yield shorter test times due to greater flexibility when determining optimal test schedules especially under tight power constraints
Thomas Edison Yu, Tomokazu Yoneda, Danella Zhao, Hideo Fujiwara
VTS3
2006 Design of a wireless test control network with radio-on-chip technology for nanometer system-on-a-chip
abstract
The continued push to smaller geometries, higher frequencies, and larger chip sizes rapidly resulted in an incompatibility between interconnect needs and projected interconnect performance. As stated in the 2003 International Technology Roadmap for Semiconductors (ITRS'03) report, revolutionary interconnect methodologies such as radio frequency (RF)/wireless will deliver the foreseen progress in semiconductor technology. Recent advances in silicon integrated circuit technique are making possible tiny low-cost transceivers to be integrated on chip, namely "radio-on-chip" (ROC) technology. This paper proposes the idea of using wireless radios to transmit test data and control signals to resolve the acerbated core accessibility problem. Three types of wireless test micronetworks are first presented, i.e., miniature wireless local area network (LAN), multihop wireless test control network (MTCNet), and distributed multihop MTCNet. Then, the test control overhead and system resource partitioning in on-chip wireless micronetworks are analyzed. Several challenging system design problems such as RF node placement, core clustering, and control routing are studied, and the test control resources (i.e., the on-chip RF nodes for intrachip communication) are properly distributed and system optimization is performed in terms of test control cost. A simulation study shows the feasibility and applicability of intrachip MTCNet.
Danella Zhao, Shambhu J. Upadhyaya, Martin Margala
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2005 A new SoC test architecture with RF/wireless connectivity
abstract
When moving into the billion-transistor era, the direct or bus interconnects in conventional SoC test control models are rather restricted in not only system performance, but also signal integrity and transmission with continued scaling of the feature size. Recent advances in silicon integrated circuit technology are making possible tiny low-cost transceivers to be integrated on chip. In this paper, we propose a new distributed multihop wireless test control network based on the recent development in "radio-on-chip" technology. Under the multilevel tree structure, the system optimization is performed on control constrained resource partitioning and distribution. Several challenging system design issues, such as RF nodes placement, clustering, and routing are studied, with the integrated resource distribution and system optimization on TAM design and test scheduling. Experimental results show that the proposed algorithm can efficiently minimize the overall testing cost.
Danella Zhao, Shambhu J. Upadhyaya, Martin Margala
ETS1
2005 Dynamically partitioned test scheduling with adaptive TAM configuration for power-constrained SoC testing
abstract
Given a system-on-chip with a set of cores and a set of test resources, and the constraints on the total power consumption during test and the maximum width on the top-level test access mechanism (TAM), it is required to optimize overall testing time of the system. To solve this problem, we first generate a power-constrained test compatibility graph and then construct a set of power-constrained concurrent test sets (PCTSs) to facilitate concurrent testing. We then handle the constrained scheduling by adaptively assigning the cores in parallel to the TAMs with variable width and efficiently utilizing the TAM bandwidth such that the tests in the same PCTS have their lengths close to each other. We concurrently schedule the test sets by dynamically partitioning and allocating the tests, and consequently constructing and updating a set of dynamically partitioned PCTSs. This reduces the test cost in terms of overall test time. Simulation study shows the productivity gained by using our integrated scheduling approach.
Danella Zhao, Shambhu J. Upadhyaya
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2003 Power Constrained Test Scheduling with Dynamically Varied TAM
abstract
In this paper we present a novel scheduling algorithm for testing embedded core-based SoCs. Given test conflicts, power consumption limitation and top level test access mechanism (TAM) constraint, we handle the constrained scheduling in a unique way that adaptively assigns the cores in parallel to the TAMs with variable width and concurrently executes the test sets by dynamic test partitioning, thus reducing the test cost in terms of the overall test time. Through simulation, we show that up to 30% of SoC testing time reduction can be achieved by using our scheduling approach.
Danella Zhao, Shambhu J. Upadhyaya
VTS1
2002 Minimizing concurrent test time in SoC's by balancing resource usage
abstract
We present a novel test scheduling algorithm for embedded core-based SoC's. Given a system integrated with a set of cores and a set of test resources, we select a test for each core from a set of alternative test sets, and schedule it in a way that evenly balances the resource usage, and ultimately reduce the test application time. Furthermore, we propose a novel approach that groups the cores and assigns higher priority to those with smaller number of alternate test sets. In addition, we also extend the algorithm to allow multiple test sets selection from a set of alternatives to facilitate testing for various fault models.
Danella Zhao, Shambhu J. Upadhyaya, Martin Margala
ACM Great Lakes Symposium on VLSI1