VLDB 2026 Research / reviewers in the wild / expert
Yingtao Jiang
dblp:98/2589
· DBLP profile ↗
76ranked-venue papers
1as first author
31since 2021 · last 2026
0000-0001-7453-9365ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 47 · 19 since 2021Artificial intelligence and machine learning · 9 · 7 since 2021Computer networks · 5Applied, interdisciplinary, general and emerging computing · 5 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-author · 2 since 2021Software engineering, systems software and programming languages · 3 · 1 since 2021Human-computer interaction and ubiquitous computing · 2Security and privacy · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Recent advances in brain tumor segmentation: A review of self-supervised learning, federated learning, and graph neural networks
Qiqi Wu, Xing Deng, Haijian Shao, Yingtao Jiang |
Eng. Appl. Artif. Intell. | 4 |
| 2026 | Online detection of hardware Trojan enabled packet tampering attack on network-on-chip: A Bayesian approach
Xiaohang Wang 0001, Ge Cao, Yingtao Jiang, Amit Kumar Singh 0002, Mei Yang 0001, Liang Wang 0020, Fen Guo |
Integr. | 4 |
| 2026 | GAS: An adaptive pruning strategy based on output feature map gradients, activations, and weight sparsity
Sixun Yan, Haijian Shao, Xing Deng, Yingtao Jiang, Ming Zhang 0033, Fei Wang 0082 |
Knowl. Based Syst. | 4 |
| 2026 | CQH-MPN: A Classical-Quantum Hybrid Prototype Network With Fuzzy Proximity-Based Classification for Early Glaucoma DiagnosisabstractGlaucoma is the second leading cause of blindness worldwide and the only form of irreversible vision loss, making early and accurate diagnosis essential. Although deep learning has revolutionized medical image analysis, its dependence on large-scale annotated datasets poses a significant barrier, especially in clinical scenarios with limited labeled data. To address this challenge, we propose a Classical-Quantum Hybrid Mean Prototype Network (CQH-MPN) tailored for few-shot glaucoma diagnosis. CQH-MPN integrates a quantum feature encoder, which exploits quantum superposition and entanglement for enhanced global representation learning, with a classical convolutional encoder to capture local structural features. These dual encodings are fused and projected into a shared embedding space, where mean prototype representations are computed for each class. We introduce a fuzzy proximity-based metric that extends traditional prototype distance measures by incorporating intra-class variability and inter-class ambiguity, thereby improving classification sensitivity under uncertainty. Our model is evaluated on two public retinal fundus image datasets-ACRIMA and ORIGA-under 1-shot, 3-shot, and 5-shot settings. Results show that CQH-MPN consistently outperforms other models, achieving an accuracy of 94.50% $\pm$1.04% on the ACRIMA dataset under the 1-shot setting. Moreover, the proposed method demonstrates significant performance improvements across different shot configurations on both datasets. By effectively bridging the representational power of quantum computing with classical deep learning, CQH-MPN demonstrates robust generalization in data-scarce environments. This work lays the foundation for quantum-augmented few-shot learning in medical imaging and offers a viable solution for real-world, low-resource diagnostic applications. Haijian Shao, Xing Deng, Yingtao Jiang |
IEEE J. Biomed. Health Informatics | 4 |
| 2025 | On Design Space Exploration of Cache System in Multi-Chiplet SystemsabstractWhile multi-chiplet based many-core systems have emerged as a viable solution for heterogeneous integration and addressing manufacturing and technological challenges in the post-Moore’s Law era, their design and optimization remain highly complex and challenging. Among the various subsystems, the cache hierarchy has significant implications for overall system performance, yet its vast design space presents substantial optimization challenges. This complexity arises from factors such as the large number of chiplets in the system, the number of cores per chiplet, memory hierarchy variations, cache size variability, caching strategies, and inter-chiplet interconnection networks. Existing design space exploration methods, such as NN-Baton and IntLP, fail to optimize cache subsystem performance or thoroughly explore the design space. To address these limitations, we propose a novel design space exploration method for cache subsystem optimization. Our approach models cache miss rates and network latency as functions of cache hierarchy and inter-/intra-chiplet interconnection network parameters. We then define an optimization problem to minimize the concurrent average memory access time (C-AMAT) under cost and power consumption constraints. This problem is addressed using a bilevel optimization algorithm, which iteratively solves two independent subproblems: (1) cache subsystem optimization, and (2) inter-chiplet interconnection network optimization. Experimental results show that our method reduces the application execution time by 39.7% and 39.2%, on average, compared to architectures similar to AMD Zen 4 and Intel Sapphire Rapids, respectively, and by $\mathbf{2 5 . 9 1 \%}$ over IntLP. These results underscore the potential of the proposed method for optimizing cache subsystems in future multi-chiplet based many-core systems. Xiaohang Wang 0001, Yingtao Jiang, Amit Kumar Singh 0002 |
DAC | 3 |
| 2025 | LEGOSim: A Unified Parallel Simulation Framework for Multi-chiplet Heterogeneous Integration
Tiantian Lin, Xiaohang Wang 0001, Ling Wang 0005, Zhulin Zheng, Yingtao Jiang, Amit Kumar Singh 0002, Jieming Yin, Sihai Qiu, Mingzhe Zhang 0005, Kui Ren 0001 |
MICRO | 6 |
| 2025 | Adaptive convolutional network pruning through pixel-level cross-correlation and channel independence for enhanced model compression
Haijian Shao, Xing Deng, Yingtao Jiang |
Eng. Appl. Artif. Intell. | 4 |
| 2025 | LUFT-CAN: A lightweight unsupervised learning based intrusion detection system with frequency-time analysis for vehicular CAN bus
Xiaohang Wang 0001, Li Lu 0008, Shuguo Zhuo, Yingtao Jiang, Amit Kumar Singh 0002, Kui Ren 0001, Mei Yang 0001, Kaiwei Wu |
J. Syst. Archit. | 5 |
| 2025 | PIDSNeRF: pose interpolation depth supervision neural radiance fields for view synthesis from challenging input
Haijian Shao, Xing Deng, Yingtao Jiang |
Multim. Tools Appl. | 4 |
| 2025 | On Task Mapping in Multi-chiplet Based Many-Core Systems to Optimize Inter- and Intra-chiplet CommunicationsabstractMulti-chiplet system design, by integrating multiple chiplets/dielets within a single package, has emerged as a promising paradigm in the post-Moore era. This paper introduces a novel task mapping algorithm for multi-chiplet based many-core systems, addressing the unique challenges posed by intra- and inter-chiplet communications under power and thermal constraints. Traditional task mapping algorithms fail to account for the latency and bandwidth differences between these communications, leading to sub-optimal performance in multi-chiplet systems. Our proposed algorithm employs a two-step process: (1) task assignment to chiplets using binary linear programming, leveraging a totally unimodular constraint matrix, and (2) intra-chiplet mapping that minimizes communication latency while considering both thermal and power constraints. This method strategically positions tasks with extensive inter-chiplet communication near interface nodes and centralizes those with predominant intra-chiplet communication. Experimental results demonstrate that the proposed algorithm outperforms existing methods (DAR and IOA) with a 37.5% and 24.7% reduction in execution time, respectively. Communication latency is also reduced by up to 43.2% and 32.9%, compared to DAR and IOA. These findings affirm that the proposed task mapping algorithm aligns well with the characteristics of multi-chiplet based many-core systems, and thus improves optimal performance. Xiaohang Wang 0001, Yingtao Jiang, Amit Kumar Singh 0002, Mei Yang 0001 |
IEEE Trans. Computers | 3 |
| 2025 | On Optimizing Inter- and Intra-Chiplet Interconnection Topologies for Robust Multi-Chiplet SystemsabstractInter- and intra-chiplet interconnection networks play a vital role in the operation of many core systems made of multiple chiplets. However, these networks are susceptible to faults caused by manufacturing defects and attacks resulting from the malicious insertion of hardware Trojans and backdoors. Unlike conventional fault-tolerant or countermeasure methods, this article focuses on optimizing network robustness to withstand both faults and attacks, while considering the constraints of chiplet area and power budget. To achieve this, this article first defines network robustness as a quantifiable measure based on various network parameters, after which an optimization problem is formulated to optimize the robustness of the network topology. To efficiently solve this problem, a reinforcement learning algorithm is proposed. Experimental results demonstrate that the proposed method is capable of generating inter- and intra-chiplet interconnection networks that are significantly more robust than existing topology generation methods. Specifically, the proposed method improves robustness over ButterDonut and Kite, respectively, by an average of 10.88% and 14.06% under random faults and by 9.37% and 7.81% under targeted attacks. These experimental results confirm that the proposed method is capable of generating robust inter- and intra-chiplet interconnection networks that can withstand both faults and attacks. By optimizing the network topology’s robustness, it provides a valuable contribution to the design and security of chiplet-based core systems. Xiaohang Wang 0001, Amit Kumar Singh 0002, Yingtao Jiang, Mei Yang 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2025 | On Improving the Performance of Intra- and Inter-chiplet Interconnection Networks in Multi-chiplet Systems for Accelerating FHE Encrypted Neural Network ApplicationsabstractFully Homomorphic Encryption (FHE) is regarded as a promising way to protect data privacy with encrypted computation. Due to high computation overhead, hardware based FHE accelerators were proposed to speed up FHE applications. To support complicated FHE-encrypted neural network applications, multi-chiplet based FHE accelerators were further proposed for scaling up system size, whereas one of the challenges is designing efficient intra- and inter-chiplet interconnection networks to accelerate data transfer. Conventional regular topologies like mesh or Kite either lead to high inter-chiplet transmission latency or excessive power consumption as these topologies assume uniform bandwidth or radix for nodes/links, ignoring the highly irregular distribution of inter-chiplet communication volumes. On the other hand, the problem of generating customized intra- and inter-chiplet interconnection networks has high complexity and previous network-on-chip topology generation works cannot efficiently improve the performance of intra- and inter-chiplet interconnection networks. In this article, the intra- and inter-chiplet interconnection optimization problem is defined, aiming to minimize the execution time of FHE applications under cost and power constraints. To efficiently solve this problem, we propose a bilevel optimization algorithm, which decomposes the problem into three sub-problems: (1) FHE parameters selection, (2) task-to-core mapping, and (3) intra-/inter-chiplet interconnection network topology generation. These sub-problems are then solved iteratively. Experimental results demonstrate that our proposed method reduces execution time by 51.66%, 43.16%, 39.44%, 43.34%, and 27.70% compared with REED and four multi-chiplet based FHE accelerators with mesh, Kite, Butterfly, and Florets as inter-chiplet interconnection networks. Therefore, the proposed method can effectively accelerate FHE applications on large-scale multi-chiplet systems. Zewei Lai, Jinhui Ye, Xiaohang Wang 0001, Zheang Fu, Amit Kumar Singh 0002, Yingtao Jiang, Kui Ren 0001, Mei Yang 0001, Sihai Qiu, Mingzhe Zhang 0005 |
ACM Trans. Embed. Comput. Syst. | 6 |
| 2025 | ImpRes: implicit residual diffusion models for image super-resolution
Shiyun Zhang, Xing Deng, Haijian Shao, Yingtao Jiang |
Vis. Comput. | 4 |
| 2024 | Toward Precise Robotic Weed Flaming Using a Mobile Manipulator with a BlowtorchabstractRobotic weed flaming is a new and environmentally friendly approach to weed removal in the agricultural field. Using a mobile manipulator equipped with a blowtorch, we design a new system and algorithm to enable effective weed flaming, which requires robotic manipulation with a soft and deformable end effector, as the thermal coverage of the flame is affected by dynamic or unknown environmental factors such as gravity, wind, atmospheric pressure, fuel tank pressure, and pose of the nozzle. System development includes overall design, hardware integration, and software pipeline. To enable precise weed removal, the greatest challenge is to detect and predict dynamic flame coverage in real time before motion planning, which is quite different from a conventional rigid gripper in grasping or a spray gun in painting. Based on the images from two onboard infrared cameras and the pose information of the blowtorch nozzle on a mobile manipulator, we propose a new dynamic flame coverage model. The flame model uses a center-arc curve with a Gaussian cross-section model to describe the flame coverage in real time. The experiments have demonstrated the working system and shown that our model and algorithm can achieve a mean average precision (mAP) of more than 76% in the reprojected images during online prediction. Di Wang 0020, Chengsong Hu, Shuangyu Xie, Joe Johnson, Hojun Ji, Yingtao Jiang, Muthukumar Bagavathiannan, Dezhen Song |
IROS | 6 |
| 2024 | Efficient filter pruning: Reducing model complexity through redundancy graph decomposition
Haijian Shao, Xing Deng, Yingtao Jiang |
Neurocomputing | 4 |
| 2023 | Digitally predicting protein localization and manipulating protein activity in fluorescence images using 4D reslicing GANabstractMOTIVATION: While multi-channel fluorescence microscopy is a vital imaging method in biological studies, the number of channels that can be imaged simultaneously is limited by technical and hardware limitations such as emission spectra cross-talk. One solution is using deep neural networks to model the localization relationship between two proteins so that the localization of one protein can be digitally predicted. Furthermore, the input and predicted localization implicitly reflect the modeled relationship. Accordingly, observing the response of the prediction via manipulating input localization could provide an informative way to analyze the modeled relationships between the input and the predicted proteins. RESULTS: We propose a protein localization prediction (PLP) method using a cGAN named 4D Reslicing Generative Adversarial Network (4DR-GAN) to digitally generate additional channels. 4DR-GAN models the joint probability distribution of input and output proteins by simultaneously incorporating the protein localization signals in four dimensions including space and time. Because protein localization often correlates with protein activation state, based on accurate PLP, we further propose two novel tools: digital activation (DA) and digital inactivation (DI) to digitally activate and inactivate a protein, in order to observing the response of the predicted protein localization. Compared with genetic approaches, these tools allow precise spatial and temporal control. A comprehensive experiment on six pairs of proteins shows that 4DR-GAN achieves higher-quality PLP than Pix2Pix, and the DA and DI responses are consistent with the known protein functions. The proposed PLP method helps simultaneously visualize additional proteins, and the developed DA and DI tools provide guidance to study localization-based protein functions. AVAILABILITY AND IMPLEMENTATION: The open-source code is available at https://github.com/YangJiaoUSA/4DR-GAN. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Lingkun Gu, Yingtao Jiang, Mo Weng, Mei Yang 0001 |
Bioinform. | 3 |
| 2023 | A graphical approach for filter pruning by exploring the similarity relation between feature maps
Haijian Shao, Shengjie Zhai, Yingtao Jiang, Xing Deng |
Pattern Recognit. Lett. | 4 |
| 2023 | Detection of Thermal Covert Channel Attacks Based on Classification of Components of the Thermal Signal FeaturesabstractIn response to growing security challenges facing many-core systems imposed by thermal covert channel (TCC) attacks, a number of threshold-based detection methods have been proposed. In this paper, we show that these threshold-based detection methods are inadequate to detect TCCs that harness advanced signaling and specific modulation techniques. Since the frequency representation of a TCC signal is found to have multiple side lobes, this important feature shall be explored to enhance the TCC detection capability. To this end, we present a pattern-classification-based TCC detection method using an artificial neural network that is trained with a large volume of spectrum traces of TCC signals. After proper training, this classifier is applied at runtime to infer TCCs, should they exist. The proposed detection method is able to achieve a detection accuracy of 99%, even in the presence of the stealthiest TCCs ever discovered. Because of its low runtime overhead ($< 0.187\%$) and low energy overhead ($< 0.072\%$), this proposed detection method can be indispensable in fighting against TCC attacks in many-core systems. With such a high accuracy in detecting TCCs, powerful countermeasures, like the ones based on dynamic voltage and frequency scaling (DVFS), can be rightfully applied to neutralize any malicious core participating in a TCC attack. Xiaohang Wang 0001, Hengli Huang, Ruolin Chen, Yingtao Jiang, Amit Kumar Singh 0002, Mei Yang 0001, Letian Huang |
IEEE Trans. Computers | 4 |
| 2023 | Modeling and Analysis of Thermal Covert Channel Attacks in Many-core SystemsabstractIn a many-core chip, thermal flux and thermal correlation among the cores can be explored to create a thermal covert channel (TCC). In this paper, we provide an analytical model to quickly determine the key TCC performance metrics, in terms of bit error rate (BER), signal to noise ratio (SNR), and channel capacity, without going through lengthy computer simulation and/or physical experiments that are normally needed in current TCC performance studies. According to our model, the TCC’s BER is proportional to the square root of the transmission frequency, which can be explored quantitatively to boost the TCC’s transmission efficiency by letting the TCC’s thermal signal be transmitted at a higher frequency. In addition, our proposed model also links the jamming noise and application of Dynamic Voltage Frequency Scaling (DVFS) to TCC’s BER performance, a feature that can be explored to design/optimize the countermeasures against the TCC attacks. The TCC performance predicted by the proposed theoretical model is found in a good agreement with that obtained from computer simulations, with an average error lower than 7%. Xiaohang Wang 0001, Yingtao Jiang, Amit Kumar Singh 0002, Mei Yang 0001, Letian Huang |
IEEE Trans. Computers | 3 |
| 2023 | RUPA: A High Performance, Energy Efficient Accelerator for Rule-Based Password Generation in Heterogenous Password Recovery SystemabstractThere has been a growing demand for energy efficient and high performance password recovery systems. As password generation and password validation are two integral components of any password recovery system, the former yet lags behind the latter in performance particularly for the case when the popular rule-based password generation method is applied to the heterogeneous CPU-FPGA system. In this paper, we thus present a high performance, energy efficient accelerator to speed up the rule functions in rule-based password generation. Dubbed RUPA, this proposed accelerator explores previously undiscovered computational features and memory access patterns for processing the rule functions. Specially, we show that the rule functions can be mapped to three distinct groups according to their character dependency graphs. Correspondingly, three kinds of datapath units, referred to as rule logic units, are created, and the rule functions from the same group will be processed in their shared rule logic unit. Compared with the state-of-the-art password recovery system built upon a CPU-GPU platform, the FPGA-based RUPA system achieves 5.3x speed improvement and is 33.1x more energy efficient. If RUPA is integrated into the popular password recovery tool John the Ripper (JtR), JtR's rule-based attack performance can soar by more than 48.7x. Peng Liu 0016, Yingtao Jiang |
IEEE Trans. Computers | 4 |
| 2022 | On Evaluation of On-chip Thermal Covert Channel AttacksabstractThermal covert channel (TCC) attacks have been a serious security concern to the use of many-core chips. Severity of these attacks is directly linked to the TCC’s transmission rate and its BER (bit error rate) performance, both of which are impacted by the transmission characteristics of thermal signals and adopted encoding, modulation, and multiplexing schemes. This paper examines, compares, and analyzes various TCCs built upon different combinations of encoding, modulation, and multiplexing. In particular, our study shows that TCC using non-return-to-zero (NRZ) line coding and frequency shift keying (FSK) modulation achieves the highest throughput of 120 bps and BER of below 10%. Jiachen Wang 0011, Xiaohang Wang 0001, Yingtao Jiang, Amit Kumar Singh 0002, Letian Huang, Mei Yang 0001 |
CASES | 3 |
| 2022 | IMSC: Instruction set architecture monitor and secure cache for protecting processor systems from undocumented instructionsabstractAbstract A secure processor requires that no secret, undocumented instructions be executed. Unfortunately, as today's processor design and supply chain are increasingly complex, undocumented instructions that can execute some specific functions can still be secretly introduced into the processor system as flaws or vulnerabilities. To address this problem that may cause potentially serious security breaches, the instruction set architecture (ISA) monitor and secure cache (IMSC) is proposed. As a lightweight solution, IMSC employs an ISA monitor to discover and correct any potential threats imposed by undocumented instructions, and it relies on a secure cache to ensure the credibility of the system. The authors’ case studies have confirmed that IMSC can effectively protect a processor system from being exploited by undocumented instructions and thus provide a trustworthy computing environment, all at low hardware and run‐time costs. Yuze Wang 0001, Peng Liu 0016, Yingtao Jiang |
IET Inf. Secur. | 3 |
| 2022 | Data streaming and traffic gathering in mesh-based NoC for deep neural network acceleration
Binayak Tiwari, Mei Yang 0001, Xiaohang Wang 0001, Yingtao Jiang |
J. Syst. Archit. | 4 |
| 2022 | CNN-Based Hidden-Layer Topological Structure Design and Optimization Methods for Image Classification
Haijian Shao, Yingtao Jiang, Xing Deng |
Neural Process. Lett. | 3 |
| 2022 | On a Consistency Testing Model and Strategy for Revealing RISC Processor's Dark Instructions and VulnerabilitiesabstractOne major security vulnerability of a microprocessor can be attributed to its underlying instruction set architecture (ISA). Generally, it is required that no secret instructions be included in the ISA or implemented in the processor micro-architecture. Such a requirement is particularly important for the reduced instruction set computing (RISC) processors that are widely used nowadays, and applying the proposed consistency testing approach is poised to ensure this requirement is met. Capable of revealing any possible dark instructions (i.e., executable instructions but without clear definitions of their behavior) in RISC processors, a consistency test comes in three phases. During the generation phase, based on the instruction set encoding rules, all the undefined instructions are generated. Even with a smaller test space, this step guarantees the test coverage needed to reveal all the dark instructions that may exist. In the next phase, all the undefined instructions obtained from the previous phase are executed on the processor under test, following a set of persistence strategies; any instruction exhibiting unusual execution result will be deemed suspicious and recorded so. During the last analysis phase, each of those recorded suspicious instructions will be checked and analyzed to decide whether it truly constitutes a dark instruction. We have applied the proposed testing model and strategy to several RISC processors and found that all of them have a few dark instructions previously unknown. The potential vulnerabilities of these processors introduced by their respective dark instructions have thus been evaluated and exposed. Yuze Wang 0001, Peng Liu 0016, Xiaohang Wang 0001, Yingtao Jiang |
IEEE Trans. Computers | 5 |
| 2022 | Performance Optimization of Many-Core Systems by Exploiting Task Migration and Dark Core AllocationabstractAs an effective scheme often adopted for performance tuning in many-core processors, task migration provides an opportunity for “hot” tasks to be migrated to run on a “cool” core that has a lower temperature. When a task needs to migrate from one processor core to another, the migration can embark on numerous modes defined by the migration paths undertaken and/or the destinations of the migration. Selecting the right migration mode that a task shall follow has always been difficult, and it can be more challenging with the existence of dark cores that can be called back to service (reactivated), which ushers in additional task migration modes. Previous works have demonstrated that dark cores can be placed near the active cores to reduce power density so that the active cores can run at higher voltage/frequency levels for higher performance. However, the existing task migration schemes neither consider the impact of dark cores on each application's performance, nor exploit performance trade-off under different migration modes. Unlike the existing task migration schemes, in this article, a runtime task migration algorithm that simultaneously takes both migration modes and dark cores into consideration is proposed, and it essentially has two major steps. In the first step, for a specific migration mode that is tied to an application whose tasks need to be migrated, the number of dark cores is determined so that the overall performance is maximized. The second step is to find an appropriate core region and its location for each application to optimize the communication latency and computation performance; during this step, focus is placed on reducing the fragmentation of the free core regions resulting from the task migration. Experimental results have confirmed that our approach achieves over 50 percent reduction in total response time when compared to recently proposed thermal-aware runtime task migration approachess. Shengyan Wen, Xiaohang Wang 0001, Amit Kumar Singh 0002, Yingtao Jiang, Mei Yang 0001 |
IEEE Trans. Computers | 4 |
| 2022 | Detection of and Countermeasure Against Thermal Covert Channel in Many-Core SystemsabstractThe thermal covert channels (TCCs) in many-core systems can cause detrimental data breaches. In this article, we present a three-step scheme to detect and fight against such TCC attacks. Specifically, in the detection step, each core calculates the spectrum of its own CPU workload traces that are collected over a few fixed time intervals, and then it applies a frequency scanning method to detect if there exists any TCC attack. In the next positioning step, the logical cores running the transmitter threads are located. In the last step, the physical CPU cores suspiciously engaging in a TCC attack have to undertake dynamic voltage frequency scaling (DVFS) such that any possible TCC trace will be essentially wiped out. Our experiments have confirmed that on average 97% of the TCC attacks can be detected, and with the proposed defense, the packet error rate (PER) of a TCC attack can soar to more than 70%, literally shutting down the attack in practical terms. The performance penalty caused by the inclusion of the proposed DVFS countermeasures is found to be only 3% for an$8\times 8$many-core system. Hengli Huang, Xiaohang Wang 0001, Yingtao Jiang, Amit Kumar Singh 0002, Mei Yang 0001, Letian Huang |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2022 | Combating Stealthy Thermal Covert Channel Attack With Its Thermal Signal Transmitted in Direct Sequence Spread SpectrumabstractMany-core systems are susceptible to attacks launched by thermal covert channel (TCC) attacks. Detection of TCC attacks often relies on the use of threshold-based approaches or variants, and a countermeasure to thwart the channel can be applied only after an attack is deemed to be present. In this article, we describe a direct sequence spread spectrum (DSSS)-based TCC, where its thermal data are modulated by a pseudo-random bit sequence. Unfortunately, such DSSS-based TCC has an extremely low signal strength that the signal is nearly indistinguishable from the noise and thus cannot be detected by any existing threshold-based detection methods. To combat this stealthy TCC, we propose a novel detection scheme that lets the received signal pass through a differential filter where irrelevant frequency components occupied mainly by the noise gets eliminated and the filtered signal is next compared against a threshold for successful detection. Experimental results show that the DSSS-based TCC can effectively survive detection by the existing detection methods with its BER as low as 4%. In contrast, with the proposed detection and countermeasure applied, the detection accuracy jumps to 89%, and the BER of the DSSS-based TCC soars to 50%, which indicates that the TCC is practically shut down. Xiaohang Wang 0001, Yingtao Jiang, Amit Kumar Singh 0002, Mei Yang 0001, Letian Huang |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2022 | Secured Data Transmission Over Insecure Networks-on-Chip by Modulating Inter-Packet DelaysabstractAs the network-on-chip (NoC) integrated into an SoC design can come from an untrusted third party, there is a growing risk that data integrity and security get compromised when supposedly sensitive data flows through such an untrusted NoC. We thus introduce a new method that can ensure secure and secret data transmission over such an untrusted NoC. Essentially, the proposed scheme relies on encoding binary data as delays between packets travelling across the source and destination pair. The maximum data transmission rate of this inter-packet-delay (IPD)-based communication channel can be determined from the analytical model developed in this article. To further improve the undetectability and robustness of the proposed data transmission scheme, a new block coding method and communication protocol are also proposed. Experimental results show that the proposed IPD-based method can achieve a packet error rate (PER) of as low as 0.3% and an effective throughput of$\boldsymbol {2.3\times 10^{5}}$b/s, outperforming the methods of thermal covert channel, cache covert channel, and circuit-based encryption and, thus, is suitable for secure data transmission in unsecure systems. Jiaen Xu, Xiaohang Wang 0001, Yingtao Jiang, Amit Kumar Singh 0002, Chongyan Gu, Letian Huang, Mei Yang 0001, Shunbin Li |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2021 | Multi-class Cardiovascular Disease Detection and Classification from 12-Lead ECG Signals Using an Inception Residual NetworkabstractA number of deep neural network (DNN)-based models have been applied to help classify and detect severe cardiovascular diseases using 12-lead electrocardiogram (ECG) signals. These models, however, suffer from poor performance in detecting one or two specific cardiac abnormalities, like ST-segment abnormalities, with their accuracy lingering at only around 60%, which in turn limits their applicability in clinical practice. In this paper, we show that the convolution layers of these DNN models can cause the diminishment of the key features of ST-segment abnormalities, making it, compared to cardiac arrhythmia and normal sinus rhythm (NSR), hard to be classified out from the ECG. Correspondingly, we, in this paper, propose a novel DNN-based model that is able to achieve high accuracy of detecting 9 classes of rhythms. In specific, the moving averages of the ECG signals are used as the second input to the proposed model so that it can help fully mine the features relevant to the ST abnormalities. In addition, as opposed to the convolution layers, our model utilizes the inception-residual layers to preserve features from the shallow layers for reuse in the deep layers. On top of these network architecture improvements, we further introduce a customized differential layer so that all the relevant features can be preserved and/or amplified for classification purposes. Trained with CPSC 2018 dataset, our proposed model is able to accurately classify the 9 rhythm classes, with an F1 score as high as 89.7%, up by merely 4.2% from the best result reported in the literature. As far as the ST segment abnormalities are concerned, the proposed model achieves an 89.6% F1 score for STD and an 80.8% F1 score for STE, which is 7.8% and 13.1% respectively higher than that of the best result. The proposed method is thus poised to become a viable solution for cardiovascular health monitoring with the increasing availability of portable and home-based ECG devices. Jian Ni, Yingtao Jiang, Shengjie Zhai, Amei Amei, Dieu-My. T. Tran, Lijie Zhai, Yu Kuang |
COMPSAC | 2 |
| 2021 | An enhanced planned obsolescence attack by aging networks-on-chip
Yinyuan Zhao, Xiaohang Wang 0001, Yingtao Jiang, Liang Wang 0020, Amit Kumar Singh 0002, Letian Huang, Mei Yang 0001 |
J. Syst. Archit. | 3 |
| 2020 | On Countermeasures Against the Thermal Covert Channel Attacks Targeting Many-core SystemsabstractAlthough it has been demonstrated in multiple studies that serious data leaks could occur to many-core systems thanks to the existence of the thermal covert channels (TCC), little has been done to produce effective countermeasures that are necessary to fight against such TCC attacks. In this paper, we propose a three-step countermeasure to address this critical defense issue. Specifically, the countermeasure includes detection based on signal frequency scanning, positioning affected cores, and blocking based on Dynamic Voltage Frequency Scaling (DVFS) technique. Our experiments have confirmed that on average 98% of the TCC attacks can be detected, and with the proposed defense, the bit error rate of a TCC attack can soar to 92%, literally shutting down the attack in practical terms. The performance penalty caused by the inclusion of the proposed countermeasures is only 3% for an 8×8 system. Hengli Huang, Xiaohang Wang 0001, Yingtao Jiang, Amit Kumar Singh 0002, Mei Yang 0001, Letian Huang |
DAC | 3 |
| 2020 | Efficient On-Chip Multicast Routing based on Dynamic Partition MergingabstractNetworks-on-chips (NoCs) have become the mainstream communication infrastructure for chip multiprocessors (CMPs) and many-core systems. The commonly used parallel applications and emerging machine learning-based applications involve a significant amount of collective communication patterns. In CMP applications, multicast is widely used in multithreaded programs and protocols for barrier/clock synchronization and cache coherence. Multicast routing plays an important role on the system performance of a CMP. Existing partition-based multicast routing algorithms all use static destination set partition strategy which lacks the global view of path optimization. In this paper, we propose an efficient Dynamic Partition Merging (DPM)-based multicast routing algorithm. The proposed algorithm divides the multicast destination set into partitions dynamically by comparing the routing cost of different partition merging options and selecting the merged partitions with lower cost. The simulation results of synthetic traffic and PARSEC benchmark applications confirm that the proposed algorithm outperforms the existing path-based routing algorithms. The proposed algorithm is able to improve up to 23% in average packet latency and 14% in power consumption against the existing multipath routing algorithm when tested in PARSEC benchmark workloads. Binayak Tiwari, Mei Yang 0001, Yingtao Jiang, Xiaohang Wang 0001 |
PDP | 3 |
| 2020 | On hardware-trojan-assisted power budgeting system attack targeting many core systems
Xiaohang Wang 0001, Yingtao Jiang, Liang Wang 0020, Mei Yang 0001, Amit Kumar Singh 0002, Terrence S. T. Mak |
J. Syst. Archit. | 3 |
| 2020 | Combating Enhanced Thermal Covert Channel in Multi-/Many-Core Systems With Channel-Aware JammingabstractAs a means to thwart thermal covert channel attack in a multi-/many-core system, a strong heat noise whose frequency band coincides with that occupied by the thermal covert channel is injected to jam the channel. However, this undiscriminating channel jamming-based countermeasure will fail if a thermal covert channel is allowed to change its transmission frequency dynamically in response to the jamming. To combat this enhanced thermal covert channel, a more advanced countermeasure is needed and thus proposed that checks the frequency spectrum and tracks any possible covert channel. Only after a channel is detected to be susceptible, a thermal noise with this channel frequency is then emitted to jam the covert channel. The communication protocols and frequency changing scheme pertaining to this enhanced thermal covert channel are described in this article. The experimental results confirm that, when the proposed countermeasure is applied, the enhanced thermal covert channel, much more resilient to jamming, suffers from an extremely high packet error rate (PER), which makes any meaningful data leakage practically impossible. As the proposed countermeasure method is poised to contain dangerous thermal covert channel attacks with an anti-jamming capability, it lends itself well to secure multi-/many-core systems. Jiachen Wang 0011, Xiaohang Wang 0001, Yingtao Jiang, Amit Kumar Singh 0002, Letian Huang, Mei Yang 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2018 | Improving the efficiency of thermal covert channels in multi-/many-core systemsabstractIn many-core chips seen in mobile computing, data center, AI, and elsewhere, thermal covert channels could be established to transmit data (e.g., passwords), supposedly to be kept secret and private. Effectiveness of a thermal covert channel, measured by its transmission rate and bit error rate (BER), is so much dependent on the thermal noise/interference imposed on the channel. In this paper, we present a few techniques to improve the capacity of thermal covert channel by overcoming the thermal interference. In particular, data in a thermal covert channel are encoded and represented following a new thermal signaling scheme where logic value, 0 or 1, modules the thermal signals duty cycle. Next, we show in this study that proper selection of transmission frequency can significantly minimize thermal interference. In addition, we propose a robust end-to-end communication protocol for reliable communications. Our experiments have confirmed that, compared to an existing thermal covert channel attack [1] [2], a thermal covert channel enhanced with all the improvements proposed in this study is seeing significant BER reduction (by as much as 75%), and transmission rate boost (by more than threefold). Building such a strong thermal covert channel is the key step towards developing robust defense and countermeasures against information leaking over thermal covert channel. Zijun Long, Xiaohang Wang 0001, Yingtao Jiang, Guofeng Cui, Terrence S. T. Mak |
DATE | 3 |
| 2018 | Effectiveness of HT-assisted sinkhole and blackhole denial of service attacks targeting mesh networks-on-chip
Xiaohang Wang 0001, Yingtao Jiang, Mei Yang 0001, Terrence S. T. Mak, Amit Kumar Singh 0002 |
J. Syst. Archit. | 3 |
| 2017 | A Scalable Parameterized NoC Emulator Built Upon Xilinx Virtex-7 FPGAabstractA number of critical design decisions, such as network topology, buffer sizes, flow control mechanism and so on so forth, have to be evaluated in any NoC the design. Designs and verifications of NoCs are based on either software simulations, which are extremely slow and inaccurate for complex models, or hardware emulations using low/mid-class FPGAs, where the scalability of the NoC system is intensively restricted by the limited on-chip resources. In this paper, we implement a parameterized NoC emulation system, capable of verifying complete functionality of routers and monitoring network performance and buffer usages in real time, on a hardware platform featuring a super large FPGA chip, Xilinx Virtex-7. This FPGA-based emulator also shall be configured to support multiple routing algorithms and packet transferring mechanisms. Compared to the existing emulators, it requires less user effort to measure the performance under various application scenarios, and it scales well to emulate large NoC designs. Currently, this emulator has been used to study NoCs with sizes of 4x4 and 8x8. For the case of 4x4 (8X8) NoC emulator, data transfers between routers can run at over 50MHz, and only occupies about 6% (25%) of the FPGA logic block resources. Yingtao Jiang, Mei Yang 0001, Louie De Luna |
ICSEng | 2 |
| 2017 | HRC: A 3D NoC Architecture with Genuine Support for Runtime Thermal-Aware Task ManagementabstractIn spite of escalating thermal challenges imposed by high power consumption, most reported 3D Network-on-chip (NoC) systems that adopt classic 3D cube (mesh) topology are unable to tackle the thermal management issues directly at the architectural level. Rather, to avoid chip being overheated, tasks running in a “hot” node have to be migrated to a “cooler” one, resulting in increased distance between communicating nodes and ultimately poor performance. In this paper, we propose a new 3D NoC architecture that genuinely supports runtime thermal-aware task management. Dubbed Hierarchical Ring Cluster (HRC), this new hierarchical 3D NoC architecture has three levels across its entire network hierarchy: 1) nodes are grouped as rings, 2) rings are then grouped into cubes, and 3) multiple cubes are connected to form the whole network. Routing in a HRC system is also performed in a hierarchical manner: Paths are set up within rings using low latency circuit switching, and data that need to cross the rings or cubes are routed following dimension-order routing supported by wormhole switching. In this organization, “hot” tasks that need to migrate can move along the rings without incurring increased communication distances. Our experimental results have confirmed that the proposed HRC architecture has a much lower network latency than other known 3D NoC architectures. When working with runtime thermal-aware task migration approaches, HRC can help reduce latency by as much as 80 percent compared to thermal-aware task migration approaches applied to 3D mesh NoC topologies. Xiaohang Wang 0001, Yingtao Jiang, Mei Yang 0001, Terrence S. T. Mak |
IEEE Trans. Computers | 2 |
| 2017 | An Adaptive PAM-4 Analog Equalizer With Boosting-State Detection in the Time DomainabstractThis paper introduces an improved adaptive analog equalizer that is required in high speed serial receivers using four-level pulse amplitude modulation signaling. By performing boosting-state detection in the time domain, the proposed adaptive analog equalizer can effectively overcome a serious problem that the received signal's eye-opening tends to be compromised by the convergence accuracy of the adaptive control loop. To suppress the pattern-dependent jitters (PDJs), an inductor-less, cross-stage feedback structure is employed in the proposed analog equalizer to help broaden its effective tuning bandwidth. Multichannel simulations have confirmed that the proposed equalizer is able to achieve a 42% improvement in eye-height opening when compared with the equalizers employing popular spectrum-comparing schemes. Trellis diagram analyses under different date rates have revealed that the bandwidth of the proposed equalizer can be extended by as much as 60%, thus effectively bringing the PDJs from 44% down to 27.5%. Shunbin Li, Yingtao Jiang, Peng Liu 0016 |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2016 | An energy-efficient scheduling scheme for time-constrained tasks in local mobile clouds
Mei Yang 0001, Yingtao Jiang |
Pervasive Mob. Comput. | 5 |
| 2016 | On Fine-Grained Runtime Power Budgeting for Networks-on-Chip SystemsabstractPower budgeting is an essential aspect of networks-on-chip (NoC) to meet the power constraint for on-chip communications while assuring the best possible overall system performance. For simplicity and ease of implementation, existing NoC power budgeting schemes treat all the individual routers uniformly when allocating power to them. However, such homogeneous power budgeting schemes ignore the fact that the workloads of different NoC routers may vary significantly, and thus may provide excess power to routers with low workloads, whereas insufficient power to those with high workloads. In this paper, we formulate the NoC power budgeting problem in order to optimize the network performance over a power budget through per-router frequency scaling. We take into account of heterogeneous workloads across different routers as imposed by variations in traffic. Correspondingly, we propose a fine-grained solution using an agile algorithm with low time complexity. Frequency of each router is set individually according to its contribution to the average network latency while meeting the power budget. Experimental results have confirmed that with fairly low runtime and hardware overhead, the proposed scheme can help save up to$50$percent application execution time when compared with the latest proposed methods. Xiaohang Wang 0001, Baoxin Zhao, Terrence S. T. Mak, Mei Yang 0001, Yingtao Jiang, Masoud Daneshtalab |
IEEE Trans. Computers | 5 |
| 2015 | Fine-grained runtime power budgeting for networks-on-chipabstractPower budgeting for NoC needs to be performed to meet limited power budget while assuring the best possible overall system performance. For simplicity and ease of implementation, existing NoC power budgeting schemes, irrespective of the fact that the packet arrival rates of different NoC routers may vary significantly, treat all the individual routers indiscriminately when allocating power to them. However, such homogeneous power allocation may provide excess power to routers with low packet arrival rates whereas insufficient power to those with high arrival rates. In this paper, we formulate the NoC power budgeting problem as to optimize the network performance over a power budget through per-router frequency scaling, taking into account of heterogeneous packet arrival rates across different routers as imposed by run time traffic dynamics. Correspondingly, we propose a fine-grained solution using an agile dynamic programming network with a linear time complexity. In essence, frequency of a router is set individually according to its contribution to the average network latency while meeting the power budget. Experimental results have confirmed that with fairly low runtime and hardware overhead, the proposed scheme can help save up to 50% application execution time when compared with the best existing methods. Xiaohang Wang 0001, Terrence S. T. Mak, Mei Yang 0001, Yingtao Jiang, Masoud Daneshtalab |
ASP-DAC | 5 |
| 2015 | An efficient runtime power allocation scheme for many-core systems inspired from auction theory
Xiaohang Wang 0001, Baoxin Zhao, Terrence S. T. Mak, Mei Yang 0001, Yingtao Jiang, Masoud Daneshtalab |
Integr. | 5 |
| 2014 | Agile frequency scaling for adaptive power allocation in many-core systems powered by renewable energy sourcesabstractAs low-power electronics and miniaturization conspire to populate the world with emerging devices, one appealing approach is to power these multi-core/many-core-based devices with energy harvested from various environments. Of the most important issues concerning these devices is how to effectively allocate power budget among the cores competing for power, which is formulated as one specific type of power-performance optimization problem in this paper. We attempt to solve this problem by proposing an Adaptive Power Allocation Technique (APAT) that uses a dynamic programming network. Our goal here is to maximize the overall system performance, taking into account a unique yet challenging fact that, available power budget might have to undergo a significant change when a renewable energy source is scavenging. APAT has a linear time complexity and low hardware overhead. Experiments have confirmed that APAT can reduce 20 ~ 30% of execution time compared to other state-of-the-art power allocation algorithms. In addition, as APAT is quite insensitive to the changing rate of the power, lending itself well for power management in many-core systems powered by energy-harvesting sources. Xiaohang Wang 0001, Mei Yang 0001, Yingtao Jiang, Masoud Daneshtalab, Terrence S. T. Mak |
ASP-DAC | 4 |
| 2014 | Adaptive power allocation for many-core systems inspired from multiagent auction modelabstractScaling of future many-core chips is hindered by the challenge imposed by ever-escalating power consumption. At its worst, an increasing fraction of the chips will have to be shut down, as power supply is inadequate to simultaneously switch all the transistors. This so-called dark silicon problem brings up a critical issue regarding how to achieve the maximum performance within a given limited power budget. This issue is further complicated by two facts. First, high variation in power budget calls for wide range power control capability, whereas most current frequency/voltage scaling techniques cannot effectively adjust power over such a wide range. Second, as the applications' behavior becomes more complicated, there is a pressing need for scalability and global coordination, rendering heuristic-based centralized or fully distributed control schemes inefficient. To address the aforementioned problems, in this paper, a power allocation method employing multiagent auction models is proposed, referred as Hierarchal MultiAgent based Power allocation (HiMAP). Tiles act the role of consumers to bid for power budget and the whole process is modeled by a combinatorial auction, whereas HiMAP finds the Walrasian equilibria. Experimental results have confirmed that HiMAP can reduce the execution time by as much as 45% compared to three competing methods. The runtime overhead and cost of HiMAP are also small, which makes it suitable for adaptive power allocation in many-core systems. Xiaohang Wang 0001, Baoxin Zhao, Terrence S. T. Mak, Mei Yang 0001, Yingtao Jiang, Masoud Daneshtalab, Maurizio Palesi |
DATE | 5 |
| 2014 | On self-tuning networks-on-chip for dynamic network-flow dominance adaptationabstractModern network-on-chip (NoC) systems are required to handle complex runtime traffic patterns and unprecedented applications. Data traffics of these applications are difficult to fully comprehend at design time so as to optimize the network design. However, it has been discovered that the majority of dataflows in a network are dominated by less than 10% of the specific pathways. In this article, we introduce a method that is capable of identifying critical pathways in a network at runtime and can then dynamically reconfigure the network to optimize for network performance subject to the identified dominated flows. An online learning and analysis scheme is employed to quickly discover the emerging dominated traffic flows and provides a statistical traffic prediction using regression analysis. The architecture of a self-tuning network is also discussed which can be reconfigured by setting up the identified point-to-point paths for the dominance dataflows in large traffic volumes. The merits of this new approach are experimentally demonstrated using comprehensive NoC simulations. Compared to the conventional network architectures over a range of realistic applications, the proposed self-tuning network approach can effectively reduce the latency and power consumption by as much as 25% and 24%, respectively. We also evaluated the configuration time and additional hardware cost. This new approach demonstrates the capability of an adaptive NoC to handle more complex and dynamic applications. Xiaohang Wang 0001, Mei Yang 0001, Yingtao Jiang, Peng Liu 0016, Masoud Daneshtalab, Maurizio Palesi, Terrence S. T. Mak |
ACM Trans. Embed. Comput. Syst. | 3 |
| 2013 | On self-tuning networks-on-chip for dynamic network-flow dominance adaptationabstractModern networks-on-chip (NoC) systems are required to handle complex run-time traffic patterns and unprecedented applications. Data traffics of these applications are difficult to be fully comprehended at design-time so as to optimize the network design. However, it has been discovered that the majority data flows in a network are dominated by less than 10% of the specific pathways. In this paper, we introduce a method that is capable of identifying critical pathways in a network at run-time and, then, can dynamically reconfigure the network to optimize for the network performance subjected to the identified dominated flows. An online learning and analysis scheme is employed to quickly discover the emerged dominated traffic flows and provides a statistical traffic prediction using regression analysis. The architecture of a self-tuning network is also discussed which can be reconfigured by setting up the identified point-to-point paths for the dominance data flows in large traffic volumes. The merits of this new approach are experimentally demonstrated using comprehensive NoC simulators. Compared to the conventional network architectures over a range of realistic applications, the proposed self-tuning network approach can effectively reduce the latency and power consumption by as much as 25% and 24%, respectively. We also evaluated the configuration time and additional hardware cost. This new approach demonstrates the capability of an adaptive NoC to handle more complex and dynamic applications. Xiaohang Wang 0001, Terrence S. T. Mak, Mei Yang 0001, Yingtao Jiang, Masoud Daneshtalab, Maurizio Palesi |
NOCS | 4 |
| 2013 | Energy Efficient Run-Time Incremental Mapping for 3-D Networks-on-Chip
Xiaohang Wang 0001, Peng Liu 0016, Mei Yang 0001, Maurizio Palesi, Yingtao Jiang, Michael C. Huang 0001 |
J. Comput. Sci. Technol. | 5 |
| 2013 | Efficient multicast schemes for 3-D Networks-on-Chip
Xiaohang Wang 0001, Mei Yang 0001, Yingtao Jiang, Maurizio Palesi, Peng Liu 0016, Terrence S. T. Mak, Nader Bagherzadeh |
J. Syst. Archit. | 3 |
| 2013 | Avoiding request-request type message-dependent deadlocks in networks-on-chips
Xiaohang Wang 0001, Peng Liu 0016, Mei Yang 0001, Yingtao Jiang |
Parallel Comput. | 4 |
| 2013 | An efficient protocol with synchronization accelerator for multi-processor embedded systems
Jiyang Yu, Peng Liu 0016, Chunming Huang, Yingtao Jiang, Qingdong Yao |
Parallel Comput. | 6 |
| 2012 | On a joint temporal-spatial multi-channel assignment and routing scheme in resource-constrained wireless mesh networks
Yingtao Jiang, Mei Yang 0001 |
Ad Hoc Networks | 3 |
| 2011 | Power-Aware Run-Time Incremental Mapping for 3-D Networks-on-Chip
Xiaohang Wang 0001, Maurizio Palesi, Mei Yang 0001, Yingtao Jiang, Michael C. Huang 0001, Peng Liu 0016 |
NPC | 4 |
| 2011 | Low latency and energy efficient multicasting schemes for 3D NoC-based SoCsabstractIn this paper, two topology oriented multicast routing algorithms, MXYZ and AL+XYZ, are proposed to support multicasting in 3D Networks on Chips (NoCs). In specific, MXYZ is a dimension order multicast routing algorithm that targets 3D NoC systems built upon regular topologies, while AL+XYZ is applicable to NoCs with irregular topologies. If the output channel found by MXYZ is not available (i.e. in the same region), an alternative output channel is used to forward/replicate the packets in AL+XYZ. MXYZ is evaluated against a path based regular topology oriented multicast routing and AL+XYZ against an irregular region oriented multiple unicast routing algorithm. Our experimental results have demonstrated that the proposed MXYZ and AL+XYZ schemes have lower latency and energy consumption than the conventional path based multicast routing and the multiple unicast routing algorithms, meriting them to be more suitable for supporting multicasting in 3D NoC systems. Xiaohang Wang 0001, Maurizio Palesi, Mei Yang 0001, Yingtao Jiang, Michael C. Huang 0001, Peng Liu 0016 |
VLSI-SoC | 4 |
| 2010 | An Efficient Technique for In-order Packet Delivery with Adaptive Routing Algorithms in Networks on ChipabstractAlthough adaptive routing algorithms promise higher communication performance, as compared to deterministic routing algorithms, they suffer from the out-of-order packet delivery problem. In the context of Network on Chip, the area and computational overhead of ordering packets at the destination is high and may reverse any gain achieved through the use of adaptivity of the routing algorithm. In this paper, we describe a novel scheme for ensuring in-order packet delivery while retaining the performance advantages of adaptive routing. The hardware architecture of a router that supports the proposed scheme is described. Although the basic idea in our proposal is topology independent we evaluate and compare the performance of our scheme with both deterministic as well as adaptive routing algorithms for 2D mesh NoC. As compared to the XY routing algorithm, our technique significantly reduces the packet delay and improves the saturation point. The impact on router area and power dissipation is also discussed. Although the power consumption of routers increase, the energy consumption per flit increases less than 2% on average, since the higher performance allows for draining more traffic during a certain time window. Maurizio Palesi, Rickard Holsmark, Xiaohang Wang 0001, Shashi Kumar, Mei Yang 0001, Yingtao Jiang, Vincenzo Catania |
DSD | 6 |
| 2010 | A power-aware mapping approach to map IP cores onto NoCs under bandwidth and latency constraintsabstractIn this article, we investigate the Intellectual Property (IP) mapping problem that maps a given set of IP cores onto the tiles of a mesh-based Network-on-Chip (NoC) architecture such that the power consumption due to intercore communications is minimized. This IP mapping problem is considered under both bandwidth and latency constraints as imposed by the applications and the on-chip network infrastructure. By examining various applications' communication characteristics extracted from their respective communication trace graphs, two distinguishable connectivity templates are realized: the graphs with tightly coupled vertices and those with distributed vertices. These two templates are formally defined in this article, and different mapping heuristics are subsequently developed to map them. In general, tightly coupled vertices are mapped onto tiles that are physically close to each other while the distributed vertices are mapped following a graph partition scheme. Experimental results on both random and multimedia benchmarks have confirmed that the proposed template-based mapping algorithm achieves an average of 15% power savings as compared with MOCA, a fast greedy-based mapping algorithm. Compared with a branch-and-bound--based mapping algorithm, which produces near optimal results but incurs an extremely high computation cost, the proposed algorithm, due to its polynomial runtime complexity, can generate the results of almost the same quality with much less CPU time. As the on-chip network size increases, the superiority of the proposed algorithm becomes more evident. Xiaohang Wang 0001, Mei Yang 0001, Yingtao Jiang, Peng Liu 0016 |
ACM Trans. Archit. Code Optim. | 3 |
| 2009 | Minimum Overlapping Layers and Its Variant for Prolonging Network Lifetime in PMRC-Based Wireless Sensor NetworksabstractThe overlapping layers (OL) scheme proposed in our previous work provides a solution to balance the load of cluster heads at different layers in the PMRC-based wireless sensor networks. However, in the OL scheme, the layer boundary and the overlap range are static through the network lifetime. The network lifetime is still limited by some nodes which have only one candidate cluster head. To overcome this limitation, in this paper, we propose the Minimum overlapping layers (MOL) scheme with gradually changed layer boundary through network lifetime and its variant, the MOL with initial overlap (MOLIO) scheme. The simulation results of the OL, MOL, and MOLIO schemes show that the MOL scheme significantly prolongs the network lifetime than the OL scheme for most transmission ranges and the MOLIO scheme achieves better results than the MOL scheme at larger transmission ranges. Qiaoqin Li, Mei Yang 0001, Yingtao Jiang, Jiazhi Zeng |
CCNC | 4 |
| 2009 | HTSMA: A Hybrid Temporal-Spatial Multi-Channel Assignment Scheme in Heterogeneous Wireless Mesh NetworksabstractA number of multi-channel assignment schemes have recently been proposed to improve the throughput of IEEE 802.11-based multi-hop wireless mesh networks (WMNs). In these schemes, channel coordination is done either through time synchronization across all the hosts, or through the use of a dedicated channel for the transmission of necessary control messages. Either way, excessive system overhead and/or waste of bandwidth resource become unavoidable, undermining the overall network throughput. To maximize the network throughput, we propose a synchronization-free, hybrid temporal-spatial multi-channel assignment scheme in a random heterogeneous network requiring only a single radio interface per host. In this scheme, the gateway is allowed to use all the available channels sequentially in a round-robin fashion. This temporal channel assignment approach ensures that all the neighboring hosts that communicate with the gateway directly shall have a fair access to the gateway. The channel assignment for the remaining wireless hosts is based on the geographical location and channel availability (a spatial approach) to avoid the interference within the communication region of each sender host in its transmission time period. Compared with another multi-channel scheme MMAC, extensive simulation results demonstrate that our proposed scheme can improve the network throughput substantially with the acceptable collision ratio. Ju-Yeon Jo, Mei Yang 0001, Yoohwan Kim, Yingtao Jiang, John Gowens |
GLOBECOM | 5 |
| 2008 | Symmetry-aware placement with transitive closure graphs for analog layout designabstractA new scheme is proposed to use transitive closure graph (TCG) to explore the full symmetry solution space in analog layout design. We define a set of TCG symmetric-feasible conditions and show that it is extremely useful in reducing the solution space. A method is presented for generating random symmetric-feasible TCGs in O(n) time preserving the TCG closure property. Experimental results have confirmed the effectiveness of the proposed symmetry-aware TCG placement algorithm. Chuanjin Richard Shi, Yingtao Jiang |
ASP-DAC | 3 |
| 2008 | Scalable and fault-tolerant network-on-chip design usingthe quartered recursive diagonal torus topologyabstractNetwork-on-a-chip (NoC) is an effective approach to connect and manage the communication between the variety of design elements and intellectual property blocks required in large and complex system-on-chips. In this paper, we propose a new NoC architecture, referred as the Quartered Recursive Diagonal Torus (QRDT), which is constructed by overlaying diagonal torus. Due to its small diameter and rich routing recourses, QRDT is determined to be well suitable to construct highly scalable NoCs. Xianfang Tan, Lei Zhang 0014, Shankar Neelkrishnan, Mei Yang 0001, Yingtao Jiang, Yulu Yang |
ACM Great Lakes Symposium on VLSI | 5 |
| 2007 | An Improved Multi-Layered Architecture and its Rotational Scheme for Large-Scale Wireless Sensor NetworksabstractIn this paper, we propose a highly scalable network architecture, named the Progressive Multi-hop Rotational Clus- tered (PMRC) structure, suitable for the construction of large- scale wireless sensor networks. In the PMRC structure, sensor nodes are partitioned into layers according to their distances (cal- culated using hop counts) to the sink node. A cluster is composed of the nodes located in the same layer and within the transmission range of the cluster head which is located in one layer up. Each cluster here actually selects two cluster heads, which makes the PMRC structure different from another multi-layered structure, MINA (4). Based on the observation that load balancing tends to help balance the energy consumption among different sensor nodes and consequently prolong the network life time, we further propose a rotational scheme functioning at two levels: 1) the two cluster heads in the same cluster rotate to receive and forward data, and 2) clusters at the same layer rotate to sense data. Exten- sive simulations have been conducted to verify the rotation scheme with two selection strategies each specially tailored for one of the two cluster heads required in the PMRC structure. These results have confirmed that the PMRC structure and its rotation scheme together can significantly prolong the node life time and reduce the number of network reconstructions compared with those obtained from a multi-layered structure with single cluster head. I. INTRODUCTION The benefits of low-cost, rapid deployment, self-organization capa- bility and cooperative data-processing have made the wireless sensor networks a practical solution for a wide range of application areas, including military, industry and commercial, environment, health and home (2), (3). The most significant challenge in sensor networks is to overcome the energy constraint since each sensor node has limited power ( ) and it is hard to replenish the power es- pecially in hazardous or hostile application scenarios. The other chal- lenge faced by sensor networks is scalability. Many applications, such as military surveillance and habitat monitoring, require the deploy- ment of large-scale sensor networks (with the number of sensor nodes in the order of hundreds or thousands, or even millions) in a large ge- ographic area, and seamless connectivity to existing infrastructures is usually required when new nodes are added. Other research work for large-scale sensor networks include (7) and (10). In (7), the SAFE protocol was proposed for data dissemination from stationary sensor nodes to mobile sink nodes in large-scale sen- sor networks. The major problems of the SAFE protocol are the large number of states to be maintained at intermediate nodes and the mul- tiple rounds of message exchanges required to set up a path. The two- tier data dissemination (TTDD) protocol (10) is another protocol for disseminating data from stationary sensor nodes to multiple mobile sinks by setting up a grid structure. However, the cost of proactively creating/maintaining the grid structure from all sources to the edge of the sensor field tends to be unbearably high for large sensor networks. In this paper, we follow the layered structure and subsequently pro- pose the Progressive Multi-hop Rotational Clustered (PMRC) struc- ture as an extension to the MINA structure (4). In a PMRC structure, a cluster is formed in the way similar to that in MINA but with a signif- icant distinction: here two cluster heads are selected for each cluster. To balance the load and the energy consumption among different sen- sor nodes, we propose a rotational scheme functioning at two levels: 1) the two cluster heads in the same cluster rotate to receive and forward data, and 2) clusters at the same layer rotate to sense data. Through simulations, we show that the PMRC structure together with its ro- tational scheme outperforms the multi-layered structure with single cluster head in node life time and hence reduce the number of network reconstructions. The rest of the paper is organized as follows. In Section II, we will describe the PMRC structure. In Section III, the problems and algo- rithms of selecting the primary cluster head and the secondary cluster head are discussed. In Section IV, simulation results are presented and discussed. Section V concludes the paper. Mei Yang 0001, Ahmed Abdelal, Yingtao Jiang, Yoohwan Kim |
CCNC | 4 |
| 2007 | SFRIC: A Secure Fast Roaming Scheme in Wireless LAN Using ID-Based CryptographyabstractIn a wireless network composed of multiple access points, a long delay during roaming from one access point to another may cause a disruption for streaming traffic. Roaming in wireless LAN is generally composed of two parts, 1) searching for a new access point and 2) performing authentication at the new access point. To reduce the second part delay, we propose an innovative lightweight authentication scheme called SFRIC (secure fast /foaming using ID-based cryptography). SFRIC employs ID-based cryptography to simplify the authentication process. It performs mutual authentication for the mobile client and AP with a 3-way handshake, then generates a PTK (pairwise transient key) directly without pre-distributing PMK (pairwise master key). It does not require contacting an authentication server or exchanging certificates. SFRIC is composed of two phases. In the first phase (the preparation phase), each mobile client obtains a temporary private key from the PKG (private key generator). In the second phase (the roaming authentication phase), mutual authentication and key distribution are performed. Our preliminary analysis indicates that SFRIC can complete the roaming authentication within a period much less than the critical 20 ms threshold, required for maintaining streaming traffic, when the cryptographic operations are performed in hardware. Yoohwan Kim, Wei Ren 0002, Ju-Yeon Jo, Yingtao Jiang, Jun Zheng 0003 |
ICC | 4 |
| 2007 | A Combinatorial Analysis of Distance Reliability in Star NetworkabstractThis paper addresses a constrained two-terminal reliability measure referred to as distance reliability (DR) between the source node u and the destination node I with the shortest distance, in an n-dimensional star network, Sn. The shortest distance restriction guarantees the optimal communication delay between processors and high link/node utilization across the network. This paper uses a combinatorial approach by limiting the number of node, link and node/link failures. For each failure model, two different cases depending on the relative positions of u and I, are analyzed to compute DR. Furthermore, DR for the antipodal communication, where every node must communicate with its antipode, is investigated as a special case. For this case, lower bound on DR of those disjoint paths is also derived. Shahram Latifi, Yingtao Jiang |
IPDPS | 3 |
| 2007 | Handover Cost Optimization in Traffic Management for Multi-homed Mobile Networks
Jianping Wang 0001, Mei Yang 0001, Xiao-chun Yun, Yingtao Jiang |
UIC | 5 |
| 2007 | Analysis of the Effect of Channel Sub-rating in Unidirectional Call Overflow Scheme for Call Admission in Hierarchical Cellular NetworksabstractKey goals of call admission control in the next generations of wireless networks are the efficient use of the limited wireless resource and the enhanced quality of service (QoS). Knowing that blocking of handoff calls is less desirable than blocking of new calls and the latter inevitably leads to the reduced resource utilization, we propose in this paper a new call admission control for balancing this tradeoff. It is based on the sub-rating strategy (SRS) implemented in conjunction with unidirectional call overflow in hierarchical cellular networks (HCN). With SRS, an occupied full-rate channel is halved to service a call in progress and a handoff call in parallel. Theoretical models based on 1-D Markov process for a microcell and 2-D Markov chain for macrocell analyses are developed to evaluate the performance of the proposed scheme. We compare the results with that obtained by Shan in W.H. Shan and P.Z. Fan, (2003). The experimental results certify higher performance in terms of the forced termination probability. Additionally, we evaluate the voice degradation quality due to subrating. Jun Zheng 0003, Emma E. Regentova, Yingtao Jiang |
VTC Spring | 4 |
| 2007 | Scheduling and optimal voltage selection with multiple supply voltages under resource constraints
Ling Wang 0004, Yingtao Jiang, Henry Selvaraj |
Integr. | 2 |
| 2006 | An Efficient Defense against Distributed Denial-of-Service Attacks using Congestion Path MarkingabstractThe Distributed Denial-of-Service (DDoS) attack is a serious threat in the Internet, and an effective method is needed for distinguishing the attack traffic from the legitimate traffic. In DDoS attacks, the large volume of attack streams cause self-induced congestion or higher utilization of the links. Based on this observation, we propose the Congestion Path Marking (CPM) scheme to identify and drop the attack packets. In this proposed scheme, we store the link utilization information in the packet header so that suspicious attack packets can be distinguished. Each router along the path records its local congestion information, and this information is accumulated to represent the overall congestion level that a packet has experienced. To enable light-weight real-time processing, we employ a RED-like random packet dropping mechanism at the victim's egress router. Through simulations, we show that when the CPM scheme is employed, most of the attack packets in excess of the link capacity are dropped while less than 4% of the legitimate packets are dropped in typical scenarios. The simulation result also shows significantly improved TCP performance when CPM is utilized. Yoohwan Kim, Ahmed Abd El Al, Ju-Yeon Jo, Mei Yang 0001, Yingtao Jiang |
ICC | 5 |
| 2006 | A multilayer perceptron-based medical decision support system for heart disease diagnosis
Yingtao Jiang, Jun Zheng 0003, Chenglin Peng, Qinghui Li |
Expert Syst. Appl. | 2 |
| 2006 | Placement Algorithm in Analog-Layout DesignsabstractAnalog macrocell placement is an NP-hard problem. This paper presents an attempt to solve this problem by using the optimization flow of a genetic algorithm (GA) enhanced by simulated annealing (SA). The bit-matrix representation is employed to improve the search efficiency. In particular, to reduce the solution space without degrading search opportunities, the technique of cell slide is deployed to transform an absolute placement to a relative placement. Following this cell-slide process, it is proved that, for an initial placement, there always exists a solution that can guarantee no occurrence of overlaps among cells and meet any applicable symmetry constraints pertaining to analog layouts. For the optimization of the algorithm parameters, the fractional factorial experiment using an orthogonal array has been conducted, and the exact parameter values are determined using a meta-GA approach. The experimental results show that, compared with the SA approach, the proposed algorithm consumes less computation time while generating higher quality layouts, comparable to expert manual placements Rabin Raut, Yingtao Jiang, Ulrich Kleine |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2006 | Scheduling and Partitioning Schemes for Low Power Designs Using Multiple Supply Voltages
Ling Wang 0004, Yingtao Jiang, Henry Selvaraj |
J. Supercomput. | 2 |
| 2006 | An automated design tool for analog layoutsabstractIn this paper, a layout synthesis tool for the design of analog integrated circuits (ICs) is presented. This tool offers great flexibility that allows analog circuit designers to bring their special design knowledge and experiences into the synthesis process to create high-quality analog circuit layouts. Different from conventional layout systems that are limited to the optimization of single devices, our layout generation tool attempts to optimize more complex modules. This tool includes a complete tool suite that covers the following three major analog physical designs stages. 1) Module Generation: designers can develop and maintain their own technology- and application-independent module generators for subcircuits using an in-house developed description language. 2) Placement: a two-stage placement technique, tailored for the analog placement design, is proposed. In particular, this placement algorithm features a novel genetic placement stage followed by a fast simulated reannealing scheme. 3) Routing: the minimum-Steiner-tree-based global routing is developed, and it is actually integrated into the placement procedure to improve reliability and routability of the placement solutions. Following the global routing, a compaction-based constructive detailed routing finally completes the interconnection of the entire layout. Several testing circuits have been applied to demonstrate the design efficiency and the effectiveness of this tool. Experimental results show that this new layout tool is capable of producing high quality layouts comparable to those manually done by layout experts but with much less design time. Ulrich Kleine, Yingtao Jiang |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2005 | On a Chaotic Neural Network with Decaying Chaotic Noise
Tian-yi Ma, Ling Wang 0004, Yingtao Jiang |
ISNN (1) | 3 |
| 2002 | A DSP-based turbo codec for 3G communication systemsabstractIn this paper, we present a high performance Turbo CODEC implemented using digital signal processor for wireless systems following recommended CDMA2000 standard. At the transmitter side, the Turbo encoder is implemented with a modified 15×13 odd-even interleaver. As modern DSP chips, like TI C64x, are designed with multiple functional units, it is important to fully explore the machine-level parallelism to maximize the usage of available computation sources. To this end, at the receiver side, compared to the algorithm used in TI [7], we have redesigned the decoding algorithms with reduced instruction count by 20%. Furthermore, by transforming a number of add/subtract operations to multiplication operations, our decoder can recycle a few functional units previously unused in TI [7]. This Turbo Codec is capable of encoding one frame in 3.1 microseconds and finishing one decoding stage in 18.1 microseconds on a C64x DSP operating at 400 Mhz. Yingtao Jiang, Yiyan Tang, Dian Zhou |
ICASSP | 1 |
| 2002 | Reduce FFT memory reference for low power applicationsabstractMemory reference is one of the major courses of power consumption incurred in a microprocessor. In this paper, we propose two matrix transformation techniques to reduce the number of memory references in the FFT computation. With the first transformation, all the butterflies sharing the same twiddle factor will be clustered and computed together to eliminate redundant memory access to load twiddle factors. With the second transformation, all remaining (N − 1) butterflies involving the twiddle factor WN0are computed using a register-based breadth-first tree traversal algorithm so that load/store operations of intermediate data arrays are minimized. The proposed two transformations together have led to a novel twiddle-factor-based FFT algorithm. The test results on the TI TMS320C62x digital signal processor show that, for a 32-point FFT, the new algorithm exhibits as much as 20% reduction in clock cycles and an average of 30% reduction in memory access than that of the conventional DIF FFT. Yiyan Tang, Yingtao Jiang |
ICASSP | 2 |
| 2001 | CAM-based label search engine for MPLS over ATM networksabstractThis paper presents a label search engine, built upon multiple hierarchically configured CAM cores, for multiprotocol label switching over ATM networks. Novel cache-based sorting logic, following a linear search algorithm with exponential insertion, is incorporated into each CAM core to efficiently explore the temporal correlations among incoming labels as well as indirect data sorting, resulting in significant power saving and throughput increase. This search engine with 1024 data entries has been designed using a 0.18 /spl mu/m CMOS technology running at 200 MHz with total power consumption of less than 2 W. Yiyan Tang, Yingtao Jiang |
GLOBECOM | 2 |