Letian Huang

dblp:147/7672 · DBLP profile ↗
← Back
33ranked-venue papers
8as first author
17since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 24 · 5 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 3 first-author · 6 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021Security and privacy · 1Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Development of a Tongue Image-Based Machine Learning Tool for the Diagnosis of Colorectal Cancer: A Prospective Multicentre Clinical Cohort Study
abstract
Colorectal cancer (CRC) remains a persistent major global health burden, with traditional diagnostic methods like colonoscopy suffering from suboptimal patient compliance rates. This study develops an intelligent diagnostic model based on tongue images to assist in CRC diagnosis, leveraging the integrative potential of traditional tongue diagnosis and modern machine learning. Between June 2023 and July 2024, we collected and processed 1,389 tongue images from CRC patients and 1,543 from non-colorectal cancer (NCRC) participants. Our methodology combines innovative image segmentation using the Segment Anything Model (SAM) with Grounding DINO, extracts both hand-crafted features (color, texture, shape) and deep learning features via Swin-Transformer, and employs feature fusion and selection techniques. The diagnostic model achieves an accuracy of 87.93% (F1-score: 0.9072) in internal validation. In an independent external cohort of 119 CRC patients and 221 NCRC participants, it demonstrates 85.18% precision (recall: 85%, F1-score: 0.8507). This non-invasive, cost-effective approach demonstrates significant potential as a complementary screening tool for CRC, particularly in regions with limited access to conventional diagnostic resources.
Xiaohe Sun, Letian Huang, Libo Qu, Xing Zeng, Zuojian Zhou, Xufeng Lang, Jie Guo 0001
IEEE J. Biomed. Health Informatics2
2026 GSReuse: Temporally Adaptive Screen-Space Reuse for Accelerating 3D Gaussian Splatting
abstract
Recent advances in 3D Gaussian Splatting (3DGS) have enabled real-time, high-fidelity novel view synthesis. However, rendering each frame independently in a video sequence leads to redundant computations, especially when adjacent frames share significant visual overlaps. This inefficiency is particularly problematic in VR applications, where high frame rates and stereoscopic rendering amplify the per-frame cost. Existing frame interpolation or reuse strategies typically rely on image-domain information and are thus not directly applicable to 3DGS rendering, which is fundamentally point-based. To address this runtime inefficiency, we propose GSReuse, a lightweight and drop-in accelerator that speeds up 3DGS rendering by reusing computations across consecutive frames. GSReuse operates in screen space and introduces only minimal modifications to existing 3DGS rendering pipelines. It also eliminates the need for retraining scene representations. Given the rendered image, depth map, and camera parameters of the current frame, GSReuse estimates reliable Gaussian splatting motion vectors for all pixels and warps reusable contents to the new view. A tile-based filtering and masking strategy is then applied to determine which regions can be safely reused, allowing the 3DGS renderer to skip redundant rendering operations. We evaluate GSReuse on multiple benchmark datasets, showing that GSReuse significantly improves rendering, while maintaining high visual fidelity. Compared to state-of-the-art video frame reuse/generation methods, GSReuse delivers better image quality with much lower latency, facilitating practical deployment of 3DGS in VR applications.
Chengzhi Tao, Jie Guo 0001, Letian Huang, Junqiu Zhu, Daoheng Wang, Yanwen Guo 0001
IEEE Trans. Vis. Comput. Graph.5
2025 360-GS: Layout-Guided Panoramic Gaussian Splatting for Indoor Roaming
abstract
3D Gaussian Splatting (3D-GS) has recently attracted great attention with real-time and photo-realistic renderings. This technique typically takes perspective images as input and optimizes a set of 3D elliptical Gaussians by splatting them onto the image planes, resulting in$2 D$Gaussians. However, applying 3D-GS to panoramic inputs presents challenges in effectively modeling the projection onto the spherical surface of 360° images using 2D Gaussians. In practical applications, input panoramas are often sparse, leading to unreliable initialization of 3D Gaussians and subsequent degradation of 3D-GS quality. In addition, due to the under-constrained geometry of texture-less planes (e.g., walls and floors), 3D-GS struggles to model these flat regions with elliptical Gaussians, resulting in significant floaters in novel views. To address these issues, we propose 360-GS, a novel layout-guided 360° Gaussian splatting for a limited set of panoramic inputs. Instead of splatting 3D Gaussians directly onto the spherical surface, 360-GS projects them onto the tangent plane of the unit sphere and then maps them to the spherical projections.
Jiayang Bai, Letian Huang, Jie Guo 0001, Wen Gong, Yuanqi Li, Yanwen Guo 0001
3DV2
2025 Spectral-GS: Taming 3D Gaussian Splatting with Spectral Entropy
abstract
Recently, 3D Gaussian Splatting (3DGS) has achieved impressive results in novel view synthesis, demonstrating high fidelity and efficiency. However, it easily exhibits needle-like artifacts, especially when increasing the sampling rate. Mip-Splatting tries to remove these artifacts with a 3D smoothing filter for frequency constraints and a 2D Mip filter for approximated supersampling. Unfortunately, it tends to produce over-blurred results, and sometimes needle-like Gaussians still persist. Our spectral analysis of the covariance matrix during optimization and densification reveals that current 3DGS lacks shape awareness, relying instead on spectral radius and view positional gradients to determine splitting. As a result, needle-like Gaussians with small positional gradients and low spectral entropy fail to split and overfit high-frequency details. Furthermore, both the filters used in 3DGS and Mip-Splatting reduce the spectral entropy and increase the condition number during zooming in to synthesize novel view, causing view inconsistencies and more pronounced artifacts. Our Spectral-GS, based on spectral analysis, introduces 3D shape-aware splitting and 2D view-consistent filtering strategies, effectively addressing these issues, enhancing 3DGS’s capability to represent high-frequency details without noticeable artifacts, and achieving high-quality realistic rendering.
Letian Huang, Jie Guo 0001, Jialin Dan, Ruoyu Fu, Yuanqi Li, Yanwen Guo 0001
SIGGRAPH Asia1
2025 TransparentGS: Fast Inverse Rendering of Transparent Objects with Gaussians
abstract
The emergence of neural and Gaussian-based radiance field methods has led to considerable advancements in novel view synthesis and 3D object reconstruction. Nonetheless, specular reflection and refraction continue to pose significant challenges due to the instability and incorrect overfitting of radiance fields to high-frequency light variations. Currently, even 3D Gaussian Splatting (3D-GS), as a powerful and efficient tool, falls short in recovering transparent objects with nearby contents due to the existence of apparent secondary ray effects. To address this issue, we propose TransparentGS, a fast inverse rendering pipeline for transparent objects based on 3D-GS. The main contributions are three-fold. Firstly, an efficient representation of transparent objects, transparent Gaussian primitives, is designed to enable specular refraction through a deferred refraction strategy. Secondly, we leverage Gaussian light field probes (GaussProbe) to encode both ambient light and nearby contents in a unified framework. Thirdly, a depth-based iterative probes query (IterQuery) algorithm is proposed to reduce the parallax errors in our probe-based framework. Experiments demonstrate the speed and accuracy of our approach in recovering transparent objects from complex environments, as well as several applications in computer graphics and vision.
Letian Huang, Dongwei Ye, Jialin Dan, Chengzhi Tao, Kun Zhou 0001, Bo Ren 0003, Yuanqi Li, Yanwen Guo 0001, Jie Guo 0001
ACM Trans. Graph.1
2025 GlossyGS: Inverse Rendering of Glossy Objects With 3D Gaussian Splatting
abstract
Reconstructing objects from posed images is a crucial and complex task in computer graphics and computer vision. While NeRF-based neural reconstruction methods have exhibited impressive reconstruction ability, they tend to be time-comsuming. Recent strategies have adopted 3D Gaussian Splatting (3D-GS) for inverse rendering, which have led to quick and effective outcomes. However, these techniques generally have difficulty in producing believable geometries and materials for glossy objects, a challenge that stems from the inherent ambiguities of inverse rendering. To address this, we introduce GlossyGS, an innovative 3D-GS-based inverse rendering framework that aims to precisely reconstruct the geometry and materials of glossy objects by integrating material priors. The key idea is the use of micro-facet geometry segmentation prior, which helps to reduce the intrinsic ambiguities and improve the decomposition of geometries and materials. Additionally, we introduce a normal map prefiltering strategy to more accurately simulate the normal distribution of reflective surfaces. These strategies are integrated into a hybrid geometry and material representation that employs both explicit and implicit methods to depict glossy objects. We demonstrate through quantitative analysis and qualitative visualization that the proposed method is effective to reconstruct high-fidelity geometries and materials of glossy objects, and performs favorably against State-of-the-Arts.
Shuichang Lai, Letian Huang, Jie Guo 0001, Bowen Pan, Xiaoxiao Long, Jiangjing Lyu, Chengfei Lv, Yanwen Guo 0001
IEEE Trans. Vis. Comput. Graph.2
2024 On the Error Analysis of 3D Gaussian Splatting and an Optimal Projection Strategy
Letian Huang, Jiayang Bai, Jie Guo 0001, Yuanqi Li, Yanwen Guo 0001
ECCV (17)1
2024 Component Dependencies Based Network-on-Chip Test
abstract
On-line test of NoC is essential for its reliability. This paper proposed an integral test solution for on-line test of NoC to reduce the test cost and improve the reliability of NOC. The test solution includes a new partitioning method, as well as a test method and a test schedule which are based on the proposed partitioning method. The new partitioning method partitions the NoC into a new type of basis unit under test (UUT) named as interdependent components based unit under test (iDC-UUT), which applies component test methods. The iDC-UUT have very low level of functional interdependency and simple physical connection, which results in small test overhead and high test coverage. The proposed test method consists of DFT architecture, test wrapper and test vectors, which can speed-up the test procedure and further improve the test coverage. The proposed test schedule reduces the blockage probability of data packets during testing by increasing the degree of test disorder, so as to further reduce the test cost. Experimental results show that the proposed test solution reduces power and area by 12.7% and 22.7% over an existing test solution. The average latency is reduced by 22.6% to 38.4% over the existing test solution.
Letian Huang, Tianjin Zhao, Ziren Wang, Junkai Zhan, Junshi Wang, Xiaohang Wang 0001
IEEE Trans. Computers1
2023 Detection of Thermal Covert Channel Attacks Based on Classification of Components of the Thermal Signal Features
abstract
In response to growing security challenges facing many-core systems imposed by thermal covert channel (TCC) attacks, a number of threshold-based detection methods have been proposed. In this paper, we show that these threshold-based detection methods are inadequate to detect TCCs that harness advanced signaling and specific modulation techniques. Since the frequency representation of a TCC signal is found to have multiple side lobes, this important feature shall be explored to enhance the TCC detection capability. To this end, we present a pattern-classification-based TCC detection method using an artificial neural network that is trained with a large volume of spectrum traces of TCC signals. After proper training, this classifier is applied at runtime to infer TCCs, should they exist. The proposed detection method is able to achieve a detection accuracy of 99%, even in the presence of the stealthiest TCCs ever discovered. Because of its low runtime overhead ($< 0.187\%$) and low energy overhead ($< 0.072\%$), this proposed detection method can be indispensable in fighting against TCC attacks in many-core systems. With such a high accuracy in detecting TCCs, powerful countermeasures, like the ones based on dynamic voltage and frequency scaling (DVFS), can be rightfully applied to neutralize any malicious core participating in a TCC attack.
Xiaohang Wang 0001, Hengli Huang, Ruolin Chen, Yingtao Jiang, Amit Kumar Singh 0002, Mei Yang 0001, Letian Huang
IEEE Trans. Computers7
2023 Modeling and Analysis of Thermal Covert Channel Attacks in Many-core Systems
abstract
In a many-core chip, thermal flux and thermal correlation among the cores can be explored to create a thermal covert channel (TCC). In this paper, we provide an analytical model to quickly determine the key TCC performance metrics, in terms of bit error rate (BER), signal to noise ratio (SNR), and channel capacity, without going through lengthy computer simulation and/or physical experiments that are normally needed in current TCC performance studies. According to our model, the TCC’s BER is proportional to the square root of the transmission frequency, which can be explored quantitatively to boost the TCC’s transmission efficiency by letting the TCC’s thermal signal be transmitted at a higher frequency. In addition, our proposed model also links the jamming noise and application of Dynamic Voltage Frequency Scaling (DVFS) to TCC’s BER performance, a feature that can be explored to design/optimize the countermeasures against the TCC attacks. The TCC performance predicted by the proposed theoretical model is found in a good agreement with that obtained from computer simulations, with an average error lower than 7%.
Xiaohang Wang 0001, Yingtao Jiang, Amit Kumar Singh 0002, Mei Yang 0001, Letian Huang
IEEE Trans. Computers6
2022 On Evaluation of On-chip Thermal Covert Channel Attacks
abstract
Thermal covert channel (TCC) attacks have been a serious security concern to the use of many-core chips. Severity of these attacks is directly linked to the TCC’s transmission rate and its BER (bit error rate) performance, both of which are impacted by the transmission characteristics of thermal signals and adopted encoding, modulation, and multiplexing schemes. This paper examines, compares, and analyzes various TCCs built upon different combinations of encoding, modulation, and multiplexing. In particular, our study shows that TCC using non-return-to-zero (NRZ) line coding and frequency shift keying (FSK) modulation achieves the highest throughput of 120 bps and BER of below 10%.
Jiachen Wang 0011, Xiaohang Wang 0001, Yingtao Jiang, Amit Kumar Singh 0002, Letian Huang, Mei Yang 0001
CASES5
2022 Detection of and Countermeasure Against Thermal Covert Channel in Many-Core Systems
abstract
The thermal covert channels (TCCs) in many-core systems can cause detrimental data breaches. In this article, we present a three-step scheme to detect and fight against such TCC attacks. Specifically, in the detection step, each core calculates the spectrum of its own CPU workload traces that are collected over a few fixed time intervals, and then it applies a frequency scanning method to detect if there exists any TCC attack. In the next positioning step, the logical cores running the transmitter threads are located. In the last step, the physical CPU cores suspiciously engaging in a TCC attack have to undertake dynamic voltage frequency scaling (DVFS) such that any possible TCC trace will be essentially wiped out. Our experiments have confirmed that on average 97% of the TCC attacks can be detected, and with the proposed defense, the packet error rate (PER) of a TCC attack can soar to more than 70%, literally shutting down the attack in practical terms. The performance penalty caused by the inclusion of the proposed DVFS countermeasures is found to be only 3% for an$8\times 8$many-core system.
Hengli Huang, Xiaohang Wang 0001, Yingtao Jiang, Amit Kumar Singh 0002, Mei Yang 0001, Letian Huang
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.6
2022 Combating Stealthy Thermal Covert Channel Attack With Its Thermal Signal Transmitted in Direct Sequence Spread Spectrum
abstract
Many-core systems are susceptible to attacks launched by thermal covert channel (TCC) attacks. Detection of TCC attacks often relies on the use of threshold-based approaches or variants, and a countermeasure to thwart the channel can be applied only after an attack is deemed to be present. In this article, we describe a direct sequence spread spectrum (DSSS)-based TCC, where its thermal data are modulated by a pseudo-random bit sequence. Unfortunately, such DSSS-based TCC has an extremely low signal strength that the signal is nearly indistinguishable from the noise and thus cannot be detected by any existing threshold-based detection methods. To combat this stealthy TCC, we propose a novel detection scheme that lets the received signal pass through a differential filter where irrelevant frequency components occupied mainly by the noise gets eliminated and the filtered signal is next compared against a threshold for successful detection. Experimental results show that the DSSS-based TCC can effectively survive detection by the existing detection methods with its BER as low as 4%. In contrast, with the proposed detection and countermeasure applied, the detection accuracy jumps to 89%, and the BER of the DSSS-based TCC soars to 50%, which indicates that the TCC is practically shut down.
Xiaohang Wang 0001, Yingtao Jiang, Amit Kumar Singh 0002, Mei Yang 0001, Letian Huang
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.6
2022 Secured Data Transmission Over Insecure Networks-on-Chip by Modulating Inter-Packet Delays
abstract
As the network-on-chip (NoC) integrated into an SoC design can come from an untrusted third party, there is a growing risk that data integrity and security get compromised when supposedly sensitive data flows through such an untrusted NoC. We thus introduce a new method that can ensure secure and secret data transmission over such an untrusted NoC. Essentially, the proposed scheme relies on encoding binary data as delays between packets travelling across the source and destination pair. The maximum data transmission rate of this inter-packet-delay (IPD)-based communication channel can be determined from the analytical model developed in this article. To further improve the undetectability and robustness of the proposed data transmission scheme, a new block coding method and communication protocol are also proposed. Experimental results show that the proposed IPD-based method can achieve a packet error rate (PER) of as low as 0.3% and an effective throughput of$\boldsymbol {2.3\times 10^{5}}$b/s, outperforming the methods of thermal covert channel, cache covert channel, and circuit-based encryption and, thus, is suitable for secure data transmission in unsecure systems.
Jiaen Xu, Xiaohang Wang 0001, Yingtao Jiang, Amit Kumar Singh 0002, Chongyan Gu, Letian Huang, Mei Yang 0001, Shunbin Li
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.6
2021 An enhanced planned obsolescence attack by aging networks-on-chip
Yinyuan Zhao, Xiaohang Wang 0001, Yingtao Jiang, Liang Wang 0020, Amit Kumar Singh 0002, Letian Huang, Mei Yang 0001
J. Syst. Archit.6
2021 ECDR$^{2}$2: Error Corrector and Detector Relocation Router for Network-on-Chip
abstract
Network-on-chip (NoC) is commonly used in modern many-core systems due to their high bandwidth and flexibility. As the manufacturing process keeps scaling, the reliability challenge in NoCs increases as well. The error correction code (ECC) is widely adopted in error correction NoCs to improve the data correctness. At the same time, extra stages are introduced in the router pipeline to improve the error correction capability. As a result, conventional error correction routers suffer from high network latency. Motivated by this limitation, i.e., we remove the extra pipeline stages delicately introduced for error correction. We propose an error correction router, called error corrector and detector relocation router (ECDR2), whose architecture optimizes the pipeline flow of the router. As a result, it can achieve both low latency and high error correction. Experimental results show that, compared with the baseline design, ECDR2obtains 13.67 and 39.4 percent less average latency under the uniform traffic pattern and Dedup benchmark, respectively, in an 8 × 8 mesh NoC. The circuit area of ECR is also 7.9 percent less than that of the baseline design under 45-nm technology.
Letian Huang, Chikun Yuan, Junshi Wang, Masoumeh Ebrahimi, Qiang Li 0021
IEEE Trans. Computers1
2021 Runtime Performance Optimization of 3-D Microprocessors in Dark Silicon
abstract
Because the increasing power density is limited by the thermal constraint, multi-core integrated systems have stepped into the dark silicon era recently, meaning not all parts of the system can be powered on at the same time. Dark silicon effects are, especially severe for 3-D microprocessors due to the even higher power density caused by the stacked structures, which greatly limit the system performances. In this article, we propose a greedy based core-cache co-optimization algorithm to optimize the performance of 3-D microprocessors in dark silicon at runtime. The new method determines many runtime settings of the 3-D system on the fly, including the active core and cache bank positions, active cache bank number, and the voltage/frequency (V/f) level of each active core, which optimizes the performance of the 3-D microprocessor under thermal constraint. Because the core-cache settings are co-optimized in the 3-D space and the power budgets are computed dynamically according to the running state of the 3-D microprocessor, the new method leads to a higher system performance compared with the existing methods. Experiments on two 3-D microprocessors show the greedy-based core-cache co-optimization algorithm outperforms the state-of-the-art 3-D dark silicon microprocessor performance optimization method by achieving a higher processing throughput with guaranteed thermal safety.
Hai Wang 0002, Wei Li 0216, Wenjie Qi, Diya Tang, Letian Huang, He Tang 0003
IEEE Trans. Computers5
2020 On Countermeasures Against the Thermal Covert Channel Attacks Targeting Many-core Systems
abstract
Although it has been demonstrated in multiple studies that serious data leaks could occur to many-core systems thanks to the existence of the thermal covert channels (TCC), little has been done to produce effective countermeasures that are necessary to fight against such TCC attacks. In this paper, we propose a three-step countermeasure to address this critical defense issue. Specifically, the countermeasure includes detection based on signal frequency scanning, positioning affected cores, and blocking based on Dynamic Voltage Frequency Scaling (DVFS) technique. Our experiments have confirmed that on average 98% of the TCC attacks can be detected, and with the proposed defense, the bit error rate of a TCC attack can soar to 92%, literally shutting down the attack in practical terms. The performance penalty caused by the inclusion of the proposed countermeasures is only 3% for an 8×8 system.
Hengli Huang, Xiaohang Wang 0001, Yingtao Jiang, Amit Kumar Singh 0002, Mei Yang 0001, Letian Huang
DAC6
2020 Combating Enhanced Thermal Covert Channel in Multi-/Many-Core Systems With Channel-Aware Jamming
abstract
As a means to thwart thermal covert channel attack in a multi-/many-core system, a strong heat noise whose frequency band coincides with that occupied by the thermal covert channel is injected to jam the channel. However, this undiscriminating channel jamming-based countermeasure will fail if a thermal covert channel is allowed to change its transmission frequency dynamically in response to the jamming. To combat this enhanced thermal covert channel, a more advanced countermeasure is needed and thus proposed that checks the frequency spectrum and tracks any possible covert channel. Only after a channel is detected to be susceptible, a thermal noise with this channel frequency is then emitted to jam the covert channel. The communication protocols and frequency changing scheme pertaining to this enhanced thermal covert channel are described in this article. The experimental results confirm that, when the proposed countermeasure is applied, the enhanced thermal covert channel, much more resilient to jamming, suffers from an extremely high packet error rate (PER), which makes any meaningful data leakage practically impossible. As the proposed countermeasure method is poised to contain dangerous thermal covert channel attacks with an anti-jamming capability, it lends itself well to secure multi-/many-core systems.
Jiachen Wang 0011, Xiaohang Wang 0001, Yingtao Jiang, Amit Kumar Singh 0002, Letian Huang, Mei Yang 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.5
2019 Online Path-Based Test Method for Network-on-Chip
abstract
A considerable amount of routers and links remains idle after each mapping application onto the Network-on-Chip based many-core systems. Online path-based test method is a kind of self-test for these idle components. In this paper, a path-based fabric for NoC is firstly proposed. A path serves as the basic component, covering one link and its associated control logic in the routers. One possibility is to apply fault detection on the idle paths, while the other paths continue to operate normally. Moreover, this paper details the hardware implementation, targeting the stuck-at and bridging faults. It suggests a good trade-off between fault coverage, hardware overhead and test time. Experimental results show that the approach achieves 93% of the stuck-at faults in control unit and cover 100% of the stuck-at and bridging faults on the global link within 256 clock cycles.
Junkai Zhan, Letian Huang, Junshi Wang, Masoumeh Ebrahimi, Qiang Li 0021
ISCAS2
2019 Testing aware dynamic mapping for path-centric network-on-chip test
Shuyan Jiang, Junkai Zhan, Junshi Wang, Masoumeh Ebrahimi, Letian Huang
Integr.7
2019 Optimized mapping algorithm to extend lifetime of both NoC and cores in many-core system
Lihuan Wang, Shuyan Jiang, Junshi Wang, Letian Huang
Integr.5
2019 Efficient Design-for-Test Approach for Networks-on-Chip
abstract
To achieve high reliability in on-chip networks, it is necessary to test the network continuously with Built-in Self-Tests (BIST) so that the faults can be detected quickly and the number of affected packets can be minimized. However, BIST causes significant performance loss due to data dependencies. We propose EsyTest, a comprehensive test strategy with minimized influence on system performance. EsyTest tests the data path and the control path separately. The data path test starts periodically, but the actual test performs in the free time slots to avoid deactivating the router for testing. A reconfigurable router architecture and an adaptive fault-tolerant routing algorithm are proposed to guarantee the access to the processing core when the associated router is under test. During the whole test procedure of the network, all processing cores are accessible, and thus the system performance is maintained during the test. At the same time, EsyTest provides a full test coverage for the NoC and a better hardware compatibility comparing with the existing test strategies. Under the PARSEC benchmark and different test frequencies, the execution time increases less than 5 percent at the cost of 9.9 percent more area and 4.6 percent more power in comparison with the execution where no test procedure is applied.
Junshi Wang, Masoumeh Ebrahimi, Letian Huang, Qiang Li 0021, Guangjun Li, Axel Jantsch
IEEE Trans. Computers3
2018 A lifetime-aware mapping algorithm to extend MTTF of Networks-on-Chip
abstract
Fast aging of components has become one of the major concerns in Systems-on-Chip with further scaling of the submicron technology. This problem accelerates when combined with improper working conditions such as unbalanced components' utilization. Considering the mapping algorithms in the Networks-on-Chip domain, some routers/links might be frequently selected for mapping while others are underutilized. Consequently, the highly utilized components may age faster than others which results in disconnecting the related cores from the network. To address this issue, we propose a mapping algorithm, called lifetime-aware neighborhood allocation (LaNA), that takes the aging of components into account when mapping applications. The proposed method is able to balance the wear-out of NoC components, and thus extending the service time of NoC. We model the lifetime as a resource consumed over time and accordingly define the lifetime budget metric. LaNA selects a suitable node for mapping which has the maximum lifetime budget. Experimental results show that the lifetime-aware mapping algorithm could improve the minimal MTTF of NoC around 72.2%, 58.3%, 46.6% and 48.2% as compared to NN, CoNA, WeNA and CASqA, respectively.
Letian Huang, Masoumeh Ebrahimi, Junshi Wang, Shuyan Jiang, Qiang Li 0021
ASP-DAC1
2018 Optimizing dynamic mapping techniques for on-line NoC test
abstract
With the aggressive scaling of submicron technology, intermittent faults are becoming one of the limiting factors in achieving a high reliability in Network-on-Chip (NoC). Increasing test frequency is necessary to detect intermittent faults, which in turn interrupts the execution of applications. On the other hand, the main goal of traditional mapping algorithms is to allocate applications to the NoC platform, ignoring about the test requirement. In this paper, we propose a novel testing-aware mapping algorithm (TAMA) for NoC, targeting intermittent faults on the paths between crossbars. In this approach, the idle links are identified and the components between two crossbars are tested when the application is mapped to the platform. The components can be tested if there is enough time from when the application leaves the platform and a new application enters it. The mapping algorithm is tuned to give a higher priority to the tested paths in the next application mapping. This leaves enough time to test the links and the belonging components that have not been tested in the expected time. Experiment results show that the proposed testing-aware mapping algorithm leads to a significant improvement over FF, NN, CoNA, and WeNA.
Shuyan Jiang, Junshi Wang, Masoumeh Ebrahimi, Letian Huang, Qiang Li 0021
ASP-DAC6
2018 Micro-Architecture Design for Low Overhead Fault Tolerant Network-on-Chip
abstract
Aggressive technology scaling results in reliability decrease of Network-on-Chips (NoCs). Error Correction Codes (ECC) is commonly used to correct error data. It is necessary to balance the reliability of transmissions and the overhead introduced by encoders and decoders. This work utilizes a mechanism reusing decoders in Network Interfaces (NIs), which is named Send-Back ECC. This paper proposes the detailed hardware implementation of Send-Back NoC after hardware overhead optimizing. The design details of the routers and NIs are described. Simulation results prove that the latency of Send-Back ECC is lower than H2H ECC when bit error rate is lower than 0.0002 and always lower than E2E ECC. The hardware overhead of Send-Back ECC is 10.6% lower than H2H ECC, while the energy consumption is also less than both E2E ECC and H2H ECC.
Chikun Yuan, Letian Huang, Junshi Wang, Qiang Li 0021
ISCAS2
2017 An energy efficient approach for C4.5 algorithm using OpenCL design flow
abstract
C4.5 is an important data mining algorithm and has been widely applied in applications including face detection, character recognition and predictive analysis. Although C4.5 is highly accurate, its training process is time-consuming because the traditional algorithm has limited parallelism and frequent data hazards. In order to address these issues, we propose an energy efficient approach for C4.5 training with parallel search, low-latency memory accesses and a folded programming structure to achieve higher acceleration and energy efficiency. This proposed approach greatly improves the C4.5 training process by using a CPU-FPGA heterogeneous platform and OpenCL design flow. Three UCI data sets including spambase, Magic04, and MiniBooNE, are used to evaluate the performance of our method. Experimental results show that our approach has achieved 58X, 63X, 380X higher energy efficiency than the basic serial C4.5 algorithm (serial SPRINT) running on an Intel i7- 3770K CPU and 10X, 1.6X, 12.7X higher energy efficiency than the existing parallel C4.5 algorithm (parallel SPRINT) running on the same CPU-FPGA heterogeneous platform. We also deliver 3X, 5X, 3X higher energy efficiency than the existing CUDA based parallel C4.5 (CUDT) in CPU-GPU heterogeneous platform when using these three datasets.
Hai Peng, Xiaofan Zhang 0001, Letian Huang
FPT3
2017 A low latency fault tolerant transmission mechanism for Network-on-Chip
abstract
Reliability of Network-on-Chip has become a critical problem because of the aggressive technology scaling. A variety of transmission mechanism to tolerant the bit errors has been proposed to achieve the best trade-off between performance and overhead. In this work, a transmission mechanism for NoC based on a novel combination of error detection, error correction, and retransmission is proposed. The light-weight error detectors are integrated into the input ports of routers to check the correctness of head flits and the decoders in Network Interfaces (NIs) are used to correct the errors in any flits of the whole packet. With a very small hardware overhead, the proposed method can guarantee high reachability of packets and greatly decrease End-to-End retransmission. Compared with Hop-to-Hop and End-to-End mechanism, the latency could be highly reduced.
Letian Huang, Xinxin Lin, Junshi Wang, Qiang Li 0021
ISCAS1
2017 Non-blocking BIST for continuous reliability monitoring of Networks-on-Chip
abstract
To achieve high reliability in on-chip networks, frequent runs of Built-in Self-Test allow the detection of and recovery from faults before they affect packets and the system functionality. However, to test routers, wrappers isolate cores from the network which leads to execution blocking and performance loss. In this paper, we propose a design-for-test reconfigurable router with two alternative bypassing channels. The router architecture allows maintaining the connection between cores and the network during the testing procedure by utilizing the bypassing channels. With the help of an adaptive routing algorithm and a testing strategy, networks can be fully tested at a high testing frequency with <;15% increase of execution time.
Junshi Wang, Letian Huang, Masoumeh Ebrahimi, Qiang Li 0021, Guangjun Li, Axel Jantsch
ISCAS2
2016 Non-Blocking Testing for Network-on-Chip
abstract
To achieve high reliability in on-chip networks, it is necessary to test the network as frequently as possible to detect physical failures before they lead to system-level failures. A main obstacle is that the circuit under test has to be isolated, resulting in network cuts and packet blockage which limit the testing frequency. To address this issue, we propose a comprehensive network-level approach which could test multiple routers simultaneously at high speed without blocking or dropping packets. We first introduce a reconfigurable router architecture allowing the cores to keep their connections with the network while the routers are under test. A deadlock-free and highly adaptive routing algorithm is proposed to support reconfigurations for testing. In addition, a testing sequence is defined to allow testing multiple routers to avoid dropping of packets. A procedure is proposed to control the behavior of the affected packets during the transition of a router from the normal to the testing mode and vice versa. This approach neither interrupts the execution of applications nor has a significant impact on the execution time. Experiments with the PARSEC benchmarks on an 8x8 NoC-based chip multiprocessors show only 3 percent execution time increase with four routers simultaneously under test.
Letian Huang, Junshi Wang, Masoumeh Ebrahimi, Masoud Daneshtalab, Xiaofan Zhang 0004, Guangjun Li, Axel Jantsch
IEEE Trans. Computers1
2015 An Efficient KNN Algorithm Implemented on FPGA Based Heterogeneous Computing System Using OpenCL
abstract
Accurate and efficient data classification techniques are of vital importance to many problems, and are rapidly developing in recent decades. K-Nearest Neighbor algorithm (KNN), as one of the most important algorithms, is widely used in text categorization, predictive analysis, data mining and image recognition, etc. To accelerate the algorithm and to optimize the parallel implementation solution are two key issues of KNN. In this paper, we propose a new solution to speed up KNN algorithm on FPGA based heterogeneous computing system using OpenCL. Based on FPGA's parallel pipeline structure, a specific bubble sort algorithm is designed to optimize KNN algorithm. The results have been shown that the efficiency of the solution in our paper is much higher than conventional GPU based KNN algorithm implementation.
Yuliang Pu, Letian Huang
FCCM3
2015 A Routing-Level Solution for Fault Detection, Masking, and Tolerance in NoCs
abstract
Faults may occur in numerous locations of a router in a NoC platform. Compared with the faults in the data path, faults in the control path may cause more severe effects which may result in crashing the entire system. Most of the current efforts in literature focus on disabling a router when a fault is detected. Considering this level of coarse-granularity, the functioning parts of a router have to be unnecessarily disabled which may severely affect the performance or functionality of the on-chip network. To cope with this problem, in this paper we propose a mechanism to tolerate faults in the control path which largely avoid disabling a router as long as the fault is not severe. This mechanism is called DMT, standing for three distinguishing characteristics of the proposed method as fault Detection, fault Masking and fault Tolerance. The proposed mechanism can efficiently detect the faults expressed as illegal turns while it has the capability to tolerate faults without a prior knowledge on where and why a fault has happened.
Xiaofan Zhang 0004, Masoumeh Ebrahimi, Letian Huang, Guangjun Li, Axel Jantsch
PDP3
2013 A Fault-Tolerant Routing Algorithm for NoC Using Farthest Reachable Routers
abstract
As technology scaling, reliability has became one of the key challenges of Network-on-Chip (NoC). Many faulttolerant routing algorithms for NoC are developed to overcome fault components and provide reliable transmission. But proposed routing algorithms do not pay enough attention to find the shortest paths, which increases latency and power consumption. In this paper, a fault-tolerant routing algorithm using new component states diffusion method based on Farthest Reachable Router (FRR) is proposed. This algorithm can reduce latency by finding the shortest paths between source and destination routers. Experiment results verify that FRR routing algorithm can tolerate 79% fault patterns within 3 × 3 and reduce latency by 16-44% compared with FON.
Junshi Wang, Xiaohang Wang 0001, Letian Huang, Terrence S. T. Mak, Guangjun Li
DASC3