EDBT 2026 Demo / reviewers in the wild / expert
Luan H. K. Duong
dblp:153/0625 · also Luan Huu Kinh Duong
· DBLP profile ↗
27ranked-venue papers
3as first author
4since 2021 · last 2023
0000-0003-4731-1896ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 26 · 3 first-author · 4 since 2021Software engineering, systems software and programming languages · 3 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | FAT: An In-Memory Accelerator With Fast Addition for Ternary Weight Neural NetworksabstractConvolutional neural networks (CNNs) demonstrate excellent performance in various applications but have high computational complexity. Quantization is applied to reduce the latency and storage cost of CNNs. Among the quantization methods, binary and ternary weight networks (BWNs and TWNs) have a unique advantage over 8 and 4-bit quantization. They replace the multiplication operations in CNNs with additions, which are favored on in-memory-computing (IMC) devices. IMC acceleration for BWNs has been widely studied. However, though TWNs have higher accuracy and better sparsity than BWNs, IMC acceleration for TWNs has limited research. TWNs on the existing IMC devices are inefficient because the sparsity is not well utilized, and the addition operation is not efficient. In this article, we propose FAT as a novel IMC accelerator for TWNs. First, we propose a sparse addition control unit, which utilizes the sparsity of TWNs to skip the null operations on zero weights. Second, we propose a fast addition scheme based on the memory sense amplifier (SA) to avoid the time overhead of both carry propagation and writing back the carry to memory cells. Third, we further propose a combined-stationary data mapping to reduce the data movement of activations and weights and increase the parallelism across memory columns. Simulation results show that for addition operations at the SA level, FAT achieves$2.00\times $speedup,$1.22\times $power efficiency, and$1.22\times $area efficiency compared with a state-of-the-art IMC accelerator ParaPIM. FAT achieves$10.02\times $speedup and$12.19\times $energy efficiency compared with ParaPIM on networks with 80% average sparsity. Shien Zhu, Luan H. K. Duong, Hui Chen 0016, Di Liu 0002, Weichen Liu 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2022 | ArSMART: An Improved SMART NoC Design Supporting Arbitrary-Turn TransmissionabstractSMART NoC, which transmits unconflicted flits to distant processing elements (PEs) in one cycle through the express bypass, is a high-performance NoC design proposed recently. However, if contention occurs, flits with low priority would not only be buffered but also could not fully utilize bypass. Although there exist several routing algorithms that decrease contentions by rounding busy routers and links, they cannot be directly applicable to SMART since it lacks the support for arbitrary-turn (i.e., the number and direction of turns are free of constraints) routing. Thus, in this article, to minimize contentions and further utilize bypass, we propose an improved SMART NoC, called ArSMART, in which the arbitrary-turn transmission is enabled. Specifically, ArSMART divides the whole NoC into multiple clusters where the route computation is conducted by the cluster controller and the data forwarding is performed by the bufferless reconfigurable router. Since the long-range transmission in SMART NoC needs to bypass the intermediate arbitration, to enable this feature, we directly configure the input and output ports connection rather than applying hop-by-hop table-based arbitration. To further explore the higher communication capabilities, effective adaptive routing algorithms that are compatible with ArSMART are proposed. The route computation overhead, one of the main concerns for adaptive routing algorithms, is hidden by our carefully designed control mechanism. Compared with the state-of-the-art SMART NoC, the experimental results demonstrate an average reduction of 40.7% in application schedule length and 29.7% in energy consumption. Hui Chen 0016, Peng Chen 0027, Luan H. K. Duong, Weichen Liu 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2022 | TAB: Unified and Optimized Ternary, Binary, and Mixed-precision Neural Network Inference on the EdgeabstractTernary Neural Networks (TNNs) and mixed-precision Ternary Binary Networks (TBNs) have demonstrated higher accuracy compared to Binary Neural Networks (BNNs) while providing fast, low-power, and memory-efficient inference. Related works have improved the accuracy of TNNs and TBNs, but overlooked their optimizations on CPU and GPU platforms. First, there is no unified encoding for the binary and ternary values in TNNs and TBNs. Second, existing works store the 2-bit quantized data sequentially in 32/64-bit integers, resulting in bit-extraction overhead. Last, adopting standard 2-bit multiplications for ternary values leads to a complex computation pipeline, and efficient mixed-precision multiplication between ternary and binary values is unavailable. In this article, we propose TAB as a unified and optimized inference method for ternary, binary, and mixed-precision neural networks. TAB includes unified value representation, efficient data storage scheme and novel bitwise dot product pipelines on CPU/GPU platforms. We adopt signed integers for consistent value representation across binary and ternary values. We introduce a bitwidth-last data format that stores the first and second bits of the ternary values separately to remove the bit extraction overhead. We design the ternary and binary bitwise dot product pipelines based on Gated-XOR using up to 40% fewer operations than State-Of-The-Art (SOTA) methods. Theoretical speedup analysis shows that our proposed TAB-TNN is 2.3× fast as the SOTA ternary method RTN, 9.8× fast as 8-bit integer quantization (INT8), and 39.4× fast as 32-bit full-precision convolution (FP32). Experiment results on CPU and GPU platforms show that our TAB-TNN has achieved up to 34.6× speedup and 16× storage size reduction compared with FP32 layers. TBN, Binary-activation Ternary-weight Network (BTN), and BNN in TAB are up to 40.7×, 56.2×, and 72.2× as fast as FP32. TAB-TNN is up to 70.1% faster and 12.8% more power-efficient than RTN on Darknet-19 while keeping the same accuracy. TAB is open source as a PyTorch Extension 1 for easy integration with existing CNN models. Shien Zhu, Luan H. K. Duong, Weichen Liu 0001 |
ACM Trans. Embed. Comput. Syst. | 2 |
| 2021 | Contention-Aware Routing for Thermal-Reliable Optical Networks-on-ChipabstractOptical network-on-chip (ONoC) architecture offers ultrahigh bandwidth, low latency, and low power dissipation for new-generation manycore systems. However, the benefits in communication performance and energy efficiency will be diminished by communication contention. The intrinsic thermal susceptibility is another challenge for ONoC designs. Under on-chip temperature variations, core functional devices suffer from significant thermal-induced optical power loss, which seriously threatens ONoCs' reliability. In this article, we develop novel routing techniques to resolve both issues for ONoCs. By analyzing the thermal effect in ONoCs, we first present a routing criterion at the network level. Combined with device-level thermal tuning, it can implement thermal-reliable ONoCs. Two routing approaches, including a mixed-integer linear programming (MILP) model and a heuristic algorithm (called CAR), are further proposed to minimize communication conflicts based on guaranteed thermal reliability, and meanwhile, maximize the communication energy efficiency in the presence of on-chip thermal variations. By applying the criterion, our approaches achieve excellent performance with largely reduced complexity of design space exploration. The evaluation results based on both synthetic traffic patterns and realistic benchmarks validate the effectiveness of our approaches with an average of 126.95% improvement in communication performance and 16.12% reduction in energy overhead compared to state-of-the-art techniques. CAR only introduces 7.20% performance difference compared to the MILP model and is more scalable to large-size ONoCs. Mengquan Li, Weichen Liu 0001, Luan H. K. Duong, Peng Chen 0027, Lei Yang 0018, Chunhua Xiao |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2020 | Person Re-Identification Via Pose-Aware Multi-Semantic LearningabstractPerson re-identification (ReID) remains an open-ended research topic, with its variety of substantial applications such as tracking, searching, etc. Existing methods mostly explore the highest-semantic feature embedding, ignoring the insights hidden among the earlier layers. Moreover, owing to the misalignment and pose variations, pose-related information is of great significance and needs to be comprehensively utilized. In this paper, we present a novel person ReID framework called Pose-aware Multi-semantic Fusion Network (PMFN). First, taking into account multiple semantics, we propose Multi-semantic Fusion Network (MFN) as the backbone, employing several shortcuts to reserve bypass feature maps for subsequent fusion. Second, to learn a pose-sensitive embedding, pose-aware clues are considered, forming the complete PMFN and investigating the well-aligned global and local body regions. Finally, the center loss is introduced for enhancing the feature discriminability. Exhaustive experiments on two large-scale person ReID benchmarks demonstrate the strengths of our approach over recent state-of-the-art works. Luan H. K. Duong, Weichen Liu 0001 |
ICME | 2 |
| 2020 | XOR-Net: An Efficient Computation Pipeline for Binary Neural Network Inference on Edge DevicesabstractAccelerating the inference of Convolution Neural Networks (CNNs) on edge devices is essential due to the small memory size and poor computation capability of these devices. Network quantization methods such as XNOR-Net, Bi-Real-Net, and XNOR-Net++ reduce the memory usage of CNNs by binarizing the CNNs. They also simplify the multiplication operations to bit-wise operations and obtain good speedup on edge devices. However, there are hidden redundancies in the computation pipeline of these methods, constraining the speedup of those binarized CNNs. In this paper, we propose XOR-Net as an optimized computation pipeline for binary networks both without and with scaling factors. As XNOR is realized by two instructions XOR and NOT on CPU/GPU platforms, XOR-Net avoids NOT operations by using XOR instead of XNOR, thus reduces bit-wise operations in both aforementioned kinds of binary convolution layers. For the binary convolution with scaling factors, our XOR-Net further rearranges the computation sequence of calculating and multiplying the scaling factors to reduce full-precision operations. Theoretical analysis shows that XOR-Net reduces one-third of the bit-wise operations compared with traditional binary convolution, and up to 40% of the full-precision operations compared with XNOR-Net. Experimental results show that our XOR-Net binary convolution without scaling factors achieves up to 135× speedup and consumes no more than 0.8% energy compared with parallel full-precision convolution. For the binary convolution with scaling factors, XOR-Net is up to 17% faster and 19% more energy-efficient than XNOR-Net. Shien Zhu, Luan H. K. Duong, Weichen Liu 0001 |
ICPADS | 2 |
| 2020 | Modeling and Analysis of Optical Modulators Based on Free-Carrier Plasma Dispersion EffectabstractSilicon photonic networks are revolutionizing computing systems by improving the energy efficiency, bandwidth, and latency of data movements. Optical modulators, such as microresonators (MRs) and Mach–Zehnder interferometers (MZIs), are the basic building blocks of silicon photonic networks. This paper proposes a SPICE-compatible electro-optical co-simulation model, basic optical switch integration model (BOSIM), to systematically study optical modulators using PN, PIN, and metal–insulator–silicon (MIS) capacitor device technologies. BOSIM holistically models both transient and steady state properties, such as switching speed, power, transmission spectrum, area, and carrier distribution. BOSIM is validated by the measured data from eight research groups and companies. Compared to MRs, BOSIM shows MZIs are fast, with a high extinction ratio and large bandwidth but in the sacrifice of loss, energy, and area. Using a PIN diode over a PN diode can save area, but retain the loss and energy, while an MIS capacitor has the shock response of carrier distribution in a narrow range and is marginalized gradually. For instance, an MZI can achieve a$2.5 {\times }$bit rate,$6.06{\times }$extinction ratio,$71.04 {\times }\,\,3$-dB bandwidth, but costs at least$1.93 {\times }$passing loss,$1.46 {\times }$energy consumption, and$16.67 {\times }$area, compared with MR. Xuanqi Chen, Yi-Shing Chang, Jiang Xu 0001, Jun Feng 0008, Peng Yang 0003, Zhehui Wang, Luan H. K. Duong |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 8 |
| 2020 | Fault-Tolerant Routing Mechanism in 3D Optical Network-on-Chip Based on Node ReuseabstractThe three-dimensional Network-on-Chips (3D NoCs) has become a mature multi-core interconnection architecture in recent years. However, the traditional electrical lines have very limited bandwidth and high energy consumption, making the photonic interconnection promising for future 3D Optical NoCs (ONoCs). Since existing solutions cannot well guarantee the fault-tolerant ability of 3D ONoCs, in this paper, we propose a reliable optical router (OR) structure which sacrifices less redundancy to obtain more restore paths. Moreover, by using our fault-tolerant routing algorithm, the restore path can be found inside the disabled OR under the deadlock-free condition, i.e., fault-node reuse. Experimental results show that the proposed approach outperforms the previous related works by maximum 81.1 percent and 33.0 percent on average for throughput performance under different synthetic and real traffic patterns. It can improve the system average optical signal to noise ratio (OSNR) performance by maximum 26.92 percent and 12.57 percent on average, and it can improve the average energy consumption performance by 0.3 percent to 15.2 percent under different topology types/sizes, failure rates, OR structures, and payload packet sizes. Pengxing Guo, Weigang Hou, Lei Guo 0005, Wei Sun 0047, Hainan Bao, Luan H. K. Duong, Weichen Liu 0001 |
IEEE Trans. Parallel Distributed Syst. | 7 |
| 2019 | Crosstalk Noise Reduction Through Adaptive Power Control in Inter/Intra-Chip Optical NetworksabstractIn recent years, optical interconnection networks have been proposed in order to achieve the ultrahigh bandwidth and low latency requirements for inter/intra-chip communication. In these optical interconection networks, series of basic optical elements are employed. Via these series of optical elements, the intrinsic crosstalk noise is generated. With a large scale of these optical elements, the signal-to-noise ratio (SNR) of an optical interconnect can be reduced by this crosstalk noise. In this paper, we utilize the adaptive power control (APC) to enhance the SNR under the crosstalk noise constraints. APC has been known to save energy and reduce power consumption. We apply this technique in one of the inter/intra-chip optical interconnect called I2CON. A new cluster design, namely the Beam cluster, is also introduced. Results have demonstrated that the APC can help to reduce crosstalk noise, hence, the overall SNR is improved. Comparison results have also indicated the further improvement of SNR in I2CON using Beam cluster when APC is applied. Luan H. K. Duong, Peng Yang 0003, Yi-Shing Chang, Jiang Xu 0001, Zhehui Wang, Xuanqi Chen |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2018 | RSON: An inter/intra-chip silicon photonic network for rack-scale computing systemsabstractThe increasing demand for more computational power from scientific computing, big data processing, and machine learning is pushing the development of HPC (high-performance computing) systems. As the basic HPC building blocks, modularized server racks with a large number of multicore nodes are facing performance and energy efficiency challenges. This paper proposes RSON, an optical network for rack-scale computing systems. RSON connects processor cores, caches, local memories, and remote memories through a novel inter/intra-chip silicon photonic network architecture. We develop a low-latency scalable channel partition and low-power dynamic path priority control scheme for RSON. Experimental results show that RSON can help rack-scale computing systems achieve up to 6.8X higher performance under the same energy consumption than state-of-the-art systems under the latest APEX (application performance at extreme scale) benchmarks. Peng Yang 0003, Zhengbin Pang, Zhehui Wang, Xuanqi Chen, Luan H. K. Duong, Jiang Xu 0001 |
DATE | 7 |
| 2017 | Modular reinforcement learning for self-adaptive energy efficiency optimization in multicore systemabstractEnergy-efficiency is becoming increasingly important to modern computing systems with multi-/many-core architectures. Dynamic Voltage and Frequency Scaling (DVFS), as an effective low-power technique, has been widely applied to improve energy-efficiency in commercial multi-core systems. However, due to the large number of cores and growing complexity of emerging applications, it is difficult to efficiently find a globally optimized voltage/frequency assignment at runtime. In order to improve the energy-efficiency for the overall multicore system, we propose an online DVFS control strategy based on core-level Modular Reinforcement Learning (MRL) to adaptively select appropriate operating frequencies for each individual core. Instead of focusing solely on the local core conditions, MRL is able to make comprehensive decisions by considering the running-states of multiple cores without incurring exponential memory cost which is necessary in traditional Monolithic Reinforcement Learning (RL). Experimental results on various realistic applications and different system scales show that the proposed approach improves up to 28% energy-efficiency compared to the recent individual-RL approach. Zhe Wang 0003, Zhongyuan Tian, Jiang Xu 0001, Rafael Kioji Vivas Maeda, Haoran Li 0002, Peng Yang 0003, Zhehui Wang, Luan H. K. Duong, Xuanqi Chen |
ASP-DAC | 8 |
| 2017 | MOCA: an Inter/Intra-Chip Optical Network for MemoryabstractThe memory wall problem is due to the imbalanced developments and separation of processors and memories. It is becoming acute as more and more processor cores are integrated into a single chip and demand higher memory bandwidth through limited chip pins. Optical memory interconnection network (OMIN) promises high bandwidth, bandwidth density, and energy efficiency, and can potentially alleviate the memory wall problem. In this paper, we propose an optical inter/intra-chip processor-memory communication architecture, called MOCA. Experimental results and analysis show that MOCA can significantly improve system performance and energy efficiency. For example, comparing to Hybrid Memory Cube (HMC), MOCA can speedup application execution time by 2.6x, reduce communication latency by 75%, and improve energy efficiency by 3.4x for 256-core processors in 7 nm technology. Zhehui Wang, Zhengbin Pang, Peng Yang 0003, Jiang Xu 0001, Xuanqi Chen, Rafael Kioji Vivas Maeda, Luan H. K. Duong, Haoran Li 0002, Zhe Wang 0003 |
DAC | 8 |
| 2017 | Energy-Efficient Power Delivery System Paradigms for Many-Core ProcessorsabstractThe design of power delivery system plays a crucial role in guaranteeing the proper functionality of many-core processor systems. The power loss suffered on power delivery has become a salient part of total power consumption, and the energy efficiency of a highly dynamic system has been significantly challenged. Being able to achieve a fast response time and multiple voltage domain control, on-chip voltage regulators (VRs) have become popular choices to enable fine-grain power management, which also enlarge the design space of power delivery systems. This paper analytically studies different power delivery system paradigms and power management schemes in terms of energy efficiency, area overhead, and power pin occupation. The analysis shows that compared to the conventional paradigm with off-chip VRs, hybrid paradigms with both on-chip and off-chip VRs are able to maintain high efficiency in a larger range of workloads, though they suffer from low efficiency at light workload. Employed with the quantized power management scheme, the hybrid paradigm can improve the system energy efficiency at light workload by a maximum of 136% compared to the traditional load balanced scheme. Besides this, the in-package (iP) hybrid paradigm further shows its advantage in reducing the physical overheads. The results reveal that at 120 W workload, it occupies only a 10.94% total footprint area or 39.07% power pins of that of the off-chip paradigm. We conclude that the iP hybrid paradigm achieves the best tradeoffs between efficiency, physical overhead, and realization of fine-grain power management. Haoran Li 0002, Xuan Wang 0001, Jiang Xu 0001, Zhe Wang 0003, Rafael Kioji Vivas Maeda, Zhehui Wang, Peng Yang 0003, Luan H. K. Duong |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 8 |
| 2016 | Inter/intra-chip optical interconnection network: opportunities, challenges, and implementationsabstractRecent advances in photonics technologies have made optical interconnection network an attractive option for computing systems from high-performance computers and data centers to automobiles and cellphones. Optical interconnection network promises ultra-high bandwidth, low latency, and great energy efficiency to alleviate the inter-rack, intra-rack, intraboard, and intra-chip communication bottlenecks in multiprocessor systems. Silicon-based photonics technologies piggyback onto developed silicon fabrication processes to provide viable and cost-effective solutions. Both industry and academia have invested significant efforts to develop and commercialize optical interconnection network technologies. This paper reviews the latest progresses and provides insights into the challenges and future developments. Peng Yang 0003, Shigeru Nakamura, Kenichiro Yashiki, Zhehui Wang, Luan H. K. Duong, Xuanqi Chen, Yuichi Nakamura 0002, Jiang Xu 0001 |
NOCS | 5 |
| 2016 | Alleviate Chip Pin Constraint for Multicore Processor by On/Off-Chip Power Delivery System CodesignabstractThe number of chip pins is limited due to the cost and reliability issues of sophisticated packages, and it is predicted that the chip pin count will be overstretched to satisfy the requirements of both power delivery and memory access. The gap between the achievable pin count and the demand will increase as the technology scales, due to the increasing computation resources and supply current. Pin reduction techniques are thus required for continued computing performance growth. In this article, we propose a chip pin constraint alleviation strategy, through on/off-chip power delivery system co-design, to effectively reduce the demand for power pins. An analytical model of a power delivery system, consisting of on/off-chip regulators and a power delivery network, is proposed to evaluate the influence of regulator design and package conduction loss. By combining this model with a multi-core processor model of performance and memory bandwidth requirements, we characterize the entire multi-core processor system to investigate the relationship between the chip pin constraint and performance in multi-core processor scaling and the effectiveness of our strategy. Experiments show that with the conventional power delivery system design, the chip pin constraint severely limits the performance growth as the technology scales. Using the on/off-chip power delivery system co-design, our strategy achieves a significant pin count reduction, for example, 31.3% at the 8nm technology node, compared to the conventional design with the same chip performance, while, provided with the same chip pin count, it is able to improve, by 35.0%, the chip performance at 8nm compared to the conventional design. For real applications of different parallelism, our strategy outperforms its counterpart, with a 23.7% performance improvement on average at the 8nm technology node. Xuan Wang 0001, Jiang Xu 0001, Zhe Wang 0003, Haoran Li 0002, Peng Yang 0003, Luan H. K. Duong, Rafael Kioji Vivas Maeda |
ACM J. Emerg. Technol. Comput. Syst. | 6 |
| 2016 | Coherent and Incoherent Crosstalk Noise Analyses in Interchip/Intrachip Optical Interconnection NetworksabstractRecently, interchip/intrachip optical interconnection networks have been proposed for ultrahigh-bandwidth and low-latency communications. These networks employ the microresonators (MRs) to modulate, direct, or detect the optical signal. However, utilized MRs suffer from intrinsic crosstalk noise and signal power loss, degrading the network efficiency via the signal-to-noise ratio (SNR). The amount of crosstalk noise and signal power loss may differ from network to network. Hence, there exists a need to systematically analyze the effect of the crosstalk noise and the power loss issues. In this paper, we have developed the analytical models considering both coherent and incoherent crosstalk for both the interchip and intrachip optical networks. The interchip/intrachip optical interconnection networks—the$\text{I}^{2}$CON—are analyzed as a case study. The quantitative results on the individual networks have demonstrated that the architectural design determines the impact of crosstalk on the SNR. We have also demonstrated that the optical interconnection networks with interchip/intrachip interconnects result in better bit error rate (BER) compared with that of only intrachip interconnect. Our analyses of the worst case can be utilized as a platform to compare the realistic performance among different optical interconnection networks via the degradation of SNR/BER and data bandwidth. Luan H. K. Duong, Zhehui Wang, Mahdi Nikdast, Jiang Xu 0001, Peng Yang 0003, Zhe Wang 0003, Rafael Kioji Vivas Maeda, Haoran Li 0002, Xuan Wang 0001, Sébastien Le Beux, Yvain Thonnart |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2016 | An Adaptive Process-Variation-Aware Technique for Power-Gating-Induced Power/Ground Noise Mitigation in MPSoCabstractPower gating (PG) is one of the most effective techniques to reduce the leakage power in multiprocessor system-on-chips (MPSoCs). However, the power-mode transition during the PG period of an individual processing unit (PU) will introduce serious power/ground (P/G) noise to the neighboring PUs. As technology scales, the P/G noise problem becomes a severe reliability threat to MPSoCs. At the same time, the increasing manufacturing process variations (PVs) also bring uncertainties to the P/G noise problem and make it difficult to predict and mitigate. To tackle this problem, in this paper, we analyze the PG-induced P/G noise in the presence of PVs and propose a hardware–software collaborated runtime technique to adaptively protect PUs from P/G noise. Sensor network-on-chip is used to gather noise information and coordinate different system components. An online PV-aware algorithm is developed to effectively decide the noise impact range and arrange protections for affected PUs based on the collected noise information. We evaluate the proposed technique through cycle-level Monte Carlo simulations of NoC-based MPSoCs in different scales. The experimental results on various realistic applications show that our technique could achieve comparable reliability to the most reliable static technique while improve on average 3.78%–29.5% the system energy efficiency and reduce 15.7%–70.4% the performance penalty on different MPSoC scales. Zhe Wang 0003, Xuan Wang 0001, Jiang Xu 0001, Haoran Li 0002, Rafael Kioji Vivas Maeda, Zhehui Wang, Peng Yang 0003, Luan H. K. Duong |
IEEE Trans. Very Large Scale Integr. Syst. | 8 |
| 2016 | A Holistic Modeling and Analysis of Optical-Electrical Interfaces for Inter/Intra-chip InterconnectsabstractWith the fast development of inter/intra-chip optical interconnects, the gap between the data rates of electrical interconnects and optical interconnects is continuously increasing. Electrical–optical (E-O) interfaces and optical–electrical (O-E) interfaces are a pair of components that convert data between parallel electrical interconnects and serial optical interconnects. This paper holistically models and analyzes E-O and O-E interfaces in terms of energy consumption, area, and latency. Traditional interfaces, where data are converted between parallel and serial ports by serializers and deserializers (SerDes), are studied. A new type of E-O and O-E interface, which serializes and deserializes data by optical weaving technologies, are proposed alongside. Traditional interfaces will become a bottleneck for the further development of optical interconnects in the near future because of the high energy consumption and large area of SerDes necessitating new technologies. Our analysis shows that optical weaving interfaces have a better overall performance than traditional interfaces. For example, if there are 64 parallel electrical interconnects and four optical wavelengths, optical weaving interfaces can achieve a 81.6% improvement in energy consumption and a 40.8% improvement in area, compared with traditional interfaces. Zhehui Wang, Jiang Xu 0001, Peng Yang 0003, Luan H. K. Duong, Xuan Wang 0001, Zhe Wang 0003, Haoran Li 0002, Rafael Kioji Vivas Maeda |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2016 | Improve Chip Pin Performance Using Optical InterconnectsabstractWith the fast development of processor chips, power-efficient, high-bandwidth, and low-latency interchip interconnects become more and more important. Studies show that the bandwidth of traditional parallel interconnects with low I/O clock frequencies will become bottlenecks in the near future. To solve this problem, two types of high-bandwidth interchip interconnects are developed. Low-swing differential electrical interconnects have widely been used in high-speed I/O designs. On the other hand, optical interconnects promise high bandwidth, low latency, and could improve the chip pin performance for manycore processors. They are becoming potential alternatives for electrical interconnects. This paper systematically models these two types of interconnects in terms of crosstalk noises, attenuation, and receiver sensitivities. Based on the proposed models, we developed optical and electrical interfaces and links (OEIL) and an analysis tool for OEIL. The OEIL can be used to analyze the energy consumption, bandwidth density, and latency of interconnects. Analytical models are verified by the results of published experiments. It shows that the optical interconnects have much higher bandwidth densities than the electrical interconnects. With this feature, the optical interconnects can significantly reduce I/O pin count compared with the electrical interconnects. For example, they can save at least 92% signal pins when connecting chips more than 25 cm (10 in) apart. The energy consumption of optical interconnects is comparable with that of electrical interconnects, and the latency of polymer waveguide-based optical interconnects is 18% less than that of electrical interconnect. Zhehui Wang, Jiang Xu 0001, Peng Yang 0003, Xuan Wang 0001, Zhe Wang 0003, Luan H. K. Duong, Rafael Kioji Vivas Maeda, Haoran Li 0002 |
IEEE Trans. Very Large Scale Integr. Syst. | 6 |
| 2015 | Alleviate chip I/O pin constraints for multicore processors through optical interconnectsabstractChip I/O pins are an increasingly limited resource and significantly affect the performance, power and cost of multicore processors. Optical interconnects promise low power and high bandwidth, and are potential alternatives to electrical interconnects. This work systematically developed a set of analytical models for electrical and optical interconnects to study their structures, receiver sensitivities, crosstalk noises, and attenuations. We verified the models by published implementation results. The analytical models quantitatively identified the advantages of optical interconnects in terms of bandwidth, energy consumption, and transmission distance. We showed that optical interconnects can significantly reduce chip pin counts. For example, compared to electrical interconnects, optical interconnects can save at least 92% signal pins when connecting chips more than 25 cm (10 inches) apart. Zhehui Wang, Jiang Xu 0001, Peng Yang 0003, Xuan Wang 0001, Zhe Wang 0003, Luan H. K. Duong, Haoran Li 0002, Rafael Kioji Vivas Maeda, Xiaowen Wu, Yaoyao Ye, Qinfen Hao |
ASP-DAC | 6 |
| 2015 | Coherent crosstalk noise analyses in ring-based optical interconnects
Luan H. K. Duong, Mahdi Nikdast, Jiang Xu 0001, Zhehui Wang, Yvain Thonnart, Sébastien Le Beux, Peng Yang 0003, Xiaowen Wu |
DATE | 1 |
| 2015 | Adaptively tolerate power-gating-induced power/ground noise under process variations
Zhe Wang 0003, Xuan Wang 0001, Jiang Xu 0001, Xiaowen Wu, Zhehui Wang, Peng Yang 0003, Luan H. K. Duong, Haoran Li 0002, Rafael Kioji Vivas Maeda |
DATE | 7 |
| 2015 | An Analytical Study of Power Delivery Systems for Many-Core Processors Using On-Chip and Off-Chip Voltage RegulatorsabstractDesign of power delivery system has great influence on the power management in many-core processor systems. Moving voltage regulators from off-chip to on-chip gains more and more interest in the power delivery system design, because it is able to provide fine-grained dynamic voltage scaling. Previous works are proposed to implement power efficient on-chip voltage regulators. It is important to analyze the characteristics of the entire power delivery system to explore the tradeoff between the promising properties and costs of employing on-chip voltage regulators, especially the on-chip buck converters. In this paper, we present a novel analysis and design optimization platform of power delivery system called power supply on-chip (PowerSoC). It employs an analytical model to provide an accurate and fast evaluation of important characteristics, e.g., power efficiency, output stability, and dynamic voltage scaling, for the entire power delivery system consisting of on-chip/off-chip buck converters and power delivery network. Based on our model, geometric programming is utilized to find the optimal design for different power delivery systems and explore the tradeoff of using on-chip converters. Compared with SPICE simulations, our model achieves a simulation time reduction of six to seven orders of magnitude within 5% model error for the characteristic evaluation of different power delivery systems. By using PowerSoC, various architectures of power delivery systems are optimized for power efficiency under constraints of output stability, area, etc. Simulation results show that the hybrid architecture, consisting of both on-chip and off-chip converters, achieves 1.0% power efficiency improvement and 66.4% area reduction of converters, compared to the conventional design. We conclude the hybrid architecture has potential for efficient dynamic voltage scaling, small area, and the adaptability of the change of power delivery network parasitic, but careful account for the overhead of on-chip converters is needed. Xuan Wang 0001, Jiang Xu 0001, Zhe Wang 0003, Kevin J. Chen, Xiaowen Wu, Zhehui Wang, Peng Yang 0003, Luan H. K. Duong |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 8 |
| 2015 | Fat-Tree-Based Optical Interconnection Networks Under Crosstalk Noise ConstraintabstractOptical networks-on-chip (ONoCs) have shown the potential to be substituted for electronic networks-on-chip (NoCs) to bring substantially higher bandwidth and more efficient power consumption in both on- and off-chip communication. However, basic optical devices, which are the key components in constructing ONoCs, experience inevitable crosstalk noise and power loss; the crosstalk noise from the basic devices accumulates in large-scale ONoCs and considerably hurts the signal-to-noise ratio (SNR) as well as restricts the network scalability. For the first time, this paper presents a formal system-level analytical approach to analyze the worst-case crosstalk noise and SNR in arbitrary fat-tree-based ONoCs. The analyses are performed hierarchically at the basic optical device level, then at the optical router level, and finally at the network level. A general 4$\,\times\,$4 optical router model is considered to enable the proposed method to be adaptable to fat-tree-based ONoCs using an arbitrary 4$\,\times\,$4 optical router. Utilizing the proposed general router model, the worst-case SNR link candidates in the network are determined. Moreover, we apply the proposed analyses to a case study of fat-tree-based ONoCs using an optical turnaround router (OTAR). Quantitative simulation results indicate low values of SNR and scalability constraints in large scale fat-tree-based ONoCs, which is due to the high power of crosstalk noise and power loss. For instance, in fat-tree-based ONoCs using the OTAR, when the injection laser power equals 0 dBm, the crosstalk noise power is higher than the signal power when the number of processor cores exceeds 128; when it is equal to 256, the signal power, crosstalk noise power, and SNR are${-}{17.3}$,${-}{11.9}$, and${-}{\rm 5.5}~{\rm dB}$, respectively. Mahdi Nikdast, Jiang Xu 0001, Luan H. K. Duong, Xiaowen Wu, Zhehui Wang, Xuan Wang 0001, Zhe Wang 0003 |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2015 | Crosstalk Noise in WDM-Based Optical Networks-on-Chip: A Formal Study and ComparisonabstractOptical networks-on-chip (ONoCs) using wavelength-division multiplexing (WDM) technology have progressively attracted more and more attention for their use in tackling the high-power consumption and low bandwidth issues in growing metallic interconnection networks in multiprocessor systems-on-chip. However, the basic optical devices employed to construct WDM-based ONoCs are imperfect and suffer from inevitable power loss and crosstalk noise. Furthermore, when employing WDM, optical signals of various wavelengths can interfere with each other through different optical switching elements within the network, creating crosstalk noise. As a result, the crosstalk noise in large-scale WDM-based ONoCs accumulates and causes severe performance degradation, restricts the network scalability, and considerably attenuates the signal-to-noise ratio (SNR). In this paper, we systematically study and compare the worst case as well as the average crosstalk noise and SNR in three well-known optical interconnect architectures, mesh-based, folded-torus-based, and fat-tree-based ONoCs using WDM. The analytical models for the worst case and the average crosstalk noise and SNR in the different architectures are presented. Furthermore, the proposed analytical models are integrated into a newly developed crosstalk noise and loss analysis platform (CLAP) to analyze the crosstalk noise and SNR in WDM-based ONoCs of any network size using an arbitrary optical router. Utilizing CLAP, we compare the worst case as well as the average crosstalk noise and SNR in different WDM-based ONoC architectures. Furthermore, we indicate how the SNR changes in respect to variations in the number of optical wavelengths in use, the free-spectral range, and the microresonators$\boldsymbol {Q}$factor. The analyses’ results demonstrate that the crosstalk noise is of critical concern to WDM-based ONoCs: in the worst case, the crosstalk noise power exceeds the signal power in all three WDM-based ONoC architectures, even when the number of processor cores is small, e.g., 64. Mahdi Nikdast, Jiang Xu 0001, Luan H. K. Duong, Xiaowen Wu, Xuan Wang 0001, Zhehui Wang, Zhe Wang 0003, Peng Yang 0003, Yaoyao Ye, Qinfen Hao |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2014 | CLAP: a crosstalk and loss analysis platform for optical interconnectsabstractBasic photonic devices in inter- and intra-chip optical networks suffer from inevitable power loss and crosstalk noise. Incoherent crosstalk introduces quick power fluctuations, while coherent crosstalk varies the optical power of the optical signal in optical interconnection networks (OINs). As a result, the accumulative crosstalk in large scale OINs considerably hurts the signal-to-noise ratio (SNR) and imposes high power penalties. In this work, we aim at studying the worst-case incoherent and coherent crosstalk in OINs at the system level. The proposed analytical models are integrated into a newly developed crosstalk and loss analysis platform, called CLAP, to facilitate the SNR analyses in arbitrary OINs. Mahdi Nikdast, Luan H. K. Duong, Jiang Xu 0001, Sébastien Le Beux, Xiaowen Wu, Zhehui Wang, Peng Yang 0003, Yaoyao Ye |
NOCS | 2 |
| 2014 | System-Level Modeling and Analysis of Thermal Effects in WDM-Based Optical Networks-on-ChipabstractMultiprocessor systems-on-chip show a trend toward integration of tens and hundreds of processor cores on a single chip. With the development of silicon photonics for short-haul optical communication, wavelength division multiplexing (WDM)-based optical networks-on-chip (ONoCs) are emerging on-chip communication architectures that can potentially offer high bandwidth and power efficiency. Thermal sensitivity of photonic devices is one of the main concerns about the on-chip optical interconnects. We systematically modeled thermal effects in optical links in WDM-based ONoCs. Based on the proposed thermal models, we developed OTemp, an optical thermal effect modeling platform for optical links in both WDM-based ONoCs and single-wavelength ONoCs. OTemp can be used to simulate the power consumption as well as optical power loss for optical links under temperature variations. We use case studies to quantitatively analyze the worst-case power consumption for one wavelength in an eight-wavelength WDM-based optical link under different configurations of low-temperature-dependence techniques. Results show that the worst-case power consumption increases dramatically with on-chip temperature variations. Thermal-based adjustment and optimal device settings can help reduce power consumption under temperature variations. Assume that off-chip vertical-cavity surface-emitting lasers are used as the laser source with WDM channel spacing of 1 nm, if we use thermal-based adjustment with guard rings for channel remapping, the worst-case total power consumption is 6.7 pJ/bit under the maximum temperature variation of 60 °C; larger channel spacing would result in a larger worst-case power consumption in this case. If we use thermal-based adjustment without channel remapping, the worst-case total power consumption is around 9.8 pJ/bit under the maximum temperature variation of 60 °C; in this case, the worst-case power consumption would benefit from a larger channel spacing. Yaoyao Ye, Zhehui Wang, Peng Yang 0003, Jiang Xu 0001, Xiaowen Wu, Xuan Wang 0001, Mahdi Nikdast, Zhe Wang 0003, Luan H. K. Duong |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 9 |