EDBT 2026 Demo / reviewers in the wild / expert
Zhaolin Li
dblp:77/2814
· DBLP profile ↗
29ranked-venue papers
2as first author
17since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 19 · 8 since 2021Artificial intelligence and machine learning · 6 · 2 first-author · 6 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021Computer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | NBCache: An Efficient and Scalable Non-Blocking Cache for Coherent Multi-Chiplet Systems
Zhirong Ye, Tao Lu 0012, Zhaolin Li, Zhiyi Yu, Mingyu Wang 0003 |
ASP-DAC | 5 |
| 2026 | LRM-GPU: Alleviating Synchronization Overhead for Multi-Chiplet GPU Architecture
Baiqing Zhong, Zhirong Ye, Haiqiu Huang, Zhaolin Li, Zhiyi Yu, Mingyu Wang 0003 |
HPCA | 6 |
| 2026 | PipeIMC: A Pipelined In-SRAM Computing Architecture
Yikai Cui, Renhao Fan, Weike Li, Mingyu Wang 0003, Zhaolin Li |
ISCA | 6 |
| 2026 | Design of Tree-Shaped Decoupling Structure Based on Generative Nonadversarial Networks for Antennas in IoT ApplicationsabstractThe miniaturization of Internet of Things (IoT) devices requires multiple antennas in a limited space, which can lead to strong mutual coupling. Traditional decoupling methods usually rely on experience-driven design or application-specific structures, which limits their effectiveness in different antenna configurations. In this article, a generalized tree-shaped structure with high flexibility is proposed for antenna decoupling in terminal devices in Internet of Things (IoT) systems. The specific geometries of the decoupling structure are optimized using generative non-adversarial networks (GNAN) according to different requirements. As the prior knowledge, quarter-wavelength branches are initialized to improve the optimization efficiency. To meet the design requirements with a smaller dataset, an evolutionary method is employed. At the same time, attention mechanism is introduced into the discriminator network, which largely improves the training speed. Following the proposed method, different tree-shaped decoupling structures are optimized and applied to three different multiple-input multiple-output (MIMO) antenna systems, including an ultra-wide band (UWB) antenna array, a narrow-band microstrip patch antenna array and an antenna array with different antenna pairs. As a result, isolations have been proved by 20 dB, 53 dB, and 45 dB, respectively. Hence, the flexible tree-shaped structure is validated to be generally effective for closely placed MIMO antennas in IoT systems. Prototypes of the transparent UWB antenna array have been fabricated and measured, with the measurement results agreeing well with the simulated ones. Hui Li 0016, Zhaolin Li, Tianxi Feng |
IEEE Internet Things J. | 4 |
| 2026 | MAX-SM: High-Utilization Dynamic SM Partitioning for Heterogeneous Workloads on Multitasking Chiplet-Based GPUs
Mingyu Wang 0003, Tao Lu 0012, Baiqing Zhong, Zhaolin Li, Zhiyi Yu |
IEEE Trans. Parallel Distributed Syst. | 6 |
| 2025 | In-context Language Learning for Endangered Languages in Speech Recognition
Zhaolin Li, Jan Niehues |
INTERSPEECH | 1 |
| 2025 | MagiCache: A Virtual In-Cache Computing EngineabstractThe rise of data-parallel applications poses a significant challenge to the energy consumption of computing architectures.In-cache computation is a promising solution for achieving high parallelism and energy efficiency because it can eliminate data movement between the cache and the processor.Existing in-cache computing architectures transform a portion of cache arrays into computing arrays, with all rows of these arrays serving as computing lines.The remaining cache arrays are used as cachelines to store the data required by computing arrays or processors.However, in these array-level in-cache computing architectures, only a few computing lines in each computing array are active at runtime while the others are idle, which incurs severe cache capacity loss and space underutilization.In addition, bursty memory accesses of data-parallel applications also cause significant in-cache data movement latency.To address these problems, we propose MagiCache, a virtual in-cache computing engine.First, we design a novel cacheline-level in-cache computing architecture in which each cache array can configure some rows as computing lines and the other rows as cachelines with negligible overhead.Second, a virtual engine is further designed on this novel architecture to dynamically allocate different rows of each array as computing lines or cachelines based on runtime computation and storage requirements, thus realizing efficient cacheline-level space management.Finally, we present an instruction chaining technique to overlap the bursty access latency by enabling asynchronous execution of computing arrays.Evaluation results show that MagiCache achieves a 1.19x-1.61xspeedup over the state-of-the-art in-cache computing architectures with 6.5 KB of additional storage.Our cacheline-level space * These authors contributed equally to this work. Renhao Fan, Yikai Cui, Weike Li, Mingyu Wang 0003, Zhaolin Li |
ISCA | 5 |
| 2024 | Speech Recognition Corpus of the Khinalug Language for Documenting Endangered LanguagesabstractAutomatic Speech Recognition (ASR) can be a valuable tool to document endangered languages. However, building ASR tools for these languages poses several difficult research challenges, notably data scarcity. In this paper, we show the whole process of creating a useful ASR tool for language documentation scenarios. We publish the first speech corpus for Khinalug, an endangered language spoken in Northern Azerbaijan. The corpus consists of 2.67 hours of labeled data from recordings of spontaneous speech about various topics. As Khinalug is an extremely low-resource language, we investigate the benefits of multilingual models for self-supervised learning and supervised learning and achieve the performance of 6.65 Character Error Rate (CER) points and 25.53 Word Error Rate (WER) points. The benefits of multilingual models are further validated through experimentation with three additional under-resourced languages. Lastly, this work conducts quality assessments with linguists on new recordings to investigate the model’s usefulness in language documentation. We observe an evident degradation for new recordings, indicating the importance of enhancing model robustness. In addition, we find the inaudible content is the main cause of wrong ASR predictions, suggesting relating work on incorporating contextual information. Zhaolin Li, Monika Rind-Pawlowski, Jan Niehues |
LREC/COLING | 1 |
| 2024 | SciEx: Benchmarking Large Language Models on Scientific Exams with Human Expert Grading and Automatic GradingabstractTu Anh Dinh, Carlos Mullov, Leonard Bärmann, Zhaolin Li, Danni Liu, Simon Reiß, Jueun Lee, Nathan Lerzer, Jianfeng Gao, Fabian Peller-Konrad, Tobias Röddiger, Alexander Waibel, Tamim Asfour, Michael Beigl, Rainer Stiefelhagen, Carsten Dachsbacher, Klemens Böhm, Jan Niehues. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024. Tu Anh Dinh, Carlos Mullov, Leonard Bärmann, Zhaolin Li, Simon Reiß, Jueun Lee, Nathan Lerzer, Jianfeng Gao 0002, Fabian Tërnava, Tobias Röddiger, Alex Waibel, Tamim Asfour, Michael Beigl, Rainer Stiefelhagen, Carsten Dachsbacher, Klemens Böhm, Jan Niehues |
EMNLP | 4 |
| 2024 | A Row Decomposition-based Approach for Sparse Matrix Multiplication on GPUsabstractSparse-Matrix Dense-Matrix Multiplication (SpMM) and Sampled Dense Dense Matrix Multiplication (SDDMM) are important sparse kernels in various computation domains. The uneven distribution of nonzeros in the sparse matrix and the tight data dependence between sparse and dense matrices make it a challenge to run sparse matrix multiplication efficiently on GPUs. By analyzing the aforementioned problems, we propose a row decomposition (RoDe)-based approach to optimize the two kernels on GPUs, using the standard Compressed Sparse Row (CSR) format. Specifically, RoDe divides the sparse matrix rows into regular parts and residual parts, to fully optimize their computations separately. We also devise the corresponding load balancing and finegrained pipelining technologies. Profiling results show that RoDe can achieve more efficient memory access and reduce warp stall cycles significantly. Compared to the state-of-the-art (SOTA) alternatives, RoDe achieves a speedup of up to 7.86× with a geometric mean of 1.45× for SpMM, and a speedup of up to 8.99× with a geometric mean of 1.49× for SDDMM; the dataset is SuiteSparse. RoDe also outperforms its counterpart in the deep learning dataset. Furthermore, its preprocessing overhead is significantly smaller, averaging only 16% of the SOTA. Peng Qu 0001, Youhui Zhang, Zhaolin Li |
PPoPP | 5 |
| 2024 | Augmenting Automatic Speech Recognition Models With Disfluency DetectionabstractSpeech disfluency commonly occurs in conversational and spontaneous speech. However, standard Automatic Speech Recognition (ASR) models struggle to accurately recognize these disfluencies because they are typically trained on fluent transcripts. Current research mainly focuses on detecting disfluencies within transcripts, overlooking their exact location and duration in the speech. Additionally, previous work often requires model fine-tuning and addresses limited types of disfluencies. In this work, we present an inference-only approach to augment any ASR model with the ability to detect open-set disfluencies. We first demonstrate that ASR models have difficulty transcribing speech disfluencies. Next, this work proposes a modified Connectionist Temporal Classification(CTC)based forced alignment algorithm from [1] to predict wordlevel timestamps while effectively capturing disfluent speech. Additionally, we develop a model to classify alignment gaps between timestamps as either containing disfluent speech or silence. This model achieves an accuracy of 81.62% and an F1-score of 80.07%. We test the augmentation pipeline of alignment gap detection and classification on a disfluent dataset. Our results show that we captured 74.13% of the words that were initially missed by the transcription, demonstrating the potential of this pipeline for downstream tasks. Robin Amann, Zhaolin Li, Barbara Bruno, Jan Niehues |
SLT | 2 |
| 2023 | FABLE: A Development and Computing Framework for Brain-inspired Learning AlgorithmsabstractSpiking neural networks (SNNs) have received extensive attention in multi-disciplinary fields, due to their rich spatiotemporal dynamics and the potential for low processing delay and high energy efficiency on neuromorphic hardware. The research on SNN learning algorithms is active and diverse, and many algorithms differ significantly from those of DNN in terms of computation model/features and weight adjustment mechanisms. This paper proposes FABLE, a multi-level framework for building and running SNN learning algorithms efficiently. Its kernel is an adaptable computation model based on synchronous data flow, which can well express the spatiotemporal parallelism of SNN and then can organize and schedule the underlying SNN-custom tensor operators (OPs) to construct optimized computing procedures. It also provides a flexible programming interface for users to design or customize their learning algorithms. In addition, the implementation of FABLE has high compatibility: It extends PyTorch's OP library, scheduler, and APIs to take advantage of the ecology and usability of the latter. To show the flexibility of the framework, we have ported five different learning algorithms, each with less programming than its original implementation. Further experiments demonstrate that FABLE outperforms all of them (up to 2.61 x) in terms of computing performance, while the original implementations are either based on PyTorch, based on some third-party tool using PyTorch, or based on GPGPU's runtime directly. Zhaolin Li, Youhui Zhang |
IJCNN | 3 |
| 2023 | MAICC : A Lightweight Many-core Architecture with In-Cache Computing for Multi-DNN Parallel InferenceabstractThe growing complexity and diversity of neural networks in the fields of autonomous driving and intelligent robots have facilitated the research of many-core architectures, which can offer sufficient programming flexibility to simultaneously support multi-DNN parallel inference with different network structures and sizes compared to domain-specific architectures. However, due to the tight constraints of area and power consumption, many-core architectures typically use lightweight scalar cores without vector units and are almost unable to meet the high-performance computing needs of multi-DNN parallel inference. To solve the above problem, we design an area- and energy-efficient many-core architecture by integrating large amounts of lightweight processor cores with RV32IMA ISA. The architecture leverages the emerging SRAM-based computing-in-memory technology to implement vector instruction extensions by reusing memory cells in the data cache instead of conventional logic circuits. Thus, the data cache in each core can be reconfigured as the memory part and the computing part with the latter tightly coupled with the core pipeline, enabling parallel execution of the basic RISC-V instructions and the extended multi-cycle vector instructions. Furthermore, a corresponding execution framework is proposed to effectively map DNN models onto the many-core architecture by using intra-layer and inter-layer pipelining, which potentially supports multi-DNN parallel inference. Experimental results show that the proposed MAICC architecture obtains a 4.3 × throughput and 31.6 × energy efficiency over CPU (Intel i9-13900k). MAICC also achieves a 1.8 × energy efficiency over GPU (RTX 4090) with only 4MB on-chip memory and 28 mm2 area. Renhao Fan, Yikai Cui, Qilin Chen, Mingyu Wang 0003, Youhui Zhang, Zhaolin Li |
MICRO | 7 |
| 2022 | Quantum Entanglement Inspired Correlation Learning for Classification
Junwei Zhang 0009, Zhao Li 0007, Shichang Hu, Jie Xiao 0005, Zhaolin Li |
PAKDD (2) | 7 |
| 2022 | Polyhedral-Based Compilation Framework for In-Memory Neural Network AcceleratorsabstractMemristor-based processing-in-memory architecture is a promising solution to the memory bottleneck in the neural network ( NN ) processing. A major challenge for the programmability of such architectures is the automatic compilation of high-level NN workloads, from various operators to the memristor-based hardware that may provide programming interfaces with different granularities. This article proposes a source-to-source compilation framework for such memristor-based NN accelerators, which can conduct automatic detection and mapping of multiple NN operators based on the flexible and rich representation capability of the polyhedral model. In contrast to previous studies, it implements support for pipeline generation to exploit the parallelism in the NN loads to leverage hardware resources for higher efficiency. The evaluation based on synthetic kernels and NN benchmarks demonstrates that the proposed framework can reliably detect and map the target operators. Case studies on typical memristor-based architectures also show its generality over various architectural designs. The evaluation further demonstrates that compared with existing polyhedral-based compilation frameworks that do not support the pipelined execution, the performance can upgrade by an order of magnitude with the pipelined execution, which emphasizes the necessity of our improvement. Jianhui Han, Zhaolin Li, Youhui Zhang |
ACM J. Emerg. Technol. Comput. Syst. | 3 |
| 2021 | Quantum Correlation Revealed by Bell State for Classification TasksabstractIn machine learning, classification algorithms often use statistical methods to build the correspondence between features (or attributes) and categories (or labels), that is, the statistical correlation between features and categories. In quantum theory, a large number of experimental results show that quantum correlation is far stronger than what can be explained by local hidden theory (i.e., classical or non-quantum theory), that is, quantum mechanics theory reveals a statistical correlation stronger than that described by classical theory. Based on this, this paper will use the strong statistical correlation revealed by Bell state to build a classification algorithm to verify the validity and superiority of the formal framework of quantum mechanics in specific classification tasks. Specifically, we use quantum joint probabilities derived from the measurement process of Bell state to model the quantum statistical correlation between features and categories. The paper first theoretically proves that the formal framework used has the ability to violate Bell inequality; moreover, a classification algorithm is implemented and verified on classic machine learning datasets. Experimental results show that the algorithm is significantly better than most mainstream machine learning algorithms. Junwei Zhang 0009, Ruifang He, Zhao Li 0007, Ji Zhang 0001, Zhaolin Li, Tianyuan Niu |
IJCNN | 6 |
| 2021 | Interactive Quantum Classifier Inspired by Quantum Open System TheoryabstractQuantum theory has attracted people's attention since it was proposed. Due to its unique advantages in information storage and processing, quantum information processing has become the most popular research field. Quantum theory also provides us with new methods or concepts for information manipulation and processing. The basic problems of classical physics are basically trying to be solved in a situation of isolation from the surrounding environment to reduce the complexity of the analysis problem, but the quantum system inevitably produces decoherence and establishes a close relationship with the environment such as entanglement, so the formal framework of quantum mechanics is inherently capable of depicting complex relationships. Based on the principle of quantum open system, a classifier under the formal framework of quantum mechanics is established to simulate the evolution process of open systems, that is, the interaction process between the target system and the environment. Specifically, we regard the features (or attributes) of the sample as environmental factors that affect the decision-making of the target system, and the target system can obtain the categories (or labels) of the sample through measurement. Based on this, we use the formal framework of quantum mechanics to establish a more natural and tighter correspondence between attributes and labels. Limited by the limitations of simulating quantum operations on classical computers, we conducted experiments on two lightweight machine learning datasets and compared them with mainstream classification algorithms. Experimental results show that the classification algorithm is better than the comparison models, and it also reveals the potential of the algorithm. Junwei Zhang 0009, Zhao Li 0007, Ruifang He, Ji Zhang 0001, Zhaolin Li, Tianyuan Niu |
IJCNN | 6 |
| 2020 | An efficient object detection framework with modified dense connections for small objects optimizationsabstractObject detection frameworks for small objects are increasingly demanded in some specific fields such as high-speed object tracking and remote sensing image recognition. In this paper, we propose an efficient object detection framework with modified dense connections for small objects. In order to improve both the detection accuracy and speed for small objects, the proposed framework constructs a convolutional neural network by using modified dense and residual cross-layer connections between multi-scale convolutional layers to extract deep features effectively. Based on the modified dense structure, a hybrid-scale feature fusion method is proposed to concatenate the multi-channel high-dimensional features and performs cross-entropy calculation and regression prediction. By using this method, this framework not only improves the detection accuracy for small objects significantly, but also improves the overall detection accuracy and optimizes the network parameters to reduce the detection time greatly. The experimental results show that the proposed framework achieves 90.6% mAP for small objects on a public ship dataset which is 25.2% more than SSD-VGGNet. Due to the detection efficiency for small objects, it improves the overall detection accuracy and detection speed by 9% and 40% respectively while about 70% network parameters are reduced. Yicong Zhang, Zhaolin Li |
CF | 3 |
| 2020 | ERA-LSTM: An Efficient ReRAM-Based Architecture for Long Short-Term MemoryabstractProcessing-in-memory (PIM) architecture based on resistive random access memory (ReRAM) crossbars is a promising solution to the memory bottleneck that long short-term memory (LSTM) faces. Based on the dataflow analysis of the LSTM computing paradigm, this article proposes to adopt the ReRAM-based analog approximate computing to conduct the LSTM-specific element-wise computation. Combined with the dot-product computation implemented with ReRAM crossbars, a new LSTM processing tile is designed to significantly reduce the demand for analog-to-digital converters (ADCs), which is the major part of power consumption of existing designs. Next, we elaborate on a mapping scheme to efficiently deploy large-scale LSTM onto multiple processing tiles. Finally, an architecture enhancement is proposed to support crossbar-friendly LSTM pruning to further improve efficiency. This overall design, named ERA-LSTM, is presented. Our evaluation shows that it can outperform two state-of-the-art FPGA-based LSTM accelerators by 103.6 and 35.9 times, respectively; compared with a state-of-the-art ReRAM-based LSTM accelerator with digital element-wise computation, it is 6.1 times more efficient. Moreover, our experiments demonstrate that the impact of hardware constraints and approximation errors on the inference accuracy can be effectively reduced by the proposed fine-tuning scheme and by optimizing the design of the approximator. Jianhui Han, Mingyu Wang 0003, Zhaolin Li, Youhui Zhang |
IEEE Trans. Parallel Distributed Syst. | 4 |
| 2017 | A Spatial and Temporal Locality-Aware Adaptive Cache Design With Network Optimization for Tiled Many-Core ArchitecturesabstractThe spatial locality and the temporal locality of workloads are the root causes for cache designs to overcome the memory wall problem. However, the real memory access behavior for each of these applications can be very different. It gives the opportunities to explore further performance improvement due to different cache organization requirements. To address this issue, a spatial and temporal locality-aware adaptive cache is proposed, which dynamically partitions the private last level cache bank as prefetch region or victim region at runtime to explore the locality characteristics. The prefetch region speculates the data blocks in subsequent addresses to exploit the spatial locality, while the victim region collects the evicted data blocks from the upper memory hierarchy to exploit the temporal locality. Fast data prefetch with prioritized dynamic buffer management and adaptive burst-aware routing is realized in the proposed hybrid burst-support network-on-chip (HBNoC). By combining the adaptive cache partition with HBNoC, the off-chip misses and the on-chip network usage are greatly reduced. Experimental results demonstrate that the proposed adaptive cache design reduces up to 25% off-chip misses and improves 11.3% performance on average compared with the prior design, respectively. Mingyu Wang 0003, Zhaolin Li |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2016 | STLAC: A spatial and temporal locality-aware cache and network-on-chip codesign for tiled many-core systemsabstractThe spatial and temporal locality of workloads are the root causes for cache designs to overcome the memory wall problem. However, few existing state-of-the-art designs exploit both the two locality features to optimize the memory hierarchies in the area of tiled many-core systems, which losses the opportunities to explore more performance improvement. To address this problem, an adaptive spatial and temporal locality-aware cache and network-on-chip (NoC) codesign (STLAC) is proposed, which dynamically partitions the last level cache (LLC) as data prefetch buffer or victim cache for locality prediction and exploits a hybrid burst-support NoC for fast data prefetch. The data prefetch buffer speculates the data blocks in subsequent addresses to exploit the spatial locality, while the victim cache collects the evicted data blocks from the upper memory hierarchy to exploit the temporal locality. By combining the proposed adaptive cache partition with the hybrid burst-support NoC, the off-chip misses and on-chip network usage are greatly reduced. Experimental results demonstrate that the proposed STLAC reduces up to 43% off-chip misses and improves 15% performance on average compared with the traditional shared LLC design. Zhaolin Li |
ASP-DAC | 2 |
| 2016 | Temperature-aware task scheduling heuristics on Network-on-ChipsabstractChip temperature becomes a critical design issue with technology scaling to nanometer-scale, especially for NoC systems with large number of cores and shrunken core size. To reduce peak temperature and balance spatial temperature distribution on NoC-based multi-cores chips, this paper proposes a temperature-aware task scheduling approach. The thermal profiles of tasks are first extracted by accurate temperature model. Then run-time task mapping heuristic is proposed considering transient core temperatures, thermal dissipation from adjacent cores, communication overheads and the thermal influence of physical position on chip. Voltage-frequency is also scaled down when timing constraint is met to reduce power consumption and core temperature. Experimental results show that the significant reduction of peak temperature and the temperature variance compared with the current approaches is achieved. Shan Cao 0001, Zoran A. Salcic, Yingtao Ding, Zhaolin Li, Shaojun Wei, Xianli Zhao |
ISCAS | 4 |
| 2016 | A hybrid SDC/SDF architecture for area and power minimization of floating-point FFT computationsabstractFloating-point fast Fourier transform (FFT) has been widely expected in scientific computing and high-resolution imaging applications due to the wide dynamic range and high processing precision. However, it suffers the high area and power overhead problem in comparison to fixed-point implementations. To address this problem, this paper presents a novel hybrid SDC/SDF architecture for area and power minimization. It minimizes the required arithmetic units and reduces the memory usage significantly in the single-path delay commutator (SDC) part and single-path delay feedback (SDF) part, respectively. By combining both the advantages on arithmetic units reducing and memory usage optimization of these two parts, the minimized total area and power are obtained without any throughput loss. Logic synthesis results in a 65nm CMOS technology show that the power consumption ranges from 35.7mW to 416.8mW for 16- to 1024-point FFTs at 400MHz, and the total hardware overhead is equivalent to 747k NAND2 gates. Zhaolin Li |
ISCAS | 2 |
| 2015 | Schedule refinement for homogeneous multi-core processors in the presence of manufacturing-caused heterogeneityabstractMulti-core homogeneous processors have been widely used to deal with computation-intensive embedded applications. However, with the continuous down scaling of CMOS technology, within-die variations in the manufacturing process lead to a significant spread in the operating speeds of cores within homogeneous multi-core processors. Task scheduling approaches, which do not consider such heterogeneity caused by within-die variations, can lead to an overly pessimistic result in terms of performance. To realize an optimal performance according to the actual maximum clock frequencies at which cores can run, we present a heterogeneity-aware schedule refining (HASR) scheme by fully exploiting the heterogeneities of homogeneous multi-core processors in embedded domains. We analyze and show how the actual maximum frequencies of cores are used to guide the scheduling. In the scheme, representative chip operating points are selected and the corresponding optimal schedules are generated as candidate schedules. During the booting of each chip, according to the actual maximum clock frequencies of cores, one of the candidate schedules is bound to the chip to maximize the performance. A set of applications are designed to evaluate the proposed scheme. Experimental results show that the proposed scheme can improve the performance by an average value of 22.2%, compared with the baseline schedule based on the worst case timing analysis. Compared with the conventional task scheduling approach based on the actual maximum clock frequencies, the proposed scheme also improves the performance by up to 12%. Zhixiang Chen 0003, Zhaolin Li, Shan Cao 0001, Jie Zhou 0001 |
Frontiers Inf. Technol. Electron. Eng. | 2 |
| 2015 | A Low-Latency and Low-Power Hybrid Scheme for On-Chip NetworksabstractNetwork-on-chip (NoC) has emerged as a vital factor that determines the performance and power consumption of many-core systems. This paper proposes a hybrid scheme for NoCs, which aims at obtaining low latency and low power consumption. In the presented hybrid scheme, a novel switching mechanism, called virtual circuit switching, is proposed to intermingle with circuit switching and packet switching. Flits traveling in virtual circuit switching can traverse the router with only one stage. In addition, multiple virtual circuit-switched (VCS) connections are allowed to share a common physical channel. Moreover, a path allocation algorithm is proposed in this paper to determine VCS connections and circuit-switched connections on a mesh-connected NoC, such that both communication latency and power are optimized. A set of synthetic and real traffic workloads are exploited to evaluate the effectiveness of the proposed hybrid scheme. The experimental results show that our proposed hybrid scheme can efficiently reduce the communication latency and power. For instance, for real traffic workloads, an average of 20.3% latency reduction and 33.2% power saving can be obtained when compared with the baseline NoC. Moreover, when compared with the NoC with virtual point-to-point connections (VIP), the proposed hybrid scheme can reduce the latency by 6.8% with the power decreasing by 11.3% averagely. Guoyue Jiang, Zhaolin Li, Shaojun Wei |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2014 | A High-Utilization Scheduling Schemeof Stream Programs on ClusteredVLIW Stream ArchitecturesabstractStream architectures have emerged as a mainstream solution for computation-intensive applications due to their rich arithmetic units. This paper proposes a multithreading technique based on a scheduling scheme of stream programs on clustered VLIW stream architecture, which aims at optimal arithmetic unit utilization without increasing energy consumption. Its principle is to exploit more kernel-level parallelism for further optimal compilation by constructing homogeneous multiple threads on stream programs. Three phases are proposed in the scheduling scheme. First, threads in stream programs are replicated for constructing homogeneous multiple threads. Second, time step assignment for homogeneous multithreaded stream programs is utilized to obtain efficient kernel combination. Third, stream segmentation is presented to optimize both memory transfers and startup overheads of kernels. A set of benchmarks are exploited to evaluate the effectiveness of the proposed technique. Experimental results show that, with exploiting kernel-level software pipeline, the proposed technique improves the performance by 20.9 percent averagely with the energy decreasing by 7.6 percent. Utilizations of adders and multipliers are up to average 77.4 and 75.8 percent, increasing 17.0 and 13.3 percent, respectively. Moreover, the proposed technique performs an average of 12.5 percent improvement over CSMT4 with the energy decreasing by 12.0 percent. Guoyue Jiang, Zhaolin Li, Shaojun Wei |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2014 | Compiler-Assisted Leakage- and Temperature- Aware Instruction-Level VLIW SchedulingabstractWith technology scaled to nanometer-scale, leakage energy consumption is accounting for a greater proportion than ever, especially for very long instruction word (VLIW) architectures with a large number of functional units (FUs). The growing energy consumption leads to an increase in chip temperature, which again brings an exponential growth in leakage current, and consequently leakage energy. However, few studies consider both leakage energy and temperature reduction during the compiling on VLIW architectures. In this paper, a leakage- and temperature-aware design flow is presented to assist the compiling of instruction-level VLIW scheduling. And two scheduling algorithms are proposed for the design flow. First, the leakage-aware rescheduling algorithm is proposed for leakage energy reduction by concentrating operations to fewer FUs and shutting more FUs down. Then, the temperature-aware workload balance algorithm is presented to reduce peak temperature by balancing the concentrated workloads among homogenous FUs. It is proved that the proposed two algorithms can reduce the leakage energy and peak temperature without performance loss. Experimental results demonstrate that the peak temperature is reduced by 15.27% and 12.84% for FU groups with three and two FUs and the leakage energy is reduced by 78.14% and 30.31% on average compared with the communication scheduling and list algorithm, respectively. Shan Cao 0001, Zhaolin Li, Shaojun Wei |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2013 | Energy-efficient stream task scheduling scheme for embedded multimedia applications on multi-issued stream architectures
Shan Cao 0001, Zhaolin Li, Guoyue Jiang, Zhixiang Chen 0003, Shaojun Wei |
J. Syst. Archit. | 2 |
| 2007 | Optimized design of a double-precision floating-point multiply-add-dused unit for data dependenceabstractThis paper presents a novel double-precision floating-point multiply-add-fused unit, which is implemented in three pipeline stages. The main improvement over the conventional design is data dependence between two consecutive floating-point instructions is considered. In the new design the intermediate computation results of the first floating-point instruction are first pretreated and then fed back to the first stage for being directly used by the second floating-point instruction if the two consecutive floating-point instructions are data dependent. In this way, floating point instructions can be executed directly following their preceding floating-point instructions without being stalled due to data dependence. 11 data dependence cases are accelerated in this paper. The experiments, which are done over four SPEC2000 benchmark programs, show that 25% performance increase can be attained at the cost of 0.27 ns time delay added to the critical path. Gongqiong Li, Zhaolin Li |
ICCD | 2 |