VLDB 2026 Research / reviewers in the wild / expert
Mengquan Li
dblp:172/2695
· DBLP profile ↗
32ranked-venue papers
9as first author
15since 2021 · last 2026
0000-0002-9385-734XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 29 · 9 first-author · 12 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | PointShuffler: Accelerating Point Cloud Neural Networks on General-Purpose GPUsabstractPoint Cloud Neural Networks (PCNNs) have emerged as a vital tool for latency-sensitive 3D perception applications, such as autonomous driving and AR/VR. However, their inherent computational redundancy—arising from excessive global sampling/search operations and repeated feature updates/aggregations caused by shared neighbors—severely constrains execution efficiency. More critically, conventional redundancy elimination methods usually introduce operations that are highly GPU-unfriendly, resulting in high memory overhead, increased branch divergence, irregular memory access, and serial dependencies, which together pose a significant challenge to PCNN acceleration. Yangfan Li 0001, Zhengjie Jin, Mengquan Li, Fengxiao Tang, Ming Zhao 0007, Cen Chen 0002 |
EuroSys | 4 |
| 2026 | PhotonRE: Photonic Redundancy Elimination Accelerator for Efficient Edge Computing
Mingkun Han, Zhaoyuan Zhang, Mengquan Li, Kenli Li 0001 |
ICIC (13) | 7 |
| 2026 | Bit-Density: A Non-Binary Photonic Accelerator Exploiting Bit-Level Sparsity
Haidong Wu, Zhaoyuan Zhang, Junyu He, Mengquan Li, Kenli Li 0001 |
ICIC (16) | 6 |
| 2025 | HIDE: Hyperspectral Imaging Dataset for Camouflaged Target Recognition
Zequn Zhang, Zhaoyuan Zhang, Zihe Chen, Yangfan Li 0001, Mengquan Li |
ICIC (17) | 8 |
| 2025 | SimDiff: Point Cloud Acceleration by Utilizing Spatial Similarity and Differential ExecutionabstractPoint cloud neural networks are gaining increasing attention in emerging 3-D computer vision applications, such as autonomous driving, robotics, and virtual reality. Many customized accelerators for 3-D point clouds have been developed to pursue superior time and energy efficiencies. In this work, we reveal that spatially adjacent points in a 3-D point cloud show similar feature values and relationships, implying substantial redundant computations and memory accesses, while which have been previously ignored. To reduce such redundancies, we propose SimDiff, an algorithm-accelerator co-design framework that boosts 3-D point cloud processing by cleverly leveraging spatial similarity toward excellent speedup and energy efficiency. On the algorithm side, we design a novel similarity-aware differential point cloud neural network (dubbed SD-PCNet). Differing from the standard flow of mainstream point cloud networks, it abstracts a brand-new execution flow for point cloud processing by utilizing spatial similarity among points and dynamic differential execution. On the accelerator side, we propose SD-PCAcc, a supporting accelerator to convert algorithm-level redundancy reductions into performance enhancements. On the deployment side, we propose efficient strategies for network-to-accelerator mapping and scheduling, high-bandwidth memory (HBM) channel allocation, and core component reconfiguration, facilitating the proposed methodologies into practical implementation. Extensive evaluation results show that, with preserved accuracy, our SimDiff gains an average of$3.2\times $speedup and$3.1\times $energy efficiency compared to the state-of-the-art competitors. Yangfan Li 0001, Mengquan Li, Cen Chen 0002, Xiaofeng Zou, Hongen Shao, Fengxiao Tang, Kenli Li 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2024 | AyE-Edge: Automated Deployment Space Search Empowering Accuracy yet Efficient Real-Time Object Detection on the Edge
Chao Wu 0006, Yifan Gong 0004, Liangkai Liu, Mengquan Li, Yushu Wu, Xuan Shen, Geng Yuan, Weisong Shi, Yanzhi Wang 0001 |
ICCAD | 4 |
| 2024 | Automated Optical Accelerator Search Toward Superior Acceleration Efficiency, Inference Robustness, and Development SpeedabstractRemarkable breakthroughs but daunting complexities of deep learning have aroused widespread interest in dedicated deep neural network (DNN) acceleration hardware, among which optical accelerators (OAs) are particularly promising thanks to their unprecedentedly high-performance-per-watt. However, the development of OAs is much slower than that of electrical accelerators due to threefold challenges. First, the OA design space is ample and discrete, making it tough for OA optimization; Second, the ecosystem that facilitates OA development is still in its infancy. Techniques to support OA design remain less explored, limiting both the achievable performance and the innovative development of OAs; and Third, OAs are highly sensitive to fabrication-induced process variations and thermal fluctuations (i.e., PTVs), which degrades OAs’ inference robustness and even renders them unusable in practice. In this article, we develop AutOAS, the first-of-its-kind framework for Automated Optical Accelerator Search, in order to jointly boost acceleration efficiency, inference robustness, and development speed. Our AutOAS comprises four enabling components: 1) a holistic OA search space, which takes full consideration of OAs’ micro-architectures (e.g., the type, shape and size of core functional units for data computation and data access), dataflow choices, DNN-to-accelerator mapping methods, memory hierarchy and PTV mitigation techniques; 2) a PTV Regulator, which can emulate the impact of PTVs on OAs’ inference accuracy based on given PTV profiles, and enables energy-efficient PTV mitigation on OAs; 3) an O-Performance Predictor, which enables accurate yet efficient predictions of an OA’s energy, throughput (latency) and chip area according to the DNN model and OA architecture parameters; and 4) two O-Search Engines (i.e., a differentiable search engine and an evolutionary search engine), which can automatically explore the large design space of OAs and identify the optimal accelerators to maximize the acceleration targets. Based on 10 DNN models widely applied in both computer vision and sequence modeling tasks, extensive experiments and ablation studies validate the effectiveness of our PTV Regulator, O-Performance Predictor, and O-Search Engines, as well as the superior performance of AutOAS-generated OAs. Mengquan Li, Kenli Li 0001, Chao Wu 0006, Gang Liu 0038, Mingfeng Lan, Yunchuan Qin, Zhuo Tang, Weichen Liu 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2023 | FIONA: Fine-grained Incoherent Optical DNN Accelerator Search for Superior Efficiency and RobustnessabstractIncoherent optical DNN accelerators (OAs) are booming thanks to unparalleled performance-per-watt and excellent scalability. To boost their innovative development, a recent work revolutionarily proposed automatic OA search. However, the robustness and the acceleration performance of the generated OAs are below expectation, because the impacts of inter-tile data transfer and fabrication process & thermal variations (i.e., PTVs) on OAs were ignored. Both hinder OA design automation from being a reality. To resolve theses challenges, we develop FIONA, a novel framework for Fine-grained Incoherent Optical DNN Accelerator search towards both superior acceleration efficiency and inference robustness. Compared against 5 state-of-the-art incoherent OAs on 9 DNN benchmarks, extensive experiments and ablation studies validate the effectiveness of FIONA, achieving up to 198.01× acceleration efficiency improvement based on guaranteed robustness. Mengquan Li, Kenli Li 0001, Mingfeng Lan, Zhuo Tang, Weichen Liu 0001 |
DAC | 1 |
| 2023 | RLAlloc: A Deep Reinforcement Learning-Assisted Resource Allocation Framework for Enhanced Both I/O Throughput and QoS Performance of Multi-Streamed SSDsabstractMulti-streamed Solid-State Disks (SSDs) have attracted increasing adoption in modern flash storage devices. Despite their excellent promise, effective flash resource allocation is still limiting both their achievable I/O performance and practical implementation. To this end, we develop the first-of-its-kind framework dubbed RLAlloc, which for the first time demonstrates deep Reinforcement Learning-assisted resource Allocation for boosting both I/O throughput and QoS performance of multi-streamed SSDs. Extensive experiments consistently validate the effectiveness of RLAlloc, improving up to 39.9% on I/O throughput and 44.0% on QoS performance over the state-of-the-art competitors. Mengquan Li, Chao Wu 0006, Congming Gao, Cheng Ji 0002, Kenli Li 0001 |
DAC | 1 |
| 2023 | An Efficient Hierarchical-Reduction Architecture for Aggregation in Route Travel Time EstimationabstractRoute travel time estimation (RTTE) is crucial in intelligent transportation systems. Performing aggregation is a fundamental operation in RTTE and is widely used in the traffic prediction and route calculation stages. Observations have revealed that aggregation operations in RTTE are influenced by the road network structure and aggregation requests, resulting in irregular data access, redundant processing, and workload imbalance. Existing architectures have not addressed these issues effectively. In this study, we begin by characterizing the execution pattern of performing aggregation operations on an Intel Core CPU. Guided by these characterizations, we propose an aggregation accelerator that utilizes a hierarchical reduction architecture (HRA) to perform aggregations in RTTE efficiently. Specifically, we construct an inverted table based on the road network and aggregation requests. Building upon the concept of hierarchical reduction, we design an HRA to accelerate aggregation operations which reduces irregular data access and eliminates redundant processing. Additionally, we introduce a reconfiguration mode for HRA to mitigate workload imbalance issues. Compared to a benchmark method executed on an Intel Core CPU, our design achieves an average$10\times$speedup. Zhao Liu 0006, Mengquan Li, Mincan Li, Kenli Li 0001 |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2022 | Contention Minimization in Emerging SMART NoC via Direct and Indirect RoutesabstractSMART (Single-cycle Multi-hop Asynchronous Repeated Traversal) Network-on-Chip (NoC), a recently proposed dynamically reconfigurable NoC, enables single-cycle long-distance communication by building single-bypass paths directly between distant communication pairs. However, such a single-cycle single-bypass path will be readily broken when contention occurs. Thus, packets will be buffered at intermediate routers with blocking latency from other contending packets, and extra router-stage latency to rebuild the remaining path when available. In this article, we propose an effective contention-minimized routing algorithm to achieve maximal bypassing. Specifically, we identify two potential routes: direct route, with which packets can reach the destination in a single bypass; and indirect route, with which packets can reach the destination in multiple bypasses via an(multiple) intermediate router(s). The novel feature is that, contrary to an intuitive approach, not the routes with minimal distance but the indirect routes via the arbitrary intermediate routers (even if they may be non-minimal) that avoid contentions yield the minimized end-to-end latency. Evaluation on realistic benchmarks demonstrates the effectiveness of the proposed routing strategy, which achieves average performance improvement by 35.48 percent in communication latency, 28.31 percent in application schedule length, and 37.59 percent in network throughput, compared with the current routing in SMART NoCs. Peng Chen 0027, Hui Chen 0016, Mengquan Li, Weichen Liu 0001, Chunhua Xiao, Yiyuan Xie, Nan Guan |
IEEE Trans. Computers | 4 |
| 2021 | O-HAS: Optical Hardware Accelerator Search for Boosting Both Acceleration Performance and Development SpeedabstractThe recent breakthroughs and prohibitive complexities of Deep Neural Networks (DNNs) have excited extensive interest in domain specific DNN accelerators, among which optical DNN accelerators are particularly promising thanks to their unprecedented potential of achieving superior performance-per-watt. However, the development of optical DNN accelerators is much slower than that of electrical DNN accelerators. One key challenge is that while many techniques have been developed to facilitate the development of electrical DNN accelerators, techniques that support or expedite optical DNN accelerator design remain much less explored, limiting both the achievable performance and the innovation development of optical DNN accelerators. To this end, we develop the first-of-its-kind framework dubbed O-HAS, which for the first time demonstrates automated Optical Hardware Accelerator Search for boosting both the acceleration efficiency and development speed of optical DNN accelerators. Specifically, our O-HAS consists of two integrated enablers: (1) an O-Cost Predictor, which can accurately yet efficiently predict an optical accelerator's energy and latency based on the DNN model parameters and the optical accelerator design; and (2) an O-Search Engine, which can automatically explore the large design space of optical DNN accelerators and identify the optimal accelerators (i.e., the micro-architectures and algorithm-to-accelerator mapping methods) in order to maximize the target acceleration efficiency. Extensive experiments and ablation studies consistently validate the effectiveness of both our O-Cost Predictor and O-Search Engine as well as the excellent efficiency of O-HAS generated optical accelerators. Mengquan Li, Zhongzhi Yu, Yongan Zhang, Yonggan Fu, Yingyan (Celine) Lin |
ICCAD | 1 |
| 2021 | Attack Mitigation of Hardware Trojans for Thermal Sensing via Micro-ring Resonator in Optical NoCsabstractAs an emerging role in new-generation on-chip communication, optical networks-on-chip (ONoCs) provide ultra-high bandwidth, low latency, and low power dissipation for data transfers. However, the thermo-optic effects of the photonic devices have a great impact on the operating performance and reliability of ONoCs, where the thermal-aware control with accurate measurements, e.g., thermal sensing, is typically applied to alleviate it. Besides, the temperature-sensitive ONoCs are prone to be attacked by the hardware Trojans (HTs) covertly embedded in the counterfeit integrated circuits (ICs) from the malicious third-party vendors, leading to performance degradation, denial-of-service (DoS), or even permanent damages. In this article, we focus on the tampering and snooping attacks during the thermal sensing via micro-ring resonator (MR) in ONoCs. Based on the provided workflow and attack model, a new structure of the anti-HT module is proposed to verify and protect the obtained data from the thermal sensor for attacks in its optical sampling and electronic transmission processes. In addition, we present the detection scheme based on the spiking neural networks (SNNs) to implement an accurate classification of the network security statuses for further high-level control. Evaluation results indicate that, with less than 1% extra area of a tile, our approach can significantly enhance the hardware security of thermal sensing for ONoC with trivial costs of up to 8.73%, 5.32%, and 6.14% in average latency, execution time, and energy consumption, respectively. Mengquan Li, Pengxing Guo, Weichen Liu 0001 |
ACM J. Emerg. Technol. Comput. Syst. | 2 |
| 2021 | Reduced Worst-Case Communication Latency Using Single-Cycle Multihop Traversal Network-on-ChipabstractThe communication latency in traditional network-on-chip (NoC) with hop-by-hop traversal is inherently restricted by the distance between source-destination communicating pairs. SMART, as one of the dynamically reconfigurable NoC architectures, enables the new feature of single-cycle long-distance communication by building a direct bypass path between distant cores dynamically at runtime. With the increasing of the number of integrated cores in multi/many-core systems, SMART has been deemed a promising communication backbone in such systems. However, SMART is generally optimized for average-case performance for best-effort traffics, not offering real-time guaranteed services for real-time traffics, and thus SMART often shows extremely poor real-time performance (e.g., schedulability). To make SMART latency-predictable for real-time traffics, by combining with the single-cycle bypass forwarding technique, in this article, we first propose a priority-preemptive scheduling to allow contending packets to be arbitrated according to predefined priorities. Based on the priority-based scheduling, for the real-time packet flows with given flow mapping and predefined priorities, we then propose a real-time communication analysis model, by considering shared virtual channels (or priority levels) and arbitrary-deadline real-time packet flows, to predict theworst-case communication latencyand validate the schedulability. Through theoretical and experimental comparison, theworst-case communication latencyof the analyzed packet flows is reduced significantly compared with that of the traditional priority-preemptive NoCs with hop-by-hop traversal and the original distance-based SMART, thus improving the schedulability. Peng Chen 0027, Weichen Liu 0001, Hui Chen 0016, Shiqing Li, Mengquan Li, Lei Yang 0018, Nan Guan |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2021 | Contention-Aware Routing for Thermal-Reliable Optical Networks-on-ChipabstractOptical network-on-chip (ONoC) architecture offers ultrahigh bandwidth, low latency, and low power dissipation for new-generation manycore systems. However, the benefits in communication performance and energy efficiency will be diminished by communication contention. The intrinsic thermal susceptibility is another challenge for ONoC designs. Under on-chip temperature variations, core functional devices suffer from significant thermal-induced optical power loss, which seriously threatens ONoCs' reliability. In this article, we develop novel routing techniques to resolve both issues for ONoCs. By analyzing the thermal effect in ONoCs, we first present a routing criterion at the network level. Combined with device-level thermal tuning, it can implement thermal-reliable ONoCs. Two routing approaches, including a mixed-integer linear programming (MILP) model and a heuristic algorithm (called CAR), are further proposed to minimize communication conflicts based on guaranteed thermal reliability, and meanwhile, maximize the communication energy efficiency in the presence of on-chip thermal variations. By applying the criterion, our approaches achieve excellent performance with largely reduced complexity of design space exploration. The evaluation results based on both synthetic traffic patterns and realistic benchmarks validate the effectiveness of our approaches with an average of 126.95% improvement in communication performance and 16.12% reduction in energy overhead compared to state-of-the-art techniques. CAR only introduces 7.20% performance difference compared to the MILP model and is more scalable to large-size ONoCs. Mengquan Li, Weichen Liu 0001, Luan H. K. Duong, Peng Chen 0027, Lei Yang 0018, Chunhua Xiao |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2020 | Contention Minimized Bypassing in SMART NoCabstractSMART, a recently proposed dynamically reconfigurable NoC, enables single-cycle long-distance communication by building single-bypass paths. However, such a single-cycle single-bypass path will be broken when contention occurs. Thus, lower-priority packets will be buffered at intermediate routers with blocking latency from higher-priority packets, and extra router-stage latency to rebuild remaining path, reducing the bypassing benefits that SMART offers. In this paper, we for the first time propose an effective routing strategy to achieve nearly contention-free bypassing in SMART NoC. Specifically, we identify two different routes for communication pairs: direct route, with which data can reach the destination in a single bypass; and indirect route, with which data can reach the destination in two bypasses via an intermediate router. If a direct route is not found, we would alternatively resort to an indirect route in advance to eliminate the blocking latency, at the cost of only one router-stage latency. Compared with the current routing, our new approach can effectively isolate conflicting communication pairs, greatly balance the traffic loads and fully utilize bypass paths. Experiments show that our approach makes 22.6% performance improvement on average in terms of communication latency. Peng Chen 0027, Weichen Liu 0001, Mengquan Li, Lei Yang 0018, Nan Guan |
ASP-DAC | 3 |
| 2020 | Lightweight Thermal Monitoring in Optical Networks-on-Chip via Router ReuseabstractOptical network-on-chip (ONoC) is an emerging communication architecture for manycore systems due to low latency, high bandwidth, and low power dissipation. However, a major concern lies in its thermal susceptibility - under onchip temperature variations, functional nanophotonic devices, especially microring resonator (MR)-based devices, suffer from significant thermal-induced optical power loss, which potentially counteracts the power advantages of ONoCs and even cause functional failures. Considering the fact that temperature gradients are typically found on many-core systems, effective thermal monitoring, performing as the foundation of thermal-aware management, is critical on ONoCs. In this paper, a lightweight thermal monitoring scheme is proposed for ONoCs. We first design a temperature measurement module based on generic optical routers. It introduces trivial overheads in chip area by reusing the components in routers. A major problem with reusing optical routers is that it may potentially interfere with the normal communications in ONoCs. To address it, we then propose a time allocation strategy to schedule thermal sensing operations in the time intervals between communications. Evaluation results show that our scheme exhibits an untrimmed inaccuracy of 1.0070 K with low energy consumption of 656.38 pJ/Sa. It occupies an extremely small area of 0.0020 mm2, reducing the area cost by 83.74% on average compared to the state-of-the-art optical thermal sensor design. Mengquan Li, Weichen Liu 0001 |
DATE | 1 |
| 2020 | Autonomous temperature sensing for optical network-on-chip
Weichen Liu 0001, Guiyu Tian, Mengquan Li |
J. Syst. Archit. | 3 |
| 2020 | Hardware-Software Collaborative Thermal Sensing in Optical Network-on-Chip-based Manycore SystemsabstractContinuous technology scaling in manycore systems leads to severe overheating issues. To guarantee system reliability, it is critical to accurately yet efficiently monitor runtime temperature distribution for effective chip thermal management. As an emerging communication architecture for new-generation manycore systems, optical network-on-chip (ONoC) satisfies the communication bandwidth and latency requirements with low power dissipation. Moreover, observation shows that it can be leveraged for runtime thermal sensing. In this article, we propose a brand-new on-chip thermal sensing approach for ONoC-based manycore systems by utilizing the intrinsic thermal sensitivity of optical devices and the inter-processor communications in ONoCs. It requires no extra hardware but utilizes existing optical devices in ONoCs and combines them with lightweight software computation in a hardware-software collaborative manner. The effectiveness of the our approach is validated both at the device level and the system level through professional photonic simulations. Evaluation results based on synthetic communication traces and realistic benchmarks show that our approach achieves an average temperature inaccuracy of only 0.6648 K compared to ground-truth values and is scalable to be applied for large-size ONoCs. Mengquan Li, Weichen Liu 0001, Nan Guan, Yiyuan Xie, Yaoyao Ye |
ACM Trans. Embed. Comput. Syst. | 1 |
| 2019 | Routing in optical network-on-chip: minimizing contention with guaranteed thermal reliabilityabstractCommunication contention and thermal susceptibility are two potential issues in optical network-on-chip (ONoC) architecture, which are both critical for ONoC designs. However, minimizing conflict and guaranteeing thermal reliability are incompatible in most cases. In this paper, we present a routing criterion in the network level. Combined with device-level thermal tuning, it can implement thermal-reliable ONoC. We further propose two routing approaches (including a mixed-integer linear programming (MILP) model and a heuristic algorithm (CAR)) to minimize communication conflict based on the guaranteed thermal reliability, and meanwhile, mitigate the energy overheads of thermal regulation in the presence of chip thermal variations. By applying the criterion, our approaches achieve excellent performance with largely reduced complexity of design space exploration. Evaluation results on synthetic communication traces and realistic benchmarks show that the MILP-based approach achieves an average of 112.73% improvement in communication performance and 4.18% reduction in energy overhead compared to state-of-the-art techniques. Our heuristic algorithm only introduces 4.40% performance difference compared to the optimal results and is more scalable to large-size ONoCs. Mengquan Li, Weichen Liu 0001, Lei Yang 0018, Peng Chen 0027, Duo Liu 0002, Nan Guan |
ASP-DAC | 1 |
| 2019 | Thermal Sensing Using Micro-ring Resonators in Optical Network-on-ChipabstractIn this paper, we for the first time utilize the micro-ring resonators (MRs) in optical networks-on-chip (ONoCs) to implement thermal sensing without requiring additional hardware or chip area. The challenges in accuracy and reliability that arise from fabrication-induced process variations (PVs) and device-level wavelength tuning mechanism are resolved. We quantitatively model the intrinsic thermal sensitivity of MRs with finegrained consideration of wavelength tuning mechanism. Based on it, a novel PV-tolerant thermal sensor design is proposed. By exploiting the hidden ‘redundancy’ in wavelength division multiplexing (WDM) technique, our sensor achieves accurate and efficient temperature measurement with the capability of PV tolerance. Evaluation results based on professional photonic component and circuit simulations show an average of 86.49% improvement in measurement accuracy compared to the state-of-the-art on-chip thermal sensing approach using MRs. Our thermal sensor achieves stable performance in the ONoCs employing dense WDM with an inaccuracy of only 0.8650 K. Weichen Liu 0001, Mengquan Li, Wanli Chang 0001, Chunhua Xiao, Yiyuan Xie, Nan Guan, Lei Jiang 0001 |
DATE | 2 |
| 2019 | Energy-Efficient Application Mapping and Scheduling for Lifetime Guaranteed MPSoCsabstractEnergy optimization is one of the most critical objectives for the synthesis of multiprocessor system-on-chip (MPSoC). Besides, to ensure a long processor lifetime and to maintain a safe chip temperature are also important for multiprocessor manufactures under deep submicrometer process technologies. This paper presents a mixed integer linear programming (MILP) model to determine the mapping and scheduling of real-time applications onto embedded MPSoC platforms, such that the total energy consumption is minimized with the lifetime reliability constraint and the temperature threshold constraint satisfied. We develop a lightweight temperature model that can be integrated in the MILP model to predict the chip temperature accurately and efficiently. By exploiting the dynamic voltage and frequency scaling capability of modern processors, processor voltage/frequency assignment is also considered in our MILP model. Extensive performance evaluations on synthetic and real-world applications demonstrate the effectiveness of the proposed approach. Our MILP model achieves an average reduction of 19.09% and 28.53% total energy in comparison with two state-of-the-art techniques on the basis of guaranteeing the safe chip temperature and system lifetime reliability. Weichen Liu 0001, Juan Yi, Mengquan Li, Peng Chen 0027, Lei Yang 0018 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2018 | User Experience-Enhanced and Energy-Efficient Task Scheduling on Heterogeneous Multi-Core Mobile SystemsabstractHeterogeneous Multi-Core Mobile Systems has been widely used to improve performance. However, it faces with the challenge of tradeoff between energy saving and user experience. ARM big. LITTLE architecture, a heterogeneous computing architecture, is a power-optimization technology. In most big. LITTLE devices, however, it still cannot achieve excellent user experience and higher energy saving. In this paper, we propose an improved task scheduling (UCES-GTS) by introducing the concept of user-centric task on big. LITTLE mobile device. In order to enhance user experience, the response time of user-centric tasks is shortened with reducing slack time of them properly. We then present a detailed algorithm to compute appropriate frequency and allocate the CPU resources to each task. The experimental evaluation results show that our improved global task scheduling model can achieve 17 % and 8 % energy saving average compared with the clustered switching scheduling and the original global task scheduling respectively. And the response time of user-centric tasks can decrease 27 % average, which means excellent user experience. Weichen Liu 0001, Mengquan Li, Peng Chen 0027, Lei Yang 0018, Chunhua Xiao, Yaoyao Ye |
ICPADS | 3 |
| 2018 | Fine-Grained Task-Level Parallel and Low Power H.264 Decoding in Multi-Core SystemsabstractIn the past few years, the extinction of Moore's Law makes people reconsider the solutions for dealing with the low computing resource utilization of applications on multicore processor systems. However, making good use of computing resources in multi-core processors systems is not easy due to the differences between single-core and multi-core architecture. Nowadays short video apps like Instagram and Tik Tok have successfully caught people's eyes by fascinating short videos, typically just 10 to 30 seconds long, uploaded by the users of apps. And almost all of these videos are recorded by their mobile devices, which are typically HD (High Definition) or FHD (Full High Definition) videos, which prefer to be encoded/decoded by H.264/AVC rather then HEVC (High Efficiency Video Coding) on mobile devices in view of the energy consumption and decoding speed. How to dive the huge potential of the computing resource on multi-core mobile devices to speed up decoding these videos while consuming low energy, is a big challenge. In our previous work [1], a relatively simple parallel framework was proposed to implement a parallel H.264/ AV C decoder. This work further proposes a more detailed systematic task-level parallel framework, together with an energy saving strategy based on this framework, to research a new H.264/AVC decoder on multi-core processor systems. The proposed parallel method is composed of a set of rules to guide parallel software programming (PSPR) and a software parallelization framework (SPF). The PSPR is applied in pre-processing steps to address the potential issues limiting the inherent parallelism, and the SPF is applied to parallelize the original serial programs. After the parallelization is successfully deployed, DVFS technique would be applied to decrease the power dissipation based on the SPF. Results show that proposed solutions make a significant improvement in decoding speed of 32% at 720p, 27% at 1080p and 29% at 2160p, and in energy savings of 25% at 720p, 25% at 1080p and 23% at 2160p on a four-core workstation running Linux, compared to the original serial H.264/ AV C decoder. The results demonstrate our methods are effective and scalable, served as a reference for future parallel software development. Wenyang Liu, Weichen Liu 0001, Mengquan Li, Peng Chen 0027, Lei Yang 0018, Chunhua Xiao, Yaoyao Ye |
ICPADS | 3 |
| 2018 | Chip Temperature Optimization for Dark Silicon Many-Core SystemsabstractIn the dark silicon era, a fundamental problem is given a real-time computation demand, how to determine if an on-chip multiprocessor system is able to accept this demand and to maintain its reliability by keeping every core within a safe temperature range. In this paper, a practical thermal model is described for quick chip temperature prediction. Integrated with the thermal model, we present a mixed integer linear programming (MILP) model to find the optimal task-to-core assignment with the minimum chip peak temperature. For the worst case where even the minimum chip peak temperature exceeds the safe temperature, a heuristic algorithm, called temperature-constrained task selection (TCTS), is proposed to optimize the system performance within chip safe temperature. The optimality of the TCTS algorithm is formally proven. Extensive performance evaluations show that our thermal model achieves an average prediction accuracy of 0.0741 °C within 0.2392 ms. The MILP model reduces chip peak temperature of ~10 °C comparing with traditional techniques. The system performance is increased by 19.8% under safe temperature limitation. Due to the satisfying scalability of our MILP formulation, the chip peak temperature is further decreased by 5.06 °C via the TCTS algorithm. The feasibility of this systematical technique is testified in a real case study as well. Mengquan Li, Weichen Liu 0001, Lei Yang 0018, Peng Chen 0027, Chao Chen 0004 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2017 | Dark silicon-aware hardware-software collaborated design for heterogeneous many-core systemsabstractARM's big. LITTLE architecture coupled with Heterogeneous Multi-Processing (HMP) has enabled energy-efficient solutions in the dark silicon era. System-level techniques activate nonadjacent cores to eliminate chip thermal hotspot. However, it unexpectedly increases communication delay due to longer distance in network architectures, and in turn degrades application performance and system energy efficiency. In this paper, we present a novel hierarchical hardware-software collaborated approach to address the performance/temperature conflict in dark silicon many-core systems. Optimizations on interprocessor communication, application performance, chip temperature and energy consumption are well isolated and addressed in different phases. Evaluation results show that on average 22.57% reduction of communication latency, 23.04% improvement on energy efficiency and 6.11°C reduction of chip peak temperature are achieved compared with state-of-the-art techniques. Lei Yang 0018, Weichen Liu 0001, Nan Guan, Mengquan Li, Peng Chen 0027, Edwin H.-M. Sha |
ASP-DAC | 4 |
| 2017 | Task Mapping on SMART NoC: Contention Matters, Not the DistanceabstractOn-chip communication is the bottleneck of system performance for NoC-based MPSoCs. SMART, a recently proposed NoC architecture, enables single-cycle multi-hop communications. In SMART NoCs, unconflicted messages can go through an express bypass and the communication efficiency is significantly improved, while conflicted messages have to be buffered for guaranteed delivery with extra delays. Therefore, that performance of SMART NoC may be seriously degraded when communication contention increases. In this paper, we present task mapping techniques to address this problem for SMART NoCs, with the consideration of communication contention, rather than inter-processor distance, by minimizing conflicts and thus maximizing bypass utilization. We first model the entire problem by ILP formulations to find the theoretically optimal solution, and further propose polynomial-time algorithms for contention-aware task mapping and message priority assignment. Communicating tasks can be mapped to distant processors in SMART NoCs as long as conflict-free communication paths can be established and bypass can be enabled. Evaluation results on real benchmarks show an average of 44.1% and 32.8% improvement in communication efficiency and application performance compared to state-of-the-art techniques. The proposed heuristic algorithms only introduce 1.9% performance difference compared to the ILP model and are more scalable to large-size NoCs. Lei Yang 0018, Weichen Liu 0001, Peng Chen 0027, Nan Guan, Mengquan Li |
DAC | 5 |
| 2017 | Quantitative Modeling of Thermo-Optic Effects in Optical Networks-on-ChipabstractOptical networks-on-chip (ONoCs) is a new promising communication paradigm that upgrades the traditional on-chip networks (NoCs) with the ultra-high communication bandwidth and low latency. Silicon microring resonators (MRRs), as a critical component of ONoCs used to implement the selection and redirection of optical signals, are inherently sensitive to the environmental temperature. The applicability of the ONoCs is essentially restricted by the performance of these optical devices that relies on the thermal conditions of the chip. In this paper, we study the thermo-optic effects of the MRRs quantitatively, build and verify the models of the MRRs based on the finite-difference time-domain (FDTD) method. We present formal relationship models between the temperature of a MRR and its optical losses and resonance wavelength. For the first time, the variation between the two types of MRRs, the parallel microring resonators (PMRs) and the crossing microring resonators (CMRs), are systematically addressed, which greatly improves the accuracy and applicability of the models. The results presented in this paper are systematically verified using professional optics methodology, and can be widely applied for accurate and efficient analysis of the thermo-optic effects in different domains of the ONoC community. Weichen Liu 0001, Mengquan Li, Yiyuan Xie, Nan Guan |
ACM Great Lakes Symposium on VLSI | 3 |
| 2017 | Hardware-software collaboration for dark silicon heterogeneous many-core systems
Lei Yang 0018, Weichen Liu 0001, Weiwen Jiang, Chao Chen 0004, Mengquan Li, Peng Chen 0027, Edwin H.-M. Sha |
Future Gener. Comput. Syst. | 5 |
| 2017 | FoToNoC: A Folded Torus-Like Network-on-Chip Based Many-Core Systems-on-Chip in the Dark Silicon EraabstractDark silicon refers to the phenomenon that a fraction of a many-core chip has to become “dark” or “dim” in order to guarantee the system to be kept in a safe temperature range and allowable power budget. Techniques have been developed to selectively activate non-adjacent cores on many-core chip to avoid temperature hotspot, while resulting unexpected increase of communication overhead due to the longer average distance between active cores, and in turn affecting application performance and energy efficiency, when Network-on-Chip (NoC) is used as a scalable communication subsystem. To address the brand-new challenges brought by dark silicon, in this paper, we present FoToNoC, a Folded Torus-like NoC, coupled with a hierarchical management strategy for heterogeneous many-core systems. On top of it, objectives of maximizing application performance, energy efficiency and chip reliability are isolated and well achieved by hardware-software co-design in several different phases, including application mapping and scheduling, cluster management and DVFS control. Evaluations on PARSEC benchmark applications demonstrate the significance of the entire strategy. Compared with state-of-the-art approaches, the proposed FoToNoC organization can achieve on average 35.4 and 35.2 percent on communication efficiency and application performance improvement, respectively, when maintaining the safe chip temperature. The hierarchical cluster-based management strategy can further reduce an average 34.6 percent of the total energy consumption with a notable reduction on the chip peak temperature. The significant achievements on system energy efficiency and the reduction on chip temperature of H.264 decoder and DSP-stone benchmarks additionally verify the effectiveness of the proposed methods. Lei Yang 0018, Weichen Liu 0001, Weiwen Jiang, Mengquan Li, Peng Chen 0027, Edwin H.-M. Sha |
IEEE Trans. Parallel Distributed Syst. | 4 |
| 2016 | FoToNoC: A hierarchical management strategy based on folded lorus-like Network-on-Chip for dark silicon many-core systemsabstractIn this dark silicon era, techniques have been developed to selectively activate nonadjacent cores in physical locations to maintain the safe temperature and allowable power budget on a many-core chip. This will result in unexpected increase in the communication overhead due to longer average distance between active cores in a typical mesh-based Network-on-Chip (NoC), and in turn reduce the system performance and energy efficiency. In this paper, we present FoToNoC, a Folded Torus-like NoC, and a hierarchical management strategy on top of it, to address this tradeoff problem for heterogeneous many-core systems. Optimizations of chip temperature, inter-core communication, application performance, and system energy consumption are well isolated in FoToNoC, and addressed in different design phases and aspects. A cluster-based hierarchical strategy is proposed to manage the system adaptively in several different control levels. Compared with mesh-based systems on a set of synthetic and real benchmarks, FoToNoC can achieve on average 39.4% performance improvement when similar temperature conditions are maintained, and the proposed strategy can further reduce the total energy consumption by up to 42.0%. Lei Yang 0018, Weichen Liu 0001, Weiwen Jiang, Mengquan Li, Juan Yi, Edwin H.-M. Sha |
ASP-DAC | 4 |
| 2016 | Application Mapping and Scheduling for Network-on-Chip-Based Multiprocessor System-on-Chip With Fine-Grain Communication OptimizationabstractNetwork-on-chip (NoC) is promising for the communication paradigm of the next-generation multiprocessor system-on-chip (MPSoC). As communication has become an integral part of on-chip computing, and even the performance bottleneck, researchers are paying much attention to its implementation and optimization. Traditional techniques that model communication inaccurately will lead to unexpected runtime performance, which is on average 90.8% worse than the predicted results based on observation, and are not suitable for the deep optimization of communication-intensive scenarios. In this paper, techniques are presented for the NoC-based MPSoCs that integrate optimization on interprocessor communications with the objective of minimizing the schedule length. A fine-grained integer-linear programming (ILP) model is proposed to properly address the communication latency with a network contention, which generates runtime scheduling with trivial performance difference from the predictions. We further propose a heuristic algorithm, unified priority-based scheduling (UPS), to effectively solve the contention problem in polynomial time by assigning priorities to messages. Evaluation results show that the solutions obtained by the ILP model outperform the state-of-the-art techniques by 31.1%, and UPS improves application performance by 34.7% and 44.4% compared with acquainted first-in-first-out (FIFO)-based and random-based methods. In addition, UPS achieves averagely 8.3% approximated results with the optimal solutions generated by ILP. A case study on H.264 high-definition television (HDTV) decoder and the digital signal processor (DSP) filter benchmarks achieves significant improvement on the performance and the results prediction accuracy, as well as the prominent reduction in the number of network contention and energy consumption. Lei Yang 0018, Weichen Liu 0001, Weiwen Jiang, Mengquan Li, Juan Yi, Edwin H.-M. Sha |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |